Open any AI governance framework written for a finance function and you will find, somewhere near the centre, a control that reads approximately: outputs are reviewed by a qualified person before use. That sentence is doing almost all the work in the framework, and the empirical literature on how humans actually behave when reviewing automated output suggests it does considerably less work than its authors believe.

Key Takeaway

Parasuraman and Manzey's 2010 review in Human Factors reports that automation bias produces both omission and commission errors when decision aids are imperfect, that it occurs in both naive and expert participants, and that it cannot be prevented by training. Automation complacency likewise occurs in both naive and expert participants and cannot be overcome with simple practice. Critically, complacency arises under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention, which means the same attentional budget funds both the productivity gain from automation and the verification that is supposed to control it. Parasuraman and Riley's 1997 taxonomy adds that disuse is commonly caused by falsely activating alarms, often because the base rate of the condition to be detected was not considered when setting the trade-off between false alarms and omissions. A control framework whose principal safeguard is a human reviewer, and whose principal mitigation is training, is resting on two propositions the literature does not support.

The Assumption Underneath Every AI Control

The reason this literature matters to a Canadian finance function is that the profession has converged on a single control pattern for AI, and that pattern has a research record.

The pattern is human review. An AI system produces a classification, a draft, a recommendation or a flag, and a qualified person checks it before it takes effect. This appears in model risk frameworks, in professional guidance, in vendor documentation and in the internal policies most firms have written over the last two years. It is intuitive, it is cheap to write down, and it allows an organisation to claim that judgment remains human.

What makes it worth examining is that human review of automated output is not a new activity invented for large language models. Aviation, medicine, process control and military command have been studying it since at least the early 1980s, and the findings are consistent, replicated and uncomfortable.

The literature does not say that human review is worthless. It says that human review degrades in specific, predictable ways, that the degradation is not a matter of individual diligence, and that the standard organisational responses to it, chiefly training and exhortation, do not work. That is a different claim from the one most AI policies implicitly make, and it has design consequences.

Use, Misuse, Disuse, Abuse

The organising framework, from what is probably the most cited paper in the field.

Parasuraman and Riley, writing in Human Factors in 1997, address theoretical, empirical and analytical studies pertaining to human use, misuse, disuse and abuse of automation technology. Use refers to the voluntary activation or disengagement of automation by human operators, and trust, mental workload and risk can influence automation use, though interactions between factors and large individual differences make prediction of automation use difficult. Misuse refers to over-reliance on automation, which can result in failures of monitoring or decision biases. Disuse, or the neglect or underutilization of automation, is commonly caused by alarms that activate falsely[1].

The fourth category, abuse, concerns the automation of functions by designers without adequate regard for the consequences for human performance, which we discuss below.

Two features of this taxonomy are immediately useful to a finance function.

It separates two failure modes that organisations habitually conflate. A team that trusts the system too much and a team that abandons it are both failing, but the causes are opposite and the remedies are opposite. An intervention that increases trust to address disuse will worsen misuse, and vice versa. Most AI adoption programmes push in one direction only.

And it identifies large individual differences that make prediction difficult[1]. That matters because it means an organisation cannot assume its control behaves uniformly across staff. The same policy, applied to two reviewers, produces different reliance behaviour, which is not something a policy document can fix by being written more clearly.

Omission And Commission

The distinction that determines what a control is actually protecting against.

Skitka, Mosier and Burdick, in a 1999 study published in the International Journal of Human-Computer Studies, found evidence of both commission and omission errors in automated monitoring tasks with non-pilot samples, in a low-fidelity part task requiring gauge and progress monitoring in addition to a tracking task[2]. Parasuraman and Manzey summarise the general finding that automation bias results in making both omission and commission errors when decision aids are imperfect[3].

The two are structurally different.

Omission errors occur when the operator fails to detect a problem because the automation did not flag it. The reviewer is not overriding the system; the reviewer never engages, because nothing prompted engagement. In a finance context this is the transaction the anomaly detection did not surface, the disclosure the drafting tool did not raise, the reconciliation difference the matching engine absorbed silently.

Commission errors occur when the operator follows an incorrect automated recommendation despite contrary evidence being available. Here the reviewer does engage and gets it wrong, accepting the system's answer over information that should have contradicted it.

Commentary summarising this literature states the empirical position bluntly: individuals are more likely to accept incorrect automated advice than to challenge it, even when contradictory information is available[4].

The design consequence, which we state as our own analysis, is that a review control addresses commission errors and does very little about omission errors. A reviewer examining what the system produced cannot, by construction, examine what it failed to produce. Yet in most finance applications the omission failure is the more dangerous one, because the missing item leaves no artefact to review.

Expertise Does Not Protect You

The finding that undermines the most common organisational mitigation.

Parasuraman and Manzey report that automation complacency is found in both naive and expert participants, and that automation bias occurs in both naive and expert participants[3].

Consider what the standard finance control does with expertise. It assigns review to a senior person on the theory that experience produces better judgment about when the system is wrong. This is the logic of a manager reviewing a junior's work, transplanted to a machine, and it is the reason firms feel comfortable deploying AI in areas where a qualified professional signs off.

The research contradicts the transplant. Expertise in the subject matter does not confer resistance to automation bias, because the bias is not a knowledge deficit. It is a pattern in how attention and trust are allocated when a source is perceived as reliable, and a domain expert allocates attention the same way a novice does once that perception is established.

There is a plausible aggravating factor worth naming, and we offer it as our own inference rather than a finding of the literature. A senior reviewer is typically the person with the least available attention, the most competing demands, and the strongest institutional pressure to clear work quickly. Against the complacency mechanism described below, that combination is unfavourable.

Neither Does Training

The finding that undermines the second most common mitigation.

Parasuraman and Manzey report that automation complacency cannot be overcome with simple practice, and that automation bias cannot be prevented by training[3]. Manzey and colleagues examined this directly, in a study titled around the misuse of automated decision aids and the impact of training experience[5].

This is the most operationally significant statement in the article, because training is the default corporate response to almost every control weakness. An organisation that discovers reviewers are accepting AI output uncritically will run a session on critical evaluation of AI output, record attendance, and consider the matter addressed.

The literature indicates that response does not work, and the reason it does not work follows from the two findings above. If the bias is not a knowledge deficit, and if expertise does not confer resistance, then transmitting more knowledge cannot be the remedy. Telling people that automation is fallible does not change how attention is allocated when the automation has been reliable for six weeks.

We would draw a careful boundary here. The literature does not say that no intervention works; it says training and simple practice do not overcome the effect. The mitigations that have research support are structural rather than educational, and they are set out below.

The immediate implication for a Canadian firm is that an AI governance framework listing training as its principal mitigation for over-reliance has identified the right risk and selected a remedy the evidence does not support.

Complacency Is An Attention Problem

The mechanism, which is what makes the effect predictable rather than mysterious.

Parasuraman and Manzey report that automation complacency occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention[3]. Their paper is subtitled an attentional integration, and its contribution is to explain complacency and automation bias as arising from a common attentional process rather than treating them as separate phenomena[6].

The word complacency invites a moral reading, as though the operator has become lazy or careless. The attentional account is different and more useful. Attention is a limited resource. When an operator has other tasks competing for it, monitoring of the automated task receives less, and detection of automation failures declines. This is a resource allocation outcome, not a character trait.

Earlier work by Molloy and Parasuraman examined monitoring of an automated system for a single failure, addressing vigilance and task complexity effects[7], and the complacency effect they document concerns detection of automation failures under differing conditions.

The predictive value for a finance function is direct. If complacency scales with competing task load, then an organisation can forecast where its AI review controls will be weakest: in the periods and roles with the highest concurrent demands. In practical terms that means month-end, year-end, filing deadlines and the reviewers carrying the most parallel responsibilities, which is to say precisely the conditions under which most finance work is performed.

The Control And The Gain Share A Budget

The implication we consider most important, offered as our own analysis built on the attentional finding.

An organisation deploys AI in a finance workflow to increase throughput. The gain is realised by having people do more, or do the same with fewer people. Either way, the concurrent task load on the remaining staff rises.

The control on that same workflow is human review, whose effectiveness, per the attentional account, declines as concurrent task load rises[3].

So the productivity benefit and the control effectiveness draw on the same finite resource, and they draw on it in opposite directions. Realising the benefit degrades the control. An organisation that deploys AI and holds headcount constant while increasing volume has, without deciding to, weakened the safeguard it wrote into its framework.

The corollary is uncomfortable and worth stating. A review control is only as strong as the unallocated attention available to the reviewer, and that attention is exactly what the business case proposed to monetise. A firm cannot claim the full efficiency gain and the full control at once.

This suggests a design question that most AI business cases never ask: what proportion of the freed capacity is being reserved to fund verification. If the answer is none, the framework's central control has been defunded at the moment of deployment.

Verification Behaviour Is Measurable

A finding that converts an abstract risk into something an organisation can actually monitor.

Bahner, Hüper and Manzey, in work published in the International Journal of Human-Computer Studies in 2008, examined the proportion of relevant parameters sampled among participants who did and did not commit a commission error[5], a figure reproduced in the Parasuraman and Manzey review[3].

The framing is what matters. Verification is not a binary act of reviewing or not reviewing. It is a quantity: how many of the relevant inputs did the operator actually consult before accepting the recommendation. And that quantity relates to whether a commission error occurred.

The transferable idea for a finance function, which is our own extension rather than a finding of that study, is that review quality can be instrumented rather than asserted. If a reviewer approves an AI-generated classification, the system can record which supporting information was actually opened, for how long, and in what sequence. That produces a measure of verification depth rather than a signature attesting that review occurred.

Most finance systems today record only the attestation. A sign-off field establishes that someone clicked approve, which is precisely the datum the research suggests is least informative, because commission errors are committed by people who did approve after looking at too little.

We would be careful about how such instrumentation is used. Measuring verification depth to discipline individuals invites gaming and misreads the research, which locates the cause in attentional conditions rather than in individual diligence. Measuring it to identify which workflows and which periods produce shallow verification is the use the evidence supports.

Disuse And The Base Rate Problem

The opposite failure, and a statistical explanation for it that deserves far more attention in finance than it receives.

Parasuraman and Riley state that disuse, or the neglect or underutilization of automation, is commonly caused by alarms that activate falsely, and that this often occurs because the base rate of the condition to be detected is not considered in setting the trade-off between false alarms and omissions[1].

The statistical point is the familiar one about rare events, and it is routinely ignored in the design of detection systems. Where the condition being detected is rare, even a highly accurate detector produces mostly false positives among the items it flags, because the large population of negatives generates false alarms that swamp the small population of true positives.

Finance applications are full of rare-event detection. Fraudulent transactions, material misstatements, duplicate payments, sanctions matches and anomalous journal entries are all rare relative to the volume screened. A detector tuned on accuracy without regard to base rate will produce an alert queue dominated by false positives.

What follows is the mechanism the literature describes: operators learn from experience that alerts are usually wrong, and the automation falls into disuse. The system remains installed, remains described in the control framework, and is functionally abandoned by the people it was built for.

The organisational danger is that disuse is invisible to governance. Utilisation looks fine because alerts are being cleared. What is not visible is that they are being cleared without genuine assessment, which converts a detection control into an administrative queue. A firm should treat a high alert volume with a low true-positive rate as a control failure in progress rather than as evidence that the system is working hard.

Out Of The Loop

The longer-run degradation, which operates on a different timescale from the errors above.

Commentary summarising the field notes that important performance consequences of complacency include a loss of situation awareness, and an elevated risk that operators fail to detect and manage automation failures in due time, and that complacency has consequently been regarded as an important factor contributing to what has been called out-of-the-loop unfamiliarity in human-automation interaction[5]. The same source records that according to Funk and colleagues, complacency belongs to the five most important issues of cockpit automation and has been identified as a contributing factor to numerous incidents and accidents in civil aviation[5].

Out-of-the-loop unfamiliarity describes a state in which the operator, having supervised rather than performed a task for an extended period, has lost the working familiarity required to take over competently when the automation fails.

The finance analogue is straightforward and, in our assessment, under-discussed. A team that has used an AI system to perform a reconciliation, a classification or a first-draft analysis for two years contains people who have supervised that process rather than performed it. Their ability to detect a subtle systematic error in the output depends on a familiarity with the underlying work that supervision does not maintain.

This matters most precisely when it is needed most. The scenario in which a firm most requires its people to catch what the system got wrong is the scenario in which the system has been quietly wrong for some time, and that is the scenario in which the supervising humans have had the longest exposure to its outputs being accepted.

The Transfer Nobody Authorised

A framing from recent commentary on this literature that captures the governance problem better than most governance documents do.

Commentary observes that as reliance intensifies, automated systems shift from decision aids to implicit arbiters of judgment through routine use rather than formal transfers of authority, and that the central danger is not that machines decide independently but that institutions organize decision-making around automated outputs in ways that weaken the capacity to question, contextualize and intervene[4].

That describes something an AI governance framework is structurally unable to detect. Frameworks govern authorised delegations: what the system may decide, what requires human approval, who signs off. The transfer being described here is not authorised and never appears in a document. It occurs because processes, timelines and staffing gradually reorganise around the system's output being right.

The observable symptoms are organisational rather than technical. Turnaround times that assume the system's answer is accepted. Downstream steps that cannot absorb an override without missing a deadline. Staffing levels premised on the review being a formality. Escalation paths that exist on paper but that nobody has used.

A practical diagnostic follows, and we offer it as our own. Ask what happens if a reviewer rejects the system's output on a routine item, and trace the consequences honestly. If rejection triggers rework that cannot be completed in the available window, or requires a conversation the reviewer would rather avoid, then the review is nominal regardless of what the policy says, and the transfer of judgment has already occurred.

Does Accountability Help?

One line of research on remedies, reported with the caution its evidentiary status requires.

Skitka, Mosier and Burdick published a study in the International Journal of Human-Computer Studies in 2000 whose stated goal was to explore the extent to which social accountability could ameliorate decision-making problems observed in highly automated work environments, and to shed further light on the psychological underpinnings of automation bias. The study reports average numbers of commission and omission errors as a function of accountability condition, and was supported by a grant from NASA Ames Research Center[8].

We have the study's design and framing but not its results, and we will not characterise findings we have not read. The honest statement is that social accountability has been investigated as a potential mitigation, that the investigation measured both error types against accountability conditions, and that a reader wanting to know whether it worked should obtain the paper.

What we can say is why the hypothesis is plausible and what it would imply. If automation bias involves an implicit reallocation of responsibility to the system, then conditions that make the human's own responsibility salient might counteract it. Commentary discussing this literature notes that heuristics lead individuals to allocate responsibility for decision outcomes to automated systems more easily, and draws an analogy to diffusion of responsibility in groups[9].

For a finance function the design question this raises is whether reviewers experience personal responsibility for an AI-assisted output in the way they would for their own work, and our observation is that the sign-off conventions many firms have adopted, in which the reviewer approves a system-generated item rather than producing and owning the analysis, may work against that.

Are Teams Better Than Individuals?

A second remedy that has been studied and that finance functions frequently assume without evidence.

Skitka, Mosier, Burdick and Rosenblatt addressed whether teams are better than individuals in the context of automation bias and errors, in work cited to the International Journal of Aviation Psychology[2].

We flag again that we have the existence and subject of this work rather than its findings, and we do not assert an outcome.

The reason it matters is that adding reviewers is the most common structural response after training. If one reviewer is insufficient, add a second; if a second is insufficient, escalate to a committee. This is the maker-checker instinct applied to AI, and it assumes that independent reviewers fail independently.

The assumption is questionable on the mechanism described in this article. If automation bias arises from a shared perception of the system's reliability and from shared attentional conditions, then two reviewers of the same organisation, looking at the same output under the same deadline, are not independent. They may correlate strongly, in which case the second review adds cost and less assurance than its cost implies. Whether that is so is an empirical question the cited work addresses and this article cannot answer.

Abuse: The Designer's Contribution

The fourth term in the taxonomy, and the one that locates responsibility outside the operator.

Parasuraman and Riley's framework includes abuse alongside use, misuse and disuse[1], referring to the automation of functions by designers without adequate regard for the consequences for human performance.

This is the category most relevant to how Canadian businesses actually acquire AI capability, because in most cases they are not designing anything. They are enabling features in software they already licence, and the decisions about what is automated, how outputs are presented, what confidence is displayed and what verification is made easy or difficult were made by a vendor optimising for adoption.

The specific design choices that matter, on the research above, include whether the interface makes relevant supporting parameters easy to sample, since verification depth relates to commission errors[5], and the saliency of automation state indicators, which Parasuraman and Riley identify among the factors affecting monitoring of automation[1].

The procurement implication, which we offer as analysis, is that a buyer should evaluate an AI feature partly on whether its interface supports verification or discourages it. An output presented with its supporting evidence one click away supports a different review behaviour from one presented as a conclusion with the evidence buried. Neither is more accurate; they produce different human error rates, and that difference belongs in the evaluation.

What Actually Follows For Control Design

Synthesising the above into design principles, which is our own analysis grounded in the cited findings.

Stop relying on training as the mitigation for over-reliance. The literature is explicit that automation bias cannot be prevented by training and complacency cannot be overcome with simple practice[3]. Training remains necessary for competence; it is not a control for this risk.

Do not assign the control to expertise alone. Both effects occur in expert participants[3], so seniority of the reviewer is not the safeguard it is treated as.

Treat attentional load as the control variable. Complacency arises under multiple-task load[3], so the design question is how much unallocated attention the reviewer has, not how well-intentioned they are.

Separate controls for omission and commission. Review addresses what the system produced. Detecting what it failed to produce requires independent sampling of unflagged population, which is a different control entirely.

Set detection thresholds against base rates, not accuracy. Otherwise false alarms drive the system into disuse[1], and the disuse will not be visible in utilisation statistics.

Instrument verification depth rather than attestation. Sampling of relevant parameters is the variable associated with commission errors[5], and a sign-off records none of it.

Maintain manual capability deliberately. Out-of-the-loop unfamiliarity develops through sustained supervision rather than performance[5], and it is not addressed by anything in a normal control framework.

A Worked Case: The Reviewer Who Checked Three Fields

A Canadian finance team using an AI system to classify and code a high volume of supplier invoices. The reconstruction illustrates the mechanisms rather than reporting a specific engagement.

The control is that a qualified analyst reviews each classification before posting. The framework describes this as human oversight. Volume has roughly doubled since deployment, headcount is unchanged, and the analyst also handles supplier queries and period-end tasks.

Three of the mechanisms above operate simultaneously. Concurrent task load is high, which is the condition under which complacency arises[3]. The analyst is experienced, which the literature indicates does not confer protection[3]. And the throughput gain the deployment was justified on has consumed the attentional slack the review depends on.

In practice the analyst checks the vendor, the amount and the account code, and accepts. That is a small proportion of the relevant parameters, which is the variable the research associates with commission errors[5]. The system records an approval, which is the only thing the control framework measures.

Meanwhile the invoices the system classified with high confidence and that were never queried constitute the omission surface, and no part of the control examines them.

Nothing here involves a careless individual or a poorly written policy. Every element follows from documented properties of human interaction with automation, which is precisely why exhortation and training do not fix it and why the remedy has to be structural.

What To Do

Audit your AI framework for the word training. Where it appears as the mitigation for over-reliance, the evidence does not support it and a structural control is needed instead.

Reserve a defined share of the efficiency gain to fund verification. If none is reserved, the control was defunded at deployment.

Map review controls against concurrent task load. Identify which reviewers and which periods carry the most competing demands, and treat those as the weak points, because the research says they are.

Add an omission control. Independent sampling of items the system did not flag, at a defined rate, because review of outputs cannot detect absent outputs.

Recalculate alert thresholds against the base rate of the condition. Then measure true-positive rate, and treat a low one as a control failure in progress rather than as diligence.

Instrument what reviewers actually opened. Use it to find shallow-verification workflows, not to discipline individuals, since the cause is attentional conditions rather than diligence.

Run the override diagnostic. Trace what happens operationally when a reviewer rejects a routine output. If the process cannot absorb it, the review is nominal.

Schedule deliberate manual performance. Periodic completion of the underlying task without the system, to maintain the familiarity that supervision does not.

The Limits Of This Analysis

Several caveats matter. The cited studies are drawn from aviation, process control and laboratory monitoring tasks, and their generalisation to finance workflows using large language models is our inference rather than a demonstrated result; the mechanisms are general but the effect sizes in a finance setting are unmeasured. We accessed these papers through abstracts, publisher summaries and secondary citation rather than in full, and characterisations should be verified against the originals; in two instances, the accountability study and the teams study, we have the design and subject but not the findings and have said so rather than assert an outcome. The literature predates generative AI, whose failure modes differ from the deterministic decision aids studied, particularly in that a language model's errors can be fluent and internally plausible in ways a malfunctioning gauge is not, which may make commission errors more likely rather than less. The productivity paradox argument, the omission control recommendation, the verification instrumentation proposal, the procurement implication and the override diagnostic are our own analysis rather than findings in the cited work. This article does not address model accuracy, evaluation methodology, explainability, professional standards on the use of AI in assurance engagements, or the Canadian regulatory position on model risk, which this publication addresses elsewhere. Nothing here is a substitute for professional advice on control design in a specific environment.

Frequently Asked Questions

What is automation bias?
A pattern in which operators over-rely on automated systems, producing both omission errors, failing to detect problems the system did not flag, and commission errors, following incorrect automated recommendations. Commentary summarising the research notes that individuals are more likely to accept incorrect automated advice than to challenge it, even when contradictory information is available.
Won't training my team to be sceptical solve it?
The evidence says no. Parasuraman and Manzey's 2010 review reports that automation bias cannot be prevented by training and that automation complacency cannot be overcome with simple practice. Training remains necessary for competence but is not a control for this risk; the mitigations with research support are structural.
Does assigning review to senior staff help?
Not on the basis of expertise alone. The same review reports that both automation complacency and automation bias occur in naive and expert participants. Since the effect is not a knowledge deficit, subject-matter seniority does not confer resistance, and senior reviewers often carry the highest competing task load, which is the condition under which complacency arises.
Why does adding volume weaken the control?
Because complacency arises under multiple-task load, when manual tasks compete with the automated task for attention. The productivity gain from AI is realised by increasing concurrent load, which is the same resource the review control depends on. A firm cannot fully monetise the efficiency gain and fully retain the safeguard.
Why do detection systems get ignored?
Parasuraman and Riley attribute disuse to falsely activating alarms, often because the base rate of the condition was not considered when setting the trade-off between false alarms and omissions. For rare events like fraud or misstatement, an accurate detector still produces mostly false positives, operators learn alerts are usually wrong, and the control becomes an administrative queue.
What should I measure instead of sign-offs?
Verification depth. Research examined the proportion of relevant parameters sampled among participants who did and did not commit commission errors, which suggests how much supporting information a reviewer actually consulted is the informative variable. Most finance systems record only that someone clicked approve, which is precisely what commission errors look like.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article states where it has a study's design but not its findings, and declines to characterise results it has not read. See References below.

References

  1. Parasuraman, R., & Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), 230–253, on the four-part taxonomy, factors influencing use, over-reliance and monitoring failures, disuse caused by false alarms, and the base rate point in setting the false alarm and omission trade-off. Accessed via publisher abstract. journals.sagepub.com/doi/10.1518/001872097778543886
  2. Skitka, L. J., Mosier, K. L., & Burdick, M. (1999). Does Automation Bias Decision Making? International Journal of Human-Computer Studies, 51(5), 991–1006, on evidence of both commission and omission errors in an automated monitoring task with non-pilot samples; also citing Skitka, Mosier, Burdick and Rosenblatt on whether teams are better than individuals. Accessed via a repository record rather than the published article. researchgate.net/publication/222507469
  3. Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410, on complacency under multiple-task load, occurrence in naive and expert participants, that it cannot be overcome with simple practice, that automation bias produces omission and commission errors with imperfect aids, and that it cannot be prevented by training. Accessed via publisher and PubMed abstracts. journals.sagepub.com/doi/10.1177/0018720810376055
  4. Commentary citing Parasuraman and Riley (1997) and Skitka et al. (1999), on individuals being more likely to accept incorrect automated advice than to challenge it even when contradictory information is available, and on automated systems shifting from decision aids to implicit arbiters of judgment through routine use rather than formal transfers of authority. Note: secondary commentary accessed through a repository record. researchgate.net/publication/222507469
  5. Bahner, E., Hüper, A.-D., & Manzey, D. (2008). Misuse of Automated Decision Aids: Complacency, Automation Bias and the Impact of Training Experience. International Journal of Human-Computer Studies, 66, 688–699, on the proportion of relevant parameters sampled by participants who did and did not commit commission errors, and, in the same publisher record, on loss of situation awareness, out-of-the-loop unfamiliarity, and Funk and colleagues on complacency among the five most important issues of cockpit automation. Accessed via publisher abstract. sciencedirect.com/science/article/abs/pii/S1071581908000724
  6. Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration, repository record, on the integrated model and attentional synthesis as a framework for mitigating such effects in automated and decision support systems. researchgate.net/publication/47792928
  7. Molloy, R., & Parasuraman, R. (1996). Monitoring an Automated System for a Single Failure: Vigilance and Task Complexity Effects. Human Factors, 38, 311–322, cited within the Parasuraman and Manzey review for the automation complacency effect for a single automation failure. researchgate.net/publication/47792928
  8. Skitka, L. J., Mosier, K., & Burdick, M. D. (2000). Accountability and Automation Bias. International Journal of Human-Computer Studies, 52, 701–717, on the study's goal of exploring whether social accountability could ameliorate decision-making problems in highly automated environments, its measurement of commission and omission errors by accountability condition, and NASA Ames grant support. Note: we accessed the design and framing but not the results, and do not assert a finding. researchgate.net/publication/222529002
  9. Anonymous. "Computer Says No": Algorithmic Decision Support and Organisational Responsibility, arXiv preprint 2110.11037, on heuristics leading individuals to allocate responsibility for decision outcomes to automated systems, citing Parasuraman and Manzey (2010) and Mosier and Skitka (1996), and the analogy to diffusion of responsibility in groups. Note: a preprint, not peer-reviewed at the version accessed. arxiv.org/pdf/2110.11037

This article discusses peer-reviewed human factors research and is provided for general informational purposes. Papers were accessed via abstracts, publisher summaries and secondary citation rather than in full, and should be verified against the originals. The cited studies concern aviation, process control and laboratory monitoring tasks; generalisation to generative AI in finance workflows is the authors' inference. Nothing here is a substitute for professional advice on control design.