Ask any Canadian firm how it manages AI risk and the answer arrives within a sentence: a qualified person reviews the output. This series has already shown that the reviewer is subject to automation bias, is losing the skill the review depends on, and cannot see certain classes of defect at all. This article addresses the design of the review itself, and reports a result that makes the standard remedy counterproductive.

Key Takeaway

A 2024 Harvard Business School study of 228 evaluators, reported at second hand, found that human reviewers given AI recommendations with clear explanations were 19 percentage points more likely to align with those recommendations than a control group, and that when the AI also provided narrative rationales, deference increased by a further 5 points. Better explainability produced worse oversight, apparently because reviewers perceive less marginal value in redoing cognitive work the system has already performed. Commentary describes the common design as a category error rather than a bad implementation: a post-hoc supervisory arrangement in which the AI produces a substantive recommendation and the human is asked to review it, described elsewhere as a warm body in the loop providing the appearance of human control without its substance. A peer-reviewed framework proposes treating agency as layered, with AI operative agency in execution and human evaluative agency in verification, steering and substitution, and focusing on high-level explanations tied to external criteria rather than on the model's internal reasoning.

The Category Error

The framing that determines whether the standard design can be repaired or must be replaced.

One paper describes the common response to difficult AI tasks as proposing human-in-the-loop designs, meaning post-hoc supervisory designs in which the AI produces a substantive recommendation and the human is asked to review, approve or reject it, observing that this sounds reasonable and that the evidence suggests it fails. It cites automation bias research documenting that when AI systems provide recommendations, human overseers tend to accept them uncritically even when the recommendations are wrong, and states that this deep automation bias is not a matter of individual laziness but a structural consequence of the oversight architecture. Its conclusion is that the problem is not bad implementation but a category error[1].

The word post-hoc identifies the defect precisely. The human enters after the substantive work is complete, and is asked to evaluate a finished artefact produced by a process they did not participate in.

That is a harder cognitive task than performing the work, not an easier one. Evaluating whether a conclusion is correct generally requires reconstructing enough of the reasoning to test it, which is most of the effort of reaching it independently, and the reviewer is given a fraction of the time on the assumption that reviewing is lighter than doing.

The distinction the paper draws, between constitutive and supervisory involvement, is the useful one. Supervisory involvement inspects a result. Constitutive involvement is part of producing it. The sections below argue that workable designs are constitutive, and that the supervisory pattern fails for reasons no amount of training or diligence addresses.

The Explainability Paradox

The central finding, and the most counterintuitive result in this literature.

One source reports that a 2024 Harvard Business School study gave 228 evaluators AI recommendations with clear explanations of the AI's reasoning, and that human reviewers were 19 percentage points more likely to align with AI recommendations than the control group. When the AI also provided narrative rationales, explaining why it made a decision, deference increased by another 5 points. The source summarises this as better explainability producing worse oversight, and describes the human in the loop as having become a rubber stamp on a form[2].

It offers a mechanism: conventional wisdom says that showing humans the AI's reasoning improves scrutiny, and the data reverse this, because when the AI does the work of explaining its reasoning, reviewers perceive less marginal value in redoing that cognitive work themselves. Its conclusion is that better explainability can substitute for oversight rather than support it[2].

We report this at second hand. The figures reach us through a commentary rather than from the study, we did not access the original, and we do not have its design, population or measure of alignment. Readers should treat the specific numbers accordingly.

The direction of the finding is nonetheless consistent with the automation bias literature this publication has already examined, in which trust in an automated aid predicts reduced independent verification. An explanation is a trust-increasing artefact, so a result in which it reduces scrutiny is coherent with the established mechanism rather than anomalous.

If it holds, it undermines the most common organisational response to concerns about AI oversight, which is to require that systems explain themselves.

A Correction To Our Own Position

This finding cuts against something this publication has recommended, and we would rather address that directly than let it stand unexamined.

The article in this series on cognitive offloading reported a study in which the number of times students checked the AI's claims against sources predicted their gains, and concluded that AI helps most when designed to question rather than to answer. From that we recommended preferring configurations that show workings and cite sources. The article on function allocation made a related recommendation about intermediate automation levels in which the system surfaces options and reasoning.

The explainability finding appears to contradict that. Reviewers given explanations deferred more, not less.

We think the two are reconcilable, and the reconciliation sharpens both, but readers should judge the argument rather than take the resolution on trust.

The offloading study measured a behaviour: checking claims against sources. The explainability study measured the effect of receiving an explanation. Those are different things, and the difference is who does the work.

A citation is a pointer to work the reviewer must still perform. Opening the source, locating the passage and comparing it to the claim is effort that the citation creates rather than removes. A narrative rationale is a finished argument the reviewer can accept or reject without going anywhere, and accepting is cheaper.

So the distinction we would now draw is between artefacts that create verification work and artefacts that substitute for it. Sources, references and pointers are the first. Explanations, rationales and confidence statements are the second. Our earlier recommendation was right about citations and would have been wrong if extended to narrative explanation, and we are stating that rather than quietly narrowing it.

Internal And External Faithfulness

A peer-reviewed framing that supports the same distinction from a different direction.

One paper argues that instead of demanding low-level explanations and controls over how a complex AI model works internally, which it terms internal reasoning faithfulness, the focus should be on high-level explanations tied to external criteria and human expert understanding, which it terms external reasoning faithfulness, and that this approach retains the system's operative agency while strengthening the human's evaluative agency[3].

The distinction maps onto the reconciliation above and generalises it usefully.

An internal explanation is the system's account of its own process. The reviewer's task becomes assessing whether that account is persuasive, which is a task at which a fluent system will reliably succeed regardless of whether the underlying output is correct.

An external explanation locates the output against criteria the reviewer already holds. In professional work those criteria exist: a standard, a policy, a rubric, a set of tests a treatment must satisfy. Evaluating against them draws on the reviewer's own expertise rather than on their assessment of the system's rhetoric.

For a Canadian finance function this is directly actionable. Rather than asking a system to explain why it classified something, ask it to state which criteria it applied and what evidence satisfies each, in a structure the reviewer's own professional judgment can test.

The difference is that the first invites the reviewer to evaluate the model, and the second invites them to evaluate the work against a standard, which is what they are qualified to do.

Accountability Diffusion

A second mechanism, and the one that operates on incentives rather than on cognition.

The same commentary identifies accountability diffusion, noting that reviewers feel psychologically safer when they can attribute adverse outcomes to the AI's recommendation rather than to their own judgment[2].

The logic is worth spelling out because it is not about laziness. A reviewer who overrides a system and is wrong has made an error that is visibly theirs. A reviewer who accepts a system's recommendation and is wrong has participated in a failure that is distributed across a process, a vendor and a tool.

The two outcomes are not equally costly to the individual, even when they are equally costly to the firm, and a rational person responds to the asymmetry.

This has a specific consequence for professional services that we would flag, and it is our own analysis. In a firm where file review focuses on errors that were made rather than on defects that were accepted, the asymmetry is institutional rather than merely psychological. Overriding carries career risk that deference does not.

The remedy is not exhortation. It is to make acceptance visible as a decision: recording who accepted what, on what basis, in the same way an override would be recorded. If deference leaves a signature and dissent leaves a signature, the asymmetry narrows.

The Warm Body In The Loop

The phrase from the legal-academic critique, which names the failure precisely.

The paper discussing the category error cites a critique that names this as a warm body in the loop, providing the appearance of human control without its substance[1].

The formulation matters because it identifies what the arrangement produces. It produces compliance evidence. The framework says a human reviews; a human did review; the control is documented as operating.

One source makes the same point in operational terms, observing that most human-in-the-loop implementations do not produce oversight but produce paperwork[2].

For a Canadian firm the risk this creates is worse than having no control, and this is our own assessment. A firm without a review control knows it is exposed and manages accordingly. A firm with a documented review control that does not function believes it is covered, allocates no further attention, and has converted an open risk into a hidden one.

That is the same structure the automation bias article described for review controls generally and the drift article described for point-in-time validation. The recurring pattern across this series is that documented controls which do not operate are more dangerous than absent ones, because they terminate the enquiry.

Time, Information, Training, Authority

The most usable definition in the material, and a test a firm can apply this week.

One source states that regulators and auditors increasingly reject implementations where a human is nominally in the process but has no practical ability to understand the AI's reasoning, no time to conduct genuine review, or no authority to change or reverse the recommendation, describing that as rubber-stamping. It reports that this does not satisfy a requirement that oversight persons have the competence, authority and resources to carry out their function, and offers a definition: an architecture in which designated human reviewers are empowered, with adequate time, information, training and authority, to understand AI outputs, detect errors or anomalies, question the recommendation, and modify or override it before that recommendation takes effect in a consequential decision[4].

We note this source is a commercial blog and that the regulatory requirement it paraphrases is foreign to Canada.

The four elements are nonetheless a serviceable checklist independent of any regime, and each is checkable.

Time. How long does a reviewer actually have per item, and is that enough to reconstruct the reasoning rather than to read the conclusion? Divide review hours by items reviewed and compare it against how long the work took before automation.

Information. Does the reviewer receive the sources, the inputs and what the model actually read, or only the output? The prompt injection article in this series argued this specifically.

Training. Does the reviewer know the failure modes of the system they are reviewing, as distinct from the subject matter?

Authority. Can they override without escalation, and what happens to them when they do?

Our observation is that most firms can satisfy the third and fourth on paper and fail the first two in practice, because the time budget was set by the business case that assumed review was quick.

The Failure Runs Two Ways

A corrective to any reading of this article as an argument for more constraint.

The peer-reviewed framework states that oversight often fails in one of two ways: either humans act as rubber stamps, approving AI outputs they do not fully understand, or systems are constrained so tightly that the AI collapses into rule-following automation, losing the very qualities that make it useful[3].

The second failure is real and under-discussed. A firm that responds to oversight concerns by narrowing what a system may do until it is effectively a decision table has not achieved safe AI; it has achieved a rules engine with additional cost and latency.

That matters for the business case. Constraint is not free, and the value of the system was in the part being constrained away.

The framework's position is that human agency and AI agency are not in a zero-sum relationship, and its proposal is to retain operative agency while strengthening evaluative agency[3].

For a Canadian firm the practical reading is that the design question is not how much to restrict the system, but where the human's contribution should sit. That is the layered framing below.

Anchoring And The Order Of Operations

A mechanism from the evaluation literature with a clean design implication.

One paper observes that common approaches to incorporating human judgment either accidentally anchor human experts, leading to rubber-stamping, or leave them unsupported in high-variance tasks. It records the pattern as the rubber-stamp effect, in which humans under time pressure tend to believe the AI rather than providing critical oversight, and notes research showing this leads to blind agreement even when the judge is wrong. It describes an alternative in which humans provide assessments on examples without seeing AI scores, characterising this as a safe way to remove anchoring that nonetheless leaves the expert completely unsupported, since assigning a quality judgment to a lengthy response for a complex task is cognitively demanding[5].

Anchoring is a sequencing problem, and sequencing is cheap to change.

A reviewer who sees the system's answer before forming their own is evaluating a proposal. A reviewer who forms a view first, then compares, is performing an independent assessment. The same person, the same information, the same time, in a different order, produces a different epistemic act.

The paper's caution is equally important: removing the anchor without providing support produces an expert doing the whole task unaided, which is expensive and defeats the purpose of the system.

That tension is what the next section resolves.

The Constitutive Pattern

The design that escapes both horns, and the most useful pattern in this literature.

The same paper describes a prototype implementing a different division of labour: humans identify what information matters, while the models handle high-volume matching of that information to system outputs, which it says plays to each party's strengths while maintaining genuine human oversight[5].

The structural insight is that the human contributes before the output exists, by defining the criteria, and the system applies those criteria at volume.

This is constitutive rather than supervisory. The human's judgment is embedded in what the system is checking for, rather than applied to what the system produced.

Three properties follow that address the failures above. There is no anchoring, because the criteria are set before any output is seen. There is no rubber stamp, because the human is not approving anything. And the human's contribution is the part requiring expertise, while the machine performs the part requiring volume.

The translation for Canadian professional work is direct. Rather than reviewing each AI-produced file, a senior practitioner defines what a correct file must demonstrate: which tests must be satisfied, which evidence must be present, which conditions require escalation. The system applies that specification across the population and surfaces what fails it.

The reviewer's time then goes to defining criteria and examining failures, both of which are judgment tasks, rather than to reading outputs that are almost always fine, which is the vigilance task the human factors literature says people perform badly.

Layered Agency

The framework that generalises the pattern.

The peer-reviewed proposal treats agency as layered: AI operative agency in task execution, and human evaluative agency in verification, steering and substitution[3].

The three verbs describe distinct interventions and are worth separating for design purposes.

Verification is confirming that an output meets the criteria, which is the activity the constitutive pattern automates against human-defined standards.

Steering is adjusting how the system operates rather than correcting a single output. It is the higher-leverage intervention and the one most review designs have no channel for, because a reviewer who notices a recurring problem has no route other than fixing each instance.

Substitution is the human taking over the task, which is the escape hatch, and it depends on the capability the deskilling article argues is eroding.

A related literature distinguishes meaningful human control, focused on strategic guidance, from effective human control, focused on operational execution, and argues that the vision is achieved only through the mutually reinforcing integration of both[6].

The practical instruction we would draw is that a firm should check which of the three verbs its design supports. Most support verification only, and steering is the one that would actually reduce error rates over time.

The Spectrum And Its Movement

The standard taxonomy, and the point that it should not be fixed.

One source sets out the spectrum: human-in-the-loop, where a person must approve or reject every decision, described as best for high-risk or high-consequence tasks including financial approvals; and human-on-the-loop, where the system operates autonomously while a person monitors and can intervene, described as suitable for moderate-risk tasks where speed matters but oversight is still required[7].

It adds that the oversight model should be reviewed regularly, since as the system matures and error rates change some workflows may move from the first to the second, while emerging failure modes may require tightening, concluding that the goal is a living process rather than a one-time setup[7].

That is the third literature in this series to reach dynamic allocation. The function allocation article found that all static allocation methods share the drawback of being static, and the drift article found that validation evidence describes a configuration rather than a system. Here the same conclusion arrives from oversight design.

The convergence is strong enough to state as a principle: any control decision that fixes the relationship between human and system at deployment will be wrong later, in a direction nobody is watching.

The practical version for a mid-sized Canadian firm is a scheduled review of the oversight level per workflow, with defined triggers for tightening, which is the same cadence argued for in the drift and incident articles and can be the same meeting.

Measuring Whether Oversight Is Real

The instrumentation, which is simpler than the problem suggests.

One source names the failure and the countermeasure in a sentence: rubber-stamping is reviewers approving everything without meaningful evaluation, and it should be combated by tracking rejection rates and reviewing a sample of approvals[7].

Both measures are available to any firm without tooling.

Rejection rate is the proportion of outputs a reviewer changes or rejects. A rate near zero is either a system that is never wrong or a control that is not operating, and the two are distinguishable only by the second measure.

Sampling approvals means independently examining a sample of items that were approved. This is the only measure that detects the specific failure of interest, because it looks at what passed rather than at what was caught.

We would add two of our own, drawn from earlier articles in this series. Record acceptance as a decision with a stated basis, which addresses accountability diffusion and cannot be done inattentively. And measure time per item against the time the work took before automation, which tests the first element of the four-part test.

The same source recommends establishing a mechanism for reviewers to report errors and patterns back to the team maintaining the system, tracking what types of errors occur and whether they are improving, and using this to refine the system, with these loops maturing into structured evaluations that measure output quality against human-defined criteria[7]. That is the steering channel the layered framework identifies and most designs lack.

When Contaminated Judgments Calibrate The Judge

A second-order effect that matters for any firm using AI to evaluate AI.

The evaluation paper notes that when judgments contaminated by the rubber-stamp effect are used to train or calibrate an automated judge, the errors compound[5].

The mechanism is a feedback loop. Human reviewers anchored on the system's output produce approvals that agree with it. Those approvals are then used as ground truth to calibrate an automated evaluator. The evaluator learns to agree with the system, and its agreement is subsequently cited as independent confirmation.

The hallucination measurement article in this series recommended calibrating an automated judge against human labels before trusting it. This finding qualifies that recommendation importantly: the human labels must be produced without the reviewer having seen the system's answer, or the calibration inherits the anchoring.

That is a small change in procedure with a large effect on what the calibration means, and it is easy to get wrong, because the natural way to produce labels is to have someone check existing outputs.

The correct procedure is the unanchored one the evaluation paper describes: assessments made on examples without seeing the system's scores[5], accepting that this is more expensive and is the only version that produces an independent signal.

Safe To Question

The organisational condition, which the frameworks name and which no amount of design substitutes for.

One review states that meaningful oversight requires more than nominal human review, depending on human-centred interface design, realistic workload management, lifecycle-oriented technical controls, and organisational cultures that make it safe to question AI outputs, with the goal of architecting interaction that amplifies rather than erodes professional judgment[8].

We note this source is a preprint explicitly labelled as not peer reviewed.

The cultural element connects to accountability diffusion. If overriding the system is treated as an inconvenience, if a reviewer who slows a process must justify the delay, or if the deployment was announced as a success before anyone reviewed anything, then the incentives point toward deference regardless of what the procedure says.

The realistic workload point is the same one the four-part test raises about time, and the human factors literature explains why it is not merely a resourcing matter: complacency arises under multiple-task load, so a reviewer given more items to review is not simply rushed but is subject to a documented degradation in the capacity the control depends on.

For a Canadian firm the diagnostic question is whether anyone has ever been thanked for rejecting an AI output. If the answer is no, the culture condition is not satisfied whatever the policy says.

The Readiness Gap

Context on how widely this is about to matter, reported at second hand.

One article cites a 2026 enterprise survey reporting that 74% of companies plan to deploy AI agents within two years, while just 21% of respondents said they have a mature model for governing AI agents, observing that this influx may be outpacing the ability to manage it, and that many organisations' answer is to implement a human-in-the-loop strategy[9].

We report the figures as stated and did not access the underlying survey, which concerns enterprises rather than the mid-sized Canadian firms this publication addresses.

The relevant observation is the sequence those two numbers describe. A large majority intend to deploy, a small minority have a governance model, and the gap is being filled by the design pattern this article's sources describe as a category error.

That is not an argument against deployment. It is an argument for treating the oversight design as a first-order engineering problem rather than as the thing you assert in the governance document after the build.

The Canadian Position

Where a Canadian reader stands, since the obligations cited are foreign.

The instruments named across these sources include a European requirement that oversight persons have competence, authority and resources, a United States state statute effective February 2026 requiring meaningful human oversight of consequential decisions about that state's residents, and federal United States guidance in credit, employment and clinical contexts[4].

None is Canadian, and Canada has no general statutory human oversight requirement for AI-assisted decisions.

The Canadian position is nonetheless not empty, and this is our own analysis. Professional obligations require that work be performed with due care and that conclusions be supported. Where a member of a professional body signs work, they are accountable for it regardless of what produced the draft, and a review that did not function is not a defence.

So the operative question for a Canadian practitioner is not whether a regulation requires meaningful oversight, but whether the review they performed would support the assertion they made if examined.

The four-part test is a reasonable way to answer that: did you have the time, the information, the knowledge of the system's failure modes, and the authority to change it. Those questions are answerable in advance and are the ones that would be asked afterwards.

A Worked Case: Two Review Designs

A Canadian firm reviewing AI-assisted classifications across a large population. The reconstruction illustrates the design choice rather than reporting a specific engagement.

Design A is the standard pattern. The system classifies each item and produces an explanation of its reasoning. A reviewer opens each item, reads the classification and the explanation, and approves or rejects.

This is post-hoc supervisory involvement[1]. The reviewer is anchored, having seen the answer before forming a view[5]. The explanation raises deference rather than scrutiny[2]. Acceptance is cheaper than override both cognitively and in accountability terms[2]. The rejection rate will be low and will be reported as evidence the system works.

Design B is constitutive. A senior practitioner specifies, before any output exists, what a correct classification must demonstrate: which criteria apply, what evidence satisfies each, and what conditions require escalation. The system applies the specification across the population and surfaces items that fail it or that fall near a boundary[5].

The practitioner's time goes to defining criteria and examining failures. There is no anchoring, because the criteria preceded the outputs. There is no rubber stamp, because nothing is being approved. The criteria are external, so the reviewer evaluates against a standard they hold rather than against the model's account of itself[3].

Both designs consume similar senior time. Only one of them produces a signal, and the firm running Design A will have better-looking control documentation.

What To Do

Move the human before the output, not after. Define criteria first and have the system apply them, rather than producing results for someone to inspect.

Be careful what you ask the system to explain. Sources and citations create verification work; narrative rationales substitute for it, and the evidence indicates the second increases deference.

Prefer external criteria to internal reasoning. Ask which standards were applied and what evidence satisfies each, in a form the reviewer's own expertise can test.

Remove the anchor where you can. A reviewer who forms a view before seeing the system's answer is performing a different act from one who does not.

Apply the four-part test. Time, information, training on failure modes, and authority to override. Most firms pass the last two and fail the first two.

Track rejection rate and sample approvals. A near-zero rejection rate is uninformative on its own, and only examining what passed detects the failure of interest.

Record acceptance as a decision with a basis. It narrows the accountability asymmetry and cannot be done inattentively.

Build a steering channel. A reviewer noticing a recurring problem needs a route other than fixing each instance.

Produce calibration labels unanchored. If human labels used to calibrate an automated judge were made while looking at the system's output, the calibration inherits the anchoring.

Review the oversight level on a schedule. Three separate literatures now agree that a fixed allocation becomes wrong in a direction nobody is watching.

The Limits Of This Analysis

Several caveats matter. The headline figures, being the 19 percentage point alignment increase and the further 5 points from narrative rationales, reach us through a commentary rather than from the 2024 study itself; we did not access the original and do not have its design, population or measure, and readers should treat the numbers as reported rather than verified. Sources are mixed: one peer-reviewed journal article, several preprints not peer reviewed at the versions accessed, one explicitly labelled as not peer reviewed, a practice library, a commercial blog and a trade publication. The enterprise survey figures on deployment and governance maturity are reported at second hand and concern large enterprises. Every regulatory instrument cited is foreign to Canada, and no Canadian general requirement for human oversight of AI-assisted decisions exists. Our reconciliation between the explainability finding and this publication's earlier recommendation on citations is an argument rather than a demonstrated result, and readers should judge it. The accountability asymmetry in professional file review, the assessment that a documented non-functioning control is worse than none, the four-part test application, the Canadian professional obligation argument and the worked case are our own analysis. This article does not address explainability techniques in technical detail, interface design methodology, professional liability, or the specific requirements of Canadian professional bodies. Nothing here is a substitute for professional advice on control design in a specific practice.

Frequently Asked Questions

Does making AI explain itself improve oversight?
The reported evidence says the opposite. In a study of 228 evaluators, reviewers given clear explanations were 19 percentage points more likely to align with AI recommendations than a control group, with narrative rationales adding another 5 points. The proposed mechanism is that when the system does the explaining, reviewers perceive less value in redoing that work themselves.
Doesn't that contradict your earlier advice about citations?
It appears to, and we address it directly rather than narrowing it quietly. The distinction we would now draw is between artefacts that create verification work and artefacts that substitute for it. A citation is a pointer to work the reviewer must still do. A narrative rationale is a finished argument they can simply accept. Our earlier recommendation holds for sources and would have been wrong extended to explanation.
What is wrong with the standard review design?
It is post-hoc and supervisory: the system produces a substantive recommendation and the human inspects it. Commentary describes this as a category error rather than a bad implementation, and the resulting arrangement has been called a warm body in the loop, providing the appearance of human control without its substance. Evaluating a finished artefact is harder than producing it, not easier.
What does a better design look like?
Constitutive rather than supervisory. A senior practitioner defines, before any output exists, what a correct result must demonstrate, and the system applies that specification at volume and surfaces what fails. There is no anchoring because criteria preceded outputs, no rubber stamp because nothing is being approved, and the human's time goes to judgment rather than to reading outputs that are almost always fine.
How do we tell whether our review is real?
Track the rejection rate and independently examine a sample of approvals. A near-zero rejection rate is either a system that is never wrong or a control that is not operating, and only sampling what passed distinguishes them. We would add recording acceptance as a decision with a stated basis, and comparing review time per item against how long the work took before automation.
Does this apply in Canada?
Not as statutory obligation. Every instrument cited in this literature is foreign, and Canada has no general requirement for human oversight of AI-assisted decisions. But a professional who signs work is accountable for it regardless of what produced the draft, and a review that did not function is not a defence, so the four-part test is worth answering in advance.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article reports a finding that cuts against a recommendation this publication made in an earlier piece, and sets out the reconciliation as an argument for readers to judge rather than as a settled resolution. See References below.

References

  1. Metis AI: The Overlooked Middle Zone Between AI-Native and World-Movers. arXiv preprint 2605.14407, on post-hoc supervisory designs and why human-in-the-loop fails, automation bias research documenting uncritical acceptance of recommendations even when wrong, deep automation bias as a structural consequence of oversight architecture rather than individual laziness, the cited critique naming a warm body in the loop, and the conclusion that this is a category error. Note: preprint; the underlying works were not accessed directly. arxiv.org/pdf/2605.14407
  2. TianPan.co. (2026, May 6). The HITL Rubber Stamp Problem, reporting a 2024 Harvard Business School study of 228 evaluators in which reviewers were 19 percentage points more likely to align with AI recommendations than a control group, with a further 5 points when narrative rationales were provided; the explainability paradox mechanism; accountability diffusion; and the observation that most implementations produce paperwork rather than oversight. Note: a commentary blog; the study was not accessed and its design, population and measure are unknown to us. tianpan.co/blog/2026-04-15-human-in-the-loop-rubber-stamp
  3. Designing Meaningful Human Oversight in AI. AI and Ethics, 2026, on the two-sided failure of rubber-stamping versus over-constraint collapsing AI into rule-following automation, the layered agency framework of AI operative agency and human evaluative agency in verification, steering and substitution, and the distinction between internal and external reasoning faithfulness. link.springer.com/article/10.1007/s43681-026-01147-7
  4. AI Buzz. (2026, June 6). Human-in-the-Loop Explained: When and How to Use It, on regulators and auditors rejecting nominal review, the competence, authority and resources requirement, the four-element definition of adequate time, information, training and authority, and the list of foreign instruments including a state statute effective February 2026 and federal guidance in credit, employment and clinical contexts. Note: a commercial blog; regulatory characterisations reported at second hand and all instruments foreign to Canada. aibuzz.blog/human-in-the-loop-explained
  5. Dietz, L. Human-in-the-Loop Nugget Annotation for Accountable LLM-as-a-Judge Evaluations. arXiv preprint 2606.29033, on approaches either anchoring experts into rubber-stamping or leaving them unsupported, the rubber-stamp effect under time pressure and blind agreement even when the judge is wrong, the compounding of errors when contaminated judgments calibrate an automated judge, the unanchored manual test set approach and its cost, and the division of labour in which humans identify what information matters while models handle high-volume matching. Note: preprint. arxiv.org/pdf/2606.29033
  6. Human-controllable AI: Meaningful Human Control. arXiv preprint 2512.04334, on the distinction between meaningful human control focused on strategic guidance and effective human control focused on operational execution, and the argument that both are required through mutually reinforcing integration. Note: preprint. arxiv.org/pdf/2512.04334
  7. Open Practice Library. Human in the Loop, on the spectrum from human-in-the-loop where a person approves or rejects every decision to human-on-the-loop where a person monitors and can intervene, rubber-stamping and the countermeasures of tracking rejection rates and reviewing a sample of approvals, feedback mechanisms maturing into structured evaluations, and the recommendation to review the oversight model regularly as a living process. openpracticelibrary.com/practice/human-in-the-loop
  8. From Human Oversight to Human-in-the-Loop, Preprints.org, posted 28 February 2026, on meaningful oversight requiring more than nominal human review and depending on human-centred interface design, realistic workload management, lifecycle-oriented technical controls and organisational cultures that make it safe to question AI outputs. Note: explicitly labelled by the platform as not peer reviewed. preprints.org
  9. TechTarget. Human-in-the-Loop Shouldn't Rubber-Stamp Decisions, reporting a 2026 enterprise survey finding that 74% of companies plan to deploy AI agents within two years while 21% said they have a mature model for governing them, and the observation that deployment may be outpacing management capacity. Note: a trade publication; the survey was not accessed and concerns large enterprises. techtarget.com/it-strategy/feature/Human-in-the-loop-shouldnt-rubber-stamp-decisions

This article discusses oversight design research and is provided for general informational purposes. Its headline figures reach us through commentary rather than from the underlying study, several sources are preprints or commercial publications, and every regulatory instrument cited is foreign to Canada. Nothing here is a substitute for professional advice on control design.