An AI expense tool flags a founder's meal charges as "excessive relative to category average" three weeks into the worst cash crisis the business has faced. The charges were real: dinners with a key client the founder was actively trying to keep from walking, spent specifically because the business could not afford to lose that account. The categorization was accurate by every measure the system was built to apply. It was also, in the moment, exactly the wrong thing for a stressed founder to see flagged as a problem, and neither fact cancels the other out.

Key Takeaway

Two well-replicated but seemingly contradictory findings coexist in the research literature: algorithm aversion, where people reject algorithmic advice after seeing it err even when it outperforms humans, and algorithm appreciation, where people prefer algorithmic judgment to human judgment. Charlotte Morewedge's 2022 framework resolves the apparent contradiction: which effect dominates depends on the ambiguity of the evaluative criteria and how identity-relevant the judgment is to the person receiving it, not on the algorithm's actual accuracy. Routine categorization tasks, where correctness is unambiguous and identity is not implicated, are exactly where algorithm appreciation and full automation both make sense. Financial judgment calls made during genuine business distress are exactly where the opposite holds, and treating both categories with the same automation policy is the mistake this article addresses.

What AI Bookkeeping Tools Actually Do Well

It is worth starting from a position of fairness to the technology, because the rest of this article argues for a real limit on it, and that argument is only credible if it does not overstate the limit. Modern AI-driven bookkeeping and expense tools are genuinely strong at optical character recognition and transaction categorization, reconciliation matching across bank feeds and ledgers, and statistical anomaly detection across large transaction volumes, tasks that are repetitive, high-volume, and governed by comparatively unambiguous rules. A receipt either matches a transaction amount or it does not. A vendor either appears in historical spend patterns or it does not. These are exactly the conditions under which automated systems reliably outperform manual human review, both in speed and, frequently, in accuracy, since humans doing repetitive categorization work are prone to exactly the kind of fatigue-driven error that a consistent algorithm does not exhibit.

Algorithm Aversion, And Its Opposite

Dietvorst, Simmons and Massey's influential 2015 study in the Journal of Experimental Psychology: General documented what they termed algorithm aversion: across a series of forecasting experiments, participants who saw an algorithm make an error lost confidence in it and reverted to their own, less accurate judgment, even when directly shown the algorithm outperformed them on average[1]. Critically, participants who saw a human forecaster make an equivalent error did not show the same loss of confidence, a human error was tolerated in a way an identical algorithmic error was not.

A separate, equally well-supported line of research finds the opposite pattern. Logg, Minson and Moore's 2019 study, "Algorithm Appreciation: People Prefer Algorithmic to Human Judgment," found that laypeople, across several judgment domains, actually favoured algorithmic advice over advice from another person, and did so more strongly than experts in the same domains did[2]. Both findings are well replicated. Neither is wrong. The honest question is not which effect is real, but what determines which one governs a given decision.

Why Domain Matters: The Uncertainty Threshold

Dietvorst and Bharti's 2020 follow-up research offers a first-pass answer: algorithm aversion correlates with the underlying uncertainty of the task itself[3]. Tasks that are genuinely stochastic, where even a perfect model cannot achieve certainty, such as predicting stock prices or human behaviour, trigger stronger aversion than tasks with a knowable, verifiable correct answer. This tracks intuitively: it is easier to trust an algorithm on a question that has a checkable right answer than on one where the algorithm's confidence and the decision-maker's own intuition are, in principle, equally uncertain guesses dressed up with different degrees of apparent authority.

A parallel and independently useful finding comes from outside finance entirely. Longoni, Bonezzi and Morewedge's 2019 research on patient resistance to medical AI found that people resist algorithmic recommendations more strongly specifically in domains they perceive as uniquely human or requiring subjective, individualized judgment, even when the algorithm's diagnostic accuracy is demonstrably equal to or better than a physician's[4]. The mechanism identified was not a rational assessment of the algorithm's competence; it was a belief that the domain itself required attention to the patient's unique, unstandardizable circumstances that a general model could not, in principle, weigh correctly.

The Domain-Specific Self-Evaluation Framework

Carey Morewedge's 2022 theoretical synthesis pulls these threads into a more precise, generalizable framework, and it is the one this article treats as the operative model[5]. Morewedge argues algorithm aversion is not domain-general, a blanket human bias against machines, but is instead driven by biased self-evaluation that is highly domain-specific, modulated by two identifiable factors: how identity-relevant the judgment is to the person making it, and how ambiguous the evaluative criteria for a correct answer are.

Applied directly to a finance function, this framework generates a testable prediction, not just an intuition. Where the evaluative criteria are unambiguous (a receipt total either matches the ledger entry or it does not) and identity-relevance is low (categorizing a supplier invoice says nothing about the reviewer's competence or judgment as a business operator), Morewedge's framework predicts, and the appreciation literature confirms, that people readily accept and even prefer automated judgment. Where the criteria are genuinely ambiguous (was this discretionary spending during a crisis wise or reckless?) and identity-relevance is high (the answer implicates the founder's own judgment about how to run their business under pressure), the same framework predicts exactly the resistance, and exactly the discomfort, that a purely accuracy-focused automation strategy will keep running into, no matter how good the underlying model becomes.

What "Feeling Panic" Actually Means For A Decision

This publication's earlier article on the neurochemistry of overspending described, in some detail, what acute financial stress actually does to a decision-maker: measurable impairment of prefrontal cortex function, the brain region responsible for weighing a decision against a broader goal, alongside heightened activity in reward and pain-anticipation circuitry around spending decisions specifically[6]. The title of this article is not merely evocative; it points at a literal, mechanistic gap. An AI categorization system optimizes its output against the query and the data it is given. It has no channel by which to detect that the human on the other end of that output is, at this specific moment, operating with measurably reduced capacity to evaluate a flagged anomaly calmly, and no mechanism to adjust its framing, tone, or urgency accordingly. A human colleague, by contrast, can often tell from a founder's voice, or from the pattern of their recent behaviour, that this is not the week to raise a minor spending flag as an urgent concern, and can choose to hold it, soften it, or contextualize it. This is not a claim that current AI systems merely execute this adjustment poorly; it is a claim that the input required to make the adjustment, a read on the human's actual psychological state, is not part of what these systems are built to perceive at all.

The Empathy Gap In Practice

This produces a specific, recurring failure pattern worth naming directly. An AI system trained to flag deviations from historical spending patterns will, correctly by its own logic, flag exactly the kind of behaviour a business under acute stress is most likely to exhibit: unusual vendor payments made to salvage a relationship, larger-than-typical client entertainment spend during a retention push, accelerated payment terms offered to a key customer to keep cash flowing both ways. These are, statistically, anomalies. They are also, in many cases, precisely the judgment calls a competent operator makes under pressure, and flagging them with the same tone and urgency used for a genuinely suspicious transaction adds a layer of algorithmic second-guessing to a decision-maker who is already operating at reduced cognitive capacity, at exactly the moment that second-guessing is least helpful and most corrosive to confidence.

Five Drivers, Not One Cause

A systematic review by Burton, Stein and Jensen identified five distinct drivers behind algorithm aversion, and disaggregating them is useful because each suggests a different fix rather than a single generic solution[8]. False expectations arise when a tool is marketed or perceived as more infallible than it actually is, so any error reads as a broken promise rather than an expected, bounded failure rate. Lack of perceived control reflects the discomfort of being unable to intervene in or adjust an automated process. Cognitive incompatibility describes situations where the algorithm's output does not fit the mental model the user already has for the task. Lack of incentives covers cases where the user has no personal stake in the algorithm succeeding, and may even be incentivized to distrust it. Divergent rationalities describes cases where the algorithm optimizes for a genuinely different objective than the one the human cares about in the moment.

Applied to the client-dinner example above, at least three of these five drivers are simultaneously in play: the founder had no perceived control over the flag being raised, no cognitive compatibility between "statistical deviation from category average" and "deliberate retention strategy," and a divergent rationality, since the system optimized for pattern consistency while the founder was optimizing for keeping a key account alive. Naming which of the five drivers is operating in a specific case points directly at the fix: perceived-control problems are addressed by giving users a lightweight override or annotation mechanism, cognitive-incompatibility problems are addressed by better framing and context in how a flag is presented, and divergent-rationality problems are addressed, as this article argues throughout, by keeping a human in the loop for exactly the judgment calls where the system's objective and the user's objective are most likely to diverge.

A Worked Case: The Flagged Expense

Return to the opening scenario. The founder, mid-crisis, received an automated notification flagging the client dinners as "23% above category average, review recommended." Nothing in the notification distinguished between "this might be fraud" and "this is unusual but plausibly a deliberate, reasonable choice," because the system had no mechanism for making that distinction, only a statistical deviation to report. The founder, already stretched thin, spent twenty minutes drafting a defensive explanation to nobody in particular, before recognizing there was no human on the other end actually asking the question.

The structural fix here is not to disable the flag, the underlying pattern-detection is genuinely useful and would catch a real problem if one existed. It is to route ambiguous-criteria, identity-relevant flags to a human for framing before they reach the founder, rather than delivering the raw statistical output directly. A bookkeeper or fractional controller reviewing the same flag can recognize the client-retention context in thirty seconds and either dismiss it or raise it with appropriate, calibrated framing, adding a layer of judgment the automated system was never built to provide, precisely at the point where Morewedge's framework predicts it is most needed.

Where Full Automation Is Actually Fine

The fair complement to the argument above is a clear statement of where it does not apply, because overcorrecting into blanket distrust of AI bookkeeping tools would be its own error, and one the algorithm aversion literature specifically warns against as economically costly. Routine transaction categorization, bank feed reconciliation, receipt-to-ledger matching, and statistical fraud screening across high volumes are low-ambiguity, low-identity-relevance tasks by Morewedge's own criteria, and the appreciation research suggests people, correctly, do not resist automation here once they trust the tool's baseline accuracy. A business owner who insists on personally reviewing every routine categorized transaction is not exercising prudent oversight; per the framework above, they are exhibiting exactly the identity-driven algorithm aversion the research describes, in a domain where it produces no benefit and real cost.

Letting People Adjust The Algorithm

A related, practically important finding from Dietvorst's own follow-up research offers a second, complementary fix. In experiments where participants were given even limited ability to modify an imperfect algorithm's forecasts, such as adjusting its output within a small, bounded range, they were substantially more willing to use the algorithm at all, even though the algorithm's underlying accuracy had not changed, and their own modifications did not, on average, improve on the algorithm's raw output[3]. The mere ability to exercise some agency over the final number was sufficient to overcome much of the aversion that a fully rigid, take-it-or-leave-it algorithmic output would have triggered.

This has an underused implication for financial software design and procurement specifically: a tool that allows an annotated override, "I reviewed this flag and here is why it's not a concern," logged and retained rather than simply silenced, is likely to see substantially higher genuine adoption than a tool that only offers accept-or-ignore, even if the two tools are functionally identical in every other respect. For a business evaluating AI bookkeeping or expense-management software, the presence of a genuine, lightweight override-and-annotate mechanism is worth weighing as seriously as the tool's raw categorization accuracy, because the research suggests it materially affects whether the tool actually gets used as intended rather than quietly worked around.

The 70% Threshold

One further, practically useful finding from the trust-in-automation literature deserves specific mention. Research on the "perfect automation schema" has identified a threshold effect: automated advice with accuracy at or below roughly 70% tends to be perceived as no longer meaningfully "near-perfect," triggering a sharper drop in trust than the accuracy decline alone would predict, while advice above that threshold retains disproportionate trust[7]. This has a direct implication for how AI tools should be introduced into a finance function: rather than presenting a new tool's occasional errors as isolated incidents to be explained away, which risks trust collapsing sharply once several errors accumulate near that threshold, it is more durable to set accuracy expectations honestly and specifically from the outset, on a per-task-category basis, so that a known, bounded error rate does not read as a broken promise the first time it is observed in practice.

A Note For Students Of Human-Computer Interaction

The coexistence of algorithm aversion and algorithm appreciation is a useful teaching case for anyone studying decision science or human-computer interaction, because it illustrates a broader methodological lesson: two well-designed, well-replicated experimental literatures can appear to contradict each other while both being correct, once the boundary conditions distinguishing them are properly specified. The mistake is not in either finding; it is in citing either one as a general truth about "how humans respond to algorithms" without specifying the domain. Financial software design sits precisely at the intersection these literatures were built to explain, spanning both the low-ambiguity, low-identity-relevance tasks where appreciation dominates and the high-ambiguity, high-identity-relevance judgment calls where aversion, appropriately, persists.

The Right Division Of Labour

Pulling the argument together into a practical rule: assign tasks to automation based on the ambiguity of their evaluative criteria and their identity-relevance to the human involved, not based on the algorithm's raw accuracy alone, since accuracy is necessary but not sufficient for appropriate automation. Categorization, reconciliation, and pattern-based fraud screening at volume belong to AI, with human spot-checks calibrated to the 70% trust threshold rather than to anxiety. Judgment calls about discretionary spending during genuine business stress, decisions that implicate the operator's own competence and carry ambiguous rather than checkable correctness, belong with a human, ideally one positioned to add context and calibrated framing before a raw statistical flag reaches a founder who may be in no state to receive it undefended.

A Word On How These Tools Get Sold

It is worth naming a tension that sits underneath much of this discussion. AI bookkeeping and finance tools are frequently marketed with language emphasizing exactly the qualities Morewedge's framework predicts will trigger the strongest resistance in ambiguous, identity-relevant domains, positioning the tool as an all-seeing, tirelessly vigilant guardian of the business's finances, capable of catching what a human would miss. This marketing is not dishonest about the tool's pattern-detection capability, which is often genuinely strong, but it invites exactly the false-expectations driver of aversion identified by Burton and colleagues: a tool sold as an infallible guardian will be judged far more harshly for the first ambiguous, context-dependent flag it gets wrong than a tool sold honestly as a pattern-detection assistant that surfaces statistical outliers for a human to interpret.

Businesses evaluating these tools are better served treating vendor marketing language as a signal of how the tool will fail trust over time, rather than as an accurate description of its capabilities. A vendor whose marketing is calibrated to the tool's actual, bounded competence, explicit about what it does and does not evaluate, is setting up a healthier long-term trust relationship than one promising comprehensive financial vigilance a pattern-matching system cannot, in principle, deliver.

The Limits Of This Analysis

Several caveats apply. Much of the foundational algorithm aversion and appreciation research was conducted using forecasting and judgment tasks in laboratory or online settings rather than live financial software used under genuine business stress, and the extension to bookkeeping and expense-flagging tools specifically is a reasonable application of a well-established general framework rather than a directly measured finding in that exact context. Morewedge's domain-specificity framework, while an influential and coherent synthesis, is one interpretation among several competing accounts of why algorithm aversion and appreciation coexist, and the field has not fully converged on a single explanatory model. The claim that current AI systems cannot perceive a user's psychological state is accurate for the categorization and flagging tools discussed throughout this article; it should not be read as a permanent, categorical claim about all future AI systems, some of which are being explicitly designed to infer user affect from behavioural signals, a development this article's central argument would need to revisit if it becomes reliable and widely deployed in financial software specifically.

Frequently Asked Questions

Is algorithm aversion or algorithm appreciation the "correct" finding?
Both are real, well-replicated findings from separate, credible research programs. Which one governs a given decision depends on the task, specifically how ambiguous the correctness criteria are and how identity-relevant the judgment is to the person involved, per Carey Morewedge's 2022 synthesis, rather than on which finding is generally "more true."
Does this mean AI bookkeeping tools are unreliable?
No. For routine categorization, reconciliation, and statistical anomaly detection at volume, these tools are genuinely strong and typically outperform manual review. The limitation described in this article is specific to ambiguous, identity-relevant judgment calls, particularly under acute stress, not to the technology's core competency.
Why would a founder resist an accurate AI flag?
Research on medical AI resistance found people resist algorithmic judgment more in domains they perceive as requiring individualized, subjective understanding of their unique circumstances, even when the algorithm is objectively accurate. A discretionary spending decision made under business stress is exactly this kind of domain: the correctness of the decision depends on context an automated flag cannot see.
What's the 70% trust threshold?
Research on trust in automated advice has found that accuracy at or below roughly 70% tends to trigger a disproportionately sharp loss of trust, since it falls below what users perceive as "near-perfect" performance. Setting honest, specific accuracy expectations for a new AI tool from the outset helps prevent trust from collapsing the first time a bounded, expected error occurs.
How should a business actually divide tasks between AI and humans in finance?
Assign high-volume, low-ambiguity, low-identity-relevance tasks (categorization, reconciliation, routine fraud screening) to automation. Route ambiguous, identity-relevant judgment calls, especially ones likely to arise during genuine business stress, to a human who can add context and calibrated framing before a raw statistical flag reaches the person it concerns.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article draws on peer-reviewed human-AI trust research; see References below.

References

  1. Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm Aversion: People Erroneously Avoid Algorithms after Seeing Them Err. Journal of Experimental Psychology: General, 144(1), 114-126.
  2. Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm Appreciation: People Prefer Algorithmic to Human Judgment. Organizational Behavior and Human Decision Processes, 151, 90-103.
  3. Dietvorst, B. J., & Bharti, S. (2020). People Reject Algorithms in Uncertain Decision Domains Because They Have Diminishing Sensitivity to Forecasting Error. Psychological Science, 31(10), 1302-1314.
  4. Longoni, C., Bonezzi, A., & Morewedge, C. K. (2019). Resistance to Medical Artificial Intelligence. Journal of Consumer Research, 46(4), 629-650.
  5. Morewedge, C. K. (2022). Preference for Human, Not Algorithm Aversion. Trends in Cognitive Sciences, 26(10), 824-826.
  6. The Insight Bureau. (2026). The Neurochemistry Of Overspending: Why Smart Founders Make Irrational Fiscal Choices Under Stress. GSH Financial.
  7. Madhavan, P., & Wiegmann, D. A. (2007). Similarities and Differences Between Human-Human and Human-Automation Trust: An Integrative Review. Theoretical Issues in Ergonomics Science, 8(4), 277-301.
  8. Burton, J. W., Stein, M.-K., & Jensen, T. B. (2020). A Systematic Review of Algorithm Aversion in Augmented Decision Making. Journal of Behavioral Decision Making, 33(2), 220-239.

This article discusses peer-reviewed human-AI trust research and is provided for general informational purposes. It is not a product recommendation or endorsement of any specific AI tool. Confirm the accuracy claims and appropriate use of any AI financial tool directly with its provider and your accountant.