Most of what a professional firm does about decision quality is aimed at bias. Almost none of it is aimed at the other half of the problem, which is measurable at far lower cost and is usually larger.
Key Takeaway
Given the same case files, the median percentage difference between quotes for any pair of underwriters is a stunningly large 55%, which is about five times as large as expected by the executives asked about this scenario in a survey[1]. The structural point: to measure bias you must know the right answer, and to measure noise you need only two people and one file. That second measurement is available to any firm immediately, and almost nobody takes it.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the decomposition of error into bias and noise. It is not a finding but a mathematical identity, and we work it below.
Grade B for the insurance audit result, which is consistently reported at the same figure across five independent sources but which we could not trace to a peer-reviewed publication.
Grade C for the specific typology of noise, which we report as a framework rather than a finding.
Our position: the most valuable thing here is not a research finding at all. It is a measurement procedure, and its value does not depend on the audit's number being right.
A Note On Method
Everything here is verified to August 2026.
We did not obtain the book. Every claim about its contents reaches us through reviews, a publisher listing and summary sites, and we flag the status of each.
Our best source is a review in a statistics periodical[1], which we prefer throughout, supported by a literary review[2] and by a publisher listing including the book's own chapter structure and index[3][4].
We located no peer-reviewed publication of the insurance noise audit and report it as the book's claim rather than as a published study.
All arithmetic is ours.
This article reviews a book about judgment. It is not audit, assurance, legal or professional standards advice, and nothing in it describes any professional body's requirements.
A Warning About Our Sources
A caution we owe the reader, in the manner of the twenty-sixth article. Ours.
This is the weakest sourcing position in this series so far. The primary object is a trade book, we did not obtain it, and most of what is written about it online is summary sites optimising for search traffic.
Three things we did about that.
We anchored on the two review venues we could identify as edited publications rather than summary sites[1][2].
We required the audit figure to appear in at least three independent sources before reporting it, and it appears in five.
And we built the article around the part that does not depend on the sourcing at all, which is the decomposition and the procedure. Those are arithmetic and method, and they survive whatever the book turns out to say.
What Noise Is
The definition.
Kahneman, Sibony and Sunstein published Noise: A Flaw in Human Judgment on 18 May 2021[5].
The concept, from a summary: noise describes the random scatter in professional judgments, under which two experts evaluating the same case may reach vastly different conclusions, and it is distinct from bias[6].
The statistics review puts the underlying observation well: "It can be natural to think of judgment in terms of mathematical functions in which the same inputs map to the same output. It turns out this isn't even remotely true in many human decision-making systems."[1]
Two observations, ours.
Bias is a systematic error: everyone leaning the same way. Noise is scatter: people leaning different ways.
And the distinction is not academic, because the two are detected by completely different procedures, which is the point of this article.
Two Felons, One Average
The illustration that makes the distinction land.
From a review: "If two felons receive sentences of three years and seven years when they should both be sentenced to five, the difference is due to noise. The average of three and seven is indeed five, but justice has quite obviously not been served!"[2]
Three observations, ours.
The average is correct and both decisions are wrong. There is no bias in this example at all, and yet two people have been badly treated.
Which means an organisation checking its aggregate numbers against a benchmark will find nothing. Average fee, average adjustment, average assessment: all fine, all concealing the scatter underneath.
And it explains why this problem persists. The measurement most firms already run is exactly the one that cannot detect it.
The Insurance Audit
The headline result, from our best source.
The statistics review states: "Given the same data (realistic but made-up information about cases), the median percentage difference between quotes for any pair of underwriters is a stunningly large 55% (so for half of the cases, it is worse than 55%), a difference about five times as large as expected by the executives asked about this scenario in a survey."[1]
Supporting detail from other sources, each flagged. One reports the executives' estimate as 10 percent or less, from a survey of 828 CEOs and senior executives[7]. Another gives the metric's definition as the median difference being 55 percent of the average of the two estimates[8]. A third adds a separate figure of 43 percent for claims adjusters[6].
And the consequence, quoted from the book by one source: "the price a customer is asked to pay depends to an uncomfortable extent on the lottery that picks the employee who will deal with that transaction."[7]
One observation, ours. We could not locate a peer-reviewed publication of this audit, and report it as the book's claim. The figure is consistent across five independent sources, which establishes that they are all reporting the same book faithfully and not that the number is right.
What Fifty-Five Percent Looks Like
Making the metric concrete. Our own arithmetic, reconstructing the measure from its stated definition; no source gives these figures.
If the difference is measured as a percentage of the average of two estimates, then a pair averaging $145,000 with a 55 percent difference is $105,125 against $184,875.
A pair averaging $50,000 is $36,250 against $63,750.
By contrast the 10 percent the executives expected, on a $145,000 average, is $137,750 against $152,250.
Two observations.
The gap between what was expected and what was found is not a matter of degree. One is a rounding difference between two professionals; the other is a different answer to the same question.
And 55 percent is the median. Half of the pairs were worse than that, a point our best source is careful to make and most summaries omit.
The Illusion Of Agreement
Why nobody had noticed.
A summary reports that many professionals maintained an illusion of agreement while in fact disagreeing in their professional judgments, and suggests one reason is that many organisations prefer consensus and harmony and have systems in place to minimise disagreements[8].
Three observations, ours.
Professionals rarely see each other's independent judgment on the same file. Work is allocated, not duplicated, precisely because duplication is expensive.
Which means the disagreement is structurally invisible. It is not concealed; there is simply no occasion on which it would appear.
And the eighth article in this series is directly relevant here. A firm that suppresses disagreement to preserve harmony has removed the only signal that would reveal the scatter, and has done so for reasons that looked like good management.
The Decomposition
The mathematical core, and the part of this article that does not depend on any source. Standard statistics; the worked table is ours.
Mean squared error in a set of judgments splits exactly into bias squared plus noise squared.
With bias 10 and noise 0, total error is 100, and none of it comes from noise.
With bias 8 and noise 6, total error is 100, of which 36 percent is noise.
With bias 7 and noise 7, total error is 98, split evenly.
With bias 6 and noise 8, total error is 100, of which 64 percent is noise.
With bias 0 and noise 10, total error is 100, and all of it is noise.
Two observations, ours.
At equal magnitudes the two contribute equally. Noise is not a secondary consideration; it enters the error in exactly the same way.
So a firm that has worked only on bias has addressed at most half the problem, and has no measurement of the other half.
The Asymmetry That Matters
The practical heart of this article. Ours, and it follows from the definitions rather than from any study.
To measure bias, you must know the right answer. Bias is systematic deviation from truth, so detecting it requires a benchmark: an outcome, a later correction, an authoritative determination.
To measure noise, you need two people and one file. Scatter is deviation between judgments, and judgments can be compared to each other without any external reference at all.
Three consequences.
A noise audit requires no outcome data. You do not have to wait for a file to resolve, an assessment to be reviewed, or a forecast to come true.
It requires no benchmark, which means it works in exactly the domains where truth is slow, contested or never established, which is most of professional judgment.
And it can be run this week, at the cost of duplicated work on a handful of files.
That asymmetry is, in our view, the most useful thing in this entire series, and it does not depend on the audit's number, on the book, or on any finding being replicable.
Three Kinds Of Noise
The typology, reported as a framework rather than a finding.
The book's index records the terms level noise, pattern noise and occasion noise[4].
We did not obtain the book's definitions, and what follows is our own characterisation from how the terms are used, flagged accordingly.
Level noise appears to describe differences in average severity between judges: one partner consistently more conservative than another.
Pattern noise appears to describe differences in how particular cases are handled: two people with the same average who disagree about which files are the difficult ones.
Occasion noise appears to describe variation within the same person: the same file assessed differently on a different day.
Two observations, ours.
The third is the uncomfortable one, and it is testable independently of anyone else: give yourself the same file twice, separated by enough time.
And the three call for different remedies. Level differences can be addressed by calibration; pattern differences require discussing specific cases; occasion variation is about the conditions of work rather than about the person.
Where Else It Has Been Found
The range, from summaries.
Reported examples include wine experts at a major US wine competition scoring only 18 percent of wines identically when tasting the same wines twice; physicians giving significantly different diagnoses when presented twice with the same case; and the observation that matching fingerprints is not nearly as clear-cut as many people think, because latent prints left at a crime scene are often very different from exemplar prints collected in a controlled environment[8][9]. Another source adds that case officers in child protective services vary in how likely they are to place children in foster care[7].
We obtained none of the underlying studies and report these as the summaries describe them.
Two observations, ours.
The wine and physician examples are occasion noise: the same person, the same input, a different answer. That is a stronger and stranger finding than disagreement between people.
And note the domains. Every one is a field with training, credentials and professional standards, which is the point. Noise is not a symptom of amateurism.
Not All Variability Is Bad
A qualification the book apparently makes and the summaries mostly bury.
One summary records: not all variability is bad; diversity of ideas in markets or creativity contexts is valuable, but in consistent-decision systems it is harmful[6].
Three observations, ours.
That distinction does a great deal of work. Disagreement in generating options is a resource, which is what the tenth article in this series found about independent idea generation.
Disagreement in applying a standard is a defect, because the standard is supposed to determine the answer.
And a firm needs to know which of the two it is running at any moment. The same partner meeting can be a valuable disagreement about strategy and an unacceptable disagreement about a technical position, in consecutive agenda items.
The Book's Own Section On Costs
Something we found in the chapter listing that almost no summary mentions.
The book's structure includes a Part VI titled Optimal Noise, containing chapters titled The Costs of Noise Reduction, Dignity, and Rules or Standards?[3]
We did not obtain these chapters and report only their titles.
Three observations, ours.
A book arguing for noise reduction that devotes a section to optimal noise is not arguing for its elimination, and the popular version of this book tends to lose that.
Dignity as a chapter title suggests the authors take seriously the objection that reducing a professional to a checklist has a cost beyond the arithmetic.
And Rules or Standards? is the live question in accounting specifically, where the tension between bright-line rules and principles-based judgment is a permanent feature of the field rather than an oversight. Noise reduction pushes toward rules, and rules have well-known failure modes of their own.
Recurrent And Singular Decisions
A distinction that determines whether any of this is usable.
A summary records the book distinguishing recurrent decisions, being repeat cases with measurable variability, from singular decisions, being unique and unrepeatable, and arguing that even singular decisions can be contaminated by analogous noise sources, being differing perspectives, moods, or contextual factors between decision makers[6].
Two observations, ours.
The recurrent case is where a noise audit works, because you need multiple comparable files.
And most professional work is more recurrent than it feels. A firm may experience each client as unique while making the same twenty categories of judgment across all of them, which is precisely the condition an audit needs.
Running One In Your Own Firm
The procedure. Ours, constructed from the definitions above rather than from the book's appendix, which we did not obtain.
Five steps.
Pick a judgment your firm makes repeatedly where the answer is a number or a category: a fee quote, a materiality threshold, a risk rating, a provision, a classification.
Assemble five to ten real past files, stripped of the conclusion that was reached and of who reached it.
Have several people answer independently, without discussion, and without knowing that anyone else is doing it or what the exercise is testing.
Compare the answers to each other, not to a correct answer, which you do not need and probably do not have.
And ask everyone beforehand how much they expect the answers to differ, because that prediction is the part that makes the result land.
What You Will Probably Find
Our expectation, stated as a prediction rather than a finding, so that it can be wrong.
Three things we would expect, none of them established by anything cited here.
The spread will be wider than anyone predicted, because that is what the audit found in insurance and because nobody in your firm has ever seen this measured either.
The disagreement will concentrate in specific categories rather than being uniform, which tells you where a standard is missing rather than where people are careless.
And the first reaction will be to defend the spread as reflecting legitimate professional judgment, which is sometimes correct and is the conversation the exercise exists to start.
The Hard Part Is Not Measurement
The obstacle, stated plainly. Ours.
The measurement is cheap. Duplicating a handful of files costs a day.
Three reasons it does not happen anyway.
It produces a number that is uncomfortable and attributable, unlike most quality metrics, and someone has to be willing to see it.
It implies that some past answers were wrong, which has professional and occasionally legal implications that a firm may prefer not to create in writing. We note that without advising on it, because it is a real consideration and pretending otherwise would be dishonest.
And it invites a response that costs more than the measurement: writing down the standard that would have produced agreement, which is the actual work and which nobody has time for.
Our own view is that the second reason is why this is rarer than it should be, and that a firm which cannot look at its own consistency has not avoided the problem, only the number.
What To Do
Separate bias from noise. One is everyone leaning the same way, the other is people leaning different ways, and they are detected by completely different procedures.
Note that your averages cannot see it. Three years and seven years average to five, and two people have still been badly treated.
Use the asymmetry. Measuring bias needs the right answer; measuring noise needs two people and one file. The second is available to you now.
Ask for the prediction first. The reported result is remarkable mainly because executives expected ten percent, and the gap between expectation and finding is what changes minds.
Test yourself against yourself. Occasion noise is variation within one person, and you can check it by assessing the same file twice with enough time between.
Distinguish generative disagreement from applied disagreement. Variation in producing options is a resource; variation in applying a standard is a defect.
Expect the remedy to cost more than the diagnosis. Writing down the standard that would have produced agreement is the real work.
Read the qualification. The book contains a section titled Optimal Noise, with chapters on the costs of reduction and on dignity, and it is not arguing for elimination.
The Limits Of This Analysis
Several caveats matter. This article reviews a book about judgment and is not audit, assurance, legal or professional standards advice; nothing in it describes any professional body's requirements. Everything is verified to August 2026. This is the weakest sourcing position in this series. We did not obtain the book, and every claim about its contents reaches us through reviews, a publisher listing and summary sites. We located no peer-reviewed publication of the insurance noise audit and report it as the book's claim; its consistency across five sources establishes that they report the same book faithfully, not that the figure is correct. We did not obtain any of the underlying studies on wine judging, medical diagnosis, fingerprints or child protective services. Our characterisations of level, pattern and occasion noise are our own, inferred from usage, not the book's definitions, which we did not obtain. The chapters in Part VI are reported by title only. All arithmetic is ours; the decomposition is standard statistics and the dollar pairs reconstruct a metric from its stated definition, with no source giving those figures. The audit procedure is our own construction from the definitions rather than the book's appendix, which we did not obtain. The section on what you will probably find is an explicit prediction, not a finding, and is offered so that it can be wrong.
Frequently Asked Questions
What is the difference between bias and noise?
How large was the reported effect?
Why can I measure noise more easily than bias?
How would I run this in my firm?
Is all variation bad?
Why doesn't every firm do this?
References
- Review of Noise: A Flaw in Human Judgment in a statistics periodical, DOI 10.1080/09332480.2024.2416879, on it being natural to think of judgment in terms of mathematical functions in which the same inputs map to the same output, and on this turning out not to be even remotely true in many human decision-making systems; and on insurance underwriting, where given the same data, being realistic but made-up information about cases, the median percentage difference between quotes for any pair of underwriters is a stunningly large 55 percent, so that for half of the cases it is worse than 55 percent, a difference about five times as large as expected by the executives asked about this scenario in a survey; and on the observation that a customer's optimal strategy is therefore to get multiple quotes. Note: a review in an edited statistics periodical; our best source and the one we prefer throughout. We obtained portions. tandfonline.com
- Review of Noise: A Flaw in Human Judgment in a literary review, on the definition that if two felons receive sentences of three years and seven years when they should both be sentenced to five, the difference is due to noise, since the average of three and seven is indeed five but justice has quite obviously not been served; and confirming the publication details as Little, Brown Spark, 2021, 464 pages. Note: an edited literary review, not peer-reviewed; used for the sentencing illustration and for publication details. lareviewofbooks.org
- Book summary site reproducing the full chapter listing of Noise: A Flaw in Human Judgment, including Part V on Improving Judgments with chapters on Better Judges for Better Judgments, Debiasing and Decision Hygiene, Sequencing Information in Forensic Science, Selection and Aggregation in Forecasting, Guidelines in Medicine, Defining the Scale in Performance Ratings, Structure in Hiring, and The Mediating Assessments Protocol; and Part VI on Optimal Noise with chapters titled The Costs of Noise Reduction, Dignity, and Rules or Standards?; and noting the book contains appendices on noise audits, decision observer checklists, and steps for correcting predictions. Note: a book summary site, not peer-reviewed. We report chapter titles only and did not obtain the chapters or the appendices. readingraphics.com
- Publisher listing for Noise: A Flaw in Human Judgment by Daniel Kahneman, Olivier Sibony and Cass R. Sunstein, reproducing the book's index terms including level noise, pattern noise, occasion noise, noise audit, objective ignorance, singular decisions, decision makers, predictive judgments, professional judgments and noise-reduction strategies; and recording author biographical details. Note: a publisher listing. We report the existence of these index terms; the book's definitions of them were not obtained, and the characterisations in this article are the authors' own. books.google.com
- Encyclopedia entry recording that Noise: A Flaw in Human Judgment is a nonfiction book by Daniel Kahneman, Olivier Sibony and Cass Sunstein, first published on 18 May 2021 by Little, Brown Spark, Hachette Book Group, ISBN 978-0-00-830899-5, concerning noise in human judgment and decision-making. Note: an encyclopedia entry, used for publication details only. en.wikipedia.org
- Book summary site, on noise describing the random scatter in professional judgments under which two experts evaluating the same case may reach vastly different conclusions, distinct from bias; on a noise audit finding a 55 percent median difference between two underwriters quoting the same risk and 43 percent for claims adjusters, against an executive estimate of only 10 percent variation, illustrating the illusion of agreement; on noise adding costs in both lost sales through overpricing and losses through underpricing; on not all variability being bad, with diversity of ideas in markets or creativity contexts being valuable while harmful in consistent-decision systems; and on the distinction between recurrent decisions, being repeat cases with measurable variability, and singular decisions, being unique and unrepeatable, with even singular decisions capable of contamination by differing perspectives, moods or contextual factors. Note: a book summary site, not peer-reviewed. bookassess.com
- Commentary newsletter on the book, on the authors having asked 828 CEOs and senior executives to guess how much variation they expected to find in judgments of insurance premiums, with the executives guessing 10 percent or less while the median difference in underwriters' judgments was 55 percent; quoting the book that for insurance claims the price a customer is asked to pay depends to an uncomfortable extent on the lottery that picks the employee who will deal with that transaction; and noting that case officers in child protective services vary in how likely they are to place children in foster care. Note: a commentary newsletter, not peer-reviewed. The figure of 828 executives comes from this source alone and we could not verify it. robkhenderson.com
- Book summary site, on the noise audit of insurance underwriters finding a median difference of about 55 percent where the company's executives had expected about 10 percent; on many professionals maintaining an illusion of agreement while in fact disagreeing in their professional judgments; on one reason being that many organisations prefer consensus and harmony and have systems in place to minimise disagreements; and on matching fingerprints not being nearly as clear-cut as many people think, because latent prints left at a crime scene are often very different from exemplar prints collected in a controlled environment. Note: a book summary site, not peer-reviewed. tosummarise.com
- Book review site, on the reported example of wine experts at a major US wine competition scoring only 18 percent of wines identically when tasting the same wines twice, usually the worst ones; on it being common to obtain significantly different diagnoses from the same physicians when presented twice with the same case; and on the noise audit having asked two randomly selected qualified underwriters to estimate a potential loss in a specific case, with the median difference being 55 percent of the average of the two estimates. Note: a book review website, not peer-reviewed. We did not obtain any of the underlying studies described. hiddenvaluegems.com
This article reviews a book about judgment and is not audit, assurance, legal or professional standards advice. The book was not obtained; every claim about its contents reaches this article through reviews, a publisher listing and summary sites. No peer-reviewed publication of the insurance noise audit was located, and it is reported as the book's claim. The characterisations of level, pattern and occasion noise are the authors' own inferences from usage, not the book's definitions. The audit procedure is the authors' own construction. All arithmetic is the authors' own.