The thirty-second article in this series described the scatter in professional judgment and said measuring it was cheap. This one is about the cheapest available remedy, which requires nobody to become better at anything and which most people decline because they believe it does not work.
Key Takeaway
"If the estimates of two judges ever fall on different sides of the truth, which we term bracketing, averaging must outperform the average judge for convex loss functions, such as mean absolute deviation." And: people "often hold incorrect beliefs about averaging, falsely concluding that the average of two judges' estimates would be no more accurate than the average judge"[1]. We worked the arithmetic ourselves. Two judges estimating a true value of 100 at 60 and 150 have an average error of 45; the error of their average is 5.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the averaging guarantee. It is not a finding but a mathematical property, stated as must in a peer-reviewed abstract, and we verified it by working cases.
Grade A that people misappreciate it, from four experiments in a leading management journal, replicated across three different task formats within the paper.
Grade B for the practical gains, which depend on independence assumptions we discuss and which our own figures illustrate rather than establish.
Our position: this is the most immediately usable thing in the series so far, because the mechanism is arithmetic, the cost is a second opinion, and the barrier is a belief that can be corrected in one demonstration.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2006 abstract verbatim from four independent sources, including two of the authors' own institutional databases and the publisher's record, and they agree word for word[1][2][3].
We obtained the bracketing definition and related findings from a peer-reviewed forecasting review[4].
We did not obtain the paper itself, and report no effect sizes, sample sizes or experimental detail, because we have none.
All arithmetic and simulation is ours.
We dropped one well-known related study from this article because we could not locate a source for it, rather than citing it from memory.
This article discusses judgment aggregation. It is not investment, valuation, audit or professional standards advice.
The Claim
The paper.
Larrick and Soll published Intuitions About Combining Opinions: Misappreciation of the Averaging Principle in Management Science, 52(1), 111–127, in January 2006, DOI 10.1287/mnsc.1060.0518, both authors at the Fuqua School of Business, Duke University[3].
The opening sentence sets the scope: "Averaging estimates is an effective way to improve accuracy when combining expert judgments, integrating group members' judgments, or using advice to modify personal judgments."[1]
Then the guarantee: "If the estimates of two judges ever fall on different sides of the truth, which we term bracketing, averaging must outperform the average judge for convex loss functions, such as mean absolute deviation (MAD)."[1]
Two observations, ours.
The word is must. This is not an empirical regularity that might fail to replicate. It is a property of the arithmetic, in the same category as the twenty-sixth and thirty-fourth articles in this series, and it cannot be overturned by future data.
And note the comparison being made. Averaging is compared to the average judge, which is what you get if you pick one of them at random. It is not a claim that averaging beats the better judge, and the distinction matters.
We Worked It Through
Four cases, because a guarantee should be checked. The figures are ours, invented, and appear in no source.
Take a true value of 100 in every case, and compare two quantities: the error of the average, against the average of the two errors.
Estimates of 80 and 130, which bracket. Their average is 105, an error of 5. The average of the two individual errors is 25. Averaging is better by 20.
Estimates of 80 and 90, both below. Average 85, error 15. Average individual error 15. Exactly equal.
Estimates of 120 and 140, both above. Average 130, error 30. Average individual error 30. Equal again.
Estimates of 60 and 150, wildly wrong but bracketing. Average 105, error 5. Average individual error 45. Averaging is better by 40.
One observation, ours: the fourth case is the striking one. Two badly wrong estimates produced an almost exact answer, because the errors pointed in opposite directions. Neither judge deserved any credit and the average was nearly right.
The Asymmetry
What those four cases establish. Ours.
A peer-reviewed forecasting review states the rule cleanly: "When their estimates bracket, the forecast generated by taking their average performs better than choosing one of the two experts at random; when the estimates do not bracket, averaging performs equally as well as the average expert."[4]
Two consequences.
Averaging wins when they bracket and ties when they do not. There is no third case in which it loses to picking at random.
Which makes it a one-way bet. You are choosing between a strategy that sometimes wins and never loses, and one that sometimes loses. That is not a close decision, and yet it is not the one most people make.
What It Buys
The magnitude. Our own simulation, with judges who are unbiased and whose errors are independent and equally sized; these assumptions are strong and the next-but-one section examines them.
With one judge, mean absolute error 7.99 in our arbitrary units.
With two, averaging gives 5.64 against 7.98 for the average judge. A reduction of 29 percent.
With three, 4.60. A reduction of 42 percent.
With five, 3.57. 55 percent.
With ten, 2.52. 68 percent.
Two observations.
The first additional judge is worth the most, and returns diminish steadily. Going from one to two buys more than going from five to ten.
And nobody in this simulation became better at anything. The individual judges are exactly as accurate throughout. The entire gain comes from combination.
The Misconception
The paper's actual finding, which is about belief rather than arithmetic.
"We hypothesized that people often hold incorrect beliefs about averaging, falsely concluding that the average of two judges' estimates would be no more accurate than the average judge. The experiments confirmed that this misconception was common across a range of tasks that involved reasoning from summary data (Experiment 1), from specific instances (Experiment 2), and conceptually (Experiment 3)."[1]
Three observations, ours.
Three different task formats, three confirmations. That is an internal replication across presentation methods, which is stronger than one experiment.
The belief is demonstrably false, not merely pessimistic. This is unusual: most of what this series has covered involves people being wrong about an empirical question. Here they are wrong about arithmetic.
And a fourth experiment identified the source: "flawed inferential rules and poor extensional reasoning abilities contributed to the misconception"[1]. We did not obtain what those flawed rules were, and do not speculate.
What Reduced It
The remedy, from the abstract, and it is unusually specific.
The misconception "decreased as observed or assumed bracketing rate increased (all three studies) and when bracketing was made more transparent (Experiment 2)"[1].
Two observations, ours.
People are not simply stubborn. Their belief responded correctly to the relevant variable. When they could see that estimates were straddling the truth, they appreciated averaging more.
And making bracketing transparent is an intervention any firm can run. It means showing people, after the fact, where their estimates fell relative to what turned out to be true. That is a record-keeping change rather than a training programme.
Why Nobody Learns This
The paper's closing observation, which explains the persistence.
The authors "conclude by describing how people may face few opportunities to learn the benefits of averaging and how misappreciating averaging contributes to poor intuitive strategies for combining estimates"[1].
Three observations, ours.
Few opportunities to learn is the crucial phrase. To learn that averaging works, you would have to average, observe the outcome, and also observe what would have happened had you picked one estimate instead. The second of those is a counterfactual you never see.
So the feedback that would teach the lesson does not exist in ordinary working life. This is the same structure the seventh article in this series found in feedback interventions and the twenty-fifth found in deadline calibration.
And it explains why the misconception survives among experienced professionals. Experience cannot correct a belief when the correcting observation is never made.
An Erratum One Month Later
A detail we found while sourcing, recorded because anyone citing the paper should know.
An erratum to the paper was published in Management Science, 52(2), 309–310, in February 2006, one month after the original[5].
We could not obtain its content and do not know what was corrected.
Two observations, ours.
We record it because a reader relying on any specific figure from the paper should check the erratum first, and because we cannot tell them whether it affects anything here. Nothing in this article depends on a numerical result from the paper, only on its abstract's statements.
And we note in passing that one institutional record renders the second author's name as "Soil" rather than "Soll" in its erratum entry[5]. This is the eleventh instance in this series of a bibliographic detail being reproduced inconsistently.
Diversity Beats Expertise
The finding that determines who to ask.
A peer-reviewed forecasting review reports that two factors influence the quality of an average forecast, individual expertise and the crowd's diversity, and quotes: "The benefits of diversity are so strong that one can combine the judgments from individuals who differ a great deal in their individual accuracy and still gain from averaging."[4]
We did not obtain the underlying paper and report the quotation as the review gives it.
Three observations, ours.
This inverts the natural instinct, which is to ask the most expert person available and then perhaps a second expert like them.
If the quotation is accurate, a less accurate but differently-minded second judge can still improve the combination, because what averaging exploits is errors pointing in opposite directions, and similar people make similar errors.
And that connects directly to bracketing. Two judges who think alike will rarely bracket, so averaging them will usually tie rather than win. The gain comes from disagreement, which is the resource the eighth article in this series argued firms suppress.
Simple Averaging Beats Clever Weighting
A result worth knowing before anyone builds a model.
The same review states: "The average point forecast also often outperforms more complicated point aggregation schemes, such as weighted combinations."[4]
We did not obtain the cited studies and note the hedge often, which is the review's.
Two observations, ours.
Weighting requires estimating how good each judge is, and those estimates are themselves noisy. A weighting scheme built on noisy weights can perform worse than equal weights, which need no estimation at all.
And commercially this is good news, because the simple version is the free one. A firm does not need to score its people before it can benefit from combining them.
The Connection To Noise
How this fits with the thirty-second article. Ours.
That article reported scatter in professional judgment and argued it could be measured cheaply because comparing judgments requires no correct answer.
Three consequences.
Averaging is the direct remedy for that scatter, and it works for the same reason the measurement did: it operates on the relationship between judgments rather than on their relationship to truth.
So a firm can reduce noise without ever establishing what the right answer was, which is the same asymmetry stated the other way round.
And the two procedures share a step. Collecting independent judgments on the same file is the measurement in one article and the remedy in this one. The same day's work produces both.
Where It Does Not Work
Being clear about scope. Ours.
Four limits.
It applies to quantities, not to choices. Averaging two views on whether to take a client produces nothing. A citing source identifies work on intuitive biases in choice versus estimation and their implications for the wisdom of crowds[7], which we did not obtain but which indicates the distinction is studied.
It does not beat the better judge, only the average one. If you reliably know which of two people is more accurate, listening to them is a different and possibly better strategy. The catch is the word reliably.
It requires genuinely independent judgments, which the next section takes seriously.
And a source identifies work finding that crowd wisdom relies on agents' ability in small groups with a voting aggregation rule[6], which indicates that with voting rather than averaging, and in small groups, ability matters more. We did not obtain it and flag the boundary rather than describing it.
The Condition Everything Rests On
The assumption that does the work, and the one most easily destroyed in practice. Ours.
Our simulation assumed independent errors. That assumption is what produced the 29 percent reduction from a second judge.
Three ways firms destroy it, all ours and untested.
The second person sees the first person's number before giving theirs, which the fifth article in this series would predict anchors them, collapsing the two judgments toward one.
Both people are trained the same way, use the same template, and read the same file notes, so their errors correlate by construction and they rarely bracket.
And the judgments are made in a meeting, where the first opinion voiced shapes the rest, which is the production-blocking and conformity problem of the tenth article.
One consequence: the procedural requirement is that the second opinion be formed before the first is heard. That is the whole design, and it is usually the step that gets dropped for convenience.
Running This In A Firm
The practical version. Ours, constructed from the principles above.
Five steps.
Identify the judgments that are quantities. A fee estimate, a provision, a time budget, a valuation range, a probability of recovery.
Have two people produce a number independently, without either seeing the other's, and without discussion first.
Average them rather than debating to a single figure, at least as the starting point.
Record where each fell relative to what actually happened, when it eventually becomes known.
And show people that record, because the paper reports the misconception decreasing when bracketing was made transparent, and that record is the transparency.
One observation. The fourth and fifth steps are the ones that make it stick, and they are the ones firms skip, because they take months to pay off and produce a record of past errors.
The Obstacle Is Not Arithmetic
Why this is harder than it looks. Ours.
Three reasons.
Averaging looks like indecision. Splitting the difference between two colleagues reads as a refusal to judge, and in a professional culture that prizes a clear view, it is a status cost.
It doubles the cost of the judgment, at least for the files where it is used, and the benefit is invisible in any single instance because the counterfactual is unobserved.
And the person who was closer will notice, and will reasonably ask why their answer was diluted by a worse one. The correct reply is that you did not know in advance who would be closer, and that is a genuinely unsatisfying thing to say to a good performer.
Our own view: this is why a demonstration matters more than an explanation. The four worked cases at the top of this article take two minutes and are more persuasive than any argument, particularly the one where two estimates 45 apart from the truth average to within 5 of it.
What To Do
Average rather than choose, when the answer is a quantity. Averaging wins when the estimates bracket the truth and ties when they do not, so against picking at random it cannot lose.
Get the second opinion before the first is heard. Independence is the condition the whole result rests on, and a second person who has seen the first number is not a second judgment.
Value a different mind over a better one. On the reported quotation, diversity benefits are strong enough that combining people who differ a great deal in accuracy still gains, because averaging exploits errors pointing opposite ways.
Do not bother weighting. Simple averaging often outperforms weighted schemes, and weights have to be estimated from data that is itself noisy.
Expect the belief to be the barrier, not the maths. Four experiments found people falsely concluding that averaging gains nothing.
Make bracketing visible. The misconception decreased when bracketing was made transparent, so record where each estimate fell relative to the eventual outcome and show people.
Note that experience will not teach this. Learning it requires observing a counterfactual you never see, which is why experienced professionals hold the belief too.
Keep it to quantities. The principle is about numbers, not about whether to take a client, and the choice case has its own literature.
The Limits Of This Analysis
Several caveats matter. This article discusses judgment aggregation and is not investment, valuation, audit or professional standards advice. Everything is verified to August 2026. We did not obtain the 2006 paper, only its abstract, verbatim from four independent agreeing sources, and we report no effect sizes, sample sizes or experimental detail because we have none; we do not know what the flawed inferential rules identified in its fourth experiment were. An erratum was published one month after the original and we could not obtain its content; nothing in this article depends on a numerical result from the paper. We did not obtain the forecasting review's underlying sources, including the paper quoted on diversity, the studies on weighted combinations, the work on choice versus estimation, or the paper on voting in small groups, and report all of them as citations or as the review characterises them. All arithmetic and simulation is ours; the worked cases use invented figures and the error-reduction percentages assume unbiased judges with independent, equally sized errors, which is a strong assumption we examine in the body. We dropped one well-known related study from this article because we could not locate a source for it rather than citing it from memory. The sections on how firms destroy independence, the firm procedure, and the obstacles are our own reasoning, untested.
Frequently Asked Questions
Why can averaging not lose?
Can you show that?
How much does it improve accuracy?
So why doesn't everyone do it?
Why doesn't experience teach it?
What is the one thing that breaks it?
References
- Larrick, R. P., & Soll, J. B. (2006). Intuitions about combining opinions: Misappreciation of the averaging principle. Management Science, 52(1), 111–127. DOI 10.1287/mnsc.1060.0518, published abstract via the lead author's institutional faculty database, on averaging estimates being an effective way to improve accuracy when combining expert judgments, integrating group members' judgments, or using advice to modify personal judgments; on averaging necessarily outperforming the average judge for convex loss functions such as mean absolute deviation if the estimates of two judges ever fall on different sides of the truth, which the authors term bracketing; on the hypothesis that people often hold incorrect beliefs about averaging, falsely concluding that the average of two judges' estimates would be no more accurate than the average judge; on the experiments confirming this misconception was common across tasks involving reasoning from summary data, from specific instances, and conceptually; on the misconception decreasing as observed or assumed bracketing rate increased and when bracketing was made more transparent; on a fourth experiment showing that flawed inferential rules and poor extensional reasoning abilities contributed to the misconception; and on the authors concluding by describing how people may face few opportunities to learn the benefits of averaging. Note: we obtained the published abstract in full but not the paper, and report no effect sizes, sample sizes or experimental detail. fds.duke.edu
- Institutional scholarly record for Larrick and Soll (2006), Management Science, 52(1), 111–127, reproducing the abstract identically and confirming the DOI, ISSN and publication date of 1 January 2006. Note: an institutional record from the authors' university; independent corroboration of the abstract text. scholars.duke.edu
- Bibliographic record reproducing the abstract identically and recording both authors' affiliation at the Fuqua School of Business, Duke University, Durham, North Carolina, with the citation as Management Science, INFORMS, vol. 52(1), pages 111–127, January 2006. Note: a bibliographic database record; a third independent confirmation of the abstract text and the source for the authors' affiliation. ideas.repec.org
- Peer-reviewed forecasting review, on Larrick and Soll (2006) defining the idea of bracketing, under which two experts can either bracket the realisation or not, with the average performing better than choosing one of the two experts at random when their estimates bracket and equally as well as the average expert when they do not; on the average point forecast also often outperforming more complicated point aggregation schemes such as weighted combinations, citing work from 2009 and 2020; on two crucial factors influencing the quality of the average point forecast being individual expertise and the crowd's diversity; and quoting that the benefits of diversity are so strong that one can combine the judgments from individuals who differ a great deal in their individual accuracy and still gain from averaging. Note: a peer-reviewed review article; we obtained the relevant passages and none of the underlying studies it cites. arxiv.org
- Institutional scholarly record for the erratum, being Larrick, R. P., and Soll, J. B. (2006), Erratum: Intuitions about combining opinions: Misappreciation of the averaging principle, Management Science, 52(2), 309–310, published 1 February 2006. Note: we could not obtain the erratum's content and do not know what was corrected; it is recorded here so that anyone citing a specific figure from the paper checks it. This record renders the second author's surname as "Soil" rather than "Soll". scholars.duke.edu
- Bibliographic record for a paper on social influence and crowd wisdom under voting, whose reference list identifies Keuschnigg, M., & Ganser, C. (2017), Crowd Wisdom Relies on Agents' Ability in Small Groups with a Voting Aggregation Rule, Management Science, 63(3), 818–828; Larrick and Soll (2006) and its erratum; Aspinall, W. (2010), A route to more tractable expert advice, Nature, 463(7279), 294–295; an experimental application of the Delphi method to the use of experts, Management Science, 9(3), 458–467; and Bonaccio, S., & Dalal, R. S. (2006), Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences, Organizational Behavior and Human Decision Processes, 101(2), 127–151. Note: citations only. We obtained none of these papers and report the voting-rule boundary as a flag rather than a described finding. ideas.repec.org
- Reference list in an academic reference work chapter on collective judgment, identifying Galton, F. (1907), Vox populi, Nature, 75, 450–451; Herzog, S. M., & Hertwig, R. (2009), The wisdom of many in one mind: improving individual judgments with dialectical bootstrapping, Psychological Science, 20(2), 231–237; Larrick, R. P., & Soll, J. B. (2006); Simmons, J. P., Nelson, L. D., Galak, J., & Frederick, S. (2011), Intuitive biases in choice versus estimation: Implications for the wisdom of crowds, Journal of Consumer Research, 38(1), 1–15; and Kerr, N. L., & Tindale, R. S. (2011), Group-based forecasting? A social psychological analysis, International Journal of Forecasting, 27(1), 14–40. Note: citations only. We obtained none of these papers; the choice-versus-estimation paper is cited in this article only to indicate that the distinction has been studied. link.springer.com
This article discusses judgment aggregation and is not investment, valuation, audit or professional standards advice. The 2006 paper was not obtained in full and no effect sizes, sample sizes or experimental detail are reported. An erratum published one month after the original could not be obtained. All arithmetic and simulation is the authors' own; the worked cases use invented figures and the error-reduction percentages assume unbiased judges with independent, equally sized errors.