The sixty-seventh article reported that simple models beat expert judgment. This one explains when they do not, and the answer is precise enough to sort your own decisions into two piles before lunch.
Key Takeaway
Kahneman and Klein conclude that "evaluating the likely quality of an intuitive judgment requires an assessment of the predictability of the environment in which the judgment is made and of the individual's opportunity to learn the regularities of that environment."[1] On our own arithmetic, a cue worth a correlation of 0.30 needs 85 observations before you could reliably detect it, which at two hires a year is 42.5 years.
The Verdict, Stated First
Five claims, in descending order of confidence.
One. The conditions for trustworthy intuition are settled and specific. Two researchers who had spent careers on opposite sides sat down together and agreed on them, which is about as strong as agreement gets in this literature.
Two. The conditions are about the work, not the person. A regular enough environment to contain learnable cues, and enough practice with clear feedback to learn them. Both are properties of the task.
Three. Confidence carries no information about which case you are in. The 2009 paper states that subjective confidence tracks the internal consistency of the evidence rather than its quality.
Four. Most business decisions do not supply enough observations. On our own arithmetic a cue of validity 0.30 requires 85 cases, and an owner making twenty such decisions a year reaches that in four years, while one making two reaches it in forty-two and a half.
Five. And the feedback you do get is selected. On our own arithmetic, hiring the top fifth on a rating worth only 0.30 shows you a 66.8 percent success rate while the people you rejected would have succeeded 45.8 percent of the time, a number you never observe.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the 2009 conclusions, from an abstract obtained verbatim from a government bibliographic database and corroborated by a second index, plus body text from a copy hosted by a university.
Grade A for the two-settings framework, from an abstract obtained verbatim from an economics database.
Grade B for the historical account of the two camps, which reaches us partly through non-academic sources, flagged at every use.
Grade A for our own arithmetic, which is standard statistics on parameters we invented and state.
Our position: this is the best-settled question in the series, because the people most likely to disagree about it went and did not.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2009 paper's abstract verbatim from a national medical library database[1], corroborated word for word by a bibliographic service[5], and selected body passages from a copy hosted by a university business school[2]. We did not obtain the paper in full.
We obtained the 2016 management paper's abstract verbatim from an economics database[3]. We did not obtain the paper.
We did not obtain the 2015 paper, the 2001 book, or the 1978 paper, and report all from citation records[4].
The historical account of the two research camps comes partly from a writer's newsletter[6], flagged at every use.
All arithmetic is ours, uses invented validity and selection figures, and assumes normal distributions throughout.
This article discusses research on judgment. It is not hiring, management or strategic advice, and nothing here bears on the lawfulness of any employment practice.
Two Camps, Both Right
The problem the 2009 paper set out to solve.
A writer's newsletter describes it: one camp, naturalistic decision making, "found that performers reliably improve with experience." The other, heuristics and biases, "found that people not only frequently do not improve with experience, sometimes they actually get worse." And: "Both camps had produced fascinating work, and were well-respected. So where was the disconnect?"[6]
This is a writer's newsletter, not an academic source, flagged here and at every use.
Four observations, ours.
The two findings are flatly contradictory as stated, and both came from serious programmes with good data.
This series has recorded that structure many times and usually the resolution is a moderator. Here the moderator turns out to be the domain itself, which is a larger and more useful answer than most.
The paper's own framing is more careful than the popular one. It starts, in its abstract's words, "from the obvious fact that professional intuition is sometimes marvelous and sometimes flawed"[1], which concedes both camps at once.
And it describes its aim as mapping "the boundary conditions that separate true intuitive skill from overconfident and biased impressions"[1], which is a question about where a line falls rather than which side is right.
The Adversarial Collaboration
The source, and the method is part of the finding.
Kahneman, D., and Klein, G. (2009), Conditions for intuitive expertise: A failure to disagree, American Psychologist, 64(6), 515–526, September, DOI 10.1037/a0016755[1].
Four observations, ours.
The title contains the result. "A failure to disagree" is what two people report when they set out to argue and could not.
Adversarial collaboration is a format this publication covered separately, and its value is exactly this. Agreement between people who wanted to disagree is worth more than agreement between people who did not.
The pairing is as adversarial as the field allows. One author spent a career documenting the failures of intuition and the other documenting its successes, and they published jointly.
And the paper appeared in the flagship journal of a national psychological association, which is where a field puts something it considers settled rather than novel.
What They Concluded
The finding, verbatim.
"They conclude that evaluating the likely quality of an intuitive judgment requires an assessment of the predictability of the environment in which the judgment is made and of the individual's opportunity to learn the regularities of that environment."[1]
Four observations, ours.
Notice what is absent from that sentence. No mention of the person's intelligence, training, seniority, track record or confidence.
The two things named are both properties of the situation: how predictable it is, and whether this person got to learn it.
That makes the question answerable without knowing anything about the judge. You can assess a decision's learnability before you assess anyone's skill at it, which is the practical move this article is built on.
And it makes the usual argument from seniority invalid rather than merely weak. Twenty years in an unpredictable domain is twenty years of not learning it, and the sentence above says so directly.
One qualification we would attach to that, because it is easy to overstate. Long tenure in an unpredictable domain does teach things, being how the industry works, who to call, what the paperwork requires and where the bodies are buried.
What it does not teach is the specific thing people cite it for. Predictive judgment about outcomes, which is the claim behind almost every sentence beginning with the number of years someone has been doing this.
Two Necessary Conditions
The paper's own statement of them, from a copy hosted by a university business school.
"In particular, we explore two necessary conditions for the development of skill: high-validity environments and an adequate opportunity to learn them."[2]
Four observations, ours.
The word "necessary" is doing the work. These are not factors that help; they are requirements, and failing either means the skill does not develop.
High-validity environment means the situation contains cues that genuinely predict outcomes. If the cues are not there, no amount of attention finds them.
Adequate opportunity to learn means enough repetitions with feedback fast and clear enough to connect action to result. That is the condition our own arithmetic below attacks.
And the two are independent, which produces four cases rather than two. A predictable domain you have not practised, and a practised domain that is not predictable, both fail, and the second is the one that produces confident experts.
Why The Camps Had Disagreed
A methodological detail in the paper that explains the whole dispute, and which we have not seen quoted.
"NDM researchers compare the performance of professionals with that of the most successful experts in their field, whereas HB researchers prefer to compare the judgments of professionals with the outcome of a model that makes the best possible use of available information."[2]
Four observations, ours.
The two camps were using different benchmarks, and neither was wrong to. One asked whether professionals approach the best humans; the other asked whether they approach the best possible use of the information.
Against the first benchmark, experience obviously helps. Against the second, it frequently does not, and the sixty-seventh article reported exactly that comparison going badly for the humans.
So a great deal of a decades-long disagreement was two questions being answered rather than one being contested. That is worth noticing because it is a common shape and rarely visible from outside a field.
And it tells a business reader which benchmark to use. Comparing your judgment to other people's tells you about the market for your skill; comparing it to a rule tells you whether the skill is doing anything.
The Two-Settings Framework
The framework the 2009 paper's conclusion matches, from its own author's later management paper.
Hogarth, R. M., and Soyer, E. (2016), Kind and Wicked Experience in Marketing Management, Journal of Marketing Behavior, 2(2–3), 81–99, December, DOI 10.1561/107.00000031[3].
Its abstract opens: "Our society venerates experience. It feels right to trust our own experience and that of others. But experience also has adverse effects. Much learning is tacit in nature and, because people are typically unaware and uncritical of the conditions in which this takes place, experience can lead to false beliefs and subsequent actions can reinforce biases."[3]
Three observations, ours.
The sentence "it feels right to trust our own experience" concedes the intuition before attacking it, which is the correct order.
The mechanism named is tacit learning under unexamined conditions, and the phrase "unaware and uncritical of the conditions" is the specific failure. Not that people learn wrongly, but that they do not check whether the setting permits learning.
And the last clause is the sharpest. "Subsequent actions can reinforce biases" means the error is self-confirming, which is the same structure the sixty-sixth article computed for a missing cell.
Kind Means Match
The technical definition, which is more precise than the popular version and better.
"We adopt a two-settings framework in which experience is conceptualized as being acquired in one setting (learning) and then applied in another (target). When information in the two-settings match, the learning environment is kind. Wicked environments are characterized by mismatches and we specify several different types."[3]
Four observations, ours.
The popular version of this idea is about feedback speed. The authors' own version is about a match between where you learned and where you are applying it, which is different and more useful.
On that definition, fast feedback in the wrong setting is not kind at all. Twenty years of quick clear lessons about a market that has since changed is a wicked environment by this test.
It also means the same person can be in a kind environment for one decision and a wicked one for another, in the same afternoon. The classification attaches to the decision rather than the career.
And we did not obtain the several types of mismatch the abstract says the paper specifies, which is the most useful missing piece in this article.
The Assumption Nobody States
The paper's diagnosis of why the mismatch goes unnoticed.
"We note that many inferential errors occur because people implicitly assume informational matches between the two settings."[3]
Four observations, ours.
The word "implicitly" is the whole problem. Nobody decides that this year resembles last year; they proceed as though it does without the thought occurring.
Which makes the remedy a question rather than a discipline. Does the setting I learned this in match the setting I am applying it to? is answerable in a sentence and almost never asked.
The abstract also states the framework has normative implications illustrated through "decision-making challenges faced by marketing managers"[3], so the authors made the commercial transfer themselves.
And we did not obtain those illustrations, so the applications later in this article are ours rather than theirs.
Why Confidence Is Not A Signal
The passage that removes the most common way people assess their own judgment.
"Subjective confidence is often determined by the internal consistency of the information on which a judgment is based, rather than by the quality of that information (Einhorn and Hogarth, 1978; Kahneman and Tversky, 1973). As a result, evidence that is both redundant and flimsy tends to produce judgments that are held with too much confidence. These judgments will be presented too assertively to others and are likely to be believed more than they deserve to be."[2]
Four observations, ours.
The mechanism is precise and worth memorising. Confidence tracks how well the evidence hangs together, not how good the evidence is.
The phrase "redundant and flimsy" names the dangerous combination. Several weak signals that all point the same way feel like strong evidence, and the redundancy is doing the work rather than the strength.
The social consequence is stated too and is unusual for a psychology paper. Overconfident judgments are "presented too assertively" and "believed more than they deserve to be", which makes this a claim about meetings rather than about heads.
And it disposes of the argument from feeling certain. If confidence tracks consistency rather than quality, then feeling sure is evidence about your inputs' agreement and about nothing else.
Two consequences follow that a business reader can act on immediately. Gathering more of the same kind of evidence raises confidence without raising accuracy, because it raises consistency, which is what a third reference from the same former employer does.
And the opposite move feels worse and is better. Evidence from an independent source will often conflict with what you have, lowering your confidence while raising the quality of the judgment, which is why people stop seeking it.
How Much Experience A Cue Requires
Putting a number on adequate opportunity to learn, which the literature names and does not quantify. Our own arithmetic, standard statistics on invented parameters.
A cue is worth learning if it correlates with the outcome. To detect a correlation of r at conventional standards, being 80 percent power and a five percent significance level, the observations required are:
At r equals 0.50: 30 observations. At 0.40: 47. At 0.30: 85. At 0.25: 124. At 0.20: 194. At 0.15: 347. At 0.10: 783.
Four observations.
These are the numbers for a researcher with clean data and a defined outcome measure, which is generous. A person learning informally from messy feedback needs more, not fewer.
The curve is brutal at the low end. Halving the validity from 0.20 to 0.10 quadruples the requirement, from 194 observations to 783.
And the validities that matter commercially sit exactly in the punishing range. The sixty-seventh article reported selection validities in the region of 0.25 to 0.40, which on this table needs between 47 and 124 cases.
Note that detecting a cue is a lower bar than using it well. These figures are what it takes to establish that a signal exists at all, not to calibrate how much weight it deserves.
Twenty Years Is Not Enough
Applying that to a working life. Ours.
At twenty such decisions a year, a cue of validity 0.30 takes four years, one of 0.20 takes ten, and one of 0.10 takes thirty-nine.
At two such decisions a year, which is closer to how often a small firm hires or changes a supplier, a cue of validity 0.30 takes forty-two and a half years.
Four observations.
The decisive variable is frequency, not tenure. Twenty years of a decision made twice a year is forty observations, which detects nothing below a validity of about 0.40.
That reframes the claim to experience precisely. The question is not how long you have been doing this but how many times you have done this particular thing, and most owners have never counted.
It also explains why some intuitions are excellent and others worthless in the same person. A restaurateur has served a hundred thousand covers and hired eleven people, and their judgment about the first should be trusted far more than about the second.
And a confident view forms long before the evidence could support one. Nobody withholds judgment for forty-two and a half years, so the belief arrives early and the data never arrives at all.
Two things about that figure deserve stating precisely, because it is the one a reader will quote. Eighty-five observations at two a year is 42.5 years, and we have rounded nothing in either direction; our own draft said forty-two and we corrected it against the computation.
And the figure is a floor rather than a forecast. It assumes every one of those eighty-five decisions is the same decision, measured the same way, in an unchanged market. Any drift in the setting resets part of the count, which is the two-settings problem arriving inside the sample-size problem and making it worse.
And You Do Not See All The Outcomes
The second problem, which is worse than the first. Our own arithmetic, invented parameters, assuming normal distributions and defining success as above-median performance.
Everything above assumed you observe what happened in every case. You do not. You observe outcomes for the people you hired, the suppliers you chose and the deals you took.
Suppose you select the top share by your own rating, and your rating has validity r against eventual performance.
Hiring the top 20 percent on a rating of validity 0.30: success rate among those hired, 66.8 percent. Among those rejected, 45.8 percent.
Top 20 percent, validity 0.50: hired 78.1 percent, rejected 43.0 percent.
Top 50 percent, validity 0.30: hired 59.7 percent, rejected 40.3 percent.
Top 50 percent, validity 0.50: hired 66.6 percent, rejected 33.4 percent.
The Number You Never See
Reading that table. Ours.
Four observations.
Take the first row, which uses a validity below what the selection literature reports for good methods. You see two thirds of your hires succeed and conclude you have a good eye.
The people you turned down would have succeeded 45.8 percent of the time. That number is not merely unknown to you; it is unknowable, because they went elsewhere.
So the visible evidence is entirely consistent with a rating that is barely better than a coin. A 66.8 percent hit rate feels like proof and is compatible with a validity of 0.30, which the sixty-seventh article's arithmetic said a simple model would beat.
And this is the sixty-sixth article's missing cell in its most expensive form. Two of the four cells are invisible by construction, and no amount of diligence recovers them.
What The Visible Number Tells You
The decisive test of whether experience can teach here. Our own arithmetic, hiring the top 20 percent throughout.
At validity 0.10: visible success 55.6 percent. At 0.20: 61.2. At 0.30: 66.8. At 0.40: 72.4. At 0.60: 83.9.
Four observations.
Even a near-worthless rating of 0.10 produces a visible success rate above half, because you are selecting from the top of a distribution and the top performs better than average whatever your rating does.
The visible number rises smoothly with validity, so it contains information in principle. But you would need to know the base rate and the selection ratio to extract it, and almost nobody records either.
Which is the formal reason experience does not teach in this setting. The observable statistic does not identify the quantity you care about, and running it for longer does not change that.
And it explains the pattern the two research camps were arguing about. A person in this situation gains confidence with every year and no accuracy at all, which is exactly what the heuristics camp found and what the naturalistic camp found absent in domains like firefighting.
The Four Cases
A framework that falls out of the two conditions being independent, which neither source we obtained draws out. Ours.
The 2009 paper names two necessary conditions: a valid environment, and an adequate opportunity to learn it. Two conditions each met or not gives four cases, and they are not equally common or equally dangerous.
Valid environment, adequate practice. Skilled intuition develops and should be trusted. Daily operational judgment for most owners sits here.
Valid environment, insufficient practice. The cues exist and you have not learned them. Honest ignorance, and it is fixable by more repetitions or by borrowing someone else's.
Invalid environment, insufficient practice. Nothing to learn and no chance to learn it. Uncomfortable, and at least it usually feels uncertain.
Invalid environment, extensive practice. This is the dangerous one, and it has no natural warning signal.
Four observations.
The fourth cell is where confidence and accuracy come apart completely. Years of repetition in a domain with no valid cues produces the feeling of expertise and none of the substance.
It is also the cell that looks most like the first from inside and from outside. Both contain a practitioner with long tenure and firm opinions, and the difference is a property of the domain that neither party has examined.
And it explains the two research camps exactly. The naturalistic camp studied the first cell and the heuristics camp studied the fourth, and both reported their cell accurately.
The practical value is that you can locate a decision in this grid without assessing anyone, which is the whole point of a criterion about environments rather than people.
Does This Prove Too Much?
The objection we would raise against our own arithmetic, because taken alone it says almost nothing is learnable. Ours.
Four observations.
Our sample-size table implies that a cue of validity 0.20 needs 194 observations, and by that standard most professional judgment could never develop at all. Something is wrong with reading it that way.
What is wrong is that the table computes what it takes to establish that one pre-specified cue exists, using a formal test. That is a researcher's problem and not a learner's.
A practitioner is doing something different and easier. They are absorbing many weak cues at once and combining them, and a bundle of correlated weak signals can reach a usable validity long before any one of them could be isolated statistically.
So the honest version of our claim is narrower than the table suggests, and we state it that way. Where feedback is fast and instances are numerous, learning happens through mechanisms our arithmetic does not model. Where instances are few and feedback is delayed, neither the formal route nor the informal one has the material to work with, and the table is a reasonable guide to which situation you are in.
What Actually Survives
Our reading, stated directly.
Five statements.
Whether experience produces skill depends on the domain, not the person. Two researchers who spent careers disagreeing published a joint paper saying so.
Two conditions are necessary: an environment containing valid cues, and an adequate opportunity to learn them.
Kind means the learning setting matches the target setting, on the framework's own definition, which is stricter than the popular version about fast feedback.
Confidence carries no information about which case you are in, because it tracks the internal consistency of your evidence rather than its quality.
And the arithmetic of learning is unforgiving. On our own figures a cue of validity 0.30 needs 85 observations, and the outcomes you observe are selected in a way that flatters any rating at all.
The Position This Puts Us In
Applying the criterion to advisory work, including our own, because it would be evasive not to. Ours.
Four observations.
Professional advice divides on exactly this line and the division is uncomfortable. Bookkeeping, compliance and reconciliation are high-frequency work with fast unambiguous feedback, and skill there is real and earned.
Forecasting, valuation and strategic recommendation are not. Few instances per adviser, outcomes arriving years later with many causes, and no observation of the path not taken.
Which means an adviser's confidence in the second category deserves the same discount as anyone else's. Seniority in a firm is accumulated tenure, and tenure is the wrong unit, as the section above sets out.
And the useful consequence for a client is a question rather than a scepticism. Ask which category a given piece of advice falls into, and ask what the adviser's own record on that category looks like, which is the seventy-third article's test arriving from a different direction.
Classifying Your Own Decisions
The practical procedure, and it takes an afternoon. Ours, untested.
Four questions per recurring decision.
How many times have I made this exact decision? Not how many years, but how many instances. Against the table above, under fifty is not enough to learn anything with a validity below 0.40.
How long until I find out, and do I find out clearly? A supplier choice you assess in a week is different from a hire you assess in two years and a strategy you never assess at all.
Do I see outcomes for the options I rejected? If not, your success rate is the flattering number computed above rather than evidence about your judgment.
Has the setting changed since I learned it? This is the two-settings test, and a mismatch makes prior experience actively misleading rather than merely useless.
Hiring
The first application, and the clearest wicked environment most owners face. Ours, and nothing here bears on the lawfulness of any employment practice.
Four points.
It fails on frequency for a small firm. A handful of hires a year cannot support learning a cue of the validity the selection literature actually reports.
It fails on feedback. You learn whether a hire worked out over a year or more, by which point the outcome has many causes and your original impression has been rewritten, which is the hindsight article's finding.
And it fails on selection hardest of all. You never observe the rejected candidates, which on our own arithmetic is what makes a 0.30 rating feel like a 0.70 one.
Which is why the sixty-seventh article's recommendation is the right one here specifically. Score the factors separately and add them, because in this environment the alternative is not intuition but the appearance of it.
Pricing And Quoting
The second application, and it is the good news. Ours.
Four points.
Quoting passes the frequency test easily for most firms. Hundreds of quotes a year is a different order of magnitude from a handful of hires, and the table above says hundreds is enough for a cue of moderate validity.
Feedback is fast and unambiguous in one direction. You learn quickly whether a quote was accepted, which is a real cue about pricing arriving in days.
But the selection problem returns in a specific form. You learn whether the job was profitable only for the quotes that were accepted, so your sense of what price wins work is well trained and your sense of what price is profitable is not.
And that split is the useful output. Trust your instinct about whether a quote will land and distrust your instinct about whether it will pay, because only the first gets clean feedback.
That asymmetry has a consequence worth stating, because it runs in a predictable direction. A pricing instinct trained only on acceptance drifts downward over time, since low quotes are accepted more often and therefore generate more of the feedback the instinct is learning from.
The correction is cheap and almost nobody does it. Record the estimated margin on every quote at the time you send it, and compare it against the realised margin on the jobs that landed, which supplies the second cue the environment withholds.
Strategy
The third application, and the least learnable thing most owners do. Ours.
Four points.
A multi-year strategic judgment fails every condition at once. Few instances, feedback delayed by years, outcomes with many causes, and no observation of the paths not taken.
The two-settings test fails too, and by definition. A strategy matters because the future differs from the past, so the setting where the experience was acquired cannot match the setting where it is applied.
This does not mean strategy should be abandoned to models, and we would not claim it. The sixty-seventh article's mechanical alternatives require a defined outcome measure, which strategy does not supply either.
What it means is narrower and more useful. Confidence in a strategic call carries no information, and should not be treated as an argument in a room, which is the finding on internal consistency applied to a meeting.
Which leaves a question this article should answer rather than dodge. If strategic judgment cannot be learned and cannot be mechanised, what is left?
Our answer is that the decision still has to be made, and the change is in how it is held rather than how it is reached. Make the call, write down what you expect and why, name what would show you were wrong, and revisit it on a date you set now. That does not make the environment kinder, and it converts an unfalsifiable conviction into something a later meeting can examine.
Where Your Intuition Is Trustworthy
The section this article needs, because it would otherwise recommend distrusting yourself entirely. Ours.
Four observations.
The 2009 paper is not a debunking of intuition and it would be a misreading to take it as one. It says skilled intuition is real and specifies where.
The conditions are met more often than a pessimistic reading suggests. Anything you do daily, with a result you see within hours, in a setting that has not changed, satisfies both.
For most owners that covers a great deal. Judging whether a job will run over, whether a customer is difficult, whether a batch is right, whether a room has gone wrong, are all high-frequency decisions with fast feedback.
And the discipline is to keep the two categories apart. The same person's intuition can be excellent about the shop floor and worthless about the acquisition, and the feeling of certainty is identical in both.
That last point is the one we would leave a reader with, because it is where the cost lands. There is no internal signal distinguishing the two, so the sorting has to be done deliberately and in advance, on the properties of the decision rather than on how the decision feels.
One further point that cuts against the pessimistic reading, and it is the authors' own. The 2009 paper's title is a failure to disagree, and what the two authors failed to disagree about includes the reality of expert intuition, not only its limits.
Making A Wicked Environment Kinder
What can actually be changed, since the classification is not a life sentence. Ours, untested.
Four observations.
You can shorten the feedback. A hire assessed at ninety days on stated criteria produces a cue years earlier than one assessed by eventual tenure, and an earlier cue is a learnable one.
You can recover some of the missing cell. Keeping notes on rejected candidates and checking where a few landed is cheap, imperfect, and more than nothing, which is the current position.
You can increase the effective sample by pooling. Three owners in the same trade comparing outcomes have three times the observations, which on the table above can move a decision from unlearnable to learnable.
And you can record the prediction, which this series has now recommended in ten articles. Without it there is no cue to learn from, because the impression you formed has been rewritten by the outcome before you compare them.
Two of those four are worth ranking, because they are not equally effective and most firms reach for the wrong one. Shortening feedback is the highest-value move, since it acts on the binding constraint: a cue that arrives in ninety days rather than two years can be learned within a working life.
Pooling is the second, and it is the only one that changes the sample size rather than the sample quality. Three firms comparing hiring outcomes turn a decision made twice a year into one observed six times a year, which on our own table moves a cue of validity 0.30 from forty-two and a half years to fourteen.
And we would put recovering the missing cell last, not because it is useless but because it is the hardest to do honestly. Checking where a few rejected candidates landed produces a small, self-selected, unmatched sample, and reading it as a validity estimate would repeat the error the sixty-sixth article warned about.
What To Do
Count instances, not years. The condition is an adequate opportunity to learn, and twenty years of a decision made twice a year is forty observations.
Check the table before trusting a cue. On our own arithmetic a validity of 0.30 needs 85 observations, 0.20 needs 194, and 0.10 needs 783.
Ask whether you see the outcomes you rejected. If not, your visible success rate flatters any rating: on our figures a 0.30 rating shows 66.8 percent success while rejected candidates would have reached 45.8.
Apply the two-settings test. Does the setting where you learned this match the setting where you are applying it? Fast feedback about a market that has changed is a wicked environment.
Stop treating confidence as evidence. The 2009 paper states that subjective confidence tracks the internal consistency of the information rather than its quality, and that redundant flimsy evidence produces the most confident judgments.
Separate your decisions into two piles and treat them differently. Daily work with fast feedback earns intuition; rare decisions with delayed outcomes do not, in the same person on the same day.
Shorten feedback where you can. A hire assessed at ninety days on stated criteria produces a learnable cue years before eventual tenure does.
Write the prediction down. Without a record there is nothing to learn from, because your original impression is rewritten by the outcome before you can compare them.
The Limits Of This Analysis
Several caveats matter. This article discusses research on judgment and is not hiring, management or strategic advice; nothing here bears on the lawfulness of any employment practice, and the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 2009 paper in full, only its abstract from two independent indexes and selected body passages from a copy hosted by a university business school, so we report none of its evidence, examples or its several suggestions for improving judgment. We did not obtain the 2016 management paper, only its abstract, which means the several types of mismatch it says it specifies are not reported here, nor are the illustrations it draws for managers; those would be the most useful parts and the applications below are ours instead. We did not obtain the 2015 paper, the 2001 book or the 1978 paper and report them from citation records. The historical account of the two research camps comes partly from a writer's newsletter, not an academic source, flagged at every use. All arithmetic is ours. The sample-size figures are standard and assume clean data, a defined outcome measure and a single pre-specified cue, all of which flatter informal learning; a person learning from messy feedback needs more observations, not fewer. The selection tables assume normal distributions, a linear relationship between rating and performance, success defined as above-median performance, and validity and selection ratios we invented; different assumptions move every figure. And the framework classifies decisions rather than people, so nothing here establishes that any particular reader's judgment is unreliable, only that certain decisions cannot supply the evidence that would establish it either way.
Frequently Asked Questions
When does experience actually produce skill?
What is a kind versus a wicked learning environment?
How many observations does learning a cue need?
Why does my track record feel better than that?
Does feeling confident mean anything?
Is my intuition ever worth trusting?
Can I make a wicked environment kinder?
References
- Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: a failure to disagree. American Psychologist, 64(6), 515–526, September, DOI 10.1037/a0016755. National medical library database record reproducing the abstract in full: on the article reporting an effort to explore the differences between two approaches to intuition and expertise often viewed as conflicting, heuristics and biases and naturalistic decision making; on the authors starting from the obvious fact that professional intuition is sometimes marvelous and sometimes flawed and attempting to map the boundary conditions that separate true intuitive skill from overconfident and biased impressions; and on their concluding that evaluating the likely quality of an intuitive judgment requires an assessment of the predictability of the environment in which the judgment is made and of the individual's opportunity to learn the regularities of that environment. The record also gives the first author's affiliation at the time as the Woodrow Wilson School of Public and International Affairs, Princeton University. Note: a national medical library database record. Our source for the abstract verbatim; we did not obtain the paper in full and report none of its evidence, examples or recommendations. pubmed.ncbi.nlm.nih.gov
- Copy of the same paper hosted by a university business school, reproducing body passages: that the authors explore two necessary conditions for the development of skill, being high-validity environments and an adequate opportunity to learn them; that naturalistic decision making researchers compare the performance of professionals with that of the most successful experts in their field, whereas heuristics and biases researchers prefer to compare the judgments of professionals with the outcome of a model that makes the best possible use of available information; that subjective confidence is often determined by the internal consistency of the information on which a judgment is based rather than by the quality of that information, citing Einhorn and Hogarth (1978) and Kahneman and Tversky (1973); and that as a result, evidence that is both redundant and flimsy tends to produce judgments held with too much confidence, which will be presented too assertively to others and are likely to be believed more than they deserve to be. Note: a copy hosted by a university business school. Our source for the two necessary conditions, the methodological difference between the two camps, and the passage on confidence; we obtained selected passages and not the paper in full. bear.warrington.ufl.edu
- Hogarth, R. M., & Soyer, E. (2016). Kind and Wicked Experience in Marketing Management. Journal of Marketing Behavior, 2(2–3), 81–99, December, DOI 10.1561/107.00000031. Economics database record reproducing the abstract in full: on society venerating experience and on it feeling right to trust our own experience and that of others; on experience also having adverse effects, with much learning being tacit in nature and, because people are typically unaware and uncritical of the conditions in which this takes place, experience leading to false beliefs and subsequent actions reinforcing biases; on the authors adopting a two-settings framework in which experience is conceptualized as acquired in one setting, learning, and applied in another, target; on the learning environment being kind when information in the two settings match, and wicked environments being characterized by mismatches of several specified types; on many inferential errors occurring because people implicitly assume informational matches between the two settings; and on the framework having normative implications illustrated through decision-making challenges faced by marketing managers. Note: an economics database record. Our source for the two-settings framework verbatim; we did not obtain the paper, so the several types of mismatch and the managerial illustrations are not reported here. ideas.repec.org
- Reference list carried on a peer-reviewed journal article on intuition in management decision-making, confirming Hogarth, R. M., Educating intuition, University of Chicago Press, Chicago, 2001; Hogarth, R. M., Intuition: A challenge for psychological research on decision making, Psychological Inquiry, 2010, 21, 338–353, DOI 10.1080/1047840X.2010.520260; and Hogarth, R. M., Lejarraga, T., and Soyer, E., The two settings of kind and wicked learning environments, Current Directions in Psychological Science, 2015, 24, 379–385, DOI 10.1177/0963721415591878. Note: a reference list carried on a peer-reviewed article; citations only. We obtained none of the works named. imrpress.com
- Bibliographic service record for the 2009 paper, reproducing the same abstract word for word as the national medical library record and confirming the citation as American Psychologist, volume 64, number 6, pages 515–526, published 1 September 2009. Note: a bibliographic service, used as an independent corroboration of the abstract text and citation. semanticscholar.org
- Writer's newsletter article on kind and wicked learning environments, recording that one research camp, naturalistic decision making, found that performers reliably improve with experience, while the other, heuristics and biases, found that people not only frequently do not improve with experience but sometimes get worse; that both camps had produced fascinating work and were well respected; that in 2009 two leading figures in those areas, Gary Klein and Daniel Kahneman, produced what Kahneman called an adversarial collaboration paper; and that they concluded whether experience alone reliably predicted exceptional performance hinged on the characteristics of the domain in question. Note: a writer's newsletter, not an academic source, flagged at every use. Our source for the historical account of the two camps only; the substantive claims attributed to the 2009 paper here are sourced from references 1 and 2 instead. davidepstein.substack.com
This article discusses research on judgment and is not hiring, management or strategic advice; nothing here bears on the lawfulness of any employment practice. Neither underlying paper was obtained in full: the 2009 paper is reported from its abstract and selected body passages, and the 2016 paper from its abstract, so the types of mismatch it specifies and the managerial illustrations it draws are not reported. One source is a writer's newsletter, flagged at every use. All arithmetic is the authors' own and assumes clean data, a defined outcome measure, normal distributions, and validity and selection figures the authors invented; informal learning needs more observations than these figures show, not fewer.