You interview a candidate, read a proposal, or assess a prospective client. You have several pieces of information and you weigh them up. This article is about a large literature saying that last step, the weighing up, is where the accuracy goes, and that a formula written on an index card would do it better.
Key Takeaway
A meta-analysis of studies in human health and behaviour reports that "on average, mechanical-prediction techniques were about 10% more accurate than clinical predictions" and that superiority was "consistent, regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined."[1] A separate meta-analysis of selection decisions reports "an improvement in prediction of more than 50%" from combining the same data mechanically[2].
The Verdict, Stated First
Five claims, in descending order of confidence.
One. This is among the most consistent findings this series has reported. Two meta-analyses, one across health and behaviour and one across selection and admissions, both favouring mechanical combination, with the first reporting the advantage held regardless of the judges' experience.
Two. The mechanism does not require the model to be smarter. On our own arithmetic, a judge who agrees with themselves only 80 percent of the time on a repeat rating can capture at most 89.4 percent of the validity of their own rule. A formula applying that same rule captures all of it.
Three. The commercial size is large and specific. The selection meta-analysis reports better than a 50 percent improvement in predicting job performance, and on our own illustration that is worth roughly eight percentage points more good hires when selecting the top ten percent of applicants.
Four. The advantage is smaller than the headline in most individual studies. The same abstract reporting a ten percent average advantage also says clinical predictions were "often as accurate." A tie is the single most common result.
Five. Almost nobody uses it, and the literature knows. A 2013 meta-analysis opens by noting holistic methods "continue to be relied upon and preferred by practitioners," and a paper title from the same field calls that reliance stubborn.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the 2000 meta-analytic result, from an abstract obtained verbatim, reporting on studies of human health and behaviour.
Grade A for the 2013 selection result, from an abstract obtained verbatim from the authors' own institutional copy.
Grade B for the 63/65/8 breakdown, which reaches us through a popular book quoted on a reader-notes site rather than from the paper, and which we flag at every use.
Grade C for the defence literature, which we can name and not report, having obtained one truncated sentence.
Grade A for our own arithmetic, which is standard psychometrics and reproducible.
Our position: the finding is unusually well supported and the typical result is a tie rather than a rout, and both halves belong in any honest summary.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2000 meta-analysis's abstract verbatim from a source reproducing it in full[1]. We did not obtain the paper, and that source is a personal blog, flagged at every use, though the citation is corroborated by four independent academic reference lists[3].
We obtained the 2013 meta-analysis's abstract verbatim from a copy hosted by the lead author's own university laboratory[2]. We did not obtain the full paper.
The 63/65/8 study breakdown comes from a popular book, quoted identically by two independent readers on a book-notes site[4]. That corroborates the book's wording and not the underlying paper, and we grade it accordingly.
We did not obtain the 1954 book, the 1979 paper, the 1989 Science article, or any of the defence literature, and report all from citation records and titles.
All arithmetic is ours. The psychometric calculations are standard; the validity figures used to illustrate the selection result are invented.
This article discusses research on judgment. It is not hiring, credit, legal or clinical advice, and nothing here bears on the lawfulness of any employment or lending practice.
A Seventy-Year-Old Question
Where this starts.
Reference lists identify the founding work as Meehl, P. E. (1954), Clinical vs. Statistical Prediction: A Theoretical Analysis and a Review of the Evidence[1], and the field's best-known summary as Dawes, R. M., Faust, D., and Meehl, P. E. (1989), Clinical versus actuarial judgment, Science, 243(4899), 1668–1674[5].
We obtained neither and report the citations.
Four observations, ours.
The question is narrower than it sounds. It is not whether humans or machines are better at everything. It is specifically about the final step: how several pieces of information get combined into one prediction.
That distinction matters commercially because the human is often irreplaceable at the earlier step. Someone has to notice what is worth measuring, and the literature's own framing concedes this, as a section below sets out.
A defence paper from 2006 records the state of play: "Despite over 50 years of one-sided research favoring formal prediction rules over human judgment, the 'clinical-statistical controversy,' as it has come to be known, remains something of a hot-button" issue, at which point our source truncates[6].
And that sentence is written by people defending clinical judgment, which makes the phrase "over 50 years of one-sided research" a concession from the side that would most like it to be otherwise.
The 2000 Meta-Analysis
The central quantitative result.
Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., and Nelson, C. (2000), Clinical versus mechanical prediction: A meta-analysis, Psychological Assessment, 12(1), 19–30, March, DOI 10.1037/1040-3590.12.1.19[3].
Its abstract opens: "The process of making judgments and decisions requires a method for combining data. To compare the accuracy of clinical and mechanical (formal, statistical) data-combination techniques, we performed a meta-analysis on studies of human health and behavior."[1]
Three observations, ours.
The first sentence is the framing to keep. Combining data requires a method, and doing it in your head is a method rather than the absence of one. The comparison is between two methods, not between rigour and instinct.
The domain is human health and behavior, which is broad and is not business. Everything we say about commercial application below is transfer, and we flag it as ours.
And we obtained the abstract from a personal blog that reproduces it in full. The citation is confirmed by four academic reference lists, but the abstract text itself rests on a non-academic source and we flag it here and at every use.
About Ten Percent More Accurate
The headline figure.
"On average, mechanical-prediction techniques were about 10% more accurate than clinical predictions. Depending on the specific analysis, mechanical prediction substantially outperformed clinical prediction in 33%-47% of studies examined."[1]
Four observations, ours.
Ten percent more accurate is a modest average and it is worth saying so plainly. This is not a claim that human judgment is worthless.
The range 33 to 47 percent is a range across analyses rather than a confidence interval, and the abstract is explicit that it depends on which analysis you run. That is unusually candid.
So on the most favourable reading, mechanical prediction substantially wins in fewer than half the studies. The word substantially is doing work, and a small win is not counted.
And the commercial reading of a ten percent accuracy gain depends entirely on the base. Ten percent more accurate on a decision you make four hundred times a year is a different proposition from ten percent on one you make twice.
The Hedge In The Same Abstract
The sentence that keeps this honest.
"Although clinical predictions were often as accurate as mechanical predictions, in only a few studies (6%-16%) were they substantially more accurate."[1]
Four observations, ours.
"Often as accurate" is the authors conceding that a tie is common, in the abstract, unprompted. This series has repeatedly praised that behaviour and it is present here.
The structure of the result is therefore three-way rather than two-way. Mechanical substantially better in 33 to 47 percent, clinical substantially better in 6 to 16 percent, and roughly the balance a tie.
The asymmetry is the finding. Mechanical wins substantially between two and eight times as often as clinical does, depending which end of each range you take.
And a tie is not a null result here. A method that ties with expert judgment while being faster, cheaper and auditable has won on every dimension except accuracy, which is a point the abstract does not make and we do.
Regardless Of Experience
The moderator finding, and it is the one that will annoy readers most.
"Superiority for mechanical-prediction techniques was consistent, regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined."[1]
Four observations, ours.
Four moderators tested, four null. Task, judge type, experience and data type all failed to identify a condition where clinical combination came out ahead.
Experience is the one that matters commercially. The obvious defence of holistic judgment is that a novice should use a checklist and a veteran should not, and this abstract says the advantage did not vary with experience.
We would enter a caution the abstract permits. A null moderator finding in a meta-analysis is weaker evidence than a positive one, because failure to detect variation may reflect limited power, and we did not obtain the analyses.
And it connects to the fifty-ninth article's finding directly. That article reported deliberate practice explaining under one percent of variance among professionals, and this one reports experience failing to moderate a judgment advantage. Two different literatures, the same uncomfortable direction.
And Worse With Interviews
The one condition the abstract does single out.
"Clinical predictions performed relatively less well when predictors included clinical interview data."[1]
Four observations, ours.
This is a moderator that did work, in a paper reporting four that did not, which makes it the most informative sentence in the abstract.
The direction is counterintuitive and important. Adding the richest, most personal source of information made human judgment relatively worse, not better.
The likely mechanism is not mysterious and is ours. An interview supplies a great deal of vivid material of unknown validity, and vivid material of unknown validity is exactly what the sixtieth and sixty-sixth articles described as displacing better evidence.
And this publication already has an article on what the selection literature supports about interviews specifically, which we would read alongside this one.
Sixty-Three, Sixty-Five, Eight
The study-by-study breakdown, from a source we grade carefully.
A widely read book on judgment reports: "A 2000 review of 136 studies confirmed unambiguously that mechanical aggregation outperforms clinical judgment. The research surveyed in the article covered a wide variety of topics, including diagnosis of jaundice, fitness for military service, and marital satisfaction. Mechanical prediction was more accurate in 63 of the studies, a statistical tie was declared for another 65, and clinical prediction won the contest in 8 cases."[4]
It adds: "These results understate the advantages of mechanical prediction, which is also faster and cheaper than clinical judgment."[4]
Three observations, ours.
This is a popular book, quoted on a book-notes site, not a paper. Two independent readers reproduce the passage identically, which corroborates the book's wording and tells us nothing about whether the book read the paper correctly.
The word "unambiguously" is the book's, not the paper's, and the paper's own abstract is considerably more hedged. We would not use it.
And the observation that the results understate the advantage because mechanical methods are faster and cheaper is the same point we made above about ties, and it is a fair one.
A Consistency Check
Testing the book's numbers against the paper's own abstract. Our own arithmetic.
The breakdown gives 63 of 136, or 46.3 percent, for mechanical; 65, or 47.8 percent, for ties; and 8, or 5.9 percent, for clinical. The three sum to 136 exactly.
The abstract gives ranges of 33 to 47 percent for mechanical and 6 to 16 percent for clinical.
Three observations.
The book's figures sit at the very top of the mechanical range and the very bottom of the clinical range, which is internally consistent and is the most favourable corner of both.
That is a genuine check that the book is describing the same paper, and it is not a check that it chose a representative analysis. The abstract says the figure depends on the specific analysis, and the book reports one.
And the ratio is worth stating either way. Mechanical wins 7.88 times as often as clinical does on those counts, which is the shape of the result whichever analysis you prefer.
The 2013 Selection Meta-Analysis
The study that moves this from health into hiring.
Kuncel, N. R., Klieger, D. M., Connelly, B. S., and Ones, D. S. (2013), Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis, Journal of Applied Psychology, 98(6), 1060–1072, November[2].
Its abstract opens with the problem: "In employee selection and academic admission decisions, holistic (clinical) data combination methods continue to be relied upon and preferred by practitioners in our field."[2]
And describes the design: "This meta-analysis examined and compared the relative predictive power of mechanical methods versus holistic methods in predicting multiple work (advancement, supervisory ratings of performance, and training performance) and academic (grade point average) criteria."[2]
Three observations, ours.
Four outcome measures, three of them work-related and objectively varied: advancement, supervisory ratings, and training performance. That breadth matters because any one of them could be criticised alone.
The opening sentence is a statement about practice, not evidence, and it is doing something deliberate. The authors are flagging a gap between what their field knows and what it does.
And we obtained this abstract from a copy hosted by the lead author's own university laboratory, which is the best provenance available short of the publisher.
More Than Fifty Percent
The result.
"In predicting job performance, the difference between the validity of mechanical and holistic data combination methods translated into an improvement in prediction of more than 50%."[2]
Four observations, ours.
More than 50 percent is far larger than the 10 percent average in the health literature, and the difference deserves an explanation we cannot give from the abstracts.
One plausible reason, ours and untested: selection decisions involve exactly the kind of rich, vivid, low-validity material, being interviews and impressions, that the 2000 abstract identified as the condition under which clinical prediction does relatively worse.
The phrase "improvement in prediction" is not the same as an improvement in outcomes, and translating one into the other requires assumptions the abstract does not supply. We make those assumptions explicitly in a section below and mark them as ours.
And we did not obtain the paper, so we cannot tell you the underlying validity coefficients, the number of studies, or the sample sizes.
Even By Experts Who Know The Job
The clause that forecloses the obvious defence.
"There was consistent and substantial loss of validity when data were combined holistically—even by experts who are knowledgeable about the jobs and organizations in question—across multiple criteria in work and academic settings."[2]
Four observations, ours.
The interpolated clause is the paper doing our work for us. The natural objection is that a general finding will not apply to someone who knows this particular job, and the abstract addresses that specifically.
"Consistent and substantial loss of validity" is strong language, and the direction is worth restating because it is easy to read past. Combining data holistically loses information that was already there.
That framing is the useful one. The mechanical method is not adding anything. It is failing to lose things, which is a different and more modest claim than machines being cleverer.
And it sets up the arithmetic that follows, which explains how information gets lost in a step that feels like careful thought.
Why The Model Wins
The explanation, which is less about intelligence than anyone expects. Ours.
A judgment has two components: the rule you are applying, meaning which factors matter and how much, and the consistency with which you apply it.
Three observations.
Most discussion of this literature assumes the contest is about the first component, and that the model must have discovered a better rule.
It usually has not. The model's advantage comes mostly from the second component, and the next section shows how much is available there.
And that reframing is what makes the finding usable rather than demoralising. You do not have to accept that a formula understands your business better than you do. You only have to accept that you have bad days.
The Reliability Ceiling
Our own arithmetic, standard psychometrics.
If a judge shown the same case twice gives correlated but not identical ratings, that self-agreement is their reliability, and a measure cannot correlate with anything else more strongly than the square root of its own reliability.
So the share of their rule's validity a judge can actually achieve:
At self-agreement 1.00: 100.0 percent, nothing lost. At 0.95: 97.5 percent. At 0.90: 94.9 percent. At 0.80: 89.4 percent. At 0.70: 83.7 percent. At 0.60: 77.5 percent. At 0.50: 70.7 percent.
Four observations.
At a self-agreement of 0.80, which is generous for human judgment on complex material, the judge captures at most 89.4 percent of the validity of their own rule. Ten and a half percent is lost to inconsistency alone.
That loss is invisible from the inside. Nobody experiences their own inconsistency, because each individual judgment feels considered.
The size is in the right neighbourhood. A ten and a half percent loss sits close to the ten percent average advantage the 2000 meta-analysis reports, which is a suggestive correspondence and not a proof, since we obtained no reliability figures from either paper.
And we chose 0.80 ourselves. No source we obtained states the test-retest reliability of the judges in either meta-analysis, so this demonstrates a mechanism rather than measuring one.
One further consequence of that table deserves stating, because it points the remedy in an unexpected direction. The relationship is not linear. Moving self-agreement from 0.50 to 0.60 recovers about seven points of achievable validity; moving from 0.90 to 1.00 recovers about five.
So the largest gains are available to the least consistent judge, which is the opposite of how these interventions are usually sold. A scoring sheet is pitched at the careful and is worth most to the person having a difficult month.
And it explains something the meta-analyses report and do not explain. If consistency is the mechanism, the advantage should hold regardless of experience, because experience improves the rule and does nothing for the wavering. That is exactly what the 2000 abstract reports on its experience moderator, and we offer the connection as our own inference rather than as anything either paper claims.
The Model Does Not Need A Better Rule
The consequence, stated as sharply as we can. Ours.
Four observations.
A model built from a judge's own past decisions applies that judge's rule perfectly. It has no better insight, no extra information and no theory. It simply never wavers.
On the arithmetic above, that alone recovers the 10.6 percent lost at a self-agreement of 0.80, without improving the rule at all.
Which produces the result that gives this article its title. A model of a person can outperform that person, and the person cannot object that the model does not understand the job, because the model learned the job from them.
And it means the useful comparison is not you against a formula. It is you on a good day against you on all days, and the formula is simply the good day written down.
Improper Linear Models
How far this goes, from a title we can name and a paper we did not obtain.
Reference lists identify Dawes, R. M. (1979), The robust beauty of improper linear models in decision making, American Psychologist, 34(7), 571–582, and Dawes, R., and Corrigan, B. (1974), Linear models in decision making, Psychological Bulletin, 81, 95–106[6].
We obtained neither paper and report the titles and citations only.
Three observations, ours.
The term improper in that title is technical. It means a model whose weights were not fitted to the data, being for instance equal weights or weights chosen by hand.
The word robust alongside it signals the paper's argument, and the pairing is why the title is famous, but a title is not a finding and we will not tell you what it reports.
What we will say is that the framing is consistent with the reliability argument above. If most of the advantage comes from consistency rather than from optimal weights, then a crude consistent model should do most of the work, which is a prediction a reader can test on their own decisions.
What Fifty Percent Buys You
Translating the selection result into hires. Our own arithmetic, and the validity figures are invented.
Suppose the reported improvement of more than 50 percent means predictive validity rising from 0.25 to 0.38. Neither number comes from any source; they are our illustration.
Hiring the top share of applicants, the proportion of hires who end up above average on the job:
Top 50 percent: 57.9 against 61.9 percent, a gain of 4.0 points.
Top 25 percent: 62.5 against 68.5 percent, a gain of 6.1 points.
Top 10 percent: 67.0 against 74.8 percent, a gain of 7.8 points.
Top 5 percent: 69.7 against 78.3 percent, a gain of 8.6 points.
Four observations.
Hiring the top ten percent, the change is worth nearly eight percentage points on every hire, from combining the same information differently.
The gain is larger when you are more selective, which follows from the arithmetic: better prediction is worth more when you are cutting finer.
And the gain is smallest when you hire nearly everyone who applies, which is worth knowing for a small firm with two candidates. The method matters most where the competition is real.
These figures rest on invented validities and standard selection mathematics. They show a shape and they are not a forecast of anyone's hiring outcomes.
The Defence
What the other side argues, reported at the strength our sourcing allows.
Reference lists identify Dana, J., and Thomas, R. (2006), In defense of clinical judgment ... and mechanical prediction, Journal of Behavioral Decision Making[6], and Einhorn, H. J. (1972), Expert measurement and mechanical combination, Organizational Behavior and Human Performance, 7, 86–106[6].
We obtained neither paper.
Four observations, ours.
The 2006 title contains an ellipsis and the word "and", which signals a both-sides position rather than a rejection, and we would very much like to read it.
The 1972 title is the more informative of the two. Expert measurement and mechanical combination names a division of labour: the human supplies the inputs, the formula combines them.
That division is the position we would defend, and it is not a compromise. Nothing in either meta-analysis suggests a formula can decide what to measure, and the 2000 abstract's own framing is about data combination specifically.
And there is an objection neither paper needs to make, which we will make ourselves in a section below, about cases where a rule is knowably wrong.
Stubborn Reliance
Why none of this has changed practice, named by the field itself.
Reference lists identify Highhouse, S. (2008), Stubborn reliance on intuition and subjectivity in employee selection, Industrial and Organizational Psychology, 1, 333–342, and by the same author Facts are stubborn things, in the same volume[7].
We obtained neither and report the titles.
Four observations, ours.
Two papers, in one volume, by one author, both with titles about stubbornness. That is a field expressing frustration rather than reporting a finding.
The 2013 meta-analysis says the same thing in flatter language: holistic methods "continue to be relied upon and preferred by practitioners in our field"[2]. Note preferred, not merely used.
We think the resistance is rational in one specific way the literature underrates. A formula makes the decision auditable, and an auditable decision can be challenged, which a holistic impression cannot.
And that suggests the obstacle is not disbelief. People who adopt a rule accept accountability for it, which is a cost the accuracy literature does not price.
The Objection We Would Make
The strongest argument against everything above, which none of our sources raises and which a business owner should hear before acting. Ours.
Both meta-analyses compare mechanical and holistic combination on the same information. That is the correct comparison for their question and it is not the situation a firm is in.
Four observations.
A holistic judge can use information nobody wrote down. The candidate mentioned something in passing, the prospect's office looked wrong, the supplier's answer arrived too quickly. None of that is in the scoring sheet because nobody knew to put it there.
If any of that material carries validity, the studies understate holistic judgment, because the human in the experiment had access to a fixed set of predictors and the human in your office does not.
Against that, the 2000 abstract reports the one condition where clinical did relatively worse was when interview data was included, which is precisely the richest, least structured input. That is evidence the extra material hurts rather than helps, and it is the strongest reply available to our own objection.
And there is a test that settles it for your own firm without settling it in general. Record the score and the impression separately, and see which predicts better over a year. If your unrecorded material is worth something, your impression will beat your own sheet, and you will have found that out cheaply.
What Actually Survives
Our reading, stated directly.
Five statements.
Mechanical combination is better on average and the typical study is a tie. About 10 percent more accurate overall, substantially better in 33 to 47 percent of studies, substantially worse in 6 to 16.
The advantage did not vary with the judges' experience, on four moderators tested, though a null moderator is weaker evidence than a positive one.
It is larger in selection than in health. More than 50 percent improvement in predicting job performance, against 10 percent on average across health and behaviour.
The mechanism is consistency rather than insight. On our own arithmetic a judge at 0.80 self-agreement loses 10.6 percent of their own rule's validity, which a formula recovers by never wavering.
And the field knows nobody uses it, describing the reliance on holistic methods as preferred by practitioners and, in two paper titles, as stubborn.
Your Hiring
The first application. Ours, untested, and not hiring or legal advice; nothing here bears on the lawfulness of any employment practice.
Four points.
The finding is not that you should stop interviewing. It is that the step where you sit back and form an overall impression is where validity is lost, and that step is separable from everything before it.
The 2000 abstract's one working moderator is directly on point. Clinical prediction did relatively worse when interview data was among the predictors, which is exactly the situation in most hiring.
The mechanical alternative is unglamorous. Score each candidate on each criterion separately, write the scores down, and add them up rather than forming a gestalt.
And the reason that works is the reliability ceiling rather than any cleverness. Separate scores recorded at the time are not subject to the drift that a single overall impression is, which is the same instrument this series has now recommended in seven articles.
Your Clients And Your Credit
The second application, which most owners will find more immediately usable. Ours, and not credit or legal advice.
Four points.
Deciding which prospects to take on, and on what terms, is a prediction problem with a measurable outcome. Did they pay, did the work go over budget, did they come back.
Most firms make it holistically, and the raw material is the same rich low-validity kind the 2000 abstract flags: a meeting, a phone manner, a sense of how organised they seem.
The mechanical version is three or four scored factors with recorded weights, applied to every prospect. The weights matter less than the consistency, on the argument above.
And the outcome is recorded for you, which is the rare advantage this application has over hiring. You already know which past clients paid late, so you can check your rule against your own history rather than trusting it.
Build One This Week
The concrete version, because this article would otherwise be a complaint. Ours, untested.
Four steps.
Pick a decision you make at least monthly with an outcome you can observe. Client acceptance, quoting, supplier selection, hiring for one role.
Write down the three to five things you actually consider, in the terms you use, and give each a scale. This is the expert measurement step and only you can do it.
Score every case on every factor before forming any overall view, and add them. That is the whole mechanical combination step and it takes a spreadsheet column.
Record your holistic judgment separately and compare both against outcomes after a year. If your judgment wins, you have a defensible reason to keep it; if it does not, you have a model.
When The Human Is Needed
The limits of all this, and one of them is decisive. Ours.
Four observations.
Somebody has to decide what goes into the model, and no formula generates its own inputs. The 1972 title names this division and it is the right one.
A rule can be knowably wrong in a specific case, on information the rule was never built to hold. If you learn a candidate's referee was their sibling, no scoring sheet accounts for it and overriding is correct.
The danger is that this exception swallows the rule, because every case feels special from the inside. The discipline is to record each override and count them, and a rule overridden a third of the time is not being used.
And rules go stale. A model built on last decade's clients encodes last decade's market, which is a maintenance obligation the literature we obtained does not discuss and a business owner cannot ignore.
And one boundary the literature cannot help with at all. Both meta-analyses concern prediction, meaning questions with an answer that arrives later and can be checked. A great many business decisions are not of that kind: what to build, whether to enter a market, how to price something novel. No scoring sheet applies where the outcome is partly created by the decision itself, and nothing in this article should be read as covering those.
What To Do
Separate collecting information from combining it. The literature concerns the second step only, and the first is where human judgment is irreplaceable.
Score factors separately and add them, rather than forming an overall impression. A 2013 meta-analysis reports better than 50 percent improvement in predicting job performance from that change alone.
Do not expect experience to exempt you. The 2000 meta-analysis reports the advantage held regardless of judges' amounts of experience, and the 2013 one says loss of validity occurred even among experts knowledgeable about the specific jobs.
Be most careful where the evidence is most vivid. Clinical prediction did relatively worse specifically when interview data was included.
Remember the model does not need a better rule than yours. On our own arithmetic a judge at 0.80 self-agreement loses 10.6 percent of their own rule's validity to inconsistency, which a formula recovers for free.
Expect a tie more often than a win. A tie is the single most common outcome, and a method that ties while being faster, cheaper and auditable has still won on three dimensions.
Count your overrides. A rule you depart from a third of the time is not a rule, and every case feels exceptional from the inside.
Check the model against your own recorded history. For client and credit decisions the outcomes already exist, which makes this testable without waiting.
The Limits Of This Analysis
Several caveats matter. This article discusses research on judgment and is not hiring, credit, legal or clinical advice; nothing here bears on the lawfulness of any employment or lending practice, and the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain either meta-analysis, only their abstracts. The 2000 abstract reaches us from a personal blog reproducing it in full, flagged at every use; its citation is corroborated by four independent academic reference lists but the abstract text itself rests on a non-academic source. The 2013 abstract comes from a copy hosted by the lead author's own university laboratory, which is better provenance, and we still did not obtain the full paper, so we report no validity coefficients, no study counts and no sample sizes from it. The 63/65/8 breakdown comes from a popular book quoted on a book-notes site, reproduced identically by two independent readers, which corroborates the book's wording and not the paper; the book's characterisation of the result as unambiguous is considerably stronger than the paper's own abstract and we decline to adopt it. We did not obtain the 1954 book, the 1974 or 1979 papers, the 1989 Science article, the 1972 paper, the 2006 defence, or either 2008 commentary, and report all from citations and titles; a title is not a finding and we have said so at each. All arithmetic is ours. The reliability ceiling is standard psychometrics but the 0.80 self-agreement figure is our own choice and no source we obtained reports the reliability of the judges in either meta-analysis. The selection figures use invented validity coefficients of 0.25 and 0.38, chosen to illustrate a reported improvement of more than 50 percent, and standard selection mathematics; they show a shape and forecast nothing. The 2000 meta-analysis covers human health and behavior, so every commercial application here is our own transfer. And a null moderator finding, such as the one on experience, is weaker evidence than a positive one, because failure to detect variation may reflect limited power in analyses we did not obtain.
Frequently Asked Questions
What is the finding?
Does the model know something I do not?
Does experience protect me?
How often does the formula actually win?
What does it buy in practice?
When should I override the rule?
Why does nobody do this?
References
- Personal blog post reproducing in full the abstract of Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., and Nelson, C. (2000), Clinical versus mechanical prediction: a meta-analysis: on the process of making judgments and decisions requiring a method for combining data; on the authors performing a meta-analysis on studies of human health and behavior to compare the accuracy of clinical and mechanical, meaning formal and statistical, data-combination techniques; on mechanical-prediction techniques being on average about 10 percent more accurate than clinical predictions; on mechanical prediction substantially outperforming clinical prediction in 33 to 47 percent of studies examined depending on the specific analysis; on clinical predictions often being as accurate as mechanical predictions but substantially more accurate in only a few studies, being 6 to 16 percent; on superiority for mechanical-prediction techniques being consistent regardless of the judgment task, type of judges, judges' amounts of experience, or the types of data being combined; and on clinical predictions performing relatively less well when predictors included clinical interview data. The same post records the founding work as Meehl's 1954 book Clinical vs. Statistical Prediction: A Theoretical Analysis and a Review of the Evidence. Note: a personal blog, not an academic source, flagged at every use. It is our only source for the abstract text, which we did not obtain from the publisher; the citation itself is corroborated by four independent academic reference lists. medium.com
- Kuncel, N. R., Klieger, D. M., Connelly, B. S., & Ones, D. S. (2013). Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis. Journal of Applied Psychology, 98(6), 1060–1072. Copy hosted by the lead author's own university laboratory, reproducing the abstract in full: on holistic, meaning clinical, data combination methods continuing to be relied upon and preferred by practitioners in employee selection and academic admission decisions; on the meta-analysis examining and comparing the relative predictive power of mechanical versus holistic methods in predicting multiple work criteria, being advancement, supervisory ratings of performance, and training performance, and the academic criterion of grade point average; on there being consistent and substantial loss of validity when data were combined holistically, even by experts knowledgeable about the jobs and organizations in question, across multiple criteria in work and academic settings; and on the difference between the validity of mechanical and holistic data combination methods translating, in predicting job performance, into an improvement in prediction of more than 50 percent. Note: a copy hosted by the lead author's own university laboratory. We obtained the abstract in full and not the paper, so we report no validity coefficients, study counts or sample sizes. goal-lab.psych.umn.edu
- Reference list carried in the 2013 meta-analysis itself, confirming Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., and Nelson, C. (2000), Clinical versus mechanical prediction: A meta-analysis, Psychological Assessment, 12, 19–30, DOI 10.1037/1040-3590.12.1.19; Dawes, R. M. (1979), The robust beauty of improper linear models in decision making, American Psychologist, 34, 571–582; Highhouse, S. (2008a), Facts are stubborn things, Industrial and Organizational Psychology, 1, 373–376; and Highhouse, S. (2008b), Stubborn reliance on intuition and subjectivity in employee selection, Industrial and Organizational Psychology, 1, 333–342. Note: a reference list carried in a peer-reviewed paper, used to confirm the 2000 citation and DOI independently of the blog reproducing its abstract. gwern.net
- Reader-notes pages on a book-cataloguing site, reproducing a passage from a widely read book on judgment: that a 2000 review of 136 studies confirmed unambiguously that mechanical aggregation outperforms clinical judgment; that the research surveyed covered a wide variety of topics including diagnosis of jaundice, fitness for military service, and marital satisfaction; that mechanical prediction was more accurate in 63 of the studies, a statistical tie was declared for another 65, and clinical prediction won the contest in 8 cases; and that these results understate the advantages of mechanical prediction, which is also faster and cheaper than clinical judgment. Two independent readers reproduce the passage identically. Note: a popular book quoted on a book-notes site, not an academic source, flagged at every use. Independent reproduction corroborates the book's wording and not the underlying paper; the book's characterisation of the result as unambiguous is stronger than the paper's own abstract, and we decline to adopt it. goodreads.com
- Academic preprint reference list confirming R. M. Dawes, D. Faust, and P. E. Meehl, Clinical versus actuarial judgment, Science, volume 243, number 4899, pages 1668–1674, 1989, JSTOR stable identifier 1703476; W. M. Grove, D. H. Zald, B. S. Lebow, B. E. Snitz, and C. Nelson, Clinical versus mechanical prediction: a meta-analysis, Psychological Assessment, volume 12, number 1, pages 19–30, March 2000; N. R. Kuncel, D. M. Klieger, B. S. Connelly, and D. S. Ones, Mechanical versus clinical data combination in selection and admissions decisions: a meta-analysis, Journal of Applied Psychology, volume 98, number 6, pages 1060–1072, November 2013; and R. M. Dawes, A case study of graduate admissions: Application of three principles of human decision making, American Psychologist, volume 26, number 2, pages 180–188, February 1971. Note: a preprint reference list, used to confirm citations, volumes, issues, pages and months independently. We obtained none of the works named. arxiv.org
- Publisher record for Dana, J., and Thomas, R. (2006), In defense of clinical judgment ... and mechanical prediction, Journal of Behavioral Decision Making, reproducing the opening of its abstract, that despite over fifty years of one-sided research favoring formal prediction rules over human judgment the clinical-statistical controversy remains something of a hot-button issue, at which point the reproduction is cut off; and carrying a reference list confirming Dawes, R. M. (1979), American Psychologist, 34(7), 571–582; Dawes, R., and Corrigan, B. (1974), Linear models in decision making, Psychological Bulletin, 81, 95–106; Dawes, R. M., Faust, D., and Meehl, P. E. (1989a), Science, 243, 1668–1674, together with a published reply by the same authors; and Einhorn, H. J. (1972), Expert measurement and mechanical combination, Organizational Behavior and Human Performance, 7, 86–106. Note: the publisher's record. Our source for the defence literature, whose abstract is truncated mid-sentence in our reproduction; we obtained none of the papers named. onlinelibrary.wiley.com
- Journal article page carrying a reference list confirming Grove, W. M., Zald, D. H., Lebow, B. S., Snitz, B. E., and Nelson, C. (2000), Clinical versus mechanical prediction, Psychological Assessment, 12, 19–30; Highhouse, S. (2002), Assessing the candidate as a whole: A historical and critical analysis of individual psychological assessment for personnel decision making, Personnel Psychology, 55, 363–396; Highhouse, S. (2008), Stubborn reliance on intuition and subjectivity in employee selection, Industrial and Organizational Psychology, 1, 333–342; and Cooksey, R. W. (1996), Judgment analysis: Theory, methods, and applications. Note: a journal article's reference list; citations only. We obtained none of the works named, and report the 2008 commentaries from their titles alone. resolve.cambridge.org
This article discusses research on judgment and is not hiring, credit, legal or clinical advice; nothing here bears on the lawfulness of any employment or lending practice. Neither meta-analysis was obtained beyond its abstract. The 2000 abstract reaches this article through a personal blog, flagged at every use, with its citation corroborated by four academic reference lists. The 63/65/8 study breakdown comes from a popular book quoted on a book-notes site, and the book's characterisation of the result is stronger than the paper's own abstract. All other works are reported from citations and titles. All arithmetic is the authors' own; the self-agreement figure and the selection validity coefficients are invented and illustrate mechanisms rather than measuring anything.