Someone in your business has explained a colleague using the Dunning-Kruger effect. It is possibly the most widely deployed piece of psychology in ordinary professional conversation. There is a peer-reviewed literature arguing that the pattern behind it is mostly a statistical artefact, and it is straightforward to demonstrate why.

Key Takeaway

A 2020 paper in Intelligence is titled The Dunning-Kruger effect is (mostly) a statistical artefact: Valid approaches to testing the hypothesis with individual differences data[1]. It follows earlier work reporting that random number simulations reveal how random noise affects the measurements and graphical portrayals of self-assessed competency[2]. We simulated 40,000 people with identical self-assessment accuracy and no metacognitive deficit at all, and reproduced the classic pattern: the bottom quartile overestimated by 13 percentile points and the top quartile underestimated by 13.

Our Grades For These Claims

Applying the scheme from the first article in this series.

That the classic pattern can be produced by statistical artefact alone is Grade A. It is demonstrable arithmetically, we demonstrate it below, and multiple independent peer-reviewed papers have made the point.

That the observed pattern in real data is mostly artefact is Grade B, resting on the 2020 paper's title and conclusion and on the earlier simulation work.

That there is no metacognitive component at all is Grade D. One peer-reviewed source records researchers concluding the effect exists as an empirical phenomenon while disputing the explanation, and we found no basis for the stronger claim.

Our position: the chart is not evidence of what it is used to claim, which is a narrower and more defensible statement than saying the effect is not real.

A Note On Method

Everything here is verified to August 2026.

We obtained titles, citations and characterisations for the original paper and for the critical literature[1][2][3]. We did not obtain any of these papers in full, and report no effect sizes and no sample sizes from any of them.

The simulation in this article is entirely our own. Its parameters are invented, it comes from no paper, and it is offered as a demonstration of a mechanism rather than as a replication of anyone's data.

Several claims reach us through blogs and commercial websites, and we flag each.

This article reviews research methods. It is not management or human resources advice, and nothing in it should be used to characterise any individual.

A Warning About Our Sources

A caution we owe the reader before proceeding. This section is our own.

The sources we located on this topic are heavily weighted toward the critical side. Several are blogs with titles announcing the effect is not real. That is a selection problem in our own research, and a reader should discount accordingly.

Three things we did to compensate.

We anchored on peer-reviewed work where possible: the 2020 paper in Intelligence[1], the simulation papers in Numeracy[2], and a 2022 paper in Frontiers in Psychology which surveys both sides[3].

We gave the opposing evidence its own section, including the original authors' follow-up and a study concluding the effect is real.

And we ran the demonstration ourselves rather than relying on anyone's account of it, which is the one part of this article that does not depend on our sources at all.

The Original Claim

What was proposed.

Kruger and Dunning published Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments in the Journal of Personality and Social Psychology, 77(6), 1121–1134, in 1999[4].

A source summarises the popular understanding: low performers tend to overestimate their ability, often rating themselves above average, while top performers tend to underestimate how much better they are than the rest[5].

The proposed mechanism is often called a double curse or dual burden: the same deficit that makes someone perform badly also prevents them from recognising that they performed badly[6].

Two observations, ours.

The mechanism is what makes the claim interesting. Without it, "people who scored badly thought they did better than they did" is a considerably less remarkable statement.

And the mechanism is specifically about the bottom of the distribution. It asserts a deficit that is worse for the unskilled, which is a testable proposition distinct from a general tendency to self-flatter.

The Chart Is Not From The Paper

A detail worth establishing early.

The curve everyone has seen, rising to a peak of confidence at low competence and then collapsing into a valley, is usually labelled with names like Mount Stupid.

A source states plainly: "The viral 'Mount Stupid' curve is an invention. It does not appear in Kruger and Dunning's 1999 paper or any of Dunning's follow-up work; it was retrofitted onto the data in mid-2000s management memes."[6]

This comes from a commercial website and we did not verify it against the paper. We report it because it is a checkable claim about a widely circulated image and because, if accurate, it means the version most people know is not a research output at all.

Two observations, ours.

The actual paper's figures are described in the critical literature as a scissors chart: perceived ability plotted against actual ability by quartile, with the two lines crossing[6]. That is a different object from the peak-and-valley curve.

And the difference matters. The scissors chart is what the data looked like. The peak-and-valley curve is a story about a journey, with a trajectory over time that the cross-sectional data could not have shown.

Twenty Years Of Critique

The lineage, as recorded in reference lists and in a peer-reviewed survey.

Krueger and Mueller (2002) raised regression to the mean[6].

Burson, Larrick and Klayman (2006), in JPSP, 90(1), 60–77, Skilled or unskilled, but still unaware of it: how perceptions of difficulty drive miscalibration in relative comparisons[7].

Krajc and Ortmann (2008), in the Journal of Economic Psychology, 29(5), 724–738, Are the unskilled really that unaware? An alternative explanation. A peer-reviewed survey records that they assumed a nonsymmetric J-distribution for the talent of the undergraduates studied, which leads to more students in the left tail of the ability distribution, resulting in the DK effect[3].

Nuhfer and colleagues (2016 and 2017), in Numeracy, the second titled How Random Noise and a Graphical Convention Subverted Behavioral Scientists' Explanations of Self-assessment Data[2]. The survey records that they used simulated random variables with correlation less than 1 and various graphical representations of the data to illustrate the DK effect[3].

Gignac and Zajenkowski (2020), in Intelligence, 80, 101449, who the survey records tested the validity of the DK effect with the Glejser test of heteroskedasticity and by nonlinear quadratic regression[3][1].

Two observations, ours.

These are independent teams in different fields: psychology, economic psychology, numeracy education and intelligence research, arriving at a similar objection by different routes.

And note the range of methods: simulation, distributional assumptions, heteroskedasticity testing and nonlinear regression. That is a stronger position than one team applying one technique.

Our Own Simulation

The demonstration. Everything in this section and the next two is ours, uses parameters we invented, and comes from no paper.

We generated 40,000 simulated people. Each has a true underlying ability drawn from a normal distribution.

Each takes a test, and the test score is their true ability plus measurement noise, because no test measures ability perfectly.

Each makes a self-assessment, which is their true ability plus a different draw of noise, plus a constant modest self-flattery applied identically to everyone.

The critical features are what we deliberately left out.

No metacognitive deficit. Every simulated person estimates their own ability with exactly the same accuracy as every other.

No differential bias. The self-flattery constant is identical for the least able and the most able.

No double curse. Nothing in the model makes low ability impair self-insight, because the whole point is to see what appears without it.

We then converted both scores to percentiles and grouped people into quartiles by test score, which is what the original analysis did.

What It Produced

The output. Our figures.

Bottom quartile: actual percentile 12.5, perceived percentile 25.7. Overestimate of 13.2 points.

Second quartile: actual 37.5, perceived 42.9. Overestimate of 5.4.

Third quartile: actual 62.5, perceived 56.8. Underestimate of 5.7.

Top quartile: actual 87.5, perceived 74.5. Underestimate of 13.0.

Three observations, ours.

That is the scissors chart. Two lines converging and crossing, with the bottom group dramatically overconfident and the top group modestly self-deprecating.

And it came out of a model that contains no psychology whatsoever. There is no metacognition in the simulation. There are no people in it. It is two correlated random variables and a sorting rule.

Which means the chart, standing alone, cannot distinguish between a world where the double curse exists and a world where it does not. Both produce the same picture.

Why That Happens

The mechanism, explained plainly. Ours.

The moment you sort people by a noisy measurement, you have selected on the noise as well as on the thing you wanted to measure.

Three consequences.

Someone in the bottom quartile by test score is there partly because they are genuinely less able and partly because they had a bad day. Their self-assessment does not know about the bad day, so it sits closer to their true ability, which is higher than their measured score.

Someone in the top quartile is there partly by ability and partly by luck. Their self-assessment again reflects their true ability, which is lower than their measured score.

So the bottom appears overconfident and the top appears underconfident, automatically, with no difference in anyone's self-insight. This is regression to the mean, which the twelfth article in this series encountered doing the same work in a completely different literature.

The Two Ingredients

Separating the two components, because they are different claims. Ours.

Our simulation contained two things, and they contribute differently.

Measurement noise produces the crossing shape: the fanning in of perceived ability toward the middle. This is the artefact.

The better-than-average constant shifts the whole perceived line upward, which is why the bottom quartile's overestimate is larger than the top quartile's underestimate in most real data.

Two observations.

The better-than-average effect is a real and separate finding in psychology, and nothing here disputes it. But it is uniform across the distribution, and a uniform bias is not the double curse.

A commercial source makes the same point with a chart caption describing simulated data with a better than average effect, not Dunning-Kruger[8], which is precisely the distinction: everybody flattering themselves equally looks like this, and it is not a claim about the unskilled specifically.

The Finding That Inverts The Story

A result that goes further than the artefact argument, reported with our sourcing stated.

Third-party text describing a study records that metacognitive sensitivity tracked performance closely, so less information was exploited by the metacognitive judgements of poor performers, but that metacognitive efficiency, being the quality of metacognitive processing itself, was unrelated to performance, and crucially that metacognitive bias was positively associated with performance, so poor performers were appropriately less confident, not more confident, than good performers. It concludes that these metacognitive factors did not cause the DKE pattern, which was driven overwhelmingly by performance scores, and that the results refute the dual-burden account[6].

We did not obtain this study, do not know its authors or sample, and report only this indexed description.

Two observations, ours.

If accurate, poor performers were less confident in absolute terms, which is the opposite of the popular story. The apparent overconfidence is entirely relative, arising from comparing their confidence to their measured score rather than to other people's confidence.

And that is a distinction worth carrying: overconfident relative to your score and confident in absolute terms are different states, and only the second matches how the effect is invoked in conversation.

The Other Side

Being fair, because our sources skew against the effect.

A peer-reviewed survey records that McIntosh and colleagues (2019) experimented with movement and memory tasks, and concluded that the DK effect exists as an empirical phenomenon. But they disagreed with the explanation that poor insight is the reason for overestimation among the unskilled[3].

The original authors also responded. The same survey's reference list records Ehrlinger, Johnson, Banner, Dunning and Kruger (2008), Why the unskilled are unaware: further explorations of (absent) self-insight among the incompetent, in Organizational Behavior and Human Decision Processes, 105, 98–121[3].

We obtained neither paper and report titles and characterisations only.

Two observations, ours.

The McIntosh position is the important one, and it is a third position: the pattern is real, and the explanation is wrong. That is different from both "it is a genuine metacognitive deficit" and "it is pure artefact".

And it is roughly where we would put ourselves, on the evidence we obtained.

What Survives

The honest residue.

A source, itself broadly critical, concedes: "A calibrated core still survives. Even the strongest critiques concede that self-assessment accuracy correlates with actual skill: the magnitude is smaller than the meme, but the signal is real."[6]

Three propositions we think survive, and this assessment is ours.

People are imperfect judges of their own ability. Uncontroversial and not in dispute.

Self-assessment accuracy correlates with actual skill, which is the calibrated core above.

And a better-than-average tendency exists, applying broadly rather than specifically to the unskilled.

What does not survive, on our reading, is the specific and interesting claim: that there is a distinct metacognitive deficit concentrated at the bottom of the distribution, demonstrated by that chart. The chart cannot demonstrate it, because noise alone produces the chart.

The Citation Gap

A number that captures the whole problem, with its sourcing flagged.

A blog reports that, as of 2022, Google Scholar showed the three critique papers with 88 citations collectively, against 7,893 for the original[9].

We did not verify these figures, they are several years old, and citation counts are a crude measure. On the numbers as given, that is a ratio of roughly 90 to 1, which is our arithmetic.

Two observations, ours.

Even discounting heavily for the source, an order-of-magnitude gap between a finding and its critiques is the pattern this series has documented repeatedly. The nineteenth article found a refuting meta-analysis cited as authority for the refuted claim.

And the same blog notes it took until 2016 for the objection to be widely made, seventeen years after publication. That is not a scandal. It is how long a simple statistical objection can go unmade when a finding is popular, memorable and flattering to the people repeating it.

What This Means For Using It

The practical translation. Ours.

Three things follow.

The phrase should leave professional conversation. Not because incompetent overconfident people do not exist, but because the research does not license the diagnosis, and the phrase functions as a way of dismissing someone with a scientific-sounding label.

It is almost always applied to other people, which is the same structure the fifteenth article in this series found in the ordinary use of groupthink: an unfalsifiable label applied retrospectively to someone you disagree with.

And it contains its own escape clause. Anyone objecting can be told that not recognising one's own incompetence is precisely the predicted symptom, which makes the claim unanswerable and therefore empty.

The Manager Who Says It

A specific case. Ours, and offered as reasoning.

A manager describes a junior colleague as a Dunning-Kruger case: confident, and worse than they think.

Three observations.

The observation may well be correct. Some people are confident and not good, and nothing here disputes that.

But the label adds an explanation the evidence does not support: that the confidence is caused by the incompetence, and therefore that feedback will not work because the deficit prevents its reception.

And that explanation has a practical consequence, which is why it matters. If the confidence is a symptom of the incompetence, there is no point telling them. If it is not, telling them clearly is the obvious first move, and the seventh article in this series has something to say about how.

The General Lesson

The transferable part, and it is not about this effect. Ours.

The rule the simulation demonstrates is this: whenever you sort people or things by a noisy measurement and then compare groups, some of what you see is the sorting.

Three business examples where the same trap operates, all ours and untested.

Your worst month. Next month will probably be better, whatever you do about it, and whatever you did will get the credit.

Your worst-performing salesperson. They are partly there by bad luck, so their next quarter will improve on average, with or without the intervention.

And the client with the highest error rate in their records, who will regress toward normal next year regardless of the process you introduced.

Two consequences. Any intervention aimed at the extremes will appear to work. And the only way to know whether it did is a comparison group that did not receive it, which is exactly what the twelfth article in this series argued rescued the social norms finding.

What To Do

Stop using the phrase as a diagnosis. The chart it rests on can be produced by noise, so it cannot establish what it is used to claim about a person.

Keep the underlying observation. People are imperfect judges of their own ability and self-assessment accuracy does correlate with skill. That much survives.

Distinguish relative from absolute overconfidence. One description of a study reports poor performers being appropriately less confident, with the apparent overconfidence arising only from comparison with their score.

Note that the popular chart is not from the paper. On one commercial source, the peak-and-valley curve was retrofitted in mid-2000s management material.

Do not conclude feedback is pointless. The label implies the deficit blocks its own detection, and that implication is the part with the least support and the most practical cost.

Watch for the same trap elsewhere. Sorting on a noisy measure and comparing groups builds regression to the mean into the result, whatever the subject.

Use a comparison group whenever you intervene at the extremes. Your worst month, worst performer and worst client will all improve on their own.

And notice the escape clause. Any claim whose predicted symptom is disagreeing with it cannot be tested, which is a reason to distrust it independent of any evidence.

The Limits Of This Analysis

Several caveats matter. This article reviews research methods and is not management or human resources advice; nothing in it should be used to characterise any individual. Everything is verified to August 2026. We did not obtain any of the papers discussed in full, including the original, and report no effect sizes and no sample sizes from any of them; we have titles, citations, and characterisations from reference lists and from a peer-reviewed survey. Our sources are heavily weighted toward the critical side, several are blogs and commercial websites, and we flag this as a selection problem in our own research rather than a feature of the literature. The simulation is entirely ours, uses invented parameters, comes from no paper, and demonstrates that the pattern can arise without a metacognitive deficit; it does not establish that the observed pattern in real data does so arise. The claim that the Mount Stupid curve is not from the paper comes from a commercial website and we did not verify it. The study reporting that poor performers were appropriately less confident reaches us as indexed third-party text; we do not know its authors or sample. The citation figures come from a blog, are several years old, and we did not verify them. We did not obtain the original authors' 2008 response or the 2019 study concluding the effect is empirically real, and report both only as a survey describes them. The section on business applications of regression to the mean is our own reasoning, untested.

Frequently Asked Questions

Is the Dunning-Kruger effect real?
The pattern is observed. Whether it demonstrates a metacognitive deficit concentrated among the unskilled is disputed, with a 2020 paper in Intelligence titled to say it is mostly a statistical artefact. Our own simulation reproduces the classic chart from a model containing no such deficit.
How can noise produce that chart?
Sorting people by a noisy test score selects partly on luck. The bottom group is there partly through a bad day, so their self-assessment sits above their measured score. The top group is there partly through a good day, so theirs sits below. That produces the crossing lines with nobody differing in self-insight.
Does that mean nobody is overconfident?
No. A better-than-average tendency is a real and separate finding, and it is in our simulation. But it applies uniformly across the distribution, and a uniform bias is not the double-curse claim, which is specifically about the unskilled.
Is there evidence on the other side?
Yes. A peer-reviewed survey records a 2019 study concluding the effect exists as an empirical phenomenon while disagreeing that poor insight explains it, and the original authors published a follow-up in 2008. We obtained neither in full, and our sources skew critical, which we flag.
What about the Mount Stupid curve?
One commercial source states it does not appear in the 1999 paper or any follow-up work and was retrofitted in mid-2000s management material. We could not verify that, but the actual figures are described as a scissors chart of perceived against actual ability, which is a different object from a journey over time.
What should I take away?
A general rule rather than a fact about people: whenever you sort by a noisy measure and compare groups, some of what you see is the sorting. Your worst month, worst salesperson and worst client will all improve on their own, and any intervention aimed at them will appear to work.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article opens by warning that its own sources are one-sided, and its central demonstration was run rather than cited, so that one part of it does not depend on those sources at all.

References

  1. Gignac, G. E., & Zajenkowski, M. (2020). The Dunning-Kruger effect is (mostly) a statistical artefact: Valid approaches to testing the hypothesis with individual differences data. Intelligence, 80, 101449. DOI 10.1016/j.intell.2020.101449, bibliographic record confirming the citation and reference list, which includes Krajc, M., & Ortmann, A. (2008), Are the unskilled really that unaware? An alternative explanation, Journal of Economic Psychology, 29(5), 724–738. Note: a bibliographic record; access to the full text is restricted and we did not obtain it. We report its title and its citation, not its data. ideas.repec.org
  2. Reference lists identifying Nuhfer, E., Cogan, C., Fleisher, S., Gaze, E., & Wirth, K. (2016), Random number simulations reveal how random noise affects the measurements and graphical portrayals of self-assessed competency, Numeracy: Advancing Education in Quantitative Literacy, 9(1); and Nuhfer, E., Fleisher, S., Cogan, C., Wirth, K., & Gaze, E. (2017), How Random Noise and a Graphical Convention Subverted Behavioral Scientists' Explanations of Self-assessment Data: Numeracy Underlies Better Alternatives, Numeracy, 10(1), 1–31. Note: citations only. We obtained neither paper and report no findings from either beyond what the peer-reviewed survey at reference 3 states. economicsfromthetopdown.com
  3. Frontiers in Psychology (2022), A Statistical Explanation of the Dunning-Kruger Effect, DOI 10.3389/fpsyg.2022.840180, on Gignac and Zajenkowski (2020) having used a sample of general community participants and tested the validity of the effect with the Glejser test of heteroskedasticity and by nonlinear quadratic regression; on Nuhfer and colleagues (2016) having used simulated random variables with correlation less than 1 and various graphical representations to illustrate the effect; on Krajc and Ortmann (2008) having assumed a nonsymmetric J-distribution for the talent of the undergraduates studied by Kruger and Dunning, leading to more students in the left tail of the ability distribution and resulting in the effect; on McIntosh and colleagues (2019) having experimented with movement and memory tasks and concluded that the effect exists as an empirical phenomenon while disagreeing with the explanation that poor insight is the reason for overestimation among the unskilled; and its reference list recording Ehrlinger, J., Johnson, K., Banner, M., Dunning, D., & Kruger, J. (2008), Why the unskilled are unaware: further explorations of (absent) self-insight among the incompetent, Organizational Behavior and Human Decision Processes, 105, 98–121. Note: a peer-reviewed open-access survey; our principal balanced source. We obtained portions. frontiersin.org
  4. Reference list confirming Kruger, J., & Dunning, D. (1999), Unskilled and unaware of it: How difficulties in recognizing one's own incompetence lead to inflated self-assessments, Journal of Personality and Social Psychology, 77(6), 1121–1134, PMID 10626367; and Burson, K. A., Larrick, R. P., & Klayman, J. (2006), Skilled or unskilled, but still unaware of it: how perceptions of difficulty drive miscalibration in relative comparisons, Journal of Personality and Social Psychology, 90(1), 60–77, PMID 16448310. Note: a reference list on a psychology commentary site. We did not obtain either paper and report citations only. psychologytoday.com
  5. Popular science reference site, summarising the effect as a metacognitive bias first described in 1999 under which low performers tend to overestimate their ability, often rating themselves above average, while top performers tend to underestimate how much better they are than the rest; and noting that recent research suggests parts of the pattern are also driven by simple statistics such as regression to the mean. Note: a popular science website, not peer-reviewed; used only for a summary of the popular understanding. scienceabc.com
  6. Commercial behavioural design website, on the viral Mount Stupid curve being an invention that does not appear in the 1999 paper or any of Dunning's follow-up work and was retrofitted onto the data in mid-2000s management memes; on Nuhfer and colleagues and Gignac and Zajenkowski having shown the iconic scissors chart can emerge from random noise plus regression to the mean; on a calibrated core still surviving, with even the strongest critiques conceding that self-assessment accuracy correlates with actual skill though the magnitude is smaller than the meme; and, in third-party indexed text, on a study reporting that metacognitive sensitivity tracked performance closely, that metacognitive efficiency was unrelated to performance, that metacognitive bias was positively associated with performance so that poor performers were appropriately less confident rather than more confident than good performers, and that these metacognitive factors did not cause the pattern, which was driven overwhelmingly by performance scores, refuting the dual-burden account. Note: a commercial website, not peer-reviewed. The Mount Stupid claim was not verified. The metacognition study reaches us as indexed third-party text and we do not know its authors or sample. yukaichou.com
  7. Commentary site summarising the convergent critical literature, identifying Krueger and Mueller (2002), Burson and colleagues (2006), Nuhfer and colleagues (2016 and 2017), Gignac and Zajenkowski (2020) and Magnus and Peresetsky (2022) as reaching the conclusion that the effect, in the specific form of the lowest-skill performers dramatically overestimating themselves while the highest-skill performers slightly underestimate themselves, is mostly a statistical artifact arising from imperfect measurement plus regression to the mean. Note: a commentary website, not peer-reviewed. We obtained none of the papers it lists and report its characterisation as such. atticusli.com
  8. Psychology commentary site, carrying a figure captioned as simulated data with a better than average effect rather than a Dunning-Kruger effect, illustrating that a uniform self-flattery applied across the distribution produces a similar graphical pattern. Note: a commentary site; used only for the conceptual distinction between a uniform better-than-average bias and a deficit specific to the unskilled. psychologytoday.com
  9. Economics blog, reporting that according to Google Scholar the three critique papers, being Nuhfer 2016 and 2017 and Gignac and Zajenkowski 2020, had 88 citations collectively, against 7,893 for Kruger and Dunning (1999); and noting that although the original was published in 1999, the objection was not fully understood until 2016. Note: a blog, not peer-reviewed. We did not verify these citation counts, they date from 2022, and citation counts are a crude measure. The ratio of roughly 90 to 1 is our own arithmetic on the figures as given. economicsfromthetopdown.com

This article reviews research methods and is not management or human resources advice. Nothing in it should be used to characterise any individual. No paper discussed was obtained in full and no effect sizes or sample sizes are reported. The sources located skew toward the critical side, several are blogs or commercial websites, and this is flagged as a limitation of the authors' own research. The simulation is the authors' own, uses invented parameters, and demonstrates that the pattern can arise without a metacognitive deficit rather than that it does so in real data.