This publication is produced by a firm that bills for expertise, employs people at varying stages of acquiring it, and spends money on developing it. That makes this article awkward in a useful way, which is the same reason the thirty-seventh article on advice discounting was worth writing.
Key Takeaway
From the abstract: "We found that deliberate practice explained 26% of the variance in performance for games, 21% for music, 18% for sports, 4% for education, and less than 1% for professions. We conclude that deliberate practice is important, but not as important as has been argued."[1] A follow-up on sports found the effect "accounted for only 1% of the variance in performance among elite-level performers"[2]. Our own translation: at one percent of variance, knowing which of two professionals practised more identifies the better performer 53.2 times in 100, against 50 for a coin.
The Verdict, Stated First
Five claims, in descending order of confidence.
One. Practice matters, and matters much less than claimed, and the size depends enormously on the domain. Twenty-six percent of variance in games is a substantial effect. Under one percent in professions is not. Both are from the same meta-analysis.
Two. The professions figure is the one that concerns a business reader and the one nobody quotes. The four numbers usually repeated are for games, music, sports and education. The fifth is for the domain most people work in.
Three. Variance explained is a bad unit for a decision, and translating it changes the impression. Our own arithmetic converts the published figures into how often practice correctly orders two people: 67 percent for games, 53.2 percent for professions. We verified the translation two ways and they agree to a tenth of a point.
Four. The effect collapses where the theory says it should be strongest. In sports the figure is 18 percent overall and 1 percent among elite performers, which is the opposite of what a theory of expert performance predicts.
Five. This is not a claim that training is worthless, and we will not let it be read that way. The meta-analysis measures how much differences between people are explained by differences in practice hours. That is a different question from whether a given person improves by practising, and the second question is not answered here.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the five headline figures. Obtained verbatim from the publishing society, the journal publisher, and two further independent reproductions.
Grade A for the elite-performer collapse in sports, obtained verbatim from the publisher's abstract.
Grade B for the interpretation of the professions figure, because we did not obtain the underlying studies, the number of them, or what counted as a profession.
Grade A for our own translation arithmetic, which is standard and which we verified by simulation.
Ungraded for the merits of the underlying dispute, since we did not obtain the 1993 paper being tested, nor any reply from its authors.
Our position: the meta-analytic numbers are solid and their translation into decisions is where nearly all the damage occurs. A figure of 26 percent has been read as vindication and a figure of 1 percent has been read as refutation, and neither reading survives converting them into something a manager could act on.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2014 meta-analysis's abstract verbatim from four independent sources, including the publishing society's own journal page[1][3][4]. We did not obtain the paper, its included studies, its sample sizes, its confidence intervals or its moderator analyses.
We obtained the 2016 sports meta-analysis's abstract verbatim[2] and not the paper.
We obtained a bibliographic record of a corrigendum to the 2014 paper[5] and could not obtain what it corrected, which is a real gap and we say so where it matters.
We did not obtain the 1993 paper that proposed the view being tested, nor any published reply by its authors to the meta-analysis. This article therefore reports one side of a dispute in detail and the other only as characterised by third parties, and a reader should weigh it accordingly.
All arithmetic is ours. The variance figures are the published ones; the translation into ordering probabilities is our own and assumes bivariate normality, which no source establishes.
This article discusses research on expertise. It is not training, professional development or human resources advice.
The Claim Being Tested
The proposition, as the meta-analysis states it.
"More than 20 years ago, researchers proposed that individual differences in performance in such domains as music, sports, and games largely reflect individual differences in amount of deliberate practice."[1]
And the reason they thought it worth testing: "This view is a frequent topic of popular-science writing, but is it supported by empirical evidence?"[1]
Four observations, ours.
The load-bearing word is "largely." The claim is not that practice matters, which nobody disputes. It is that differences between people mostly come from differences in practice, and that is a quantitative claim with a testable magnitude.
The second word to notice is "individual differences." This is a claim about variation across people, not about whether practice improves a given person, and the two are constantly conflated in the popular version.
The sentence about popular-science writing is unusually direct for an abstract, and it tells you the authors regarded the gap between the claim's circulation and its evidence as the motivation for the work.
And note the domains named in the original proposal: music, sports, and games. Professions are not in that list, which becomes important below.
What Counts As Deliberate Practice
The definition, which does a great deal of work.
The meta-analysis gives it as "engagement in structured activities created specifically to improve performance in a domain."[1]
Three observations, ours.
Two conditions are stated: activities must be structured and created specifically to improve performance. Ordinary work does not qualify. Doing your job is not deliberate practice under this definition, however many hours it consumes.
That has a direct consequence for the professions figure. A professional accumulating years of client work is accumulating experience, not deliberate practice, and the two are being measured separately even when the underlying studies use hours as a proxy.
And the definition is precisely where a separate critique lands. A paper by an overlapping group is described as documenting "critical inconsistencies in the definition of deliberate practice" along with "apparent shifts in the standard for evidence"[6], which we return to.
The Meta-Analysis
The study.
Macnamara, B. N., Hambrick, D. Z., and Oswald, F. L. (2014), Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis, Psychological Science, 25(8), 1608–1618, DOI 10.1177/0956797614535810, published online 1 July 2014[1][5].
Its scope: "we conducted a meta-analysis covering all major domains in which deliberate practice has been investigated."[1]
Three observations, ours.
All major domains is an ambitious scope claim and it is what makes the professions figure meaningful rather than incidental. Professions were included because the review aimed to be comprehensive.
The journal is the flagship empirical outlet of a major psychological society, which is worth noting given how often this series reports findings from smaller venues.
And we did not obtain the paper. We have the abstract from four independent reproductions and nothing about how many studies fed each domain estimate, which is a limitation we return to when grading the professions figure.
The Five Numbers
The result, quoted exactly.
"We found that deliberate practice explained 26% of the variance in performance for games, 21% for music, 18% for sports, 4% for education, and less than 1% for professions."[1]
And the conclusion the authors drew: "We conclude that deliberate practice is important, but not as important as has been argued."[1]
A description of the paper's own figure confirms the metric: "Percentage of variance in performance explained... and not explained... by deliberate practice within each domain studied. Percentage of variance explained is equal to r-squared times 100."[4]
Four observations, ours.
The spread across domains is a factor of more than twenty-six, from 26 percent to under 1 percent. Any summary that reports a single overall figure has discarded the most informative feature of the result.
The conclusion sentence is notably restrained. It says important but not as important as argued, which is neither the vindication nor the demolition the finding has been used for.
The confirmation that the metric is r-squared times 100 matters more than it appears, and the next two sections are about why.
And the ordering of the domains is worth sitting with. The effect is largest in games and smallest in professions, which is an ordering by how closed, rule-bound and measurable the domain is.
The One That Concerns Us
The figure this publication's readers should care about, and the one that travels least. Ours.
Four observations.
Less than one percent. Of all the variation in professional performance across people, differences in deliberate practice account for under a hundredth of it, on this estimate.
Note that the original proposal named music, sports, and games, and professions were added by the reviewers. So the theory is being tested outside the domains it was built for, which is a legitimate extension and also a reason for care.
We did not obtain the underlying studies, so we cannot tell you how many there were, what professions they covered, how performance was measured, or how wide the confidence interval was. We grade the interpretation B for that reason and would not build a decision on this figure alone.
And there is an obvious candidate explanation the abstract does not address and we cannot resolve. Professional performance is harder to measure than a chess rating, and unreliable outcome measurement attenuates any correlation. A small figure could reflect a small effect or a noisy dependent variable, and nothing we obtained distinguishes them.
Variance Explained, Translated
Because percentage of variance is close to useless as a decision input, and almost nobody converts it. Our own arithmetic, standard, and the variance figures are the published ones.
Variance explained is r-squared. The correlation is its square root, and the correlation is what does the predicting.
Games: 26 percent becomes r = 0.51. Music: 21 percent becomes r = 0.46. Sports: 18 percent becomes r = 0.42. Education: 4 percent becomes r = 0.20. Professions: 1 percent becomes r = 0.10.
Three observations.
The square root compresses the apparent difference between domains. A twenty-six-fold gap in variance explained is a five-fold gap in correlation, and readers who know only one of the two units will disagree about the same data.
Both units are correct, which is the difficulty. Reporting r-squared makes an effect sound small and reporting r makes it sound larger, and a writer choosing between them is choosing an impression.
And neither answers the question a manager has, which is what the number lets you do. That requires a third translation, and it is the next section.
Fifty-Three Times In A Hundred
The translation we think should always accompany a correlation. Our own arithmetic, computed two independent ways.
The question: take two people at random, pick the one who practised more, and how often is that the better performer?
For jointly normal variables this is one half plus the arcsine of r divided by pi. We also ran a simulation of two million paired draws per domain. The two methods agree to within a tenth of a percentage point in every case, which is the check that made us confident enough to publish it.
Games: 67.0 percent. Music: 65.2 percent. Sports: 63.9 percent. Education: 56.4 percent. Professions: 53.2 percent. Elite sport: 53.2 percent.
Four observations.
Read the professions row. Knowing which of two professionals practised more identifies the better one about 53 times in 100, against 50 for a coin toss. That is the entire predictive content of the relationship in that domain.
Read the games row and be fair to the finding. 67 percent is a genuinely useful signal, well above chance, and the same meta-analysis produced it. This is not a paper reporting that practice does not matter.
The translation is more honest than either variance or correlation, because it states a decision you could actually make and how often it would be right.
And we would apply this generally. Any correlation reported to a business audience should carry this conversion, because the number that sounds like something and the number that does something are frequently very far apart.
What "Largely" Would Require
Testing the original claim against the arithmetic. Ours.
The proposal was that individual differences largely reflect differences in deliberate practice. So what would "largely" have to look like?
If practice explained 50 percent of variance: r = 0.71, correct ordering 75 percent of the time.
At 70 percent: r = 0.84, correct ordering 82 percent.
At 90 percent: r = 0.95, correct ordering 90 percent.
Three observations.
The observed figures for professions and for elite sport are 1 percent. On any reading of "largely" that a reader would accept, the gap between the claim and the finding is enormous, and that gap is the whole dispute.
But the games figure of 26 percent is not nothing, and it is not "largely" either. The honest summary is that the strong claim fails everywhere and the weak claim holds in some domains, which is what the authors' own restrained conclusion says.
And this is the value of forcing a vague word into a number. "Largely" cannot be argued about productively. Seventy-five percent correct ordering can.
The Sports Follow-Up
A second meta-analysis, narrower and in some ways more damaging.
Macnamara, B. N., Moreau, D., and Hambrick, D. Z. (2016), The Relationship Between Deliberate Practice and Performance in Sports: A Meta-Analysis, Perspectives on Psychological Science[2].
Its framing: "Why are some people more skilled in complex domains than other people? According to one prominent view, individual differences in performance largely reflect individual differences in accumulated amount of deliberate practice."[2]
Its headline: "Overall, deliberate practice accounted for 18% of the variance in sports performance."[2]
Two observations, ours.
The 18 percent matches the 2014 paper's sports figure exactly, which is reassuring about consistency across the two reviews though they share authors and likely studies.
And the paper's contribution is not the headline. It is the breakdown by skill level, which the earlier paper did not report and which is the next section.
One Percent Among Elites
The finding that runs directly against the theory's central prediction.
"However, the contribution differed depending on skill level. Most important, deliberate practice accounted for only 1% of the variance in performance among elite-level performers. This finding is inconsistent with the claim that deliberate practice accounts for performance differences even among elite performers."[2]
Four observations, ours.
The effect collapses from 18 percent to 1 percent exactly where the theory of expert performance says it should be doing the most work.
The authors flag it themselves with "Most important", which is unusual emphasis inside an abstract and signals they regarded this as the paper's contribution.
There is a technical reason to expect some attenuation here and it deserves stating even though no source we obtained raises it. Elite performers are a restricted range on both variables, and restriction of range mechanically reduces correlations. That is our own observation, it does not appear in anything we obtained, and it is a reason for caution rather than dismissal.
But the finding still bites, because the theory's claim is specifically about what separates the very best. If the effect is only visible when you include beginners, it is explaining the gap between novices and experts rather than differences among experts, and those are different claims.
And They Did Not Start Earlier
A second finding in the same abstract, easy to miss and quite specific.
"Another major finding was that athletes who reached a high level of skill did not begin their sport earlier in childhood than lower skill athletes."[2]
Three observations, ours.
Starting age is the most direct behavioural implication of an accumulation theory. If total hours drive expertise, an earlier start is a straightforward advantage.
The finding is stated as a null, and we report it at that strength. Higher-skill athletes did not start earlier; the abstract does not claim they started later.
And the commercial analogue is direct. If accumulated hours drove professional expertise, entering the profession younger would be an advantage, and this finding is at least a reason to doubt the parallel assumption.
A Corrigendum, Four Years Later
A detail we would rather report than omit, and one we cannot resolve.
A national medical literature database records a Corrigendum: Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis, indexed separately from the original and dated 2018, four years after the paper[5].
We could not obtain the corrigendum's contents. We know it exists, we know when, and we do not know what it corrected.
Four observations, ours.
This is a real gap in this article and we are placing it prominently rather than in the limits section, because every figure quoted above comes from the paper it corrects.
Corrigenda range from trivial to substantive, and we cannot tell you which this is. A reader who intends to rely on these numbers should obtain it.
The five headline figures appear identically in every reproduction we found, including sources published after 2018, which is weak evidence that the headline results were not what changed.
And this is the correct behaviour of a scientific record. A correction issued four years later, indexed alongside the original, is the system working, and the fiftieth article in this series found that most corrections it encountered were already published and simply had not travelled.
The Definitional Critique
A separate line of attack, on the theory's testability rather than its magnitude.
A paper by an overlapping author group is summarised as documenting "critical inconsistencies in the definition of deliberate practice" along with "apparent shifts in the standard for evidence concerning deliberate practice," and considering the impact of these on the field "focusing on the empirical testability and falsifiability of the deliberate practice view."[6]
We did not obtain that paper and report a bibliographic service's summary of it.
Three observations, ours.
Falsifiability is a much stronger charge than a small effect size. It says the theory can be adjusted after the fact so that no result would count against it.
This is the same structure the fifty-third article found in the dual-process exchange, where defenders replied that the disputed features were never definitional. A theory whose definition moves is a theory whose refutations can always be deflected, and both disputes turn on that point rather than on data.
And we grade this ungraded, because a summary of a critique is not evidence and because the authors overlap with the meta-analysis, so it is not independent.
What The Other Side Says
The part of this article we are least able to write, and we would rather say so than paper over it. Ours.
Three points.
We did not obtain the 1993 paper proposing the deliberate practice view, and we did not obtain any published reply from its authors to the 2014 meta-analysis. We know from citing sources that replies exist.
The obvious counter-argument is visible even from what we have, and it concerns the definition. If the meta-analysis counted hours of practice that were not structured and not designed to improve performance, it measured something broader than the theory claims. A defender would say the reviews aggregated the wrong quantity, and nothing we obtained settles that.
And a related point about measurement runs the same way. Most included studies almost certainly used retrospective self-reported hours, which the forty-third article in this series would predict are reconstructed toward known outcomes. That is our own inference, not a claim from any source, and it would bias the estimate in an unclear direction.
A Correlation Is Not A Threshold
The number everyone actually knows is ten thousand hours, and it makes a different kind of claim from anything the meta-analyses measure. Our own simulation, using the published music figure and invented parameters for everything else.
A citing source states the origin: the 1993 authors "theorized that the acquisition of musical expertise unfolds over approximately 10,000 h of deliberate practice distributed across years of training."[7]
That is a threshold claim. Cross the line and expertise follows. What the meta-analyses report is a correlation, and a correlation of 0.46 does not produce a line.
We simulated music at the published 21 percent of variance, gave hours a mean of 8,000 and a spread of 3,000, defined experts as the top ten percent of performers, and looked at their hours.
Among those experts, the 10th percentile logged 6,891 hours. The 25th: 8,561. The median: 10,395. The 75th: 12,252. The 90th: 13,869.
And 18.9 percent of the experts had below-average hours.
Four observations, and one is a correction to ourselves.
Our first description of that last figure said roughly a third. It is 18.9 percent, closer to a fifth, and we caught it by reading the output rather than our sentence about it. That is the fourth time in this series that method has caught an error and it remains the only one that reliably does.
The median expert sits near 10,400 hours, which is close to the famous figure and is the reason the threshold claim looks supported. The median of a wide distribution is not a threshold.
The interquartile range among experts alone spans 8,561 to 12,252, and the tenth to ninetieth spans 6,891 to 13,869, twice as wide. Experts do not cluster above a line; they are spread across many thousands of hours.
And nearly one in five reached the top decile on below-average practice. A threshold claim has no room for that group, and a correlation of 0.46 predicts it. The two framings are not versions of the same finding.
We Tested Our Own Objection
Earlier we raised a defence of the deliberate practice view that no source we obtained makes: that professional performance is measured badly, and that unreliable measurement depresses any correlation. That objection is quantifiable, so we quantified it. Our own arithmetic; all reliabilities are invented and no source states any.
The standard correction divides the observed correlation by the square root of the product of the two measures' reliabilities. So the question is: what true correlation would produce an observed 0.10?
With both measures at a reliability of 0.90: true r of 0.111, or 1.2 percent of variance.
At 0.80 and 0.70: true r 0.134, 1.8 percent.
At 0.70 and 0.60: 0.154, 2.4 percent.
At 0.60 and 0.50: 0.183, 3.3 percent.
At 0.50 and 0.30, which would be poor measurement by any standard: 0.258, 6.7 percent.
Three observations.
The correction moves the professions figure from about 1 percent to somewhere between 2 and 7 percent across a wide range of pessimistic assumptions. That is a real difference and it is not a rescue.
Even at 6.7 percent, the implied ordering accuracy is about 58 in 100 rather than 53, which is better and is nowhere near what the strong claim requires.
And the objection we raised does not survive being taken seriously, which is the point of taking it seriously.
How Bad The Measurement Would Have To Be
Running the same correction backwards, which is the sharper version. Ours.
Instead of asking what a plausible reliability implies, ask what reliability the strong claim would require.
For the true correlation to reach 0.30, being 9 percent of variance, both measures would need reliabilities of about 0.33.
For 0.50, being 25 percent: about 0.20.
For 0.71, being 50 percent and the lowest figure that could be called "largely": both measures at about 0.14.
Four observations.
A reliability of 0.14 means roughly 86 percent of what the instrument records is noise. No published study would use a measure like that, and no reviewer would accept one.
So the attenuation defence, pushed to the magnitude it would need, becomes a claim that the entire professions literature measured almost nothing at all. That is self-defeating, because a literature that measured nothing cannot support the original claim either.
We think this is the correct way to handle an objection you raise yourself. State it, bound it, and report that it fails at the size required, rather than leaving it as an unresolved caveat that quietly protects the conclusion you prefer.
And the residual is honest: attenuation explains part of the gap between domains and cannot explain the gap between 1 percent and "largely." Games are measured well and professions are not, so some of the spread across the five figures is measurement rather than substance, and the arithmetic above bounds how much.
And Is Still Being Asserted
A note on where the field sits, from a recent citing source.
A 2026 paper writes: "Ericsson et al. (1993) theorized that the acquisition of musical expertise unfolds over approximately 10,000 h of deliberate practice distributed across years of training, a figure whose precision has been contested in subsequent meta-analyses, though the broader finding that structured deliberate practice drives expertise acquisition remains well-supported."[7]
Three observations, ours.
That sentence concedes the number and keeps the claim. The precision is contested; the broader finding remains well-supported. It is a reasonable position and it is not obviously consistent with a professions figure under one percent.
It also demonstrates the transmission pattern this series keeps documenting. A meta-analysis is cited as having contested "precision", when what it reported was that the effect ranges from substantial to negligible depending on domain, which is a different objection.
And we report it because it is evidence the dispute is live in 2026, not settled, and because presenting only the meta-analysts' side would misrepresent the state of the field.
What Actually Survives
Our reading, stated directly.
Five statements.
Practice explains a substantial minority of performance differences in closed, rule-bound domains. Twenty-six percent for games is real and useful.
It explains almost nothing about differences between professionals, on this estimate, though the measurement problems there are severe and unresolved.
The strong version of the claim fails everywhere. No domain approaches what "largely" would require, and the largest figure would order two people correctly 67 percent of the time.
The effect vanishes among elites in the one domain where it was tested by skill level, which is where the theory most needs it, and range restriction is a partial but incomplete explanation.
And none of this measures whether an individual improves by practising, which is a different question, is not addressed by any source here, and is the question most readers will think they are asking.
Your Professional Development Budget
The application. Ours, untested, and not training or human resources advice.
Four points.
The finding does not say training is worthless, and we want to close that reading off firmly. It says differences between professionals are poorly explained by differences in their practice hours, which is compatible with training helping everyone and with the resulting gains being swamped by other sources of variation.
What it does challenge is a specific and common budgeting logic: that the way to produce better professionals is more hours of structured development. On the professions figure, hours are close to uninformative about who performs well.
The definitional point is where we would actually intervene. Almost nothing a professional firm calls development meets the definition of structured activities created specifically to improve performance. Attending an update seminar is not it; doing more client work is certainly not it.
And the honest question to put to any development spend is the one the sixth article in this series asked of selection. What outcome would tell you it worked, and are you measuring it? The literature here cannot answer that for you and does not claim to.
Hiring For Years Of Experience
The second application, and the more directly actionable one. Ours.
Four points.
Years of experience is an accumulation measure, and it is a looser one than deliberate practice, since it counts all hours rather than structured ones.
If the tighter measure explains under one percent of professional performance variance, the looser one is unlikely to do better, and that is our own inference rather than a finding from any source we obtained.
Put in the terms of our own translation: preferring the candidate with more years would identify the better performer somewhere near 53 times in 100 if the relationship matched the professions figure. That is a real edge and it is a very small one against how heavily experience is usually weighted.
And this converges with a finding this publication has already reported. The sixth article covered a substantial revision of the selection literature in which structured assessment of relevant capability outperformed proxies, and years of experience is a proxy.
Why The Domains Rank As They Do
An observation about the pattern across the five figures, which no source we obtained makes and which we think is the most useful thing to carry away. Ours.
The order is games, music, sports, education, professions, from 26 percent down to under 1. That is not a random ordering.
Three observations.
It tracks how closed and rule-bound the domain is. A game has fixed rules, a fixed objective and an unambiguous result. A profession has none of those, and what counts as good performance is itself contested.
It also tracks how well the outcome can be measured, which is the attenuation point above, and the two are hard to separate because closed domains are precisely the ones with clean scoring.
And it suggests a rule we would apply beyond this literature. The more a domain resembles a game, the more a practice-hours account of performance will explain, and professional work is at the far end of that scale from a game.
One consequence is worth stating for a reader deciding where to apply any of this. Findings established in closed domains should be expected to shrink as they move toward open ones, and this literature supplies an unusually clean example of how far the shrinkage can go: a factor of more than twenty-six, within a single review, using a single method.
What Else Matters
The question the lead author put to her own field, which we think is the right one to end on.
Quoted in her institution's release: "It is just less important than has been argued... For scientists, the important question now is, what else matters?"[8] This is a press release, not a peer-reviewed source, and we flag it as such.
Three observations, ours.
One candidate appears in the surrounding literature. A related study is described as finding that deliberate practice is "necessary but not sufficient" to explain differences in piano sight-reading skill, with working memory capacity accounting for further variance[9]. We did not obtain it.
We are deliberately not extending that into a claim about what determines professional performance, because nothing we obtained supports one and this series has criticised others for exactly that move.
And the honest position is that the field has established what does not explain most of the variation and has not established what does. A negative result of that clarity is genuinely useful and it is not an answer.
What To Do
Convert every variance figure before acting on it. Twenty-six percent of variance sounds decisive and orders two people correctly 67 times in 100. One percent orders them correctly 53 times in 100.
Notice which domain a figure comes from. The same meta-analysis reports numbers differing by a factor of more than twenty-six across domains, and a single summary figure discards its most informative feature.
Weight years of experience less than you do. On our own inference from the professions figure, it is a very weak signal relative to how heavily it is usually weighted, and structured assessment of the actual capability is the alternative.
Check whether your development activity meets the definition. Structured, and created specifically to improve performance. Most professional development is neither.
Do not read this as evidence that training fails. The measure is variation between people explained by variation in their practice, which is a different question from whether a person improves.
Ask what a claim would look like if it were false. A critique of this literature charges that its central concept shifts, and a theory whose definition moves cannot be refuted.
Get the corrigendum before relying on the numbers. A correction was issued four years after publication and we could not obtain its contents.
Treat the dispute as live. A 2026 paper still asserts the broader finding is well-supported, and we could not obtain the original authors' reply to any of this.
The Limits Of This Analysis
Several caveats matter, and this article's sourcing is one-sided in a way we want stated plainly. This article discusses research on expertise and is not training, professional development or human resources advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 2014 meta-analysis, only its abstract from four independent reproductions, so we report no sample sizes, no number of included studies per domain, no confidence intervals and no moderator analyses. We did not obtain the 2016 sports meta-analysis beyond its abstract. A corrigendum was issued to the 2014 paper in 2018 and we could not obtain its contents, which bears on every figure quoted here; we place that in the body rather than only here. We did not obtain the 1993 paper being tested, nor any published reply from its authors, so this article reports one side of a live dispute in detail and the other only through third-party characterisations, and a reader should weigh it accordingly. We did not obtain the definitional critique, the working memory study, or any primary study in this literature. Two sources are not peer-reviewed: a press release and a bibliographic service's summaries, both flagged at every use. All arithmetic is ours. The variance figures are published; the conversion to correlations and to ordering probabilities is our own, assumes bivariate normality which no source establishes, and was verified against a two-million-draw simulation agreeing to within a tenth of a percentage point. Our observations about range restriction among elite performers and about retrospective self-reported practice hours are our own and appear in nothing we obtained. And the central caution bears repeating: this literature measures how much variation between people is explained by variation in their practice, which is not the question of whether practice improves a given person, and nothing here answers the second.
Frequently Asked Questions
Does practice not matter?
What does 1 percent of variance actually mean?
Why did the effect vanish among elite athletes?
Should I stop investing in development?
What about hiring on years of experience?
Is there anything you could not check?
Is the dispute settled?
References
- Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis. Psychological Science, 25(8), 1608–1618, DOI 10.1177/0956797614535810, published online 1 July 2014. Publishing society's own journal page reproducing the abstract in full: on researchers having proposed more than 20 years ago that individual differences in performance in domains such as music, sports and games largely reflect individual differences in amount of deliberate practice, defined as engagement in structured activities created specifically to improve performance in a domain; on this view being a frequent topic of popular-science writing and the authors asking whether it is supported by empirical evidence; on the authors having conducted a meta-analysis covering all major domains in which deliberate practice has been investigated; on the finding that deliberate practice explained 26 percent of the variance in performance for games, 21 percent for music, 18 percent for sports, 4 percent for education, and less than 1 percent for professions; and on the authors concluding that deliberate practice is important but not as important as has been argued. Note: the publishing society's own page. We obtained the abstract only; we did not obtain the paper, its included studies, sample sizes, confidence intervals or moderator analyses. psychologicalscience.org
- Macnamara, B. N., Moreau, D., & Hambrick, D. Z. (2016). The Relationship Between Deliberate Practice and Performance in Sports: A Meta-Analysis. Perspectives on Psychological Science, DOI 10.1177/1745691616635591. Publisher record reproducing the abstract: on asking why some people are more skilled in complex domains than others; on one prominent view holding that individual differences in performance largely reflect individual differences in accumulated amount of deliberate practice; on the authors investigating the relationship between deliberate practice and performance in sports; on deliberate practice having accounted overall for 18 percent of the variance in sports performance; on the contribution differing depending on skill level, with deliberate practice accounting for only 1 percent of the variance in performance among elite-level performers; on this finding being inconsistent with the claim that deliberate practice accounts for performance differences even among elite performers; and on another major finding being that athletes who reached a high level of skill did not begin their sport earlier in childhood than lower skill athletes. Note: the publisher's record. We obtained the abstract only and not the paper. journals.sagepub.com
- Journal publisher's record for the 2014 meta-analysis, reproducing the abstract identically to reference 1 and recording the corresponding author's institutional email; together with its reference list identifying a companion 2014 paper by Hambrick, D. Z., Oswald, F. L., Altmann, E. M., Meinz, E. J., Gobet, F., and Campitelli, G. Note: a second independent reproduction of the abstract, used to confirm the five headline figures verbatim. journals.sagepub.com
- Bibliographic service record for the 2014 meta-analysis, carrying the caption of the paper's own figure 3, recording that it plots the percentage of variance in performance explained and not explained by deliberate practice within each domain studied, and that percentage of variance explained is equal to r-squared multiplied by 100; and carrying summaries of related works including a statement that the authors concluded deliberate practice does correlate positively with expert performance but explains only a fraction of the variance in ability, at 26 percent for games, 21 percent for music, 18 percent for sports, 4 percent for education and 1 percent for professions. Note: a bibliographic service record. Our source for confirming that the reported metric is r-squared, which the conversions in this article depend on. Note that this source renders the professions figure as 1 percent where the abstract says less than 1 percent. semanticscholar.org
- National medical literature database record for a Corrigendum: Deliberate Practice and Performance in Music, Games, Sports, Education, and Professions: A Meta-Analysis, PMID 29733772, indexed separately from and linked to the original article, which is recorded as Macnamara BN, Hambrick DZ, Oswald FL, Psychological Science, 2014 August, 25(8), 1608–18, DOI 10.1177/0956797614535810, Epub 1 July 2014, PMID 24986855. Note: a bibliographic record establishing that a corrigendum exists and was issued four years after publication. We could not obtain its contents, and this bears on every figure quoted in this article. pubmed.ncbi.nlm.nih.gov
- Bibliographic service record carrying summaries of works related to the 2014 meta-analysis, including a paper documenting critical inconsistencies in the definition of deliberate practice along with apparent shifts in the standard for evidence concerning deliberate practice, and considering the impact of these issues on progress in the field of expertise with a focus on the empirical testability and falsifiability of the deliberate practice view; and a further summary stating that evidence indicates working memory capacity is highly general, stable and heritable, such that the view that expert performance is solely a reflection of deliberate practice is called into question. Note: a bibliographic service's summaries of papers we did not obtain. The authors overlap with the meta-analysis, so these critiques are not independent of it, and we grade them ungraded. semanticscholar.org
- Repository page for the 2014 meta-analysis carrying recent scholarly citing text, reproducing the abstract in full and recording a 2026 citing statement that Ericsson and colleagues (1993) theorized that the acquisition of musical expertise unfolds over approximately 10,000 hours of deliberate practice distributed across years of training, a figure whose precision has been contested in subsequent meta-analyses citing Macnamara and colleagues (2014), though the broader finding that structured deliberate practice drives expertise acquisition remains well-supported. Note: a repository page reproducing text from a paper citing the meta-analysis. Reported as evidence that the dispute remains live in 2026; we did not obtain the citing paper. researchgate.net
- Science news release from the publishing society reporting the 2014 meta-analysis, quoting the lead author that deliberate practice is just less important than has been argued, and that for scientists the important question now is what else matters; and giving the study citation as Macnamara, B. N., Hambrick, D. Z., and Oswald, F. L., Psychological Science, 2014, DOI 10.1177/0956797614535810. Note: a press release, not a peer-reviewed source, flagged as such in the body. Used only for the author's own characterisation of her finding. sciencedaily.com
- Reference list carried on the publisher's record for the 2016 sports meta-analysis, identifying Meinz, E. J., and Hambrick, D. Z. (2010), Deliberate practice is necessary but not sufficient to explain individual differences in piano sight-reading skill: The role of working memory capacity, Psychological Science, 21, 914–919; and Macnamara, B. N., Hambrick, D. Z., and Oswald, F. L. (2014), Psychological Science, 25, 1608–1618. Note: a reference list; the 2010 title is quoted for the claim it makes and we did not obtain the paper. journals.sagepub.com
This article discusses research on expertise and is not training, professional development or human resources advice. Neither meta-analysis was obtained beyond its abstract. A corrigendum to the 2014 paper was issued in 2018 and its contents could not be obtained, which bears on every figure quoted. The 1993 paper being tested and any reply from its authors were not obtained, so one side of a live dispute is reported in detail and the other only through third-party characterisations. All arithmetic is the authors' own; the conversion of variance to ordering probabilities assumes bivariate normality, which no source establishes. This literature measures how much variation between people is explained by variation in their practice, which is not the question of whether practice improves a given person.