The eighty-fifth article found a number that was not in the paper it was credited to. This one concerns a dataset that was not analysed at all, and a conclusion drawn from it for seventy-five years.

Key Takeaway

The published abstract states: "Existing descriptions of supposedly remarkable data patterns prove to be entirely fictional." It also states: "We do find more subtle manifestations of possible Hawthorne effects."[1] On our own arithmetic, a twelve-week pilot on the worst of ten teams gains 16 to 22 percent from a learning curve and regression alone, before the intervention does anything.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The received account of the founding study is false, and this is stated in a published abstract. The phrase used is "entirely fictional," in the American Economic Journal.

Two. The authors did not dismiss the effect. The same abstract reports finding more subtle manifestations of possible Hawthorne effects, and any account of this that omits that sentence is doing what it accuses others of.

Three. A separate reanalysis of a different Hawthorne study reached a compatible conclusion, attributing the productivity variation to incentive pay and ordinary learning rather than to observation.

Four. The practical problem is not observation at all. On our own arithmetic, a pilot programme improves substantially from learning and from regression to the mean, both of which happen without any intervention.

Five. Which means a pilot without a control group cannot tell you whether your change worked, regardless of what you believe about being watched.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the 2011 findings, from a published abstract obtained verbatim from the journal's own association, and a fuller framing from the working paper hosted by the bureau that issued it.

Grade A for the fact that the data exists and is public, from a repository record giving its own DOI.

Grade C for the 1992 reanalysis, which reaches us through a guide site's description rather than from the paper.

Grade C for the historical detail about the original write-up, from an educational site quoting two secondary sources.

Grade D for the 2014 systematic review, whose citation we have and whose conclusion our reproduction cuts off before reaching.

Grade A for our own arithmetic, which is exact and which contains one overclaim we caught and corrected, documented below.

A Note On Method

Everything here is verified to August 2026.

We obtained the published abstract verbatim from the publishing association[1], and the working paper's fuller framing from the research bureau that issued it[2].

We obtained a data repository record confirming the replication data is publicly archived with its own DOI[3].

Two sources are non-academic: a research guide site[4] and an educational psychology site[5], both flagged at every use.

We did not obtain the 2011 paper itself, the 1992 reanalysis, the 2014 systematic review, or any original Hawthorne document.

All arithmetic is ours and every parameter in it is invented. It demonstrates what can produce a pilot result, not what did.

This article discusses research methodology. It is not management advice.

The Claim Being Tested

What the received account says, stated by the authors who tested it.

The working paper records: "The 'Hawthorne effect,' a concept familiar to all students of social science, has had a profound influence both on the direction and design of research over the past 75 years."[2]

And the specific claim: "Both academics and popular writers commonly summarize the results as showing that every change in light, even those that made the room dimmer, had the effect of increasing productivity."[2]

Four observations, ours.

That is a very strong claim and it is what makes the story memorable. Not that observation helps, but that every change helped, including changes that should have hurt.

The strength is the point. A dramatic pattern is what made this the most-taught study in management, and a modest one would have been forgotten.

The authors are careful to attribute the summary to both academics and popular writers, which locates the claim in the scholarly literature and not only in business books.

And seventy-five years is the span they give for its influence on research design, which is the reason a correction warranted a paper in a general economics journal.

Two consequences of that span, ours. The concept shaped how studies were designed, not merely what people believed, so a false premise propagated into methodology across several fields.

And it acquired the status that makes checking feel unnecessary. A finding taught in every introductory course is one nobody expects to have to verify, which is the condition under which an unverified claim survives longest.

The 2011 Paper

The source.

Levitt, S. D., and List, J. A. (2011), Was There Really a Hawthorne Effect at the Hawthorne Plant? An Analysis of the Original Illumination Experiments, American Economic Journal: Applied Economics, 3(1), 224–238, January, DOI 10.1257/app.3.1.224[1]. An earlier version circulated as a working paper[2].

Four observations, ours.

The venue is an economics journal, and the subject is the founding study of industrial psychology, which is worth noting: the correction came from outside the field that owned the finding.

The title is a question, which is a more modest framing than the abstract turns out to require.

The published version and the working paper carry entirely different subject classification codes[1][2], which we return to in the bibliographic note.

And the first author is widely known outside economics, which cuts both ways. It gave the reanalysis reach it might not otherwise have had, and popular reach is not evidence.

The Sentence In The Abstract

The finding, in the published abstract's own words.

"The data from the first and most influential of these studies, the 'Illumination Experiment,' were never formally analyzed and were thought to have been destroyed. Our research has uncovered these data. Existing descriptions of supposedly remarkable data patterns prove to be entirely fictional."[1]

Four observations, ours.

"Entirely fictional" is not hedged, not qualified, and appears in a published abstract in a journal of the American Economic Association. That is about as blunt as the literature gets.

The phrase attaches to descriptions of data patterns, not to the concept and not to the researchers' honesty. It says the pattern everyone describes is not in the data.

"Never formally analyzed" is the more remarkable clause and gets less attention. The founding evidence for a concept taught in every management course had never been subjected to statistical analysis at all.

And a summary that stops at this sentence would be doing precisely what the paper criticises, which is why the next section exists.

Two things the sentence does not say, ours, because both are commonly read into it. It does not say the original researchers fabricated anything. The fiction is in the descriptions, which accumulated later.

And it does not say the plant workers were unaffected by the experiment. It says the pattern that everyone describes, in which every lighting change raised output, is not present in the numbers.

What They Did Find

The part that a fair account has to include.

The abstract continues: "We do find more subtle manifestations of possible Hawthorne effects. We also propose a new means of testing for Hawthorne effects based on excess responsiveness to experimenter-induced variations relative to naturally occurring variation."[1]

We did not obtain the paper, so we cannot describe what those subtle manifestations were or how large.

Four observations, ours.

The authors did not conclude the effect does not exist. They concluded the famous description of it is false and that something smaller may be present.

The words "subtle" and "possible" are both doing work, and neither survives in the popular retelling of the debunking.

The proposed test is the constructive part and it is genuinely clever. Compare responsiveness to changes an experimenter makes against responsiveness to changes that happen anyway, and the excess is attributable to the experiment rather than to the change.

And that test is available to any business running a pilot, which we return to below. It requires only that you also measure what happens when something changes for ordinary reasons, which most firms have in their history and never look at.

One reason that proposal is more valuable than the debunking, ours. A correction tells you a belief was wrong; a test tells you how to find out for yourself, and only the second survives being applied to a new situation.

Data Thought Destroyed

The provenance, which is a story in itself.

The abstract records the data "were thought to have been destroyed" and that "our research has uncovered these data."[1]

A repository record confirms the replication dataset is now publicly archived with its own DOI, deposited by the journal's publisher and distributed by a university research consortium[3].

Four observations, ours.

The trajectory is worth stating plainly. Never analysed, believed destroyed, found, analysed, and now downloadable by anyone with an internet connection.

Which means the claim in this article is checkable by a reader in a way that almost nothing else in this series is. The data is there.

That is worth an explicit invitation, ours. A reader who doubts this article can download the underlying figures and look, which is a standard almost no business writing meets and which we would rather point at than paraphrase around.

It also raises a question the paper does not answer and we cannot. How did a claim survive seventy-five years without anyone establishing that the underlying data existed?

And the answer is presumably mundane and is the theme of the last several articles. Nobody checked, because the citation looked solid and the story was good.

Two features of this case make it worse than the ones the last three articles documented, ours. Those concerned a qualifier lost in transmission from a real analysis. Here there was no analysis to transmit from.

And the gap was seventy-five years rather than fifteen or twenty, across a period in which the study was taught continuously to people whose profession was the evaluation of evidence.

A Few Paragraphs In A Trade Journal

The historical detail that makes the rest intelligible.

An educational psychology site quotes a 1982 article recording that "the original [illumination] research data somehow disappeared," and quotes a 2004 source more starkly: "[T]hese particular experiments were never written up, the original study reports were lost, and the only contemporary account of them derives from a few paragraphs in a trade journal."[5]

This is an educational website quoting secondary sources, not research, flagged here and at every use. We obtained neither of the works quoted.

Four observations, ours.

If that description is accurate, the evidential base for the most influential study in management was a few paragraphs in a trade journal.

That would explain how the description became dramatic. A short secondhand summary is exactly the format in which qualifications disappear, and the eighty-third and eighty-fifth articles both documented the same process operating over shorter timescales.

We are relying here on a non-academic site quoting two sources we did not obtain, which is thin, and we grade it accordingly.

But it is consistent with what the 2011 abstract says directly. "Never formally analyzed" is the journal's own phrasing[1], and a study never formally analysed was necessarily described from something other than an analysis.

The Other Reanalysis

An earlier and independent challenge, to a different Hawthorne study.

A research guide site records: "Stephen R. G. Jones (1992) examined the relay-assembly-test-room records in a paper published in the American Journal of Sociology."[4]

This is a guide site, not research, flagged at every use, and we did not obtain the 1992 paper.

Four observations, ours.

This concerns the relay assembly test room, a different experiment from the illumination one, so it is an independent line of evidence rather than the same challenge repeated.

It is nineteen years earlier than the 2011 paper, which means the concerns were in the literature well before the illumination data was found.

The venue is a leading sociology journal, so this was not a fringe objection.

And its proposed explanation is the one that matters for a business, which the next section sets out.

One point about the sequence, ours. The sociology reanalysis preceded the economics one by nineteen years and did not settle the matter, which suggests that a challenge published in a good journal is not by itself enough to dislodge a story people like.

Incentive Pay And Learning Curves

What the 1992 reanalysis attributed the results to, on our source's description.

It records that Jones found that "once the introduction of a new incentive pay scheme and ordinary learning-curve effects over the course of the study were accounted for, there was very little productivity variation left to attribute to 'being observed' as a mechanism in its own right."[4]

Again a guide site, and we did not obtain the paper.

Four observations, ours.

Two ordinary explanations account for the result on this description. They paid people more, and people got better at the job over time.

Neither requires any psychology. Incentive pay changing output is economics, and a learning curve is what happens when anybody repeats a task, and both would have operated whether or not anyone was watching.

Which reframes the entire episode. The founding study of the psychology of observation may be a demonstration of piece rates and practice, on this account, which we could not verify directly.

And that reframing is what makes the rest of this article possible, because both of those confounds are present in every pilot programme a business has ever run.

One caution before we build on it, ours. This reaches us as a guide site's one-sentence summary of a paper we did not read, and a summary can drop qualifications exactly as this article has documented elsewhere.

What makes us willing to build on it is that the arithmetic below does not depend on the 1992 paper being right. Learning curves and regression to the mean operate whether or not they explained Hawthorne, and we compute them directly.

The Systematic Review

The modern assessment, whose conclusion we do not have.

Our source cites McCambridge, J., Witton, J., and Elbourne, D. R. (2014), Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects, Journal of Clinical Epidemiology, 67(3), 267–277[5], describing it as asking whether the effect exists at all, under what conditions, and how large.

Our reproduction states that "following a systematic review of the evidence, pooling findings across many studies rather than relying on one, the researchers concluded" and then cuts off.[5]

Four observations, ours.

We do not have the conclusion, and this is the largest gap in this article. A systematic review is the strongest form of evidence available on this question.

The title itself carries information, and it is the informative part we do have. "New concepts are needed to study research participation effects" implies the existing concept was found inadequate.

The venue is clinical epidemiology rather than management, which is where the practical stakes are highest, since a trial confounded by participation effects can misprice a treatment.

And a reader who wants this answer can obtain the paper in a few minutes. We could not, and say so rather than characterising a conclusion we have not read.

What Actually Survives

Our reading, stated directly.

Five statements.

The famous description of the illumination experiment is false, on the published finding of a reanalysis of the recovered data.

The data was never formally analysed before that, which the same abstract states plainly.

The authors nonetheless report subtle manifestations of possible Hawthorne effects, and an account omitting this is not a fair one.

An independent reanalysis of a different Hawthorne experiment attributed its results to incentive pay and learning, on a description we obtained rather than from the paper.

And a systematic review exists whose conclusion we could not obtain, which is the honest state of our knowledge.

One thing that follows from all five together, ours. The concept is contested rather than refuted, and a business reader should treat it as an open question rather than as either a law or a myth.

Which is fortunate, because the practical problem below does not depend on resolving it. The confounds that will contaminate your pilot are arithmetic and would operate in an empty room.

Your Pilot Programme

Why any of this matters to a business. Ours.

Four observations.

Every firm runs pilots. Try the new process on one crew, the new script on one salesperson, the new supplier on one line, and see whether the numbers improve.

The usual worry, if there is one, is that the pilot team knew they were being watched. That is the Hawthorne worry and it is the one people have heard of.

The 1992 reanalysis suggests the worry is misdirected. The confounds that mattered there were pay and practice, both of which are present in every pilot and neither of which anyone thinks about.

So the useful question is not whether observation inflates a pilot. It is how much a pilot would improve if you changed nothing at all, and that has an arithmetic answer.

What A Learning Curve Alone Produces

Our own arithmetic, on invented figures, using the standard formulation in which unit time falls by a fixed percentage each time cumulative output doubles.

A team producing 100 units a week, starting from 2,000 cumulative units, over a twelve-week pilot:

At a 95 percent learning rate: 3.5 percent productivity gain. At 90: 7.4. At 85: 11.7. At 80 percent: 16.3 percent.

Four observations.

An 85 percent learning curve produces an 11.7 percent gain over twelve weeks with no intervention whatever.

Learning rates in that range are ordinary rather than exceptional in production settings, though we are supplying the number rather than measuring it and it varies enormously by task.

The mechanism requires nothing of anybody. People repeating a task get faster at it, and the curve was running before your pilot started and will run after it ends.

And it means a pilot showing an eleven percent gain has shown nothing at all unless something rules this out.

Two ways to rule it out, ours, and both are cheap. Look at the same team's trend in the twelve weeks before the pilot, which shows the curve you were already climbing.

Or compare against a team that did not receive the change, which is subject to the same curve and differences it out. The first is available to any firm with records; the second needs a second team.

The Newer The Team, The Bigger The Illusion

The property that makes this worse in practice. Ours, same invented figures at an 85 percent learning rate.

Twelve-week gain, by cumulative output at the start:

At 500 units: 33.2 percent. At 1,000: 20.3. At 2,000: 11.7. At 5,000: 5.2. At 10,000: 2.7. At 20,000: 1.4 percent.

Four observations.

A new team improves 33 percent in twelve weeks from practice alone, and an experienced one improves 1.4 percent.

Which produces a perverse incentive nobody intends. Pilot a change with a new team and it will look spectacular, and the newer the team the better it will look.

That is also the team a manager is most likely to choose, for reasons that seem sensible. New teams have no established way of doing things and are willing to try something, which makes them the natural pilot site and the worst measurement.

And the effect is largest exactly where the numbers are most convincing, which is the general shape of every problem in this series.

One further consequence worth naming, ours. The same arithmetic says a change that genuinely works will look weakest on your most experienced team, because there is little learning left to add to it.

Which means the two errors run opposite ways and can cancel confusingly. A good change piloted on a veteran crew may show almost nothing, and be discarded, while a worthless one piloted on a new crew shows thirty percent and gets rolled out.

And If You Picked The Worst Team

The second confound. Ours, invented figures.

Teams vary around an average. If you pick the team that performed worst last quarter, part of that shortfall was bad luck, and bad luck does not repeat.

The worst of N teams sits, in expectation, this far below average: at 3 teams, 0.85 standard deviations. At 5: 1.16. At 10: 1.54. At 20: 1.87 standard deviations.

With performance varying by 10 percent, the worst of ten teams sits 15.4 percent below average as observed.

Three observations.

Some of that shortfall will reverse next quarter whatever you do, which is the free rebound a pilot on that team collects.

How much reverses depends on how much of the gap was luck, which is the question the next section addresses.

And it is the question we initially got wrong.

Before that, one clarification on what regression to the mean is and is not, ours. It is not a force pulling teams toward average, and nothing causes it in the way a change causes an effect.

It is a selection artefact of the kind the eighty-first article described. You selected on an extreme observation, and extreme observations contain more than their share of luck, which does not repeat because it never was a property of the team.

A Correction To Our Own Figure

Documented here rather than removed, in this series' usual habit.

We first computed the free rebound as the whole 15.4 percent. That is wrong, and it is wrong in the direction that flatters our own argument.

The full shortfall reverses only if every team is genuinely equally good and the ranking is pure luck. If some of the gap is real, the worst team is really worse and stays worse.

The correct expression uses reliability, being the share of observed variation that is real rather than noise. The expected rebound is one minus reliability, times the observed shortfall.

On a 15.4 percent observed shortfall: at reliability 0.0, rebound 15.4 percent, which was our first figure. At 0.2: 12.3. At 0.5: 7.7. At 0.8: 3.1. At 1.0: zero, no rebound at all.

Four observations, ours.

Our first figure was the extreme case presented as the answer, which is the same error we documented in the eighty-fourth article and have now made in a third consecutive article.

A plausible middle assumption of 0.5 reliability gives 7.7 percent, which is half what we first printed.

We do not know the real reliability of team performance in any business, and neither does the business, which is itself worth knowing.

And the correction matters for honesty rather than for the conclusion, as the next section shows.

Two notes on why we keep publishing these, ours. An article arguing that pilots overstate their results has an obvious incentive to overstate the confounds, and a reader is entitled to know we noticed ours running that way.

And the pattern is now consistent enough to name as a working rule. Our errors cluster at the extreme case presented as the typical one, which is a specific thing to check rather than a general resolution to be careful.

Restacking It Honestly

The combined figure, with the correction applied. Ours, invented parameters throughout: a twelve-week pilot on the worst of ten teams, an 85 percent learning curve, 2,000 cumulative units, 10 percent performance variation.

At reliability 0.3: learning +11.7, regression +10.8, total +22.4 percent.

At reliability 0.5: +11.7 and +7.7, total +19.3 percent.

At reliability 0.7: +11.7 and +4.6, total +16.3 percent.

Four observations.

Even at reliability 0.7, where most of the team gap is genuinely real, the pilot gains 16.3 percent before the intervention does anything.

The conclusion survives our own correction, which is the test we applied in the previous article and which we would rather run than skip.

A fifteen percent pilot result is fully accounted for at every reliability assumption we tested, and we have invoked observation nowhere.

And every parameter here is invented. This shows what can produce a pilot result, not what did, and a firm wanting the real figures would have to measure its own.

One thing the arithmetic does establish regardless of the parameters, ours. Both confounds push in the same direction, upward, which means they add rather than partially cancelling.

And there is no corresponding pair of confounds pushing down. Nothing about running a pilot makes a team systematically worse, so the bias in pilot measurement is one-directional and cumulative.

Why Pilots Almost Always Succeed

The consequence, and it explains something every owner has noticed. Ours.

Four observations.

Pilots succeed at a rate that should be suspicious, and the arithmetic above supplies a sufficient explanation without anybody being dishonest.

The selection compounds it. Pilots run on willing teams, with attention from management, with the newest processes, on the site with most room to improve, and each of those independently produces improvement.

And the rollout is where it shows. A change that gained fifteen percent in pilot and nothing across the business has not failed at rollout; it never worked, and the pilot measured something else.

Which is the expensive part, because the rollout is where the money is spent, and the pilot was supposed to be the cheap way of avoiding that.

Two costs follow from that and only one gets counted, ours. The direct cost of a rollout that does not work appears in the accounts and is painful and visible.

The larger one usually does not. The change you rejected because its pilot looked flat may have been the one that worked, tested on a veteran crew with no learning left, and nothing in any report will ever show you that.

The Third Confound Nobody Lists

Two are computed above. The third resists arithmetic and is probably the largest. Ours.

Four observations.

A pilot is run by whoever proposed it, and that person wants it to work. Not dishonestly, but with the ordinary energy anyone brings to their own idea.

That energy is itself an intervention, and it is not the one being tested. The pilot receives management attention, faster problem-solving and better inputs, none of which will be present at rollout.

Which produces the specific failure the previous section described. The rollout is run by people who inherited somebody else's idea, without the attention that made the pilot work, and it underperforms for reasons the pilot could never have shown.

And we cannot compute this one, which is worth admitting rather than skipping. The two confounds we quantified are the tractable ones and probably not the biggest, and a firm that fixed only those would still be misled.

One partial remedy that does not require measuring it, ours. Have the rollout run by whoever will actually own it, during the pilot, rather than by the person whose idea it was.

That makes the pilot a test of the change under realistic ownership, which is the condition it will face. A pilot that only works when its author is running it has told you something useful, and it is not that the change works.

What A Readable Pilot Looks Like

Assembling the requirements. Ours, and not management advice.

Four points.

A comparison that changes nothing, subject to the same season, market and learning curve, so that the difference between them is attributable to the change.

Assignment you did not choose, because every basis on which you would choose introduces one of the confounds above.

A baseline long enough to show the existing trend, so that a rising line is not mistaken for a step change.

And a stated threshold before you start. Deciding in advance what result would count as success prevents the outcome from setting the standard, which is the seventy-ninth article's point about deciding after the fact.

One addition to that threshold, ours, which follows from everything above. State it net of what you would expect anyway, so that a target of fifteen percent becomes fifteen percent above the trend the baseline showed.

What A Control Group Costs

The remedy, and it is cheaper than it sounds. Ours, and not management advice.

Four points.

A control group means running the pilot on one team and measuring a comparable team that changes nothing. Both are subject to the learning curve, the season, and the market.

The cost is not money. It is that you must measure a team you are not helping, which feels wasteful and is the entire value of the exercise.

Assignment matters more than most firms realise. Choosing the pilot team for enthusiasm or for weakness introduces exactly the confounds above, and choosing at random removes them, which is why randomisation exists.

And for a small firm the practical version is modest. Two crews, two branches, two salespeople, assigned by coin flip, is a real experiment and costs nothing beyond the discipline of not choosing.

Two objections owners raise, ours, and both have answers. The teams are not identical, which is true and is why assignment is random rather than matched: randomness makes the differences unsystematic rather than absent.

And it is unfair to withhold an improvement from one team, which assumes you already know it is an improvement. If you knew that, you would not be piloting it, and the team you withheld it from can have it in twelve weeks.

When You Genuinely Cannot Have One

Because many small businesses have one team and no comparison available. Ours.

Four points.

The 2011 paper's own proposal is usable here, and it is the constructive contribution people skip. Compare responsiveness to changes you made against responsiveness to changes that happened anyway.[1]

In a small business that means looking at your own history. What did output do in the twelve weeks after the last supplier change, the last new hire, the last software switch, none of which were pilots and all of which were changes.

If output rose after those too, you have learned that your numbers rise after changes generally, which is the excess-responsiveness test applied to a firm rather than a laboratory.

And the cheap alternative is a longer baseline. Measure for twelve weeks before you start, which will show you the learning curve you were already on and is free.

One further single-team design worth knowing, ours. Switch the change on, off, and on again, which most firms never consider and which is available whenever a change is reversible.

A learning curve does not reverse when you switch something off. If output falls back during the off period and rises again after, that pattern is hard to explain by practice, and it is the strongest evidence a single team can give you.

The Lesson That Is Not About Observation

Where we land. Ours.

Four observations.

The Hawthorne effect is taught as a lesson about people behaving differently when watched, and that may well be true in some degree.

But the episode itself teaches something else and larger. A dramatic pattern was described from data nobody had analysed, and it survived seventy-five years in textbooks.

And the reanalyses point at confounds that have nothing to do with psychology. Pay, practice and selection, which are the same three things that will contaminate your pilot next quarter.

So the practical inheritance from Hawthorne is not a caution about being watched. It is a caution about measuring change without a comparison, which is what the original researchers did and what almost every business still does.

And that inheritance is more useful than the one usually taken, ours. A caution about observation offers nothing to do, since you cannot run a business without anybody noticing.

A caution about comparison offers a specific practice, which is the difference between a fact people repeat and a fact people use.

Bibliographic Note

The series keeps a count.

The working paper version carries subject classification codes A0, C91, C92, C93, D03 and L22[2]. The published version carries C90, J24, J28, M12, M54 and N32[1]. Not one code is common to both.

Three observations, ours.

This is not an error by anyone. Classification codes are routinely revised between working paper and publication, and both sets are defensible for this paper.

We record it because it has a practical consequence for anyone searching by subject. A researcher browsing either code set would not find the other version, and the working paper is the freely available one.

That brings the running count of bibliographic variants across this series to thirty-one.

And it is the mildest kind we have recorded, being nobody's mistake at all. We include it because the count is a record of what a careful reader encounters, and a search that silently misses the free version of a paper is a real obstacle whoever is at fault.

What To Do

Assume your pilot will improve without your change. On our own arithmetic a twelve-week pilot on the worst of ten teams gains 16 to 22 percent from a learning curve and regression alone.

Do not pilot on the newest team. At an 85 percent learning rate a team at 500 cumulative units gains 33 percent in twelve weeks from practice, against 1.4 percent for a team at 20,000.

Do not pilot on the worst team either, because part of what made it worst was luck and luck does not repeat.

Assign by coin flip if you possibly can. Two crews, two branches or two salespeople chosen at random is a real experiment, and choosing removes the thing that makes it one.

Measure a baseline first if you cannot have a control. Twelve weeks of before-data shows you the trend you were already on and costs nothing.

Look at what happened after changes you did not pilot. That is the 2011 paper's own proposed test, applied to your history, and most firms have the data.

Treat a rollout that underperforms its pilot as a measurement failure, not an execution failure. The usual conclusion is that people stopped trying, and the arithmetic offers a simpler one.

And do not conclude the effect is nothing. The paper that called the famous description fictional also reported finding subtle manifestations of possible Hawthorne effects, and both sentences are in the same abstract.

The Limits Of This Analysis

Several caveats matter. This article discusses research methodology and is not management advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 2011 paper, only its published abstract from the publishing association and the working paper's framing from the issuing bureau; we therefore cannot describe the subtle manifestations of possible Hawthorne effects that the authors report finding, and have not characterised their size. We did not obtain the 1992 reanalysis, and every statement about it reaches us through a research guide site's description; we did not obtain the 2014 systematic review, and our reproduction of its account cuts off immediately before its conclusion, which is the largest gap here since a systematic review is the strongest evidence available on this question. We did not obtain any original Hawthorne document. The historical claim that the illumination experiments were never written up and survive only in a few paragraphs of a trade journal comes from an educational website quoting two secondary sources we did not obtain, and is graded accordingly, though it is consistent with the published abstract's own statement that the data were never formally analysed. Two of our five sources are non-academic and are flagged at every use. All arithmetic is ours and every parameter in it is invented: the learning rate, the cumulative output, the team count, the performance variation and the reliability are all supplied by us and measured by nobody, so the tables show what can produce a pilot result rather than what did. The learning-curve formulation assumes a constant rate, which real production rarely obeys. And this article contains one documented correction to our own working: we first computed a regression rebound of 15.4 percent, which assumes team differences are entirely noise, and the honest range given a reliability parameter is 3.1 to 12.3 percent.

Frequently Asked Questions

What did the reanalysis find?
The published abstract states that the illumination experiment data were never formally analysed and were thought destroyed, that the authors recovered them, and that existing descriptions of supposedly remarkable data patterns prove to be entirely fictional. It also reports finding more subtle manifestations of possible Hawthorne effects.
So the Hawthorne effect does not exist?
That is not what the paper says, and an account stopping at the word "fictional" would be doing what the paper criticises. The authors report subtle manifestations of possible effects. What was found false is the dramatic description, in which every lighting change including dimmer ones raised productivity.
What explains the original results then?
A separate 1992 reanalysis of a different Hawthorne experiment attributed the productivity variation to a new incentive pay scheme and to ordinary learning-curve effects, leaving very little to attribute to being observed. We obtained that through a guide site's description rather than from the paper.
Why does this matter for my business?
Because pay, practice and selection confound every pilot programme, and those are the confounds the reanalyses point at. On our own arithmetic a twelve-week pilot on the worst of ten teams gains 16 to 22 percent from a learning curve and regression to the mean before your change does anything.
Which team should I pilot on?
Not the newest, because at an 85 percent learning rate a team at 500 cumulative units gains 33 percent in twelve weeks from practice alone. Not the worst, because part of what made it worst was luck. On our own arithmetic, a team chosen at random is the only one that gives you a readable answer.
What if I only have one team?
Measure a longer baseline before you start, which shows the trend you were already on. And apply the 2011 paper's own proposed test to your history: look at what output did after changes you did not pilot, such as a supplier switch or a new hire. If it rose after those too, you have learned something.
My rollout underperformed the pilot. What happened?
The usual conclusion is that people stopped trying once the attention moved on. The arithmetic offers a simpler possibility: the change never worked, and the pilot measured a learning curve, a rebound from a bad quarter, or a team selected for enthusiasm. That is a measurement failure rather than an execution failure.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article contains one correction to its own arithmetic, made in the direction that had flattered our argument.

References

  1. Publishing association record for Levitt, S. D., & List, J. A. (2011), Was There Really a Hawthorne Effect at the Hawthorne Plant? An Analysis of the Original Illumination Experiments, American Economic Journal: Applied Economics, 3(1), 224–238, January, DOI 10.1257/app.3.1.224, reproducing the abstract in full: that the Hawthorne effect draws its name from a landmark set of studies conducted at the Hawthorne plant in the 1920s; that the data from the first and most influential of these, the Illumination Experiment, were never formally analysed and were thought to have been destroyed; that the authors' research has uncovered these data; that existing descriptions of supposedly remarkable data patterns prove to be entirely fictional; that the authors do find more subtle manifestations of possible Hawthorne effects; and that they propose a new means of testing for Hawthorne effects based on excess responsiveness to experimenter-induced variations relative to naturally occurring variation. The record carries the subject classification codes C90, J24, J28, M12, M54 and N32. Note: the publishing association's own record. Our source for the abstract. We did not obtain the paper, and therefore cannot describe the subtle manifestations the authors report finding, or their size. aeaweb.org
  2. Research bureau record for the working paper version of the same study, number 15016, May 2009, reproducing its framing: that the Hawthorne effect, a concept familiar to all students of social science, has had a profound influence both on the direction and design of research over the past seventy-five years; that it is named after a landmark set of studies conducted at the Hawthorne plant in the 1920s, the first and most influential of which is known as the Illumination Experiment; that both academics and popular writers commonly summarise the results as showing that every change in light, even those that made the room dimmer, had the effect of increasing productivity; and that the data from the illumination experiments were never formally analysed and were thought to have been destroyed, but that the authors' research has uncovered them. The record carries the subject classification codes A0, C91, C92, C93, D03 and L22. Note: the issuing research bureau's record for the freely available working paper version. Our source for the precise claim being tested. Recorded also as a bibliographic variant: none of its classification codes appear in the published version's set. nber.org
  3. Data repository record confirming that the replication data for the 2011 study is publicly archived under its own DOI 10.3886/E113776V1, deposited by the publishing association and distributed by a university research consortium, and reproducing the summary that the data from the Illumination Experiment were never formally analysed and were thought to have been destroyed. Note: a data repository record. Recorded because it establishes that the recovered data is now publicly available, so a reader can check the claims in this article against the underlying figures in a way that is possible for almost nothing else in this series. We did not download or analyse it. openicpsr.org
  4. Research guide site describing the reanalysis history, recording that Stephen R. G. Jones (1992) examined the relay-assembly-test-room records in a paper published in the American Journal of Sociology, and found that once the introduction of a new incentive pay scheme and ordinary learning-curve effects over the course of the study were accounted for, there was very little productivity variation left to attribute to being observed as a mechanism in its own right; and recording that Levitt and List located the original illumination-study data and published a formal reanalysis, whose central finding was that output did not track the lighting manipulations in the clean, dramatic pattern usually described. Note: a research guide site, NOT an academic source, flagged at every use. Our only source for the 1992 reanalysis, which we did not obtain, and every statement about it here is this site's characterisation. casrai.org
  5. Educational psychology website summarising the Hawthorne literature, quoting a 1982 article that the original illumination research data somehow disappeared, and a 2004 source that these particular experiments were never written up, the original study reports were lost, and the only contemporary account of them derives from a few paragraphs in a trade journal; describing the Levitt and List reanalysis as finding the dramatic patterns largely absent with at most modest and inconsistent traces of an effect; and citing McCambridge, J., Witton, J., & Elbourne, D. R. (2014), Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects, Journal of Clinical Epidemiology, 67(3), 267–277, as asking whether the effect exists at all, under what conditions and how large, with our reproduction cutting off immediately before that review's conclusion. It also cites McCarney, R., and colleagues (2007), The Hawthorne Effect: a randomised, controlled trial, BMC Medical Research Methodology, 7(1), 1–8. Note: an educational website, NOT an academic source, flagged at every use. Our source for the historical detail about the original write-up, which it quotes from two secondary works we did not obtain. Our reproduction of the systematic review's conclusion is truncated and we therefore do not report it. simplypsychology.org

This article discusses research methodology and is not management advice. The 2011 paper was not obtained, only its abstract, so the subtle effects its authors report finding are not described or sized here. The 1992 reanalysis and the 2014 systematic review were not obtained; the review's conclusion is truncated in our reproduction and is not reported. Two of the five sources are non-academic and are flagged at every use. All arithmetic is the authors' own and every parameter in it is invented. This article contains one documented self-correction.