Thirty articles into a series about what behavioural research supports, this one turns the same question on the business books themselves. The answer is not that they are dishonest. It is that the method most of them use cannot produce the answer they claim to have found.
Key Takeaway
The halo effect describes the basic human tendency to make specific inferences on the basis of a general impression[1]. Applied to companies: when a company's sales and profits are up, people often conclude that it has a brilliant strategy, a visionary leader, capable employees, and a superb corporate culture, and when performance falters they conclude the reverse, though in fact, little may have changed[2]. Of the companies identified as Excellent in a famous 1982 study, over the next five years, only one-third grew faster than the overall stock market[1].
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the halo effect itself. It dates to 1920, has a dedicated Psychological Bulletin review titled Ubiquitous Halo[1], and is among the longest-standing findings in applied psychology.
Grade B for the argument that it contaminates business research. It is a methodological argument published in a peer-reviewed management journal, and it is an argument rather than an experiment.
Grade C for the specific performance figure, for a reason we discovered by checking it and which has its own section.
Our position: the methodological argument is stronger than the headline statistic supporting it, which is an awkward thing to report about a critic we largely agree with.
A Note On Method
Everything here is verified to August 2026.
We obtained portions of the critic's own peer-reviewed journal article[1], which we prefer throughout to his book, which we did not obtain and which we cite only through its publisher description[2].
We obtained Thorndike's own published abstract[3] but not his paper.
We found two conflicts between sources on the 1920 study, including one in the critic's own footnote, and report both rather than resolving them.
We tested the article's central statistic ourselves and the result is in its own section.
Several supporting works are reported as citations only, obtained not at all.
This article reviews research methods. It is not investment or management advice, and nothing in it is a comment on any named company.
The 1920 Finding
The origin.
Thorndike published A constant error in psychological ratings in the Journal of Applied Psychology in 1920[3][8].
From his own abstract: "In a study made in 1915 of employees of two large industrial corporations, it appeared that the estimates of the same man in a number of different traits such as intelligence, industry, technical skill, reliability, etc., were very highly correlated and very evenly correlated."[3]
Two observations, ours.
The phrase very evenly correlated is the tell. It is not merely that good people were rated well across the board, which could be true. It is that the correlations between quite different traits were uniformly high, which is what you would see if raters were producing one judgment and repeating it.
And note the date of the data: 1915. This is one of the oldest quantitative findings in applied psychology, and it concerns exactly the activity every manager performs at review time.
Two Conflicts In Our Sources
Two discrepancies we found while sourcing, reported rather than resolved.
The page numbers. Most sources give the 1920 paper as Journal of Applied Psychology, 4(1), 25–29[3][4]. The critic's own footnote in his peer-reviewed article gives it as 4 (1920): 469–477[1].
The sample. Thorndike's own abstract describes employees of two large industrial corporations in 1915[3]. Two secondary sources instead describe army commanding officers rating soldiers on physique, intelligence, leadership and personal qualities[5][6].
Two observations, ours.
The paper may well contain more than one study, which would reconcile the second conflict. We do not know, because we did not obtain it, and we are not going to assume.
And the page discrepancy sits in a footnote in a peer-reviewed journal article by an author whose entire subject is errors in business research. We report that without comment beyond noting it is the eighth time in this series that a famous paper's details have been recorded inconsistently by careful sources.
What Thorndike Concluded
His explanation, which is more specific than the modern shorthand.
"It consequently appeared probable that those giving the ratings were unable to analyze out these different aspects of the person's nature and achievement and rate each in independence of the others."[3]
Three observations, ours.
The claim is about inability, not laziness. Thorndike is not saying raters could not be bothered to distinguish traits. He is saying they could not.
The modern gloss is that the halo effect is the tendency to make specific inferences on the basis of a general impression[1], which is compatible but flatter. Thorndike's version names a specific failure: the traits could not be analyzed out from one another.
And that matters for every appraisal form ever designed. A form with eight separate rating scales assumes the rater can produce eight separable judgments, and this finding says that assumption was tested in 1915 and failed.
The Business Version
The extension to companies.
Rosenzweig published The Halo Effect: ...and the Eight Other Business Delusions That Deceive Managers in 2007[2], and a companion article, Misunderstanding the Nature of Company Performance, in California Management Review, 49(4), in the same year[1].
We did not obtain the book and use the journal article throughout, preferring the peer-reviewed venue to the trade one.
The book's publisher description states the argument: "Much of our business thinking is shaped by delusions, errors of logic and flawed judgments that distort our understanding of the real reasons for a company's performance," affecting "the business press and academic research, as well as many bestselling books that promise to reveal the secrets of success"[2].
The Mechanism, Stated Plainly
The central passage, quoted.
"When a company's sales and profits are up, people often conclude that it has a brilliant strategy, a visionary leader, capable employees, and a superb corporate culture. When performance falters, they conclude that the strategy was wrong, the leader became arrogant, the people were complacent, and the culture was stagnant. In fact, little may have changed. Company performance creates a Halo that shapes the way we perceive strategy, leadership, people, culture, and more."[2]
Three observations, ours, and the third is the whole methodological point.
The direction of inference is backwards from the intended one. Researchers want to learn what causes performance. What they gather is descriptions produced by people who already know the performance.
Which means the data is contaminated at source, and no amount of statistical care downstream can decontaminate it. If the input measure of culture already contains the output measure of profit, a correlation between them is guaranteed.
And it explains why such studies always find something. This series has now encountered several literatures where the striking result was built into the design, in the groupthink case by selecting on the outcome and in the Dunning-Kruger case by sorting on a noisy measure.
In Search Of Excellence
The worked example, from the journal article.
The critic describes the method: the authors studied companies including Hewlett-Packard, IBM, Johnson & Johnson, McDonald's, Procter & Gamble, and 3M, then "gathered data from archival sources, press accounts, and from interviews. Based on these data, Peters and Waterman identified eight practices that appeared to be common to the Excellent companies, including 'a bias for action,' 'staying close to the customer,' 'stick to the knitting,'" and simultaneous loose and tight properties[1].
Three observations, ours.
Archival sources, press accounts, and interviews. Every one of those was produced by someone who knew the company was successful. There is no uncontaminated input in the list.
The companies were selected because they were excellent, which is selecting on the outcome, the exact objection the fifteenth article in this series made against groupthink.
And the eight practices are unfalsifiable as stated. No company describes itself as having a bias against action or as ignoring its customers, so the practices could not have distinguished the sample from anybody.
The Five Years After
What happened next, in the critic's words.
"Excellent companies regressed sharply. Over the next five years, only one-third of the Excellent companies grew faster than the overall stock market."[1] The promised "blueprint of long-term prosperity" turned out, he writes, to be a delusion[1].
Two observations, ours.
The word regressed is doing real work and it is the right word. Companies selected for being at the top of a distribution will, on average, be closer to the middle later, for the reasons the twenty-sixth article in this series set out.
And one-third is the number a reader remembers. So we checked it.
Our Own Check On That Figure
What we found. The simulation below is entirely ours, uses invented parameters, and comes from no source.
One third sounds damning. But individual company returns are positively skewed: a small number of very large winners pull the average up above the median. Which means fewer than half of randomly chosen companies beat the average, with no selection effect at all.
We simulated 200,000 companies with lognormal five-year returns and an equal-weighted index, at three levels of dispersion.
At moderate dispersion, 39 percent of companies beat the index.
At high dispersion, 33 percent.
At very high dispersion, 27 percent.
Two observations.
One third is squarely inside the range a random selection would produce. The figure is not evidence that the Excellent companies did unusually badly.
We should be clear about what our simulation does and does not establish. It uses invented parameters and a specific distributional assumption, and it does not tell you what the real base rate was in that period. What it establishes is that the figure cannot be interpreted without the base rate, and the article we obtained does not supply one.
What The Check Does To The Argument
The important part, and it goes the other way from what you might expect. Ours.
Our check weakens the statistic and strengthens the thesis.
Three steps.
If one-third is roughly what random selection gives, then the Excellent companies did not collapse. They performed as though the selection had contained no information about the future at all.
That is exactly what the halo argument predicts. The claim was never that excellence is fatal. It was that the excellence was identified from performance that had already happened, and therefore carried no forward-looking content.
So the honest version of the finding is less dramatic and more damaging: not that the blueprint backfired, but that following it was equivalent to choosing at random.
We note that this is the fourth time in this series we have asked whether a striking number beats the naive baseline. The sixteenth article found experts losing to simple extrapolation, the twenty-fourth found peak-end possibly matched by a simple average, the twenty-sixth found a famous chart reproducible from noise. It would have been inconsistent to apply that test only to people we disagree with.
The Delusion Of Absolute Performance
The second delusion, and the one with the clearest practical bite.
The publisher's description states it: "The Delusion of Absolute Performance: Company performance is relative to competition, not absolute, which is why following a formula can never guarantee results."[2]
Three observations, ours.
This is a point about logic, not evidence, and it does not depend on any study. If your performance depends on what competitors do, then a formula everyone can follow cannot deliver above-average results to everyone who follows it.
Which means the entire genre of secrets-of-success writing is self-defeating at scale. A widely adopted best practice becomes table stakes rather than an advantage.
And it is a useful filter. Any promised improvement that could be adopted by all your competitors tomorrow is not a source of relative advantage, whatever its absolute merits.
Cargo Cult Science
The characterisation, reported with its source flagged.
An encyclopedia summary of the book records the argument that much business writing is what Feynman called cargo cult science, having the superficial trappings of science but operating at the level of story-telling[7]. The same source records that the book also considers some more scientific business research, whose conclusions are more rigorous but do not promise a simple recipe for success[7].
We did not obtain the book and report this characterisation as the encyclopedia gives it.
Two observations, ours.
The second sentence matters more than the first, and it is usually left out. The argument is not that business research is worthless. It is that the rigorous work does not sell, because rigour and a simple recipe are close to incompatible.
That is an uncomfortable structural point rather than a complaint about any author. The market selects for the recipe.
The Romance Of Leadership
A related literature, from the critic's own footnotes.
His references include Meindl, J. R., Ehrlich, S. B., and Dukerich, J. M. (1985), The Romance of Leadership, Administrative Science Quarterly, 30(1), 78–102; and Meindl, J. R., and Ehrlich, S. B. (1987), The Romance of Leadership and the Evaluation of Organizational Performance, Academy of Management Journal, 30(1), 91–109[1].
We obtained neither and report titles and citations only.
Two observations, ours.
The word romance in a 1985 journal title is a deliberate provocation, and it suggests the argument that leadership is over-credited predates the popular version by two decades.
And the second title, and the Evaluation of Organizational Performance, points at the same circularity: how performance is evaluated depends on beliefs about leadership, and beliefs about leadership depend on performance.
The Methodological Root
The paper underneath the whole argument, which is older than the book by thirty-two years.
The critic cites Staw, B. M. (1975), Attribution of "Causes" of Performance: A General Alternative Interpretation of Cross-Sectional Research on Organizations, in Organizational Behavior and Human Performance, 13, 414–432[1].
We did not obtain it and report the title only.
Two observations, ours.
The title is a complete statement of the problem: cross-sectional research on organisations has a general alternative interpretation, namely that the supposed causes are attributions made after the fact.
And it was published in 1975, seven years before In Search of Excellence. The methodological objection was on the record before the genre it undermines existed.
Forewarned Is Not Forearmed
The finding that limits what any of this can do for you.
Reference lists identify Nisbett, R. E., and Wilson, T. D. (1977), The Halo Effect: Evidence for Unconscious Alteration of Judgments, Journal of Personality and Social Psychology, 35(4), 250–256; and, four years later, The Halo Effect Revisited: Forewarned Is Not Forearmed, in the Journal of Experimental Social Psychology, 17(4), 427–439[4].
We obtained neither paper and report titles only. But the second title states a conclusion.
Two observations, ours.
If warning people does not protect them, then reading this article does not immunise you, and we would rather say that than imply otherwise.
Which points at the same conclusion the fifth article in this series reached about anchoring: where awareness does not help, the remedy has to be procedural rather than mental. You change what information reaches the judgment, not how hard you try when making it.
How To Read A Business Book
The practical translation. Ours.
Four questions to put to any claim that some practice causes company success.
Were the companies selected because they succeeded? If so, the sample is chosen on the outcome and cannot show what distinguishes success from failure.
Was the evidence about their practices gathered after the results were known? Press accounts, interviews and case studies all fail this, and that is most of the genre.
Could any company plausibly claim the opposite practice? If not, the practice cannot discriminate.
And is the claimed advantage available to competitors? If it is, it cannot produce relative outperformance once adopted.
One observation. A book failing all four may still be worth reading for its examples, its vocabulary or its provocations. What it cannot supply is evidence that following it will work.
Turning It On This Series
The obligation this article creates for us. Ours.
Thirty articles in, the same four questions apply here.
Do we select on the outcome? Partly. We choose topics where there is something interesting to say, and a finding that quietly replicated is less interesting than one that did not. That is a selection effect in our topic choice and we have not corrected for it.
Do we gather evidence after knowing the result? Structurally, yes. We read abstracts knowing what the finding was, and our reading of a method is not blind to its conclusion.
Are our criticisms falsifiable? We have tried to make them so, by stating what we could not obtain, running our own simulations where the claim was checkable, and reporting figures that cut against our own line.
And is our advice available to everyone? Yes, which by the second delusion means it cannot be a source of advantage. What it can be is a defence against paying for something that is not one.
We record this because a series that spent thirty articles applying a standard to others and never to itself would have earned the same objection.
What To Do
Assume performance shaped the description. When you read that a successful company has a strong culture, ask whether anyone would have said that had the numbers been worse.
Check whether the sample was selected on success. A study of winners cannot tell you what distinguishes winners from losers.
Ask whether any company would claim the opposite. A practice nobody disavows cannot explain differences between companies.
Discount formulas available to competitors. Performance is relative, so a widely adopted practice becomes a baseline rather than an advantage.
Ask for the base rate before believing a percentage. One third of a group beating the market sounds bad and may be exactly what random selection produces.
Do not rely on being forewarned. One paper's title states that warning does not protect, which means the remedy is procedural rather than a matter of trying harder.
Design appraisal forms accordingly. Thorndike's 1915 data suggests raters could not analyse traits out from one another, which is what a multi-scale form assumes.
Read the rigorous work anyway. The critic's own point is that better business research exists and simply does not promise a recipe.
The Limits Of This Analysis
Several caveats matter. This article reviews research methods and is not investment or management advice; nothing in it is a comment on any named company. Everything is verified to August 2026. We did not obtain the book and rely on its publisher description and on the author's own peer-reviewed journal article, of which we obtained portions. We did not obtain Thorndike's paper, only his published abstract. We found two conflicts in our sources, on the 1920 paper's page numbers and on its sample, and report both without resolving them. We did not obtain the Staw paper, either Romance of Leadership paper, either Nisbett and Wilson paper, the Cooper review, or In Search of Excellence, and report titles and citations only. Our simulation is entirely ours, uses invented parameters and a lognormal assumption with an equal-weighted index, and does not establish the real base rate in the relevant period; what it establishes is that the figure cannot be interpreted without one, which the source does not supply. The four reading questions, the analysis of why the check strengthens the thesis, and the section applying the argument to this series are our own reasoning, not findings. One source we consulted claims that higher IQ scorers were more susceptible to halo error; we could not verify it and do not report it as a finding.
Frequently Asked Questions
What is the halo effect in a business context?
Why does that undermine business research?
Did the Excellent companies really collapse?
Does that weaken the critique?
Does knowing about this protect me?
So are business books useless?
References
- Rosenzweig, P. M. (2007). Misunderstanding the Nature of Company Performance: The Halo Effect and Other Business Delusions. California Management Review, 49(4), Summer 2007, hosted copy, on the halo effect as described by Edward Thorndike in 1920 being the basic human tendency to make specific inferences on the basis of a general impression; on Peters and Waterman having studied companies including Hewlett-Packard, IBM, Johnson & Johnson, McDonald's, Procter & Gamble and 3M, gathering data from archival sources, press accounts and interviews, and identifying eight practices common to the Excellent companies including a bias for action, staying close to the customer, and stick to the knitting; on the Excellent companies having regressed sharply, with only one-third growing faster than the overall stock market over the next five years; on the promised blueprint of long-term prosperity turning out to be a delusion; and its footnotes citing Thorndike, A Constant Error in Psychological Ratings, Journal of Applied Psychology, 4 (1920): 469–477, Cooper, W. H., Ubiquitous Halo, Psychological Bulletin, 90/2 (1981): 218–244, Staw, B. M., Attribution of "Causes" of Performance: A General Alternative Interpretation of Cross-Sectional Research on Organizations, Organizational Behavior and Human Performance, 13 (1975): 414–432, and Meindl, J. R., Ehrlich, S. B., & Dukerich, J. M., The Romance of Leadership, Administrative Science Quarterly, 30/1 (1985): 78–102, with Meindl and Ehrlich, The Romance of Leadership and the Evaluation of Organizational Performance, Academy of Management Journal, 30/1 (1987): 91–109. Note: the author's own peer-reviewed journal article; we obtained portions and prefer it throughout to his book. Its page reference for Thorndike conflicts with every other source we located. cebma.org
- Publisher description for Rosenzweig, P. (2007), The Halo Effect: ...and the Eight Other Business Delusions That Deceive Managers, Free Press, on much business thinking being shaped by delusions, errors of logic and flawed judgments that distort understanding of the real reasons for a company's performance; on these affecting the business press and academic research as well as bestselling books promising to reveal the secrets of success; on such books claiming to be based on rigorous thinking while operating mainly at the level of storytelling; on the most pervasive delusion being the halo effect, under which observers conclude a company has brilliant strategy, visionary leadership, capable employees and superb culture when sales and profits are up and the reverse when performance falters, though in fact little may have changed, because company performance creates a halo that shapes perception of strategy, leadership, people and culture; on examples drawn from Cisco Systems, IBM, Nokia and ABB; on the argument undermining bestsellers from In Search of Excellence to Built to Last and Good to Great; and on the Delusion of Absolute Performance, being that company performance is relative to competition rather than absolute, which is why following a formula can never guarantee results. Note: a publisher description on a book cataloguing site, not the book. We did not obtain the book. goodreads.com
- Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29. DOI 10.1037/h0071663, published abstract via a bibliographic record, on a study made in 1915 of employees of two large industrial corporations in which the estimates of the same man in a number of different traits such as intelligence, industry, technical skill and reliability were very highly correlated and very evenly correlated; and on it consequently appearing probable that those giving the ratings were unable to analyze out these different aspects of the person's nature and achievement and rate each in independence of the others. Note: we obtained the published abstract only, not the paper. semanticscholar.org
- Reference lists identifying Thorndike, E. L. (1920), A constant error in psychological ratings, Journal of Applied Psychology, 4(1), 25–29, DOI 10.1037/h0071663; Nisbett, R. E., & Wilson, T. D. (1977), The Halo Effect: Evidence for Unconscious Alteration of Judgments, Journal of Personality and Social Psychology, 35(4), 250–256; Nisbett, R. E., & Wilson, T. D. (1981), The Halo Effect Revisited: Forewarned Is Not Forearmed, Journal of Experimental Social Psychology, 17(4), 427–439; Cooper, W. H. (1981), Ubiquitous halo, Psychological Bulletin, 90(2), 218–244; and Murphy, K. R., & Anhalt, R. L. (1992), Is halo error a property of the rater, ratees, or the specific behaviors observed?, Journal of Applied Psychology, 77(4), 494–500. Note: citations only. We obtained none of these papers and report no findings from any of them; the 1981 title is quoted because it states a conclusion. ethicsunwrapped.utexas.edu
- Commentary website describing the 1920 study as involving army commanding officers rating soldiers on several supposedly independent traits including physical appearance, intelligence, leadership ability and personal qualities, with officers who rated a soldier highly on one dimension tending to rate that soldier highly across nearly every other dimension. Note: a commentary website, not peer-reviewed. This description conflicts with Thorndike's own abstract at reference 3, which describes employees of two industrial corporations; we report the conflict without resolving it. psychologynoteshq.com
- Blog describing the 1920 study as an experiment in which commanding officers rated subordinates on aspects including physique, intelligence and leadership, with those rated positively for appearance scoring higher in unrelated realms; and describing Rosenzweig's application of the halo effect to company performance, using Cisco in the late 1990s as an example of success being attributed to corporate culture, leadership and strategy. Note: a blog, not peer-reviewed; recorded as a second source for the conflicting description of the 1920 sample. ecotalker.wordpress.com
- Encyclopedia entry on Rosenzweig's book, on it criticising pseudoscientific tendencies in the explanation of business performance; on it targeting specific books offering secrets of guaranteed business success and academic research published by business schools; on it outlining nine delusions, being mistakes of reasoning that undermine these recipes; on the argument that much business writing is what Feynman called cargo cult science, having the superficial trappings of science but operating at the level of story-telling; on the book also considering more scientific business research whose conclusions are more rigorous but do not promise a simple recipe for success; and on publication details including the February 2007 date and the 232-page length. Note: an encyclopedia entry, not peer-reviewed. We did not obtain the book and report this characterisation as the entry gives it. en.wikipedia.org
- Design research organisation article on the halo effect, describing it as a well documented social-psychology phenomenon causing people to be biased in their judgments by transferring their feelings about one attribute of something to other, unrelated attributes; and confirming the citations for Thorndike (1920) and for Rosenzweig (2007), The Free Press. Note: a professional research organisation's summary; used for the general definition and citation confirmation. nngroup.com
This article reviews research methods and is not investment or management advice. Nothing in it is a comment on any named company. The book discussed was not obtained; the author's own peer-reviewed journal article is used instead and was obtained only in part. Thorndike's paper was not obtained, only its abstract, and two conflicts between sources on that paper are reported rather than resolved. The simulation checking the central statistic is the authors' own, uses invented parameters, and does not establish the real base rate for the period.