The fiftieth article looked back at the first half of this series[1]. This one looks back at all of it, and applies the same treatment we have given to Maslow, to Dunbar, to the learning pyramid and to ninety-six other things.

Key Takeaway

Zero of the first fifty articles reached 8,000 words. Forty-eight of the last forty-nine did. The standard this series enforces was adopted halfway through and applied to nothing before it, which means the earlier half is measurably weaker by our own rule and we did not go back and fix it. That is the most important finding in this audit, and it is about us.

What This Article Is

Not a summary. Ours.

Four observations.

A summary of ninety-nine articles would be useless, since each already has a verdict stated in its first section and a list of what to do at the end.

What is worth doing instead is an audit: measuring the series as an object, the way we have measured other people's work.

Every figure below was computed from the published files[6], not from our recollection, which is the discipline the ninety-ninth article had to invoke on itself after we miscited our own article number.

And the audit is unflattering in places. We have reported those places rather than the flattering ones only, which is the whole argument of the series applied to itself.

Two reasons that is worth doing rather than performing, ours. A publication that never reports against itself has given a reader no way to calibrate it, and every article here has asked readers to calibrate somebody else's work.

And the counts are checkable. Anybody with the published files and a text editor can reproduce every figure in this article, which is the standard we have asked of the studies we examined.

The Series, Measured

Our own count, taken from the published files in August 2026.

Ninety-nine articles. 657,558 words. A mean of 6,642 words each, a shortest of 4,159 and a longest of 9,689.

746 reference entries, a mean of 7.5 per article.

The phrase "did not obtain" appears 1,073 times. The phrase "our own arithmetic" appears 434 times. Sources are explicitly flagged as not academic 26 times in reference notes.

Four observations.

657,558 words is longer than most reference works, and we note it not as an achievement but because length is a cost borne by readers.

The mean of 6,642 conceals the thing this audit is really about, which the next section takes apart.

1,073 admissions of not having obtained something is roughly eleven per article, which is either scrupulous or an indictment depending on how you read it, and we think it is both.

It is scrupulous because the alternative was to say nothing and nobody would have known, and an indictment because a series that had obtained more would have needed to say it less.

And 434 instances of flagging our own arithmetic is the count we are most comfortable with. Every number we generated is marked as ours, which was the rule from the first article and appears to have held.

We would qualify that slightly, ours, since a count cannot establish it. 434 occurrences of the phrase does not prove that no unmarked calculation slipped through, only that the practice was heavily used.

The Standard Rose Halfway Through

The finding that matters most, and it is about our own inconsistency. Our own count.

Mean word count by quarter of the series: articles 1 to 25 average 5,660 words. Articles 26 to 50 average 4,913. Articles 51 to 75 average 8,010. Articles 76 to 99 average 8,039.

And the cleanest form of it: 48 of the 99 articles reach 8,000 words. Zero of them are in the first fifty.

Four observations.

The eight-thousand-word floor was adopted around the fifty-first article and has held essentially without exception since.

The second quarter is the weakest of the four, at 4,913 words, which is below even the first, so the trend was not steady improvement.

We have no account of why it dipped, ours, and would rather say that than construct one. A plausible story is available and we have not tested it, which is the position several articles here criticised other people for filling in.

The change is a step, not a slope, which tells you it was a decision rather than a drift.

We can date it fairly precisely from the files, ours. No article before number 51 reaches the floor and almost every one after it does, which is the signature of a rule being adopted rather than of gradually improving practice.

And a decision to raise a standard is a good thing that has an uncomfortable corollary. Everything published before it was published below the standard we now enforce, and we left it there.

Two defences are available and neither is complete, ours. Rewriting fifty articles would have cost the time of roughly twenty new ones, which is a real trade-off.

And a reader can see the publication dates, so nothing is concealed. But the articles carry no notice saying they predate the standard, which they easily could.

Which Means The First Half Is Worse

Stated plainly, because the previous section could be read as a boast. Ours.

Four observations.

By our own current standard, forty-nine articles in this series are inadequate. Not wrong, not badly sourced, but below the length at which we now think a topic can be treated properly.

We would defend some of that. Length is a proxy and a poor one, and a 5,000-word article on a simple finding may be complete where an 8,000-word one on a complex finding is not.

But we do not get to have it both ways. If length is a poor proxy, then the floor we enforce on every recent article is arbitrary, and if the floor is right then the early work is short.

Our honest position, ours: the floor exists because it forces us to find the third and fourth thing worth saying, and the early articles frequently stop after the second. That is a real deficiency and we did not go back.

One thing the floor did not fix is worth naming too, ours. Length forces breadth and does not force depth, and a long article can be a short one with more examples, which some of ours are.

Seven And A Half References

Our own count, and the picture is more complicated than a mean suggests.

746 reference entries across 99 articles, a mean of 7.5. By quarter: 7.9, then 7.5, then 9.1, then 5.5.

Four observations.

A reference entry in this series is not a citation in the usual sense. Each carries a note describing what the source is, what we took from it, and what we did not, which is why they run to a paragraph.

Those notes are where the sourcing discipline actually lives. A reference list without them would report the same 746 items and disclose nothing.

Which is a small transferable idea for anybody who cites things, ours. A citation tells a reader where you looked; a note tells them what you found and what you did not, and only the second is useful.

The rise to 9.1 in the third quarter reflects articles built from many partial sources, which is what you do when no single source is good.

That quarter also carries the highest "did not obtain" rate at 14.9 per article, ours, and the two facts are the same fact. More sources and more admissions of not having them describe a period of assembling findings from fragments.

And the fall to 5.5 in the last quarter is the part we examine next, because it looks like a decline and we do not think it is.

The Reference Count Fell

Examined honestly, with both readings given. Ours.

Four observations.

The pessimistic reading is available and we will state it first. Fewer references means less checking, and a series that gets tired would look exactly like this.

The alternative reading is that the source quality rose. The last quarter includes two primary documents obtained in full: a first author's own repudiation of her famous paper[2], and the man who named the Pareto principle explaining that he named it wrongly[3].

An article built on a document you have read in full needs fewer supporting sources than one built on an abstract and five partial confirmations, which is the honest mechanism.

We cannot fully distinguish the two readings from a count, ours, and a reader is entitled to the less charitable one. What we can point to is that the "did not obtain" rate also fell in that quarter, from 14.9 per article to 8.8, which is consistent with having obtained more.

One Thousand And Seventy-Three

The most characteristic number in the series. Our own count.

The phrase "did not obtain" appears 1,073 times across ninety-nine articles, roughly eleven times each, peaking at 14.9 per article in the third quarter.

Four observations.

That number is the honest measure of what this series is. It is largely built on abstracts, and it says so, repeatedly, in every article.

The alternative was available and we declined it. An article can be written from an abstract without mentioning that, and most business writing about research is.

It has a cost we should name. An abstract is a summary written by authors to attract readers, and building on one means inheriting whatever emphasis they chose.

And it sets the ceiling on what any of these articles can claim. We can report what a paper says about itself and cannot assess whether it is right, which is a much weaker position than the confident tone of a business article usually implies.

That mismatch between tone and evidential position is the thing we would most want to fix, ours. The verdict-first format this series uses reads as authority, and it sits on top of a sourcing base that is frequently a single abstract.

We chose it because a reader deciding something needs an answer near the top. The cost is that the format is more confident than the evidence, and the limits section at the end is a poor remedy for that.

Fifteen Calculations We Threw Away

The practice we are proudest of, and it is measurable. Our own count.

Fifteen articles report a calculation we attempted and withdrew: numbers 1, 4, 11, 21, 40, 50, 55, 61, 77, 84, 88, 90, 92, 94 and 99.

Four observations.

Each withdrawal is reported in its own section in the article body, with the reason, rather than in a footnote or not at all.

Placement was a deliberate choice, ours. A withdrawal in a footnote is a disclosure nobody reads, and one in the body interrupts the argument, which is the point of it.

The reasons cluster. Most were withdrawn because completing them required inventing data: a joint distribution nobody has, a contingency table we did not possess, a prior nobody has established.

One was withdrawn for a different and more embarrassing reason. A table in the ninetieth article had two identical columns because one had been computed from the other, which is an error rather than a limit, and we said so.

And fifteen out of ninety-nine is a rate of about one in seven, ours. We do not know whether that is high or low, because nobody else reports this, which is rather the point.

What we can say is that the rate did not fall, ours. Five of the fifteen are in the last twenty-five articles, so the practice held rather than lapsing once it stopped being novel.

Fifty-Four Variants

The running count, and what it turned out to mean. Our own tally.

Across ninety-nine articles we recorded fifty-four bibliographic variants: the same paper cited with different years, volumes, page ranges, article numbers, titles or author orders by different sources.

Four observations.

We did not go looking for a single one of them. They turned up because we checked more than one record for each paper, which we did to confirm abstracts rather than to audit citations.

Roughly one every two articles is a rate we found genuinely surprising. We expected the occasional typo and found a systematic condition.

The worst were not the obvious ones. A wrong page number sends you to the right paper; a wrong article number in an electronic-only journal sends you nowhere, and a reversed author order defeats a search entirely.

The pattern in which ones are dangerous is worth carrying, ours. An error that fails loudly costs a reader a minute; one that fails silently costs them the paper, and the silent ones are the ones nobody notices are errors.

And the one we would single out is from the ninety-fourth article. A researcher miscited the replication of her own study, in the document repudiating her own study, which tells you this is not about carelessness.

The generalisable form is simple and we would keep it, ours. Chase the identifier rather than the year, since a DOI resolves to one item and a year, a volume and a page range each fail differently.

The First Pattern

What recurred across the ninety-nine, stated as we found it. Ours.

A qualification published in the source and absent from the reception.

Four instances, all from the last dozen articles.

Mehrabian's scope condition was published by the author himself: unless a communicator is discussing feelings or attitudes, the equations do not apply. The 7-38-55 rule travels without it.

Dunbar's interval was 100 to 230, and its author called the result exploratory because of the large error measure[4]. Only the point estimate travelled.

The marshmallow finding included a sentence saying most of the variation came from being able to wait at least twenty seconds. Both the enthusiastic and the debunking receptions dropped it.

And Juran spent from 1950 to 1974 explaining that he had misnamed the Pareto principle[3], and wrote that the confession changed nothing.

What unites the four is not that anybody was dishonest, ours. In every case the careful version was published and the careless version travelled, and the researchers are on the losing side of that as much as anybody.

The Second Pattern

The one that surprised us most. Ours.

The numbers were frequently added by somebody other than the researcher, and often were never measured at all.

Four instances.

Maslow never drew the pyramid, on a peer-reviewed abstract's flat statement, and the triangular form asserts that the base is nine times the apex, which no list of needs claims.

The learning pyramid's percentages are all multiples of five across three versions that disagree with each other about how many tiers exist.

Juran's own statement of his principle contains no numbers: a relative few of the contributors account for the bulk of the effect. Eighty and twenty were attached later.

And the specific figures everybody quotes about the pioneer advantage paper appear nowhere in its abstract, and we could not verify them.

The common structure is worth stating, ours. A qualitative finding is more defensible and less usable, so somewhere between the paper and the boardroom a number gets attached, and the number is what licenses a decision.

The Third Pattern

The methodological one, and the most useful commercially. Ours.

Who assigned the categories, and could the assignment depend on the outcome?

Four instances.

The pioneer advantage literature asked surviving firms whether they had pioneered, and on our own model a twenty percent over-claim rate produces a two-and-a-half-fold advantage from an association of exactly zero.

Personality typologies cut a continuum and then treat the sides as kinds, which on our own arithmetic reassigns one person in seven at an excellent reliability.

The open plan study is the counter-example that proves it. It used wearable sensors precisely so that nobody had to be asked, and found the opposite of what surveys report[5].

And the general test is one sentence. Could a respondent's answer have been different if their outcome had been different?

We would rank this the most practically useful of the four patterns, ours, because it applies to material a business owner reads weekly: benchmark surveys, award lists, case studies and any research where firms described themselves.

The Fourth Pattern

The one about why corrections lose. Ours.

The properties that make a claim usable are the properties that make it unreliable.

Four observations.

A point estimate is repeatable and an interval is not. A category is actionable and a score is not. A round number is memorable and a measured one is not.

One of the researchers we quoted said it better than we have. The number is widely spread because it is easy to understand, and the claim that no number can be calculated is not quite as entertaining[4].

This is not a story about the public misunderstanding science. That same quotation notes the number circulates among researchers too.

And it explains the shape of this whole series, ours. Every article is longer than the finding it examines, because the finding is a sentence and the reason it was misread is not.

The Anatomy Of A Business Belief

Assembling the four patterns into one account, because they turn out to be stages. Ours.

Four stages.

A researcher publishes a careful, qualified finding in a specific domain, with an interval, a scope condition or an explicit limitation.

Somebody writing for practitioners strips the qualification, because it does not fit the format, and frequently attaches a number the original did not contain.

The numbered version acquires a shape or a name, which makes it repeatable and gives it an authority the underlying work would not have claimed.

And the correction, when it comes, cannot compete, because it is longer, less usable, and offers nothing to replace what it removes.

We would note that no stage requires anybody to act badly, which is why the pattern is so robust. Each step is a reasonable thing for the person taking it to do.

What Being Careful Costs

An observation about why this is hard, rather than a complaint. Ours.

Four observations.

The careful version of any claim in this series is longer, weaker and harder to act on than the version it replaces, without exception.

That is not a rhetorical failure on the careful side. It is what accuracy costs, since a claim with its conditions attached applies to fewer situations and says less about each.

Which means a manager choosing the careful version is accepting a real reduction in usable guidance, and pretending otherwise would be dishonest.

What we would say in its favour is narrow and we think sufficient, ours. Less guidance that is right beats more guidance that is wrong, but only if you actually then do the work the guidance no longer does for you.

And that condition is the one most readers will fail, ours, including us. Removing a false rule and replacing it with nothing leaves a firm worse off than it was, which is why every article here ends with something to do rather than something to stop.

What We Got Wrong

The section this article exists for. Ours.

Four admissions.

We raised our own standard halfway through and did not go back. Forty-nine articles sit below the floor we now enforce, and republishing them properly would have been the consistent thing to do.

We have built too much on abstracts. Eleven admissions per article of not having obtained something is honest and is also a description of a limitation we mostly accepted rather than fixed.

We miscited our own series, in the ninety-ninth article, referring to article 85 as article 81, and caught it only by checking the file.

And we have used invented figures more than we would like, which the section after next examines properly rather than in a clause.

A fifth admission belongs here and is harder to state, ours. We do not know how many errors remain undetected, and the fifteen withdrawals and three corrections tell you only about the ones we caught.

The Corrections We Published

Because the record should include them. Ours.

Four observations.

The most substantial was a correction to our own back catalogue on loss aversion, published as an article rather than as an edit, after we concluded that earlier articles in this series had overstated a finding.

The ninetieth article withdrew a table mid-build because two of its columns were identical, one having been computed from the other, and reported the tell in the body: matching columns mean one came from the other.

The ninety-ninth corrected its own cross-reference in a numbered section rather than editing it out.

And the reason we publish these is not virtue, ours. A reader who finds an error we disclosed learns that we disclose them, which is the only evidence about a publication that is worth anything.

The reverse also holds and is worth saying, ours. A publication with no visible corrections has either made none or reported none, and a reader has no way to tell which.

The Findings That Held

Because a series that only ever found things wanting would be as unreliable as one that never did. Ours.

Four observations.

The self-report effect in power posing replicated. People who adopt expansive postures do report feeling more powerful; the hormonal claim is what failed.

Unequal contribution is real. Juran's qualitative principle, that a relative few contributors account for the bulk of an effect, is not in dispute; the specific ratio and the action it licenses are.

Hemispheric specialisation is real, affirmed in the first sentence of the paper that demolished the popular version. Functions are lateralized; people are not.

And active learning probably does beat passive learning. The ordering in the learning pyramid is not seriously disputed. It is the magnitudes that were never measured, and the magnitudes are the only part a budget can use.

The shape those four share is worth naming, ours. In each case a real phenomenon acquired a false precision, and stripping the precision leaves something true and much less useful.

Four Hundred And Thirty-Four Calculations

What we actually contributed, as distinct from what we reported. Ours.

The phrase "our own arithmetic" appears 434 times across the series.

Four observations.

Almost none of it is difficult mathematics. It is division, compounding, standard statistical identities and small simulations, all of which a business owner with a spreadsheet could reproduce.

That was deliberate. An argument a reader cannot check is a claim to authority, and the series has tried to avoid making any.

Several calculations produced results we did not expect, which is the only real evidence that the arithmetic was doing work. The self-similarity claim in the Pareto article checked out exactly when we expected to find an error, and the survivorship decomposition in the ninety-eighth explained only about half the gap when we expected it to explain all of it.

And every one of the 434 is marked as ours in the sentence that carries it, so no reader should ever have mistaken one for a published finding.

That mattered more than any other rule we adopted, ours. The single most damaging thing a business article can do is present its own construction as somebody's research, and we would rather be tedious about the distinction than risk it.

The Invented Figures Problem

The methodological weakness we would most want a critic to attack. Ours.

Four observations.

A great many of our calculations rest on invented business figures: a firm of twenty, a rent per square foot, a loaded hourly cost, a fixed cost to serve.

The defence is that the structure survives any reasonable substitution, and we have tried to say so each time, and frequently to show where the answer flips.

The honest criticism is that invented figures can be chosen to produce a conclusion, and a reader has only our word that they were not.

What we would offer against that is the break-even form, ours, which several later articles adopted. Reporting the value at which a decision reverses is harder to rig than reporting a result, because the reader supplies the comparison themselves.

Two examples of it working, ours. The Pareto article's break-even is simply the contribution margin, which any accountant would recognise and could have derived without us.

And the open plan article's threshold is a third of loaded hourly cost, which asks a reader only whether an hour of their staff talking is worth more or less than that.

The Correction We Published As An Article

The largest of our own corrections deserves its own section rather than a mention. Ours.

Four observations.

Partway through the series we concluded that earlier articles here had overstated the loss aversion literature, and published that conclusion as a numbered article rather than editing the originals.

That choice has a cost we should name. A reader who finds one of the earlier articles first gets the overstated version, and there is no notice on it pointing forward.

It also has a benefit that we think is larger. An edit is invisible and an article is not, so the correction reached the people who follow the series rather than only those who happened to reread an old piece.

And the honest position is that we should have done both, ours. Publish the correction and mark the originals, which is what the running errata we recommend to ourselves below would have supplied.

Two things that episode taught us about our own process, ours. The error was ours and had been repeated across several articles, which is how an assumption behaves once it is in a series.

And we found it by writing another article on an adjacent topic, not by reviewing. Nobody was auditing the back catalogue, and nobody is now.

Which is the argument for the running errata in one sentence, ours. Errors are found by accident unless somebody is looking on purpose, and across a hundred articles nobody was.

What We Did Not Cover

The coverage gap, stated so a reader is not misled by the number one hundred. Ours.

Four observations.

One hundred articles sounds comprehensive and is a small fraction of the relevant literature. We covered what we could find good sources for and what seemed commercially consequential.

That selection has a shape. We are drawn to findings that failed, which is more interesting to write and produces a distorted picture of the field.

A series with the opposite selection would be equally misleading and would be easy to write. Plenty of behavioural findings have replicated well, and a reader who took this series as a survey of the field would badly underrate it.

And there are whole areas we did not touch, ours. Nothing here on negotiation, almost nothing on pricing psychology, and nothing on the economics of attention, all of which matter to a business and none of which we examined.

A reader wanting a complete picture should therefore treat this as one input, ours. Ninety-nine articles selected for having something wrong with them is not a survey, and we would not want it read as one.

What To Distrust In This Series

Written for a critical reader, which is the reader we want. Ours.

Four things.

Distrust every dollar figure. Almost all of them are invented, and while we mark them, a reader skimming will absorb the number and not the flag.

Distrust our characterisations of prior literature. Where a paper criticises earlier work, we have generally reported one side's account of the other side and did not read the other side.

Distrust the tone. A verdict stated first reads as confidence, and the confidence is often carried by an abstract we could not get behind.

And distrust the pattern-finding above all, including the four patterns in this article. We chose the ninety-nine topics and are now reporting what they have in common, which is a finding about our own choices at least as much as about the literature.

That is the ninety-eighth article's mechanism applied to us, ours. We classified the cases and are now reporting the association, which is exactly the design we criticised, and we have no defence against it except naming it.

Why Stop At One Hundred

Because a round number is exactly the thing this series has spent ninety-nine articles being suspicious of. Ours.

Four observations.

One hundred is not a natural stopping point, and we would be embarrassed to defend it as one after the ninety-fifth article's argument about round figures.

The honest reason is that it was set as a target at the beginning, and targets shape behaviour whether or not they are principled, which is a finding several articles here examined.

A better rule would have been to stop when the topics stopped being interesting, ours. By that rule we would have stopped somewhere in the eighties or continued past a hundred, and we cannot tell you which.

And we record the inconsistency rather than dressing it up. The series ends on a round number for the same reason most things end on round numbers, which is that somebody chose it in advance and then met it.

One thing in its favour, ours, and it is the same thing article ninety-four found. A target set in advance and then met is at least not a target adjusted after seeing the results, which is the worse version of the same problem.

The Eight Tests

What ninety-nine articles distil to, if a reader keeps nothing else. Ours.

Ask what the interval is. A point estimate without one is a claim without a strength, and the interval is dropped first at every step of transmission.

Ask who assigned the categories, and whether the assignment could depend on the outcome being studied.

Ask who is missing, and whether they are missing for a reason connected to what you want to learn.

Ask what the qualification was. Almost every finding we examined carried one in the original, and almost none of them travelled.

Look at the last digit. A full set of round numbers means somebody chose them rather than measured them.

Ask what the diagram is claiming. A shape asserts magnitudes and orderings that the words beside it may not.

Move to each end of the range and see whether your decision changes. If it does, the estimate did not support the decision.

And measure your own firm. A noisy measurement of your own people beats a clean published figure about somebody else's, because yours is the firm you are deciding about.

What Actually Survives

Our reading of the whole series, stated directly.

Five statements.

Most of the famous business findings we examined were not fabricated. They were real results generalised past what they supported, which is a different and more difficult problem.

The qualification was usually published. In most cases the researcher said the careful thing and the careful thing did not travel.

The numbers were frequently added later, by somebody writing for business rather than by the researcher.

The design fault was usually about who got counted, rather than about how they were measured.

That last one surprised us most, ours, since we began expecting to write about measurement error and ended up writing mostly about sampling and classification, which are the faults nobody teaches business students to look for.

And the corrections lose because they are less usable, which is a property of how information travels and not of anybody's honesty.

Those five are what we would defend across the whole series, ours, and the first is the one most likely to be misread. This was not an exercise in debunking, and a reader who takes it as evidence that management research is worthless has drawn the wrong conclusion from it.

What We Would Do Differently

If we started again. Ours.

Four things.

Set the length standard at the first article rather than the fifty-first, or abandon it and defend a different one, but not both in the same series.

Obtain more full texts, which would mean fewer articles, and we think that trade would have been worth making.

Report a running errata rather than correcting inside individual articles, so a reader can find every correction in one place instead of discovering them.

And pick fewer topics where the answer was predictable. Several of these articles found what anybody familiar with the replication literature would have expected, and the interesting ones are the ones where the arithmetic went somewhere we did not plan.

A fifth thing, and it is the one we would most want to have done, ours. Test something ourselves. Ninety-nine articles of examining other people's measurements and not one original measurement of our own is a real gap.

A Note To The Students

Part of the stated audience for this publication, and the group this series was most useful for. Ours.

Four observations.

Most of what you will be taught about management contains at least one of the four patterns above, and your lecturers will frequently not know which, because they learned it the same way.

The remedy is not scepticism, which is cheap and produces nothing. It is the habit of finding the original, which is usually one search away and almost never says what the textbook says it says.

You will find that the original is more careful, more qualified and more interesting than the version you were given, which is the consistent experience of writing these ninety-nine.

And you will occasionally find that the original does not exist, as with a training chart whose research was reportedly lost, which is a thing worth knowing about your own field.

One practical suggestion, ours, and it costs an afternoon. Take one claim from your current course and chase it to its source, all the way, and write down what you find at each hop.

The Thing We Could Not Count

The limitation this whole audit cannot reach. Ours.

Four observations.

Everything measured above is a property of text: length, reference counts, phrase frequencies, withdrawal counts.

None of it establishes the thing that matters, which is whether any individual article is correct. A long, heavily referenced, scrupulously flagged article can be wrong throughout.

We have no way to assess that from inside, ours, and neither does any publication about itself. The only test is somebody outside checking, which is what we have asked of every source we examined.

So the honest closing position is this. We have shown our working, disclosed our limits, and marked what is ours, which makes the series checkable and does not make it right.

That is a smaller claim than a hundred articles might seem to warrant, ours, and it is the one the evidence supports.

If You Read Only One Section

The practical residue of 657,558 words. Ours, and not advice on any specific decision.

Four points.

Almost nothing in the management literature is precise enough to set a threshold with. Use it to generate hypotheses about your own firm, not to set caps, ratios or rules.

Your own data is worse than published research and more relevant, and the second usually wins. Forty of your own customers beat a study of five hundred brands for deciding about your customers.

With one caution the eighty-eighth article established, ours. Forty of anything is a small sample and will be noisy, so the right posture is to look repeatedly rather than to act on one reading.

Write down the decision rule before you look, whether that is a sample size, a stopping point, a definition or a threshold. Almost every fault in this series traces to a decision made after seeing the data.

And say the real reason. The open plan article's finding is only embarrassing to a firm that claimed collaboration; one that said cost had nothing to be embarrassed about.

That last point is the one with the widest application, ours. A justification chosen for how it will land commits you to a claim you will never test, and most of the beliefs in this series survived on exactly that mechanism.

The Limits Of This Analysis

Several caveats matter, and they are unusual for this article because its subject is itself. Every figure here was computed from the published files in August 2026 and is a count of text rather than a measure of quality: word counts, reference-entry counts and phrase frequencies say something about form and very little about whether any individual article is right. Word count is a poor proxy for thoroughness, as we say in the body, and our use of it to criticise our own first half is therefore itself open to criticism. The quarter-by-quarter comparisons are arbitrary groupings we chose, and different boundaries would give different means, though the step at article fifty-one is large enough to survive any grouping. The count of fifteen withdrawn calculations comes from searching for the word withdraw and its variants, so it may include articles that use the word for another purpose and may miss withdrawals described differently; we did not read all ninety-nine to check. The tally of fifty-four bibliographic variants is our own running count kept during writing, not an independent audit, and nobody has verified it. The four patterns we identify are our own reading, selected by us from our own work, which is the most obvious selection problem available and one a reader should weigh heavily: we chose the topics, we chose what to emphasise in each, and we are now reporting the pattern in our own choices. Our account of why the reference count fell in the last quarter offers two readings and cannot distinguish them. The list of findings that held is not exhaustive and was assembled from memory of the series rather than by systematic search. And this article, like the ninety-nine before it, is not advice on any specific business, financial, legal or medical decision.

Frequently Asked Questions

What is the single most important finding across the series?
That most famous business findings were not fabricated but generalised past what they supported, and that the qualification limiting them was usually published by the researcher and simply did not travel. That pattern appeared in a large majority of the articles here.
Is the first half of this series worse than the second?
By our own current standard, yes. Zero of the first fifty articles reach 8,000 words and 48 of the last 49 do. The floor was adopted around the fifty-first article and we did not go back and rewrite what came before, which is an inconsistency we report rather than defend.
How much of this is built on abstracts rather than full papers?
A great deal. The phrase "did not obtain" appears 1,073 times across ninety-nine articles, roughly eleven per article. That is honest disclosure and it is also a real ceiling on what any of these articles can claim: we can report what a paper says about itself, not assess whether it is right.
What are the fifty-four bibliographic variants?
Instances where the same paper is cited with different years, volumes, page ranges, article numbers, titles or author orders by different sources. We did not go looking for them; they turned up because we checked more than one record per paper, at a rate of roughly one every two articles.
Did anything survive intact?
Yes. The self-report effect in power posing replicated. Unequal contribution is real as a qualitative claim. Hemispheric specialisation is real. Active learning probably does beat passive learning. In each case what failed was a specific magnitude, mechanism or ratio attached to a real underlying phenomenon.
Why does this series use so many invented figures?
Because published findings rarely come with the business quantities a decision needs. The weakness is that invented figures can be chosen to produce a conclusion, and a reader has only our word that they were not. The break-even form adopted in later articles is harder to rig, because it reports the value at which a decision reverses and lets the reader supply the comparison.
What should a business owner take from all of it?
That almost nothing in the management literature is precise enough to set a threshold with, that your own data is worse than published research and more relevant, that decision rules should be written down before you look, and that the real reason for a decision should be stated out loud rather than replaced with a more popular one.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This is the hundredth and final article in the Behavioural Finance and Psychology series. It reports that forty-nine of its predecessors fall below the standard the series now enforces, which is the most useful thing an audit of one's own work can find.

References

  1. Article 50 in this series, Fifty In: What The First Half Actually Found, published in this publication, being the halfway audit of the first fifty articles. Note: our own earlier work, cited as the precedent for this article's form. At 4,159 words it is the shortest article in the series and sits below the standard adopted immediately after it. gshfinancial.com
  2. Article 94 in this series, The Author Wrote It Down: Power Posing, which obtained in full a document published by the first author of the 2010 power posing paper on her own faculty page, stating that she does not believe the effects are real and listing the procedures used, including running subjects in chunks and checking the effect along the way. Note: our own work, cited for the primary document it obtained. The underlying source is a self-published statement by the researcher, not a peer-reviewed paper. gshfinancial.com
  3. Article 96 in this series, The Bottom Eighty Percent Pay The Rent, which obtained in full Juran's 1974 paper The Non-Pareto Principle; Mea Culpa, in which he states that he mistakenly applied the wrong name to the principle, that this confession changed nothing, that the first exposition of the universal was his own, and that Pareto's models were not intended to be applied to other fields. Note: our own work, cited for the primary document it obtained. That document is a self-archived first-person recollection written decades after the events it describes. gshfinancial.com
  4. Article 93 in this series, The Interval Nobody Quotes: Dunbar's Number, which records the 2021 reanalysis reporting 95 percent confidence intervals of 4 to 520 and 2 to 336 and concluding that specifying any one number is futile, the encyclopedia account that the 1992 original predicted 148 with an interval of 100 to 230 and was considered exploratory by its author, and an author's remark that the number spreads because it is easy to understand while the claim that no number can be calculated is not quite as entertaining. Note: our own work. The 1992 figures within it reach us through an encyclopedia and were not verified against the paper, which that article flags as its weakest link. gshfinancial.com
  5. Article 99 in this series, They Stopped Talking To Each Other, which records the 2018 field studies finding that face-to-face interaction decreased by approximately 70 percent in two corporate headquarters after they moved to open plan, measured using wearable devices and electronic communication servers rather than surveys. Note: our own work. That article obtained the study's abstract only, not its full text, and did not survey the wider literature on office design. gshfinancial.com
  6. The complete Behavioural Finance and Psychology series index, listing all one hundred articles with their categories and summaries. Every count in this article was computed from these published files in August 2026: 99 articles, 657,558 words, 746 reference entries, 1,073 occurrences of the phrase "did not obtain", 434 occurrences of "our own arithmetic", and 48 articles at or above 8,000 words, none of them in the first fifty. Note: our own published corpus, and the source of every figure in this audit. These are counts of text and say nothing about whether any individual article is correct. gshfinancial.com

This article is an audit of this publication's own series and is not advice on any specific business, financial, legal or medical decision. Every figure in it was computed from the published files and is a count of text rather than a measure of quality. The four patterns identified are the authors' own reading of their own topic selection, which is a selection problem a reader should weigh heavily.