The eighty-sixth article was about measuring a change without a comparison, and the eighty-seventh about what promotion costs. This one is about a measurement that a great many small businesses are told to take, and about what it can and cannot tell them.

Key Takeaway

A 2007 replication in the Journal of Marketing, using 21 firms and over 15,500 interviews, reports that "the research fails to replicate his assertions regarding the 'clear superiority' of Net Promoter."[1] On our own arithmetic, four wholly different customer bases all score exactly zero, and at 100 responses a small business cannot reliably detect a real half-point improvement with any metric.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The superiority claim failed a published replication. Twenty-one firms, over fifteen thousand interviews, in the discipline's leading journal, using the industries the original author cited as exemplars.

Two. The academic assessment is negative but not unanimous, and a 2021 review states both halves of that carefully, which we quote rather than select from.

Three. The metric discards information by construction. On our own arithmetic four very different customer bases produce identical scores, and the mean of the same answers distinguishes them.

Four. The bucketing also costs statistical power, though on our own simulation the loss is consistent and modest rather than dramatic, and we report the modest figure.

Five. And the finding that should concern a small business is neither of those. At the sample sizes small firms actually collect, no version of this measurement reliably detects a realistic improvement, and that is true of the mean as well.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the 2007 replication, from an abstract obtained verbatim from the publisher, though our reproduction truncates mid-word at the end.

Grade A for the 2021 assessment, from body text obtained verbatim from the publisher, including the quotations it carries from two other papers.

Grade A for the 2008 clarification, from an abstract obtained verbatim, which is the source that keeps this article honest.

Grade C for the cross-cultural point, from a journal article's discussion rather than from a tested finding we obtained.

Grade D for the supportive 2021 study, which reaches us through a non-academic blog's characterisation.

Grade A for our own arithmetic, which is exact, includes a four-thousand-trial simulation per row, and contains one comparison we made and withdrew, documented below.

A Note On Method

Everything here is verified to August 2026.

We obtained the 2007 abstract verbatim from the publisher[1]. Our reproduction cuts off mid-word in its final sentence, which we flag where it arises.

We obtained body text of a 2021 review verbatim from its publisher[2], including quotations it carries from two other papers we did not obtain.

We obtained the 2008 clarification's abstract verbatim[3], and citation confirmations from two reference lists[4].

One source is a non-academic blog[6], flagged at every use and cited only for a characterisation of a paper we did not obtain.

We did not obtain the 2003 original, the 2007 replication itself, or any of the studies discussed within our sources.

All arithmetic is ours; the distributions are invented, and the power figures come from four thousand simulated trials per row. This article discusses research on customer metrics and is not marketing advice.

The Claim

What is being asserted, as the replicating authors state it.

Their abstract records: "Managers have widely embraced and adopted the Net Promoter metric, which noted loyalty consultant Frederick Reichheld advocates as the single most reliable indicator of firm growth compared with other loyalty metrics, such as customer satisfaction and retention."[1]

The source of the claim: Reichheld, F. F. (2003), The One Number You Need to Grow, Harvard Business Review, 81(12), 46–55[4]. We did not obtain it.

Four observations, ours.

The claim is comparative, not absolute. It is not that the metric is useful but that it beats satisfaction and retention, which is a much stronger and more testable proposition.

The word "single" is doing a great deal of work, and the article's title makes it explicit: one number, and by implication you may stop collecting the others.

The venue was a business magazine rather than a peer-reviewed journal, which is not a criticism and is worth knowing when weighing it against what follows.

And the authors describing it are the ones who tested it, so this is a characterisation by people with a position, which we flag, though their description matches the original article's own title.

Two reasons we are content to rely on it anyway, ours. The 2003 article is titled "The One Number You Need to Grow," which asserts singularity in five words and needs no interpretation.

And the same authors are on record elsewhere defending the metric against a different critique, which we cover below. They are not people who set out to find against it, and that materially raises what their characterisation is worth.

The Replication

The source that tested it.

Keiningham, T. L., Cooil, B., Andreassen, T. W., and Aksoy, L. (2007), A Longitudinal Examination of Net Promoter and Firm Revenue Growth, Journal of Marketing, 71(3), 39–51[1].

Its design, from the abstract: "This article (1) employs longitudinal data from 21 firms and 15,500-plus interviews from the Norwegian Customer Satisfaction Barometer to replicate the analyses used in Net Promoter research and (2) compares Reichheld and colleagues' findings with the American Customer Satisfaction Index."[1]

Four observations, ours.

Longitudinal is the important word. The claim is that the metric predicts growth, which is a claim about the future, and only a study following firms over time can test it.

Over 15,500 interviews across 21 firms is a substantial base, drawn from a national satisfaction barometer rather than assembled by the authors, which removes one obvious avenue for selection.

The comparison against a second national index is a further check. Two independent data sources rather than one, which is more than most replications in this series managed.

And the journal is the leading peer-reviewed outlet in its field, which is the appropriate venue for testing a claim that was made in a business magazine.

One asymmetry worth naming, ours, because it recurs throughout this series. The claim reached millions of managers through a business magazine and the test reached a few thousand academics through a journal.

That is nobody's fault and it is the reason a correction rarely catches its original. The two publications are not competing for the same readers, and only one of them was ever going to shape practice.

What It Found

The result, and the abstract is careful about scope.

It records: "Using industries Reichheld cites as exemplars of Net Promoter, the research fails to replicate his assertions regarding the 'clear superiority' of Net Pro", at which point our reproduction is cut off mid-word.[1]

Four observations, ours.

The test was run on the industries the original author himself cited as exemplars, which is the strongest possible ground for the claim and the fairest place to test it.

The phrase "fails to replicate his assertions regarding the clear superiority" is specific about what failed. Not that the metric is worthless, but that its claimed advantage over other measures was not reproduced.

That distinction matters and gets lost in both directions. An advocate can say the metric still works; a critic can say it was debunked, and the finding supports neither reading cleanly.

And we did not obtain the paper, only this abstract, so we report no effect sizes, no correlations and no comparison figures from it.

One consequence of that gap, ours. We cannot tell you how close the metric came, and the difference between narrowly failing to beat satisfaction and being useless is large.

The abstract's own wording suggests the former rather than the latter. What failed was an assertion of clear superiority, and a metric can fail that test while remaining perfectly serviceable.

The 2021 Assessment

The state of the field, from a review in a leading marketing journal, quoted at length because selecting from it would misrepresent it.

It records: "As well as the methodological critiques cited above, empirical studies aiming to replicate Reichheld's results have generally failed to do so, and many have found that NPS has no impact on sales growth."[2]

And then, immediately: "Furthermore, although studies by van Doorn et al. (2013) and Pingitore et al. (2007) did find that NPS can predict sales growth to a certain extent, even these authors appear skeptical of NPS."[2]

Four observations, ours.

"Generally failed" rather than "failed," and "many" rather than "all," which are the qualifications a careful review includes and a summary drops.

The review then does something better than most. It names the studies that found in favour, rather than omitting them, which is what makes the passage usable as evidence.

And it characterises those supportive authors' own attitude, which is the part we found most striking and which the next section covers.

The venue is a review article in a leading peer-reviewed marketing journal, so this is the field describing its own state rather than an outsider's summary.

Two things that makes the passage worth more than a critique would be, ours. A review has no thesis to defend about this particular metric, and its job is to characterise a literature accurately.

And it was written by people who would be embarrassed by a colleague pointing out an omitted study, which is a real discipline. The incentive in a review runs toward completeness, which is the opposite of the incentive in a business article.

Even The Supportive Studies

The detail that settles the weight of evidence, in the review's own words.

It records that "Pingitore et al. (2007) for example called their study 'The Single Question Trap,'" and that "van Doorn et al. (2013, p. 317) concluded that 'the predictive capability of customer metrics, such as NPS, for future sales growth […] is limited.'"[2]

And it concludes: "Overall, it is fair to say that despite some limited support for the predictive value of NPS, the academic perception of NPS is predominantly negative."[2]

Four observations, ours.

A study titled "The Single Question Trap" that nonetheless found the metric predicts growth to some extent is unusually informative, because the title tells you what its authors made of their own result.

The second quotation is the supportive study's own conclusion, and the word it chooses is "limited."

The review's summary sentence is scrupulous in both directions: "some limited support" and "predominantly negative" in the same sentence.

And we did not obtain either of the studies quoted, only the review's quotations of them, which is a chain of two we would rather have shortened.

The Same Authors Defend It Elsewhere

The source that stops this article being one-sided, and we were glad to find it.

The authors of the 2007 replication published a clarification the following year: Keiningham, T. L., Aksoy, L., Cooil, B., and Andreassen, T. W. (2008), Net Promoter, Recommendations, and Business Performance: A Clarification on Morgan and Rego, Marketing Science, 27(3), 531–532[3].

Its abstract records: "One of the most controversial findings in Morgan and Rego (2006) was that two widely advocated loyalty metrics, 'Net Promoter' and 'Number of Recommendations,' have little or no value in predicting the financial outcomes of firms. We argue that neither measure was actually examined and that conclusions about the predictive value of these measures cannot be drawn from their analysis."[3]

Four observations, ours.

The same authors who failed to replicate the superiority claim argue that a different critical paper did not actually test the metric. That is the behaviour of people following evidence rather than pursuing a target.

Their stated objection is technical and specific: "the measures used in Morgan and Rego (2006) do not adequately adjust for the presence of neutral word-of-mouth activity."[3]

They also credit the paper they are criticising, which we would note: "we are unaware of another longitudinal study that examines the predictive value of satisfaction and loyalty metrics in such a comprehensive way."[3]

And this creates a complication for anyone citing the critical literature, ourselves included. One of the two studies the 2021 review names as finding no impact is one whose authors' peers argue did not measure the thing, and we did not obtain that dispute's resolution.

And A Condition Under Which It Works

A further counterpoint, from our weakest source and reported as such.

A non-academic blog characterises a 2021 study as arguing that firms have sought value in the metric under particular conditions, and showing that it "can effectively predict sales growth in certain conditions."[6]

This is a blog's characterisation of a paper we did not obtain, flagged here and at every use, and we do not know what those conditions are.

Three observations, ours.

A metric that works under specified conditions is a normal scientific outcome and not a defeat, and it is a considerably more useful claim than a universal one.

But it is also a substantial retreat from the original position. "The one number you need" and "works in certain conditions" are different claims, and a business adopting the first should know it has been narrowed to the second.

And we would want the conditions before recommending anything. We do not have them, which is why this article ends in arithmetic rather than in a verdict on the metric's usefulness.

What Actually Survives

Our reading, stated directly.

Five statements.

The comparative superiority claim failed a published longitudinal replication on the original author's own exemplar industries.

The academic assessment is predominantly negative, on a leading review's own summary, with some limited support acknowledged.

The supportive studies' authors are themselves sceptical, on the evidence of one's title and the other's stated conclusion.

The critical literature is not unanimous either, and the replication's own authors argue that one prominent critical paper did not measure the metric.

And none of this tells a small business whether to run the survey, which is a question about their own numbers rather than about the literature.

What The Metric Actually Does

The construction, because everything below follows from it. Ours.

Customers answer on a 0 to 10 scale. Scores of 9 and 10 are promoters, 7 and 8 are passives, and 0 through 6 are detractors. The score is the promoter percentage minus the detractor percentage, and the passives are discarded.

Four observations.

An eleven-point scale is collapsed into three buckets, and then one bucket is thrown away entirely.

The remaining two are subtracted rather than combined, which produces a figure that can run from minus one hundred to plus one hundred.

The bucket widths are wildly unequal. Promoters span two points, passives two, and detractors seven.

And every step of that discards information that the raw answers contained. The next three sections compute how much.

One point in the construction's defence before we take it apart, ours. Discarding information is what a summary statistic is for, and a mean discards a great deal too.

The question is never whether a metric loses information but whether it loses the information you needed, and the four-businesses table below is our answer to that for this one.

Four Businesses, One Score

Our own arithmetic, on invented customer distributions, each with 100 respondents.

A business where everyone answers 8: score 0, mean 8.0, spread 0.0.

A business where half answer 10 and half answer 0: score 0, mean 5.0, spread 5.0.

A business where half answer 10 and half answer 6: score 0, mean 8.0, spread 2.0.

A business where everyone answers 7: score 0, mean 7.0, spread 0.0.

Four observations.

All four score exactly zero. Their mean scores run from 5.0 to 8.0 and their spreads from 0.0 to 5.0.

The second and the first are not similar businesses in any respect. One has universal mild approval and the other has half its customers actively hostile, and the metric reports them as identical.

The mean of the same answers separates them immediately, at no additional cost. The information was collected and then discarded in the calculation, not missing from the survey.

And the polarised case is the one a business most needs to detect. Half your customers giving zeroes is an emergency, and it produces the same headline number as a placid, unremarkable book.

Two further points about that case, ours. A polarised book is not merely a worse version of a placid one; it is a different business problem, usually meaning one segment is being served well and another badly.

And it is the more fixable of the two, because a bimodal distribution points at a specific group, while a uniform mediocre score points at everything at once.

A Six Is A Detractor

The bucket boundary, which does most of the damage. Ours.

Four observations.

A customer answering 6, a customer answering 3, and a customer answering 0 are counted identically. All three are detractors and all three subtract equally from the score.

On any ordinary reading those are different customers. A six is somebody who would probably use you again and would not enthuse about it; a zero is somebody warning their friends away.

Meanwhile a 6 and a 7 are one point apart and fall on opposite sides of a boundary, one counting fully against you and the other not counting at all.

And the practical consequence is a perverse target. Moving a customer from 6 to 7 improves your score exactly as much as moving one from 0 to 7, though only one of those represents real work.

Two defences of the boundary exist and we would give them their due, ours. Sharp thresholds are easy to explain, and a metric nobody understands is not used.

And there is a substantive argument that only genuine enthusiasm produces referrals, so treating a 7 as neutral rather than positive may reflect something real about behaviour. We did not obtain evidence either way, and note that the argument justifies the promoter boundary rather than the seven-point detractor bucket.

The Sample Size Problem

What the construction costs in precision. Our own arithmetic, invented figures: a business with 40 percent promoters and 20 percent detractors, a score of 20.

Because the score is a difference of two proportions, its margin of error at 95 percent confidence is:

At 30 responses: plus or minus 26.8 points. At 50: 20.7. At 100: 14.7. At 200: 10.4. At 400: 7.3. At 1,000: 4.6 points.

Four observations.

With 100 responses, a reported score of 20 could truly be anywhere from 5 to 35.

Which means the quarterly movement most firms report on is noise. A score going from 18 to 26 at that sample size is not a change, and a business that celebrates it is celebrating sampling variation.

At 30 responses, which is what a small firm typically gets back, the margin is wider than the gap between a good score and a bad one.

And this is arithmetic about proportions rather than a criticism of the metric's construction, so the same problem afflicts any percentage-based measure at these sample sizes.

One property specific to this metric is worth separating out, ours. Because it is a difference of two proportions rather than one, its variance carries both, which makes it noisier than a single percentage from the same sample.

So the sample-size problem is general and this metric sits at the worse end of it. Both facts are true and only the first is usually mentioned.

A Correction To Our Own Comparison

Documented rather than removed, in this series' usual habit.

We first compared the score's margin of error as a share of its 200-point range against the mean's margin as a share of a 10-point scale, and concluded the mean was more precise.

That comparison is meaningless. Those are different rulers, and expressing each error as a percentage of its own arbitrary range can be made to favour either metric by rescaling.

Three observations, ours.

The error is a familiar one and this series has criticised it in others. Comparing two quantities by normalising each to its own range is the move that makes any two things comparable and therefore compares nothing.

We withdrew it and ran the honest test instead, which is whether each metric detects a real improvement, using the same survey and the same answers.

And the honest test produced a smaller advantage than the bad comparison had implied, which is the third time in four articles that checking our own work has cost us a stronger claim.

Which we would rather note than let pass, since the pattern now has enough instances to mean something. Every one of those errors would have strengthened the argument we were making, and none was caught by anything except running the check.

Which Metric Finds A Real Improvement

Our own simulation, 4,000 trials per row, on invented parameters: a genuine improvement raising the underlying average from 7.5 to 8.0 on a ten-point scale, with a spread of 2.0, tested at 95 percent confidence.

How often each metric detects it:

At 30 responses: the score 16 percent of the time, the mean 16 percent.

At 50: 21 against 23. At 100: 37 against 41. At 200: 61 against 67. At 400: 88 against 93. At 800: 100 against 100 percent.

Four observations.

The mean detects the same real improvement more often at every sample size where it matters, using the identical survey and the identical answers.

But the advantage is modest: four to six percentage points of power in the middle of the range, and we report that rather than the dramatic version.

At the extremes the two converge. At 30 responses both fail almost always and at 800 both succeed, so the choice of metric only matters in between.

And that finding is worth more than it first appears, because it says the metric choice is a second-order problem. The first-order problem is the next section.

One reason we report the modest figure rather than the striking one, ours. We could have chosen a latent distribution that made the gap look larger, since the advantage depends on where the bucket boundaries fall relative to the customers.

A distribution clustered around 6 and 7 would exaggerate the difference enormously, and one clustered at 9 would erase it. We used a plausible middle and report what it gave us, which is a small consistent advantage.

The Finding That Matters Most

The number a small business should take from this article. Ours.

Four observations.

At 100 responses, neither metric detects a genuine half-point improvement more than about 40 percent of the time. The improvement is real, it is in the data, and the survey misses it three times in five.

Which reframes the entire question. Arguing about which customer metric to use is arguing about the second-order term while the measurement itself is underpowered.

And 100 responses is more than most small firms collect. At 30 responses both metrics find the improvement 16 percent of the time, which is barely above the rate at which they would report a change that was not there.

So the honest position for a small business is uncomfortable and worth stating. A quarterly customer score computed from a few dozen responses is not measuring your customers; it is measuring which few dozen answered.

Two consequences follow, ours, and the first is liberating rather than discouraging. You can stop agonising over which metric to adopt, because at your sample size the choice does not determine the answer.

And the second is where the effort should go instead. Getting from thirty responses to two hundred changes what you can learn far more than any change of formula, and it is a question of asking more customers rather than of analysis.

What Happens When It Becomes A Target

The consequence the arithmetic predicts once anyone is measured on the number. Ours.

Four observations.

The bucket boundaries create a specific and cheap way to move the score without improving anything. Every customer sitting on a 6 is worth a full point of score if nudged to a 7.

Which is why the request to reconsider a score exists, and anyone who has bought a car has met it. The staff member is not being dishonest; they are responding correctly to the measure they are judged on.

The same pressure operates on who gets surveyed. Sending the survey to satisfied customers is the cheapest available improvement, and at the sample sizes above it works completely.

And this is the seventy-seventh article's finding arriving in a new place. A measure used as a target stops measuring, and a measure with sharp arbitrary boundaries stops measuring faster, because the boundaries tell you exactly where to push.

Comparing Against Someone Else's Score

A use we would treat with particular caution. Ours, and not marketing advice.

Four observations.

The metric's genuine advantage is comparability, since everyone computes it identically and published benchmarks exist by industry.

But comparability requires that the other figures were produced the same way, and the arithmetic above lists the ways they might not be. Different sample sizes, different response rates, different selection of who was asked, and different markets.

A published benchmark rarely states any of those. An industry average of 35 says nothing about how many responses it rests on, and on our own figures a score computed from thirty responses carries a margin wider than most benchmark gaps.

And the cross-cultural observation below sharpens it further, because a benchmark drawn from a different market may be measuring something systematically different.

One practical rule we would apply, ours. Compare your score against your own previous score, computed the same way from a similar sample, which controls most of what an external benchmark does not.

That is a weaker comparison than a benchmark appears to offer and a stronger one than it actually delivers. Your own trend is the only comparison whose method you can verify, because you ran it.

The Question Asks Two Things

A separate problem with the wording, from a journal article's discussion.

It observes that customer experience surveys usually use a single stimulus, being the company or its product, but that the recommendation question introduces two: the product and the friend[5].

It reports this as an explanation for consistently low scores in some markets, suggesting respondents may hold a positive view of a company and still answer low because they would not risk relationships by recommending[5].

This is a discussion passage rather than a tested finding we obtained, and we grade it accordingly.

Three observations, ours.

If it holds, the question is measuring willingness to spend social capital as well as satisfaction, and those come apart.

That matters for any business with customers across different cultures, and for any firm comparing its score against a published benchmark drawn from a different market.

And it is a specific instance of a general caution. A question that asks about two things at once returns an answer about neither cleanly, which is elementary survey design and is nonetheless the most famous customer question in business.

One test a firm can run on its own data, ours. Ask the recommendation question and a plain satisfaction question in the same survey, and look at how far apart they place the same customers.

If they agree closely, the two-stimuli concern does not bite in your market and you can use either. If they disagree systematically, you have found something specific about your customers, which is more valuable than either score.

Why It Persists Anyway

Our reading, offered as reasoning rather than evidence.

Four observations.

One number is genuinely easier to manage than several, and the demand the metric met is real even if the claim behind it did not survive.

It is comparable across firms, which a mean satisfaction score is not, because everyone computes it the same way and published benchmarks exist.

The single-question format has a real practical advantage nobody should dismiss. Response rates fall as surveys lengthen, and a one-question survey gets answered.

And the arithmetic above does not touch any of that. We have shown what the metric discards and how imprecise it is at small samples, and neither addresses why a manager wants a single comparable figure, which is a real need.

One further reason worth naming honestly, ours. A great deal of money and organisational commitment now sits behind the metric, in software, in training, and in bonus schemes, and that is a cost of abandoning it independent of whether it works.

Which is not an argument for keeping it and is a reason to expect it to persist. The eighty-first article's point about sunk commitment applies to measurement systems as readily as to projects.

What To Use Instead

Ours, untested, and not marketing advice.

Four points.

Keep the raw answers. Whatever headline you report, storing the 0 to 10 responses costs nothing and lets you compute the mean, the spread and the score from the same data.

Report the distribution alongside any single figure, because the four-businesses table above is only invisible if you look at one number.

Put a margin of error next to it. At 100 responses that is plus or minus roughly 15 points on our own figures, and printing it prevents the quarterly celebration of noise.

And pool across periods rather than reporting quarterly, if your volume is low. Four quarters of 30 responses is 120 responses, and the annual figure is the one that means something.

One caution on pooling, ours. It buys precision at the cost of timeliness, and a problem that started in March will not show in a figure that also contains January.

Which is a real trade and not an objection. A precise annual number and a stream of comments read weekly covers both requirements better than an imprecise quarterly number covers either.

Not An Argument For Measuring Nothing

Because this article could be read that way and should not be. Ours.

Four observations.

Asking customers what they think is worth doing, and the case against a particular arithmetic operation is not a case against the question.

The verbatim comments are frequently worth more than the number, and they are immune to every problem in this article. A customer explaining why they gave a six is telling you something no score can.

Small samples are less of a problem for that purpose. Thirty comments will not give you a reliable trend and may well give you the specific fault, which is a different and often more useful thing.

And we would put the point plainly. The metric's weakness is as a tracked number; its survey's strength is as a prompt for people to tell you things, and a firm that keeps the second and stops reporting the first has lost nothing.

One qualification on that, ours, because comments have their own selection problem. The people who write comments are not a random sample of the people who answered, and they skew toward the extremes at both ends.

That makes comments good for finding faults and bad for judging how common a fault is. Use them to generate the hypothesis and the numbers to size it, which is the division of labour each is actually suited to.

Bibliographic Note

The series keeps a count.

The 2007 replication is cited by our sources as 71(3), 39–51[1], as 71 (March), 39–51[5], and as 71, July, 39–51[4].

Three observations, ours.

March and July cannot both be right for the same issue of the same journal, and the volume, pages and title agree across all three.

The pages are consistent everywhere, so this would not defeat a lookup, which is why we record it as a variant rather than an error of consequence.

That brings the running count of bibliographic variants across this series to thirty-two.

One observation on the shape of this one, ours. It is a disagreement between two peer-reviewed reference lists, which is the category we found most often in the first fifty articles and less often lately.

What To Do

Keep the raw 0 to 10 answers whatever you report, because the score can be computed from them and they cannot be recovered from the score.

Look at the distribution, not just the headline. On our own arithmetic four wholly different customer bases all score exactly zero, and their means run from 5.0 to 8.0.

Print a margin of error. At 100 responses a score of 20 could truly be anywhere from 5 to 35 on our own figures.

Stop reporting quarterly if your volume is low, and pool instead. At 30 responses both metrics detect a real half-point improvement only 16 percent of the time.

Do not treat a movement as a result unless it exceeds the margin of error you just printed.

Read the comments. They are immune to every arithmetic problem in this article and thirty of them may give you the actual fault.

Be careful comparing against published benchmarks, particularly across markets, given the observation that the recommendation question asks about two things at once.

And do not conclude the metric is worthless. The replication that failed was of a comparative superiority claim, a review reports some limited support, and the replication's own authors elsewhere defend the metric against a different critique.

The Limits Of This Analysis

Several caveats matter. This article discusses research on customer metrics and is not marketing advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 2003 original article, and our statement of its claim comes from the replicating authors' characterisation of it, which is a characterisation by people with a position, though it matches the original's own title. We did not obtain the 2007 replication itself, only its abstract, whose final sentence cuts off mid-word in our reproduction; we therefore report no effect sizes, correlations or comparison figures from it. We did not obtain any of the studies discussed inside the 2021 review, so the quotations from Pingitore and colleagues and from van Doorn and colleagues reach this article through a chain of two sources. We did not obtain the 2021 study reported as finding the metric works under certain conditions, which reaches us through a non-academic blog and whose conditions are unknown to us; that is the largest gap here, since a specified condition would be more useful than any general verdict. The cross-cultural observation is a discussion passage rather than a tested finding we obtained. All arithmetic is ours and every parameter in it is invented: the four distributions, the promoter and detractor proportions, the spread of 2.0 and the half-point improvement are all supplied by us; the power figures come from 4,000 simulated trials per row and would differ under a different latent distribution. The simulation assumes responses are drawn independently from a stable distribution, which real surveys with self-selected respondents are not. And this article contains one withdrawn comparison: we first compared each metric's error as a share of its own range, which is meaningless, and replaced it with a direct test that produced a smaller advantage.

Frequently Asked Questions

Did the Net Promoter claim fail to replicate?
A 2007 study in the Journal of Marketing, using 21 firms and over 15,500 interviews and testing the industries the original author cited as exemplars, reports that it fails to replicate his assertions regarding the clear superiority of the metric. That is a failure of the comparative claim rather than a finding that the metric is worthless.
Is the academic verdict unanimous?
No, and a 2021 review says so carefully: despite some limited support for its predictive value, the academic perception is predominantly negative. The same review notes that two studies did find it predicts growth to a certain extent, while observing that even those authors appear sceptical of it.
What does the metric discard?
An eleven-point scale is collapsed into three unequal buckets, one bucket is thrown away, and the other two are subtracted. On our own arithmetic four wholly different customer bases all score exactly zero, with mean scores from 5.0 to 8.0 and spreads from 0.0 to 5.0.
Is a score of 6 really counted as negative?
Yes. The detractor bucket runs from 0 to 6, so a customer answering 6 counts identically to one answering 0, while a 6 and a 7 are one point apart and fall on opposite sides of the boundary. Moving a customer from 6 to 7 improves the score exactly as much as moving one from 0 to 7.
How precise is the score at my sample size?
On our own arithmetic, for a business at 40 percent promoters and 20 percent detractors, the 95 percent margin is plus or minus 26.8 points at 30 responses, 14.7 at 100 and 4.6 at 1,000. At 100 responses a reported score of 20 could truly be anywhere from 5 to 35.
Would using the average be better?
Slightly, and we report the modest figure. On our own simulation of a real half-point improvement, the mean detects it 41 percent of the time at 100 responses against 37 for the score, and 67 against 61 at 200. The advantage is consistent and it is four to six points, not a transformation.
So what should a small business actually do?
Keep the raw answers, print a margin of error, pool across periods rather than reporting quarterly at low volume, and read the comments. On our own simulation neither metric detects a genuine half-point improvement more than about 40 percent of the time at 100 responses, so the first-order problem is sample size rather than metric choice.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article withdraws one of its own comparisons and reports the smaller advantage the honest test produced.

References

  1. Publisher record for Keiningham, T. L., Cooil, B., Andreassen, T. W., & Aksoy, L. (2007), A Longitudinal Examination of Net Promoter and Firm Revenue Growth, Journal of Marketing, 71(3), 39–51, reproducing the abstract: that managers have widely embraced and adopted the Net Promoter metric, which loyalty consultant Frederick Reichheld advocates as the single most reliable indicator of firm growth compared with other loyalty metrics such as customer satisfaction and retention; that there has recently been considerable debate about whether the metric is truly superior; that the article employs longitudinal data from 21 firms and more than 15,500 interviews from the Norwegian Customer Satisfaction Barometer to replicate the analyses used in Net Promoter research, and compares Reichheld and colleagues' findings with the American Customer Satisfaction Index; and that using industries Reichheld cites as exemplars, the research fails to replicate his assertions regarding the clear superiority of Net Promoter, at which point our reproduction cuts off mid-word. Note: the publisher's record. Our source for the replication. We obtained the abstract and not the paper, and report no effect sizes, correlations or comparison figures from it. journals.sagepub.com
  2. Publisher record for a 2021 review article in the Journal of the Academy of Marketing Science, DOI 10.1007/s11747-021-00790-2, reproducing body text: that as well as the methodological critiques cited, empirical studies aiming to replicate Reichheld's results have generally failed to do so, and many have found that the metric has no impact on sales growth, citing Keiningham and colleagues (2007) and Morgan and Rego (2006); that although studies by van Doorn and colleagues (2013) and Pingitore and colleagues (2007) did find the metric can predict sales growth to a certain extent, even these authors appear sceptical of it; that Pingitore and colleagues titled their study "The Single Question Trap"; that van Doorn and colleagues concluded the predictive capability of customer metrics such as this one for future sales growth is limited; and that overall it is fair to say that despite some limited support for its predictive value, the academic perception is predominantly negative. Note: the publisher's record for a peer-reviewed review article. Our source for the state of the field. The quotations from Pingitore and from van Doorn reach this article through this review rather than from those papers, which we did not obtain. link.springer.com
  3. Economics database record for Keiningham, T. L., Aksoy, L., Cooil, B., & Andreassen, T. W. (2008), Net Promoter, Recommendations, and Business Performance: A Clarification on Morgan and Rego, Marketing Science, 27(3), 531–532, reproducing the abstract: that one of the most controversial findings in Morgan and Rego (2006) was that two widely advocated loyalty metrics, Net Promoter and Number of Recommendations, have little or no value in predicting the financial outcomes of firms; that the authors argue neither measure was actually examined and that conclusions about their predictive value cannot be drawn from that analysis; that a primary problem is that the measures used do not adequately adjust for the presence of neutral word-of-mouth activity; and that nevertheless Morgan and Rego provide important information regarding other common customer metrics and firm financial outcomes, the authors being unaware of another longitudinal study examining the predictive value of satisfaction and loyalty metrics in such a comprehensive way. Note: an economics database record reproducing a publisher abstract. Recorded because these are the same authors as the 2007 replication, here defending the metric against a different critical paper, which is the source that keeps this article from being one-sided. ideas.repec.org
  4. Reference lists carried on two peer-reviewed articles discussing the metric, confirming Reichheld, F. F. (2003), The one number you need to grow, Harvard Business Review, 81(12), 46–55; Keiningham, T. L., Cooil, B., Andreassen, T. W., and Aksoy, L. (2007), Journal of Marketing, 71, 39–51; Kristensen, K., and Eskildsen, J. (2014), Is the NPS a trustworthy performance measure?, The TQM Journal, 26(2), 202–214; Morgan, N. A., and Rego, L. L. (2006), The value of different customer satisfaction and loyalty metrics in predicting business performance, Marketing Science, 25(5), 426–439; and Keiningham, T. L., Gupta, S., Aksoy, L., and Buoye, A. (2014), The High Price of Customer Satisfaction, MIT Sloan Management Review, 55(3), 37–46. Note: reference lists; citations only. Recorded also as a source of a bibliographic variant: one of these gives the 2007 paper's month as July where another source gives March. We obtained none of the works named here. link.springer.com
  5. Journal article in the Asia Marketing Journal on consistently low Net Promoter scores in cross-cultural research, whose discussion observes that customer experience surveys usually utilise a single stimulus such as the company or its products and services, but that the recommendation question involves two stimuli, namely the company's product or service and the influence of friends; and suggests that respondents in some markets may hold a positive attitude towards a company while giving low scores because they would not risk ruining relationships with friends. Its reference list gives the 2007 replication as Journal of Marketing, 71 (March), 39–51. Note: a peer-reviewed article, but this is a discussion passage rather than a tested finding we obtained, and is graded accordingly. Recorded also for a bibliographic variant in the month given for the 2007 paper. koreascience.kr
  6. Non-academic blog post discussing whether the metric predicts growth, recording that empirical studies to replicate the 2003 work found it has no impact on sales growth, citing Morgan and Rego (2006) and Keiningham and colleagues (2007); and characterising a 2021 study by Baehre, O'Dwyer, O'Malley and Lee as arguing that firms have sought value in the metric under a particular condition, addressing how practitioners can best use it, and showing that it can effectively predict sales growth in certain conditions. Note: a non-academic blog, flagged at every use in this article. Cited solely for its characterisation of a 2021 study we did not obtain and whose conditions are therefore unknown to us. This is the weakest source here and supports the article's main counterpoint, which we would rather have had directly. numberanalytics.com

This article discusses research on customer metrics and is not marketing advice. The 2003 original and the 2007 replication were not obtained, only abstracts and characterisations; the 2007 abstract cuts off mid-word in our reproduction. Quotations from two supportive studies reach this article through a review rather than from those papers. All arithmetic is the authors' own and every parameter in it is invented; power figures come from 4,000 simulated trials per row. This article contains one withdrawn comparison.