Every firm that has ever built a dashboard has been warned that metrics get gamed. That warning is correct and it is about the wrong problem. The finding in this article concerns people who are not gaming anything, who believe they are pursuing the objective, and who have quietly stopped being able to tell the difference between the objective and its measurement.

Key Takeaway

From the paper that named it: "Ideally, managers see measures for what they are, imperfect proxies for intangible strategic constructs. However, managers may fail to fully appreciate the fact that measures are merely representations of the strategic constructs, and act as though the measures are the construct of interest, a phenomenon we label surrogation."[1] Our own simulation found that incentivising a proxy can make it a better signal or a much worse one depending entirely on which kind of effort it buys, and that the number on the dashboard rises in every case.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. Surrogation and metric gaming are different failures requiring opposite remedies. Gaming requires knowing the measure is not the goal and exploiting the gap. Surrogation is the loss of that knowledge. Monitoring catches the first and is useless against the second, because there is nothing to catch.

Two. The evidence is experimental, published in the field's leading journals, and narrow. Two experiments supporting two hypotheses, in The Accounting Review, using a video game as the task. We report that last detail because it matters.

Three. Incentives do not simply corrupt a measure. They amplify whatever effort they buy. On our own simulation, a proxy correlating 0.70 with its construct rose to 0.877 when incentivised effort raised both, and fell to 0.447 when the same effort raised only the measure. The crossover sits near an even split.

Four. The measured number goes up in every scenario. That is why the failure is invisible from the dashboard, and it is the single most useful thing in this article.

Five. The paper's finding that multiple measures reduce surrogation is derivable, and it inherits a condition nobody states. Averaging several proxies improves the composite only if their errors are independent. On our own model, three measures drawn from a common biased source recover almost none of the benefit.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A that the construct is defined and published as described. The definition is quoted verbatim from the publisher's record of a paper in The Accounting Review.

Grade B that incentive compensation exacerbates surrogation, and Grade B that single measures do so more than multiple measures. Both are supported by two experiments reported in the abstract, and we obtained no effect sizes, no sample sizes and no results section.

Grade C for generalisation beyond the laboratory. The experimental task was a video game, a citing source confirms this, and no field evidence reached us.

Grade A for Goodhart's original formulation, which we obtained verbatim, and Grade B for its popular rephrasing, whose attribution we verified but whose source we did not obtain.

Grade A for our own arithmetic, which is a derivation rather than a finding and which anyone can reproduce.

Our position: the construct is real, well specified, and its evidence base is thinner than its usefulness. This is the rare case in this series where we would act on a Grade B finding, because the mechanism is checkable inside your own firm without believing the experiments.

A Note On Method

Everything here is verified to August 2026.

We obtained the 2012 paper's abstract verbatim from the publisher, plus a variant of the same abstract from a working paper repository[1][2]. We did not obtain the paper, its experiments, its samples, its statistics or any effect size, and we report none.

We obtained one sentence of the 2013 paper's abstract before our source truncated[3].

The detail that the experimental task was a video game comes from a separate paper citing it[4], not from the paper itself.

Goodhart's 1975 sentence is quoted from a reference work and confirmed against a second source[5][6]. We did not obtain the 1975 article.

Two sources used here are not academic: a reference encyclopedia and an editorial website, both flagged at every use.

All arithmetic and simulation is ours, uses invented parameters throughout, and demonstrates mechanisms rather than measuring anything.

This article discusses research on performance measurement. It is not compensation design, management accounting or human resources advice.

A Phenomenon We Label Surrogation

The definition, from the paper that introduced the term.

Choi, J. (Willie), Hecht, G. W., and Tayler, W. B. (2012), Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review, 87(4), 1135–1163, DOI 10.2308/accr-10273[1].

The abstract opens: "To facilitate managers' decision-making, firms develop strategic performance measurement systems that translate strategy into performance measures. Ideally, managers see measures for what they are, imperfect proxies for intangible strategic constructs. However, managers may fail to fully appreciate the fact that measures are merely representations of the strategic constructs, and act as though the measures are the construct of interest, a phenomenon we label surrogation."[1]

A citing source states it more plainly: "The study demonstrates that managers may focus on performance measures to the point of overlooking the strategic constructs they represent, a behavior termed surrogation."[4]

And another gives the operational version: "Strategy surrogation occurs when managers focus on the measures in the SPMS on which they are compensated and completely or partially lose focus on the overall strategic objectives of the organization."[7]

Four observations, ours.

The word doing the work is "act as though." This is not a claim that managers say the measure is the goal; it is a claim about behaviour proceeding as if it were.

Note the phrase "fail to fully appreciate." That is a description of a cognitive limitation, not of a choice, and it is what separates this from every account of metric gaming.

"Completely or partially" in the third quotation matters. Surrogation is presented as a matter of degree rather than a binary state, which makes it harder to detect and harder to deny.

And the framing is unusually careful about what a measure is. Imperfect proxies for intangible strategic constructs concedes in advance that the measure was never going to be the thing, which is the correct starting point and one most dashboards do not take.

What They Predicted

The hypotheses, which are specific and directional.

"In accordance with the attribute substitution framework (Kahneman and Frederick 2002), we predict incentive compensation exacerbates surrogation, and that this effect is more prevalent when managers are compensated on a single measure of a strategic construct than when managers are compensated on multiple measures of a strategic construct."[1]

Three observations, ours.

Two claims are being made, and they are separable. First, paying on a measure worsens the substitution. Second, paying on one measure worsens it more than paying on several.

The second is the actionable one, because a firm can act on it directly. It is also the one we can derive independently, which we do below.

And the prediction is theoretically grounded rather than exploratory, drawing on an established framework from the judgment literature. That is the correct way to structure an experiment and it is worth crediting.

And What They Found

The result, in one sentence, which is all we have.

"Via two experiments, we find support for these hypotheses."[1]

The paper positions the contribution: "More generally, we identify a by-product of contracting on imperfect performance measures not previously considered in extant literature, and establish when consideration of costs of this by-product are likely to be critical."[1] A repository version adds: "given that surrogation can be detrimental to the implementation of strategy, our study furthers academic and practitioner understanding of factors that affect the extent to which firms incur this cost."[2]

Four observations, ours, and they are mostly about what we cannot tell you.

We obtained no effect sizes, no sample sizes, no statistics and no results section. Two experiments supporting two hypotheses is the entire empirical content available to us, and we grade the findings B on that basis.

This is a real limitation and we want to be direct about it. The fifty-sixth article in this series demonstrated what underpowered experimental literatures do to reported effects, and we have no way to assess whether that applies here.

The framing as a cost is deliberate and correct. Surrogation is presented as a by-product of contracting, meaning a price paid for the benefits of incentive compensation rather than an argument against it.

And "not previously considered in extant literature" is a strong claim for a phenomenon that resembles a well-known adage. The next sections are about whether it is justified, and we think it is.

The Task Was A Video Game

A methodological detail we obtained from elsewhere, and which we report because it bears on generalisation.

A separate paper describing experimental practice in the field records: "our experiment follows a long line of experimental accounting research that uses abstract tasks to test psychological mechanisms and then infers that the results pertaining to these psychological mechanisms generalizes to real-world accounting tasks. For other examples of this type of research see Choi, Hecht, and Tayler (2012) who use a video game for their experimental task."[4]

Three observations, ours.

The defence offered is explicit and reasonable: abstract tasks test mechanisms, and mechanisms are assumed to transfer. That is a defensible research strategy and it is stated openly rather than hidden.

It is also an assumption rather than a demonstration, and this series has repeatedly found the gap between a laboratory mechanism and a field outcome to be where effects vanish.

And a video game is arguably a favourable setting for surrogation, since a score is unusually salient and the underlying objective unusually abstract. A design that maximises the chance of observing an effect is the right way to test whether it exists at all, and the wrong way to estimate how large it is in practice.

Attribute Substitution

The theoretical basis, which connects this article to two earlier ones.

The prediction is made "in accordance with the attribute substitution framework"[1], a general account under which a person facing a difficult question answers an easier one instead, without noticing the substitution.

Three observations, ours.

The fit is exact. "Is our strategy working?" is hard. "Is the number up?" is easy. Attribute substitution predicts the second question gets answered while the first is believed to have been.

We did not obtain the framework's source paper, and report only that the 2012 paper cites it as its theoretical basis.

And the framework sits inside the dual-process tradition, which the fifty-third article in this series examined at length, reporting a published exchange in which critics called the two-type typology a convenient and seductive myth and five researchers replied that the disputed features were never definitional. Attribute substitution does not depend on the contested part, since it requires only that some processing is automatic, but a reader should know the surrounding theory is under argument.

This Is Not Goodhart's Law

The distinction this article exists to make. Ours, though a non-academic source states something similar.

An editorial site puts it: "Goodhart's Law is the broad principle that targeting a proxy corrupts it. Surrogation is a specific, quieter mechanism within it: the cognitive slip where well-intentioned people" lose the distinction, with our source truncating[6]. This is an editorial website, not an academic source, and we flag it as such.

Four observations, ours, and this is the analytical core.

Gaming requires knowing the measure is not the goal. A call centre agent who hangs up to shorten average handling time knows perfectly well that the objective was serving the customer. The knowledge is what makes the behaviour possible.

Surrogation is the absence of that knowledge. A manager who has surrogated is not exploiting a gap between measure and goal; they have stopped perceiving one. They would fail a lie detector about it, because they are not lying.

The two therefore demand opposite remedies. Gaming responds to monitoring, audit and consequence, because there is a concealed intention to detect. Monitoring does nothing against surrogation, because the behaviour is sincere and there is nothing concealed.

And a firm that diagnoses surrogation as gaming will install controls, find nothing, conclude the problem is solved, and continue to miss its objective while every number improves.

What Goodhart Actually Wrote

The original, which is narrower than its reputation.

A reference work records the 1975 formulation, from an article on United Kingdom monetary policy: "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes."[5]

A second source gives the context: the Bank of England "had found stable relationships between certain money-supply measures and inflation, but when it started targeting those measures to control inflation, the relationships broke down as banks and markets adapted."[6]

We did not obtain the 1975 article and report the sentence from a reference work, confirmed against a second source.

Three observations, ours.

The original is a claim about statistical regularities, not about human motivation. It says a relationship collapses under control pressure, and says nothing about why.

That is a considerably more general and more careful statement than the version in circulation, and it covers both gaming and surrogation without committing to either mechanism.

And the mechanism in the monetary case was adaptation by banks and markets, which is neither cheating nor cognitive substitution. Goodhart's law has at least three distinct mechanisms and the popular version names none of them.

The Sentence Everyone Quotes

The famous formulation, and its provenance.

A reference work attributes the adage "When a measure becomes a target, it ceases to be a good measure" to a 1997 work by the anthropologist Marilyn Strathern[5], and a widely circulated book describes her as having "rephrased it clearly and concisely" because "Goodhart's original formulation is a bit opaque"[8].

The same reference work records an earlier generalisation, from a 1996 book chapter: "'Goodhart's Law' – That every measure which becomes a target becomes a bad measure – is inexorably, if ruefully, becoming recognized as one of the overriding laws of our times."[5]

Three observations, ours.

The 1996 formulation predates the 1997 one and says substantially the same thing. The sentence that travelled was not the first to say it.

We did not obtain either the 1997 or the 1996 source, and report both from a reference work, which is weaker sourcing than this series prefers for a load-bearing quotation.

And note what the rephrasing did. Goodhart described a statistical collapse; Strathern's version describes a measure becoming bad. Those are not identical claims, and the second implies an evaluative judgment the first does not.

And It May Not Even Be First

A provenance note, because attribution in this area is unusually tangled.

The same reference work records that "Campbell's law likely has precedence, as Jeff Rodamar has argued, since various formulations date to 1969."[5]

We did not obtain Campbell's work or the argument for its precedence and report the claim as the reference work states it.

Two observations, ours.

A principle named after a 1975 economist, popularised by a 1997 anthropologist, generalised by a 1996 historian, and possibly anticipated in 1969 is a principle whose attribution is itself an example of a measure detaching from what it measures.

And we report it because this series keeps a running record of bibliographic inconsistency, and this topic supplies more of it than most. The next relevant instance appears below and concerns the authors' own papers.

Three Mechanisms, One Adage

Developing the point left open above, because Goodhart's law is routinely treated as one phenomenon and behaves as at least three. Ours.

The first is gaming. A person knows the measure is a proxy, knows what the proxy misses, and acts in the gap deliberately. The examples an editorial source gives are of this kind: call centres measured on call duration where staff learn to hang up, hospitals measured on emergency-room waiting times, schools measured on test scores[6]. The intention is concealed, which means it can in principle be uncovered.

The second is surrogation, which is the subject of this article. The person no longer distinguishes measure from construct. Nothing is concealed because nothing is known, and the behaviour would survive any amount of scrutiny because it is sincere.

The third is adaptation, and it involves nobody inside the organisation at all. In the monetary case a second source describes, the relationships broke down "as banks and markets adapted"[6]. Nobody at the Bank of England gamed anything and nobody surrogated. The regularity collapsed because external parties changed their behaviour in response to it being targeted, which is exactly what Goodhart's own sentence describes and neither of the other two mechanisms covers.

Four observations, ours.

The three have different detection signatures. Gaming leaves evidence of intent. Surrogation leaves a divergence between measure and outcome with no evidence of intent. Adaptation leaves a change in the behaviour of people you do not employ.

They have different remedies. Gaming responds to monitoring and consequence. Surrogation responds to keeping the construct visible and measured separately. Adaptation responds to neither, because the adapting party is outside your control and is behaving reasonably; the only defence is to expect the relationship to decay and to re-estimate it.

And they can occur together, which is the practical difficulty. A sales target can simultaneously be gamed by one person, surrogated by another, and adapted around by customers, producing a single degraded number with three causes and no single fix.

Which is why we would resist the usual advice to pick better metrics. The metric is rarely the problem. Gaming is about available actions, surrogation is about attention, and adaptation is about other people, and a better-chosen number addresses none of the three on its own.

What Incentivising A Proxy Does

The arithmetic. Our own simulation, 200,000 runs per condition, invented parameters throughout, and it measures nothing about any real firm.

We built a measure M that proxies a construct C, calibrated so their correlation before any incentive is 0.699, which is a good proxy by most standards.

Then we attached pay to M and let effort arrive. Some effort raises both M and C, which is what a well-chosen measure is supposed to produce. Some raises M alone. Total effort is held constant, so the only thing varying is where it goes.

With none of the effort going to M alone: correlation 0.877.

At 25 percent: 0.814. At 40 percent: 0.749. At 50 percent: 0.695. At 60 percent: 0.635. At 80 percent: 0.525. At 100 percent: 0.447.

The Result We Did Not Expect

Reading that table, which says something more interesting than we set out to demonstrate.

Four observations, ours.

Incentives do not simply corrupt a proxy. When the effort they buy raises both the measure and the construct, the measure becomes a better signal than it was before anyone was paid on it, moving from 0.699 to 0.877.

When the effort raises only the measure, the proxy collapses to 0.447, which is worse than the uncontracted baseline by a third.

So the honest statement is not that measurement corrupts. It is that incentives amplify a measure in whichever direction the effort goes, and the sign of the effect depends entirely on what is available to be done.

And this reframes the design problem usefully. The question is not whether to incentivise a measure. It is whether your people have easy access to actions that move the measure without moving the objective, and that is a question about the design of the work, not about anyone's character.

We Got The Reference Point Wrong Twice

Continuing a practice from the previous five articles, because this one took three attempts.

First error. We wanted a proxy correlating 0.70 with its construct, and built both variables with a loading of 0.70 on a shared factor. That produces a correlation of 0.49, not 0.70, because the correlation is the product of the loadings. The required loading is the square root, about 0.837.

Second error. Having fixed that, we compared every incentivised run against the uncontracted 0.70 and reported the zero-gaming case as a 25 percent loss of proxy quality. It is not a loss. Effort that raises both variables adds shared variance and lifts the correlation, which is why the zero-gaming case reads 0.877. We had mislabelled an improvement as a degradation.

Three observations, ours.

Both errors were caught by checking a baseline against what it should have been. The first because 0.488 is not 0.70; the second because a column of negative percentages next to rising correlations makes no sense.

The second error is the more instructive, because it would have produced a more quotable article. The wrong version said incentives always degrade a proxy, which is the conclusion a reader expects and which we would have been pleased to report.

And that is the specific danger this series keeps documenting. An error that confirms what you expected is much harder to notice than one that contradicts it, and the only defence that has reliably worked is checking the arithmetic against a case where you already know the answer.

The Crossover

The practically useful feature of that table. Ours.

The correlation passes back through its uncontracted level of about 0.70 when roughly half the incentivised effort is going into the measure alone.

Three observations.

Below that point, incentivising the measure makes it a better proxy than it was. Above it, worse. The crossover is where the contract stops helping.

And here is the fact that makes surrogation dangerous rather than merely undesirable. The measured number rises in every single row of that table. At 100 percent gaming, M is higher than it has ever been and tells you almost nothing.

So the failure is invisible from the dashboard by construction. Every diagnostic a firm normally runs, comparing the metric to target, to prior period, to peers, will report success in exactly the case where the measure has stopped meaning anything.

Why Multiple Measures Help

Deriving the paper's second hypothesis independently, which we can do because it is arithmetic. Our own simulation, invented parameters.

Suppose several measures each read the same construct with independent error, each correlating 0.70 with it individually. Average them.

With one measure: 0.700. With two: 0.810. With three: 0.862. With four: 0.891. With six: 0.924. With eight: 0.940.

Four observations, ours.

Moving from one measure to three lifts the composite from 0.70 to 0.86, which is a substantial improvement for a change most firms could make in an afternoon.

This is the averaging principle from the thirty-sixth article in this series, applied to performance measurement rather than to forecasts. Independent errors cancel; the signal does not.

It gives an arithmetic reason for the paper's prediction that compensating on a single measure exacerbates surrogation more than compensating on several. A single measure is both a worse proxy and an easier target, and the two problems compound.

And the returns diminish sharply. The move from one to three captures most of the available benefit; going from six to eight buys very little.

And When They Do Not

The condition nobody states, and it is the one that fails in practice. Ours.

All of the above requires the measurement errors to be independent. We tested what happens when they are not, using three measures throughout.

With independent errors: 0.863.

With 30 percent of the error shared across all three: 0.801.

With 60 percent shared: 0.752.

With 90 percent shared: 0.712, which is barely better than using one measure alone.

Four observations, and this is the most practically important passage in the article.

Three measures with a common bias are worth almost exactly one measure. The dashboard looks balanced and is not.

And shared error is the normal case, not the exception. Measures drawn from the same system, the same period, the same reporting chain, or the same self-reports inherit that source's distortions together.

The thirty-sixth article established this in a different setting: averaging only works on independent inputs, and the forty-first showed what happens when the independence is destroyed by sequence.

And the practical test follows directly. If a single failure could move all your measures the same way, they are not multiple measures. They are one measure reported several times.

The Second Paper

The follow-up, of which we obtained one sentence.

Choi, J., Hecht, G. W., and Tayler, W. B. (2013), Journal of Accounting Research, 51(1), 105–133, DOI 10.1111/j.1475-679X.2012.00465.x[3].

Its abstract opens: "Strategic performance measurement systems operationalize firm strategy with a set of performance measures. A consequence of such alignment is the tendency for managers to lose sight of the strategic" constructs, with our source truncating[3].

Two observations, ours.

The truncation falls at the worst possible point, immediately before the paper's contribution. We can tell you the topic and not the finding, which is the position the forty-third article described as this series' most common frustration.

What we can say is that the sentence frames surrogation as "a consequence of such alignment", meaning a cost of doing performance measurement properly rather than of doing it badly. That is consistent with the 2012 paper's framing of a by-product.

Four Versions Of Two Titles

A bibliographic note, recorded because this series keeps count and because this instance is unusual.

The publisher of The Accounting Review gives the 2012 title as Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation[1]. One of the authors' own university faculty pages gives it as Lost in Translation: An Examination of Surrogation in Strategic Performance Measurement Systems, with the same journal, volume, issue and page range[9].

The 2013 paper is given by its publisher as Strategy Selection, Surrogation, and Strategic Performance Measurement Systems[3], and on the same faculty page as Strategy Selection, Compensation, and Surrogation[9].

Separately, a bibliography lists the 2013 paper as (2012) and the 2012 paper as (2011), both shifted a year, while giving correct volumes, pages and DOIs[10].

Three observations, ours.

None of this changes anything. The DOIs, volumes, issues and page ranges are consistent across every source, so any reader can find the papers.

But the variants originate on an author's own institutional page, listing his own two papers under titles the journals did not print. That is not a citation error propagating through third parties; it is upstream of them.

And it is a small joke at the expense of this article's subject. A paper about measures detaching from what they represent is itself recorded under a title that has detached from the one published. That brings this series' running count of bibliographic variants to nineteen.

What Actually Survives

Our reading, stated directly.

Five statements.

The construct is well defined and worth having. Surrogation names something real that gaming does not cover, and the distinction changes which remedy applies.

The experimental evidence is thin. Two experiments, a video game task, no effect sizes or samples available to us, and no field evidence obtained.

The mechanism is derivable without the experiments. Everything in our arithmetic follows from the definition of a proxy and requires no psychology at all.

Incentives amplify rather than corrupt. The sign depends on whether actions exist that move the measure without moving the objective.

And the multiple-measures remedy is real and conditional. It works on independent errors and does almost nothing on correlated ones, and correlated is the normal case.

Your Own Dashboard

The application. Ours, untested, and not compensation design advice.

Four points.

Write down what each measure is a proxy for. If nobody can state the construct in a sentence that is not the measure itself, surrogation has already happened at the design stage.

List the actions that raise the measure without raising the construct. This is the single most useful exercise implied by our arithmetic, because the sign of the incentive effect depends on whether such actions are easy or hard to reach.

Check whether your measures share a source. Three metrics from the same system with the same reporting chain recover almost none of the benefit of having three, on our own model.

And stop expecting the dashboard to reveal the problem. In every row of our table the number rises. A measure that has stopped meaning anything looks, from the report, exactly like a measure that is working.

And Ours

Turning it inward, since this publication is produced by a firm that sells advisory work and measures itself. Ours.

Three examples we would not want to be asked about carelessly.

Realisation and utilisation rates are proxies for whether client work is profitable and whether capacity is deployed. Both can be raised by actions that do not improve either.

Turnaround time proxies responsiveness, and can be improved by narrowing what counts as the work.

And this publication's own article count, which the fiftieth article audited, is a proxy for research output and would be a poor thing to optimise. We noted there that the durable residue of fifty articles was procedural rather than topical, and the count itself was never the point.

The Diagnostic Question

The single question we would put to any measure, derived from the definition rather than the experiments. Ours.

"If this number improved and the thing it measures did not, would anyone notice?"

Four observations.

If the answer is no, you have no defence against surrogation, because the failure is undetectable by construction.

If the answer is yes, ask who would notice and how. The mechanism has to be named or it does not exist.

Note that the question does not require anyone to admit anything, which matters. Surrogation is sincere, so a question that presumes bad faith gets an honest denial and learns nothing.

And it works equally against gaming, which is the useful accident. A firm that can detect a measure moving without its construct has covered both failures, without needing to diagnose which one it has.

What To Do

Separate surrogation from gaming before choosing a remedy. Gaming knows the measure is not the goal and exploits the gap. Surrogation has lost the distinction. Monitoring catches the first and does nothing about the second.

Name the construct behind every measure, in writing. If it cannot be stated except by restating the measure, the substitution has already occurred.

Enumerate the measure-only actions. On our own arithmetic the sign of an incentive's effect turns on whether such actions are easily available.

Use three measures rather than one. On our own model that lifts a composite from about 0.70 to about 0.86, and the paper predicts and reports lower surrogation with multiple measures.

Check the three are not one. Measures sharing a source, a period or a reporting chain share their errors, and on our own model that recovers almost none of the benefit.

Do not read a rising number as evidence of anything. In every scenario we simulated, including total corruption of the proxy, the measure went up.

Ask whether anyone would notice the divergence. If nobody would, the measure is unmonitored regardless of how closely it is watched.

Treat the experimental evidence as suggestive, not settled. Two experiments with a video game task, no effect sizes available to us, and no field evidence obtained.

The Limits Of This Analysis

Several caveats matter. This article discusses research on performance measurement and is not compensation design, management accounting or human resources advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the 2012 paper, only its abstract from the publisher and a variant from a working paper repository; we report no effect size, no sample size, no statistical result and nothing from its results section, and grade its findings B on that basis. We obtained one truncated sentence of the 2013 paper's abstract and cannot report its finding. We did not obtain the attribute substitution framework's source paper. The detail that the experimental task was a video game comes from a separate paper citing it, not from the paper itself. We did not obtain Goodhart's 1975 article, the 1997 rephrasing, the 1996 generalisation, or any of Campbell's work, and report all of these from a reference encyclopedia confirmed where possible against a second source. Two sources here are not academic: a reference encyclopedia and an editorial website, both flagged at every use, and the surrogation-versus-Goodhart distinction as stated by the latter is our own analysis in its developed form. All arithmetic and simulation is ours, uses invented parameters throughout including a proxy correlation of 0.70 and an effort magnitude stated by no source, and demonstrates mechanisms rather than measuring any real firm. We made two errors building that simulation, using a loading where a squared loading was required and then comparing incentivised runs against the wrong baseline; both are corrected and described in the body, and the second would have produced a more quotable and wrong conclusion. And the arithmetic cannot confirm or refute the experimental findings; it shows why the predictions are plausible, which is a different thing.

Frequently Asked Questions

What is surrogation?
The paper that named it describes managers failing to fully appreciate that measures are merely representations of strategic constructs, and acting as though the measures are the construct of interest. It is a cognitive substitution, not a strategic choice.
How is that different from metric gaming?
Gaming requires knowing the measure is not the goal and exploiting the gap. Surrogation is losing that knowledge. The difference decides the remedy: monitoring works against gaming because there is a concealed intention to detect, and does nothing against surrogation because the behaviour is sincere.
Do incentives always corrupt a measure?
No, and this surprised us. On our own simulation, a proxy correlating 0.70 with its construct rose to 0.877 when incentivised effort raised both, and fell to 0.447 when the same effort raised only the measure. Incentives amplify in whichever direction the available actions point.
Why is this hard to detect?
Because in every scenario we simulated, including complete corruption of the proxy, the measured number went up. Comparing the metric to target, to prior period or to peers reports success in exactly the case where the measure has stopped meaning anything.
Does using more measures fix it?
Partly, and only under a condition rarely stated. On our own model three independent measures lift a composite from 0.70 to 0.86. But with 90 percent of the error shared across them, three measures are worth 0.712, which is barely better than one. Measures from a common source share their distortions.
What single question should I ask?
If this number improved and the thing it measures did not, would anyone notice? If no, the failure is undetectable by construction. If yes, name who and how. The question works against both surrogation and gaming, and it does not require anyone to admit anything.
How strong is the evidence?
Thinner than the idea's usefulness. Two experiments in a leading accounting journal supporting two hypotheses, with a video game as the task, and we obtained no effect sizes, no sample sizes and no results section. The arithmetic in this article is independent of those experiments and stands on its own.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article took three attempts to build its own simulation correctly, and the version we would have published after the second attempt was more quotable and wrong.

References

  1. Choi, J. (Willie), Hecht, G. W., & Tayler, W. B. (2012). Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation. The Accounting Review, 87(4), 1135–1163, DOI 10.2308/accr-10273, published 1 July 2012 by the American Accounting Association. Publisher record reproducing the abstract: on firms developing strategic performance measurement systems that translate strategy into performance measures to facilitate managers' decision-making; on managers ideally seeing measures for what they are, being imperfect proxies for intangible strategic constructs; on managers nonetheless possibly failing to fully appreciate that measures are merely representations of the strategic constructs, and acting as though the measures are the construct of interest, a phenomenon the authors label surrogation; on the paper investigating whether and how the use of strategically linked performance measures for compensation purposes affects managers' propensity to exhibit surrogation; on the authors predicting, in accordance with the attribute substitution framework attributed to Kahneman and Frederick (2002), that incentive compensation exacerbates surrogation and that this effect is more prevalent when managers are compensated on a single measure of a strategic construct than on multiple measures; on the authors finding support for these hypotheses via two experiments; and on the paper identifying a by-product of contracting on imperfect performance measures not previously considered in extant literature. Note: the publisher's record. We obtained the abstract only; we did not obtain the paper, its experiments, samples, statistics or any effect size, and report none. publications.aaahq.org
  2. Working paper repository record for the same study, listing it as an American Accounting Association 2010 Management Accounting Section meeting paper dated 24 May 2011, and reproducing a variant of the abstract: on the authors predicting, in accordance with the attribute substitution framework, that the tendency is most prevalent when managers are compensated on a single measure of a strategic construct and less prevalent when compensated on multiple measures; on support for these hypotheses being found via two experiments; and on surrogation being potentially detrimental to the implementation of strategy, such that the study furthers academic and practitioner understanding of factors affecting the extent to which firms incur this cost. Note: a working paper repository record for an earlier version. The abstract differs in wording from the published version at reference 1; we quote the published wording where the two diverge. doi.org
  3. Choi, J., Hecht, G. W., & Tayler, W. B. (2013). Strategy Selection, Surrogation, and Strategic Performance Measurement Systems. Journal of Accounting Research, 51(1), 105–133, DOI 10.1111/j.1475-679X.2012.00465.x. Publisher record reproducing the opening of the abstract, on strategic performance measurement systems operationalizing firm strategy with a set of performance measures, and on a consequence of such alignment being the tendency for managers to lose sight of the strategic constructs, our source truncating; together with the article's reference list confirming the 2012 paper's citation as The Accounting Review 87 (2012), 1135–63. Note: the publisher's record. Our source truncates immediately before the paper's contribution; we can report its topic and not its finding. onlinelibrary.wiley.com
  4. Repository page for the 2012 paper carrying scholarly citing text, recording that the citing authors' experiment follows a long line of experimental accounting research using abstract tasks to test psychological mechanisms and then inferring that results pertaining to these mechanisms generalize to real-world accounting tasks, and naming Choi, Hecht and Tayler (2012) as an example of this type of research, who use a video game for their experimental task; and separately recording that the study demonstrates that managers may focus on performance measures to the point of overlooking the strategic constructs they represent, a behavior termed surrogation. Note: a repository page reproducing text from papers citing the 2012 study. Our source for the video game detail, which does not come from the paper itself. academia.edu
  5. Reference encyclopedia entry on Goodhart's law, recording the adage as stated in the form that when a measure becomes a target it ceases to be a good measure, attributed to a 1997 work by Marilyn Strathern; recording that the law is named after the British economist Charles Goodhart, credited with expressing the core idea in a 1975 article on monetary policy in the United Kingdom in the words that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes; recording that Campbell's law likely has precedence, as Jeff Rodamar has argued, since various formulations date to 1969; recording a 1996 book chapter by Keith Hoskin stating that Goodhart's Law, that every measure which becomes a target becomes a bad measure, is inexorably if ruefully becoming recognized as one of the overriding laws of our times; and recording Mario Biagioli's application of the concept to citation impact measures, that all metrics of scientific evaluation are bound to be abused because people start to game them. Note: a reference encyclopedia, not an academic source, flagged at every use. We did not obtain the 1975 article, the 1997 rephrasing, the 1996 chapter, or any of Campbell's work. en.wikipedia.org
  6. Editorial website explainer on Goodhart's law, recording Charles Goodhart as an advisor to the Bank of England who made the observation in 1975 about monetary policy; recording that the Bank had found stable relationships between certain money-supply measures and inflation, and that when it began targeting those measures to control inflation the relationships broke down as banks and markets adapted; recording that the anthropologist Marilyn Strathern later generalized the point in 1997; recording examples of call centres measured on call duration, hospitals measured on emergency-room wait times, and schools measured on test scores; and stating that Goodhart's Law is the broad principle that targeting a proxy corrupts it while surrogation is a specific, quieter mechanism within it, being the cognitive slip where well-intentioned people lose the distinction, our source truncating. Note: an editorial website, not an academic source, flagged at every use. Used to corroborate the 1975 context and for one framing of the surrogation distinction; the developed form of that distinction in this article is our own analysis. whennotesfly.com
  7. Repository page for the 2012 paper carrying further scholarly citing text, recording that strategy surrogation occurs when managers focus on the measures in the strategic performance measurement system on which they are compensated and completely or partially lose focus on the overall strategic objectives of the organization, citing Choi and colleagues (2012, 2013); and recording that subsequent work has contributed to strategy surrogation research by examining institutional factors that may inhibit or exacerbate surrogation. Note: a repository page reproducing citing text. We obtained none of the citing studies. researchgate.net
  8. Reader-recorded excerpt from a published trade book on data skepticism, stating that the problem is canonized in a principle known as Goodhart's law; that while Goodhart's original formulation is a bit opaque, the anthropologist Marilyn Strathern rephrased it clearly and concisely as the statement that when a measure becomes a target it ceases to be a good measure; and that in other words, if sufficient rewards are attached to some measure, people will find ways to increase their scores one way or another, and in doing so will undercut the value of the measure for assessing what it was originally designed to assess. Note: a reader-submitted excerpt from a trade book on a book cataloguing site, not an academic source. Used only to corroborate the attribution and characterisation of the rephrasing; we did not obtain the book or Strathern's original. goodreads.com
  9. University faculty publication listing for the second author, recording Choi, J., Hecht, G., and Tayler, W. (2013), Strategy Selection, Compensation, and Surrogation, Journal of Accounting Research, 51(1), 105–133; and Choi, J., Hecht, G., and Tayler, W. (2012), Lost in Translation: An Examination of Surrogation in Strategic Performance Measurement Systems, The Accounting Review, 87(4), 1135–1163. Note: an author's own institutional page. Both titles differ from those printed by the respective journals, while journal, volume, issue and page ranges match. Recorded as a bibliographic variant originating upstream of any third party. sitefinitygies-campustheme.itpartners.illinois.edu
  10. Independent bibliography on surrogation, listing Choi, Hecht, and Tayler (2012), Strategy Selection, Surrogation, and Strategic Performance Measurement Systems, Journal of Accounting Research, 51(1), 105–133, DOI 10.1111/j.1475-679X.2012.00465.x; and Choi, Hecht, and Tayler (2011), Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review, 87(4), 1135–1163, DOI 10.2308/accr-10273. Note: an independent bibliography. Both years are shifted relative to the publishers' records, the 2013 paper being listed as 2012 and the 2012 paper as 2011, while titles, volumes, pages and DOIs are correct. Recorded as a second bibliographic variant. theinexactsciences.github.io

This article discusses research on performance measurement and is not compensation design, management accounting or human resources advice. Neither primary paper was obtained beyond its abstract, and no effect size, sample size or statistical result from either is reported. Goodhart's 1975 article, its 1997 rephrasing and Campbell's work were not obtained and are reported from a reference encyclopedia. Two sources are non-academic and flagged at every use. All arithmetic and simulation is the authors' own, uses invented parameters throughout, and demonstrates mechanisms rather than measuring any real firm; two errors made in its construction are corrected and described in the body.