A Canadian firm wants to know whether its AI investment is working. It asks the people using the tools. They report saving time, the finance function multiplies hours by rate, and a return appears in a board pack. Every step is standard practice, and the best available evidence indicates the first step produces an answer that can be wrong in direction.
Key Takeaway
METR conducted a randomised controlled trial in which 16 experienced open-source developers completed 246 real tasks on repositories they had contributed to for an average of five years, with each task randomly assigned to allow or disallow AI. Developers forecast that AI would reduce completion time by 24%. They actually took 19% longer, with a confidence interval from +2% to +39%. After completing the study and experiencing the slowdown, they still estimated that AI had reduced completion time by 20%. Roughly a quarter saw improved performance and three quarters saw reduced performance, and the one developer with more than fifty hours of tool experience saw positive speedup, suggesting a high skill ceiling. METR's own February 2026 follow-up found some evidence of speedup, with confidence intervals crossing zero, and announced a change of experiment design because developers increasingly refuse to work without AI, meaning the study systematically misses the most optimistic participants.
The Measurement Question
This article closes the human factors sequence in this series by asking how an organisation would know any of it.
The preceding articles described mechanisms: reviewers accepting incorrect output, experienced practitioners losing capability, knowledge migrating out of the firm, and automation creating work nobody counted. Each is an argument that the effects of AI deployment are not what they appear.
If that is right, then measurement matters more than usual, and the measurement method most organisations use is the one most likely to be affected. Asking people how much time a tool saved them is a self-report about a counterfactual, which is among the hardest things a person can be asked to estimate accurately.
The evidence base here is unusually good and unusually narrow. There is one prominent randomised controlled trial in expert knowledge work, and it is in software development rather than finance. This article reports it carefully, including the follow-up that partly complicates its headline, and then addresses what a Canadian finance function should do given that the direct evidence for its own domain does not exist.
The Trial
The design, which is the reason the result carries weight.
METR conducted a randomised controlled trial to understand how AI tools at the February to June 2025 frontier affect the productivity of experienced open-source developers, with 16 developers of moderate AI experience completing 246 tasks in mature projects on which they had an average of five years of prior experience, each task randomly assigned to allow or disallow usage of the tools[1].
Further detail: developers were recruited from large open-source repositories averaging over 23,000 stars and a million lines of code that they had contributed to for multiple years, and they provided lists of real issues that would be valuable to the repository, being bug fixes, features and refactors that would normally be part of their regular work. When AI was allowed they could use any tools they chose, primarily one commercial coding assistant with frontier models of the period; when disallowed they worked without generative assistance[2]. Tasks averaged 2.0 hours[3].
Commentary emphasises what distinguishes this from most evidence in the area: it was not a survey, not a self-reported productivity estimate and not a blog post, but a randomised controlled trial, the same methodology used in clinical drug trials[4]. Another notes the design was unusually close to real work, since developers were not asked to solve toy problems or timed coding puzzles but worked on issues from repositories they already knew well[5].
Two features matter for transferring anything to a professional services setting. Randomisation was at the task level within each developer, so each person served as their own control, which removes the confound of comparing enthusiastic adopters against reluctant ones. And the tasks were the participants' own real work rather than a constructed exercise.
The Result
What was found, with the precision the finding deserves.
The authors state that they initially were broadly expecting to see positive speedup, and that surprisingly they found that when developers use AI tools they take 19% longer than without[2]. The trial reports that tasks take 19% more time when participants can use the tools[3].
METR's own later publication gives the interval: the early 2025 study found the use of AI causes tasks to take 19% longer, with a confidence interval between +2% and +39%[6].
The interval is worth pausing on, because reporting the point estimate alone overstates the precision. The data are consistent with a slowdown as small as two percent and as large as thirty-nine percent. What the interval excludes, at the stated confidence, is a speedup.
The authors' prior expectation is also worth recording. This was not a study designed to find a negative result; the researchers expected the opposite and reported what they found. That is the ordinary standard of research conduct and it is worth noting in a field where most published measurement is produced by parties selling something.
Wrong About The Direction
The finding that matters most for how a firm assesses its own return.
Before starting tasks, developers forecast that allowing AI would reduce completion time by 24%. After completing the study, developers estimated that allowing AI reduced completion time by 20%[1].
One summary states the significance directly: even after completing their tasks, even after experiencing the slowdown firsthand, developers still estimated that AI had sped them up[4]. Another describes it as the perception-reality gap persisting even after task completion, indicating that subjective impressions of AI-assisted productivity are systematically miscalibrated[7]. A third calls the inverted self-assessment the durable result, noting that developers were wrong in direction about their own productivity under controlled conditions[8].
The distinction between miscalibrated magnitude and wrong direction is the whole point. A firm could work with estimates that are too optimistic by a predictable factor; you would discount them. An estimate with the wrong sign cannot be corrected by discounting, because no adjustment converts a reported gain into a measured loss.
Note also that the post-hoc estimate was barely different from the forecast. Experiencing the effect moved the estimate from 24% to 20%, in the wrong direction from reality by roughly forty percentage points on both occasions. Direct experience supplied almost no corrective information.
One commentary draws the conclusion for organisations: the trial challenges the common assumption that perceived speedup is a reliable substitute for measured speedup, and in this case it was not[5].
This is the same shape as the calibration finding this publication reported for models, where self-assessments moved upward after poor performance. The uncomfortable synthesis is that on this evidence neither the human nor the system reliably knows how well the human and system are doing.
Where The Time Went
The explanation offered, which connects directly to the function allocation argument in this series.
One source attributes the slowdown largely to the overhead of reviewing and integrating generated output[8].
That is the substitution myth measured. The preceding article in this series set out the argument that inserting a technology creates new tasks for the human, including entering inputs and monitoring, and that business cases count the work removed and not the work created. Here is a randomised trial in which the created work exceeded the work removed.
The specific form matters for a finance function. Reviewing and integrating generated output is exactly the shape of the work an AI-assisted professional workflow produces: a draft memorandum to check against sources, a classification to verify, an extraction to reconcile to the document. The activity that consumed the time in the trial is the activity a Canadian firm's AI workflow generates most of.
And the review burden scales with the volume of generated output rather than with the difficulty of the underlying problem, which means a system that produces more, faster, can increase total time even while producing correct work quickly.
The Metacognitive Load
A theoretical account of why the overhead exists, which links this article to the offloading discussion.
One source records that Tankelevitch and colleagues argued in 2024 that generative AI imposes substantial metacognitive demands on users, who must formulate prompts, evaluate outputs, and continuously calibrate their reliance on the tool[7].
Each of those three is cognitive work that did not exist before. Formulating a prompt is specifying a task precisely enough for a system that cannot ask for clarification the way a colleague would. Evaluating outputs is the verification burden. And continuously calibrating reliance is the metacognitive judgment the offloading article identified as itself a capability.
The reason this matters for measurement is that none of the three appears in a task inventory. They are not steps in the work; they are the cost of directing a system that performs the steps. A time study that measures the task will capture them only as unexplained duration.
Our observation is that this predicts a specific pattern a firm can look for: AI assistance helping most where the task is easy to specify and the output is cheap to verify, and helping least where specification is hard and verification is expensive. Much valuable professional work sits in the second category, which is a reason to expect finance results to differ from the simpler benchmarks.
A Quarter Got Faster
A distribution finding that the headline number conceals.
One account records that a quarter of the participants saw increased performance and three quarters saw reduced performance[2].
An average of minus nineteen percent across a group in which a quarter improved is a different fact from a uniform nineteen percent slowdown, and it has a direct implication for how a firm should measure.
A firm-level average tells you very little about whether to deploy, because it aggregates people the tool helps with people it hinders. If the same distribution held in a Canadian practice, a firm measuring only the average would conclude the tool does not work while a quarter of its people were benefiting, or conclude it works while most were slowed.
The practical instruction is to measure per person and examine the distribution before acting on the mean. That also permits the more useful question, which is what distinguishes the people the tool helps, and the next section suggests one candidate answer.
The Skill Ceiling
The observation that may reframe the entire result, reported with the tentativeness its source uses.
One account notes that one of the top performers for AI was also someone with the most previous experience of the tool, and that the paper acknowledges this: positive speedup was observed for the one developer with more than fifty hours of experience with the assistant, so it is plausible that there is a high skill ceiling for using it, such that developers with significant experience see positive speedup[2].
This is a single participant and the paper's own language is that it is plausible, so nothing is established. It is nonetheless the most interesting hypothesis in the study.
If tool proficiency is the moderating variable, then the trial measured a population in the middle of a learning curve, and the result describes a transition cost rather than a steady state. The participants were described as having moderate AI experience[1], which is consistent with that reading.
Two implications for a Canadian firm, offered as our own analysis. Measuring during the first months of adoption may capture the learning curve rather than the capability, so an early negative result should not be treated as final. And if there is a high skill ceiling, then the return depends on deliberate skill development in tool use, which is a different investment from the tool licence and is rarely budgeted.
We would resist the opposite over-reading, which is to dismiss any negative measurement as insufficient training. Fifty hours is a substantial investment per person, and a firm should establish whether it is prepared to make it rather than assume the curve will be climbed incidentally.
What The Trial Does Not Cover
The boundaries, which METR states explicitly and which commentary sometimes omits.
One summary records that METR is explicit about scope limits: the trial studied experienced contributors in mature, high-quality codebases using early-2025 tools, and does not apply to junior developers, greenfield projects, enterprise teams, or later AI generations. METR itself labels the result historical[9]. METR's own framing describes the result as a snapshot of early-2025 AI capabilities in one relevant setting, as these systems continue to rapidly evolve[2].
Every one of those limits matters for a Canadian finance function reading across.
Experienced contributors on familiar material is close to the profile of a senior practitioner on recurring work, which is favourable for transfer. Mature high-quality codebases has no clean analogue. Junior staff are excluded, and the effect may well differ for them in either direction. And the tools are from a specific window that has passed.
Our position is that the result should be treated as strong evidence that measured and perceived productivity can diverge in direction in expert knowledge work, and as weak evidence about the magnitude or sign of the effect in Canadian professional services today. The first proposition is the one this article rests on, and it is the one least sensitive to the scope limits.
The Follow-Up That Complicates It
The development that a fair treatment has to report, because it partly cuts against the headline.
METR published new data in February 2026 on the productivity impact of late-2025 tools[2]. In that update it states that raw results show some evidence for speedup: for the subset of the original developers who participated in the later study, it now estimates a speedup of minus 18%, with a confidence interval between minus 38% and plus 9%, and among newly-recruited developers the estimated speedup is minus 4%, with a confidence interval between minus 15% and plus 9%[6].
In METR's notation a negative figure indicates a reduction in completion time, so both estimates point toward AI making tasks faster, reversing the direction of the earlier finding.
Both confidence intervals cross zero, which means neither estimate is statistically distinguishable from no effect at the stated confidence. The honest reading is that the later data suggest improvement and do not establish it.
The direction of travel is nonetheless meaningful. Between the two studies the tools changed generation, and the estimate moved from a measured slowdown to a possible speedup. That is consistent with the view that the original result was a snapshot of a particular capability level rather than a durable property of AI assistance.
A firm citing the nineteen percent slowdown as current evidence about today's tools would be misusing it, and METR's own labelling of the result as historical is the appropriate caution.
The Selection Spiral
The most methodologically interesting thing in the follow-up, and a genuine problem for the whole field.
METR reports two effects of wider adoption on its study: recruitment and retention of developers has become more difficult, and an increased share of developers say they would not want to do 50% of their work without AI, even though the study pays them fifty dollars an hour to work on tasks of their own choosing. It concludes that the study is thus systematically missing developers who have the most optimistic expectations about AI's value, and that the true speedup could be much higher among the developers and tasks which are selected out of the experiment[6].
Read the mechanism carefully. The randomised design requires participants to work without AI on half their tasks. As reliance grows, the people who rely most decline to participate, at a rate that payment does not overcome.
So the sample is filtered by willingness to work without the tool, which correlates with the effect being measured. The population most likely to show large gains is the population least likely to enrol.
METR's candour here is notable: it is reporting a bias that runs against its own headline result, and stating that the true effect may be larger than it measured. That is the behaviour that makes the rest of its reporting credible.
When The Method Itself Degrades
The broader implication, offered as our own analysis.
Randomisation is the gold standard because it removes selection. The selection spiral describes a case where randomisation cannot be applied without selecting on the outcome, because assignment to the control condition is itself something participants will refuse.
That is a structural problem rather than a design flaw, and it worsens as adoption rises. The better the tools become and the more embedded they are, the harder it becomes to run the experiment that would establish how much they help.
METR's response is to change the experiment design, and the update discusses alternatives including continued observation of its developer pool, questionnaires, and fixed-task experiments, noting that its original study was novel in letting developers choose their own tasks before randomisation was applied. On surveys it says that despite the difficulty of interpreting self-reported productivity counterfactuals and their potential biases, careful choice of survey questions along with time-use studies could produce useful signals[6].
That last sentence deserves attention from anyone who reads this article as saying self-report is worthless. The researchers who demonstrated the direction error are not abandoning surveys; they are saying the instrument needs care and should be paired with time-use data.
For a Canadian firm the transferable lesson is that the same refusal dynamic will affect internal measurement. Asking practitioners to complete comparable work without the tool for measurement purposes will meet resistance, and the people who resist most are the ones whose data would be most informative.
Two Narratives That Cannot Both Be Right
The contrast one commentary draws, which describes the information environment a Canadian finance leader is operating in.
The commentary sets a series of claims against each other: controlled experiments finding slowdowns and skill loss on one side, and executives on earnings calls claiming revolutionary gains on the other, including a report that one platform company's co-chief executive told analysts his best developers had not written a line of code since a given month and only generate code and supervise it. Its conclusion is that both cannot be entirely right[4].
The same source reports that researchers at one AI developer found AI assistance impairs conceptual understanding, code reading and debugging, with no significant efficiency gains on average[4]. We report that at second hand and have not accessed the underlying work, so it should be treated as a claim rather than an established finding.
Our reading is that the two narratives are not straightforwardly contradictory, and that the reconciliation is instructive. Executive claims typically concern what people now do rather than how long it takes: a developer who supervises rather than writes has changed activity, which is a real change and not a measured productivity gain. The controlled experiments measure completion time on comparable tasks, which is a narrower and harder question.
A Canadian firm should notice which kind of claim it is making internally. "Our people now work differently" is usually demonstrable. "Our people are thirty percent more productive" is a measurement claim, and the evidence above indicates it cannot be supported by asking them.
Productivity Is One Dimension
The qualification that keeps this article from overreaching, and it comes from the commentary rather than from us.
One source notes that productivity is only one dimension of tool use, since a person might use AI because it makes work more pleasant, helps with exploration, lowers friction, teaches unfamiliar material, or makes some tasks less tedious, and that those benefits may matter even when a narrow timing measure does not improve[5].
That is right and it deserves weight rather than a passing mention. A tool that leaves completion time unchanged while making work less unpleasant may be worth its cost, particularly in a Canadian professional services market where retention is a live constraint.
The problem is not organisations valuing those benefits. It is organisations valuing those benefits and describing them as a productivity gain, because productivity is what a business case format expects.
The recommendation that follows is to state what you are buying. If the purchase is speed, measure speed. If it is tolerability, capability extension into unfamiliar areas, or reduced friction on tedious work, say so and evaluate it on those terms. Both are defensible; conflating them produces a business case that cannot be verified and a disappointment that cannot be diagnosed.
What A Finance Function Should Measure
Translating the evidence into a workable internal method, offered as our own analysis.
Cycle time on comparable work, recorded rather than recalled. The direction error concerns estimates. Elapsed time on a defined unit of work is an observation, and most professional firms already capture time.
Define the unit before deployment. A comparable unit of work has to be specified while the pre-deployment version still exists, which is the baseline problem discussed below.
Measure per person and look at the distribution. A quarter of participants improved while three quarters did not; the mean would have concealed both.
Do not measure only at the start. The skill ceiling hypothesis implies early measurement may capture a learning curve, so measure repeatedly.
Capture rework separately. The mechanism identified was the overhead of reviewing and integrating generated output, so a measure that stops at first draft will miss where the time went.
Pair any survey with time data. The researchers who demonstrated the direction error still see value in careful survey design combined with time-use studies.
Expect resistance to unassisted comparison. Participants declined to work without the tool even when paid, and the same dynamic will operate internally, so the design should minimise how much unassisted work is required.
The Baseline You No Longer Have
The practical obstacle most Canadian firms have already walked into, and it is our own observation.
Measuring a change requires a before. Most firms deployed AI capability without recording a pre-deployment baseline, because deployment was incremental, because features appeared inside software already in use, and because nobody anticipated needing the comparison.
The result is that a firm asking today whether AI improved its productivity generally cannot answer, and the reason is not analytical difficulty but the absence of the earlier measurement. This is the same structural point the incident response article made about logging: the decision that determines whether a question can be answered was taken before anyone asked it.
Two things can still be done. A firm can establish a baseline now, for the current configuration, so that future changes are measurable even though the original transition is not. And it can use the trial's within-person design on a small scale, asking a few practitioners to complete a defined unit of comparable work without the tool, accepting that this is expensive and will meet resistance.
What it should not do is substitute recollection for the missing baseline, since the evidence indicates recollection of this specific quantity can be wrong in direction.
A Worked Case: The Survey That Proved Nothing
A Canadian firm evaluating its AI investment after twelve months. The reconstruction illustrates the reasoning rather than reporting a specific engagement.
The firm surveys its professionals, asking how much time the tools save on a typical engagement. The average response is a meaningful percentage. Multiplied by chargeable hours and rates, the return is comfortable, and the board approves continued investment.
Every element of that exercise is standard, and the evidence indicates the input is unreliable in a specific way. In the trial, participants estimated a twenty percent reduction after experiencing a nineteen percent increase[1], an error not of magnitude but of sign, and one that direct experience did not correct[4].
The firm also has no baseline, so it cannot check the survey against elapsed time on comparable work, and the pre-deployment version of that work no longer exists in its records.
What the firm does have is real: people report preferring the tools, work is described as less tedious, and nobody wants to go back. Those are the benefits the commentary identifies as mattering independently of a timing measure[5], and they would support a decision to continue on stated grounds.
The defect is not the decision. It is that the board approved a productivity return the firm has not measured and, on this evidence, cannot infer from what it asked.
What To Do
Stop treating self-reported time savings as measurement. The one randomised trial in expert knowledge work found the estimate wrong in direction, before and after experiencing the effect.
Establish a baseline now, even though the original one is gone. It makes the next change measurable, which is the best available position.
Measure elapsed time on a defined comparable unit. Recorded, not recalled, and specified precisely enough to be repeatable.
Look at the distribution, not the mean. A quarter improving while three quarters did not is a different situation from a uniform effect, and requires a different response.
Re-measure over time. The skill ceiling hypothesis means an early result may describe a learning curve rather than a capability.
Budget for tool proficiency, not just licences. If fifty hours of experience is the threshold at which returns appear, that is an investment decision to make deliberately.
Include review and rework in the measure. The identified mechanism was the overhead of reviewing and integrating output, which sits after the point most measures stop.
Say what you are buying. Speed, tolerability, capability extension and friction reduction are all legitimate purchases, and only the first is a productivity claim requiring measurement.
Do not cite the nineteen percent figure as current. The researchers label it historical, and their own later data point the other way without establishing it.
The Limits Of This Analysis
Several caveats matter. The central study concerns software development, not finance, and its authors are explicit that it does not apply to junior developers, greenfield projects, enterprise teams or later AI generations; the transfer of any magnitude to Canadian professional services is not supported, and we rest only on the narrower proposition that measured and perceived productivity can diverge in direction. The sample is 16 developers and 246 tasks, and the confidence interval on the headline figure runs from +2% to +39%. The skill ceiling observation rests on a single participant and the paper's own language is that it is plausible. The February 2026 follow-up estimates point toward speedup but both confidence intervals cross zero, so nothing is established in that direction either, and METR has changed its experimental design because the data became hard to interpret. We accessed the study through its abstract, METR's published summaries and secondary commentary rather than the full paper, and several supporting items, including the Tankelevitch metacognitive framing and the reported finding from researchers at one AI developer, reach us at second hand and were not independently verified. One cited source is a personal blog and several are commercial publications. The finance measurement programme, the baseline argument, the specification and verification prediction, the reading of the two narratives and the worked case are our own analysis. This article does not address cost accounting for AI, licence and inference economics, or benefit realisation methodology, which this publication treats separately. Nothing here is a substitute for professional advice on investment appraisal.
Frequently Asked Questions
What did the randomised trial find?
Why does the direction error matter more than the size?
Where did the extra time go?
Does this mean AI does not help?
What is the selection spiral?
What should we measure instead of surveying?
References
- Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv 2507.09089, abstract, on the randomised controlled trial design, 16 developers with moderate AI experience completing 246 tasks in mature projects with an average of five years of prior experience, the 24% forecast speedup and 20% post-study estimate. Note: accessed via abstract rather than the full paper. arxiv.org/abs/2507.09089
- METR. (2025, July 10). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, on recruitment from repositories averaging 23,000 stars and a million lines, the 246 real issues, the random assignment design, the tools used, the authors' prior expectation of positive speedup, the 19% slowdown, the characterisation as a snapshot of early-2025 capabilities, and the February 2026 follow-up publication. metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study
- METR. Early 2025 AI Experienced OS Devs Study, paper, on tasks taking 19% more time and the 246 tasks averaging 2.0 hours on repositories averaging 23,000 stars. metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf
- Let's Data Science. (2026, February 20). Developers Thought AI Made Them Faster, The Data Said Otherwise, on the randomised controlled trial as the gold standard of scientific evidence, the persistence of the belief after experiencing the slowdown, the reported finding from researchers at one AI developer on impaired conceptual understanding and no significant efficiency gains, and the contrast with executive claims. Note: a commercial publication; the AI developer finding is reported at second hand and unverified. letsdatascience.com/blog/developers-thought-ai-made-them-faster-the-data-said-otherwise
- ScienceBlog. (2026, July 20), on the study design being unusually close to real work, the challenge to the assumption that perceived speedup substitutes for measured speedup, and the observation that productivity is only one dimension of tool use alongside pleasantness, exploration, friction and learning. scienceblog.com/t-a-randomized-trial-by-metr-found...
- METR. (2026, February 24). We Are Changing Our Developer Productivity Experiment Design, on the +2% to +39% confidence interval for the original 19% figure, the later estimates of minus 18% and minus 4% with intervals crossing zero, the difficulty of recruitment and retention, developers declining to work without AI at fifty dollars an hour, the systematic loss of the most optimistic participants, and the position on questionnaires and time-use studies. metr.org/blog/2026-02-24-uplift-update
- Using Biometrics to Understand AI-Assisted Coding Performance and its Perception. arXiv preprint 2606.20598, on the perception-reality gap persisting after task completion and subjective impressions being systematically miscalibrated, and on Tankelevitch and colleagues (2024) arguing that generative AI imposes substantial metacognitive demands including formulating prompts, evaluating outputs and calibrating reliance. Note: preprint; the Tankelevitch work reported at second hand. arxiv.org/pdf/2606.20598
- The Substrate Collapse: AI Code Generation Invalidates Authorship-Based Knowledge Metrics. arXiv preprint 2606.20882, on the slowdown being traced largely to the overhead of reviewing and integrating generated output, and on the inverted self-assessment as the durable result with developers wrong in direction under controlled conditions. Note: preprint. arxiv.org/pdf/2606.20882
- Thrumos. METR Study: AI Tools Slow Developers 19% Despite 20% Speedup Perception, on METR's explicit scope limits excluding junior developers, greenfield projects, enterprise teams and later AI generations, the labelling of the result as historical, and the caution against vibes-based adoption metrics. Note: a commercial publication. thrumos.com/insights/ai-makes-developers-19-percent-slower-metr-rct
- Willison, S. (2025, July 12). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, on METR's institutional background, the quarter of participants seeing increased performance against three quarters seeing reduced performance, and the positive speedup for the one developer with more than fifty hours of tool experience with the paper's acknowledgement of a plausible high skill ceiling. Note: a personal blog. simonwillison.net/2025/Jul/12/ai-open-source-productivity
This article discusses productivity measurement research and is provided for general informational purposes. The central study concerns software development rather than finance, its authors state it does not apply to junior staff, new projects, enterprise teams or later AI generations, and they label the result historical. Later data from the same researchers point toward speedup with confidence intervals crossing zero. Nothing here is a substitute for professional advice on investment appraisal.