Nineteen articles into this series, the recurring shape has been a finding that did not hold, a dispute that never closed, or a number nobody could trace. This one is different. Two researchers who had published opposite answers sat down together and worked out why.

Key Takeaway

"Do larger incomes make people happier? Two authors of the present paper have published contradictory answers." They engaged in an adversarial collaboration to search for a coherent interpretation of both studies. The reanalysis confirmed the flattening pattern only for the least happy people, while happiness increases steadily with log(income) among happier people, and even accelerates in the happiest group. The conflict arose from mislabeling of the dependent variable and the incorrect assumption of homogeneity, both described as practices that are standard in social science but should be questioned more often[1][2].

A Note On Scope

Stated first, because of what this article touches.

This is a description of research about income and reported wellbeing. It is not financial, medical or psychological advice, and nothing here should be used to draw conclusions about any individual's circumstances, including the reader's.

One passage of the research quoted below concerns bereavement and clinical depression. We report it because it is the authors' own framing of their most important finding, and because omitting it would misrepresent the result. It is not a clinical statement and neither is this article.

If you are struggling, speak to a qualified professional. No published study is a substitute for that, and the research below says so more clearly than most.

Our Grades For These Claims

Applying the scheme from the first article in this series.

Grade A for the resolution. Two research teams pooled raw data from two large datasets, appointed a neutral facilitator, and published a joint reanalysis explaining both prior results. We have not encountered a stronger design for settling a disagreement.

Grade A for the underlying relationship between income and reported wellbeing, which both original studies found and which the joint paper confirms in the same direction.

Grade B for the specific shape, including the acceleration in the happiest group, because a subsequent paper challenges the functional form.

Any specific dollar threshold is Grade C. The famous figure comes from a 2010 study whose authors now say their measures could not discriminate among degrees of happiness.

A Note On Method

Everything here is verified to August 2026.

We obtained the published abstract in full from the journal itself and from PubMed, plus the opening of the paper[1][2], and an announcement from one author's own institution[3].

We did not obtain the full text of any of the three papers, and report no effect sizes from any of them.

One characterisation of the effect's magnitude, and one sample figure, reach us through commercial websites, and we flag both as unverified where they appear.

We also located a subsequent paper challenging the resolution, and report it, because a series about evidence should not present a paper titled "conflict resolved" as the end of a literature.

All arithmetic on logarithms is ours.

The Two Findings

The disagreement, in the joint paper's own words.

"Using dichotomous questions about the preceding day, [Kahneman and Deaton, 2010] reported a flattening pattern: happiness increased steadily with log(income) up to a threshold and then plateaued."[2]

"Using experience sampling with a continuous scale, [Killingsworth, 2021] reported a linear-log pattern in which average happiness rose consistently with log(income)."[2]

The two papers are Kahneman and Deaton (2010), High income improves evaluation of life but not emotional well-being, and Killingsworth (2021), Experienced well-being rises with income, even above $75,000 per year, both in PNAS[1].

Two observations, ours.

The second paper's title is a direct contradiction of the first. That is unusually explicit for academic publishing.

And notice that the joint paper's abstract leads with the measurement difference, before stating either result. Dichotomous questions about yesterday against experience sampling on a continuous scale. That placement is a signal about where the answer turned out to be.

The Measurement Difference

The two instruments, and why the difference is not a detail. This explanation is ours.

The 2010 study asked dichotomous questions about the preceding day[2]. Did you experience a lot of happiness yesterday, yes or no. It drew on more than 450,000 responses[3].

The 2021 study used experience sampling with a continuous scale[2]: an app pinging people at random moments asking how they feel right now, on a slider.

Three consequences.

A yes-or-no question about a whole day cannot register degrees. Someone mildly content and someone delighted both answer yes.

It also relies on recall of a whole day, whereas the sampling method asks about the present moment.

And the first is bounded in a way the second is not. Once you are answering yes, there is no further yes available, which is a point the joint paper returns to and which becomes the crux of the whole resolution.

What Log Income Actually Means

A point that changes how the whole debate reads, and it is almost always dropped. Our own arithmetic.

Both studies model happiness against log(income), not income[2].

A log-linear relationship means each doubling of income adds the same amount. So on this model, moving from $25,000 to $50,000 buys the same increment as moving from $400,000 to $800,000.

Two consequences.

"Happiness keeps rising with income, with no plateau" and "each additional dollar buys steadily less" are both true simultaneously, in the same finding. They are not in tension, and a great deal of popular commentary treats them as though they were.

Which means the practically relevant question for anyone is not whether there is a ceiling. It is how many doublings away you are, and for most people the answer is that the next one is a long way off.

What They Did Instead Of Arguing

The procedure, which is the reason this article exists.

The joint paper states: "We engaged in an adversarial collaboration to search for a coherent interpretation of both studies."[2] And separately: "We engaged in an adversarial collaboration and asked Barbara Mellers to be the facilitator."[1]

The resulting paper is Killingsworth, Kahneman and Mellers (2023), Income and emotional well-being: A conflict resolved, PNAS, 120(10), e2208661120[2].

Four features, and this assessment is ours.

Both parties bring their raw data. Neither is arguing from the other's published summary.

A neutral third party facilitates. Mellers is a co-author of the resulting paper, not an outside reviewer of it.

The analysis is agreed in advance of seeing whose side it favours, which is the whole point and is the same protection that preregistration provides in the replication literature.

And the result is published jointly, so neither party can later characterise it selectively.

Compare that with the disputes in this series' fourth, ninth and eighteenth articles, each of which produced a comment, a response, and no resolution.

The Resolution

What they found.

"A reanalysis of Killingsworth's experienced sampling data confirmed the flattening pattern only for the least happy people. Happiness increases steadily with log(income) among happier people, and even accelerates in the happiest group. Complementary nonlinearities contribute to the overall linear-log relationship."[2]

An institutional summary adds that they examined the least happy 20 percent of the sampling respondents, and found the flattening where the Kahneman-Deaton study would have anticipated finding it, adjusted for inflation[3].

Three observations, ours.

Both prior findings were right about different people. The plateau is real, in a minority. The continued rise is real, in the majority.

The phrase complementary nonlinearities explains the whole puzzle. Two opposite curvatures in different subgroups average out to something that looks straight, which is why one study saw a line and the other saw a bend.

And accelerates in the happiest group is a finding neither original study reported. The collaboration produced something new rather than merely adjudicating.

Why The First Study Overstated It

The diagnosis of the 2010 paper, delivered by one of its own authors.

The joint paper states: "We then explain why Kahneman and Deaton overstated the flattening pattern and why Killingsworth failed to find it."[2]

On the first: "their measures could not discriminate among degrees of happiness because of a ceiling effect."[2]

Two observations, ours.

Recall the yes-or-no instrument. Once enough people are saying yes, additional happiness has nowhere to register. The measure stops moving before the thing it measures does, and the result looks like a plateau in the world when it is a plateau in the questionnaire.

And note that this is the same problem the eighteenth article in this series described being raised against the endowment effect experiments: ceiling effects caused by short and easy tests producing apparent interactions. Two entirely separate literatures, the same measurement failure.

The Reframe That Fixes It

The most elegant sentence in the paper.

"We suggest that Kahneman and Deaton might have reached the correct conclusion if they had described their results in terms of unhappiness rather than happiness."[2]

The institutional summary puts the corrected reading plainly. Instead of "happiness rises with income, but there is no further progress beyond $75,000", the data more accurately showed that an increase of income staved off unhappiness for a while but not after the flattening point[3].

Three observations, ours.

The data were fine. The measurement was fine for what it could do. What was wrong was the label on the variable, and therefore the sentence written about it.

A yes-or-no question about a bad day is good at detecting misery and bad at detecting degrees of contentment. So it was always an unhappiness instrument being reported as a happiness instrument.

And that is a remarkably specific and humble correction for an author to make about his own most-cited finding, thirteen years later.

The Sentence That Matters Most

The substantive result, quoted rather than paraphrased.

The authors, as reported by an institutional summary: "This income threshold may represent the point beyond which the miseries that remain are not alleviated by high income. Heartbreak, bereavement, and clinical depression may be examples of such miseries."[3]

Three observations, ours, offered carefully.

This is the finding, not a caveat to it. The plateau is real and it is located precisely: it is where money stops helping with the kinds of suffering money cannot reach.

It is more useful than the popular version. "Money stops buying happiness above seventy-five thousand" is both wrong and unactionable. "Above a certain income, the remaining sources of misery are ones income does not address" is correct and tells you something.

And we note the obvious: the examples given are matters for professional help, not financial planning. That is the authors' point and this article does not extend it.

Why The Second Study Missed It

The other half of the diagnosis.

The paper attributes the conflict to two causes: "The mislabeling of the dependent variable and the incorrect assumption of homogeneity."[2]

The first, mislabeling, explains the 2010 paper. The second explains the 2021 one.

Explaining it, and this explanation is ours. Assuming homogeneity means treating the sample as one population in which the relationship has a single shape. Fit one curve to everybody.

Two consequences.

If the relationship genuinely differs between subgroups, an average curve describes nobody. It is the arithmetic mean of two different truths.

And the flattening in the least happy fifth was present in the 2021 data all along. It was not detected because nobody looked separately. That is not a data problem; it is an analysis choice, and it is the same choice almost every study makes by default.

Two Standard Practices

The generalisation the authors draw, which is why this paper matters beyond its topic.

"The mislabeling of the dependent variable and the incorrect assumption of homogeneity were consequences of practices that are standard in social science but should be questioned more often. We flag the benefits of adversarial collaboration."[2]

Three observations, ours.

Neither cause is misconduct, error or incompetence. Both are ordinary practice, applied by careful people, producing a decade-long public contradiction.

Which is the same conclusion the first article in this series reached about the replication crisis: the failures were generated by normal method rather than by bad actors.

And this series has now hit the homogeneity problem repeatedly, in the choice overload literature, in the conflict literature, in growth mindset, in social norms, and here. An average across two populations moving differently is the single most recurrent trap in the whole field.

This Has A Tradition

Adversarial collaboration is not a one-off, and the reference list makes that clear.

The paper's references include Mellers, Hertwig and Kahneman (2001), Do frequency representations eliminate conjunction effects? An exercise in adversarial collaboration; Gilovich, Medvec and Kahneman (1998), Varieties of regret: A debate and partial resolution; Kahneman and Klein (2009), Conditions for intuitive expertise: A failure to disagree; Cowan and colleagues (2020), How do scientific views change? Notes from an extended adversarial collaboration; and Ellemers and colleagues (2020), Adversarial alignment enables competing models to engage in cooperative theory building toward cumulative science[4].

We obtained none of these and report titles and citations only.

Two observations, ours.

Look at the titles. "A debate and partial resolution." "A failure to disagree." These are honest about producing incomplete answers, which is itself a mark of the format.

And one name recurs across four decades of them. Whatever else is true, the practice had a persistent advocate, which is probably what a method needs to survive.

Resolved Is Not Final

Something a series about evidence has to report.

A subsequent working paper is titled Income and emotional well-being: Evidence for well-being plateauing around $200,000 per year. It states that using linear regression, Killingsworth and colleagues (2023) validates the findings of Killingsworth (2021), that emotional well-being increases monotonically with log-income without reaching a plateau, and then argues for a plateau at a different level[5].

We did not obtain this paper beyond its abstract and take no view on it.

Two observations, ours.

A paper titled "A conflict resolved" already has a challenger. That is not a criticism of it; it is how the process works.

And it is a caution against this article's own framing. We have presented an adversarial collaboration as a model, and it is. But a model procedure does not produce a final answer, and treating "resolved" in a title as the end of a literature would repeat exactly the error this series keeps documenting.

How Big Is The Effect

The question a reader will want answered, and our honest position on it.

We did not obtain effect sizes from any of the three papers and state none.

We located one characterisation on a commercial website, which describes money's effect on happiness as real but modest, holding that a four-fold income difference matters less than a headache, about as much as being a caregiver, and twice as much as being married[6]. The same source gives the joint paper's sample as 33,391 working US adults with a median household income of $85,000[6].

Both are unverified. They come from a commercial website, not from the papers, and we could not check either.

Two observations, ours.

If the comparison is even roughly right, it is the most useful calibration in the topic, because it puts income alongside things a reader can price against their own experience.

And a four-fold difference is two doublings, on our own arithmetic. So the claim being made is that quadrupling your income moves your day less than a headache does. We would want to see that in the paper before repeating it, and we are not repeating it as fact.

For A Business Owner

The application, and it is our own extension of research conducted on employed adults.

Four observations.

The relationship is logarithmic, so the wellbeing return on the next increment of income depends entirely on where you are starting. A business owner deciding whether a year of additional effort is worth it is choosing a fraction of a doubling, not a doubling.

The plateau, where it exists, is about unhappiness rather than happiness. Income relieves the miseries income can relieve and then stops. That is a statement about which problems money solves, not about a ceiling on satisfaction.

The finding that happiness accelerates in the happiest group should be read with care. It is a description of a sample, not a promise, and it says nothing about what any individual gains from any particular decision.

And the studies concern income, not business ownership, wealth, exit proceeds, or the experience of running a firm. Nothing in this literature was measured on business owners, and this section is extrapolation.

Why This Is The Model

The reason this article closes the first fifth of the series. Ours.

Nineteen previous articles have documented a consistent set of failures. Findings that did not replicate. Meta-analyses that reached opposite conclusions and stayed there. Effect sizes nobody could trace to a paper. A study cited as authority for the claim it refuted.

Four things this paper did differently.

The parties identified their disagreement precisely rather than talking past each other.

They pooled raw data instead of arguing over published summaries.

They appointed a referee before knowing the answer.

And they diagnosed the cause as ordinary practice rather than as the other side's error, which is what allowed both prior papers to be substantially correct at once.

Our own view: this is what the disputes in the fourth, ninth and eighteenth articles of this series should have become, and did not. The obstacle is not intellectual. It is that the format requires each party to accept in advance that they might be shown to be wrong, in print, jointly, with their own data.

What To Do

Drop the seventy-five thousand dollar figure. Its own author now says the measure could not discriminate among degrees of happiness because of a ceiling effect.

Read "no plateau" correctly. The relationship is logarithmic, so happiness continuing to rise and each dollar buying less are both true at once.

Ask whether an average describes anybody. The conflict existed because two subgroups moved differently and one curve was fitted to both, which is the most recurrent trap this series has found.

Check what a measure can register. A yes-or-no question cannot detect degrees, so a flat line may be a flat instrument.

Check what the variable is actually called. The 2010 result was correct about unhappiness and reported as being about happiness, and that single label caused a decade of dispute.

When you disagree with someone, pool the data. The procedure is available to any two people who both have records and a willingness to be shown wrong.

Appoint the referee before you see the answer. That is the part that makes it work, and it is the part that is uncomfortable.

Do not treat "resolved" as final. This paper already has a challenger, and that is the process functioning rather than failing.

Note which problems income does not solve. The authors name heartbreak, bereavement and clinical depression, and those are matters for professional support rather than for planning.

The Limits Of This Analysis

Several caveats matter. This article describes research on income and reported wellbeing. It is not financial, medical or psychological advice, and nothing in it should be used to draw conclusions about any individual, including the reader. Anyone struggling should speak to a qualified professional. Everything is verified to August 2026. We obtained the published abstract of the 2023 paper in full and its opening, but did not obtain the full text of it or of either 2010 or 2021 paper, and we state no effect sizes from any of them. The characterisation of the effect's magnitude and the sample figure both come from a commercial website, are unverified, and we do not rely on either. The institutional announcement we cite is from an author's own university and is not a neutral source. We obtained none of the adversarial collaboration precedents and report titles and citations only. We located a subsequent paper challenging the resolution and obtained only its abstract, taking no view on it. All arithmetic on logarithms is ours. The research concerns income among employed adults, largely in the United States; nothing in it was measured on business owners, and the section applying it to them is our own extrapolation. The reading of what the measurement difference implies, the account of why homogeneity assumptions fail, and the comparison with earlier articles in this series are our own reasoning, not findings.

Frequently Asked Questions

Does money stop buying happiness above seventy-five thousand dollars?
No. The joint reanalysis found the flattening only among the least happy people, with happiness continuing to rise with log income among happier people and even accelerating in the happiest group. One of the original authors now says their measure could not discriminate among degrees of happiness because of a ceiling effect.
So the plateau was wrong?
It was real but mislabelled. The authors suggest the original conclusion would have been correct if described in terms of unhappiness rather than happiness: income staved off unhappiness up to a point and then stopped. The remaining miseries, they name heartbreak, bereavement and clinical depression, are not ones income addresses.
What is an adversarial collaboration?
Researchers who have published contradictory findings pool their raw data, appoint a neutral facilitator, agree the analysis before seeing whose side it favours, and publish jointly. Here the facilitator was Barbara Mellers, who is a co-author of the resulting paper rather than a reviewer of it.
How did two careful studies contradict each other?
Through two ordinary practices: mislabelling the dependent variable, and assuming the relationship had the same shape for everyone. The authors describe both as standard in social science but deserving to be questioned more often. Neither is misconduct or error.
Is the question settled now?
No. We located a subsequent paper arguing for a plateau at a different income level. A paper titled "A conflict resolved" already has a challenger, which is the process working rather than failing, and treating a title as the end of a literature is the error this series keeps documenting.
What does this mean for me as a business owner?
Less than you might hope, honestly. The relationship is logarithmic, so what matters is how many doublings away the next increment is. And nothing in this research was measured on business owners, so applying it to a decision about your own firm is extrapolation.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This is the twentieth article in a hundred-part series and the first to describe a scientific dispute being settled properly. It closes by noting that the paper titled "A conflict resolved" already has a challenger.

References

  1. Killingsworth, M. A., Kahneman, D., & Mellers, B. (2023). Income and emotional well-being: A conflict resolved. Proceedings of the National Academy of Sciences, 120(10), e2208661120. DOI 10.1073/pnas.2208661120, journal record, reproducing the abstract and the opening of the paper, on the authors having engaged in an adversarial collaboration and asked Barbara Mellers to be the facilitator; and giving full citations for Kahneman, D., & Deaton, A. (2010), High income improves evaluation of life but not emotional well-being, PNAS 107, 16489–16493, and Killingsworth, M. A. (2021), Experienced well-being rises with income, even above $75,000 per year, PNAS 118, e2016976118. Note: the journal's own record; we obtained the abstract and the opening, not the full text. pnas.org
  2. Killingsworth, Kahneman and Mellers (2023), published abstract via PubMed, on two authors of the paper having published contradictory answers to whether larger incomes make people happier; on Kahneman and Deaton having used dichotomous questions about the preceding day and reported a flattening pattern in which happiness increased steadily with log(income) up to a threshold and then plateaued; on Killingsworth having used experience sampling with a continuous scale and reported a linear-log pattern in which average happiness rose consistently with log(income); on the authors engaging in an adversarial collaboration to search for a coherent interpretation of both studies; on a reanalysis of the experience sampling data confirming the flattening pattern only for the least happy people, with happiness increasing steadily with log(income) among happier people and even accelerating in the happiest group; on complementary nonlinearities contributing to the overall linear-log relationship; on the authors explaining why Kahneman and Deaton overstated the flattening pattern and why Killingsworth failed to find it; on the suggestion that Kahneman and Deaton might have reached the correct conclusion had they described their results in terms of unhappiness rather than happiness, their measures being unable to discriminate among degrees of happiness because of a ceiling effect; and on the mislabeling of the dependent variable and the incorrect assumption of homogeneity being consequences of practices that are standard in social science but should be questioned more often, with the authors flagging the benefits of adversarial collaboration. Note: the published abstract in full; we did not obtain the full text. pubmed.ncbi.nlm.nih.gov
  3. Princeton University, Kahneman-Treisman Center for Behavioral Science and Public Policy, announcement on the 2023 paper, on Kahneman and Deaton having published a landmark 2010 paper showing a rise in income increased well-being only to a ceiling of $75,000, drawing on more than 450,000 responses; on Killingsworth a decade later showing well-being rises with income even above $75,000; on the new adversarial collaboration with Barbara Mellers; on the more accurate reading being that an increase of income staved off unhappiness for a while but not after the flattening point; on the team examining the least happy 20 percent of Killingsworth's respondents and finding the flattening where Kahneman and Deaton would have anticipated it, adjusted for inflation; and quoting the authors that this income threshold may represent the point beyond which the miseries that remain are not alleviated by high income, with heartbreak, bereavement and clinical depression given as possible examples. Note: an institutional announcement from one author's own university; not a neutral source. behavioralpolicy.princeton.edu
  4. Author manuscript of Killingsworth, Kahneman and Mellers (2023), reference list, identifying Mellers, B., Hertwig, R., & Kahneman, D. (2001), Do frequency representations eliminate conjunction effects? An exercise in adversarial collaboration, Psychological Science 12, 269–275; Gilovich, T., Medvec, V. H., & Kahneman, D. (1998), Varieties of regret: A debate and partial resolution, Psychological Review 105, 602–605; Kahneman, D., & Klein, G., Conditions for intuitive expertise: A failure to disagree, American Psychologist 64; Cowan, N., and colleagues (2020), How do scientific views change? Notes from an extended adversarial collaboration, Perspectives on Psychological Science 15, 1011–1025; and Ellemers, N., Fiske, S. T., Abele, A. E., Koch, A., & Yzerbyt, V. (2020), Adversarial alignment enables competing models to engage in cooperative theory building toward cumulative science, PNAS 117, 7561–7567. Note: a reference list only. We obtained none of these works and report titles and citations with no findings. pnas.org
  5. Working paper, Income and emotional well-being: Evidence for well-being plateauing around $200,000 per year, on Kahneman and Deaton (2010) having analyzed a survey of more than 450,000 US citizens and found emotional well-being increasing linearly with household log-income up to a threshold around $75,000; on a similar conclusion in Jebb and colleagues (2018); on Killingsworth (2021) finding emotional well-being increases monotonically with log-income without reaching a plateau; on Killingsworth and colleagues (2023) being a notable instance of adversarial collaboration; and on that paper, using linear regression, validating the findings of Killingsworth (2021). Note: a working paper challenging the resolution. We obtained its abstract only and take no view on it; it is included because a paper titled "A conflict resolved" should not be presented as the end of a literature. arxiv.org
  6. Commercial website summary of the 2010, 2021 and 2023 papers, characterising money's effect on happiness as real but modest, with a four-fold income difference mattering less than a headache, about as much as being a caregiver, and twice as much as being married; giving the 2023 paper's sample as 33,391 working US adults with a median household income of $85,000; and stating that for around 80 percent of people happiness keeps rising past $100,000 while only the unhappiest 15 to 20 percent see a plateau near $100,000, being the original $75,000 adjusted for inflation. Note: a commercial website, not peer-reviewed. Every figure in this entry is unverified; we could not check any of them against the papers and this article does not rely on them. unfinishedman.com

This article describes research on income and reported wellbeing. It is not financial, medical or psychological advice, and nothing in it should be used to draw conclusions about any individual. Anyone struggling should speak to a qualified professional. No paper discussed was obtained in full and no effect sizes are reported. The magnitude characterisation and sample figure come from a commercial website and are unverified. One source is an institutional announcement from an author's own university. A subsequent paper challenging the resolution is noted. Nothing in this research was measured on business owners.