This article is about a finding that would matter enormously to any firm running a compliance function, and about a literature that has been unusually honest regarding its own weaknesses. The two facts are related.
Key Takeaway
A meta-analysis of 91 studies and 7,397 participants estimates the moral licensing effect at Cohen's d of 0.31, and reports that "published studies tend to have larger moral licensing effects than unpublished studies"[1]. One of its authors writes that "we actually think the effect we document is still an overestimation as the power of licensing studies was extremely low"[2]. A 2024 registered replication of the founding study found the effect "close to zero" in one scenario and "in the opposite direction" in another[3].
The Verdict, Stated First
Five claims, in descending order of confidence.
One. Two independent meta-analyses agree on the magnitude, and it is small. One reports d = 0.31 across 91 studies; another reports d = 0.319 across 106 effect sizes. Agreement between independent syntheses is genuine evidence and we do not dismiss it.
Two. The literature's own authors say the number is inflated, and our arithmetic says they are right to. At the average study size we can infer from the meta-analysis, a study had roughly 29 percent power to detect its own literature's effect. On our simulation, a significant result at that size overstates a true d of 0.31 by about 84 percent.
Three. The founding paradigm did not survive a registered replication. Close to zero in the ethnicity scenario, and reversed in the gender scenario at d between negative 0.50 and negative 0.38. A reversal is not a null.
Four. The best-known moderator did not survive either. A registered report with 5,091 participants tested whether conceptual abstraction determines consistency versus licensing and found the interaction was not replicated.
Five. For a compliance function, the practical answer is not to worry about licensing and not to relax either. The reason is the same one the thirty-eighth article in this series established for a different literature: an intervention that moves a measure has not been shown to move behaviour, and this is a case where the second link is exactly what is unestablished.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A that a small effect exists on average across the studied paradigms. Two independent meta-analyses converging at d of about 0.31 is strong evidence for a small effect, whatever the problems below.
Grade A for the publication bias finding, which the meta-analysis reports directly and which our own arithmetic independently predicts.
Grade A for the failure of the founding credential paradigm, from a registered report published in a peer-reviewed journal with effect sizes we obtained.
Grade A for the failure of the abstraction moderator, from a registered report with 5,091 participants.
Grade D for any specific claim about when licensing occurs. One meta-analysis found no evidence for theorised moderators; another found culture and comparison type mattered. Those cannot both be the last word.
Our position: this literature is more honest about itself than most this series has covered, and the honesty is what makes the picture look bad. A field that publishes its own failed replications and flags its own overestimates is behaving correctly.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2015 meta-analysis's abstract verbatim from the publisher, plus portions of its reference list[1]. We did not obtain the paper, its funnel plots, its moderator analyses or its list of included studies.
We obtained an author's own plain-language commentary on that meta-analysis from his personal academic page[2], which is not peer-reviewed and is flagged as such at every use.
We obtained the 2024 registered replication's abstract and results passages verbatim from three sources, including the journal itself[3][4].
We obtained the 2024 abstraction replication's abstract verbatim[5] and not its results.
We obtained the second meta-analysis's abstract verbatim[6] and not the paper.
We did not obtain the 2001 founding study, the 2014 failed replications, or any primary licensing experiment, and report all from abstracts, citation records and citing descriptions.
All arithmetic and simulation is ours. The per-study sample size used in it is our own inference from two reported totals and is not a figure any source states.
This article discusses research on ethical behaviour. It is not compliance, legal, audit or human resources advice.
What Is Being Claimed
The proposition, in the meta-analysis's own words.
"Moral licensing refers to the effect that when people initially behave in a moral way, they are later more likely to display behaviors that are immoral, unethical, or otherwise problematic."[1]
A second meta-analysis frames it in terms of self-image: "Moral licensing is a cognitive bias, which enables individuals to behave immorally without threatening their self-image of being a moral person."[6]
And a review describes the mechanism: "Past good deeds can liberate individuals to engage in behaviors that are immoral, unethical, or otherwise problematic, behaviors that they would otherwise avoid for fear of feeling or appearing" immoral, with our source truncating[7].
Three observations, ours.
The claim is causal and sequential, which is what makes it testable: a first act changes the probability of a second. That is a cleaner structure than most constructs in this series.
It is also the opposite of the intuitive prediction, which is that people behave consistently. That contrast is the reason the finding travelled, and it is the reason the literature contains a genuine debate rather than a consensus.
And note the mechanism named in the third quotation: fear of feeling or appearing immoral. The licence works on the self-image, not on the incentive, which is why the effect is expected to survive even when nothing material changes.
The Founding Study
The experiment the field is built on, described by the team that later replicated it.
The replication describes it: "Monin and Miller (2001, Study 2) found that participants who initially had an opportunity to hire a job candidate from disadvantaged groups (vs. those without such an opportunity) subsequently indicated preferences that were more likely to be perceived as prejudiced."[4]
The citation is Monin, B., and Miller, D. T. (2001), Moral credentials and the expression of prejudice, Journal of Personality and Social Psychology, 81, 33–43[8].
The replication also offers a real-world illustration of the mechanism drawn from the literature: "turning down a White couple who did not comply with dress code regulations provided the front desk with a moral credential that later helped them be confident that their decision would not be considered prejudiced (but fair) as they decided to turn down a Black couple."[4]
We did not obtain the 2001 paper and report it entirely through the replication's description.
Three observations, ours.
The design is elegant and the logic is clear. An opportunity to demonstrate non-prejudice, then a later choice that could be read as prejudiced. If the first act licenses the second, the manipulation should move the second.
The term moral credentials is worth keeping distinct from moral licensing generally. A credential is evidence about what kind of person you are; it is a narrower mechanism than a general good-deed balance.
And note that the outcome is an expressed preference rather than a behaviour with consequences. This series has repeatedly found that gap to be where effects disappear.
Where It Has Been Claimed
The spread of the literature, from the replication's own introduction.
It records that moral licensing "has received empirical support from both experiments and field studies and across a wide variety of contexts, such as hiring, environmental conservation, charitable giving, and volunteering."[4]
It also records two extensions. First, anticipation: "when people anticipate doing something morally dubious, they seem to strategically establish moral credentials in advance by demonstrating, if not exaggerating, their good morals." Second, transfer: "there is also evidence that people can be morally licensed not only by their own good behaviors but also by those of their ingroup members, a phenomenon called vicarious moral licensing."[4]
We obtained none of the studies referenced in those passages.
Three observations, ours.
The anticipatory version is the commercially alarming one, because it describes establishing a credential in order to spend it. That is a strategic account rather than a cognitive one.
The vicarious version is more alarming still for an organisation, because it implies a firm's ethical reputation could license individuals who contributed nothing to it.
And we grade neither, because we obtained no primary evidence for either and both are reported through a single citing paragraph.
Ninety-One Studies
The synthesis.
Blanken, I., van de Ven, N., and Zeelenberg, M. (2015), A Meta-Analytic Review of Moral Licensing, Personality and Social Psychology Bulletin, 41(4), 540–558[1].
The abstract: "We provide a state-of-the-art overview of moral licensing by conducting a meta-analysis of 91 studies (7,397 participants) that compare a licensing condition with a control condition. Based on this analysis, the magnitude of the moral licensing effect is estimated to be a Cohen's d of 0.31. We tested potential moderators and found that published studies tend to have larger moral licensing effects than unpublished studies. We found no empirical evidence for other moderators that were theorized to be of importance. The effect size estimate implies that studies require many more participants to draw solid conclusions about moral licensing and its possible moderators."[1]
Four observations, ours.
d = 0.31 places this near the middle of the distribution the fiftieth article in this series assembled from its own corpus, where the median reported effect was 0.41.
The comparison structure matters. These are studies that compare a licensing condition with a control condition, which is the right design and means the estimate is of a manipulation's effect rather than a correlation.
"We found no empirical evidence for other moderators that were theorized to be of importance" is a substantial negative finding stated in one clause. A field with many theorised moderators and no evidence for them is a field that does not know when its effect occurs.
And the last sentence is a methodological instruction to the field from its own synthesis: studies require many more participants. We take that instruction literally below.
Published Studies Show Bigger Effects
The finding inside the finding.
The only moderator that reached significance was publication status: "published studies tend to have larger moral licensing effects than unpublished studies."[1]
Three observations, ours.
This is the signature of publication bias, and it is being reported by the authors about their own literature. The reason the meta-analysis could detect it at all is that it included unpublished work, including the authors' own unpublished data, which appears in its reference list[1].
The direction is unambiguous and the mechanism is not mysterious. Studies finding larger effects are more likely to be published, so a synthesis restricted to published work would report a larger number.
And it means the 0.31 estimate is already a corrected figure, produced by a synthesis that went looking for the unpublished studies. A reader should not treat it as a naive average of the published literature, and should notice that the authors nonetheless think it is too high.
The Authors' Verdict On Their Own Number
The most striking statement in this article, and we flag its source carefully.
On his personal academic page, one of the meta-analysis's authors writes of it: "we conducted a meta-analysis on all the studies on moral licensing we could get our hands on. the analysis suggests a small moral licensing effect, the idea that after doing a good deed people more easily allow themselves to do something immoral. however, note that we actually think the effect we document is still an overestimation as the power of licensing studies was extremely low."[2]
And summarising the programme: "All in all, our meta-analysis and replication suggests a modest effect, but also that it might be less reliable/robust than often thought."[2]
This is a personal academic web page, not a peer-reviewed publication, and it is the author's own plain-language gloss on his published work. We cite it because the statement is unusually candid and because it is the author speaking about his own paper. A reader should weigh it accordingly.
Three observations, ours.
Researchers publicly stating that their own headline estimate is probably still too high is rare, creditable, and worth more than most peer-reviewed hedging.
The reason given is specific and checkable: the power of licensing studies was extremely low. That is not a vague caveat; it is a claim with arithmetic consequences, and we test it below.
And it converts the meta-analysis's own summary sentence from a suggestion into a diagnosis. Studies require many more participants is what you write when you have concluded that the existing ones were too small to trust.
And They Failed To Replicate It Themselves
The other half of the same research programme.
The same team published Blanken, I., van de Ven, N., Zeelenberg, M., and Meijers, M. H. C. (2014), Three attempts to replicate the moral licensing effect, Social Psychology, 45(3), 232–238[1][9].
A description records: "The present work includes three attempts to replicate the moral licensing effect by Sachdeva, Iliev, and Medin (2009). The original authors found that writing about positive traits led to lower" subsequent prosocial behaviour, with our source truncating[7].
We did not obtain the 2014 paper and report its title, citation and this partial description.
Three observations, ours.
The same four researchers published three failed replications in 2014 and a meta-analysis finding a small positive effect in 2015. That is not a contradiction; it is what an honest research programme looks like when individual studies are underpowered and a synthesis is not.
Their own failed replications are included in the meta-analysis, appearing in its reference list. They did not exclude their own inconvenient results.
And this is the correct relationship between a replication and a meta-analysis, which this series has not often been able to report. A single failed replication of a small effect is close to uninformative on its own. Three of them, plus a synthesis, plus a statement about power, is a position.
A Second Meta-Analysis Agrees
Independent corroboration of the magnitude.
A separate meta-analysis published in Management Review Quarterly reports: "Based on a random effects model, the point estimate for the generalized effect size Cohen's d is 0.319 (SE = 0.046; N = 106)."[6]
Its framing is commercial: it investigates the phenomenon "in a cross-cultural marketing context" and addresses "(i) how big moral licensing effects typically are and (ii) which factors systematically influence the size of this effect."[6]
Three observations, ours.
0.319 against 0.31, from a different team with a different inclusion set, is close agreement. This is the strongest evidence in the article that a real effect exists at some small size, and we weight it accordingly.
The reported standard error of 0.046 gives an approximate interval of about 0.23 to 0.41 by our own calculation, which is a reasonably tight bound on a small effect.
And the agreement is about magnitude only. The two syntheses disagree sharply about what governs the effect, which is the subject of the next section.
And Disagrees About Moderators
The disagreement that matters most for anyone trying to apply this.
The 2015 meta-analysis: "We found no empirical evidence for other moderators that were theorized to be of importance."[1]
The second meta-analysis: "Results of a meta-regression advance theory, by showing for the first time that both cultural background and type of comparison explain a substantial amount of the total variation of the effect size of moral licensing."[6]
Four observations, ours.
One synthesis found no moderators; the other found two explaining substantial variation. Both cannot be the last word.
They are not strictly contradictory, because the moderators tested may differ, and we did not obtain either moderator analysis so we cannot check the overlap. That is a real gap in this article.
But the 2015 paper's closing sentence anticipates exactly this problem: studies require many more participants to draw solid conclusions about moral licensing and its possible moderators. Moderator analyses need far more power than main effects, and a literature that cannot establish its main effect reliably has no business ranking its moderators.
And we grade any specific claim about when licensing occurs D on this basis. The two best syntheses available disagree, and one of them says the data cannot support the question.
What d Equals 0.31 Requires
Taking the meta-analysis's instruction literally. Our own arithmetic, standard two-group power calculation, five percent significance, two-tailed.
To detect an effect of d = 0.31:
At 50 percent power: 80 per group, 160 total.
At 80 percent power, the conventional standard: 164 per group, 328 total.
At 90 percent: 219 per group, 438 total.
At 95 percent: 271 per group, 542 total.
Three observations.
328 participants for a single two-condition comparison. That is the requirement to have a four-in-five chance of finding the effect if it is there at the meta-analytic size.
Small effects are expensive in a way that is not intuitive. The requirement scales with the inverse square of the effect size, so halving the effect quadruples the sample.
And a moderator analysis, which asks whether the effect differs between two subgroups, requires substantially more than this again. That is the arithmetic behind the meta-analysis's closing sentence.
What These Studies Had
The comparison, and this is our own inference rather than a reported figure. We flag that clearly.
The meta-analysis covers 91 studies and 7,397 participants[1]. Dividing gives an average of about 81 participants per study, or roughly 41 per condition in a two-condition design.
This is a crude average and no source states it. Study sizes will vary considerably around it, some designs have more than two conditions, and the figure should be read as an order of magnitude rather than a measurement.
Taking it at face value, the power of a study with 41 per group to detect d = 0.31 is 28.9 percent.
At 20 per group: 16.5 percent. At 30: 22.5 percent. At 50: 34.1 percent. At 80: 50.0 percent.
Three observations.
The average study in this literature had under a one-in-three chance of detecting its own literature's effect. Roughly seven studies in ten would miss it even though it is there.
This is what the author's phrase "extremely low" refers to, and it is not an exaggeration.
And the consequence is not merely that studies fail to find things. It is what happens to the ones that do find something, which is the next section and the most important arithmetic in this article.
The Winner's Curse
What low power does to the studies that succeed. Our own simulation, 300,000 simulated studies at each size, with the true effect fixed at the meta-analytic estimate of 0.31.
When power is low, a study only reaches significance if random noise happens to push the observed effect upward. So the studies that get published are selected on having been lucky, and their reported effects are inflated.
At 20 per group: significant studies report an average d of 0.773, an inflation of 149 percent.
At 30 per group: 0.650, inflation 110 percent.
At 41 per group, our inferred literature average: 0.570, inflation 84 percent.
At 80 per group: 0.436, inflation 41 percent.
At 164 per group, the properly powered size: 0.348, inflation 12 percent.
Four observations.
At the inferred average study size, a significant published result overstates the true effect by about 84 percent. Nothing dishonest has occurred anywhere in that process.
This is a purely statistical consequence of selecting on significance, and it happens to any literature with small studies and small effects, regardless of the topic or the researchers.
Note the last row. Even a properly powered study inflates by 12 percent when you condition on significance, which is why individual significant findings should never be read as unbiased estimates of magnitude.
And the inputs are ours. The true effect of 0.31 is the meta-analytic estimate; the per-study sample is our own inference; the simulation is our construction. It demonstrates a mechanism and does not measure this literature.
Which Explains The Publication Bias Finding
Putting the two together, because they are the same phenomenon seen from two directions. Ours.
Three points.
The meta-analysis reports, as an empirical result, that published studies show larger effects than unpublished ones. Our simulation predicts exactly that, from nothing but sample size and a selection rule.
So the empirical finding and the arithmetic corroborate each other. The observed pattern is what you would expect if the literature consisted of small studies of a small effect, filtered by significance.
And it substantiates the author's statement that the meta-analytic estimate is still an overestimate. A synthesis that includes some unpublished work corrects part of the selection, but only the part it managed to find, and unfound unpublished studies are by construction the ones with the smallest effects.
The Registered Replication Of The Founding Study
The most consequential recent result, and it is a registered report.
Xiao, Q., Li, L. C., Au, Y. L., Tan, S. N., Chung, W. T., and Feldman, G. (2024), Licensing via Credentials: Replication Registered Report of Monin and Miller (2001) with Extensions Investigating the Domain-Specificity of Moral Credentials and the Association Between the Credential Effect and Trait Reputational Concern, International Review of Social Psychology, DOI 10.5334/irsp.945[3].
The design was pre-specified in a published study design table, whose research question reads: "Do previous moral behaviors that give one moral credentials make people more likely to engage in morally questionable behaviors later?" with the stated hypothesis "Moral credentials make people more likely to engage in subsequent morally questionable acts."[3]
The paper records that "This Registered Report has been endorsed by Peer Community In Registered Reports" and that "All materials, data, and analysis scripts are shared" at a public repository[3].
Three observations, ours.
A registered report is accepted for publication on the basis of its design, before results exist. That removes the selection mechanism our simulation just modelled, which is precisely why this study's result carries weight the original paradigm's supporting studies cannot.
The team is the same one whose status quo bias replication the fifty-fifth article in this series reported, and it includes Gilad Feldman, whose group also produced the outcome bias replication in the forty-fourth.
And the extensions were pre-specified too, testing domain-specificity of credentials and an association with trait reputational concern, both of which are theoretically motivated additions rather than exploratory analyses.
The Reversal
The result, and it is worse for the paradigm than a null.
"We found no support for a consistent moral credential effect: the effect was close to zero in a scenario where participants indicated their preferences to hire from different ethnicities (d = 0.02 to 0.08, depending on inclusion criteria), and was in the opposite direction in a scenario where they indicated preferences for different genders (d = −0.50 to −0.38)."[4]
And on the extensions: "With two extensions to the original study design, we found no evidence that domain-inconsistent moral credentials are less effective in licensing than domain-consistent moral credentials and that moral credentials moderate the association between reputational concern and expressing potentially prejudiced preferences."[4]
Four observations, ours.
d between 0.02 and 0.08 in the ethnicity scenario is not a small effect; it is indistinguishable from nothing, in the paradigm that founded the field.
The gender scenario is the striking one. d between negative 0.50 and negative 0.38 is a moderate effect running the other way, which describes moral consistency: having demonstrated non-prejudice, participants were subsequently less likely to express a prejudiced preference, not more.
A reversal of that size is larger in absolute terms than the meta-analytic licensing effect itself, which is worth stating plainly.
And the ranges are given "depending on inclusion criteria", which is the authors reporting their own analytic sensitivity rather than choosing the most favourable specification. That is the correct practice and it is why the interval is a range.
Five Thousand And Ninety-One
The other registered report, testing the field's best-known moderator.
Martuza, J., and Kim, O., Does Conceptual Abstraction Moderate Whether Past Moral Deeds Motivate Consistency or Compensatory Behavior? A Registered Replication and Extension of Conway and Peetz (2012), Personality and Social Psychology Bulletin, DOI 10.1177/01461672241238420[5].
Its abstract: "A long-standing debate in psychology concerns whether doing something good or bad leads to more of the same or the opposite. Conway and Peetz proposed that conceptual abstraction moderates if past moral deeds lead to consistent or compensatory behavior. Although cited 384 times across disciplines, we did not find any direct replications... A large-scale experiment (N = 5,091) in the registered report format tested Conway and Peetz's original hypothesis. The hypothesized interaction was not replicated: conceptual abstraction did not moderate the effect of recalling moral vs. immoral behavior on prosocial intentions."[5]
Four observations, ours.
N = 5,091. Against our own power calculation this is enormous, roughly fifteen times what is needed for the main effect and enough to detect a moderator with real confidence.
The moderator tested is the field's most cited answer to the central question. A review had summarised it as: individuals show consistency when they focus abstractly on the connection between their behaviour and their values, and licensing when they think concretely about what they accomplished[7].
"Cited 384 times across disciplines, we did not find any direct replications." That sentence describes the condition this series has documented repeatedly: influence accumulating without verification.
And the result is stated without qualification. The hypothesised interaction was not replicated.
What Actually Survives
Our reading, stated directly, and it is more balanced than the last two sections alone would suggest.
Five statements.
A small average effect across the studied paradigms is well supported. Two independent meta-analyses at 0.31 and 0.319 is real evidence and we do not discount it.
The estimate is probably still high, on the authors' own statement and on our own winner's curse arithmetic, which independently predicts the publication bias they observed.
The founding credential paradigm did not survive a registered replication, coming out near zero in one scenario and reversed in another.
Nobody can currently say when it happens. One synthesis found no moderators, another found two, and the most cited moderator failed in a 5,091-participant registered report.
And a small average effect with unknown moderators is not actionable. An average across paradigms tells a firm nothing about its own situation if the conditions governing the effect are unestablished.
Your Compliance Programme
The application this article exists for. Ours, untested, and not compliance or legal advice.
Four points.
The worry is legitimate in principle: completing an ethics module is a moral act, and the literature's claim is that moral acts license subsequent lapses. If true at any meaningful size, this would mean compliance training carries a cost as well as a benefit.
The evidence does not support acting on that worry. The effect is small on average, its estimate is described by its own authors as inflated, the founding paradigm reversed under registered replication, and no one can say when it occurs. Redesigning a compliance function around it would be building on the least secure part of the literature.
But the evidence does not support complacency either, and this is the part we would emphasise. The reason to doubt licensing is not that ethics training works. Those are separate questions, and this article says nothing about the second.
And the honest position for a firm is the uncomfortable one. You do not know whether your ethics training changes behaviour, in either direction, and the licensing literature is not the reason you do not know.
And Your CSR Reporting
The organisational version, which the second meta-analysis frames in marketing terms. Ours.
Three points.
The vicarious claim, that people can be licensed by the good behaviour of their ingroup, is the one with organisational implications, since a firm's published sustainability record is a group-level moral credential.
We grade that claim ungraded, because we obtained no primary evidence for it and it reaches this article through a single citing sentence.
And we would note the asymmetry that makes the topic attractive to write about and hard to use. The story is memorable, the mechanism is intuitive, and the evidence for the specific organisational version does not exist in anything we obtained. That combination is where this series has most often found overstatement.
The Same Broken Chain
The structural point, which connects this to two earlier articles. Ours.
Three observations.
The thirty-eighth article reported a network meta-analysis of 492 studies in which interventions changed a target measure while the change did not carry through to behaviour. The forty-sixth reported a replication in which a manipulation made a task feel harder while the downstream judgment effect did not appear.
Moral licensing research has the same two-link structure. The first link is that a moral act changes self-perception. The second is that the changed self-perception changes subsequent behaviour. Most of the outcomes in this literature, including the founding study's, are expressed preferences rather than consequential acts.
And the general principle is the one worth carrying out of all three articles. An intervention working through an intermediate step inherits the product of both stages, so a chain with a strong first link and an unestablished second link is an unestablished chain.
What To Do
Do not redesign a compliance function around moral licensing. The effect is small on average, its own authors call the estimate inflated, and the founding paradigm reversed under registered replication.
Do not treat that as evidence your ethics training works. Those are separate questions and this literature addresses only the first.
Ask what any programme actually measures. If the outcome is an expressed preference or a survey response rather than a consequential act, the second link is untested.
Treat a small effect with unknown moderators as unusable. Two syntheses disagree about what governs this effect and the most cited moderator failed in a 5,091-participant registered report.
Learn the winner's curse and apply it everywhere. On our own simulation, at a realistic study size a significant published result overstates a true effect of 0.31 by about 84 percent, with nobody behaving dishonestly.
Weight registered reports above ordinary publications. Acceptance on design rather than results removes the selection mechanism that produces the inflation above.
Notice when a citation count outruns its verification. A finding cited 384 times with no direct replications is a pattern, not an accident.
Give credit where a field polices itself. The team that produced the main meta-analysis also published three failed replications, included their own unpublished data, and stated publicly that their headline number is probably too high.
The Limits Of This Analysis
Several caveats matter. This article discusses research on ethical behaviour and is not compliance, legal, audit or human resources advice; the organisational sections are our own reasoning and are untested. Everything is verified to August 2026. We did not obtain the 2015 meta-analysis, only its abstract and part of its reference list, so we report no funnel plot, no heterogeneity statistic, no confidence interval on its headline estimate, and no list of included studies. We did not obtain the second meta-analysis beyond its abstract, and could not check the overlap between the two moderator analyses, which is a real gap given that they disagree. We did not obtain the 2001 founding study and report it entirely through its replicators' description. We did not obtain the 2014 failed replications, only their citation and a partial description. We did not obtain the abstraction replication's results, only its abstract. We obtained no primary licensing experiment of any kind. One key source is a personal academic web page, cited for an author's own gloss on his own published work, flagged at every use, and not peer-reviewed. All arithmetic and simulation is ours. The power calculations use a standard two-group formula at five percent two-tailed significance. The figure of roughly 41 participants per condition is our own inference from dividing two reported totals, is not stated by any source, ignores variation between studies and designs with more than two conditions, and should be read as an order of magnitude. Our winner's curse simulation takes the meta-analytic estimate as the true effect, which is an assumption rather than a fact, and demonstrates a statistical mechanism rather than measuring this literature. And we note that two independent meta-analyses agreeing at approximately 0.31 is genuine evidence for a small effect; nothing in this article should be read as a claim that no effect exists.
Frequently Asked Questions
What is moral licensing?
Is the effect real?
Why does low power inflate published effects?
What happened in the 2024 replication?
When does licensing happen and when does consistency?
Should I change my ethics training because of this?
Does this literature come out of it badly?
References
- Blanken, I., van de Ven, N., & Zeelenberg, M. (2015). A Meta-Analytic Review of Moral Licensing. Personality and Social Psychology Bulletin, 41(4), 540–558, DOI 10.1177/0146167215572134. Publisher record reproducing the abstract in full: on moral licensing referring to the effect that when people initially behave in a moral way they are later more likely to display behaviors that are immoral, unethical, or otherwise problematic; on the authors providing a state-of-the-art overview by conducting a meta-analysis of 91 studies with 7,397 participants comparing a licensing condition with a control condition; on the magnitude of the effect being estimated at a Cohen's d of 0.31; on the authors having tested potential moderators and found that published studies tend to have larger moral licensing effects than unpublished studies; on the authors having found no empirical evidence for other moderators theorized to be of importance; and on the effect size estimate implying that studies require many more participants to draw solid conclusions about moral licensing and its possible moderators. The record also reproduces part of the reference list, which includes the authors' own unpublished raw data from 2012 and their own 2014 replication attempts. Note: the publisher's record. We obtained the abstract and part of the reference list only; we did not obtain the paper, its funnel plots, its heterogeneity statistics, any confidence interval on the headline estimate, or its list of included studies. journals.sagepub.com
- Personal academic web page of one of the meta-analysis's authors, describing his own published work: on the authors having conducted a meta-analysis on all the studies on moral licensing they could get their hands on; on the analysis suggesting a small moral licensing effect, being the idea that after doing a good deed people more easily allow themselves to do something immoral; on the authors nonetheless thinking the effect they document is still an overestimation, as the power of licensing studies was extremely low; and on the overall conclusion that their meta-analysis and replication suggest a modest effect but also that it might be less reliable or robust than often thought; together with the citations for the 2015 meta-analysis and the 2014 replication attempts. Note: a personal academic web page, not a peer-reviewed publication. Cited because it is an author's own plain-language gloss on his own published work; flagged as such at every use in the body. A reader should weigh it accordingly. nielsvandeven.nl
- Xiao, Q., Li, L. C., Au, Y. L., Tan, S. N., Chung, W. T., & Feldman, G. (2024). Licensing via Credentials: Replication Registered Report of Monin and Miller (2001) with Extensions Investigating the Domain-Specificity of Moral Credentials and the Association Between the Credential Effect and Trait Reputational Concern. International Review of Social Psychology, DOI 10.5334/irsp.945, published 20 May 2024, authors at the Faculty of Psychology, University of Vienna, and the Department of Psychology, University of Hong Kong. Full-text repository copy reproducing the abstract, keywords and study design table: on the moral credential effect being the phenomenon where an initial behavior that presumably establishes one as moral licenses the person to subsequently engage in morally questionable behaviors; on the pre-specified research question of whether previous moral behaviors that give one moral credentials make people more likely to engage in morally questionable behaviors later, with the stated hypothesis that moral credentials make people more likely to engage in subsequent morally questionable acts; on the two extensions finding no evidence that domain-inconsistent moral credentials are less effective in licensing than domain-consistent ones, nor that moral credentials moderate the association between reputational concern and expressing potentially prejudiced preferences; on all materials, data and analysis scripts being shared publicly; and on the Registered Report having been endorsed by Peer Community In Registered Reports. Note: a full-text repository copy of a registered report. We obtained the abstract, keywords and design table. ncbi.nlm.nih.gov
- Journal record for the same 2024 registered report, reproducing its results and introductory passages: on the authors having found no support for a consistent moral credential effect, with the effect close to zero in a scenario where participants indicated their preferences to hire from different ethnicities, at d of 0.02 to 0.08 depending on inclusion criteria, and in the opposite direction in a scenario where they indicated preferences for different genders, at d of negative 0.50 to negative 0.38; on moral licensing having received empirical support from both experiments and field studies across contexts such as hiring, environmental conservation, charitable giving and volunteering; on people who anticipate doing something morally dubious appearing to strategically establish moral credentials in advance by demonstrating, if not exaggerating, their good morals; on evidence that people can be morally licensed by the good behaviors of their ingroup members, termed vicarious moral licensing; on an illustration in which turning down a White couple who did not comply with dress code regulations provided a front desk with a moral credential that later supported confidence that turning down a Black couple would be considered fair rather than prejudiced; and on Monin and Miller (2001, Study 2) having found that participants who initially had an opportunity to hire a job candidate from disadvantaged groups subsequently indicated preferences more likely to be perceived as prejudiced. Note: the journal's own record. Our source for the replication's effect sizes and for the description of the 2001 founding study, which we did not obtain. The extension studies referenced in the introduction were likewise not obtained. rips-irsp.com
- Martuza, J., & Kim, O. Does Conceptual Abstraction Moderate Whether Past Moral Deeds Motivate Consistency or Compensatory Behavior? A Registered Replication and Extension of Conway and Peetz (2012). Personality and Social Psychology Bulletin, DOI 10.1177/01461672241238420, corresponding author at the Department of Strategy and Management, Norwegian School of Economics. Publisher record reproducing the abstract: on a long-standing debate in psychology concerning whether doing something good or bad leads to more of the same or the opposite; on Conway and Peetz having proposed that conceptual abstraction moderates whether past moral deeds lead to consistent or compensatory behavior; on that proposal having been cited 384 times across disciplines while the present authors found no direct replications; on it having been unclear how increases or decreases from one's baseline prosociality might underlie the effect; on a large-scale experiment with 5,091 participants in the registered report format having tested the original hypothesis; and on the hypothesized interaction not having been replicated, with conceptual abstraction not moderating the effect of recalling moral versus immoral behavior on prosocial intentions. Note: the publisher's record. We obtained the abstract and not the results. journals.sagepub.com
- Publisher record for a culture-moderated meta-analysis of moral licensing published in Management Review Quarterly, reproducing its abstract: on moral licensing being a cognitive bias which enables individuals to behave immorally without threatening their self-image of being a moral person; on the authors investigating the phenomenon in a cross-cultural marketing context and addressing how big moral licensing effects typically are and which factors systematically influence the size of the effect; on the authors approaching these questions by conducting a meta-analysis and a meta-regression; on the point estimate for the generalized effect size Cohen's d being 0.319, with a standard error of 0.046 and N of 106, based on a random effects model; and on results of the meta-regression showing for the first time that both cultural background and type of comparison explain a substantial amount of the total variation of the effect size. Note: the publisher's record. We obtained the abstract only, and could not compare this paper's moderator analysis against the 2015 meta-analysis's, which reported no evidence for theorized moderators. link.springer.com
- Bibliographic service record for the 2015 meta-analysis carrying summaries of related works: on past good deeds being able to liberate individuals to engage in behaviors that are immoral, unethical, or otherwise problematic, behaviors they would otherwise avoid for fear of feeling or appearing immoral, our source truncating; on a review of the literature on moderators of moral consistency versus licensing effects revealing that individuals are more likely to exhibit consistency when they focus abstractly on the connection between their initial behavior and their values, whereas they are more likely to exhibit licensing when they think concretely about what they have accomplished; on a demonstration that threats to moral identity can increase how definitively people think they have previously proven their morality; and on the 2014 work including three attempts to replicate the moral licensing effect reported by Sachdeva, Iliev and Medin in 2009, whose original authors found that writing about positive traits led to lower subsequent behaviour, our source truncating. Note: a bibliographic service record reproducing summaries of related papers. We obtained none of the works summarised. semanticscholar.org
- Publisher record for Conway, P., and Peetz, J. (2012), When does feeling moral actually make you a better person? Conceptual abstraction moderates whether past moral deeds motivate consistency or compensatory behavior, Personality and Social Psychology Bulletin, 38, 907–919, DOI 10.1177/0146167212442394; carrying a reference list confirming Monin, B., and Miller, D. T. (2001), Moral credentials and the expression of prejudice, Journal of Personality and Social Psychology, 81, 33–43. Note: a publisher record used to confirm the citation details of the 2001 founding study and the 2012 moderator paper. We obtained neither paper. journals.sagepub.com
- Reference list carried on a publisher record for a 2024 registered replication, independently confirming Blanken, I., van de Ven, N., Zeelenberg, M., & Meijers, M. H. (2014), Three attempts to replicate the moral licensing effect, Social Psychology, 45, 232–238, DOI 10.1027/1864-9335/a000189; together with Blanken, I., van de Ven, N., & Zeelenberg, M. (2015), Personality and Social Psychology Bulletin, 41, 540–558; and Brown, R. P., and colleagues (2011), Moral credentialing and the rationalization of misconduct, Ethics and Behavior, 21, 1–12. Note: a reference list; citations only, used to confirm the 2014 paper's journal, volume, pages and DOI independently. We did not obtain the 2014 paper. journals.sagepub.com
This article discusses research on ethical behaviour and is not compliance, legal, audit or human resources advice. Neither meta-analysis was obtained beyond its abstract, so no funnel plot, heterogeneity statistic or confidence interval is reported. The 2001 founding study, the 2014 failed replications, and every primary licensing experiment were not obtained. One source is a personal academic web page, cited for an author's own gloss on his own published work and flagged at every use. All arithmetic and simulation is the authors' own; the figure of roughly 41 participants per condition is an inference from dividing two reported totals and is stated by no source. Two independent meta-analyses agreeing at approximately 0.31 is genuine evidence for a small effect, and nothing here should be read as a claim that no effect exists.