Every article in this series so far has had to reconstruct what happened to a finding from the outside. This one does not. The first author wrote it down and published it herself.
Key Takeaway
In a document on her own faculty page, the paper's first author writes: "As such, I do not believe that 'power pose' effects are real." She then lists what was done, including that subjects were run in chunks with the effect checked along the way, and that "The self-report DV was p-hacked."[1] Our own simulation of the checking procedure she describes puts the false positive rate at 10.5 percent rather than the nominal 5.
The Verdict, Stated First
Five claims, in descending order of confidence.
One. The first author has publicly repudiated the finding, in a document we obtained in full from her own institutional page, and the language is not hedged.
Two. She discloses specific procedures, including checking the data as it came in and selecting among p-values, and she names one of them p-hacking herself.
Three. A larger replication found no effect on hormones or risk taking, and reported that its power to detect the original effect exceeded 95 percent.
Four. On our own simulation, the procedure she describes roughly doubles the false positive rate, from the nominal 5 percent to about 10.5.
Five. The self-report effect did replicate, which is the part most accounts of this episode omit and which we treat as the honest residue.
That fifth claim is the one we would most defend against our own side of the argument, ours. A series that only ever found things wanting would be as unreliable as one that never did, and here something survived.
Our Grades For These Claims
Applying the scheme from the first article in this series. This is the best-sourced article we have published.
Grade A-plus for the author's own account, obtained in full as a document she hosts herself, which is a primary source of a kind we have not had before in this series.
Grade A for the replication's design and stated power, obtained from the paper's own methods text.
Grade B for the subsequent dispute, which reaches us through a professional body's review article rather than through the papers themselves.
Grade A for our own simulation, which reproduces the nominal 5 percent exactly when run as a single test, and which is our construction from her description.
One calculation withdrawn, reported in its own section, because we could not compute it honestly.
The grade spread here is unusually narrow, ours, and it changes what this article can do. We are not weighing thin sources against each other; we are reading a first-hand account and doing arithmetic on it, which is the position this series has been trying to reach for ninety-four articles.
A Note On Method
Everything here is verified to August 2026.
We obtained the author's position document in full from her university faculty page[1], and every quotation attributed to her comes from it.
We obtained the replication paper's methods text, including its statement of statistical power[2], though not the paper in full.
We did not obtain the original 2010 paper's full text, only its abstract opening and its bibliographic record.
The later dispute reaches us through a professional psychological body's review article[3], which is an institutional publication rather than a peer-reviewed one and is flagged.
All simulation and arithmetic is ours. The chunk sizes and the p-values are hers; every figure derived from them is our own construction.
This article discusses research methods and is not advice about interview or presentation preparation.
The Claim As Used
What the finding became in business, before we look at what it was.
A summary of the original describes participants who briefly adopted expansive postures as having sought more risk, had higher testosterone levels, and had lower cortisol levels than those adopting contractive ones, in a sample of 42[6].
Four observations, ours.
In business it became a preparation ritual. Two minutes in a confident posture before an interview, a pitch or a difficult meeting.
The appeal is that it is free, private and immediate, which is an unusual combination and explains the uptake better than the evidence does.
And the mechanism claimed was hormonal, which is what gave it authority. Advice to stand up straight is folk wisdom; advice that measurably changes testosterone is science.
That distinction matters for what follows, because the hormonal claim and the feeling claim have had completely different fates, and only one of them failed.
Which we flag now because most retellings of this episode collapse them, ours. The story is usually told as a finding that vanished, and that is not what the record shows.
The Document
The source, and it is unusual enough to describe carefully.
It is titled "My position on 'Power Poses'" and is hosted as a PDF on the author's faculty page at her university[1]. It concerns Carney, Cuddy and Yap (2010), of which she is first author.
She writes: "Reasonable people, whom I respect, may disagree. However since early 2015 the evidence has been mounting suggesting there is unlikely any embodied effect of nonverbal expansiveness", and then: "As evidence has come in over these past 2+ years, my views have updated to reflect the evidence. As such, I do not believe that 'power pose' effects are real."[1]
And later: "The evidence against the existence of power poses is undeniable."[1]
Four observations, ours.
This is a primary source of a kind this series has not had before. We are not reconstructing what happened from an abstract; we are reading the account of the person who did it.
The document is self-published rather than peer reviewed, which cuts both ways: nobody vetted it, and nobody could have stopped it either.
Its evidential status is unusual and worth stating precisely. Statements against interest are generally credible, and there is no version of this document that benefits its author.
And she states her position on the finding in five numbered points, including "I discourage others from studying power poses" and "I do not teach power poses in my classes anymore."[1]
Two notes on how it became known, ours. It was reported by a national broadcaster at the time[5], so this was not a quiet posting nobody noticed.
And yet the reporting did not displace the advice. The statement is nearly a decade old and the posture is still recommended, which is the gap between correction and circulation that this series keeps measuring.
The List Of Facts
The section of the document headed "Here are some facts."
Among the items she lists[1]: "The data are real." "The sample size is tiny." "The data are flimsy. The effects are small and barely there in many cases."
And: "Some subjects were excluded on bases such as 'didn't follow directions.' The total number of exclusions was 5. The final sample size was N = 42."[1]
Four observations, ours.
"The data are real" is the first item, and it is doing important work. This is a disclosure about analysis, not about fabrication, and the distinction should not be blurred.
The exclusion detail is the sort that never reaches a paper's reader. Five exclusions from 47 is more than one in ten, and the stated basis is a judgement call.
We would note that excluding participants who did not follow directions is a normal and defensible practice, and the problem is not the exclusions but that the decision was available to be made after seeing the data.
And the ordering of the list is itself notable. She puts the sample size and the flimsiness before the specific procedures, which is the order a critic would use.
One more item belongs here and we have held it back until now, ours. She writes that the self-report measure was p-hacked, in that many different power questions were asked and those chosen were the ones that worked[1].
That matters more than it first appears, because the self-report effect is the one that later replicated, and its measurement in the original was selected after the fact.
Ran In Chunks
The item that made this article possible, quoted in full.
She writes: "Initially, the primary DV of interest was risk-taking. We ran subjects in chunks and checked the effect along the way. It was something like 25 subjects run, then 10, then 7, then 5. Back then this did not seem like p-hacking. It seemed like saving money (assuming your effect size was big enough and p-value was the only issue)."[1]
Four observations, ours.
This describes optional stopping: collecting data, testing, and continuing only if the test has not yet come out. It is one of the best understood problems in statistics and one of the least intuitive.
The chunk sizes add to 47, and with five exclusions that gives the reported 42, so the arithmetic in her account is internally consistent.
We checked that ourselves rather than taking it on trust, ours, and it is a small point in the document's favour. An account invented after the fact does not usually reconcile to the published sample size.
Her explanation is the one that makes this common rather than exotic. It seemed like saving money, which is exactly what it looks like from inside, and is why the practice was widespread.
And the parenthesis is the whole error in eleven words. Assuming your effect size was big enough and p-value was the only issue is precisely the assumption that fails, because the checking changes what the p-value means.
Two things make this the most useful disclosure in the document, ours. It is specific enough to simulate, which almost no admission of this kind ever is.
And it describes a practice that is still standard in business, where running a trial and watching the numbers weekly is considered diligence rather than a methodological problem.
What That Does, Simulated
Our own simulation. The chunk sizes are hers; the model and every figure below are ours.
We simulated two groups with no real difference between them at all, tested at 25 subjects, again at 35, again at 42 and again at 47, stopping the moment a test came out below the conventional threshold. Two hundred thousand runs.
Testing once at the end, at 47 subjects: a false positive rate of 5.0 percent, which is what it should be and which confirms our code is behaving.
Testing four times as described: a false positive rate of 10.5 percent, a multiplier of 2.11.
Four observations.
The nominal five percent is a promise about a single test, and it is the only thing the threshold guarantees. Four looks at the same accumulating data is four chances, and the promise does not survive it.
The doubling is not enormous and that is the point. Optional stopping does not usually produce absurd results; it produces a rate about twice what was advertised, which is small enough to be invisible in any single study.
We would emphasise what the simulation does and does not show. It shows that the procedure inflates false positives under the null, and it says nothing about whether this particular result was one.
And it is worth noting that our figure is a floor rather than a ceiling, because it assumes only this one procedure. Her account describes several.
Two limits on the simulation we want stated plainly, ours. Our model compares two groups of normally distributed numbers, which is not what the original analysed, and the exact figure would move under a different model.
What would not move is the direction and the rough magnitude. Repeated testing on accumulating data inflates the false positive rate, and with four looks the inflation is on the order of a doubling, which is a standard result and not something we discovered.
The Two P-Values
The second procedure, in her words.
She writes: "For the risk-taking DV: One p-value for a Pearson chi square was .052 and for the Likelihood ratio it was .05. The smaller of the two was reported despite the Pearson being the more ubiquitously used test of significance for a chi square. This is clearly using a 'researcher degree of freedom.'"[1]
She adds: "I had found evidence that it is more appropriate to use 'Likelihood' when one has smaller samples and this was how I convinced myself it was OK."[1]
Four observations, ours.
The gap between the two values is 0.002, and what turns on it is the entire published claim.
The justification offered was not fabricated, on her own account. She had a real reason to prefer the likelihood ratio at small samples, and that is what makes this the difficult case rather than the easy one.
The problem is the sequence. A reason discovered after seeing which test gave the better answer is not the same reason as one adopted before, even when the reason itself is sound.
And the phrase "this was how I convinced myself it was OK" is the most useful sentence in the document, because it describes an internal experience anybody who has analysed data will recognise.
We would put the general form of it plainly, ours. The mind supplies a justification for whichever choice produced the better answer, and it supplies it sincerely, which is why the practice is not experienced as dishonest by the person doing it.
A Calculation We Withdrew
Reported here because this series reports the ones that fail.
We intended to simulate the second procedure as we had the first: how much does picking the smaller of two test statistics raise the false positive rate?
Four observations on why we could not.
The two tests are a Pearson chi square and a likelihood ratio computed on the same table, so they are not independent, and their correlation depends on the table's dimensions and cell counts.
We do not have the table. Without it, any number we produced would depend on a contingency table we had invented, and the answer would be a function of our invention rather than of her disclosure.
We could have picked a plausible table and reported a figure. It would have looked like a result and been an assumption, which is the failure mode this series exists to point out.
So what we can say is the part that needs no simulation. Selecting the smaller of two p-values cannot lower a false positive rate, and the author calls it a researcher degree of freedom herself.
And the withdrawal has a use of its own, ours. It marks the boundary between what her disclosure supports and what it does not, and a reader can now see exactly where our arithmetic stops.
What A Sample Of 42 Can Detect
Our own simulation, on the sample size she calls tiny.
Forty-two participants split between two conditions gives about 21 per cell. We simulated what such a study finds, at the conventional threshold, for a range of true effect sizes.
A true effect of 0.2 standard deviations is detected 10 percent of the time. 0.4: 25 percent. 0.5: 35 percent. 0.8: 71 percent. 0.9: 81 percent.
Four observations.
To be found four times in five at this sample size, the true effect must be around 0.9 standard deviations, which is a very large effect in this field.
So a study of this size is only equipped to find large effects, and a real but moderate effect would be missed three times in four.
The consequence people find least intuitive is the reverse one. If a small study does report a significant effect, that effect is necessarily estimated as large, because a smaller one would not have cleared the threshold, which means small studies systematically overstate whatever they find.
And that is a general point about your own pilot programmes rather than about this paper, which we return to below.
One more comparison worth having, ours. At about 100 per cell, which is roughly what the replication used, an effect of 0.4 standard deviations is detected 81 percent of the time, against 25 percent at the original's size.
That is the whole difference between the two studies stated as arithmetic. One could see moderate effects and the other could not, and no amount of care in the smaller study would have changed that.
The Replication
The study that changed her mind, and she says so.
Ranehill, Dreber, Johannesson, Leiberg, Sul and Weber (2015), Assessing the Robustness of Power Posing: No Effect on Hormones and Risk Tolerance in a Large Sample of Men and Women, Psychological Science, 26(5), 653–656[3].
Its own methods text describes a conceptual replication using "a substantially larger sample (N = 200) and a design in which the experimenter was blind to condition."[2]
A review article records the sample as 98 female and 102 male participants, against the original's less even mix, and a longer posture duration of six minutes[3]. A professional body's publication, flagged.
Four observations, ours.
The design improvement that matters most is experimenter blinding, and the original document lists the absence of it as a confound in her own words[1].
The sample is roughly five times larger, which moves the study from being able to detect only very large effects to being able to detect moderate ones.
Calling it a conceptual replication is honest and is also the opening for the dispute that followed, since the departures from the original are real.
And the author of the original was a reviewer on it. She writes that she signed her reviews and was "strongly in favor in the Ranehill case."[1]
Powered To Find It
The sentence that settles the most common objection.
The replication's methods state: "Our statistical power to detect an effect of the magnitude reported by Carney et al. was more than 95%."[2]
Four observations, ours.
The standard objection to a null result is that the study was too small to find the effect. This sentence forecloses it in advance, and the authors evidently anticipated it.
Ninety-five percent power means that if the original effect were real and of the reported size, this study would have found it nineteen times out of twenty.
Our own figures agree with that being achievable. At about 100 per cell, an effect of 0.4 standard deviations is detected 81 percent of the time and 0.5 is detected 94 percent, and the original's reported effects were larger than these.
What that does not establish is that no effect exists. A study powered to find a large effect can miss a small one, and the finding is properly read as ruling out an effect of the advertised size.
Which is the correct and unsatisfying form of most replication results, ours. They bound the effect rather than abolishing it, and a bound is what the evidence can support.
What Did Survive
The part most accounts omit, and we think it is the honest residue.
A review article records that the replication successfully replicated self-reported feelings of power while failing to produce significant results for the behavioural and hormonal measures[3]. A professional body's publication, flagged.
Four observations, ours.
So the finding split cleanly. People who adopt expansive postures report feeling more powerful; their hormones and their risk taking did not move.
The surviving effect is the least surprising of the three, and it is also the one most vulnerable to participants guessing the hypothesis, which is a limitation rather than a dismissal.
It is nonetheless not nothing. Feeling more confident before a difficult meeting is a real thing to want, and if a posture reliably produces it then the practice has a defensible basis.
What it does not have is the hormonal mechanism, which is what made the advice sound like science rather than like standing up straight.
Two consequences of losing the mechanism, ours. The advice loses its claim to work regardless of belief, since a hormonal effect would operate whether or not you expected it and a feeling need not.
And it loses its transferability. A hormonal effect would be expected to hold across people and settings; a feeling produced partly by expectation would not.
The Confounds She Lists
A separate section of her document, and one of the items is genuinely clever.
She writes that "The experimenters were both aware of the hypothesis"[1].
And on the testosterone result: participants were told immediately whether they had "won" the risk task, with a small extra prize, and she notes that research shows winning increases testosterone. "Thus, effects observed on testosterone as a function of expansive posture may have been due to the fact that more expansive postured subjects took the 'risk' and you can only 'win' if you take the risk."[1]
Four observations, ours.
That confound is a chain: posture leads to risk taking leads to winning leads to testosterone, which would produce the observed correlation with no direct hormonal effect of posture at all.
It is the kind of thing that is invisible until stated and obvious afterwards, and she says as much, describing the confounds as clear only in hindsight.
She also notes that gender was not handled appropriately in the testosterone analyses, which matters because baseline testosterone differs enormously between men and women.
And we would observe that this section is doing something rarer than the repudiation. She is explaining how the result could have arisen honestly, which is more useful than declaring it wrong.
The winning confound also has a direct business analogue worth naming, ours. Any measurement taken after an intervention has produced an outcome can be reflecting the outcome rather than the intervention.
A firm measuring morale after a sales team hits its target has the same chain. The programme may have raised morale, or the target may have, and the design cannot separate them.
The Paper She Regrets
An episode with a lesson about responding to criticism.
She writes of the 2015 review and summary paper her team published in response to the replication: "What I regret about writing that 'summary' paper is that it suggested people do more work on the topic which I now think is a waste of time and resources."[1]
And with unusual candour about how it was used: "Ultimately, this summary paper served its intended purpose because it offered a reasonable set of studies for a p-curve analysis which demonstrated no effect."[1]
Four observations, ours.
The summary reviewed 33 studies and proposed moderators for future work[6], which is the standard response to a failed replication and is usually reasonable.
The twist is that assembling the supporting literature in one place made it possible to analyse the literature as a whole, and the analysis went against it.
There is a general principle in that, ours. A defence that assembles all the favourable evidence is also an inventory of the favourable evidence, and inventories can be audited.
And her regret is specific rather than general. She does not regret defending the work; she regrets directing other people's time toward it, which is a distinction worth preserving.
Two things that distinction protects, ours. Defending your own work under challenge is what a researcher is supposed to do, and a norm against it would be worse than the problem it solved.
And the cost she identifies falls on other people, which is the right thing to weigh, since the researchers who pursued the moderators spent years on it.
The Dispute That Continued
Because this did not end with the document, and a reader should know that.
A p-curve analysis of the literature was published as Simmons and Simonsohn (2017), Power Posing: P-Curving the Evidence, Psychological Science, 28(5), 687–693[3].
A reply followed: Cuddy, Schultz and Fosse (2018), P-curving a more comprehensive body of research on postural feedback, Psychological Science, 29, 656–666[4], which on its title reports clear evidential value for postural feedback effects.
Four observations, ours.
We obtained neither paper, only their bibliographic records and titles, so we describe the dispute rather than adjudicating it.
The disagreement is substantially about which studies belong in the analysis, on the titles alone, one being described as more comprehensive than the other.
That is a real methodological question and not a rhetorical one. Which studies count is the central decision in any evidence synthesis, and reasonable people do disagree about it.
And a co-author continuing to defend a line of work is ordinary scientific conduct, not misconduct, which we state plainly because this episode is often narrated otherwise.
One further reason for that care, ours. The two authors are not making the same claim: one has repudiated the hormonal effects specifically, and the continuing defence concerns postural feedback more broadly.
Those can both be right. A narrow claim can fail while a broader one survives, and treating the dispute as a personal contest obscures that it is partly a disagreement about scope.
What Actually Survives
Our reading, stated directly.
Five statements.
The first author does not believe the effects are real and has said so in a document she publishes herself.
The data collection involved checking as it went, which on our own simulation roughly doubles the false positive rate.
A larger, experimenter-blind replication found no hormonal or behavioural effect, and reported power above 95 percent to detect the original.
The self-report effect did replicate, so people who adopt the posture do report feeling more powerful.
And the dispute over the wider literature continued after the repudiation, which we did not adjudicate.
Those five are what we would defend, ours, and the first two rest on a document we obtained in full from the author who wrote it, which is the strongest evidential position this series has occupied.
Credit Where It Is Due
Because the tone of most retellings is wrong, and ours will not be. Ours.
Four observations.
Publishing this document was voluntary and costly. No process compelled it, and it damages the author's own most cited work.
She also reviewed the failed replication in favour of publication and signed her review[1], which is the opposite of obstruction.
And the practices she describes were normal at the time, on her own account, which is the important context. She writes that back then the chunk checking did not seem like p-hacking.
So the correct reading is not that somebody behaved badly. It is that a field's ordinary practice was capable of producing a result nobody could reproduce, and one participant in it wrote down exactly how.
And that is why we graded this article's sourcing above every previous one, ours. Every other article here reconstructs a failure from outside; this one has the account from inside, which is available only because somebody chose to write it.
Not An Indictment Of Research
The obvious misreading, addressed. Ours.
Four observations.
This is a story about research working, on a longer timescale than anybody would like. A finding was published, challenged, replicated without success, and repudiated by its own author within seven years.
The self-correction is the part that could not happen in business, where a strategy adopted on a weak finding is rarely revisited at all, and where nobody publishes a document explaining what they did to the numbers.
What the episode does indict is the speed of transmission relative to the speed of correction. The advice reached millions; the correction reached a faculty webpage.
And that asymmetry is the ninety-third article's finding restated. A memorable claim travels and a correction does not, and the correction here is a PDF with a filename containing a space.
Your Own Pilot Programmes
The application to a firm's own decisions, and it is direct. Ours, and not statistical advice.
Four points.
Most business pilots are smaller than 42. A new process tried on one team, a script tested on twenty calls, a layout tried in one location.
At those sizes, on our own figures, you can only detect very large effects, and a real but moderate improvement will usually look like nothing.
The reverse risk is worse and less obvious. Anything a small pilot does show will be overstated, because only large apparent effects clear any threshold at that sample size, and the firm then budgets for the overstated figure.
And the checking behaviour is nearly universal in business. Running a trial and looking at the numbers every week is exactly the procedure the author describes, done for exactly her reason, which was to avoid spending more than necessary.
Two reasons the business version is worse, ours. There is no threshold at all in most firms, so the stopping decision is made on impression rather than on a test, which removes even the nominal guarantee.
And the person checking usually wants the pilot to succeed, having proposed it, which is the blinding problem the author lists as her first confound.
Set The Stopping Rule First
The one practice that follows directly. Ours.
Four observations.
Decide the sample size and the decision rule before the trial starts, and write them down where you cannot revise them quietly.
That single discipline removes the problem our simulation measures, because the inflation comes entirely from the option to stop, and an option you have foreclosed cannot be exercised.
It costs something real and we will not pretend otherwise. You will sometimes run a trial for longer than you needed to, which is the saving the author was reaching for.
And where you genuinely must look early, the answer is to plan the looks in advance and set a stricter threshold for each, which is standard practice in clinical trials and is available to anyone willing to look it up.
One honest caveat on all of this, ours. Most business decisions do not warrant this machinery, and a firm that formalises every trial will formalise itself into paralysis.
The discipline is worth its cost where the decision is expensive and hard to reverse, and is not worth it for the rest, which is a judgement only the owner can make.
Bibliographic Note
The series keeps a count, and this article produced two, one of them remarkable.
The original paper's page range is given as 1363–1368 by the journal's own record[6] and as 1363–1369 in a reference work's list[4].
And in the reference list of her own repudiation document, the author cites the replication as "Psychological Science, 33, 1-4"[1], where the correct record is 26(5), 653–656[3].
Three observations, ours.
The second is the most striking variant we have recorded. The author miscites the replication of her own study, in the document repudiating her own study, and the volume number is wrong by seven.
We record it without any suggestion that it matters to her argument, because it plainly does not, and it is evidence chiefly of how easily citations degrade even at the source.
If anything it strengthens the document's credibility, ours. A carefully lawyered statement would have had its references checked, and this one reads as written quickly by somebody who had decided to say the thing.
The first is the ordinary kind we have seen many times, and it comes from a reference work rather than from a journal, which is where most of them originate.
That brings the running count of bibliographic variants across this series to forty-three.
What To Do
Do not present the hormonal claim as established. The first author does not believe the effects are real, and a blind replication with power above 95 percent found nothing on hormones or risk taking.
You may keep the posture if it helps you. The self-report effect replicated, and feeling more confident before a difficult meeting is a legitimate thing to want.
Set your sample size and stopping rule before any pilot begins. On our own simulation, checking four times as data accumulates roughly doubles the false positive rate.
Assume a small pilot cannot find a moderate effect. At about 21 per group, an effect of 0.4 standard deviations is found a quarter of the time.
And discount whatever a small pilot does find. Only large apparent effects clear a threshold at that size, so the estimate you get is systematically too high.
Blind whoever runs the trial where you can. The absence of blinding is the first confound the author lists in her own account.
Notice when you are choosing among analyses. The difference between the two tests here was 0.002 and it decided everything.
And write down your reason for a methodological choice before you see what it does, which is the only reliable defence against convincing yourself it was fine.
The Limits Of This Analysis
Several caveats matter. This article discusses research methods and is not advice about interview or presentation preparation, and not statistical advice; the applications to business pilots are our own reasoning. Everything is verified to August 2026. We did not obtain the original 2010 paper's full text, only its bibliographic record and a summary of its findings, so every characterisation of what it did comes either from its first author's later account or from secondary descriptions. The author's position document is self-published and not peer reviewed; we treat it as credible chiefly because it is a statement against interest, and a reader who weighs it differently would reach different conclusions. We obtained the replication's methods text and not the full paper, so we cannot describe its analyses. The detail that the self-report effect replicated, the participant gender split, the six-minute posture duration, and the bibliographic records for the later dispute all reach us from a professional body's review article, flagged, which is an institutional publication rather than a peer-reviewed one. We obtained neither paper in the p-curve dispute, only their titles and records, and we do not adjudicate it. All simulation is ours. The chunk sizes, the two p-values and the sample size are the author's; the false positive rates, the multiplier and the detection probabilities are our own constructions, and our simulation assumes a simple two-group comparison of normally distributed data, which is not what the original analysed. One calculation was withdrawn, on the two competing tests, because it would have depended on a contingency table we do not have and would have had to invent. And our figure of 10.5 percent is a floor rather than an estimate of what actually occurred, since it models one of several disclosed procedures in isolation.
Frequently Asked Questions
Did the first author really repudiate her own study?
What is optional stopping and why does it matter?
Was the replication big enough to find the effect?
So is any of it real?
What can a sample of 42 actually detect?
Why does this matter for a business pilot?
Does this mean the researchers behaved badly?
References
- Carney, D. R., My position on "Power Poses", PDF hosted on the author's faculty page at the Haas School of Business, University of California, Berkeley, regarding Carney, Cuddy & Yap (2010). Obtained in full. States that reasonable people whom she respects may disagree, but that since early 2015 the evidence has been mounting suggesting there is unlikely any embodied effect of nonverbal expansiveness on internal or psychological outcomes; that as evidence has come in over the past two years her views have updated to reflect the evidence, and as such she does not believe that power pose effects are real; that her lab is conducting no research on the topic; that the evidence against the existence of power poses is undeniable; that she continues to review failed replications and re-analyses, signing her reviews as in the Ranehill case, almost always in favour of publication, and was strongly in favour in that case. Under a heading of facts, states that the data are real, that the sample size is tiny, that the data are flimsy with effects small and barely there in many cases, that subjects were run in chunks with the effect checked along the way at something like 25 subjects then 10 then 7 then 5, that back then this did not seem like p-hacking but like saving money, that some subjects were excluded on bases such as not following directions with five exclusions total and a final sample of 42, that for the risk-taking measure a Pearson chi square gave .052 and a likelihood ratio gave .05 with the smaller reported despite Pearson being more ubiquitously used, that this is clearly a researcher degree of freedom and that she had found evidence favouring likelihood at smaller samples and this was how she convinced herself it was acceptable, and that the self-report measure was p-hacked in that many power questions were asked and those chosen were the ones that worked. Under a heading on confounds, states that both experimenters were aware of the hypothesis; that participants were told immediately whether they had won an extra prize on the risk task, that research shows winning increases testosterone, and that the testosterone effect may therefore have been a winning effect rather than a posture effect; and that gender was not dealt with appropriately in the testosterone analyses. States her position in five points including that she does not have faith in the effects, does not study them, discourages others from studying them, and no longer teaches them. Also states regret that the 2015 summary paper suggested others do more work on the topic, which she now considers a waste of time and resources, while noting it served its purpose by assembling a set of studies for a p-curve analysis that demonstrated no effect. Note: a self-published document by the paper's first author, NOT peer reviewed, and the source of every statement attributed to her in this article. We treat it as credible principally because it is a statement against interest. faculty.haas.berkeley.edu
- Methods text of Ranehill, E., Dreber, A., Johannesson, M., Leiberg, S., Sul, S., & Weber, R. A. (2015), Assessing the Robustness of Power Posing: No Effect on Hormones and Risk Tolerance in a Large Sample of Men and Women, Psychological Science, obtained as a PDF hosted on the first-named author of the original study's own faculty page. Records that the original examined whether posing affected levels of hormones such as testosterone and cortisol, financial risk taking and self-reported feelings of power in a sample of 42 participants randomly assigned to hold poses suggesting either high or low power; that the replication used a similar methodology but a substantially larger sample of 200 and a design in which the experimenter was blind to condition; and that statistical power to detect an effect of the magnitude reported by the original authors was more than 95 percent. Note: the replication's own methods text and our source for the sample size, the blinding and the stated power. We obtained the methods section and not the paper in full. faculty.haas.berkeley.edu
- Review article published by a national professional body for psychologists, surveying a decade of power posing research. Records that the replication was a conceptual replication which successfully replicated self-reported feelings of power while failing to produce significant results for the behavioural and hormonal measures; that it featured 200 participants with a more even gender mix than the original at 98 female and 102 male; that the original authors responded the following month in the same journal listing methodological departures including a longer posture duration of six minutes; and giving bibliographic records for Ranehill et al. (2015), Psychological Science, 26(5), 653–656, and Simmons, J. P. & Simonsohn, U. (2017), Power posing: P-curving the evidence, Psychological Science, 28(5), 687–693. Note: a professional body's publication, NOT a peer-reviewed article, flagged at every use. Our source for the surviving self-report effect, the sample composition and the bibliographic records of the later dispute. bps.org.uk
- Reference work entry on power posing in an encyclopedia of personality and individual differences, defining high and low power poses and carrying a reference list including Carney, D. R., Cuddy, A. J. C., & Yap, A. J. (2010), Psychological Science, 21, 1363–1369; Carney, Cuddy & Yap (2015), Review and summary of research on the embodied effects of expansive (vs. contractive) nonverbal displays, Psychological Science, 26, 657–663; and Cuddy, A. J., Schultz, S. J., & Fosse, N. E. (2018), P-curving a more comprehensive body of research on postural feedback reveals clear evidential value for power-posing effects: Reply to Simmons and Simonsohn (2017), Psychological Science, 29, 656–666. Note: a reference work entry, recorded principally for the bibliographic records of the 2015 summary and the 2018 reply, and as the source of one bibliographic variant in the original paper's page range. link.springer.com
- National broadcaster's news report on the first author's statement, reporting that she says she now has no faith in the embodied effects of power poses, and recording that the idea came from a 2010 study co-authored by three researchers then at Columbia and Harvard. Note: a news report, NOT an academic source, flagged. Used only to establish that the statement was publicly reported at the time and not merely posted quietly. npr.org
- Research aggregator record carrying the abstract of Simmons and Simonsohn, Power Posing: P-Curving the Evidence, which states that in a well-known article Carney, Cuddy and Yap (2010) documented the benefits of power posing; that in their study participants, N equals 42, randomly assigned to briefly adopt expansive powerful postures sought more risk, had higher testosterone levels and had lower cortisol levels than those assigned contractive postures; and that in their response to a failed replication by Ranehill et al. (2015), the original authors reviewed 33 successful studies investigating the effects of expansive versus contractive posing to identify possible moderators. Note: an aggregator record, used for the description of the original study's reported findings and for the count of studies in the 2015 summary. We did not obtain the p-curve paper itself. core.ac.uk
This article discusses research methods and is not advice about interview or presentation preparation, and not statistical advice. The original 2010 paper was not obtained in full. The first author's position document is self-published and not peer reviewed. All simulation and arithmetic is the authors' own, built from figures disclosed in that document, and one calculation was attempted and withdrawn.