A blog on statistics, methods, philosophy of science, and open science. Understanding 20% of statistics will improve 80% of your inferences.

Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Monday, May 25, 2026

Evaluating Dr. Cuddy’s Claim that the Debunking of Power Posing is a Myth

In this blog post I will analyse the arguments that Dr. Amy Cuddy provided in a LinkedIn post “The "Power Posing Was Debunked" Myth: What the Research Actually Shows — and Why Scientific Discourse Matters” on February 26. You can find the LinkedIn post here:

https://www.linkedin.com/pulse/power-posing-debunked-myth-what-research-actually-shows-amy-cuddy-t6lnc

In the post, Cuddy says she was “effectively silenced” by an “attempt to shut down this line” of research. She credits “the courage of the individual scientists who kept going despite enormous pressure not to” for the fact that she can still summarize “what the evidence now shows”.

Power posing has two categories of claimed effects. The first effect is on self-reported feelings. For example, if we instruct people to stand in a constricted versus an expanded posture, they will self-report feeling more powerful. There is an ongoing debate about whether, or how much, this effect is caused by a demand effect (i.e., people report what they think the investigator wants them to say, not what they actually feel). A meta-analysis has shown this self-report effect is larger in within-subject designs, and in studies without a cover story (Körner et al., 2022). The second effect is on physiological or behavioral outcomes. This is the contested area, and the research outcome that Cuddy is mainly trying to defend in her blog post. If you want to explore a meta-analysis on these two categories of effects, you can do so at https://metaanalyses.shinyapps.io/bodypositions/ (made by Körner et al., 2022). I would especially recommend exploring the QRP/Publication bias tab for the physiological and behavioral outcomes.

At the end of the post, Cuddy writes that she is thankful that not everyone stopped doing research on power poses, because then: “We would not know what we now know — which is that these effects are real, that they matter, and that the story people were told was wrong.”  She concludes with: “The evidence is there. It has been there for years. All I am asking is that people look at it.”

I am happy to do so. Let’s go.

Trying to find the references

I tried to look up the references cited by Cuddy in her post. However, this reference:

Andolfi, V. R., & Antonietti, A. (2020). Contractive vs. expansive body posture effects on convergent-integrative thinking tasks. Journal of Creative Behavior, 54(4), 871–880.

does not exist in literature databases, and the authors (who do exist) do not list this paper on their own websites. An inspection of the journal’s website shows that a different article was published in volume 54, issue 4 on these pages. This raises questions about how this reference was generated, with generation by AI being a plausible candidate (also in view of the 4 malformed references I will point out below). The reference appears in the following sentence in Cuddy’s LinkedIn post:

Andolfi and Antonietti (2020, Journal of Creative Behavior) provided further evidence that contractive postures specifically benefited convergent-integrative thinking tasks. That level of specificity — where the direction of the effect depends on the type of cognitive task — is exactly the kind of finding that emerges when a field matures.

When Cuddy says ‘The evidence is there’, this is not correct for the Andolfi and Antonietti article, which does not seems to exist in the scholarly record.

There are an additional 4 references that suggest that the literature review may have in part been generated by automated tools, but for these 4 references, there are papers that match the content discussed in the literature review in the LinkedIn post.

 

Reference in LinkedIn post

Actual Reference

Michinov, E., & Michinov, N. (2020). Creativity connected with body posture: The effects of expansive and contractive postures on creative performance. Psychology of Aesthetics, Creativity, and the Arts, 14(1), 116–127

Michinov, N., & Michinov, E. (2022). Do open or closed postures boost creative performance? The effects of postural feedback on divergent and convergent thinking. Psychology of Aesthetics, Creativity, and the Arts, 16(3), 504–518. https://doi.org/10.1037/aca0000306

 

Wainio-Theberge, S., Bhatt, M., Bhattacharyya, K., et al. (2025). Neural correlates of power-related postures and their behavioural consequences: A preliminary electrophysiological investigation. Social Cognitive and Affective Neuroscience, 20(1), nsaf03

 

Wainio-Theberge, S., & Armony, J. L. (2025). Neural correlates of power-related postures and their behavioural consequences: A preliminary electrophysiological investigation. Social Cognitive and Affective Neuroscience, 20(1), nsaf036. https://doi.org/10.1093/scan/nsaf036

 

Elkjær, E., Mikkelsen, M. B., Michalak, J., Mennin, D. S., & O'Toole, M. S. (2023). Using bodily displays to facilitate approach action outcomes within the context of a personally relevant task. Frontiers in Psychology, 14, 1147printing

Elkjær, E., Mikkelsen, M. B., Tramm, G., Michalak, J., Mennin, D. S., & O’Toole, M. S. (2022). Using bodily displays to facilitating approach action outcomes within the context of a personally relevant task. Brain and Behavior, 13(1), e2855. https://doi.org/10.1002/brb3.2855

 

Körner, R., Köhler, H., & Schütz, A. (2020). Powerful and confident children through expansive body postures? A preregistered test of the effects of power posing on children. School Psychology International, 41(4), 315–330.

 

Körner, R., Köhler, H., & Schütz, A. (2020). Powerful and confident children through expansive body postures? A preregistered study of fourth graders. School Psychology International, 41(4), 315–330. https://doi.org/10.1177/0143034320912306

 

 

We see all these references that are incorrect refer to the later literature, and summarize the research of the people who ‘kept going’. These references are at the core of the argument Cuddy is making.

Evaluating the evidence: Three examples

Cuddy wrote a narrative review, which requires that the validity of the conclusions, and the strength of the evidence, needs to be evaluated for every study. Let’s carefully examine some of the papers she cited and evaluate the evidence. Cuddy writes about a first study:

Wainio-Theberge and colleagues (2025, Social Cognitive and Affective Neuroscience) published the first EEG study of power posing, finding significant effects on arousal and valence, with suggestive differences in frontal brain activity between expansive and contractive postures. A new neural methodology for a question people said was already answered.

From the description in Cuddy’s blog, you might assume the “significant effects on arousal and valence, with suggestive differences in frontal brain activity between expansive and contractive postures” would support the hypothesis. But this is not the case. The significant effects were actually in the opposite direction of the hypothesis. This is not mentioned in the abstract of the Wainio-Theberge et al. article, and one would need to read the paper to get this information:

We found no significant posture differences in the EEG spectral exponent (t(101) = 1.01, P = .32). In contrast, a significant posture effect was observed for frontal asymmetry (t(101) = −2.63, P = .01); however, post hoc t-tests in each group separately (‘Models 1c and 1e’) revealed that the effect was in the opposite direction as hypothesized (see Discussion). Namely, we observed a significant right-lateralized frontal alpha asymmetry (FAA) in the contractive group (t(45) = 2.17, P = .04) and a left-lateralized FAA in the expansive one which failed to reach significance (t(55) = −1.63, P = .11).

Cuddy writes in the blog that she responded to journalists skeptical about power posing: “I spent more than ten hours responding — reviewing the literature, pulling citations, writing carefully, anticipating distortions” In this case, her review of the literature presented a finding as providing support for power posing, when in fact the effect was in the opposite direction of the hypothesis.

As a second paper, let’s take Barel and colleagues (2024). First, I want to thank the authors for sharing their data, after I tried to access it by clicking the google drive link in the article. All numbers were reproducible. Cuddy cites the paper as follows:

“As other researchers began testing that broader construct, using different measures in different populations, they found effects consistently: action orientation (Huang et al., 2011, Psychological Science), […], and risk-taking itself, partially (Barel et al., 2024, BMC Psychology).

It is unclear what is meant by 'partially', as the authors are clear that they found that power posing did not affect risk-taking: "There was no statistically significant distribution in risk-taking between high and low power conditions [χ2 = 0.00, p > 0.99]." The risk-taking outcome that Cuddy cites the study for is a clear null result.

The basis for "partially" is presumably a separate analysis reported in the paper: in a logistic regression predicting risk-taking, the authors found a significant interaction between power condition and cortisol change and write that they "did partially replicate an effect of changes in cortisol levels on risk-taking." But note that they claim an effect of cortisol changes on risk, not of power posing on risk. For power posing to affect risk through cortisol, power posing would first have to change cortisol, and it did not: the authors report no main effects of time or power on cortisol. With that first link missing, the high-power participants whose cortisol fell are not a subgroup of people for whom the power pose worked, as their cortisol would have moved the same way without any pose. The significant effect is a within-group association between two measures, which can't be attributed to the power posing manipulation.

To their credit, the authors themselves never claim power posing affected risk-taking. This framing comes from Cuddy, who presents the paper as a partial replication of a risk-taking effect after a power posing manipulation, which the study did not support.

When discussing a third paper, Cuddy writes: “Körner, Köhler, and Schütz (2020, School Psychology International) conducted a preregistered study of 108 German fourth graders — children — and found that expansive postures increased self-esteem, positive feelings, feelings of power, and even children's perceptions of their relationship with their teacher. The strongest effects were on school-related self-esteem. This is exactly the kind of applied, developmentally informed research that matters — taking findings from the lab and asking whether they help real children in real classrooms.”

The Körner et al study was preregistered: https://aspredicted.org/blind.php?x=sn4su9 with 4 t-tests to examine 4 dependent variables of interest. Of the 4 tests, 2 are significant (p = 0.04 and p = 0.013), but neither survive a correction for multiple comparisons (0.05/4 = 0.0125) which was necessary in this analysis.

The blog by Cuddy states “The strongest effects were on school-related self-esteem.” But the biggest effect is actually on the student-teacher relationship:

Finally, there was a significant difference between the two groups regarding the pictures related to the student–teacher relationship: high power posers more frequently chose the picture showing a good student–teacher relationship than low power posers, Χ²(1) = 11.181, p = .001, φ = –.322.

But there is a problem with this finding. Students spent months building a relationship with their teacher. Then, as part of the experiment, the students posed for 60 seconds and self-reported on that relationship, without any further interaction with the teacher.  There is no possible causal mechanism for the power pose to impact the relationship with teachers. Although unintended, this question is an excellent probe for demand effects. As the power pose can’t change history and impact the actual relationship between students and teachers, the observed effect can only be caused by a demand effect. Neither Cuddy nor the original authors realized this. Cuddy instead concludes: “This is exactly the kind of applied, developmentally informed research that matters — taking findings from the lab and asking whether they help real children in real classrooms.”

Evaluating the Research Line

Evaluating evidence is effortful and messy. Single studies always have weaknesses, and the reader might reasonably wonder whether I’m cherry-picking a few bad apples from an otherwise strong set. I don’t think I am, and I will explain the more general pattern I observed when reading all the cited papers.

Exploratory claims

The Körner et al (2020) study above was preregistered, and therefore we were able to evaluate that the claims were not severely tested, as they would not survive the required correction for multiple comparisons (Lakens, 2019). But most claims in the papers that Cuddy cites are based on exploratory analyses. The studies all have many dependent variables, and a large number of tests can be performed. These studies observe a mix of significant and non-significant results, but the significant results have a high probability of being Type 1 errors and can’t be presented as evidence. If researchers in this field would perform more direct replication studies, and would preregister their studies more, they could address this problem. Some preregistered their studies, which is excellent, but some don't, even though they work in a highly contested research area, and the significant results primarily come from exploratory analyses.

Researchers in the field are often honest about this, but especially in a narrative summary, it is easy to lose track of the fact that most of the authors of studies cited by Cuddy do not consider their own findings to be strong evidence. For example, Metzler et al (2023) write “Finally, it is important to transparently report on the level of evidence this study provides for power pose effects on low-level social behavior. This requires mentioning its exploratory nature [...] we are convinced that the medium effect sizes, given our sample size, would require replication before strong conclusions can be drawn”. I would say this is especially important given that the main result was a 3-way interaction with a p-value of 0.03: “the predicted three-fold interaction suggested that this effect of emotion on action choices (more avoidance for anger than fear) changed between sessions as a function of adopted pose (OR = 1.19, 95% CI[1.02, 1.38], z = 2.18, p = .029)”.

Another example comes from Elkjær et al (2022). The main finding is: “Concerning approach tendencies, the 2 × 3 interaction analysis on DAT “approach threat 1” was significant (F(1, 87) = 3.27, p = .043, ηp2 = .07). Regarding DAT avoid threat (1 + 2), the overall 2 × 3 interaction analysis was significant (F(1, 87) = 6.39, p = .003, ηp2 = .13).” The study was preregistered (https://aspredicted.org/blind.php?x=9j3b38) which allows us to see that the preregistered predictions are not supported. The authors predicted significant effects for the expansive condition compared to both the constricted condition and the control condition. However, they did not find effects compared to the control condition. Such patterns of mixed results are present in many studies in the literature. On the one hand, this is part of normal research, especially early on in research lines, when researchers have not figured out how to reliably produce the effect they are examining. On the other hand, power posing has been studied since 2010, and a research line can never get a strong basis if it does not move beyond a literature where all significant results are based on exploratory partial confirmations.

If you want to see the exploration of data in action, I would recommend looking at the OSF repository related to the paper by Michinov and Michinov (2024): www.osf.io/c9mzh, and see which variables and ways of computing variables are reported in the final paper, and which are not.

Underpowered studies and selection for significance

The sample sizes in the studies cited by Cuddy are often small – especially for key sub-group analyses, when the total sample size might be distributed across cells in a 2x3 design. This would not be problematic if the effects of power posing were known to be large. But even the self-report effect where participants indicate they feel more or less powerful has a rather small effect size of only g = 0.37 (see https://metaanalyses.shinyapps.io/bodypositions/). Less direct effects, for example on behavior, are likely to have a much smaller effects (unless researchers can propose strong theoretical arguments why more indirect effects would be larger, see Anvari et al., 2023). In one-tailed independent t-tests, 80% power would require 184 participants (92 per condition), but none of the studies are close to achieving such sample sizes.

The research area of power posing is also characterized by the selective reporting of significant results. This combination of underpowered studies and selection for significance leads to highly inflated effect sizes. We can see these effects in Andolfi et al (2017):

The effect sizes of an open or closed posture simply can’t be in the range of d = 1.22, or even d = 0.69 (for examples of realistic effect sizes to expect based on group differences, see DataColada 18). The effects are inflated, and there is no way of knowing what the true effect sizes are. They might be zero, as many replication studies of exactly such implausibly large effects based on studies with tiny samples have turned out to be.

The study by Michinov and Michinov similarly shows effects for significant tests that are too large. Adopting a posture for a few minutes can’t plausibly influence creative tasks with effects such as d = 0.634. When you evaluate evidence, thinking about selective reporting and inflated effects should be part of the evaluation.

 

Quality of the design and analysis

I could not help noticing that there is a lot of room to improve the quality of the study design and analysis, as reported in papers in this literature. This in itself does not mean that the evidence is unreliable, but it does not make it easier for a research field to generate high quality evidence. For example, Elkjær et al (2022) report the following power analysis:

“Based on a priori power calculations, using a repeated-measures ANOVA interaction analysis, 2 (time; before vs. after the manipulation) × 3 (condition; EXP, CON, N), 90 participants were required to detect a small effect size (d = 0.34), with an alpha of .05 and a beta of .20.”

At first sight, this looks like best practice. They acknowledge power posing effects are small (d = 0.34 is very much in line with the meta-analysis they published in the same year). Regrettably, what the authors actually did was enter an f = -.34, not a d, as you can see in the screenshot below, which leads to a sample size that is much lower than what they would actually have needed to achieve high power, according to their own meta-analytic effect size estimate:

This means that despite the power analysis, the study was still massively underpowered. The sample size justifications in all studies cited by Cuddy are problematic. This is probably true for many research lines, but it is especially problematic for a research line where researchers are still trying to establish if the basic effect exists or not.

While reading the articles, I also noticed many of the issues that we often see in other literatures when research teams lack statistical expertise. There are often small inconsistencies in the correct degrees of freedom, incorrectly performed statistical tests, an overreliance on p-values despite underpowered studies, and misinterpretations of non-significant results. I don’t want to single out more examples, but it would probably be good for the field if researchers would enlist some methodological and statistical expertise if they want to generate reliable evidence.  

Tools to evaluate claims

Cuddy writes: “When people are told that research is fake — without being given the tools to evaluate that claim — it doesn't just affect one researcher or one line of work. It feeds a broader cynicism: that science can't be trusted, that findings are arbitrary, that expertise is performance.” I strongly agree. This is why I have created a free textbook, Improving Your Statistical Inferences, to learn how to evaluate the actual evidence in scientific papers. Here are three decent heuristics to follow when you evaluate the evidence in a research line:

  1. If a finding shows what you want to be true, be extra skeptical.
  2. If you have a strong conflict of interest, be extra skeptical.
  3. Studies with low power due to too small sample sizes, lack of preregistration, no direct replications, strong indications of selective reporting, low methodological quality, repeating limitations in discussion sections without addressing them, implausibly large effect sizes, a lack of impact on other research areas, significant claims that mainly come from exploratory analyses, continued uncertainty about the basic effect after more than a decade and dozens of studies, and the research community disengaging with a literature are all signs of a lack of evidence.

According to Cuddy, she “live[s] inside a false narrative” where power posing is incorrectly believed to be a ‘myth’, and she believes that “none of this would have happened if the methods guys, and the journalists who trusted them without doing proper research, hadn't created the conditions that made it happen.”

 

Scientific criticism is a cornerstone of a healthy science

When I read Cuddy’s LinkedIn post, I was highly skeptical of the claim that there was evidence for effects of power posing on measures other than self-report, and that the debunking was a 'myth'. But my first response was to ignore the post. I did not want to examine the evidence behind the claims Cuddy made, because I am clearly one of the “method guys” who, according to Cuddy “manufactured the "debunked" narrative and aimed it, with great precision, at a single researcher”. If I would criticize her post, would I be seen as contributing to “the bullying I was subjected to”, as Cuddy writes?

But I care about criticism in science. And I think it is important that we can criticize scientific claims. My decision to not follow up on examining the claims in the blog post kept nagging me. Cuddy has 900,000 followers on LinkedIn who have read the very strong statement that it is a “myth” that power posing was debunked. If the evidence she presented was overstated – as I feared – scientific criticism would be needed to correct the record. I think it is essential to increase social safety in academia, while being able to criticize each other. I do not want bullying and scientific criticism to become conflated. Scientific criticism is too important for a healthy science to shy away from it, for fear of being called a bully. 

I think scientific criticism is a cornerstone of a reliable science. We have a responsibility to criticize public claims that we believe to be incorrect (either because they are AI generated, miscitations, or overstate the evidence). When I asked whether criticism like this should be voiced publicly (here, here, here, and here), most of the people in my network remarked that such criticisms should be voiced publicly. Others thought I should share these issues privately. In a way, I always have found it comforting to do things which you know will upset some scientists either way. It makes it easier to act on my own principles. And I believe it is essential for a science that aims to contribute to society to maintain a healthy culture of public scientific criticism.

 

 

Thanks to Nina, Sajedeh, Nick and Lisa for feedback on this blog post.

 

 

References

Andolfi, V. R., Di Nuzzo, C., & Antonietti, A. (2017). Opening the mind through the body: The effects of posture on creative processes. Thinking Skills and Creativity, 24, 20–28. https://doi.org/10.1016/j.tsc.2017.02.012

Anvari, F., Kievit, R., Lakens, D., Pennington, C. R., Przybylski, A. K., Tiokhin, L., Wiernik, B. M., & Orben, A. (2023). Not All Effects Are Indispensable: Psychological Science Requires Verifiable Lines of Reasoning for Whether an Effect Matters. Perspectives on Psychological Science, 18(2), 503–507. https://doi.org/10.1177/17456916221091565

Barel, E., Shahrabani, S., Mahagna, L., Massalha, R., Colodner, R., & Tzischinsky, O. (2024). The effects of power posing on neuroendocrine levels and risk-taking. BMC Psychology, 12(1), 726. https://doi.org/10.1186/s40359-024-02194-7

Elkjær, E., Mikkelsen, M. B., Tramm, G., Michalak, J., Mennin, D. S., & O’Toole, M. S. (2022). Using bodily displays to facilitating approach action outcomes within the context of a personally relevant task. Brain and Behavior, 13(1), e2855. https://doi.org/10.1002/brb3.2855

Körner, R., Röseler, L., Schütz, A., & Bushman, B. J. (2022). Dominance and prestige: Meta-analytic review of experimentally induced body position effects on behavioral, self-report, and physiological dependent variables. Psychological Bulletin, 148(1–2), 67–85. https://doi.org/10.1037/bul0000356

Lakens, D. (2019). The value of preregistration for psychological science: A conceptual analysis. Japanese Psychological Review, 62(3), 221–230. https://doi.org/10.24602/sjpr.62.3_221

Metzler, H., Vilarem, E., Petschen, A., & Grèzes, J. (2023). Power pose effects on approach and avoidance decisions in response to social threat. PLOS ONE, 18(8), e0286904. https://doi.org/10.1371/journal.pone.0286904

Michinov, N., & Michinov, E. (2024). Can Sitting Postures Influence the Creative Mind? Positive Effect of Contractive Posture on Convergent-Integrative Thinking. Creativity Research Journal, 36(1), 58–69. https://doi.org/10.1080/10400419.2022.2072557

Sunday, September 28, 2025

Type S and M errors as a “rhetorical tool”

Update 30/09/2025: I have added a reply by Andrew Gelman below my original blog post. 

We recently posted a preprint criticizing the idea of Type S and M errors (https://osf.io/2phzb_v1). From our abstract: “While these concepts have been proposed to be useful both when designing a study (prospective) and when evaluating results (retroactive), we argue that these statistics do not facilitate the proper design of studies, nor the meaningful interpretation of results.”

In a recent blog post that is mainly on p-curve analysis, Gelman writes briefly about Type S and M errors, stating that he does not see them as tools that should be used regularly, but that they mainly function as a ‘rhetorical tool’:

I offer a three well-known examples of statistical ideas arising in the field of science criticism, three methods whose main value is rhetorical:

[…]

2. The concepts of Type M and Type S errors, which I developed with Francis Tuerlinckx in 2000 and John Carlin in 2014. This has been an influential idea–ok, not as influential as Ioannidis’s paper!–and I like it a lot, but it doesn’t correspond to a method that I will typically use in practice. To me, the value of the concepts of Type M and Type S errors is they help us understand certain existing statistical procedures, such as selection on statistical significance, that have serious problems. There’s mathematical content here for sure, but I fundamentally think of these error calculations as having rhetorical value for the design of studies and interpretation of reported results.

The main sentence of interest here is that Gelman says this is not a method he would use in practice. I was surprised, because in their article Gelman and Carlin (2014) recommend the calculation of Type S and M errors more forcefully: “We suggest that design calculations be performed after as well as before data collection and analysis.” Throughout their article, they compare design calculations where Type S and M errors are calculated to power analyses, which are widely seen as a requirement before data collection of any hypothesis testing study. For example, in the abstract they write “power analysis is flawed in that a narrow emphasis on statistical significance is placed as the primary focus of study design. In noisy, small-sample settings, statistically significant results can often be misleading. To help researchers address this problem in the context of their own studies, we recommend design calculations”.

They also say design calculations are useful when interpreting results, and that they add something to p-values and effect sizes, which again seems to suggest they can complement ordinary data analysis: “Our retrospective analysis provided useful insight, beyond what was revealed by the estimate, confidence interval, and p value that came from the original data summary.” (Gelman & Carlin, 2014, p. 646). In general, they seem to suggest design analyses are done before or after data analysis: “First, it is indeed preferable to do a design analysis ahead of time, but a researcher can analyze data in many different ways—indeed, an important part of data analysis is the discovery of unanticipated patterns (Tukey, 1977) so that it is unreasonable to suppose that all potential analyses could have been determined ahead of time. The second reason for performing postdata design calculations is that they can be a useful way to interpret the results from a data analysis, as we next demonstrate in two examples.” (Gelman & Carlin, 2014, p. 643).

One the other hand, in a single sentence in the discussion, they also write: “Our goal in developing this software is not so much to provide a tool for routine use but rather to demonstrate that such calculations are possible and to allow researchers to play around and get a sense of the sizes of Type S errors and Type M errors in realistic data settings.”

Maybe I have always misinterpreted Gelman and Carlin, 2014, in that I took it as a paper that recommended the regular use of Type S and M errors, and I should have understood that the sentence in the discussion made it clear that this was never their intention. If the idea is to replace Type 1 and 2 errors, and hence, replace power analysis and the interpretation of data, design analysis should be part of every hypothesis testing study. Sentences such as “the requirement of design analysis can stimulate engagement with the existing literature in the subject-matter field” seemed to suggest to me that design analyses could be a requirement for all studies. But maybe I was wrong.

 

Or maybe I wasn’t.

 

In this blog post, Gelman writes: “Now, one odd thing about my paper with Carlin is that it gives some tools that I recommend others use when designing and evaluating their research, but I would not typically use these tools directly myself! Because I am not wanting to summarize inference by statistical significance.” So, here there seems to be the idea that others routinely use Type S and M errors. And in a very early version of the paper with Carlin, available here, the opening sentence also suggests routine use: “The present article proposes an ideal that every statistical analysis be followed up with a power calculation to better understand the inference from the data. As the quotations above illustrate, however, our suggestion contradicts the advice of many respected statisticians. Our resolution of this apparent disagreement is that we perform retrospective power analysis in a different way and for a different purpose than is typically recommended in the literature.”

Of course, one good thing about science is that people change their beliefs about things. Maybe Gelman one time thought Type S and M errors should be part of ‘every statistical analysis’ but now sees the tool mainly as a ‘rhetorical device’. And that is perfectly fine. It is also good to know, because I regular see people who suggest that Type S and M error should routinely be used in practice. I guess I can now point them to a blog post where Gelman himself disagrees with that suggestion.

As we explain in our preprint, the idea of Type S errors is conceptually incoherent, and any probabilities calculated will be identical to the Type 1 error in directional tests, or the false discovery rate, as all that Type S errors do is remove the possibility of an effect being 0 from the distribution, but this probability is itself 0. We also explain how other tools are better to educate researchers about effect size inflation in studies selected for significance (for which Gelman would recommend Type M errors), and we actually recommend p-uniform for this, or just teaching people about critical effect sizes.

Personally, I don’t like rhetorical tools. Although in our preprint we agree that teaching the idea of Type S and M errors can be useful in education, there are also conceptually coherent and practically useful statistical ideas that we can teach instead to achieve the same understanding. Rhetorical tools might be useful to convince people who do not think logically about a topic, but I prefer to have a slightly higher bar for the scientists that I aim to educate about good research practices, and I think they are able to understand the problem of low statistical power and selection bias without rhetorical tools.


--Reply by Andrew Gelman--


Hi, Daniel.  Thanks for your comments.  It's always good to see that people are reading our articles and blog posts.  I think you are a little bit confused about what we wrote, but ultimately that's our fault for not being clear, so I appreciate the opportunity to clarify.

So you don't need to consider this comment as a "rebuttal" to your post.  For convenience I'll go through several of your statements one by one, but my goal is to clarify.

First, I guess I should've avoided the word "rhetorical."  In my post, I characterized Ioannidis's 2005 claim, type M and S errors, and multiverse analysis as "rhetorical tools" that have been been useful in the field of science criticism but which I would not use in my own analyses.  I could've added to this many other statistical methods including p-values and Bayes factors.

When I describe a statistical method as "rhetorical" in this context, I'm not saying it's mathematically invalid or that it's conceptually incoherent (to use your term), nor am I saying these methods should not be used!  All these tools can be useful; they just rely on very strong assumptions.  P-values and Bayes factors are measures of evidence relative to a null hypothesis (not just an assumption that a particular parameter equals zero, but an entire set of assumptions about the data-generating process) that is irrelevant in the science and decision problems I've seen--but these methods are clearly defined and theoretically justified, and many practitioners get a lot out of them.  I very rarely would use p-values or Bayes factors in my work because I'm very rarely interested in this sort of discrepancy from a null hypothesis.

A related point comes up in my paper with Hill and Yajima, "Why we (usually) don't have to worry about multiple comparisons" (https://sites.stat.columbia.edu/gelman/research/published/multiple2f.pdf).  Multiple comparisons corrections can be important, indeed I've criticized some published work for misinterpreting evidence by not accounting for multiple comparisons or multiple potential comparisons--but it doesn't come up so much in the context of multilevel modeling.

Ioannidis (2005) is a provocative paper that I think has a lot of value--but you have to be really careful to try to directly apply such an analysis to real data.   He's making some really strong assumptions!  The logic of his paper is clear, though.  O'Rourke and I discuss the challenges of moving from that sort of model to larger conclusions in our 2013 paper (https://sites.stat.columbia.edu/gelman/research/published/GelmanORourkeBiostatistics.pdf).

The multiverse is a cool idea, and researchers have found it to be useful.  The sociologists Cristobal Young and Erin Cumberworth recently published a book on it (https://www.cambridge.org/core/books/multiverse-analysis/D53C3AB449F6747B4A319174E5C95FA1).  I don't think I'd apply the method in my own applied research, though, because the whole idea of the multiverse is to consider all the possible analyses you might have done on a dataset, and if I get to that point I'm more inclined to fit a multilevel model that subsumes all these analyses.  I have found multiverse analysis to be useful in understanding research published by others, and maybe it would be useful for my own work too, given that my final published analyses never really include all the possibilities of what I might have done.  The point is that this is yet another useful method that can have conceptual value even if I might not apply it to my own work.  Again, the term "rhetorical" might be misleading, as these are real methods that, like all statistical methods, are appropriate in some settings and not in others.

So please don't let your personal dislike of the term "rhetorical tools" to dissuade you from taking seriously the tools that I happen to have characterized as "rhetorical," as these include p-values, multiple comparisons corrections, Bayesian analysis with point priors, and all sorts of other methods that are rigorously defined and can be useful in many applied settings, including some of yours!

OK, now on to Type M and Type S errors.  You seem to imply that at some time I thought that these "should be part of ‘every statistical analysis,'" but I can assure you that I have never believed or written such a thing.  You put the phrase "every statistical analysis," but this is your phrase, not mine.

One very obvious way to see that I never thought Type M and Type S errors "should be part of ‘every statistical analysis'" is that, since the appearance of that article in 2014, I've published dozens of applied papers, and in only very few of these did I look at Type M and Type S errors.

What is that?  Why is it that my colleagues and I came up with this idea that has been influential, and which I indeed think can be very useful and which I do think should often be used by practitioners, but I only use it myself?

The reason is that the focus of our work on Type M and Type S errors has been to understand selection on statistical significance (as in that notorious estimate that early childhood intervention increases adult earnings by 42% on average, but with that being the result of an inferential procedure that, under any reasonable assumptions, greatly overestimates the magnitude of any real effect; that is, Type M error).  In my applied work it's very rare that I condition on statistical significance, and so this sort of use of Type M and S errors is not so relevant.  So it's perfectly coherent for me to say that Type M and S error analysis is valuable in a wide range of settings that that I think these tools should be applied very widely, without believing that they should be part of "every statistical analysis" or that I should necessarily use them for my own analyses.

That said, more recently I've been thinking that Type M and S errors are a useful approach to understanding statistical estimates more generally, not just for estimates that are conditioned on statistical significance.  I'm working with Erik van Zwet and Witold Więcek on applying these ideas to Bayesian inferences as well.  So I'm actually finding these methods to be more, not less, valuable for statistical understanding, and not just for "people who do not think logically about a topic" (in your phrasing).  Our papers on these topics are published in real journals and of course they're intended for people who <em>do</em> think logically about the topic!  And, just to be clear, I believe that you're thinking logically in your post too; I just think you've been misled by my terminology (again, I accept the blame for that), and also you work on different sorts of problems than I do, so it makes sense that a method that I find useful might not be so helpful to you.  There are many ways to Rome, which is another point I was making in that blog post.

Finally, a few things in your post that I did not address above:

1.  You quote from my blog post, where I wrote, “Now, one odd thing about my paper with Carlin is that it gives some tools that I recommend others use when designing and evaluating their research, but I would not typically use these tools directly myself! Because I am not wanting to summarize inference by statistical significance.”  That's exactly my point above!  You had it right there.

2.  You wrote, "Maybe I have always misinterpreted Gelman and Carlin, 2014, in that I took it as a paper that recommended the regular use of Type S and M errors, and I should have understood that the sentence in the discussion made it clear that this was never their intention."  So, just to clarify, yes in our paper we recommended the regular use of Type M and S errors, and we still recommend that!

3.  You write that our "sentences such as 'the requirement of design analysis can stimulate engagement with the existing literature in the subject-matter field' seemed to suggest to me that design analyses could be a requirement for all studies."  That's right--I actually do think that design analysis should be done for all studies!

OK, nothing is done all the time.  I guess that some studies are so cheap that there's no need for a design analysis--or maybe we could say that in such studies the design analysis is implicit.  For example, if I'm doing A/B testing in a company, and they've done lots of A/B tests before, and I think the new effect will be comparable to previous things being studied, then maybe I just go with the same design as in previous experiments, without performing a formal design analysis.  But one could argue that this corresponds to some implicit calculation.

In any case, yeah, in general I think that a design analysis should come before any study.  Indeed, that is what I tell students and colleagues:  never collect data before doing a simulation study first.  Often we do fake-data simulation after the data come in, to validate our model-fitting strategies, but for a while I've been thinking it's best to do it before.

This is not controversial advice in statistics, to recommend a design analysis before gathering data!  Indeed, in medical research it's basically a requirement.  In our paper, Carlin and I argue--and I still believe--that a design analysis using Type M and S errors is more valuable than the traditional Type 1 and 2 errors.  But in any case I consider "design analysis" to be the general term, with "power analysis" being a special case (design analysis looking at the probability of attaining statistical significance).  I don't think traditional power analysis is useless--one way you can see this is that we demonstrate power calculations in chapter 16 of Regression and Other Stories, a book that came out several years after my paper with Carlin--; I just think it can be misleading, especially if it is done without consideration of Type M and S errors.

Thanks again for your comments.  It's good to have an opportunity to clarify my thinking, and these are important issues in statistics.

P.S.  If you see something on our blog that you disagree with, feel free to comment there directly, as that way you can also reach readers of the original post.

--

References: 

Lakens, D., Cristian, Xavier-Quintais, G., Rasti, S., Toffalini, E., & Altoè, G. (2025). Rethinking Type S and M Errors. OSF. https://doi.org/10.31234/osf.io/2phzb_v1

Gelman, A., & Carlin, J. (2014). Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors. Perspectives on Psychological Science, 9(6), 641–651. https://doi.org/10.1177/1745691614551642