A blog on statistics, methods, philosophy of science, and open science. Understanding 20% of statistics will improve 80% of your inferences.

Showing posts with label scientific norms. Show all posts
Showing posts with label scientific norms. Show all posts

Monday, May 25, 2026

Evaluating Dr. Cuddy’s Claim that the Debunking of Power Posing is a Myth

In this blog post I will analyse the arguments that Dr. Amy Cuddy provided in a LinkedIn post “The "Power Posing Was Debunked" Myth: What the Research Actually Shows — and Why Scientific Discourse Matters” on February 26. You can find the LinkedIn post here:

https://www.linkedin.com/pulse/power-posing-debunked-myth-what-research-actually-shows-amy-cuddy-t6lnc

In the post, Cuddy says she was “effectively silenced” by an “attempt to shut down this line” of research. She credits “the courage of the individual scientists who kept going despite enormous pressure not to” for the fact that she can still summarize “what the evidence now shows”.

Power posing has two categories of claimed effects. The first effect is on self-reported feelings. For example, if we instruct people to stand in a constricted versus an expanded posture, they will self-report feeling more powerful. There is an ongoing debate about whether, or how much, this effect is caused by a demand effect (i.e., people report what they think the investigator wants them to say, not what they actually feel). A meta-analysis has shown this self-report effect is larger in within-subject designs, and in studies without a cover story (Körner et al., 2022). The second effect is on physiological or behavioral outcomes. This is the contested area, and the research outcome that Cuddy is mainly trying to defend in her blog post. If you want to explore a meta-analysis on these two categories of effects, you can do so at https://metaanalyses.shinyapps.io/bodypositions/ (made by Körner et al., 2022). I would especially recommend exploring the QRP/Publication bias tab for the physiological and behavioral outcomes.

At the end of the post, Cuddy writes that she is thankful that not everyone stopped doing research on power poses, because then: “We would not know what we now know — which is that these effects are real, that they matter, and that the story people were told was wrong.”  She concludes with: “The evidence is there. It has been there for years. All I am asking is that people look at it.”

I am happy to do so. Let’s go.

Trying to find the references

I tried to look up the references cited by Cuddy in her post. However, this reference:

Andolfi, V. R., & Antonietti, A. (2020). Contractive vs. expansive body posture effects on convergent-integrative thinking tasks. Journal of Creative Behavior, 54(4), 871–880.

does not exist in literature databases, and the authors (who do exist) do not list this paper on their own websites. An inspection of the journal’s website shows that a different article was published in volume 54, issue 4 on these pages. This raises questions about how this reference was generated, with generation by AI being a plausible candidate (also in view of the 4 malformed references I will point out below). The reference appears in the following sentence in Cuddy’s LinkedIn post:

Andolfi and Antonietti (2020, Journal of Creative Behavior) provided further evidence that contractive postures specifically benefited convergent-integrative thinking tasks. That level of specificity — where the direction of the effect depends on the type of cognitive task — is exactly the kind of finding that emerges when a field matures.

When Cuddy says ‘The evidence is there’, this is not correct for the Andolfi and Antonietti article, which does not seems to exist in the scholarly record.

There are an additional 4 references that suggest that the literature review may have in part been generated by automated tools, but for these 4 references, there are papers that match the content discussed in the literature review in the LinkedIn post.

 

Reference in LinkedIn post

Actual Reference

Michinov, E., & Michinov, N. (2020). Creativity connected with body posture: The effects of expansive and contractive postures on creative performance. Psychology of Aesthetics, Creativity, and the Arts, 14(1), 116–127

Michinov, N., & Michinov, E. (2022). Do open or closed postures boost creative performance? The effects of postural feedback on divergent and convergent thinking. Psychology of Aesthetics, Creativity, and the Arts, 16(3), 504–518. https://doi.org/10.1037/aca0000306

 

Wainio-Theberge, S., Bhatt, M., Bhattacharyya, K., et al. (2025). Neural correlates of power-related postures and their behavioural consequences: A preliminary electrophysiological investigation. Social Cognitive and Affective Neuroscience, 20(1), nsaf03

 

Wainio-Theberge, S., & Armony, J. L. (2025). Neural correlates of power-related postures and their behavioural consequences: A preliminary electrophysiological investigation. Social Cognitive and Affective Neuroscience, 20(1), nsaf036. https://doi.org/10.1093/scan/nsaf036

 

Elkjær, E., Mikkelsen, M. B., Michalak, J., Mennin, D. S., & O'Toole, M. S. (2023). Using bodily displays to facilitate approach action outcomes within the context of a personally relevant task. Frontiers in Psychology, 14, 1147printing

Elkjær, E., Mikkelsen, M. B., Tramm, G., Michalak, J., Mennin, D. S., & O’Toole, M. S. (2022). Using bodily displays to facilitating approach action outcomes within the context of a personally relevant task. Brain and Behavior, 13(1), e2855. https://doi.org/10.1002/brb3.2855

 

Körner, R., Köhler, H., & Schütz, A. (2020). Powerful and confident children through expansive body postures? A preregistered test of the effects of power posing on children. School Psychology International, 41(4), 315–330.

 

Körner, R., Köhler, H., & Schütz, A. (2020). Powerful and confident children through expansive body postures? A preregistered study of fourth graders. School Psychology International, 41(4), 315–330. https://doi.org/10.1177/0143034320912306

 

 

We see all these references that are incorrect refer to the later literature, and summarize the research of the people who ‘kept going’. These references are at the core of the argument Cuddy is making.

Evaluating the evidence: Three examples

Cuddy wrote a narrative review, which requires that the validity of the conclusions, and the strength of the evidence, needs to be evaluated for every study. Let’s carefully examine some of the papers she cited and evaluate the evidence. Cuddy writes about a first study:

Wainio-Theberge and colleagues (2025, Social Cognitive and Affective Neuroscience) published the first EEG study of power posing, finding significant effects on arousal and valence, with suggestive differences in frontal brain activity between expansive and contractive postures. A new neural methodology for a question people said was already answered.

From the description in Cuddy’s blog, you might assume the “significant effects on arousal and valence, with suggestive differences in frontal brain activity between expansive and contractive postures” would support the hypothesis. But this is not the case. The significant effects were actually in the opposite direction of the hypothesis. This is not mentioned in the abstract of the Wainio-Theberge et al. article, and one would need to read the paper to get this information:

We found no significant posture differences in the EEG spectral exponent (t(101) = 1.01, P = .32). In contrast, a significant posture effect was observed for frontal asymmetry (t(101) = −2.63, P = .01); however, post hoc t-tests in each group separately (‘Models 1c and 1e’) revealed that the effect was in the opposite direction as hypothesized (see Discussion). Namely, we observed a significant right-lateralized frontal alpha asymmetry (FAA) in the contractive group (t(45) = 2.17, P = .04) and a left-lateralized FAA in the expansive one which failed to reach significance (t(55) = −1.63, P = .11).

Cuddy writes in the blog that she responded to journalists skeptical about power posing: “I spent more than ten hours responding — reviewing the literature, pulling citations, writing carefully, anticipating distortions” In this case, her review of the literature presented a finding as providing support for power posing, when in fact the effect was in the opposite direction of the hypothesis.

As a second paper, let’s take Barel and colleagues (2024). First, I want to thank the authors for sharing their data, after I tried to access it by clicking the google drive link in the article. All numbers were reproducible. Cuddy cites the paper as follows:

“As other researchers began testing that broader construct, using different measures in different populations, they found effects consistently: action orientation (Huang et al., 2011, Psychological Science), […], and risk-taking itself, partially (Barel et al., 2024, BMC Psychology).

It is unclear what is meant by 'partially', as the authors are clear that they found that power posing did not affect risk-taking: "There was no statistically significant distribution in risk-taking between high and low power conditions [χ2 = 0.00, p > 0.99]." The risk-taking outcome that Cuddy cites the study for is a clear null result.

The basis for "partially" is presumably a separate analysis reported in the paper: in a logistic regression predicting risk-taking, the authors found a significant interaction between power condition and cortisol change and write that they "did partially replicate an effect of changes in cortisol levels on risk-taking." But note that they claim an effect of cortisol changes on risk, not of power posing on risk. For power posing to affect risk through cortisol, power posing would first have to change cortisol, and it did not: the authors report no main effects of time or power on cortisol. With that first link missing, the high-power participants whose cortisol fell are not a subgroup of people for whom the power pose worked, as their cortisol would have moved the same way without any pose. The significant effect is a within-group association between two measures, which can't be attributed to the power posing manipulation.

To their credit, the authors themselves never claim power posing affected risk-taking. This framing comes from Cuddy, who presents the paper as a partial replication of a risk-taking effect after a power posing manipulation, which the study did not support.

When discussing a third paper, Cuddy writes: “Körner, Köhler, and Schütz (2020, School Psychology International) conducted a preregistered study of 108 German fourth graders — children — and found that expansive postures increased self-esteem, positive feelings, feelings of power, and even children's perceptions of their relationship with their teacher. The strongest effects were on school-related self-esteem. This is exactly the kind of applied, developmentally informed research that matters — taking findings from the lab and asking whether they help real children in real classrooms.”

The Körner et al study was preregistered: https://aspredicted.org/blind.php?x=sn4su9 with 4 t-tests to examine 4 dependent variables of interest. Of the 4 tests, 2 are significant (p = 0.04 and p = 0.013), but neither survive a correction for multiple comparisons (0.05/4 = 0.0125) which was necessary in this analysis.

The blog by Cuddy states “The strongest effects were on school-related self-esteem.” But the biggest effect is actually on the student-teacher relationship:

Finally, there was a significant difference between the two groups regarding the pictures related to the student–teacher relationship: high power posers more frequently chose the picture showing a good student–teacher relationship than low power posers, Χ²(1) = 11.181, p = .001, φ = –.322.

But there is a problem with this finding. Students spent months building a relationship with their teacher. Then, as part of the experiment, the students posed for 60 seconds and self-reported on that relationship, without any further interaction with the teacher.  There is no possible causal mechanism for the power pose to impact the relationship with teachers. Although unintended, this question is an excellent probe for demand effects. As the power pose can’t change history and impact the actual relationship between students and teachers, the observed effect can only be caused by a demand effect. Neither Cuddy nor the original authors realized this. Cuddy instead concludes: “This is exactly the kind of applied, developmentally informed research that matters — taking findings from the lab and asking whether they help real children in real classrooms.”

Evaluating the Research Line

Evaluating evidence is effortful and messy. Single studies always have weaknesses, and the reader might reasonably wonder whether I’m cherry-picking a few bad apples from an otherwise strong set. I don’t think I am, and I will explain the more general pattern I observed when reading all the cited papers.

Exploratory claims

The Körner et al (2020) study above was preregistered, and therefore we were able to evaluate that the claims were not severely tested, as they would not survive the required correction for multiple comparisons (Lakens, 2019). But most claims in the papers that Cuddy cites are based on exploratory analyses. The studies all have many dependent variables, and a large number of tests can be performed. These studies observe a mix of significant and non-significant results, but the significant results have a high probability of being Type 1 errors and can’t be presented as evidence. If researchers in this field would perform more direct replication studies, and would preregister their studies more, they could address this problem. Some preregistered their studies, which is excellent, but some don't, even though they work in a highly contested research area, and the significant results primarily come from exploratory analyses.

Researchers in the field are often honest about this, but especially in a narrative summary, it is easy to lose track of the fact that most of the authors of studies cited by Cuddy do not consider their own findings to be strong evidence. For example, Metzler et al (2023) write “Finally, it is important to transparently report on the level of evidence this study provides for power pose effects on low-level social behavior. This requires mentioning its exploratory nature [...] we are convinced that the medium effect sizes, given our sample size, would require replication before strong conclusions can be drawn”. I would say this is especially important given that the main result was a 3-way interaction with a p-value of 0.03: “the predicted three-fold interaction suggested that this effect of emotion on action choices (more avoidance for anger than fear) changed between sessions as a function of adopted pose (OR = 1.19, 95% CI[1.02, 1.38], z = 2.18, p = .029)”.

Another example comes from Elkjær et al (2022). The main finding is: “Concerning approach tendencies, the 2 × 3 interaction analysis on DAT “approach threat 1” was significant (F(1, 87) = 3.27, p = .043, ηp2 = .07). Regarding DAT avoid threat (1 + 2), the overall 2 × 3 interaction analysis was significant (F(1, 87) = 6.39, p = .003, ηp2 = .13).” The study was preregistered (https://aspredicted.org/blind.php?x=9j3b38) which allows us to see that the preregistered predictions are not supported. The authors predicted significant effects for the expansive condition compared to both the constricted condition and the control condition. However, they did not find effects compared to the control condition. Such patterns of mixed results are present in many studies in the literature. On the one hand, this is part of normal research, especially early on in research lines, when researchers have not figured out how to reliably produce the effect they are examining. On the other hand, power posing has been studied since 2010, and a research line can never get a strong basis if it does not move beyond a literature where all significant results are based on exploratory partial confirmations.

If you want to see the exploration of data in action, I would recommend looking at the OSF repository related to the paper by Michinov and Michinov (2024): www.osf.io/c9mzh, and see which variables and ways of computing variables are reported in the final paper, and which are not.

Underpowered studies and selection for significance

The sample sizes in the studies cited by Cuddy are often small – especially for key sub-group analyses, when the total sample size might be distributed across cells in a 2x3 design. This would not be problematic if the effects of power posing were known to be large. But even the self-report effect where participants indicate they feel more or less powerful has a rather small effect size of only g = 0.37 (see https://metaanalyses.shinyapps.io/bodypositions/). Less direct effects, for example on behavior, are likely to have a much smaller effects (unless researchers can propose strong theoretical arguments why more indirect effects would be larger, see Anvari et al., 2023). In one-tailed independent t-tests, 80% power would require 184 participants (92 per condition), but none of the studies are close to achieving such sample sizes.

The research area of power posing is also characterized by the selective reporting of significant results. This combination of underpowered studies and selection for significance leads to highly inflated effect sizes. We can see these effects in Andolfi et al (2017):

The effect sizes of an open or closed posture simply can’t be in the range of d = 1.22, or even d = 0.69 (for examples of realistic effect sizes to expect based on group differences, see DataColada 18). The effects are inflated, and there is no way of knowing what the true effect sizes are. They might be zero, as many replication studies of exactly such implausibly large effects based on studies with tiny samples have turned out to be.

The study by Michinov and Michinov similarly shows effects for significant tests that are too large. Adopting a posture for a few minutes can’t plausibly influence creative tasks with effects such as d = 0.634. When you evaluate evidence, thinking about selective reporting and inflated effects should be part of the evaluation.

 

Quality of the design and analysis

I could not help noticing that there is a lot of room to improve the quality of the study design and analysis, as reported in papers in this literature. This in itself does not mean that the evidence is unreliable, but it does not make it easier for a research field to generate high quality evidence. For example, Elkjær et al (2022) report the following power analysis:

“Based on a priori power calculations, using a repeated-measures ANOVA interaction analysis, 2 (time; before vs. after the manipulation) × 3 (condition; EXP, CON, N), 90 participants were required to detect a small effect size (d = 0.34), with an alpha of .05 and a beta of .20.”

At first sight, this looks like best practice. They acknowledge power posing effects are small (d = 0.34 is very much in line with the meta-analysis they published in the same year). Regrettably, what the authors actually did was enter an f = -.34, not a d, as you can see in the screenshot below, which leads to a sample size that is much lower than what they would actually have needed to achieve high power, according to their own meta-analytic effect size estimate:

This means that despite the power analysis, the study was still massively underpowered. The sample size justifications in all studies cited by Cuddy are problematic. This is probably true for many research lines, but it is especially problematic for a research line where researchers are still trying to establish if the basic effect exists or not.

While reading the articles, I also noticed many of the issues that we often see in other literatures when research teams lack statistical expertise. There are often small inconsistencies in the correct degrees of freedom, incorrectly performed statistical tests, an overreliance on p-values despite underpowered studies, and misinterpretations of non-significant results. I don’t want to single out more examples, but it would probably be good for the field if researchers would enlist some methodological and statistical expertise if they want to generate reliable evidence.  

Tools to evaluate claims

Cuddy writes: “When people are told that research is fake — without being given the tools to evaluate that claim — it doesn't just affect one researcher or one line of work. It feeds a broader cynicism: that science can't be trusted, that findings are arbitrary, that expertise is performance.” I strongly agree. This is why I have created a free textbook, Improving Your Statistical Inferences, to learn how to evaluate the actual evidence in scientific papers. Here are three decent heuristics to follow when you evaluate the evidence in a research line:

  1. If a finding shows what you want to be true, be extra skeptical.
  2. If you have a strong conflict of interest, be extra skeptical.
  3. Studies with low power due to too small sample sizes, lack of preregistration, no direct replications, strong indications of selective reporting, low methodological quality, repeating limitations in discussion sections without addressing them, implausibly large effect sizes, a lack of impact on other research areas, significant claims that mainly come from exploratory analyses, continued uncertainty about the basic effect after more than a decade and dozens of studies, and the research community disengaging with a literature are all signs of a lack of evidence.

According to Cuddy, she “live[s] inside a false narrative” where power posing is incorrectly believed to be a ‘myth’, and she believes that “none of this would have happened if the methods guys, and the journalists who trusted them without doing proper research, hadn't created the conditions that made it happen.”

 

Scientific criticism is a cornerstone of a healthy science

When I read Cuddy’s LinkedIn post, I was highly skeptical of the claim that there was evidence for effects of power posing on measures other than self-report, and that the debunking was a 'myth'. But my first response was to ignore the post. I did not want to examine the evidence behind the claims Cuddy made, because I am clearly one of the “method guys” who, according to Cuddy “manufactured the "debunked" narrative and aimed it, with great precision, at a single researcher”. If I would criticize her post, would I be seen as contributing to “the bullying I was subjected to”, as Cuddy writes?

But I care about criticism in science. And I think it is important that we can criticize scientific claims. My decision to not follow up on examining the claims in the blog post kept nagging me. Cuddy has 900,000 followers on LinkedIn who have read the very strong statement that it is a “myth” that power posing was debunked. If the evidence she presented was overstated – as I feared – scientific criticism would be needed to correct the record. I think it is essential to increase social safety in academia, while being able to criticize each other. I do not want bullying and scientific criticism to become conflated. Scientific criticism is too important for a healthy science to shy away from it, for fear of being called a bully. 

I think scientific criticism is a cornerstone of a reliable science. We have a responsibility to criticize public claims that we believe to be incorrect (either because they are AI generated, miscitations, or overstate the evidence). When I asked whether criticism like this should be voiced publicly (here, here, here, and here), most of the people in my network remarked that such criticisms should be voiced publicly. Others thought I should share these issues privately. In a way, I always have found it comforting to do things which you know will upset some scientists either way. It makes it easier to act on my own principles. And I believe it is essential for a science that aims to contribute to society to maintain a healthy culture of public scientific criticism.

 

 

Thanks to Nina, Sajedeh, Nick and Lisa for feedback on this blog post.

 

 

References

Andolfi, V. R., Di Nuzzo, C., & Antonietti, A. (2017). Opening the mind through the body: The effects of posture on creative processes. Thinking Skills and Creativity, 24, 20–28. https://doi.org/10.1016/j.tsc.2017.02.012

Anvari, F., Kievit, R., Lakens, D., Pennington, C. R., Przybylski, A. K., Tiokhin, L., Wiernik, B. M., & Orben, A. (2023). Not All Effects Are Indispensable: Psychological Science Requires Verifiable Lines of Reasoning for Whether an Effect Matters. Perspectives on Psychological Science, 18(2), 503–507. https://doi.org/10.1177/17456916221091565

Barel, E., Shahrabani, S., Mahagna, L., Massalha, R., Colodner, R., & Tzischinsky, O. (2024). The effects of power posing on neuroendocrine levels and risk-taking. BMC Psychology, 12(1), 726. https://doi.org/10.1186/s40359-024-02194-7

Elkjær, E., Mikkelsen, M. B., Tramm, G., Michalak, J., Mennin, D. S., & O’Toole, M. S. (2022). Using bodily displays to facilitating approach action outcomes within the context of a personally relevant task. Brain and Behavior, 13(1), e2855. https://doi.org/10.1002/brb3.2855

Körner, R., Röseler, L., Schütz, A., & Bushman, B. J. (2022). Dominance and prestige: Meta-analytic review of experimentally induced body position effects on behavioral, self-report, and physiological dependent variables. Psychological Bulletin, 148(1–2), 67–85. https://doi.org/10.1037/bul0000356

Lakens, D. (2019). The value of preregistration for psychological science: A conceptual analysis. Japanese Psychological Review, 62(3), 221–230. https://doi.org/10.24602/sjpr.62.3_221

Metzler, H., Vilarem, E., Petschen, A., & Grèzes, J. (2023). Power pose effects on approach and avoidance decisions in response to social threat. PLOS ONE, 18(8), e0286904. https://doi.org/10.1371/journal.pone.0286904

Michinov, N., & Michinov, E. (2024). Can Sitting Postures Influence the Creative Mind? Positive Effect of Contractive Posture on Convergent-Integrative Thinking. Creativity Research Journal, 36(1), 58–69. https://doi.org/10.1080/10400419.2022.2072557

Wednesday, August 9, 2023

How I tried to get a paper that I own retracted: A journey and call to action against Prime Scholars

This is a Guest Post by Noah van Dongen



Right after the replication emergency in mental science, numerous clinicians have embraced rehearses that support the vigor and straightforwardness of the logical cycle, including preregistration, information sharing, code sharing, and enormous reproducibility studies
 
not Noah van Dongen, but somebody else (2022)


This is a true story about how I tried to get a paper that I own retracted. The copyright of the paper in question is attributed to me. Several requests and demands for removal were issued, but the journal has not yet fully complied. Of course, this might be due to the fact that I did not write the paper, the paper is an incoherent mess of academic gibberish, and Prime Scholars is the publisher of the journal. This post tells my story. It ends with a call to action against Prime Scholars and similar outfits.

For those of you who are unfamiliar with Prime Scholar, this is what Wikipedia has to say about them:




As Bishop notes, some famous, though long dead, authors are counted among the people that publish in Prime Scholars’ journals.1





Apparently, I can now count myself among the illustrious company of Hesse, Bronte, and Whitman. For I too, without my knowledge and against my wishes, am now a scholar that published a word soup essay in a Prime Scholar journal.

The paper that I never wrote

 
I will not keep you in suspense any longer and tell you which journal and publication I am talking about. Acta Psychopathologica is one of the 56 journals owned by Prime Scholars. It claims to have a Journal Impact Factor of 2.15 (or 2.4 according to their about page) and has 524 citations reported on Google Scholar (or 447 according to Google Scholar). The journal has published 9 volumes, with some as many as 9 issues, and a total of 9 special issues. The academic trainwreck attributed to me was published in May 2022, as part of the fourth issue of volume 8. There are four other publications in this issue. For as far as I can tell, all of their authors are existing and living scientists. And all publications are of similar quality: syntactically it looks like English, but it reads like a bad acid trip (or the ramblings of a drunk philosopher).

And now, the word vomit in question. The literary disaster is titled “Phenomena can be Characterized as General Patterns in Observations of Psychology Theories.” and is attributed to a single author, Noah Van Dongen. The efficiency of the review process is astounding:

  • 1 April 2022: submission (nice touch on April fools)

  • 4 April 2022: editor assigned

  • 18 April 2022: reviewed

  • 25 April 2022: revised

  • 2 May 2022: published

One month and a day between submission and publication! Receiving and correcting proofs from a journal’s copy editor usually takes me more than a month! It is also very neat that, apart from the submission to editor assignment, everything else was done in exactly a week; like clockwork! I wish my colleagues and I could work this organized and systematic. The terrible text(ual) trifle2 cannot be found on Google Scholar, it does not show up in the Google Scholar page of the journal, and its DOI does not exist. Services, like Google Scholar and Researchgate, who typically notify me when something academic is published under my name, have not picked up on this addition to my oeuvre. The only reason I know of its existence is that a colleague came across it in a Google search for the definition of “phenomena”.

About the content of the linguistic barf, just like the other papers in Issue 4 of Volume 8, it consists of a collection of English sentences that appear to approach syntactic correctness, though devoid of any clear meaning. The only thing it has going for it, is that it is blessedly short. It is good to read that there were no conflicts of interest, but it’s too bad that they misspelled my name. In the Netherlands, our surnames can have prefixes, like “van” or “van der” (which translates to “from”), which are not capitalized. My surname is spelled “van Dongen” not “Van Dongen”. Just pointing this out. However, considering the quality of the commentary, it is surprising they got so close to getting my personal information correct. Yes, this verbose vacuity is published as a commentary, though for the life of me I cannot figure out what it is supposed to be commenting on.

Trying to make sense of the senseless drivel, it seems like the true authors have taken part of the introduction from a preprint (that I did write) and ran it through a thesaurus to avoid being flagged for plagiarism, creating what is called tortured phrases. The paper that they used as the mold seems to be an early version of Productive Explanation: A Framework for Evaluating Explanations in Psychological Science, which was posted on PsyArXiv on 13 April 20223. For comparison, here are the first sentences of the original and the forgery, respectively:

In the wake of the replication crisis in psychological science, many psychologists have adopted practices that bolster the robustness and transparency of the scientific process, including preregistration (Chambers, 2013), data sharing (Wicherts et al., 2006), code sharing, and massive reproducibility studies (e.g., Aarts et al., 2015; Walters, 2020).

Right after the replication emergency in mental science, numerous clinicians have embraced rehearses that support the vigor and straightforwardness of the logical cycle, including preregistration, information sharing, code sharing, and enormous reproducibility studies.

I did not take the time to figure out which other sentences they used for the figurative butchery (or is ‘literal’ more appropriate here?) and what kind of procedure they used to end up with this collection of sentences. It looks like they just selected a certain part of the introduction, but I am not sure.

The process of getting the paper retracted and succeeding partially

 
I think this is enough for setting the stage. Let me now tell you about my journey of getting the verbal catastrophe retracted, of which I explicitly own the copyrights. This story started in November of 2022 when a colleague stumbled across the linguistic salad. Ironically enough, we were working on the original, wanting to improve our definition of phenomena, and searching for definitions by others that we might be able to use. It is quite a surreal experience to see your credentials on work you know you didn’t produce while simultaneously searching your memory for what you could have done to make this happen.

As a true academic, I acted immediately, ten days later. My first attempt was emailing the journal requesting the removal of the textual horror. This was on 8 December 2022.4 I gave the journal ample time to respond (or I forgot about this problem due to other stuff that was going on at the moment). But, after three months without reply, I contacted the legal team of the University of Amsterdam. They were very understanding and wanted to be of assistance.

On 11 2023, the UvA’s legal team emailed Acta Psychopathologica demanding the retraction of the atrocious article and threatened legal action if they did not comply. Acta Psychopathologica did not respond to this email either. On 20 April, the legal team sent a reminder, to which they also did not reply. At the moment, Acta Psychologica has not responded to any of our messages.

The UvA’s legal team had also advised me to report the identity theft to the Rijksdienst Identiteitsgegevens (RVIG; translation: the identity data safety services of the Dutch government), which I did on 11 April 2023. Conveniently, the RVIG has an online form for this. However, identity theft is usually about the illegitimate use of your credit cards or passports. It was more than a bit awkward to write about how you have been impersonated to publish shoddy scientific work in (a website that calls itself a) journal. Nine days after I submitted the form, the RVIG contacted me. They were very understanding, but told me there was nothing they could do for me. They advised me to report the crime to the police.

Again, other responsibilities got in the way. About a month later, I called the police on 1 June 2023. The central operator noted down the specifics of my predicament and told me that the department responsible for identity theft would contact me to make an appointment, which they did on Saturday 3 June 2023. As it turns out, reporting identity theft must be done in person at the police station. Maybe this is to make sure that it is actually you that is reporting the theft of your identity. On 13 June 2023, I went to my appointment to report the crime, which took about an hour. The officer taking my report was friendly and understanding. Actually, she was very understanding considering the curiousness of the situation. Noteworthy about this experience, is that she wrote the report from my (first person) perspective, but in her own words. There are many new experiences I am gaining throughout this journey, and this was one of the strange ones. Reading something as if you said it, correct in terms of content, though not in a way that you would say, is surreal to say the least. You are instantly aware of your own idiolect. Or at least, that is what I experienced.

A week later, I informed the UvA’s legal team of the police report. They promptly sent another email to Acta Psychopathologica requesting them again to remove the semantic puree, though this time adding that the crime had been reported to the police and that legal action would follow.

The UvA’s legal team also contacted the Editors-in-Chief of Acta Psychopathologica to request them to remove the language pit stain. They received two replies to this request, which can be summarized as: I did not actually do anything for this journal and I would like to be removed from the editorial board.

On 7 July 2023 the horrendous word swirl was no longer visible on the website. There are now only four papers in Volume 8 Issue 4 instead of five. However, the pdf version of the paper can still be reached and I am still listed as an author on the Prime Scholars website.

What is next?

 
My paper is not the only instance of identity theft in Acta Psychopathologica and other Prime Scholars journals. Currently, I am contacting the other authors in Acta Psychopathologica one by one to ask if they wrote the paper that is attributed to them but this is a slow process. My aim is to make them aware of the fraud and start requesting the traction of the fraudulent papers.

I’ve also come across this article in The Times Higher Education, which mentions legal actions being undertaken. I’m trying to find out if legal actions against Prime Scholars are indeed in the works. If so, I hope I can join them. If not, I want to get them started.

What you can do to help

  1. Check if you or people you know have papers attributed to them in a Prime Scholar journal.
  2. Contact me if you are also a victim of identity theft and want to join legal actions against Prime Scholars.
  3. Share this story and (ask people to) take action against Prime Scholars.
Generating academic articles that appear authentic has become much easier now large language models like ChatGPT have arrived on the scene. I think it is safe to assume that we don’t want fake papers to start invading our academic corpusses; devaluing our work and eroding the public’s trust in science. This does not seem to be a problem that will solve itself. We need to take active steps against these practices and we need to act now!

One last thought about publishers


Don’t you think that ‘respectable’ academic publishers should accept some responsibility for this predicament we find ourselves? If outfits like Prime Scholars are making a mockery of their business and start poisoning the well, should they not also undertake (legal) steps to protect their profession? 
 
 
1 Also, see this article on Retraction Watch↩
2 For the interested, a trifle is a layered dessert of English origin. In Friends’ episode 9 of season 6, Rachel tries to make this dessert for Thanksgiving and accidentally adds beef sautéed with peas and onion to the dessert. The result still sounds better than the paper in question.↩
3 The astute reader realized that this is 12 days after it was supposedly submitted to Acta Psychopathologica. For everybody else is now also aware due to this informative footnote.↩
4 I know you Americans like to put the month in front of the date, that just looks wrong.↩

Sunday, October 31, 2021

Not All Flexibility P-Hacking Is, Young Padawan

During a recent workshop on Sample Size Justification an early career researcher asked me: “You recommend sequential analysis in your paper for when effect sizes are uncertain, where researchers collect data, analyze the data, stop when a test is significant, or continue data collection when a test is not significant, and, I don’t want to be rude, but isn’t this p-hacking?”

In linguistics there is a term for when children apply a rule they have learned to instances where it does not apply: Overregularization. They learn ‘one cow, two cows’, and use the +s rule for plural where it is not appropriate, such as ‘one mouse, two mouses’ (instead of ‘two mice’). The early career researcher who asked me if sequential analysis was a form of p-hacking was also overregularizing. We teach young researchers that flexibly analyzing data inflates error rates, is called p-hacking, and is a very bad thing that was one of the causes of the replication crisis. So, they apply the rule ‘flexibility in the data analysis is a bad thing’ to cases where it does not apply, such as in the case of sequential analyses. Yes, sequential analyses give a lot of flexibility to stop data collection, but it does so while carefully controlling error rates, with the added bonus that it can increase the efficiency of data collection. This makes it a good thing, not p-hacking.

 

Children increasingly use correct language the longer they are immersed in it. Many researchers are not yet immersed in an academic environment where they see flexibility in the data analysis applied correctly. Many are scared to do things wrong, which risks becoming overly conservative, as the pendulum from ‘we are all p-hacking without realizing the consequences’ swings back to far to ‘all flexibility is p-hacking’. Therefore, I patiently explain during workshops that flexibility is not bad per se, but that making claims without controlling your error rate is problematic.

In a recent podcast episode of ‘Quantitude’ one of the hosts shared a similar experience 5 minutes into the episode. A young student remarked that flexibility during the data analysis was ‘unethical’. The remainder of the podcast episode on ‘researcher degrees of freedom’ discussed how flexibility is part of data analysis. They clearly state that p-hacking is problematic, and opportunistic motivations to perform analyses that give you what you want to find should be constrained. But they then criticized preregistration in ways many people on Twitter disagreed with. They talk about ‘high priests’ who want to ‘stop bad people from doing bad things’ which they find uncomfortable, and say ‘you can not preregister every contingency’. They remark they would be surprised if data could be analyzed without requiring any on the fly judgment.

Although the examples they gave were not very good1 it is of course true that researchers sometimes need to deviate from an analysis plan. Deviating from an analysis plan is not p-hacking. But when people talk about preregistration, we often see overregularization: “Preregistration requires specifying your analysis plan to prevent inflation of the Type 1 error rate, so deviating from a preregistration is not allowed.” The whole point of preregistration is to transparently allow other researchers to evaluate the severity of a test, both when you stick to the preregistered statistical analysis plan, as when you deviate from it. Some researchers have sufficient experience with the research they do that they can preregister an analysis that does not require any deviations2, and then readers can see that the Type 1 error rate for the study is at the level specified before data collection. Other researchers will need to deviate from their analysis plan because they encounter unexpected data. Some deviations reduce the severity of the test by inflating the Type 1 error rate. But other deviations actually get you closer to the truth. We can not know which is which. A reader needs to form their own judgment about this.

A final example of overregularization comes from a person who discussed a new study that they were preregistering with a junior colleague. They mentioned the possibility of including a covariate in an analysis but thought that was too exploratory to be included in the preregistration. The junior colleague remarked: “But now that we have thought about the analysis, we need to preregister it”. Again, we see an example of overregularization. If you want to control the Type 1 error rate in a test, preregister it, and follow the preregistered statistical analysis plan. But researchers can, and should, explore data to generate hypotheses about things that are going on in their data. You can preregister these, but you do not have to. Not exploring data could even be seen as research waste, as you are missing out on the opportunity to generate hypotheses that are informed by data. A case can be made that researchers should regularly include variables to explore (e.g., measures that are of general interest to peers in their field), as long as these do not interfere with the primary hypothesis test (and as long as these explorations are presented as such).

In the book “Reporting quantitative research in psychology: How to meet APA Style Journal Article Reporting Standards” by Cooper and colleagues from 2020 a very useful distinction is made between primary hypotheses, secondary hypotheses, and exploratory hypotheses. The first consist of the main tests you are designing the study for. The secondary hypotheses are also of interest when you design the study – but you might not have sufficient power to detect them. You did not design the study to test these hypotheses, and because the power for these tests might be low, you did not control the Type 2 error rate for secondary hypotheses. You can preregister secondary hypotheses to control the Type 1 error rate, as you know you will perform them, and if there are multiple secondary hypotheses, as Cooper et al (2020) remark, readers will expect “adjusted levels of statistical significance, or conservative post hoc means tests, when you conducted your secondary analysis”.

If you think of the possibility to analyze a covariate, but decide this is an exploratory analysis, you can decide to neither control the Type 1 error rate nor the Type 2 error rate. These are analyses, but not tests of a hypothesis, as any findings from these analyses have an unknown Type 1 error rate. Of course, that does not mean these analyses can not be correct in what they reveal – we just have no way to know the long run probability that exploratory conclusions are wrong. Future tests of the hypotheses generated in exploratory analyses are needed. But as long as you follow Journal Article Reporting Standards and distinguish exploratory analyses, readers know what the are getting. Exploring is not p-hacking.

People in psychology are re-learning the basic rules of hypothesis testing in the wake of the replication crisis. But because they are not yet immersed in good research practices, the lack of experience means they are overregularizing simplistic rules to situations where they do not apply. Not all flexibility is p-hacking, preregistered studies do not prevent you from deviating from your analysis plan, and you do not need to preregister every possible test that you think of. A good cure for overregularization is reasoning from basic principles. Do not follow simple rules (or what you see in published articles) but make decisions based on an understanding of how to achieve your inferential goal. If the goal is to make claims with controlled error rates, prevent Type 1 error inflation, for example by correcting the alpha level where needed. If your goal is to explore data, feel free to do so, but know these explorations should be reported as such. When you design a study, follow the Journal Article Reporting Standards and distinguish tests with different inferential goals.

 

1 E.g., they discuss having to choose between Student’s t-test and Welch’s t-test, depending on wheter Levene’s test indicates the assumption of homogeneity is violated, which is not best practice – just follow R, and use Welch’s t-test by default.

2 But this is rare – only 2 out of 27 preregistered studies in Psychological Science made no deviations. https://royalsocietypublishing.org/doi/full/10.1098/rsos.211037 We can probably do a bit better if we only preregistered predictions at a time where we really understand our manipulations and measures.