Psychology is not mostly fine
Astral Codex endorses justice delayed on the replication crisis
Last week, maybe the most prominent blogger to have touched the replication crisis, and author of one of its classics, Scott Alexander, wrote a piece that surprised some of his audience. “Psychology Research Is Mostly Fine” was true to its name. In the follow-up, he addressed the apparent backlash and said:
“...just like in economics, medicine, sociology, political science, or even some of the hard sciences. All of these encountered the replication crisis at different times, all of them adjusted in different ways, and they’ve all risen to the occasion in some cases and failed to confront others.”
This doesn’t really mean anything, but to the extent that it does, it’s wrong. Alexander doesn’t have anything to cite here because nearly everything about psychology and beyond says that fields have not: 1. Tested themselves at all, 2. Retested to show improvement, 3. Done much of anything to address the crisis or even admit that word crisis.
On psychology, he goes on, “Psychology is no different. It has flaws, but most of it consists of good ideas and trustworthy research.”
I’ll present an argument against a fortiori using a textbook I happened to have to read recently. Other ways of making the same argument are here and here.
Fields have not taken the first step, education, and they’re not doing well with the others. Education is a great place to start between the plentiful large-scale, roughly 50% findings, because it is so miserably universal. There is good reason to believe that schools don’t teach the replication crisis, updated methods, or even the old methods properly according to longstanding complaints that predate the crisis.
One might assume, for instance, in the field most aware of the problems, psychology, that this 2024 textbook “Methods in Behavioral Research” by Paul Cozby would encounter and adjust.
No. Instead, what we get is a speedrun of every hypothesis-testing misnomer there is. The authors encourage the exact problem that caused the replication crisis, that “Conceptual replications are even more important than exact replications in furthering our understanding of behavior,” that we gain more certainty not with independent results or more observations, but by testing a phenomenon in different ways. Critically, they endorse the granddaddy phenomenon that will cover all the retrying up: “Many rejected papers are submitted to other journals and eventually accepted for publication, but much research is never published. This is not necessarily bad; it simply means that selection processes separate high-quality research from that of lesser quality.”
They don’t mention the danger in multiple methods, p-hacking and publication bias. This doesn’t mean that the authors have a mistaken view of p-values as a mark of quality, but we can say for certain that they do have that interpretation because they say so:
“The logic of the null hypothesis is this: If we can determine that the null hypothesis is incorrect, then we accept the research hypothesis as correct. Acceptance of the research hypothesis means that the independent variable had an effect on the dependent variable.”
Even though this is more or less the classic misinterpretation of significance testing, and the worst one, let’s give them the benefit of the doubt and say they mean in dichotomous and exhaustive cases where the prior is also very high and excuse the causal language. They paint over the problem of priors by casually assuming the hypothesis is always true:
“you are most likely to obtain significant results when you have a large sample size because larger sample sizes provide better estimates of true population values.”
Not to be left out, the classic mistake that random error is the only other explanation for the effect:
“Statistical inference begins with a statement of the null hypothesis and a research (or alternative) hypothesis. The null hypothesis is simply that the population means are equal—the observed difference is due to random error.”
And the one about the true value probably being within the confidence interval:
“A news story might report that a sample of people in your state found that 55% support a measure that will increase education funding in your state; 45% oppose the measure. The report then says that these results are accurate to within 3 percentage points, with a 95% confidence level. This means that the researchers are very (95%) confident that, if they were able to study the entire population rather than a sample, the actual percentage who support the education measure would be between 58% and 52% and the percentage opposing the measure would be between 48% and 42%.”
In a word, this book is absolute shit. Whatever lip service it gives the crisis and possible remedies with one hand, it undoes them with the other. The furthest it goes is to say that the result from the famous 2015 study of the replication rate in psychology was disappointing.
The twist
Of course we have to go where the data leads us and in this case the data almost made me fall off my chair. Paul Cozby wrote another “Methods in Behavioral Research” book that corrects all of these issues and more. It was written with Raymond Mar and it’s the Canadian edition of this book.
It’s not just good. It’s pixel perfect:
“Research out of the University of Guelph also found that almost 90 percent of introductory psychology textbooks that define the p-value fail to define it correctly (Cassidy et al., 2019).” Indeed. Indeed.
So the only thing to do is to re-examine Alexander’s statement. It’s true that fields have confronted the replication crisis at different, wildly inappropriate and late, times. And they’ve adjusted in different ways at different times like the different ways that exist on one side of the Canadian border or the other, or one side of Paul Cozby’s desk or the other. The same ISBN is still being sold as 2026 “evergreen” by the publisher.
I don’t know how to put a finer point on this other than to say that the readers of this book will grow up to make false statements to our government for money. In Canada, maybe you’ll be okay. Just make sure you buy the right one.


