As the throughput of scientific instruments has risen over the last twenty years, the discipline required of researchers has risen too. If someone sees a publishable result flash across their computer screen, they have to admit it, along with all the unpublishable numbers they saw too. Science requires counting the spaghetti on the wall and the spaghetti on the floor.
Arguably, the conditions that led some to conclude that half of published research is false in the early 2000s are more acute today. The spaghetti-throwing machines are too powerful. We can tell they are because, in the textbook case of genomics, a statistical guardrail was put in place so that researchers could throw no more than a specific number of spaghetti strands at the wall, and that number was 1 million.
What happens if you throw 1 million-and-one strands, or 2 million, or 20 million? The answer is that you can get published anyway. From there, The New York Times.
The study in question is one of many, but it’s particularly illustrative because it combines every statistical dimension that can be abused: a giant search space of what to test, a great number of tests actually performed, and tiny associations treated as important.
The spaghetti tower
Schwaba et al. has 138 authors and 16 categories of analysis. All of the analyses have some level of undisclosed flexibility and we are left to imagine how many more numbers flashed across all those screens. How much spaghetti is on the floor is hard to detect. Luckily, there’s a limit to how obvious you can be about these things and this paper barely waves from the window as it shoots through that limit.
The authors tacitly admitted they were testing beyond what the publishing threshold balances out, writing in the peer review file that they counted only the 5 analyses they would need to test the Big Five personality traits. “We use a p-value of .01, which corrects within test dividing .05/5 for the Big Five. We did not correct p-values otherwise.” In this analysis alone there were 90 tests, not 5. The full database, called Add Health, has hundreds more variables available that could have been tested too. Across the laptops and high-performance computing clusters of 138 authors, untold numbers of tests could have been run, and any novel, coherent finding could have been its own paper with failed associations dropped.
The situation with genomics today is not, as some in the community believe, that the low threshold covers anything researchers would like to do. The threshold is static, the 1 million tests, while the number of tests run is flexible.
This may explain why the paper highlights seemingly disparate associations like the one between neurotic genomes and “having pulled a gun or knife on someone,” and a less “attractive personality,” but not with the many, many variables we would also expect to be associated with such neurotic tendencies.
We can see some lack of coherence even in the variables the authors admitted to testing, but didn’t discuss: Neuroticism is associated with the participant’s father having been imprisoned, and pulling a gun or knife on someone, but not with lying to one’s parents, drunk driving, arrest, or the participant being imprisoned themselves. Neuroticism is associated with having been suspended, but not expelled from school.

The “1,000 variants linked to personality” reported by The Times is almost certainly too strong. And there’s another problem that the threshold didn’t fix: small effects. In one of the more extreme findings, people with a moderately high genetic score for neuroticism among at-risk participants still only had about 7 percentage points higher chance of having an imprisoned father than those with moderately low neurotic scores. Other effects are much smaller. The sample size meant that an association could make it into the headline even if the difference was minuscule, less than half a percentile.
The larger point
As it was when the math behind irreplicable “candidate gene” studies was described in the early 2000s, there’s a larger point than candidate genes, or even genomics. As capacity grew, more and more tests became possible, which begs for discipline that can be limiting to one’s career and to the careers of one’s collaborators. Perhaps yet more discipline is required when 138 researchers have to agree to limit their chances of getting into The New York Times, or when a single peer reviewer has to tell them they miscounted.
The success of dividing by 1 million, called the Genome-Wide Significance Threshold, prevented many candidate genes being erroneously associated with disease, and many, many more genes being erroneously associated as computing power and throughput rose. But it’s not magic. It was, in a way, a band-aid. The threshold could be considered one era behind even when it was implemented. Worse, the search space has expanded, making the odds that you’ve truly found something vanishingly small. The “number of gene variants” that the 1 million threshold acknowledges has become the “number of sets of variants” more recently, a combinatorial explosion.
The forefront of science is unpredictable. One of the signs of poor science is that it’s predictable. A result cannot be statistically surprising and unsurprising at the same time. Dalton Conley, cofounder of sociogenomics who was interviewed for the article, implied that discoveries like this can be both predictable and fruitful. “The next study that has double the sample size will find 3,000 variants, no surprise…” He can be sure of that not because he knows there are 2,000 more real and important variants to find this way, but because he can do the math. The authors probably could too. As it was before the first replication crisis, a discovery claim is guaranteed with a large enough sample size, enough variables to test, and ways to test them. Historically, genomics is just the field that makes it obvious.


This research was inspired by the xkcd "paper ' jelly beans cause acne.