On the occasion of AI overtaking humans in math, and those humans admitting it, it may be useful to reflect on what else that means. It means AI has already overtaken us in the simpler math found in most research papers and it means the public has access to how poor that math often is. It’s now conceivable that an ordinary person could debate the greatest scientists and it could cost those scientists their reputation.
In a world where reputation is not mechanized and mechanically inflated, this wouldn’t matter too much. Maybe AI would bruise a few egos and life would continue as usual. But we don’t live in that world. We live in a world where reputation can be hacked.
Reputation in science is thought of as reputation amongst peers transferred to the masses, but direct public scrutiny can be decisive. A claim that ESP is real helped touch off the replication crisis, not because other researchers doubted it — they privately doubt a lot of papers — but because the public did too. The methods used in that paper were routine and still are.

So it may be no coincidence that what the public doesn’t understand, research has done little to reform. That would be bad enough, but consider how easy it would have been to prepare for the day that the public has a PhD in their pocket. In the case of the most prevalent and least understood problem, a disclaimer would have cost nothing but a little embarrassment: “This paper’s result is contingent on a p-value. P-values should not be interpreted as the probability that any hypothesis is true. Additionally, without a detailed preregistration, p-values may have been selected for publishability rather than plausibility.” An approach with a similar goal was proposed by prominent reformers in 2012, the “21 word solution.” It went nowhere.
The last possible moment
Some mathematicians have been forthright, saying AI’s superiority in math has become “impossible to deny,” that early resistance was wishful thinking, and that AI has “gone from mediocre PhD student to very good PhD student to experienced mathematician level to top mathematician level.”

Almost all at once, anyone can use AI for simple validation of a scientific news article or paper. Much of the time, simple checks are all it takes to push papers over. Alzheimer’s was pushed over by duplicate images. Social science can be pushed over just by running the code and comparing the output to the claim in the paper (25% of it). Many papers can be invalidated by running two scripts that were available pre-AI. Many can be shown to be improbable before the experiment takes place. The list goes on, from citation errors to plagiarism (18%) to outright fraud.
In 2024, a paper claiming black spatulas cause cancer was debunked by a human. Then someone decided to see if AI could have found the same error. It did, and that was back when LLMs couldn’t do math well.
Of course, armchair debunking is not all that is going to happen. Researchers will run papers through AI themselves. They will get help writing papers and peer reviewing them. Many papers will be inoculated to start with. On the other hand, more papers will be written as the value of each paper drops.
With any luck, though, we’ll get everything scientific reform has been asking for since 2011. At the last possible moment, but nonetheless.
If the AI then kills us, it really will be the absolute last possible moment. Those of us wearing convincing paperclip costumes will be left to write this down as the quintessential case of procrastination.
