
AI Makes Generating Analysis So Easy That Knowing What to Trust Is Now the Hard Part
AI made rigorous-looking analysis cheap to produce. A viral math claim, disproved within a day, shows why verification can't depend on the right expert happening to look at the right moment.
AI has made it cheap and quick to produce analysis that appears rigorous, whether that work is a data analysis, a research memo, or a claimed scientific breakthrough. The hard part has shifted from the generation of the work to assessing whether the finished work is actually correct and reliable.
Bad research and overclaiming have been around for as long as research itself. What has changed is the volume and the speed. Claims that once took weeks to assemble can now be generated by the dozens in days.
A small but telling example
Recently, Greg Brockman, president of OpenAI, shared a striking claim on X: a paper used GPT-5.6 Pro to disprove a mathematical conjecture that had stood for decades. He called it evidence of "a renaissance of scientific, medical, and mathematical discovery."
Within a day of this claim being shared, the mathematician Tony Feng found a gap that undercut the paper's central argument. The author conceded that the claim was wrong.
The claimant (who had already made and withdrawn one earlier big claim) was working outside his usual area of math. Brockman is not an expert in the topic either, so neither person was in a position to check what they were sharing. The false claim was persuasive (and complex) enough that it was easily amplified and spread widely. And nobody really noticed that it was wrong until Feng, who has the relevant experience and interest, happened to check it.
The field owes Feng thanks, but relying on heroes to catch errors one at a time doesn't scale with the coming volume of AI-generated research. A system that relies on the serendipity of the right expert noticing the right claim at exactly the right moment will be overwhelmed.
Three implications
AI is already better than humans at catching many kinds of errors, such as inconsistent logic or a claim that doesn't follow from its assumptions. But most researchers have not paired generation with an equally good process for verification, so they are far from taking advantage of these capabilities, and we see errors like this one.
Verification cannot depend on home-grown processes. Researchers can and should build verification into their own workflows. And readers should be skeptical. But verification is too important to be left to individual researchers. For example, the researcher who made the false claim in this case had done his best to verify his work with frontier models.
The checking step has to live in shared tools and norms. Science already did something like this once when it slowly built peer review and replication into everyday practice over many decades. This time the same kind of building has to happen, faster.
The public conversation matters just as much as the research world does. Some people who make big claims want them to survive scrutiny, while others mainly want them to travel far and win attention. Mathematical and scientific discourse is already paying the price for that second group, with grand claims being made, attracting attention, and then falling apart.
What has to change
Journals need to rethink peer review in a world where anyone can generate a convincing research paper in minutes. But this isn't just an academic problem. Investment research, legal analysis, policy work, and many other fields now depend on documents that are far easier to generate than to verify.
Terence Tao, in a talk about math in the age of AI at this year's International Congress of Mathematicians, said that we've entered a turbulent period where there's essentially a crisis in the foundations of how science checks its own work. He also offered a hopeful thought: once a field thoroughly examines and codifies those foundations, its community will emerge stronger and more resilient than before.
The stakes are high in both directions, because good verification practices, used well, will create rails that direct AI energy toward immense progress, which in turn will move research faster than ever. If verification falls behind generation, that same speed will mostly produce noise, in the form of retracted claims, wasted attention, and a public with good reason to trust research less.
Where we come in
This is the problem we started Refine to solve. We built an independent verification layer for research documents. Refine checks documents for accuracy, reasoning errors, and internal inconsistencies. Catching a mistake shouldn't depend on the chance that one person happens to notice it at the right moment.
Have a document you'd like us to review?
Sobre el autor

Co-founder
Co-founder of Refine and professor of economics at Northwestern University.
Más artículos

Your Document Doesn't Exist in a Vacuum. Now Your Refine Report Doesn't, Either.
With Context, attach referee reports, response letters, data outputs, and prior analyses to any Refine review — and have your document checked against everything it depends on, up to 100,000 words.

Refine has completed its SOC 2 Type 1 examination
Refine has completed its SOC 2® Type 1 examination and received its first SOC 2 report — an important step in our ongoing commitment to protecting researchers and their work.

A Structured Benchmark for AI Paper Review
Refine won 90.4% of 1,349 head-to-head matches against single-shot LLM reviewers and scaffolded review systems on 150 economics preprints.