
A Retrospective Math Benchmark
Mathematician Daniel Litt ran 20 of his papers through Refine. Of the 278 comments it generated, 97.8% pointed to a real issue that coauthors, referees, and editors had missed.
As a mathematician who earned his PhD from Stanford and is a professor at the University of Toronto, Daniel Litt knows something most people do not: it's scary to put your past work under a microscope.
At the frontier of math and science, virtually every research article contains mistakes, and until the mistakes are flagged, we cannot know whether they are small or big. Some could unravel a piece of work a researcher had been proud of.
Litt was brave enough to do the retrospective audit anyway. He fed 20 of his papers into Refine — all but one already published, the last accepted for publication — which generated 278 comments.
272 of 278 comments across 20 papers, per Litt's own audit.
The verdict? For Refine, pretty good: nearly all of the comments pointed to legitimate issues and areas worth reviewing. According to Litt's audit:
Of the 278 detailed comments, 248 were correct, 24 were partially correct, and 6 were incorrect. Thus 272 comments (97.8%) identified a real issue, although the proposed explanation or repair sometimes needed revision.
What the comments found
Most of the issues Refine turned up were small and easy to fix — 217 were correctable errors confined to a statement, proof step, formula, citation, or hypothesis. A few were bigger: two substantial theorem-preserving defects and five cases requiring a technical correction to a main result.

Litt's summary of the impact of corrections on his research papers.
Litt's evaluation showcases a key strength of Refine. The papers he used had already been through peer review. Coauthors, referees, and journal editors had gone over every one of them. And yet Refine returned a significant volume of comments with a 98% accuracy rate.
Litt also had constructive criticism for Refine's proposed corrections, noting that they can veer toward slop or be overzealous — often suggesting a paragraph's worth of correction when changing a phrase would do. In this audit, Refine proved better at finding errata-worthy issues than at correcting them: the humans really are the subject-matter experts, and fixing what it catches is where they should stay in the loop.
The verdict on Litt's work
Reassuring: the most significant issues required technical adjustments to statements of results, but in no case was an important message of a paper overturned.
In our view, science proceeds better when distinguished scientists put their work through the wringer and are rewarded with more confidence in that body of work.
Check out Litt's website, where he posted the reviews, his audits, and auto-generated errata. Then try the same experiment on your own work:
About the author

Co-founder
Co-founder of Refine and professor of economics at Northwestern University.
More posts

Refine partners with the American Economic Association and the Econometric Society
Two leading publishers in economics now use Refine's AI-assisted technical verification as part of their publication processes.

AI Makes Generating Analysis So Easy That Knowing What to Trust Is Now the Hard Part
AI made rigorous-looking analysis cheap to produce. A viral math claim, disproved within a day, shows why verification can't depend on the right expert happening to look at the right moment.

Your Document Doesn't Exist in a Vacuum. Now Your Refine Report Doesn't, Either.
With Context, attach referee reports, response letters, data outputs, and prior analyses to any Refine review — and have your document checked against everything it depends on, up to 100,000 words.