Skip to main content
Litt's paper review archive showing the overall audit — 278 detailed comments, with validity and impact breakdowns.
3 min read

A Retrospective Math Benchmark

Mathematician Daniel Litt ran 20 of his papers through Refine. Of the 278 comments it generated, 97.8% pointed to a real issue that coauthors, referees, and editors had missed.

As a mathematician who earned his PhD from Stanford and is a professor at the University of Toronto, Daniel Litt knows something most people do not: it's scary to put your past work under a microscope.

At the frontier of math and science, virtually every research article contains mistakes, and until the mistakes are flagged, we cannot know whether they are small or big. Some could unravel a piece of work a researcher had been proud of.

Litt was brave enough to do the retrospective audit anyway. He fed 20 of his papers into Refine — all but one already published, the last accepted for publication — which generated 278 comments.

97.8%
of comments identified a real issue

272 of 278 comments across 20 papers, per Litt's own audit.

The verdict? For Refine, pretty good: nearly all of the comments pointed to legitimate issues and areas worth reviewing. According to Litt's audit:

Of the 278 detailed comments, 248 were correct, 24 were partially correct, and 6 were incorrect. Thus 272 comments (97.8%) identified a real issue, although the proposed explanation or repair sometimes needed revision.

What the comments found

Most of the issues Refine turned up were small and easy to fix — 217 were correctable errors confined to a statement, proof step, formula, citation, or hypothesis. A few were bigger: two substantial theorem-preserving defects and five cases requiring a technical correction to a main result.

Litt's impact-after-audit table: most comments were local correctable errors; none was a fundamental failure

Litt's summary of the impact of corrections on his research papers.

Litt's evaluation showcases a key strength of Refine. The papers he used had already been through peer review. Coauthors, referees, and journal editors had gone over every one of them. And yet Refine returned a significant volume of comments with a 98% accuracy rate.

Litt also had constructive criticism for Refine's proposed corrections, noting that they can veer toward slop or be overzealous — often suggesting a paragraph's worth of correction when changing a phrase would do. In this audit, Refine proved better at finding errata-worthy issues than at correcting them: the humans really are the subject-matter experts, and fixing what it catches is where they should stay in the loop.

The verdict on Litt's work

Reassuring: the most significant issues required technical adjustments to statements of results, but in no case was an important message of a paper overturned.

In our view, science proceeds better when distinguished scientists put their work through the wringer and are rewarded with more confidence in that body of work.

Check out Litt's website, where he posted the reviews, his audits, and auto-generated errata. Then try the same experiment on your own work:

About the author

More posts