
3 min de lectura
A Retrospective Math Benchmark
Mathematician Daniel Litt ran 20 of his papers through Refine. Of the 278 comments it generated, 97.8% pointed to a real issue that coauthors, referees, and editors had missed.

Mathematician Daniel Litt ran 20 of his papers through Refine. Of the 278 comments it generated, 97.8% pointed to a real issue that coauthors, referees, and editors had missed.

AI made rigorous-looking analysis cheap to produce. A viral math claim, disproved within a day, shows why verification can't depend on the right expert happening to look at the right moment.

Refine won 90.4% of 1,349 head-to-head matches against single-shot LLM reviewers and scaffolded review systems on 150 economics preprints.