
Refine 5 beats GPT-6 Astra and Claude Fable 5.1 at paper review
On 108 papers in ten fields, from macroeconomics to genomics and statistical physics, Refine 5 won 197 of 215 matches.
Our latest reviewer, Refine 5, is now live on refine.ink, available to all users at the standard price for a limited time. Refine uses frontier AI models to perform substantive technical diligence on research documents and other technical work.
Frontier models are improving quickly, and they can now create agentic workflows suited to a variety of tasks, reviewing included. So it's reasonable to ask whether prompting one of these models for a review gives good results. To answer the question, we benchmarked Refine 5 against that approach. We used a strong single-shot prompt on the very strongest publicly available models, GPT-6 Astra (OpenAI) and Claude Fable 5.1 (Anthropic). Both of these were run at xhigh reasoning, on a sample of 108 papers across ten research areas.
96 of 108 against GPT-6 Astra; 101 of 107 against Claude Fable 5.1.

The ten research areas in the benchmark, with Refine 5's wins against both models combined.
Summary
The benchmark covers 108 papers, including 72 papers in six STEM fields:
- Clinical epidemiology and biostatistics
- Computational genomics
- Control theory and dynamical systems
- Environmental and ecological modelling
- Machine learning and statistical learning
- Statistical physics and soft matter
We also used 36 economics papers in four areas (macro, econometrics, applied micro and theory), since many of Refine's most active users are in these fields.
In each match, Refine 5's review of a paper is assessed against one model's single-shot review of the same paper. We use a panel of frontier-model judges to decide which review is likely to be more useful to the paper's author. The judges follow a procedure, described below, designed to reduce bias and make the comparison as objective as possible.
The bottom line: Against GPT-6 Astra, Refine 5 won 96 of 108 matches (88.9%), with 10 losses and 2 ties. Against Claude Fable 5.1, it won 101 of 1071 (94.4%), with 5 losses and 1 tie.2 Pooled, that is 197 of 215 (91.6%).
The advantage is broad: Refine 5's performance looks similar across all ten areas.
Matches are decided by a procedure we developed for the benchmark we published in June. The appendix documents the papers, the baselines, the procedure and the full results.
The baselines: the strongest publicly available models, at xhigh reasoning
GPT-6 Astra and Claude Fable 5.1 are the very strongest publicly available models, and we ran both with reasoning effort set at xhigh (extra-high). Each got one pass per paper. The prompt, reproduced in full in the appendix, asks the model to act as an expert referee for a selective journal in the paper's field. It requires an itemized list of substantive problems, each tied to a specific place in the manuscript, with a proposed fix for each (if possible).
Results against each model
Every match pits Refine 5's review of a paper against a frontier model's single-shot review of that same paper.

Counts are win–loss–tie from Refine 5's side: 108 matches against GPT-6 Astra, 107 against Claude Fable 5.1.
Refine 5 won 96 of its 108 matches against GPT-6 Astra (88.9%), losing 10 and tying 2. Against Claude Fable 5.1 it did better still: 101 of 107 (94.4%), with 5 losses and 1 tie. (Separate results, not published in this blog post, show that the Astra reviews tend to beat Fable reviews quite decisively.)
Each judge gives a match a score between 0 and 1, with 1 meaning Refine 5's review is better, and we average the judges' scores. Above 0.5 counts as a Refine 5 win, exactly 0.5 as a tie, and anything lower as a loss.3 We keep ties in the denominator, so the percentages count wins out of every match played.
The judging procedure works as follows: Judges see each review's paper-grounded, substantive criticisms, minus any spurious concerns, and also excluding the points both reviews made. We first review the substantive results and then give more detail on judging below.
The advantage holds across ten research areas
The ten areas are four in economics, with 9 papers each, and six STEM fields, with 12 papers each. The STEM papers come from six corpora of 50 public papers each, half published and half preprint-only, dated 2022 to 2026. We drew papers from prestigious journals and canonical preprint servers in each field.

Against both models, Refine 5's performance looks broadly similar across the ten areas, and it won every match in several of them. In the four economics areas it won 30 of 36 matches against GPT-6 Astra and 31 of 36 against Claude Fable 5.1; in the six STEM fields, 66 of 72 and 70 of 71. With nine to twelve papers per area, we do not test for differences between areas.
An example from statistical physics
A 2026 paper in Physical Review E asks whether liquid copper and aluminium retain some crystal-like order, and answers by comparing predicted heat capacities with measured ones. The predictions rely on one number in particular: the liquid's thermal-expansion coefficient, or how much its volume grows per degree of heating, since a measured heat capacity includes the energy spent on expanding. Refine 5 recomputed that coefficient from the paper's own formula and inputs and got about 4.2×10⁻⁵ per kelvin for copper at 1350 K. The paper's table says 1.02×10⁻⁴, roughly 2.4 times as much. A second quantity built from the same inputs matches the table to within 1 percent, so the problem sits in this one step. GPT-6 Astra's review talked about how the coefficient is derived, but it never recomputed the tabulated values or spotted the gap. The judges cited Refine 5's catch in every verdict on this match.
The paper calculates the heat capacity of copper separately under three models of the liquid (partial local crystal-like order, long-range crystalline order, and no crystalline order at all) and finds the first fits the data best, within about 2 percent from 2500 K up. The authors rely on that fit twice: as evidence for partial order, and as a check on their method before they turn to electrical resistivity, the paper's main subject. When we correct the paper's apparent error, every prediction drops by 14 to 15 percent. The partial-order model then comes in about 16 percent below the measurements above 2500 K, and the long-range-order model fits best instead, at every temperature in the paper's table. This may change the paper's conclusion, and is certainly something the authors would want to look at.
How a match is decided
We now give a bit more detail on the judging. Judging is not a totally straightforward matter. If we simply ask a model which of two referee reports is better, we run into some well-known biases. Judges tend to prefer the longer report, or the one with more criticisms, whether or not the criticisms hold up.
Each match follows a procedure that boils down to three steps:
- Split each review into concerns. Each review is split into separate concerns, and each concern is tied to a specific passage, equation, table or figure. Concerns with no anchor in the paper get dropped.
- Remove shared points. When both reviews make the same point, the two versions are matched and taken out. The comments that remain reflect differences between the reviews, excluding their overlap.
- Compare what remains. A panel of frontier models (Gemini 3.1 Pro and GPT-5.6 Sol among them, all at xhigh reasoning) reads the paper (figures included) and checks every remaining concern against it. Each judge is asked to assess which review's distinctive concerns would do more to help the author. Verifiable errors count most: a broken step in a derivation or a table inconsistent with the text. Mistaken concerns count against a review. Concerns are weighted by importance, so that one real problem with a main result is more valuable than a long list of minor quibbles. If neither review has a clear edge, the judge declares the match tied. We average the judges' scores to decide the match.
The appendix describes each step in detail.
Try Refine 5
Refine 5 won 197 of 215 matches against the very strongest publicly available models, run at xhigh reasoning, on papers across ten research areas in economics and STEM. It is available to all users at the standard price for a limited time. Upload a paper and read what it finds.
Footnotes
-
One environmental-modelling match against Claude Fable 5.1 did not finish judging within the time limit and is not counted, so that area is scored out of 11 matches against that model. ↩
-
In four matches against Claude Fable 5.1, three in computational genomics and one in machine learning, the model failed to return a review. We score these as Refine 5 wins. ↩
-
LLM judges can lean toward whichever review they read first. In our runs, when judges saw the same pair of reviews in both orders, they picked the same review 91–95% of the time, so the order of presentation barely moves these results. ↩
More posts

A Retrospective Math Benchmark
Mathematician Daniel Litt ran 20 of his papers through Refine. Of the 278 comments it generated, 97.8% pointed to a real issue that coauthors, referees, and editors had missed.

Refine partners with the American Economic Association and the Econometric Society
Two leading publishers in economics now use Refine's AI-assisted technical verification as part of their publication processes.

AI Makes Generating Analysis So Easy That Knowing What to Trust Is Now the Hard Part
AI made rigorous-looking analysis cheap to produce. A viral math claim, disproved within a day, shows why verification can't depend on the right expert happening to look at the right moment.
