Refine 5 benchmark
Appendix
Supporting material for Refine 5 beats GPT-6 Astra and Claude Fable 5.1 at paper review across ten research fields.
This appendix documents the benchmark behind the post: the papers, the configuration of the two baseline reviewers and the referee prompt given to them, the match procedure, the judging rule and the full results by research area.
Contents
- Papers
- Baseline configuration
- Refine 5 in the benchmark
- Match procedure
- Classification terms
- Judging and scoring
- Each judge alone
- Full results
- The example
Papers
The benchmark uses 108 papers in ten research areas.
- Economics, 36 papers: 9 each in macroeconomics, econometrics, applied microeconomics and economic theory.
- STEM, 72 papers: 12 each from six fields: clinical epidemiology and biostatistics; computational genomics and quantitative molecular biology; control theory and dynamical systems engineering; environmental and ecological modelling; machine learning and statistical learning; and statistical physics and soft condensed matter.
The STEM papers were drawn from six corpora assembled for Refine's benchmarks, one per field. Each corpus contains 50 public papers, 25 published and 25 preprint-only, sampled with quotas on the prominence of the authoring groups (17 papers from prominent groups, 17 from mid-tier groups and 16 from ordinary groups), dated 2022 to 2026. From each corpus the benchmark uses the 12 longest papers by word count extracted from the PDF, capped at 50,000 words. The set was fixed before any review ran.
Baseline configuration
The two baselines are GPT-6 Astra (OpenAI) and Claude Fable 5.1 (Anthropic), the very strongest publicly available models, each run at xhigh reasoning effort. Each model was given the full paper: the PDF together with a markdown conversion of its text, or, for the three largest PDFs, which exceeded the 32 MB request limit, the text alone. Each model reviewed each paper in a single pass with the prompt below and returned one itemized referee report, which entered the match as written.
The prompt is shown verbatim. It does not specify a field.
You are an expert referee for a selective journal in this field. You have been assigned to suggest improvements to the attached manuscript. Produce an itemized report that matches the rigor of a careful reviewer: thorough and fair, specific rather than generic, and constructive about how the paper could be improved. The goal is to flag all significant issues that would require correction if published as is.
## Given
- You are given a high-quality conversion of the paper in markdown format, a pdf of the paper, and any supplementary materials that are relevant in pdf form.
## Stance
- Assume the authors are competent and assume good faith/read their words as a reasonable reader would.
- Focus on substantive flaws that compromise correctness, clarity, or both, rather than stylistic choices.
- Anchor every critique to a specific paragraph, equation, table, or figure.
- Do not pad with summary. Summarize only as much as is needed to ground a critique.
- Where you identify a problem, propose a specific potential revision that would address it.
## Format
- Return your feedback as an itemized list, each item tied to a specific place in the document.
- Aim to prioritize issues, from more significant and load-bearing to less
Refine 5 in the benchmark
Refine 5 received each paper through Refine's standard upload, the way a user submits it, together with any supplementary PDF. Its review of each paper entered the match in full, and the extraction stage reduced it to atomic concerns exactly as it did the baseline reviews.
Match procedure
Each match pairs Refine 5's review of a paper with one baseline's review of the same paper and runs the same eight stages as the June benchmark. The unit throughout is the paper-grounded atomic concern: a single, self-contained issue a review raises, recorded together with the place in the paper it points to. The stages are:
- Extract. Convert each free-text review into atomic concerns and record the paper feature, if any, that each concern points to.
- Classify. Label each concern by where it can be evaluated, how much it matters, whether the author can act on it, and whether outside factual knowledge is required.
- Anchor-check. Check that the concern points to something in the paper, without judging whether the critique is correct.
- Align. Match concerns the two reviews share, by content rather than wording.
- Harmonize. Equalize the significance of matched concerns when two reviewers raised the same issue with different urgency.
- Diff. Remove high-confidence overlaps and out-of-scope positioning comments, leaving each side's residual list.
- Rank. Order the residual concerns within priority buckets so the judge sees the most useful catches first.
- Judge. Ask a panel of frontier models to compare the anchored, substantive residual concerns.
Six of the eight stages are model-assisted; Harmonize and Diff are deterministic. The judges see only concerns anchored to the paper and classified as threatening a main result or as substantive and local; cosmetic points, out-of-scope positioning comments and high-confidence shared concerns are removed first.
Classification terms
The classification stage labels concerns along several axes. In reader-facing terms:
- Where the issue can be evaluated: some concerns can be checked from the paper alone; others require outside literature, field positioning, or generic review norms.
- How much the issue matters: some concerns threaten a central result or identification step; some would improve the manuscript locally; some are only typography, formatting, or prose polish.
- Whether the author can act on it: some concerns specify a concrete fix or check; others raise an issue without enough remediation detail.
- Whether outside facts are needed: some concerns hinge on institutional, empirical, or historical facts not verifiable from the paper alone.
- Whether the target exists: the anchor check asks only whether the cited equation, claim, table, proof step, or structural gap is really present in the paper. A critique can point to a real target and still be wrong; correctness is left to the final judge.
For example, "Equation 12 drops a discount factor used in the previous display" is internal to the paper, concrete, actionable, and checkable against the page. "The contribution is incremental relative to a nearby literature" requires outside comparison. "More robustness checks would help" may be useful but is generic unless it identifies the specific claim, variable, or design choice at issue.
Judging and scoring
The judges are frontier models, among them GLM-5.3 Flash, Gemini 3.1 Pro and GPT-5.6 Sol, run at xhigh reasoning effort. Alongside the paper's text, the judges receive the paper's figures, extracted from the PDF and labelled with page and caption. The judge prompt tells them to consult a figure only to test a concern already on the table, never to originate one.
Each judge's verdict is scored between 0 and 1 from Refine 5's side: 1 is a preference for Refine 5's residual list, 0 a preference for the baseline's, and 0.5 no preference. The panel score is the mean of the judges' scores, so a score of 1 means every judge preferred Refine 5. A panel score above 0.5 is a Refine 5 win, exactly 0.5 a tie, and below 0.5 a loss. Ties stay in the denominator of every win rate in the post.
Each judge alone
We also scored the 72 STEM matches against GPT-6 Astra with each judge alone, using that judge's score in place of the panel mean. Win–loss–tie counts are from Refine 5's side.
| Scoring | Refine 5 W-L-T, 72 matches |
|---|---|
| Full panel | 66-6-0 |
| Gemini 3.1 Pro alone | 65-6-1 |
| GPT-5.6 Sol alone | 62-5-5 |
| GLM-5.3 Flash alone | 64-3-5 |
Each judge alone reaches the same conclusion as the panel.
Full results
Win–loss–tie counts are from Refine 5's side; n is the number of matches in the cell.
| Research area | vs GPT-6 Astra W-L-T |
n | vs Claude Fable 5.1 W-L-T |
n |
|---|---|---|---|---|
| Economics: macro | 7-2-0 | 9 | 7-2-0 | 9 |
| Economics: econometrics | 9-0-0 | 9 | 9-0-0 | 9 |
| Economics: applied micro | 7-1-1 | 9 | 8-1-0 | 9 |
| Economics: theory | 7-1-1 | 9 | 7-1-1 | 9 |
| Economics subtotal | 30-4-2 | 36 | 31-4-1 | 36 |
| Clinical epidemiology & biostatistics | 12-0-0 | 12 | 12-0-0 | 12 |
| Computational genomics | 12-0-0 | 12 | 12-0-0a | 12 |
| Control theory & dynamical systems | 12-0-0 | 12 | 12-0-0 | 12 |
| Environmental & ecological modelling | 12-0-0 | 12 | 11-0-0b | 11 |
| Machine learning & statistical learning | 10-2-0 | 12 | 12-0-0a | 12 |
| Statistical physics & soft matter | 8-4-0 | 12 | 11-1-0 | 12 |
| STEM subtotal | 66-6-0 | 72 | 70-1-0 | 71 |
| All areas | 96-10-2 | 108 | 101-5-1 | 107 |
a Claude Fable 5.1 failed to return a review in these matches (three in computational genomics, one in machine learning and statistical learning); they are scored as Refine 5 wins. b One match did not finish judging within the time limit and is not counted.
Pooled over both baselines, Refine 5 won 197 of 215 matches (91.6%), with 15 losses and 3 ties.
The example
The example in the post is Inferring partial crystalline order in liquids from electrical resistivity, Physical Review E, 2026, one of the twelve statistical physics and soft condensed matter papers. Refine 5's concern, as extracted from its review:
The thermal-expansion coefficients reported in Tables II and V appear not to follow from the displayed formulas. For copper at 1350 K, the printed inputs give approximately C̄V=3.22, μ=95.8, and γG=1.70, hence αVQH≈4.2×10−5 K−1; the stated quasiparticle transformation lowers this slightly, rather than producing the tabulated 1.02×10−4 K−1.
GPT-6 Astra's review did not raise the tabulated coefficients. Every judge verdict on the match cited this concern as pivotal, and Refine 5 won the match with a panel score of 1.