Refine 5 benchmark

Appendix

Supporting material for Refine 5 beats GPT-6 Astra and Claude Fable 5.1 at paper review across ten research fields.

This appendix documents the benchmark behind the post: the papers, the configuration of the two baseline reviewers and the referee prompt given to them, the match procedure, the judging rule and the full results by research area.

Contents

Papers

The benchmark uses 108 papers in ten research areas.

The STEM papers were drawn from six corpora assembled for Refine's benchmarks, one per field. Each corpus contains 50 public papers, 25 published and 25 preprint-only, sampled with quotas on the prominence of the authoring groups (17 papers from prominent groups, 17 from mid-tier groups and 16 from ordinary groups), dated 2022 to 2026. From each corpus the benchmark uses the 12 longest papers by word count extracted from the PDF, capped at 50,000 words. The set was fixed before any review ran.

Baseline configuration

The two baselines are GPT-6 Astra (OpenAI) and Claude Fable 5.1 (Anthropic), the very strongest publicly available models, each run at xhigh reasoning effort. Each model was given the full paper: the PDF together with a markdown conversion of its text, or, for the three largest PDFs, which exceeded the 32 MB request limit, the text alone. Each model reviewed each paper in a single pass with the prompt below and returned one itemized referee report, which entered the match as written.

The prompt is shown verbatim. It does not specify a field.



You are an expert referee for a selective journal in this field. You have been assigned to suggest improvements to the attached manuscript. Produce an itemized report that matches the rigor of a careful reviewer: thorough and fair, specific rather than generic, and constructive about how the paper could be improved. The goal is to flag all significant issues that would require correction if published as is.

## Given
- You are given a high-quality conversion of the paper in markdown format, a pdf of the paper, and any supplementary materials that are relevant in pdf form.

## Stance
- Assume the authors are competent and assume good faith/read their words as a reasonable reader would.
- Focus on substantive flaws that compromise correctness, clarity, or both, rather than stylistic choices.
- Anchor every critique to a specific paragraph, equation, table, or figure.
- Do not pad with summary. Summarize only as much as is needed to ground a critique.
- Where you identify a problem, propose a specific potential revision that would address it.

## Format
- Return your feedback as an itemized list, each item tied to a specific place in the document.
- Aim to prioritize issues, from more significant and load-bearing to less

Refine 5 in the benchmark

Refine 5 received each paper through Refine's standard upload, the way a user submits it, together with any supplementary PDF. Its review of each paper entered the match in full, and the extraction stage reduced it to atomic concerns exactly as it did the baseline reviews.

Match procedure

Each match pairs Refine 5's review of a paper with one baseline's review of the same paper and runs the same eight stages as the June benchmark. The unit throughout is the paper-grounded atomic concern: a single, self-contained issue a review raises, recorded together with the place in the paper it points to. The stages are:

  1. Extract. Convert each free-text review into atomic concerns and record the paper feature, if any, that each concern points to.
  2. Classify. Label each concern by where it can be evaluated, how much it matters, whether the author can act on it, and whether outside factual knowledge is required.
  3. Anchor-check. Check that the concern points to something in the paper, without judging whether the critique is correct.
  4. Align. Match concerns the two reviews share, by content rather than wording.
  5. Harmonize. Equalize the significance of matched concerns when two reviewers raised the same issue with different urgency.
  6. Diff. Remove high-confidence overlaps and out-of-scope positioning comments, leaving each side's residual list.
  7. Rank. Order the residual concerns within priority buckets so the judge sees the most useful catches first.
  8. Judge. Ask a panel of frontier models to compare the anchored, substantive residual concerns.

Six of the eight stages are model-assisted; Harmonize and Diff are deterministic. The judges see only concerns anchored to the paper and classified as threatening a main result or as substantive and local; cosmetic points, out-of-scope positioning comments and high-confidence shared concerns are removed first.

Classification terms

The classification stage labels concerns along several axes. In reader-facing terms:

For example, "Equation 12 drops a discount factor used in the previous display" is internal to the paper, concrete, actionable, and checkable against the page. "The contribution is incremental relative to a nearby literature" requires outside comparison. "More robustness checks would help" may be useful but is generic unless it identifies the specific claim, variable, or design choice at issue.

Judging and scoring

The judges are frontier models, among them GLM-5.3 Flash, Gemini 3.1 Pro and GPT-5.6 Sol, run at xhigh reasoning effort. Alongside the paper's text, the judges receive the paper's figures, extracted from the PDF and labelled with page and caption. The judge prompt tells them to consult a figure only to test a concern already on the table, never to originate one.

Each judge's verdict is scored between 0 and 1 from Refine 5's side: 1 is a preference for Refine 5's residual list, 0 a preference for the baseline's, and 0.5 no preference. The panel score is the mean of the judges' scores, so a score of 1 means every judge preferred Refine 5. A panel score above 0.5 is a Refine 5 win, exactly 0.5 a tie, and below 0.5 a loss. Ties stay in the denominator of every win rate in the post.

Each judge alone

We also scored the 72 STEM matches against GPT-6 Astra with each judge alone, using that judge's score in place of the panel mean. Win–loss–tie counts are from Refine 5's side.

Scoring Refine 5 W-L-T, 72 matches
Full panel 66-6-0
Gemini 3.1 Pro alone 65-6-1
GPT-5.6 Sol alone 62-5-5
GLM-5.3 Flash alone 64-3-5

Each judge alone reaches the same conclusion as the panel.

Full results

Win–loss–tie counts are from Refine 5's side; n is the number of matches in the cell.

Research area vs GPT-6 Astra
W-L-T
n vs Claude Fable 5.1
W-L-T
n
Economics: macro 7-2-0 9 7-2-0 9
Economics: econometrics 9-0-0 9 9-0-0 9
Economics: applied micro 7-1-1 9 8-1-0 9
Economics: theory 7-1-1 9 7-1-1 9
Economics subtotal 30-4-2 36 31-4-1 36
Clinical epidemiology & biostatistics 12-0-0 12 12-0-0 12
Computational genomics 12-0-0 12 12-0-0a 12
Control theory & dynamical systems 12-0-0 12 12-0-0 12
Environmental & ecological modelling 12-0-0 12 11-0-0b 11
Machine learning & statistical learning 10-2-0 12 12-0-0a 12
Statistical physics & soft matter 8-4-0 12 11-1-0 12
STEM subtotal 66-6-0 72 70-1-0 71
All areas 96-10-2 108 101-5-1 107

a Claude Fable 5.1 failed to return a review in these matches (three in computational genomics, one in machine learning and statistical learning); they are scored as Refine 5 wins. b One match did not finish judging within the time limit and is not counted.

Pooled over both baselines, Refine 5 won 197 of 215 matches (91.6%), with 15 losses and 3 ties.

The example

The example in the post is Inferring partial crystalline order in liquids from electrical resistivity, Physical Review E, 2026, one of the twelve statistical physics and soft condensed matter papers. Refine 5's concern, as extracted from its review:

The thermal-expansion coefficients reported in Tables II and V appear not to follow from the displayed formulas. For copper at 1350 K, the printed inputs give approximately C̄V=3.22, μ=95.8, and γG=1.70, hence αVQH≈4.2×10−5 K−1; the stated quasiparticle transformation lowers this slightly, rather than producing the tabulated 1.02×10−4 K−1.

GPT-6 Astra's review did not raise the tabulated coefficients. Every judge verdict on the match cited this concern as pivotal, and Refine 5 won the match with a panel score of 1.