In preparation2026

Scorer-Induced Partial Identification of Benchmark Comparisons: An Audit-Design Approach

Rajveer Singh Pall

Once a scorer’s error rate is audited rather than assumed, audit uncertainty outweighs sampling uncertainty in all 120 well-defined comparisons.

The discovery in one figure

WHAT SETS THE WIDTH OF A BENCHMARK COMPARISON?Sampling widthIdentification width, from the audited scorer6.10×For the median pair among the 120 comparable model pairs on MATH-Hard.EVERY ONE OF THE 120 RATIOS IS ABOVE 1×equal widthsmin 2.15×5.74× self-consistency checkmedian 6.10×The harness’s two live scorers disagree on 3.63% of 22,508 responses (95% CI 3.39% to 3.88%).

Swipe to see the whole figure

What is not known about the scorer, not sampling error, sets the width of a benchmark comparison.

The paper in five minutes

An automated verifier grades every free-text answer on a math benchmark, and it can disagree with a careful human reader. Using a real human audit of that verifier, this work derives the range of true accuracies consistent with a reported score. It shows that how little we know about the verifier’s own error matters more than the sampling error that leaderboards usually report.

The research question

Given a benchmark’s observed accuracy and an audited estimate of its scorer’s false-credit and false-miss rates, what can still be concluded about a comparison between two models?

How it works

Classical misclassification bounds are adapted to benchmark comparison and applied to MATH-Hard, using a human audit of the standard boxed-answer comparator (400 items submitted, 350 usable across 27 models). The work then formalises how further audit budget should be allocated and tests the textbook allocation rule on the real, sparse audit data.

  1. 01
    Audit the scorer400 items submitted, 350 usable across 27 models
  2. 02
    Bound true accuracyclassical misclassification bounds adapted to benchmarks
  3. 03
    Compare widthsidentification width against sampling width, pair by pair
  4. 04
    Design the next auditwhere extra human labels help most

Experimental results

The median ratio of identification width to sampling width is 6.10×, dominant in all 120 comparable pairs; a self-consistency check gives 5.74×, and the smallest ratio is 2.15×. The harness’s two live scorers disagree on 3.63% of responses (95% CI 3.39% to 3.88%, n = 22,508). The textbook Neyman allocation of audit effort, in its plug-in form, is dominated by uniform allocation on this sparse real data.

Identification width ÷ sampling width
Median ratio
6.10×
Self-consistency check
5.74×
Smallest ratio
2.15×

Across 120 well-defined comparisons among 17 qualifying models. Every ratio is above 1.

  • 6.10×median identification-to-sampling width ratio
  • 120 / 120comparable pairs where audit uncertainty dominates
  • 3.63%responses where two live scorers disagree

Stated honestlyOne analysis waits on a pre-registered 458-row human audit of the harness’s second scorer; the manuscript marks every number that depends on it as pending instead of estimating it.

What this changes

Before arguing that one model beats another, a benchmark should measure its own scorer. The work documents a failure mode of standard audit design and tests a correction.

Resources

  • Benchmark Evaluation
  • NLP · LLMs

Citation

Manuscript in preparation. Reach me at rajveerpall04@gmail.com.