Scorer-Induced Partial Identification of Benchmark Comparisons: An Audit-Design Approach
Once a scorer’s error rate is audited rather than assumed, audit uncertainty outweighs sampling uncertainty in all 120 well-defined comparisons.
The discovery in one figure
Swipe to see the whole figure
The paper in five minutes
An automated verifier grades every free-text answer on a math benchmark, and it can disagree with a careful human reader. Using a real human audit of that verifier, this work derives the range of true accuracies consistent with a reported score. It shows that how little we know about the verifier’s own error matters more than the sampling error that leaderboards usually report.
The research question
Given a benchmark’s observed accuracy and an audited estimate of its scorer’s false-credit and false-miss rates, what can still be concluded about a comparison between two models?
How it works
Classical misclassification bounds are adapted to benchmark comparison and applied to MATH-Hard, using a human audit of the standard boxed-answer comparator (400 items submitted, 350 usable across 27 models). The work then formalises how further audit budget should be allocated and tests the textbook allocation rule on the real, sparse audit data.
- 01Audit the scorer400 items submitted, 350 usable across 27 models
- 02Bound true accuracyclassical misclassification bounds adapted to benchmarks
- 03Compare widthsidentification width against sampling width, pair by pair
- 04Design the next auditwhere extra human labels help most
Experimental results
The median ratio of identification width to sampling width is 6.10×, dominant in all 120 comparable pairs; a self-consistency check gives 5.74×, and the smallest ratio is 2.15×. The harness’s two live scorers disagree on 3.63% of responses (95% CI 3.39% to 3.88%, n = 22,508). The textbook Neyman allocation of audit effort, in its plug-in form, is dominated by uniform allocation on this sparse real data.
6.10×5.74×2.15×Across 120 well-defined comparisons among 17 qualifying models. Every ratio is above 1.
- 6.10×median identification-to-sampling width ratio
- 120 / 120comparable pairs where audit uncertainty dominates
- 3.63%responses where two live scorers disagree
Stated honestlyOne analysis waits on a pre-registered 458-row human audit of the harness’s second scorer; the manuscript marks every number that depends on it as pending instead of estimating it.
What this changes
Before arguing that one model beats another, a benchmark should measure its own scorer. The work documents a failure mode of standard audit design and tests a correction.
Citation
Manuscript in preparation. Reach me at rajveerpall04@gmail.com.