Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Sho Kawano∗ Zehang Richard Li
Paul A. Parker
arXiv:2609.20758v1 [stat.ML] 17 Sep 2026
Department of Statistics, University of California, Santa Cruz, CA, USA
September 18, 2026
Abstract Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain’s own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain’s prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage. At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator’s error far more accurately.
Keywords: disaggregated evaluation; prediction-powered inference; small area estimation; cross-validation; survey sampling
1
Introduction
AI systems built on large language models (LLMs) are rapidly being deployed in a wide range of settings, increasing the number of systems and versions to evaluate. A single headline score can mask heterogeneous performance. Evaluators therefore report performance separately for each task domain, product line or user segment the system serves. This practice is called disaggregated evaluation (Barocas et al., 2021). Disaggregated evaluation can be expensive because observing an outcome may require querying the system, grading its response, or both. An evaluation unit could be a single task in an LLM benchmark or a single customer-support conversation with a deployed system. Reporting dimensions (e.g., task category) partition the population into domains. The outcome for each unit is a grade for how well the task was completed or a binary indicator of successful completion. The target is the mean outcome over all units in a given domain. Obtaining an outcome takes two steps: querying, in which the system carries out the task; and grading, in which the attempt is scored. We distinguish two settings by the grading step. ∗
Corresponding author: [email protected]
1
• Verifiable. Each task has a known answer, so grading can be done programmatically. • Open-ended. Correctness cannot be checked against a known answer, so the gold standard often needs to be produced by a human grader. For verifiable tasks, grading is automatic and the cost lies in querying. Evaluating a single LLM on HELM, a major benchmark suite, can cost $9,337 in API credits (Liang et al., 2022). For open-ended tasks, human assessment creates an additional cost. In monitoring the traffic of a deployed AI system, the interactions between human and AI have already occurred, so the remaining cost is only in reviewing and grading them. A benchmark of open-ended tasks, such as replicating a research paper (Starace et al., 2025) or answering a medical question at length (Singhal et al., 2023), requires both tokenintensive querying and expert grading. As the number of units, systems, and model versions grows, exhaustive evaluation becomes impractical (Wu et al., 2026). Evaluation therefore increasingly relies on a sample of units or tasks (Maia Polo et al., 2024; Yauney et al., 2026). In the benchmarking literature, Fogliato et al. (2024b) and Fisch et al. (2024) draw the questions to score at random within strata, and estimate one overall score. Wu et al. (2026) propose an adaptive sampling method under a fixed query budget, paired with a prediction from a Bayesian factor model fit on other LLMs’ historical results. For traffic monitoring, Dobi et al. (2026) estimate the prevalence of policy-violating content by domain from a daily random sample. Although benchmarking and traffic monitoring have developed separately, they both require estimating performance over a larger population using a subset of samples. We formulate this common problem using finite-population survey sampling. We treat the complete evaluation set as a finite population and target the mean outcome within each reporting domain. This perspective also makes the goal of uncertainty quantification concrete: the true domain mean is a fixed population quantity, and a valid interval should cover it at the stated rate over repeated samples, while remaining as narrow as possible. Because this design-based interpretation does not require assumptions on a superpopulation distribution, it stays meaningful under distribution shift, which is common in AI evaluation as systems and their users change (Wu et al., 2026). A common approach in AI evaluation is prediction-powered inference (PPI; Angelopoulos et al., 2023a). Under the finite-population view, PPI is equivalent to the difference estimator of classical survey sampling (Särndal et al., 1992), a correspondence first noted by Fogliato et al. (2024b) and formalized by Mozer (2026). Difference estimators use unit-level auxiliary information, which is data recorded on every population unit that may help predict the outcome. PPI takes the auxiliary as a prediction of the outcome to make the estimate more precise. PPI++ (Angelopoulos et al., 2023b), applied to disaggregated evaluation by Emmenegger et al. (2026), further tunes how much of the prediction to use. Its survey sampling equivalent is the generalized regression (GREG) estimator (Mozer, 2026). The need to estimate many domain means from a limited sample leads naturally to small area estimation (Rao and Molina, 2015). Developed for surveys with large overall samples but few observations within particular domains, small area estimation uses models to borrow strength across domains. Using the terminology from small area estimation, both PPI and GREG estimate a domain mean from that domain’s data alone, and thus fall in a category called ‘direct estimators’. For disaggregated evaluation, the sampling budget is spread across many domains. Thus, direct estimators are imprecise for domains with few labels. We therefore focus on so-called area-level models, which combine the domain direct estimates and their sampling variances, without requiring a model of the underlying unit-level outcomes. The standard area-level model is the Fay–Herriot model (Fay and Herriot, 1979), which shrinks each direct estimate toward a shared regression on domain-level covariates, with greater shrinkage applied to less precise estimates.
2
Recent work in disaggregated AI evaluation has begun to incorporate shrinkage in estimators as well. Empirical Bayes and hierarchical models that shrink subgroup estimates toward a shared fit have been proposed for benchmark task types and for subgroups in algorithmic fairness (Fogliato et al., 2024a; Herlihy et al., 2024; Li and Ignatiadis, 2025). Fogliato et al. (2024a) note that the Fay–Herriot model is a linear special case of their approach. These developments suggest that shrinkage and smoothing can improve domain estimators, but raise more practical questions: which smoothing model to use and how much smoothing should a model introduce? Choosing between model-based estimates requires formal model validation. We can compare direct estimators based on their sampling variance, since they are typically design-unbiased, whereas a modelbased estimator often introduces bias in exchange for lower variance and must be judged on its total error. In AI evaluation, cross-validation has been used to tune an estimator (Herlihy et al., 2024; Emmenegger et al., 2026) rather than to choose among them, and Fogliato et al. (2024a) identify model selection as an open question. In small area estimation, the comparison is often made by an information criterion, which requires a likelihood and cannot score direct estimators. Out-of-sample validation has only recently appeared in small area estimation (Kawano et al., 2026a; Dong and Li, 2026). Specifically, design-based cross-validation (DB-CV; Dong and Li, 2026) is pertinent here as it allows comparisons across direct and model-based estimators. We develop an integrated workflow for estimation and validation in disaggregated AI evaluation through the lens of small area estimation. For estimation, we introduce prediction-powered smoothing (PP-S), a Bayesian Fay–Herriot model applied to the GREG estimate of each domain. Its base form shrinks across all domains, and an extension (PP-TS) borrows strength along a nested hierarchy of the domains. For validation, we propose novel refinements to DB-CV, replacing a conservative bias bound with an approximately unbiased score, for reliable comparison and error reporting. We examine the proposed methods on one instance of each setting: a curated benchmark with verifiable grading, and deployed agent traffic graded by humans. Outcomes are observed for every unit in both datasets, allowing us to assess every estimator and the validation score itself against the oracle. We show that smoothing improves on the direct estimators in both point and interval estimation. We also show that the debiased score matches selection on an independent validation sample while estimating the chosen estimator’s error far more accurately. The remainder of the paper is organized as follows. Section 2 makes the case for a probability sample as the foundation for valid inference. Section 3 sets up the problem and defines the direct estimators, the smoothing models and the validation score. Section 4 studies them on the benchmark and on the traffic data, and Section 5 concludes.
2
The Case for Probability Sampling
In this paper, we assume that the evaluated units are obtained from a probability sample (i.e., every unit’s nonzero probability of being selected is known). In a benchmark, the evaluator chooses which tasks to query and grade and can therefore implement such a design directly. Yet in practice, benchmarks are often scored on a curated subset, with the questions chosen by a model fit to other LLMs’ results and the full score predicted from it (Maia Polo et al., 2024). Zhang et al. (2025) find that none of these methods consistently beats the mean of a random subset once the LLM under evaluation is more accurate than those the predictor was fit on. For ranking LLMs using a benchmark, Yauney et al. (2026) find that consistently ordering two of similar accuracy takes a subset large enough that random sampling is competitive with the selection methods. In deployed traffic, the interactions that reach a human grader are often selected by a complaint, a flag, or a reviewer’s attention, and such a process is usually not a probability sample. Without a probability design, domain estimates can be biased unless
3
Bias
Coverage 1.0
0 −1
0.8
−2
0.6
sample mean PPI
−3 0.4 5
10
20
30
5
10
20
30
sampling budget (%) Figure 1: Mean satisfaction estimated without weights from labels sampled with a bias toward rejected responses (PRISM, 63 domains, 100 replications). The sample bias puts the correlation between being labeled and the outcome at 0.029 in magnitude at the 5% budget. Reported bias is in points of the 1 to 100 satisfaction scale, averaged over domains. Coverage is of nominal 95% intervals. Dotted lines mark zero bias and the nominal level. the selection process is modeled explicitly (Pfohl et al., 2025). However, the modeling assumptions on the selection process cannot be validated from the labeled data (Molenberghs et al., 2008). We illustrate the consequence of the sampling design on PRISM (Kirk et al., 2024), a dataset of conversations with LLMs in which the users themselves graded every response for satisfaction. To mimic a nonprobability sample, we draw samples in which responses that were rejected by the participant are selected at about twice the rate of the others, and then estimate as if the sample had been drawn at random (see Appendix C.1 for the full details). The resulting bias of the sample mean matches that of a typical online opt-in survey panel (Pew Research Center, 2023; Kawano et al., 2026b). Figure 1 compares the estimates with the truth. The moderate bias persists as the sampling budget grows, while the coverage of the 95% intervals falls well below nominal. This is an instance of the Big Data Paradox of Meng (2018): “The bigger the data, the surer we fool ourselves.” PPI does not fix this problem either, as without accounting for the selection process, the residuals from the labeled sample remain biased for the residuals in the full population. Since neither the selection nor its correction can be checked from the labels alone, the rest of the paper treats a probability sample as the foundation. Moreover, we recommend its use for disaggregated evaluation more broadly.
3
Methods
3.1
Direct Estimation with Auxiliary Information
The units making up an evaluation form a finite population U = {1, . . . , N }, such as the questions of an LLM benchmark or the customer-support conversations a deployed system handled over an evaluation window. Reporting dimensions partition U into m domains U1 , . . . , Um with known sizes N1 , . . . , Nm . The outcome label yij is a grade of the evaluated system’s performance on unit j of domain i. The
4
estimand for domain i is the finite-population mean 1 X yij , θi = Ni
i = 1, . . . , m.
j∈Ui
Note that θi is the population mean grade or rating for a given domain Ui . When the given grade is binary, θi is the accuracy in domain i. A sampling budget of n ≪ N labels is spent by drawing a probability sample S ⊂ U, in which every unit has a known probability of being labeled. The sampled units in domain i are Si = S ∩ Ui , of size P ni = |Si |, where n = i ni . Unit j of domain i is sampled with probability πij = P (j ∈ S). The problem is to estimate every θi from the n labels, with an interval whose coverage holds over repeated −1 draws of the sample. The simplest estimator of θi weights each labeled outcome by wij = πij , θ̂iHT =
1 X wij yij . Ni j∈Si
This is the Horvitz–Thompson (HT) estimator (Horvitz and Thompson, 1952), which we mark with the superscript HT and compare the other estimators against. Under simple random sampling (SRS) within a domain, it is the sample mean. The HT estimator is unbiased over repeated samples. It is also design-consistent, i.e., it converges to the true domain mean as labels accumulate, p
θ̂iHT − θi −→ 0
as ni → ∞,
where the probability is over which sample is drawn. Dobi et al. (2026) describe an estimator of this kind deployed in real-world traffic monitoring at Pinterest. Now suppose that every unit includes auxiliary information, recorded on all N units, and write fij for its value on unit j of domain i. Examples are the score of an LLM judge prompted to grade each interaction (Zheng et al., 2023), or a question’s historical pass rate among previously tested models. Prediction-powered inference (PPI; Angelopoulos et al., 2023a) treats the auxiliary as a prediction of the label, 1 X 1 X θ̂iPPI = fij + wij {yij − fij }. (1) Ni Ni j∈Ui
j∈Si
The first term is a known constant, and the second is a correction term made of weighted errors. The weighted errors term is itself an HT estimate of the population error. The correction term removes any systematic error in the auxiliary, so the estimator is design-consistent regardless of fij . The more predictive fij is of the outcome, the more variance this approach can reduce. Rearranging (1) allows another way to understand PPI given by 1 X X 1 θ̂iPPI = θ̂iHT + fij − wij fij , Ni Ni j∈Ui
j∈Si
the HT estimator plus a correction for how far the sampled units sit from the population on the auxiliary’s scale. The generalized regression (GREG) estimator extends this approach and is given by 1 X X 1 θ̂iGREG = θ̂iHT + λ̂ fij − wij fij , Ni Ni j∈Ui
j∈Si
where λ̂ is a tuning parameter which we call the tuned correction slope. It is the coefficient from one weighted least-squares regression of y on f , pooled across domains. With several unit-level covariates, 5
the regression has one coefficient per covariate and λ̂ becomes the vector of those coefficients. Setting λ = 1 recovers θ̂iPPI , and λ = 0 recovers the HT estimator. GREG is also design-consistent (Särndal et al., 1992). The power-tuned PPI (PPI++; Angelopoulos et al., 2023b) takes the same form as GREG but requires the auxiliary to be on the outcome’s scale, whereas GREG fits its prediction by regression and so accepts any unit-level covariate, such as the content of a conversation in traffic. Because of this, we refer to estimators of this type (e.g., PPI++) as GREG. GREG’s linear fit corrects the auxiliary’s level and scale but not nonlinear miscalibration. The correction slope protects GREG from a poor auxiliary, since one with no predictive value drives λ̂ toward zero. In AI evaluation, PPI has been applied with stratified sampling of the units to a single overall score on a benchmark, using an LLM judge as the prediction (Fisch et al., 2024) or predicted correctness based on a classifier (Fogliato et al., 2024b). The prediction has also been fit to other LLMs’ results on the same questions, under random sampling (Zhang et al., 2025) and adaptive sampling (Wu et al., 2026). Emmenegger et al. (2026) apply GREG to disaggregated evaluation.
3.2
Prediction-Powered Smoothing
PPI and GREG improve each domain’s estimates using auxiliary correction, but they do not explicitly model the domain means jointly. Each uses only its own domain’s labels, so where labels are few the estimates remain imprecise, with a variance that scales as n−1 i . Small area estimation addresses this imprecision with area-level models, which borrow information across domains. The inputs into an area-level model are the design-consistent estimates zi of each θi and their variance di , which are typically plugged in and treated as known (Rao and Molina, 2015). The foundational arealevel model is the Fay–Herriot model (Fay and Herriot, 1979), ind
zi | θi ∼ N (θi , di ) , θi = x⊤ i β + vi ,
(2)
iid vi ∼ N 0, σ 2 ,
where xi is a vector of domain-level covariates, β a coefficient vector shared by all domains, and vi a domain random effect. The model rests on two assumptions, both of which matter most in the small domains that borrow the most. The sampling model in (2) is a central limit approximation for the direct estimate, and it is weakest in the smallest domains. The linking model in (2) is a modeling assumption, and misspecifying it loses precision, again in the small domains. Conditional on β and σ 2 , the posterior mean of θi is a precision-weighted compromise between the input and the regression component, EM θi | zi , β, σ 2 = γi zi + (1 − γi ) x⊤ i β,
γi =
σ2 , σ 2 + di
(3)
where EM denotes expectation under the model. When di is large relative to σ 2 , the weight γi is near zero and the estimate is close to the regression component, so a noisy domain is estimated largely by borrowing strength from the other domains, to the extent that the linking model describes them well. When di is small relative to σ 2 , the weight is near one and a precise domain yields a model-based estimate that is close to its direct estimate. Since di vanishes as the domain sample size grows, the estimate converges to zi and is design-consistent whatever the linking model (Rao and Molina, 2015, Section 6.1). Taking the HT estimate as the input, zi = θ̂iHT , gives the classical Fay–Herriot (FH) estimator. To
6
r=1 benchmark
r=2 task type (domain i)
···
BBH
causal judgement
date understanding
sports understanding
··· 20 more
MATH
algebra
number theory
prealgebra
advanced subjects
Figure 2: An example taxonomy from the Open LLM Leaderboard benchmark dataset: 34 task types, the domains at level r = 2, nested in five benchmarks at level r = 1. Two of the five benchmarks are shown. better incorporate auxiliary information, we take the GREG estimate as the input zi , i.e., ind
θ̂iGREG | θi ∼ N (θi , di ) , θi = x⊤ i β + vi ,
iid vi ∼ N 0, σ 2 ,
(4)
where di is now the variance of thehGREG estimator. We call this approach prediction-powered smoothing i GREG PP-S . The same two-stage construction appears in survey (PP-S) and denote θ̂i = EM θi | θ̂ statistics as the smoothed model-assisted estimator of Gao and Wakefield (2024), who smooth a GREG on the logit scale for binary outcomes. Under PP-S, the auxiliary can enter the smoother in two places and act on different terms of the weighted compromise shown in (3). At the unit level, it enters through the GREG estimator. A good unit-level auxiliary reduces the variance di , so the smoothing can happen over a more precise input. At the domain level, the auxiliary can enter through xi . A good domain-level auxiliary improves the regression component the input is pulled toward, which matters most in the small domains. Under the linking model in (4), the domain effects are modeled a priori as independent draws from a single distribution. However, evaluation domains often provide more structure, with a domain sitting inside nested reporting categories. We call such a nested partition of the domains a taxonomy, borrowing the term AI evaluation uses for the hierarchical category schemes its benchmarks are organized by (Vidgen et al., 2024; Li et al., 2024). Figure 2 shows an example taxonomy from the Open LLM Leaderboard benchmark dataset. We index the taxonomy levels by r = 1, . . . , R from coarsest to finest, with level R being the domains themselves. Write ar (i) for the level-r category of domain i, so that aR (i) = i. The single random effect can be replaced by a sum of effects per level, θi =
x⊤ i β+
R X (r) var (i) ,
(r) iid vk ∼ N 0, σr2 ,
r=1
independently across levels, with the sampling model in (4) unchanged. We refer to this linking model as prediction-powered taxonomy smoothing (PP-TS). Under PP-TS, domains share the effects of every category above them, so information is borrowed most strongly among the domains the taxonomy places closest together. Each level’s variance σr2 is estimated from the data, so a level whose categories do not differ beyond the covariates receives a variance near zero. When applied to the HT estimator instead of GREG, the taxonomy linking model is the Fay–Herriot model with two-fold (Torabi and Rao, 2014) or three-fold (Marcis et al., 2023) random effects, which we label TFH. 7
We fit all models using Bayesian inference, with a flat prior p(β, σ 2 ) ∝ 1 for FH and PP-S. For TFH and PP-TS, we assign a diffuse normal prior on β and weakly informative inverse-gamma priors on the variances σr2 to stabilize estimation.
3.3
Design-Based Validation
The usefulness of the auxiliary information and smoothing is unknown in advance, so no candidate estimator is uniformly preferable. We therefore develop a design-based procedure for comparing the disaggregated estimators. We split the sampled units of each domain at random into K folds so that HT,(k) each fold is itself a probability sample. Write θ̂i for the HT estimate from fold k alone, with the weights multiplied by K, since the fold holds one K-th of the sample. A candidate estimator M fitted M,(−k) without fold k is θ̂i . We evaluate each candidate estimator in terms of its mean squared error under the cross-validation design, h 2 i M,(−k) MSEi (M ) = E θ̂i − θi , (5) where the expectation is over both the sampled units and the random splitting into K folds. For the choice of K and its effect on validation, see McAlinn and Takanashi (2025). A naive CV score estimating the target MSE can be computed by comparing the candidate estimator with the held-out HT estimate, K
Snaive (M ) i
1 X M,(−k) HT,(k) 2 = θ̂i − θ̂i . K k=1
HT,(k)
Because the score uses θ̂i in place of θi , the naive score is biased for the target error (5). The first source is the fold-level bias, which arises from the split of the sample into folds. The held-out HT estimate for each fold is noisy and correlated with the fit on the other folds, and both shift every candidate’s score. The second source is the sample-level bias; for a given sample, every fold inherits the realized error of the HT estimator, θ̂iHT − θi , even though the HT estimator is unbiased over repeated samples. Dong and Li (2026) remove the fold-level bias, applying the correction naive Sadj (M ) − i (M ) = Si
K
K
k=1
k=1
2 M,(−k) 2 X HT,(k) 1 X HT,(k) θ̂i θ̂i − θ̄iHT + − θ̄iHT θ̂i − θ̄iM , K K HT,(k)
M,(−k)
and θ̂i over the K folds. The two correction terms where θ̄iHT and θ̄iM are the averages of θ̂i estimate the fold-level bias, with full details in Appendix A. This adjusted score does not account for the sample-level bias. Left uncorrected, this bias favors candidates that track the HT estimate on the given sample over those that track the true domain mean (Appendix A.2). Dong and Li (2026) propose a bound for the remaining sample-level bias instead. Assuming that each fit without one fold centers on the full-sample fit up to a small remainder, the sample-level bias can be estimated to give a debiased score h i adj HT M Sdeb (M ) = S (M ) − d + 2 Cov θ̂ , θ̂ , (6) i i i i i where di here is the variance of the HT estimator, the same for every candidate M . We refer to validation with this score as DB-CV. The covariance in (6) is available in closed form for every candidate estimator considered here. For PPI = y − f and the HT estimator, the covariance is the variance di . For PPI and GREG, write rij ij ij 8
GREG = y − λ̂f for the residual, with λ̂ held fixed as in the GREG variance. Their covariance with rij ij ij the HT estimator is estimated by the bivariate form of the variance estimator that gives di , P 2 n j∈Si wij (1 − πij ) (yij − ȳi )(rij − r̄i ) i HT GREG d , Cov θ̂i , θ̂i = 2 P ni − 1 j∈Si wij
P P with ȳi = j∈Si wij yij / j∈Si wij the weighted mean of the labels and r̄i the same weighted mean of the residuals. For the smoothers, the covariance is h i h i VarM [ θi | z ] Cov θ̂iHT , θ̂iM ≈ Cov θ̂iHT , zi , Var[ zi ] where VarM [ θi | z ] is the posterior variance of θi under the model. The approximation is exact when the variance parameters are known and holds to first order when the posterior integrates over them. The value of hVar[ zi ] depends on the input direct estimator. When the input is the HT estimate, Var[ zi ] i HT and Cov θ̂i , zi both equal di and the covariance reduces to the posterior variance. When the input h i is the GREG estimate, Var[ zi ] is the GREG variance and Cov θ̂iHT , zi is the covariance defined for GREG. We derive each of these in Appendix A.4.
4
Results
We study the methods on one instance of each setting distinguished in Section 1. The first is a curated benchmark with verifiable grading, the Open LLM Leaderboard as compiled by Wu et al. (2026). Here, a sample from a benchmark is used to estimate a single LLM’s performance across domains. The second is PRISM (Kirk et al., 2024), which we treat as the traffic of a deployed AI system that routes each conversation to one of many LLMs. Each response is graded by the user who received it, and the domains are the combinations of conversation type and LLM. We use the two studies for different purposes. The benchmark study is a comprehensive comparison of every estimator we discuss in this paper against the oracle. The traffic monitoring study demonstrates a workflow using a single sample and evaluates the DB-CV score. The dataset at hand is treated as a finite population, held fixed throughout, so every θi is known exactly and the only randomness is which units are sampled. Table 1 summarizes the two studies. In both studies, labels are drawn by stratified SRS without replacement, with domains as strata and allocation proportional to Ni . We set the sampling budget at 10% of each population throughout. The appendix repeats the results at larger budgets, and similar conclusions are found. Sampling at each budget is replicated S = 100 times. The sample is a sizable share of each population, so the variance of every direct estimate includes a finite-population correction. For each estimator, we report a point estimate θ̂i and√a 95% interval [li , ui ] for each domain. For direct estimators, we consider the Wald interval θ̂i ± 1.96 di and for model-based estimators, we use the posterior credible interval. Errors are aggregated over domains with weights qi that sum to one. On the benchmark every task type is reported on its own, so qi = 1/m. In traffic an error matters in proportion to the traffic it affects, so qi = Ni /N , the domain’s traffic share. Point estimate accuracy is summarized by the root mean squared error (RMSE), RMSE =
m X
qi (θ̂i − θi )2
i=1
9
1/2
.
Table 1: The two studies, with the population, estimand and auxiliary information of each. Open LLM Leaderboard
PRISM agent traffic
Setting
verifiable: graded against a known answer
open-ended: graded by a human user
Population
9,324 questions from five benchmarks
68,371 rated LLM responses from 1,500 participants
Outcome
microsoft/phi-4’s grade (binary)
the participant’s satisfaction rating (1 to 100)
Domains
m = 34 task types nested in the 5 benchmarks
m = 63 crossings of 21 LLMs and 3 conversation types
Estimand θi
phi-4’s accuracy
mean satisfaction
Auxiliary information
answers from two historical models; historical difficulty
an LLM judge’s score; content covariates
Labels at 10%
932, about 27 per domain
6,837, about 108 per domain
Study scope
assessing the accuracy of every estimator
demonstrating the workflow and assessing its validation step
For the intervals we report the empirical coverage of the nominal 95% level, the weighted fraction of domains whose interval contains θi . Coverage checks calibration alone, and an estimator can game it with wide intervals, so we also report the interval score (Gneiting and Raftery, 2007), a proper scoring rule that penalizes width and missed coverage together, ISα =
m n o X 2 2 qi (ui − li ) + (li − θi ) I{θi < li } + (θi − ui ) I{θi > ui } , α α i=1
where I{·} is the indicator function, and smaller values indicate better interval estimates. We take α = 0.05 in our analysis. We report the means of these metrics across the S replicate samples.
4.1
Estimator Comparison: Open LLM Leaderboard
We use the question-by-model record compiled from the Open LLM Leaderboard by Wu et al. (2026), containing data for about 4,400 models. This compilation is a valuable resource for evaluation research since the original records from the leaderboard are scattered across thousands of files. The LLM under evaluation is microsoft/phi-4, and the estimand is its accuracy on each of the 34 task types. Each task type belongs to one of five benchmarks (BBH, GPQA, IFEval, MATH and MuSR), and this two-level structure is the taxonomy shown in Figure 2. We use three auxiliaries, all built from the leaderboard’s historical records on every question, which contain neither phi-4 nor any model fine-tuned from it. Two of the auxiliaries are the answers of a single earlier model, scored correct or incorrect: a math specialist (Qwen2.5-Math-7B-Instruct) and a small generalist (Qwen2.5-1.5B-Instruct). The third is historical difficulty, the share of earlier models that answered the question correctly. The answers of the two historical models are weak auxiliaries and difficulty is a strong one, correlating 0.31, 0.33 and 0.61 with the outcome at the unit level. The comparison covers the three direct estimators (HT, PPI and GREG) and the four smoothers: 10
Table 2: RMSE and IS of the seven estimators at the 10% budget under each auxiliary, averaged over 100 replications (Open LLM Leaderboard, 34 domains). Each smoother carries the auxiliary’s domain mean in its linking model. Lower is better and the best value in each column is bold. The HT estimator uses no auxiliary, so its row is the same under all three. Standard errors of the means are at most 0.001 for RMSE and 0.01 for IS. Specialist model
Generalist model
Historical difficulty
Estimator
RMSE
IS
RMSE
IS
RMSE
IS
θ̂iHT θ̂iPPI θ̂iGREG
0.082 0.106 0.080
0.418 0.498 0.393
0.082 0.105 0.082
0.418 0.503 0.409
0.082 0.074 0.074
0.418 0.358 0.358
θ̂iFH θ̂iPP-S
0.078 0.075
0.377 0.357
0.072 0.071
0.362 0.355
0.067 0.061
0.335 0.301
θ̂iTFH θ̂iPP-TS
0.074 0.072
0.366 0.350
0.071 0.070
0.359 0.351
0.065 0.059
0.326 0.295
FH and TFH on the HT input, and PP-S and PP-TS on the GREG input. Every smoother uses the auxiliary’s domain mean as its only domain-level covariate. Table 2 reports RMSE and IS at the 10% budget under each auxiliary. First we see that the prediction-powered smoothers outperform their baselines under every auxiliary on both RMSE and IS. The gap is larger under historical difficulty, which is a strong auxiliary at both the unit and domain levels. Notably, PP-TS outperforms all estimators, direct or smoothing, for every auxiliary. Every estimator’s coverage lies between 0.93 and 0.95 against the nominal 0.95. The full results at every budget, including coverage, are given in Appendix B. The value of GREG’s tuning is apparent, especially under the weak auxiliaries. PPI has no such tuning, so under the weak auxiliaries its RMSE rises above that of HT. PPI is designed for an auxiliary that is itself a prediction of the outcome and a weak prediction can still make it perform worse than the HT estimator. GREG tunes the correction slope down to about a third and stays at the level of HT. Under historical difficulty, GREG’s tuned slope is close to one and the two estimators agree. Figure 3 breaks down the components of PP-TS one at a time and shows where its gain over HT comes from. We see that pooling on its own adds little: the Fay–Herriot model with intercept-only linking on the HT input uses no auxiliary, and it barely improves on HT. Most of the gains come from the other components. Replacing the HT input with the GREG input helps most under the strong auxiliary, where GREG performs well. The biggest gains generally come from the domain-level covariate. It is clear that an auxiliary’s value at the domain level cannot be determined from its unit-level accuracy. The two weak auxiliaries have nearly the same unit-level correlation with the outcome, yet the generalist’s domain mean lowers RMSE far more than the specialist’s. These findings likely apply to auxiliary information provided by LLM judges in general: two judges of equal accuracy can differ in what they contribute as a covariate. Taxonomy smoothing picks up structure that the covariate leaves behind. The specialist’s domain mean errs unevenly (the specialist does well on the MATH task types but poorly on IFEval) and this pattern is exactly what taxonomy smoothing absorbs. PP-TS therefore gains most under the specialist, and it brings the two weak auxiliaries close to even. How much an auxiliary helps at the domain level, and how much taxonomy smoothing adds, both vary with the data at hand. Thus, the linking structure of the smoothing model must be validated.
11
specialist model (r = 0.31)
generalist model (r = 0.33)
0.082
0.082 0.080
0.080
historical difficulty (r = 0.61) 0.082
0.080
0.080
0.080
0.078 0.075
RMSE
0.075 0.072
0.072
0.071 0.070
0.070
0.065 0.061 0.060
F
HT t only input riate nomy p G ova taxo rce RE +c + inte + G , H
HT t only input riate nomy p G ova taxo rce RE +c + inte + G , H
F
0.059
HT t only input riate nomy p G ova taxo rce RE +c + inte + G , H
F
Figure 3: RMSE at the 10% budget as the components of PP-TS are added one at a time, under each auxiliary (Open LLM Leaderboard, 100 replications). Components accumulate left to right from the HT estimator: the Fay–Herriot model on the HT input with intercept-only linking, the GREG input (PP-S with xi = 1), the auxiliary’s domain mean as linking covariate (PP-S), and taxonomy smoothing (PP-TS). The panel strips give each auxiliary’s unit-level correlation r with the outcome. The printed values are entries of Table 2 and of the tables in Appendix B.
4.2
Workflow and Validation: PRISM Agent Traffic
A real-world evaluation of a deployed AI system has three steps: design and collection of the sample, fitting candidate estimators, and validation without an oracle. To demonstrate this, we first walk through a single-sample analysis that compares multiple candidate estimators using the DB-CV score of Section 3.3. We then assess DB-CV against alternative validation procedures over repeated samples. We treat the PRISM dataset as the traffic of a single deployed service that routes conversations across many LLMs and evaluate that service as a whole. The estimand is the mean satisfaction rating in each of the 63 domains, the crossings of 21 LLMs with 3 conversation types. The two-level structure of routed LLMs each with the 3 conversation types provides the taxonomy. We use gpt-5-nano at low reasoning effort as a judge to predict each participant’s own rating based on the content of each conversation (see Appendix C.2). Its score correlates 0.40 with the ratings at the unit level and 0.93 between domain means. As before, we report the 10% budget, which gives 6,837 labels or about 108 per domain. Single-sample workflow. In our first example, the analyst is given one sample and considers seven candidate estimators: the HT estimator along with variants of GREG, PP-S and PP-TS. GREG, PPS, and PP-TS are each fit with the judge or with both the judge and content covariates from the conversation. The content covariates include the length of the response and other information about the conversation (see Appendix C.1). When an auxiliary is used, it is used both in the GREG at the unit level and in the smoothing as domain means. DB-CV uses K = 5 folds within strata, and its scores are aggregated over domains with the traffic-share weights qi = Ni /N and reported on the RMSE scale. 12
Table 3: The candidate field on one 10% sample (replicate 1). The DB-CV score, on the RMSE scale, and the mean 95% interval width are computed from the sample alone. The last column is the candidate’s oracle RMSE, weighted by traffic share, over 100 replications. The full result across replications is given in Appendix C.3. candidate HT GREG (judge) GREG (judge + content) PP-S (judge) PP-S (judge + content) PP-TS (judge) PP-TS (judge + content)
DB-CV score
interval width
oracle RMSE (hidden from analyst)
2.90 2.75 2.39 2.43 1.09 2.07 1.72
10.47 10.01 8.87 8.97 5.59 8.57 7.64
2.505 2.418 2.148 2.253 1.527 1.943 1.720
Table 3 shows the candidate field, where every column but the last is computed from the sample alone. DB-CV selects PP-S with the judge and content covariates, and orders all seven candidates exactly as the oracle RMSE does. The content covariates clearly help, and through validation, the analyst can see this without the oracle. That is, every candidate that includes content covariates has a lower score and a narrower interval compared to its counterpart without content covariates. With the judge alone, PP-TS beats PP-S on both the score and the oracle RMSE, but adding the content covariates reverses the order. Figure 4 shows the interval widths of four candidates on the left and the estimates of the selected candidate on the right, with domains ordered by the standard error of the HT estimator. Smoothing narrows the intervals most where HT is least precise, to about half of HT’s width in the sparsest quarter of domains. These sparse domains also happen to have the lowest satisfaction, so precision is gained where monitoring may be needed most. Smoothing may add more operational value in monitoring rare outcomes, where the direct estimates can be even less precise. Dobi et al. (2026) address this with larger samples and by pooling over weeks. Validation at a fixed budget. In the single-sample analysis, the rankings of DB-CV and the oracle RMSE agree. We now turn to the evaluation of the model comparison procedure over repeated samples. We compare the proposed DB-CV procedure with three other validation procedures at the same total sampling budget over 100 replications. The first is the naive CV score of Section 3.3, with no bias correction. The other two split the sampling budget between two independent samples: the candidates are fit on the first and scored against the second, and the chosen candidate is then refit on the pooled labels. We test an even 50/50 split and the 80/20 split conventional for a held-out set. We also consider the adjusted CV score in Dong and Li (2026), but in this task, the procedure abstains on every pairwise comparison due to the conservative bounds, so its results are not reported here. Table 4 grades each procedure on three dimensions. The first is the oracle RMSE of the candidate it chooses across samples. On this dimension, the two CV scores and the 50/50 split choose roughly equally well, and the 80/20 split chooses worse. The second is the rank correlation between scores for a particular validation procedure and the oracle ordering. Here the two CV scores order the candidates better than either split. The third is the error each procedure reports, relative to the oracle RMSE of its chosen candidate. This is where DB-CV differs most from the others. The splits and naive CV overstate the chosen candidate’s error several times over, because the held-out estimates they score 13
interval width, four of the candidates
selected: PP−S (judge + content) 80
mean satisfaction
95% interval width
20 15 10 5
60
40
20
0 0
20
40
60
0
20
40
60
domain, ordered by the HT standard error HT
GREG (judge)
GREG (judge + content)
PP−S (judge + content)
Figure 4: The single-sample analysis (PRISM, 10%, replicate 1). Left: interval widths of four candidates over the 63 domains. Right: the domain estimates and 95% intervals of the selected candidate, PP-S with judge and content covariates, against the estimated traffic-wide mean (dashed). Domains are ordered by the standard error of the HT estimator, so the right end of each panel is where smoothing has the most to do. In the sparsest quarter of domains the intervals average 15 points for HT, 12 for GREG with all covariates and 7 for PP-S. against carry sampling error of their own (Section 3.3). Debiasing removes this, and the benefit has practical significance, as the reported error is what tells the analyst whether the chosen estimates can be trusted and whether the sampling budget was enough. To conclude, the proposed DB-CV score can select as well as two independent samples at the same budget and it is the only procedure whose reported error is close to the oracle.
5
Conclusion
This paper builds on ideas from small area estimation to develop an integrated workflow for estimation and validation in disaggregated AI evaluation. We apply it to a benchmark with verifiable grading and to deployed agent traffic graded by humans, two settings that to our knowledge have not been treated as one estimation problem. The proposed PP-S approach, smoothing a difference estimator across domains, improves both point and interval estimates over PPI and the other direct estimators in both settings. Coverage of the interval estimates stays near nominal throughout. In the benchmark study, we found that shrinkage across domains on its own accounts for little of the improvement. The auxiliary information used at the unit and domain levels, and the reporting taxonomy, provide essential structure for the improvements from smoothing. The structure that helps differs by dataset, so validation is a critical part of the workflow. The debiased score we propose makes this choice from a single sample and, at equal budget, selects just as well as having two independent samples (one for training and the other for validation). The proposed score also estimates the chosen estimator’s error accurately, while the alternatives overstate it several times over. One immediate extension concerns AI evaluation in unseen domains with no observed outcomes (Saxena et al., 2024). The smoothing models may still be used to estimate performance in a domain with no labeled units whenever its domain-level covariates are available. Determining when such zero-
14
Table 4: Four validation procedures at the same total sampling budget of 10%, 100 replications. Rank correlation is Kendall’s τ between each procedure’s scores and the oracle ordering of the seven candidates; the last column is the reported RMSE of the chosen candidate as a multiple of its oracle RMSE. The 20% budget is in Appendix C.3. validation procedure
RMSE of chosen
rank corr.
reported / oracle RMSE
1.543 1.528 1.557 1.678
0.89 0.88 0.78 0.62
1.07× 4.00× 2.59× 3.54×
DB-CV (one sample) naive CV (one sample) 50/50 (two samples) 80/20 (two samples)
shot domain evaluation is reliable remains an open question. More broadly, there is untapped potential in the deeper integration of AI evaluation and survey methodology. For example, a grade for an openended task is a judgment, and two graders may not agree on a single number. The survey literature treats this disagreement as measurement error and carries it into the uncertainty of the estimate (Hansen et al., 1961). Another example is the large nonprobability samples that deployed systems generate, such as conversations flagged for review. Combining such data with a probability sample guards against its selection bias (Kawano et al., 2026b). Uncertainty quantification is essential to AI evaluation, and it is where the integration of these two fields could be most fruitful.
Data Availability Code reproducing every figure and table in this paper is available at https://github.com/shokawano/disagg-ai-eval. The repository includes the processed Open LLM Leaderboard questionby-model tables used in the benchmark study, derived from the public release of Wu et al. (2026), and the precomputed LLM judge scores and frozen judge prompt used in the traffic study. The PRISM dataset (Kirk et al., 2024) is publicly available on the Hugging Face Hub at https://huggingface. co/datasets/HannahRoseKirk/prism-alignment.
Acknowledgements We used Anthropic’s Claude models (Opus 4.8 and 5, Fable 5 and 5.1) to assist with coding, literature search, review of derivations, LaTeX formatting, and manuscript editing.
Funding and Conflict of Interest None declared.
15
References Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zrnic, T. (2023a). Prediction-powered inference. Science, 382(6671):669–674. Angelopoulos, A. N., Duchi, J. C., and Zrnic, T. (2023b). PPI++: Efficient prediction-powered inference. arXiv:2311.01453. Barocas, S., Guo, A., Kamar, E., Krones, J., Morris, M. R., Vaughan, J. W., Wadsworth, D., and Wallach, H. (2021). Designing disaggregated evaluations of AI systems: Choices, considerations, and tradeoffs. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES). ACM. Dobi, A., Manickavasagam, A., Yang, X., Thompson, B., and Farooq, F. (2026). Measuring the prevalence of policy-violating content with ML-assisted sampling and LLM labeling. arXiv:2602.18518. Dong, Q. and Li, Z. R. (2026). Design-based cross-validation for comparing small area estimators. arXiv:2604.23464. Efron, B. (2011). Tweedie’s formula and selection bias. Journal of the American Statistical Association, 106(496):1602–1614. Emmenegger, N., Stahler, E., and Podimata, C. (2026). Prediction-powered inference across many tasks for AI evaluation & social science research. arXiv:2605.29249. Fay, R. E. and Herriot, R. A. (1979). Estimates of income for small places: An application of James–Stein procedures to census data. Journal of the American Statistical Association, 74(366):269–277. Fisch, A., Maynez, J., Hofer, R. A., Dhingra, B., Globerson, A., and Cohen, W. W. (2024). Stratified prediction-powered inference for hybrid language model evaluation. arXiv:2406.04291. Fogliato, R., Patil, P., Akpinar, N.-J., and Monfort, M. (2024a). Precise model benchmarking with only a few observations. arXiv:2410.05222. Fogliato, R., Patil, P., Monfort, M., and Perona, P. (2024b). A framework for efficient model evaluation through stratification, sampling, and estimation. arXiv:2406.07320. Gao, P. A. and Wakefield, J. (2024). Smoothed model-assisted small area estimation of proportions. Canadian Journal of Statistics, 52:337–358. Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378. Hansen, M. H., Hurwitz, W. N., and Bershad, M. A. (1961). Measurement errors in censuses and surveys. Bulletin of the International Statistical Institute, 38(2):359–374. Herlihy, C., Truong, K., Chouldechova, A., and Dudı́k, M. (2024). A structured regression approach for evaluating model performance across intersectional subgroups. arXiv:2401.14893. Horvitz, D. G. and Thompson, D. J. (1952). A generalization of sampling without replacement from a finite universe. Journal of the American Statistical Association, 47(260):663–685. Kawano, S., Parker, P. A., and Li, Z. R. (2026a). On data thinning for model validation in small area estimation. arXiv:2604.04141. 16
Kawano, S., Vedensky, D., Dong, Q., Pawl, E., Wang, Q., Parker, P. A., Li, Z. R., and Holan, S. H. (2026b). Nonprobability samples for small area estimation: A review and comparative simulation study. arXiv:2608.13673. Kirk, H. R., Whitefield, A., Röttger, P., Bean, A. M., Margatina, K., Mosquera-Gomez, R., Ciro, J., Bartolo, M., Williams, A., He, H., et al. (2024). The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. In Advances in Neural Information Processing Systems, volume 37. arXiv:2404.16019. Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J. (2024). SALAD-bench: A hierarchical and comprehensive safety benchmark for large language models. arXiv:2402.05044. Li, S. and Ignatiadis, N. (2025). Prediction-powered adaptive shrinkage estimation. arXiv:2502.14166. Liang, P. et al. (2022). Holistic evaluation of language models. arXiv:2211.09110. Maia Polo, F., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M. (2024). tinyBenchmarks: Evaluating LLMs with fewer examples. arXiv:2402.14992. Marcis, L., Morales, D., Pagliarella, M. C., and Salvatore, R. (2023). Three-fold Fay–Herriot model for small area estimation and its diagnostics. Statistical Methods & Applications, 32:1563–1609. McAlinn, K. and Takanashi, K. (2025). Determining the K in K-fold cross-validation. arXiv:2511.12698 [stat]. Meng, X.-L. (2018). Statistical paradises and paradoxes in big data (i): Law of large populations, big data paradox, and the 2016 US presidential election. The Annals of Applied Statistics, 12(2):685–726. Molenberghs, G., Beunckens, C., Sotto, C., and Kenward, M. G. (2008). Every missingness not at random model has a missingness at random counterpart with equal fit. Journal of the Royal Statistical Society Series B: Statistical Methodology, 70(2):371–388. Mozer, R. (2026). PPI is the difference estimator: Recognizing the survey sampling roots of predictionpowered inference. arXiv:2603.19160. Pew Research Center (2023). Comparing two types of online survey samples. Technical report, Pew Research Center. Pfohl, S. R., Harris, N., Nagpal, C., Madras, D., Mhasawade, V., Salaudeen, O., Dieng, A., Sequeira, S., Arciniegas, S., Sung, L., Ezeanochie, N., Cole-Lewis, H., Heller, K., Koyejo, S., and D’Amour, A. (2025). Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness. arXiv:2506.04193. Rao, J. N. K. and Molina, I. (2015). Small Area Estimation. Wiley, Hoboken, NJ, 2nd edition. Särndal, C.-E., Swensson, B., and Wretman, J. (1992). Model Assisted Survey Sampling. Springer, New York. Saxena, R., Kim, T., Mehra, A., Baek, C., Kolter, Z., and Raghunathan, A. (2024). Predicting the performance of foundation models via agreement-on-the-line. Advances in Neural Information Processing Systems, 37:31854–31906.
17
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., ColeLewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., Agüera y Arcas, B., Webster, D., Corrado, G. S., Matias, Y., Chou, K., Gottweis, J., Tomasev, N., Liu, Y., Rajkomar, A., Barral, J., Semturs, C., Karthikesalingam, A., and Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620:172–180. Starace, G., Jaffe, O., Sherburn, D., Aung, J., Chan, J. S., Maksin, L., Dias, R., Mays, E., Kinsella, B., Thompson, W., Heidecke, J., Glaese, A., and Patwardhan, T. (2025). PaperBench: Evaluating AI’s ability to replicate AI research. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 56843–56873. Torabi, M. and Rao, J. N. K. (2014). On small area estimation under a sub-area level model. Journal of Multivariate Analysis, 127:36–55. Vidgen, B. et al. (2024). arXiv:2404.12241.
Introducing v0.5 of the AI safety benchmark from MLCommons.
Wu, S., Nair, Y., and Candès, E. J. (2026). Efficient evaluation of LLM performance with statistical guarantees. arXiv:2601.20251. Yauney, G., Warraich, S. S., and Swayamdipta, S. (2026). How reliable is language model microbenchmarking? In International Conference on Learning Representations (ICLR). arXiv:2510.08730. Zhang, G., Dorner, F. E., and Hardt, M. (2025). How benchmark prediction from fewer data misses the mark. arXiv:2506.07673. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36, Datasets and Benchmarks Track.
18
A
Details for Design-Based Validation
Here we provide the derivations of the debiased score of Section 3.3. For ease of notation, we work in one domain and suppress its index. Here θ is the domain mean and θ̂HT the full-sample HT estimator, with variance d. Our objective is to validate a candidate M , where θ̂M is its estimator computed from the full sample. We divide our sample into folds indexed by k. For each fold, we construct an HT estimate θ̂HT,(k) using the held-out fold k, which is used for validation. We also construct the estimator θ̂M,(−k) based on the remaining K − 1 folds. We denote the averaged value of these quantities across folds as θ̄HT and θ̄M . Two sources of randomness are in play, which sample S is drawn and how it is split into folds. Every expectation below operates over both sources.
A.1
The Bias of the Naive Score
Write eM,(−k) = θ̂M,(−k) − θ,
eHT,(k) = θ̂HT,(k) − θ
for the errors of the candidate and held-out HT estimate at fold k. The target is the candidate’s mean squared error averaged over the folds, MSE(M ) =
K 1 X h M,(−k) 2 i E (e ) . K k=1
The folds are drawn at random, so the error has the same distribution at every fold. The naive score replaces the unknown θ by the held-out HT estimate, whose discrepancy from the candidate is the difference of the two errors, Snaive (M ) =
K K 2 2 1 X M,(−k) 1 X M,(−k) θ̂ − θ̂HT,(k) = e − eHT,(k) . K K k=1
k=1
We first seek to show that the bias for this naive score is: K K i i h h 2 X 1 X E Snaive (M ) − MSE(M ) = Var θ̂HT,(k) − Cov θ̂HT,(k) , θ̂M,(−k) . K K k=1
k=1
Proof. Given the definition of the naive score as a difference of the two errors, we have i i h h h i h 2 i E eM,(−k) − eHT,(k) = E (eM,(−k) )2 + E (eHT,(k) )2 − 2 E eM,(−k) eHT,(k)
(7)
where this is a per-fold quantity. Note the second term is simply the variance of the held-out HT estimate, since it is unbiased. The last cross term can be simplified to h i h i h i h i E eM,(−k) eHT,(k) = Cov eM,(−k) , eHT,(k) + E eM,(−k) E eHT,(k) h i h i h i = Cov θ̂M,(−k) , θ̂HT,(k) + E eM,(−k) E eHT,(k) h i h i = Cov θ̂HT,(k) , θ̂M,(−k) + E eM,(−k) · 0. The second equality is due to θ being a constant and the third equality results from the unbiasedness of the held-out HT estimate. Then (7) can be simplified to h h i h i h i 2 i E eM,(−k) − eHT,(k) = E (eM,(−k) )2 + Var θ̂HT,(k) − 2 Cov θ̂HT,(k) , θ̂M,(−k) . 19
Thus, the expectation of the naive score is
naive
E S
K 2 i 1 X h M,(−k) (M ) = E e − eHT,(k) K k=1 K h i h i 1 X h M,(−k) 2 i E (e ) + Var θ̂HT,(k) − 2 Cov θ̂HT,(k) , θ̂M,(−k) = K
k=1
= MSE(M ) +
K K h i h i 1 X 2 X Var θ̂HT,(k) − Cov θ̂HT,(k) , θ̂M,(−k) . K K k=1
k=1
Note that this bias breaks into two components when conditioned on the realized sample S. The variance of the held-out estimate splits as h i h h ii h h ii + Var E θ̂HT,(k) | S Var θ̂HT,(k) = E Var θ̂HT,(k) | S h h ii h i = E Var θ̂HT,(k) | S + Var θ̂HT h h ii = E Var θ̂HT,(k) | S + d. The first equality is the law of total variance. The second holds because, given the sample, each sampled unit lands in fold k with probability 1/K, which cancels the factor K on its weight, so the held-out estimate has conditional mean θ̂HT .1 The covariance with the candidate splits in the same way, h i h h ii Cov θ̂HT,(k) , θ̂M,(−k) = E Cov θ̂HT,(k) , θ̂M,(−k) | S h h i h ii + Cov E θ̂HT,(k) | S , E θ̂M,(−k) | S h h ii h h ii = E Cov θ̂HT,(k) , θ̂M,(−k) | S + Cov θ̂HT , E θ̂M,(−k) | S , by the law of total covariance and the same conditional mean. Averaging the last term over the folds, K h h ii h i 1 X Cov θ̂HT , E θ̂M,(−k) | S = Cov θ̂HT , E θ̄M | S K k=1 h i h h ii = Cov θ̂HT , θ̄M − E Cov θ̂HT , θ̄M | S h i = Cov θ̂HT , θ̄M .
The first equality is linearity, the second is the law of total covariance for θ̄M , and the third holds 1
For the Hájek estimator, which divides by the sum of the weights rather than the domain size, this conditional mean and the fold average used in Appendix A.2 hold to first order, and exactly under simple random sampling within the domain with equal fold sizes.
20
because θ̂HT is fixed given the sample. Substituting the three results into the bias of the naive score, E Snaive (M ) − MSE(M ) K K h i h i 1 X 2 X = Var θ̂HT,(k) − Cov θ̂HT,(k) , θ̂M,(−k) K K =
1 K
k=1 K n X
k=1
h h ii o E Var θ̂HT,(k) | S +d
k=1
−
K h ii h h i io 2 Xn h E Cov θ̂HT,(k) , θ̂M,(−k) | S + Cov θ̂HT , E θ̂M,(−k) | S K k=1
=
K X
K h ii h ii 1 2 X h E Var θ̂HT,(k) | S − E Cov θ̂HT,(k) , θ̂M,(−k) | S K K k=1 k=1 {z } | fold-level bias i h + d − 2 Cov θ̂HT , θ̄M . | {z }
h
(8)
sample-level bias
In the last equality, we group the terms into fold-level and sample-level bias components. The fold-level bias comes from the random split of a fixed sample. The sample-level bias comes from the draw of the sample itself. Every held-out estimate centers on the full-sample HT estimate rather than on θ, so the spread across the folds carries no information about it. We first examine the adjusted score of Dong and Li (2026) that removes the fold-level bias.
A.2
The Adjusted Score
The fold-level bias above has two terms, the average conditional variance of the held-out estimates and their average conditional covariance with the candidate. The adjusted score of Dong and Li (2026) estimates them by correcting the naive score, Sadj (M ) = Snaive (M ) − υ̂ + 2 ĉM , where υ̂ and ĉM are defined by υ̂ =
K 2 1 X HT,(k) θ̂ − θ̄HT , K
ĉM =
k=1
K 1 X HT,(k) θ̂ − θ̄HT θ̂M,(−k) − θ̄M . K k=1
We seek to show that E[ υ̂ ] =
K h ii 1 X h E Var θ̂HT,(k) | S , K k=1
K h ii 1 X h E ĉM = E Cov θ̂HT,(k) , θ̂M,(−k) | S . K k=1
Proof. The proof uses two facts about the held-out HT estimates, which Pfollow from how the folds are (k) HT,(k) −1 built. Write S for the sampled units in fold k, so that θ̂ =N j∈S (k) K wj yj . The first fact is that the held-out estimates average to the full-sample estimate, θ̄
HT
K 1 X 1 X 1 X = K w j yj = wj yj = θ̂HT , K N N (k) k=1
j∈S
j∈S
21
(9)
where the second equality cancels the factor K and uses that each sampled unit sits in exactly one fold. The second fact is that, given the sample, each held-out estimate has conditional mean equal to the full-sample estimate, h i 1 X 1 X E θ̂HT,(k) | S = K wj yj P j ∈ S (k) | S = wj yj = θ̂HT , N N j∈S
(10)
j∈S
where the first equality takes the expectation over the split with the weights and labels fixed, and the second uses that each sampled unit lands in fold k with probability 1/K. Given the sample, the expectation of υ̂ is E[ υ̂ | S ] =
K i 2 1 X h HT,(k) E θ̂ − θ̄HT | S K k=1
K i 2 1 X h HT,(k) = E θ̂ − θ̂HT | S K
=
1 K
k=1 K X
h i Var θ̂HT,(k) | S .
k=1
The first equality is linearity of conditional expectation, and the second uses (9). The third uses (10), since a squared deviation from the conditional mean has conditional expectation equal to the conditional variance. The expectation of ĉM takes two steps. First, the candidate’s fold average θ̄M drops out, " # K M 1 X HT,(k) θ̂ − θ̂HT θ̂M,(−k) − θ̄M | S E ĉ | S = E K k=1 X K K X 1 HT,(k) HT M,(−k) HT,(k) HT M 1 =E θ̂ − θ̂ θ̂ − θ̂ S θ̂ − θ̄ K K =
1 K
k=1 K X h
E
k=1
i
θ̂HT,(k) − θ̂HT θ̂M,(−k) | S .
k=1
The first equality uses (9), and the second expands the product and takes θ̄M , which does not depend on k, outside the second sum. The third holds because the second sum is θ̄HT − θ̂HT , which is zero by (9). h i Second, each summand is a conditional covariance. Write µ(k) = E θ̂M,(−k) | S for the conditional mean of the candidate’s fit, which is fixed given the sample. Then h i h i E θ̂HT,(k) − θ̂HT θ̂M,(−k) | S = E θ̂HT,(k) − θ̂HT θ̂M,(−k) − µ(k) | S h i = Cov θ̂HT,(k) , θ̂M,(−k) | S . The first equality subtracts µ(k) , which changes nothing because µ(k) is fixed given the sample and θ̂HT,(k) − θ̂HT has conditional mean zero by (10). The second holds because, by (10), both factors are deviations from their conditional means.
22
Finally, we average the two conditional expectations over samples, E[ υ̂ ] = E[ E[ υ̂ | S ] ] =
K h ii 1 X h E Var θ̂HT,(k) | S , K k=1
K h ii 1 X h E ĉM = E E ĉM | S = E Cov θ̂HT,(k) , θ̂M,(−k) | S . K k=1
In each line, the first equality is the law of total expectation, and the second substitutes the conditional expectation computed above. The adjusted score is therefore left with the sample-level bias, h i h i E Sadj (M ) − MSE(M ) = d − 2 Cov θ̂HT , θ̄M = d (1 − 2ρ), h i where ρ = Cov θ̂HT , θ̄M /d. The bias is zero only at ρ = 1/2. The HT estimator has ρ = 1, so its adjusted score is too low by d. A candidate that uses no labels has ρ = 0, so its adjusted score is too high by d.
A.3
The Debiased Score
h i The sample-level bias is d − 2 Cov θ̂HT , θ̄M , which involves the average of the candidate’s K training fits, while the debiased score of Section 3.3 uses the full-sample fit θ̂M . We assume that, given the sample, each training fit centers on the full-sample fit up to a remainder of order n−1 , with n the size of the whole sample, h i E θ̂M,(−k) | S = θ̂M + Op n−1 . (11) Dong and Li (2026) derive this approximation for the difference of two candidates from a linearization condition, and we assume it for a single candidate. Under this assumption, h i h i d − 2 Cov θ̂HT , θ̄M = d − 2 Cov θ̂HT , E θ̄M | S h i = d − 2 Cov θ̂HT , θ̂M + Op n−1 h i ≈ d − 2 Cov θ̂HT , θ̂M . The first equality holds because θ̂HT is fixed given the sample, and the second averages (11) over the folds. Subtracting this sample-level bias from the adjusted score gives the debiased score (6), h i Sdeb (M ) = Sadj (M ) − d + 2 Cov θ̂HT , θ̂M .
A.4
The Covariance for Each Candidate
h i The debiased score needs d and the covariance Cov θ̂HT , θ̂M , and only the covariance depends on the candidate. Both are computed from the full sample. HT. When the candidate is the HT estimator itself, the covariance is its variance d, so the debiased score adds −d + 2d = d to the adjusted score. 23
PPI and GREG.
Both estimators are a known constant plus a weighted mean of residuals, 1 X 1 X θ̂M = pj + wj rj , N N j∈U
j∈S
where pj is the prediction, equal to fj under PPI and λ̂fj under GREG, and rj = yj − pj is the residual. With λ̂ treated as fixed, the first term is a constant, so h i X X X 1 1 1 Cov θ̂HT , θ̂M = Cov wj yj , pj + wj rj N N N j∈S j∈U j∈S 1 X 1 X wj yj , wj rj . = Cov N N j∈S
j∈S
The first equality writes out both estimators, and the second drops the constant. The result is the covariance between the weighted means of the labels and of the residuals. It is not the candidate’s own variance, which would have the residuals in place of the labels in the first argument. Both d and this covariance are estimated from the full sample by the linearization estimator for a weighted mean, P 2 2 n j∈S wj (1 − πj ) (yj − ȳ) dˆ = , 2 P n−1 w j j∈S P 2 (1 − π ) (y − ȳ)(r − r̄) w j j j j∈S j d θ̂HT , θ̂M = n Cov , 2 P n−1 j∈S wj P P wherePn is the domain’s sample size, ȳ = j∈S wj yj / j∈S wj is the weighted mean of the labels, and P r̄ = j∈S wj rj / j∈S wj is the weighted mean of the residuals. The covariance estimator replaces one of the two centered labels in dˆ by a centered residual. The factor 1 − πj is the finite-population correction, and n/(n − 1) is the small-sample factor. Smoothers. A smoother uses every domain, so we bring back the domain index and write zi for the input of domain i. Once β and the variance parameters are fixed (σ 2 under FH and PP-S, and σr2 under TFH and PP-TS), each of the four smoothers is Az plus a term that does not depend on z, for a matrix A that we treat as fixed. Then # " m i h X Cov θ̂iHT , θ̂iM = Cov θ̂iHT , Ail zl l=1
=
m X
i h Ail Cov θ̂iHT , zl
l=1
h i = Aii Cov θ̂iHT , zi . The first equality writes out row i of Az and drops the constant term, and the second is bilinearity with A fixed. The third holds because sampling errors of the direct estimates are independent across domains. The coefficient Aii has a closed form. Stacking the domains, the four smoothers of Section 3.2 share the form z | θ ∼ N (θ, D) , θ = Xβ + v, v ∼ N (0, G) , 24
where D = diag{Var[ z1 ] , . . . , Var[ zm ]} holds the sampling variances of the inputs, X stacks the covariates, and G is the covariance of the random effects. Given β, the sampling error z − θ is independent of θ, so CovM [ θ, z ] = G and VarM [ z ] = G + D, and conditioning on z gives A = G(G + D)−1 , VarM [ θ | z ] = G − G(G + D)−1 G = G(G + D)−1 (G + D) − G(G + D)−1 G = G(G + D)−1 G + D − G = AD. Because D is diagonal, the i-th diagonal entry gives Aii = VarM [ θi | z ] / Var[ zi ], so h i h i Cov θ̂iHT , zi . Cov θ̂iHT , θ̂iM = VarM [ θi | z ] Var[ zi ] Under FH and TFH the input is the HT estimate, so the ratio is one and the covariance is the posterior variance. Under PP-S and PP-TS the input is the GREG estimate. The numerator of the ratio is then the GREG covariance above, and the denominator is the GREG variance already supplied to the smoother. The derivation above holds the variance parameters fixed. The implementation instead fits each smoother by Markov chain Monte Carlo (MCMC), which integrates over them, so the smoothed estimate is no longer of the form Az plus a constant. A first-order expansion of θ̂iM in the inputs takes the place of that linear form, and the covariance chain above becomes m h i X h i ∂ θ̂iM Cov θ̂iHT , θ̂iM ≈ Cov θ̂iHT , zl ∂zl =
l=1 ∂ θ̂iM
∂zi
h i Cov θ̂iHT , zi .
This is the earlier chain with Ail replaced by ∂ θ̂iM /∂zl , and the equality again uses the independence of the sampling errors across domains. With the variance parameters fixed, ∂ θ̂iM /∂zi = Aii and the two chains coincide. To find the derivative, note that θ̂iM = EM [ θi | z ] is a posterior mean, and apply Tweedie’s formula (Efron, 2011). Under the normal sampling model and any prior on θ, including one that integrates over the variance parameters, the formula and its companion for the posterior variance are θ̂iM = zi + Var[ zi ]
∂ℓ(z) , ∂zi
VarM [ θi | z ] = Var[ zi ] + Var[ zi ]2
∂ 2 ℓ(z) , ∂zi2
where ℓ(z) is the log density of the inputs under the model. The first follows from differentiating ℓ once in zi , which brings down the posterior mean of θi − zi . Differentiating a second time brings down its posterior second moment, and subtracting the square of the first derivative leaves the posterior variance, which gives the second. Differentiating the first, ∂ θ̂iM ∂ 2 ℓ(z) = 1 + Var[ zi ] ∂zi ∂zi2 2 1 2 ∂ ℓ(z) = Var[ zi ] + Var[ zi ] Var[ zi ] ∂zi2 VarM [ θi | z ] = , Var[ zi ] 25
where the second equality factors out 1/ Var[ zi ], and the third recognizes the term in parentheses as the posterior variance from the companion formula. Substituting into the first-order chain gives h i Var [ θ | z ] h i M i Cov θ̂iHT , θ̂iM ≈ Cov θ̂iHT , zi , Var[ zi ] the same formula as in the fixed case, with the posterior variance now computed from the retained draws.
26
B
Additional Details for the Benchmark Simulation
We merge the MATH task types on advanced subjects into one domain and exclude one task type on which phi-4 answers every question correctly, leaving 34 task types with Ni from 146 to 670, median 250. Labels are drawn by SRS without replacement within each domain, so πij = ni /Ni and the variance of the HT estimator is estimated by ni s2i di = 1 − , Ni ni with s2i the sample variance of the labels in Si . PPI and GREG use the same expression on their residuals. Every smoother is fit by Gibbs sampling with one chain, keeping 2,000 draws after a burn-in of 1,000 iterations. FH and PP-S are fit on the original scale under the flat prior on β and σ 2 . For TFH and PP-TS the inputs zi are standardized by the mean and standard deviation of the observed inputs, with di scaled to match, and each coefficient in β then has a N 0, 104 prior and each level variance σr2 an inverse-gamma prior with shape 3 and scale 0.6. The point estimate is the posterior mean, and the 95% interval runs from the 2.5% to the 97.5% posterior quantile. The compilation of Wu et al. (2026) sorts about 4,400 models by release date and takes the earlier half, about 2,200 models, as historical. All three auxiliaries are built from that half alone, and phi-4 and every model fine-tuned from it fall in the newer half. The same 100 samples serve every estimator and every auxiliary, so the rows of the estimators that use no auxiliary repeat exactly across the three auxiliary blocks. A domain whose sampled labels all agree has an estimated variance of zero and is left out of every estimator’s metrics in that replication; this affects 0.6 domains on average across the 100 replications at 10%, 0.1 at 20% and none at 30%. Tables 5, 6 and 7 report RMSE, interval score and coverage at all three sampling budgets and under all three auxiliaries, with each smoother twice, once with intercept-only linking and once with the auxiliary’s domain mean as covariate. At a matched linking model the smoother on the GREG input is never worse than the same smoother on the HT input, under any auxiliary at any budget and paired within replication. Table 5: RMSE over the 34 domains, averaged over 100 replications, at each sampling budget and under each auxiliary. Lower is better and the best value in each column is bold. Specialist model
Generalist model
Historical difficulty
Estimator
10%
20%
30%
10%
20%
30%
10%
20%
30%
θ̂iHT θ̂iPPI θ̂iGREG
0.0819 0.1056 0.0800
0.0555 0.0700 0.0541
0.0408 0.0525 0.0398
0.0819 0.1047 0.0816
0.0555 0.0702 0.0550
0.0408 0.0526 0.0406
0.0819 0.0736 0.0739
0.0555 0.0496 0.0498
0.0408 0.0359 0.0360
θ̂iFH , intercept θ̂iFH , covariate θ̂iPP-S , intercept θ̂iPP-S , covariate
0.0804 0.0776 0.0782 0.0753
0.0550 0.0539 0.0537 0.0526
0.0405 0.0400 0.0395 0.0390
0.0804 0.0724 0.0797 0.0714
0.0550 0.0522 0.0545 0.0517
0.0405 0.0392 0.0402 0.0389
0.0804 0.0675 0.0722 0.0606
0.0550 0.0497 0.0493 0.0448
0.0405 0.0382 0.0357 0.0336
θ̂iTFH , intercept θ̂iTFH , covariate θ̂iPP-TS , intercept θ̂iPP-TS , covariate
0.0781 0.0737 0.0768 0.0720
0.0539 0.0521 0.0531 0.0512
0.0402 0.0393 0.0394 0.0384
0.0781 0.0707 0.0780 0.0702
0.0539 0.0512 0.0537 0.0511
0.0402 0.0388 0.0399 0.0385
0.0781 0.0652 0.0712 0.0594
0.0539 0.0485 0.0488 0.0441
0.0402 0.0377 0.0356 0.0334
27
Table 6: Interval score of the nominal 95% intervals, averaged over 100 replications, at each sampling budget and under each auxiliary. Lower is better and the best value in each column is bold. Specialist model
Generalist model
Historical difficulty
Estimator
10%
20%
30%
10%
20%
30%
10%
20%
30%
θ̂iHT θ̂iPPI θ̂iGREG
0.4181 0.4982 0.3926
0.2723 0.3244 0.2593
0.1974 0.2441 0.1887
0.4181 0.5035 0.4093
0.2723 0.3287 0.2643
0.1974 0.2476 0.1939
0.4181 0.3583 0.3580
0.2723 0.2362 0.2363
0.1974 0.1685 0.1681
θ̂iFH , intercept θ̂iFH , covariate θ̂iPP-S , intercept θ̂iPP-S , covariate
0.3860 0.3766 0.3650 0.3567
0.2632 0.2607 0.2506 0.2493
0.1926 0.1911 0.1836 0.1823
0.3860 0.3622 0.3805 0.3551
0.2632 0.2502 0.2550 0.2444
0.1926 0.1881 0.1897 0.1846
0.3860 0.3354 0.3404 0.3008
0.2632 0.2377 0.2296 0.2129
0.1926 0.1824 0.1653 0.1588
θ̂iTFH , intercept θ̂iTFH , covariate θ̂iPP-TS , intercept θ̂iPP-TS , covariate
0.3805 0.3661 0.3642 0.3497
0.2561 0.2503 0.2470 0.2415
0.1882 0.1856 0.1820 0.1777
0.3805 0.3591 0.3770 0.3515
0.2561 0.2454 0.2503 0.2404
0.1882 0.1852 0.1863 0.1817
0.3805 0.3257 0.3388 0.2954
0.2561 0.2317 0.2250 0.2085
0.1882 0.1792 0.1628 0.1566
Table 7: Empirical coverage of the nominal 95% intervals, averaged over 100 replications, at each sampling budget and under each auxiliary. The value closest to nominal in each column is bold. Specialist model Estimator
10%
θ̂iHT θ̂iPPI θ̂iGREG
20%
30%
Generalist model 10%
10%
20%
30%
0.935 0.941 0.952 0.935 0.937 0.947 0.951 0.939 0.941 0.942 0.946 0.941
0.941 0.952 0.935 0.941 0.948 0.943 0.945 0.950 0.942
0.941 0.943 0.942
0.952 0.951 0.953
θ̂iFH , intercept θ̂iFH , covariate θ̂iPP-S , intercept θ̂iPP-S , covariate
0.941 0.942 0.946 0.945
0.937 0.938 0.942 0.943
0.951 0.941 0.951 0.940 0.946 0.942 0.948 0.943
0.937 0.940 0.947 0.945
0.951 0.951 0.951 0.952
0.941 0.937 0.941 0.940 0.943 0.945 0.944 0.945
0.951 0.950 0.954 0.955
θ̂iTFH , intercept θ̂iTFH , covariate θ̂iPP-TS , intercept θ̂iPP-TS , covariate
0.932 0.934 0.932 0.936
0.933 0.935 0.938 0.940
0.949 0.949 0.944 0.949
0.933 0.936 0.939 0.942
0.949 0.932 0.933 0.951 0.946 0.937 0.948 0.931 0.942 0.949 0.942 0.943
0.949 0.951 0.952 0.954
0.932 0.938 0.933 0.940
28
20%
30%
Historical difficulty
C
Additional Details for the Agent Traffic Analysis
C.1
The PRISM Data
We use the PRISM dataset (Kirk et al., 2024), which records human ratings on every response to an LLM prompt. The dataset is large enough to sample from, and we treat it as the traffic of a deployed AI system. In PRISM, a participant’s opening prompt is sent to up to four LLMs, and the participant rates every response and chooses one LLM to continue the conversation with. Each later turn returns two responses from the chosen LLM, which the participant again rates before choosing one. Every rated response is a unit, so one conversation contributes several units to the population. The content covariates are recorded for every response, whether or not its rating is sampled: • the log of the response’s length in characters, • an indicator that the response answers the opening prompt, • an indicator that the participant chose the response. The traffic analysis also uses the score of an LLM judge, described in the next subsection. In the traffic analysis of the Results, ratings are drawn by stratified SRS, so the chance that a response is sampled depends only on its domain. Whether the participant chose a response therefore plays no part in which ratings are sampled, and the chosen indicator enters only as a covariate. The demonstration in Section 2 instead uses informative selection, in which the chance that a response is sampled depends on something related to its rating. Participants rate the responses they choose far higher, 81.7 on average against 54.1 for the rest. Each response is sampled independently with probability proportional to exp(tz), where z is the standardized chosen indicator and t = −0.30, scaled so that the expected sample size meets the budget. A chosen response is therefore sampled at 0.54 times the rate of a rejected one, which puts the correlation between being sampled and the rating at about 0.03 in magnitude at the 5% budget. The judge is nearly blind to this choice. The judge’s scores differ by about 5 points between chosen and rejected responses, while the participants’ ratings differ by about 28.
C.2
The Judge
The traffic judge is gpt-5-nano at low reasoning effort, prompted zero-shot to predict the participant’s own 1 to 100 rating of the final response. Each conversation is submitted in full, with the response to score marked and no participant attributes included, and the judge returns a score and a confidence as structured output. The prompt was written once and frozen before any scoring: You will see a conversation between a user and an AI assistant. The conversation ends with an assistant response marked for scoring. Immediately after that response, the user rated it on a scale from 1 to 100, where higher means they were more satisfied with it. Predict the score the user gave. Judge the response from that user’s perspective in the context of the conversation, not against an external standard of quality. Reply with JSON: {"score": <integer 1-100>, "confidence": <integer 0-100>} where confidence is how sure you are of the prediction. Scoring all 68,371 responses cost about six dollars. 29
C.3
Results at the 20% Budget
Tables 8 and 9 repeat the traffic study’s comparisons at the 20% budget, 13,674 ratings, alongside the 10% columns of the main text. Domain sizes Ni run from 423 to 1,809, median 1,129. Every conclusion carries over: the oracle ordering of the candidates is unchanged, the two one-sample scores and the even split remain equivalent on the choice, and only DB-CV reports a calibrated error. Table 8 also reports each candidate’s interval score and coverage against the oracle. Coverage sits between 0.943 and 0.964 against the nominal 95% level for every candidate, and the interval-score gains track the RMSE gains. Table 8: The candidate field against the oracle at both budgets, weighted by traffic share over 100 replications: RMSE, the interval score, and the empirical coverage of the nominal 95% intervals. HT is the Horvitz–Thompson estimator. The best RMSE, the best interval score, and the coverage closest to nominal in each budget are bold. 10% candidate
RMSE
20%
IS coverage RMSE
HT 2.505 11.45 GREG (judge) 2.418 11.03 GREG (judge + content) 2.148 9.93 PP-S (judge) 2.253 10.23 PP-S (judge + content) 1.527 7.03 PP-TS (judge) 1.943 9.01 PP-TS (judge + content) 1.720 8.10
IS coverage
0.950 1.684 7.71 0.950 1.615 7.41 0.949 1.451 6.57 0.945 1.563 7.13 0.943 1.190 5.43 0.964 1.408 6.53 0.964 1.267 5.83
0.948 0.949 0.951 0.946 0.945 0.957 0.959
Table 9: The four validation procedures at both budgets, 100 replications, columns as in the main text. budget
validation procedure
RMSE of chosen
rank corr.
reported / oracle RMSE
10%
DB-CV (one sample) naive CV (one sample) 50/50 (two samples) 80/20 (two samples)
1.543 1.528 1.557 1.678
0.89 0.88 0.78 0.62
1.07× 4.00× 2.59× 3.54×
20%
DB-CV (one sample) naive CV (one sample) 50/50 (two samples) 80/20 (two samples)
1.209 1.190 1.232 1.247
0.82 0.84 0.74 0.60
1.11× 3.67× 2.37× 3.46×
30