BayesAME: Bayesian Active Model Evaluation Paula Cordero Encinar1 , Taylan Cemgil2 , Arnaud Doucet2 , Virginia Aglietti*2 and Silvia Chiappa*2
arXiv:2607.27023v1 [cs.LG] 29 Jul 2026
1 Imperial College London (work conducted while at Google DeepMind), 2 Google DeepMind
Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically targeting automatic determination of the coreset size. BayesAME models performance as a random variable by defining a latent ability for each group of items sharing the same historical model performances, with a joint prior distribution encoding the belief that the target model behaves similarly to these historical models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. Moreover, we propose a multi-target extension that actively captures performance correlations across multiple target models to further reduce the coreset size. Through extensive experiments across diverse benchmarks, we demonstrate that BayesAME consistently outperforms sequential adaptations of existing methods. Crucially, our comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection, and demonstrating that one can reliably improve upon the random sample mean baseline. Additionally, we highlight that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.
1. Introduction The massive scale and frequency of generating responses render large generative model evaluation across benchmarks a time-consuming and computationally expensive process (Liang et al., 2023). To mitigate this, a growing number of methods estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Such methods mostly require the practitioner to input a coreset size (Berrada et al., 2025; Bowyer et al., 2026; Fisch et al., 2026; Hsu and Shekhar, 2026; Huang et al., 2026; Kipnis et al., 2025; Kossen et al., 2021; Liao et al., 2025; Liu et al., 2026; Perlitz et al., 2024; Saranathan et al., 2025; Vivek et al., 2024; Zhang et al., 2025). While this is most suitable when facing strict computational limits, in scenarios where reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. In addition, the predominant approach of leveraging evaluations from historical models (reference models) has largely operated within the regime in which the models being evaluated (target models) fall within the performance range of these reference models (interpolation regime). Zhang et al. (2025) underscored the importance of also testing efficient model evaluation methods in regimes where the target models significantly deviate from this range (extrapolation regimes) to obtain a more accurate assessment of a method’s robustness. In particular, through a study that considered both regimes, the authors questioned the utility of non-random coreset selection and highlighted the difficulty of outperforming the mean of a random coreset as a performance estimate in extrapolation. In this work, we focus on efficient model evaluation leveraging reference models, specifically targeting automatic determination of the coreset size. We propose BayesAME, a sequential Bayesian method that models
*Senior co-authors. © 2026 Google DeepMind. All rights reserved
BayesAME: Bayesian Active Model Evaluation
performance as a random variable via underlying abilities whose joint prior distribution encodes the belief that the target model behaves similarly to the reference models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. In the setting where multiple target models are being evaluated, we capture correlation among them using a linear model of coregionalization (Álvarez et al., 2012; Journel and Huijbregts, 1978) with the goal of further reducing the coreset size. To validate our approach, we comprehensively evaluate BayesAME against sequential adaptations of existing methods. Our main contributions are summarized as follows: • We formulate efficient model evaluation with automatic coreset-size determination as a Bayesian sequential process that augments the coreset until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. • We develop a computationally efficient method, equipped with robust performance estimators possessing well-defined properties. • We design active selection strategies that identify benchmark items maximizing information gain. • We extend the method to capture correlations across multiple target models to further reduce the coreset size. • We conduct an extensive empirical analysis across varying numbers of reference models, interpolation/extrapolation regimes, and different degrees of target model correlation, showing that BayesAME outperforms sequential adaptations of existing methods designed to automatically determine the coreset size. We further show that BayesAME remains advantageous even when the coreset size is predetermined. • Crucially, this analysis overturns recent claims in the literature by demonstrating that non-random coreset selection does improve performance and that several methods can reliably outperform the random sample mean baseline. • We demonstrate that leveraging available response log-likelihoods, rather than using binary scores, greatly improves performance estimation.
2. Problem Setup Consider a benchmark 𝐸 = {( 𝑥 𝑖 , 𝑦𝑖 )} 𝑖𝑁=1 , where 𝑥 𝑖 is the 𝑖-th input instance and 𝑦𝑖 is the corresponding groundtruth response. Let T = {𝑡1 , . . . , 𝑡 𝐾𝑡 } be a set of 𝐾𝑡 target models that we wish to evaluate, and let 𝑠𝑡𝑖 be the score of target model 𝑡 ∈ T on item ( 𝑥 𝑖 , 𝑦𝑖 ). We consider the problem of estimating the full benchmark performance Í 𝑅★𝑡 = 𝑁1 𝑖𝑁=1 𝑠𝑡𝑖 by evaluating the target model 𝑡 only on a subset of items C 𝑡 ⊆ 𝐸, which we refer to as a coreset. We focus on the setting where the full benchmark scores for a set of 𝐾𝑟 reference models {𝑟1 , . . . , 𝑟 𝐾𝑟 } are available. Although our framework is applicable to scores 𝑠𝑡𝑖 ∈ ℝ, in our experiments we consider either binary scores (𝑠𝑡𝑖 ∈ {0, 1}), indicating whether the model’s response matches 𝑦𝑖 , or bounded continuous scores, where, e.g., 𝑠𝑡𝑖 ∈ [0, 1] represents the probability assigned by the model to the ground-truth response 𝑦𝑖 . Our goal is to address the setting in which the user prioritizes reliable performance estimation over minimizing cost, without a priori knowing the coreset size required to achieve this. Therefore, we focus on methods that not only operate with a user-specified coreset size, but can also automatically determine a coreset size that reflects this priority. In the remainder of the paper, with a slight abuse of notation, we use 𝐸 and C 𝑡 to indicate the indices of the benchmark items rather than the items themselves.
3. BayesAME BayesAME is a sequential Bayesian method that assigns a latent ability to each group of benchmark items sharing the same reference model performances. It then treats the item scores as noisy versions of these 2
BayesAME: Bayesian Active Model Evaluation
abilities and models the full benchmark performance as a random variable defined as the average of the scores, Í 𝑅𝑡 = 𝑁1 𝑖𝑁=1 𝑠𝑡𝑖 (where, with a slight abuse of notation, we use 𝑠𝑡𝑖 to denote both the score random variable and its realization). The abilities are assigned a joint Gaussian prior distribution to encode the belief that, if the reference models achieved similar scores on two groups of items, the target model likely exhibits similar behavior. At each iteration, the posterior distribution of the latent abilities is used to derive performance estimators, quantify performance uncertainty, and select an item to add to the coreset via an information gain criterion. The sequential process terminates when the performance fluctuation and uncertainty fall below their respective user-defined thresholds.
3.1. Single-Target Setting We first consider the setting of a single target model. For the remainder of this section, we omit the target 𝑟 model index. Let ( 𝑠𝑟𝑖 1 , . . . , 𝑠𝑖 𝐾𝑟 ) be the vector of scores achieved by the 𝐾𝑟 reference models on item ( 𝑥 𝑖 , 𝑦𝑖 ). Suppose that the reference models yield 𝐵 unique score vectors across the 𝑁 benchmark items, which we denote as the set {𝜙1 , . . . , 𝜙𝐵 }. This partitions the items into 𝐵 disjoint buckets, such that all items assigned to bucket 𝑏 share the same reference score vector 𝜙𝑏 . We consider a random vector 𝜃 = ( 𝜃1 , . . . , 𝜃𝐵 ), where 𝜃𝑏 represents the target model’s underlying ability on items belonging to bucket 𝑏, with a Gaussian prior distribution 𝑝 ( 𝜃) = N ( 𝜇 𝜃 , Σ𝜃 ), defined as: 1 ⊤ 𝛽 𝜃 2 𝜙𝑏 1 𝐾𝑟 , Σ𝑏,𝑏 , (1) 𝜇 𝜃𝑏 = ′ = 𝛼 exp − 𝐾 ∥ 𝜙𝑏 − 𝜙𝑏′ ∥ 𝑟 𝐾𝑟
where 1 𝐾𝑟 denotes a 𝐾𝑟 -dimensional column vector of ones, 𝛼 and 𝛽 are parameters to be learned, and ∥ 𝜙𝑏 − 𝜙𝑏′ ∥ 2 is the squared Euclidean distance. This form of the covariance encodes the belief that buckets with similar reference scores yield highly correlated latent abilities for the target model1 . Applying this latent variable approach at the bucket level, rather than the item level, reduces overreliance on prior information when reference models behave similarly and thus a large number of items map to the same 𝜙𝑏 , which most commonly manifests when 𝐾𝑟 is small or when the scores are binary. For any item ( 𝑥 𝑖 , 𝑦𝑖 ) belonging to bucket 𝑏, we model the target model’s score 𝑠𝑖 as a random variable with conditional distribution 𝑝 ( 𝑠𝑖 | 𝜃𝑏 ) = N ( 𝜃𝑏 , 𝑁𝑏 𝜎2 ), where 𝑁𝑏 is the number of items in bucket 𝑏. Let 𝐻 denote an 𝑁 × 𝐵 binary indicator matrix, where 𝐻𝑖,𝑏 = 1 if the 𝑖-th item belongs to bucket 𝑏, and 0 otherwise. The target model’s score vector 𝑆 = ( 𝑠1 , . . . , 𝑠𝑁 ) is then modeled as a random variable with Í conditional distribution 𝑝 ( 𝑆 | 𝜃) = N ( 𝐻𝜃, 𝐷), where 𝐷 is an 𝑁 × 𝑁 diagonal noise matrix such that 𝐷𝑖,𝑖 = 𝜎2 𝑏𝐵=1 𝐻𝑖,𝑏 𝑁𝑏 . Note that, when all items map to unique reference vectors ( 𝐵 = 𝑁 ), 𝐷𝑖,𝑖 = 𝜎2 . This bucket-size-dependent noise accounts for the inherent heterogeneity within larger buckets. Even though items in a large bucket share identical reference scores, the target model’s performance on each individual item might vary. By scaling the individual item variance by 𝑁𝑏 , we ensure that the variance score remains constant at 𝜎2 , regardless of ofÍthe average bucket 1 1 Í how many items the bucket contains: Var 𝑁𝑏 𝑖 ∈bucket 𝑏 𝑠𝑖 = 𝑁 2 𝑖 ∈bucket 𝑏 Var( 𝑠𝑖 ) = 𝑁12 ( 𝑁𝑏 · 𝑁𝑏 𝜎2 ) = 𝜎2 . While 𝑏
𝑏
a more principled approach to modeling binary and bounded continuous scores would explicitly model their support, exploring this direction revealed that the approximations required not only introduced substantial computational overhead but also resulted in inferior performance (see Appendix B.3). In contrast, the Gaussian model admits closed-form posterior distributions, delivering superior scalability and performance. We therefore adopt a Gaussian model as the default choice. Let 𝑠 C denote the subvector of 𝑠 corresponding to the items in C, and let 𝜇 𝜃| C and Σ𝜃| C be the posterior mean and covariance of 𝜃 conditioned on 𝑠 C , such that 𝑝 ( 𝜃 | 𝑠 C ) = N ( 𝜇 𝜃| C , Σ𝜃| C ). The posterior mean and variance of Í the full benchmark performance 𝑅 = 𝑁1 𝑖𝑁=1 𝑠𝑖 are given by: ! 𝐵 ∑︁ ∑︁ 1 ©∑︁ ª 1 ∑︁ 𝜃 𝔼[ 𝑅 | 𝑠 C ] = 𝑠𝑖 + 𝔼[ 𝑠 𝑗 | 𝑠 C ] ® = 𝑠𝑖 + ( 𝑁𝑏 − 𝑛𝑏 ) 𝜇 𝑏 | C , 𝑁 𝑁 𝑖 ∈ C 𝑖 ∈ C 𝑏 =1 𝑗 ∈ 𝐸 \C « ¬ 1When given metadata for the reference and target models, we could employ a weighted mean and covariance, where the weights capture the similarity between the target and reference models, see Appendix B.2.
3
BayesAME: Bayesian Active Model Evaluation
Var( 𝑅 | 𝑠 C ) =
1 ∑︁ 𝑁2
𝐵
∑︁
Cov( 𝑠 𝑗 , 𝑠 𝑗′ | 𝑠 C ) =
𝑗 ∈ 𝐸\C 𝑗′ ∈ 𝐸\C
1 ∑︁ 𝑁 2 𝑏=1
2
( 𝑁𝑏 − 𝑛𝑏 ) 𝜎 𝑁𝑏 +
𝐵 ∑︁
! 𝜃 ( 𝑁𝑏′ − 𝑛𝑏′ ) Σ𝑏,𝑏 ′|C
,
(2)
𝑏′ =1
where 𝑛𝑏 is the number of evaluated items in bucket 𝑏. As more items are evaluated, the posterior mean increasingly relies on the actual scores, converging to the true performance 𝑅★ when the entire benchmark has been evaluated (𝑛𝑏 = 𝑁𝑏 for all 𝑏), at which point the posterior variance vanishes.
Performance Estimator We construct an estimator 𝑅ˆ of 𝑅★ by adapting our approach to the number of unique reference score vectors in the benchmark, introducing a small margin 𝑁0 > 0 to allow for slight deviations from uniqueness. In the nearly-unique reference regime 𝐵 ≥ 𝑁 − 𝑁0 , the natural and optimal performance estimator is the Bayesian posterior expectation: ! 𝐵 ∑︁ 1 ∑︁ 𝜃 ˆBayes = 𝔼[ 𝑅 | 𝑠 C ] = (3) 𝑠𝑖 + ( 𝑁𝑏 − 𝑛𝑏 ) 𝜇 𝑏 | C . 𝑅 𝑁
𝑖∈ C
𝑏=1
However, using Eq. (3) in the non-unique reference regime 𝐵 < 𝑁 − 𝑁0 presents a potential issue: if a heavily populated bucket 𝑏 has a large number of unevaluated items (𝑛𝑏 ≪ 𝑁𝑏 ), its posterior mean 𝜇 𝜃𝑏| C exerts a high weight on the overall estimate, amplifying any estimation errors. In this regime, we instead use the following estimator: 𝐵 𝐵 ∑︁ ∑︁ 𝑁𝑏 𝑁𝑏 1 ∑︁ 𝑛𝑏 1 ∑︁ 𝑛𝑏 ˆCV = 𝑠𝑖 + 𝑠𝑖 + 𝑅 − 𝔼[ 𝜃𝑏 | 𝑠 C ] = − 𝜇 𝜃𝑏| C , (4) 𝑛
𝑖∈ C
𝑏=1
𝑁
𝑛
𝑛
𝑖∈ C
𝑏=1
𝑁
𝑛
where 𝑛 indicates the size of the coreset C. This estimator serves as an in-sample proxy for the unbiased control variate estimator given in Appendix A.2. We can unify both regimes into a single estimator: ˆ=𝑤 𝑅
∑︁ 𝑖∈ C
𝑠𝑖 +
𝐵 ∑︁ 𝑁𝑏 𝑏=1
𝑁
− 𝑛𝑏 𝑤 𝜇 𝜃𝑏| C ,
(5)
where 𝑤 = 1𝑛 if 𝐵 < 𝑁 − 𝑁0 and 𝑤 = 𝑁1 if 𝐵 ≥ 𝑁 − 𝑁0 .
Selection Strategy At each iteration of the sequential process, we use an information-gain criterion to select the next item on which to evaluate the target model. Specifically, we choose an item that maximally reduces the expected uncertainty of the model’s performance 𝑅, measured via differential entropy: 𝛼IG ( 𝑖) = H ( 𝑅 | 𝑠 C ) − 𝔼𝑠𝑖 [H ( 𝑅 | 𝑠 C , 𝑠𝑖 )]. Assuming candidate item ( 𝑥 𝑖 , 𝑦𝑖 ) belongs to bucket 𝑏, due to the Gaussian nature of the posterior distribution, this expected information gain simplifies to: 2 Í𝐵 𝜃 2 ′ ′ 1 Var( 𝑅 | 𝑠 C ) 1 𝑏′ =1 ( 𝑁𝑏 − 𝑛𝑏 ) Σ𝑏,𝑏′ | C + 𝜎 𝑁𝑏 © ª 𝛼IG ( 𝑖) = log = − log 1 − ®. 𝜃 2𝑁 ) 2 Var( 𝑅 | 𝑠 C , 𝑠𝑖 ) 2 𝑁 2 Var( 𝑅 | 𝑠 C ) ( Σ𝑏,𝑏 + 𝜎 𝑏 | C « ¬
(6)
Note that all items in the same bucket have the same expected information-gain value. In the non-unique reference regime with imbalanced bucket sizes, purely active selection can degrade performance. This occurs because the target model may exhibit fine-grained variations in behavior that are not captured by the coarse grouping of the reference models. Consequently, when 𝐵 < 𝑁 − 𝑁0 , we default to random selection, which ensures that the expected number of sampled items from each bucket remains proportional to its size. The margin 𝑁0 > 0 allows the method to use active sampling whenever the reference score vectors are nearly unique. Furthermore, when dealing with continuous scores, unique reference score vectors may become non-unique when rounded to a specified tolerance. Under such conditions, we similarly default to random selection. We also explored alternative selection strategies that operate on batches of items rather than individual items, including non-myopic approaches. However, these provided similar performance (see Appendix B.1).
4
BayesAME: Bayesian Active Model Evaluation
Stopping Criterion One of our key objectives is to determine the coreset size, rather than requiring it as input from the user. Intuitively, we would like to stop evaluating additional items once they are unlikely to meaningfully change our estimate of the target model’s performance. To formalize this, we monitor the uncertainty of 𝑅 as items are evaluated. We √︁ quantify this uncertainty using the width of the 95% credible interval of 𝑅: 𝑞97.5 ( 𝑅 | 𝑠 C ) − 𝑞2.5 ( 𝑅 | 𝑠 C ) = 2𝑧0.975 Var( 𝑅 | 𝑠 C ), where 𝑧0.975 denotes the standard normal 𝑧-score and Var( 𝑅 | 𝑠 C ) is given in Eq. (2). We dynamically augment the coreset until a ˆ two-condition stopping criterion is satisfied: ( 𝑖) the performance estimate 𝑅 √︁ stabilizes, defined as its range over a window of 𝑊 iterations falling below a threshold 𝜖1 ; and ( 𝑖𝑖) 2 𝑧0.975 Var( 𝑅 | 𝑠 C ) ≤ 𝜖2 . The threshold 𝜖2 represents the maximum acceptable width of this uncertainty interval, translating directly into a user’s tolerated margin of error. By treating 𝜖1 and 𝜖2 as user-defined parameters, BayesAME enables the user to decide the precision required. For additional discussion, see Appendix A.3.
Posterior Distribution To optimize the cost required to compute 𝜇 𝜃| C and Σ𝜃| C , throughout the sequential process we adapt our formulation as detailed below. Let 𝐻 C and 𝐷 C denote the submatrices of 𝐻 and 𝐷 formed by the rows corresponding to C. Let C𝑏 denote the subset of 𝑛𝑏 items in the coreset C that are in bucket 𝑏. Í Finally, let ¯𝑠 C1:𝐵 = (¯𝑠 C1 , . . . , ¯𝑠 C𝐵 ), where ¯𝑠 C𝑏 = 𝑛1𝑏 𝑖 ∈ C𝑏 𝑠𝑖 (for notational simplicity we assume 𝑛𝑏 > 0 for all 𝑏). As detailed in Appendix A.1, the posterior distribution of 𝜃 satisfies 𝑝 ( 𝜃 | 𝑠 C ) = 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ). Although mathematically equivalent, computing 𝑝 ( 𝜃 | 𝑠 C ) requires inverting an 𝑛 × 𝑛 matrix, whereas computing 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ) requires inverting at most a 𝐵 × 𝐵 matrix. Indeed, 𝑝 ( 𝜃 | 𝑠 C ) is Gaussian with mean 𝜇 𝜃| C and covariance Σ𝜃| C given by: 𝜃 ⊤ −1 𝜃 𝜇 𝜃| C = 𝜇 𝜃 + Σ𝜃 𝐻 ⊤ C ( 𝐻C Σ 𝐻C + 𝐷C ) (𝑠C − 𝐻C 𝜇 ), 𝜃 ⊤ −1 𝜃 Σ𝜃| C = Σ𝜃 − Σ𝜃 𝐻 ⊤ C ( 𝐻C Σ 𝐻C + 𝐷C ) 𝐻C Σ ,
(7)
while, starting from 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ), the mean 𝜇 𝜃| C and covariance Σ𝜃| C can be expressed as: 𝜇 𝜃| C = 𝜇 𝜃 + Σ𝜃 ( Σ𝜃 + Δ) −1 (¯𝑠 C1:𝐵 − 𝜇 𝜃 ) , Σ𝜃| C = Σ𝜃 − Σ𝜃 ( Σ𝜃 + Δ) −1 Σ𝜃 ,
(8)
where Δ is a diagonal noise matrix with entries Δ𝑏,𝑏 = 𝜎2 𝑁𝑏 /𝑛𝑏 . To achieve the best computational cost, we use Eq. (7) when 𝑛 < 𝐵 and Eq. (8) when 𝑛 ≥ 𝐵. We employ Cholesky updates across both regimes, leveraging the fact that the prior covariance Σ𝜃 remains fixed between hyperparameter updates, which occur every 𝐹 iterations. This yields a computational complexity of O (min( 𝑛, 𝐵) 2 ) per iteration.
Algorithm We summarize the method in Algorithm 1. Starting with an empty coreset C, BayesAME iteratively selects the benchmark item maximizing the expected information gain 𝛼IG ( 𝑖) when 𝐵 ≥ 𝑁 − 𝑁0 (or at random otherwise). The target model is evaluated on this item, its index is added to C, and the posterior mean 𝜇 𝜃| C and covariance Σ𝜃| C are updated, along with the performance estimate 𝑅ˆ. Every 𝐹 iterations, the covariance hyperparameters 𝛼 and 𝛽 are optimized by minimizing the negative log-marginal likelihood L ( 𝛼, 𝛽 ) of the coreset scores (see Appendix A.4). This sequential process terminates once the two-condition stopping criterion is satisfied: ( 𝑖) the range of the performance estimate 𝑅ˆ over a window of 𝑊 iterations falls below threshold 𝜖1 ; and ( 𝑖𝑖) the width of the 95% credible interval falls below threshold 𝜖2 .
3.2. Multi-Target Setting In the case of multiple target models T = {𝑡1 , . . . , 𝑡 𝐾𝑡 }, a standard approach would be to simply apply the single-target method of Section 3.1 to each model independently. However, jointly modeling target models that exhibit correlated performance behaviors can improve estimation accuracy and reduce the required coreset size. We model the targets jointly using a linear model of coregionalization (Álvarez et al., 2012; Journel and Huijbregts, 1978).
5
BayesAME: Bayesian Active Model Evaluation
Algorithm 1 BayesAME Single-target Require: Benchmark 𝐸; noise variance 𝜎2 ; margin 𝑁0 ; stopping criterion thresholds 𝜖1 , 𝜖2 and window-length 𝑊 ; Optimizer, hyperparameter initial values 𝛼0 , 𝛽0 and update frequency 𝐹 Initialize coreset C ← ∅ Initialize hyperparameters 𝛼 ← 𝛼0 , 𝛽 ← 𝛽0 for 𝑗 ∈ {1, . . . , 𝑁 } do if 𝐵 ≥ 𝑁 − 𝑁0 then Select item via expected information gain: 𝑖∗ = arg max𝑖 ∈ 𝐸\C 𝛼IG ( 𝑖) ⊲ Eq. (6) else Select item via random sampling: 𝑖∗ ∼ Uniform( 𝐸 \ C) end if Obtain score 𝑠𝑖∗ and add 𝑖∗ to the coreset: C ← C ∪ { 𝑖∗ } Compute 𝜇 𝜃| C and Σ𝜃| C ⊲ Eq. (7) or Eq. (8) Compute performance estimate 𝑅ˆ ⊲ Eq. (5) Set 𝑅ˆ 𝑗 ← 𝑅ˆ if 𝑗%𝐹 = 0 then Update hyperparameters: 𝛼, 𝛽 ← Optimizer(L ( 𝛼, 𝛽 ) , 𝛼, 𝛽 ) ⊲ Eq. (11) end if √︁ if 𝑗 ≥ 𝑊 and max𝑙 ∈ [ 𝑗 −𝑊 +1, 𝑗 ] ( 𝑅ˆ𝑙 ) − min𝑙 ∈ [ 𝑗 −𝑊 +1, 𝑗 ] ( 𝑅ˆ𝑙 ) ≤ 𝜖1 and 2 𝑧0.975 Var( 𝑅 | 𝑠 C ) ≤ 𝜖2 then break end if end for Return 𝑅ˆ We capture cross-model correlations by assuming that their abilities are driven by 𝐿 shared, independent latent random vectors 𝑢𝑙 ∈ ℝ 𝐵 . Specifically, let 𝜃𝑏𝑡 represent the latent ability of target model 𝑡 on bucket 𝑏, and let 𝜃𝑡 = ( 𝜃1𝑡 , . . . , 𝜃𝑡𝐵 ). We model 𝜃𝑡 as: 𝐿 ∑︁ 𝑡 𝜃𝑡 = 𝜇 𝜃 + 𝑤𝑙𝑡 𝑢𝑙 , 𝑙 =1 𝑡
where 𝜇 𝜃 is defined as in Eq. (1), 𝑤1𝑡 , . . . , 𝑤𝑡𝐿 are model-specific weights to be learned, and 𝑢𝑙 ∼ N (0, Σ𝜃,𝑙 ), where Σ𝜃,𝑙 is a covariance matrix specific to latent vector 𝑙. This induces a Gaussian prior distribution over 𝜃 = ( 𝜃𝑡1 , . . . , 𝜃𝑡 𝐾𝑡 ) given by: ! 𝐿 ∑︁ 𝜃 ⊤ 𝜃,𝑙 𝑝 ( 𝜃) = N 𝜇 , 𝑤𝑙 𝑤𝑙 ⊗ Σ , 𝑙 =1 𝑡 ( 𝑤𝑙𝑡1 , . . . , 𝑤𝑙 𝐾𝑡 ) ⊤
𝑡
𝑡𝐾
where ⊗ denotes the Kronecker product, 𝑤𝑙 = and 𝜇 𝜃 = ( 𝜇 𝜃 1 , . . . , 𝜇 𝜃 𝑡 ). This formulation provides a flexible yet computationally efficient way to capture correlation across both buckets and 𝜃,𝑙 𝑡 models by combining bucket-specific (Σ𝑏,𝑏 ′ ) and model-specific (𝑤𝑙 ) terms. To provide greater modeling flexibility, we use two covariance families with complementary inductive biases: the squared exponential covariance as in Eq. (1), which imposes strict smoothness, and a Matérn-3/2 kernel of the form defined √ √ 𝜃,𝑙 Σ𝑏,𝑏′ = 𝛼𝑙 1 + 3 𝛽𝑙 ∥ 𝜙𝑏 − 𝜙𝑏′ ∥ exp − 3 𝛽𝑙 ∥ 𝜙𝑏 − 𝜙𝑏′ ∥ with 𝛼𝑙 , 𝛽𝑙 > 0, which imposes weaker smoothness. Because the degree of correlation between the target models is typically unknown a priori, we set 𝐿 = 𝐾𝑡 . This enables the framework to dynamically adapt: it captures strong correlations by sharing latent components across models, yet it retains the ability to recover the fully independent case. The scores are then modeled similarly to the single-target setting described in Section 3.1.
Selection Strategy In contrast to the single-target setting, the multi-target selection step requires simultaneously selecting an item and a target model to evaluate on that item, substantially expanding the search space. Furthermore, relying on a purely greedy information-gain criterion across this joint item-model space can bias performance estimates, particularly during early stages when correlations across target models are still being learned. To mitigate this bias and establish a robust initial estimate of the correlation structure, we depart
6
BayesAME: Bayesian Active Model Evaluation
from the empty coreset initialization used in the single-target setting. Instead, we randomly select 10% of the benchmark items and evaluate all target models on them. After this initialization step, we deploy the following selection strategy, which maintains some degree of exploration. We iterate deterministically through the target models. For a given model 𝑡 , we select the item that maximizes 𝑡 the expected information gain: 𝛼𝑡IG ( 𝑖) = H ( 𝑅𝑡 | 𝑠 C ) − 𝔼𝑠𝑡𝑖 [H ( 𝑅𝑡 | 𝑠 C , 𝑠𝑡𝑖 )], where 𝑠 C = { 𝑠𝑡C1𝑡1 , . . . , 𝑠 C𝐾𝑡𝑡𝐾𝑡 }. This approach substantially reduces the search space, maintaining a selection complexity identical to that of the single-target setting. We iteratively augment the coresets {C 𝑡1 , . . . , C 𝑡 𝐾𝑡 } until all target models have met the two-condition stopping criterion detailed above. We also investigated an alternative average marginalized information-gain strategy. At each iteration, we first select the item that maximizes the expected information gain averaged across all 𝐾𝑡 target models: Í 𝐾𝑡 H ( 𝑅𝑡 | 𝑠 C ) − 𝔼𝑠𝑡𝑖 [H ( 𝑅𝑡 | 𝑠 C , 𝑠𝑡𝑖 )] . After selecting an item, we choose the target model uniformly 𝛼IG ( 𝑖) = 𝐾1𝑡 𝑡=1 at random. We present these results in Figure 15 of Appendix C.4.
4. Related work The issue of computational inefficiency in model evaluation is primarily tackled from two, sometimes intersecting, directions: constructing smaller, static versions of existing benchmarks (or developing methods for reducing benchmarks) (Bean et al., 2025; Kipnis et al., 2025; Zhao et al., 2025), and estimating performance without fully evaluating the target model (Berrada et al., 2025; Bowyer et al., 2026; Fisch et al., 2026; Hsu and Shekhar, 2026; Huang et al., 2026; Kipnis et al., 2025; Kossen et al., 2021; Liao et al., 2025; Liu et al., 2026; Polo et al., 2024; Saranathan et al., 2025; Vivek et al., 2024; Wang et al., 2025; Zhang et al., 2025).
Automatic Determination of Coreset Size With the exception of Cer-Eval (Wang et al., 2025), all methods in the second category assume that the evaluation budget is specified by the user. Cer-Eval differs substantially from our method as it does not utilize reference models, relying instead exclusively on the responses from the target model. It repeats the following steps until a strict confidence condition is met: ( 𝑖) partitioning the benchmark based on previously evaluated items to isolate regions of low variance in the target model’s scores, ( 𝑖𝑖) computing summary statistics for each partition, ( 𝑖𝑖𝑖) identifying the partition that yields the greatest expected reduction in evaluation uncertainty; and ( 𝑖𝑣) sampling an item from that optimal partition. Extending methods designed for user-specified coreset sizes to automatically determine the coreset size poses fundamental challenges. This task requires striving for accurate estimation across all coreset sizes, which, as we show in our experiments, is not always achievable. Furthermore, establishing a reliable stopping criterion requires confidence intervals that accurately reflect the true remaining uncertainty, a property that existing formulations often lack.
Performance Modeling and Estimation From a modeling perspective, many works that leverage reference model information rely on Item Response Theory (IRT) to infer the target model’s profile (Kipnis et al., 2025; Liao et al., 2025; Polo et al., 2024). However, obtaining reliable parameter estimates with IRT requires on the order of hundreds (Jiang et al., 2026) or thousands (Kipnis et al., 2025) of reference models. Another line of work (Bowyer et al., 2026; Zhang et al., 2025) frames the problem as a regression task. ProEval (Huang et al., 2026) introduces a Gaussian process approach that shares similarities with our Bayesian modeling. However, we depart critically from their formulation to address automatic coreset size determination and more robustly handle the lack of prior information about the target model.
Multi-Target Setting Finally, to the best of our knowledge, the only other work that leverages performance correlations across multiple target models rather than treating them independently is Fisch et al. (2026). This approach frames model evaluation as a matrix completion problem. While the method could be run
7
BayesAME: Bayesian Active Model Evaluation
sequentially by randomly sampling the next item and model to evaluate, its confidence intervals capture only the sampling uncertainty around the mean performance, assuming the benchmark is a finite sample drawn from an infinite population of prompts. Consequently, even when all items in the finite benchmark have been evaluated, the confidence interval does not shrink to zero. In contrast, our Bayesian credible intervals provide a direct probability statement about the true, finite-sample benchmark performance given the observed data, yielding an intuitively appealing and actionable stopping criterion.
5. Experiments We evaluate our approach across seven standard benchmarks collected from three primary sources: the Open LLM Leaderboard (Fourrier et al., 2024), evaluations compiled by Zhang et al. (2025), and the HELM Lite benchmark (Liang et al., 2023). From the Open LLM Leaderboard, we extracted scores for GPQA (Rein et al., 2024), MMLU-Pro (Hendrycks et al., 2021; Wang et al., 2024), BBH (Suzgun et al., 2023), ARC-Challenge (Clark et al., 2018), and MuSR (Sprague et al., 2024). Beyond standard binary scores, we leveraged the leaderboard’s raw log-likelihoods to obtain continuous scores via a softmax transformation. These continuous scores allow us to assess the impact of exploiting richer information and mitigate the issue of identical reference model scores. To ensure a representative distribution of model capabilities, we partitioned the available models into five quantiles based on accuracy. We then sampled evenly by selecting the most popular models from each quantile, resulting in an initial set of 300 models. After extracting the data for these selections, we filtered out any records containing NaN values caused by incomplete leaderboard entries, formatting discrepancies, or evaluation timeouts. This curation process yielded the following final model counts per benchmark: GPQA: 102, MMLU-Pro: 165, BBH: 279, ARC-Challenge: 125, and MuSR: 288. From Zhang et al. (2025), we extracted binary scores for IFEval (Zhou et al., 2023) across 448 models. Finally, from HELM Lite (up to version 1.13.0), we extracted the F1 scores for the Natural QA Openbook long-answer scenario (Kwiatkowski et al., 2019) across 44 models.
5.1. Single-Target Setting Baselines Since all existing approaches that leverage reference models assume a predefined coreset size, we adapt the most effective methods to operate in the same sequential manner as √ Algorithm 1. Specifically, C is iteratively augmented until max𝑙 ∈ [ 𝑗 −𝑊, 𝑗 ] ( 𝑅ˆ𝑙 ) − min𝑙 ∈ [ 𝑗 −𝑊, 𝑗 ] ( 𝑅ˆ𝑙 ) ≤ 𝜖1 and 2 𝑧0.975 Var𝑅 ≤ 𝜖2 where 𝑅ˆ and Var𝑅 are defined below. 1. BayesAME-RS: A simpler version of BayesAME that selects items at random. 2. Seq-APW: An extension of the anchor points weighted method from Vivek et al. (2024), in which C is augmented by choosing items that maximize space covering, as measured by correlations between the reference models’ performances. Benchmark items are partitioned into clusters by assigning each item to Í its nearest anchor in C. 𝑅ˆ = 𝑖 ∈ C 𝑤𝑖 𝑠𝑖 , where 𝑤𝑖 is a weight proportional to the cluster size, while Var𝑅 is given by the following estimator of Var( 𝑅): ! 2 Í ∑︁ 𝑖 ∈ C 𝑤𝑖 2 2 𝑁 − 𝑛eff Var𝑅 = 𝑠 𝑤𝑖 , 𝑛eff = Í , (9) 2 𝑖∈ C
𝑁
𝑖 ∈ C 𝑤𝑖
where 𝑠 denotes the sample standard deviation. By construction, 𝑅ˆ recovers the true performance 𝑅★ once the entire benchmark has been evaluated. Moreover, Var𝑅 captures the sampling variance arising from the randomness in the coreset selection. Due to the finite-population correction factor, this variance decreases as the coreset grows and converges to zero when the full benchmark is evaluated. 3. RS-Mean: The simplest baseline that augments C by selecting items at random and estimates the full benchmark performance as the empirical average performance on C. Var𝑅 is given by Eq. (9) using ˆ and Var𝑅 discussed for Seq-APW apply. 𝑤𝑖 = 1/𝑛. The same properties regarding 𝑅 8
BayesAME: Bayesian Active Model Evaluation
4. Bayes-AIPW: A Bayesian ridge regression extension of AIPW (Zhang et al., 2025), obtained by placing 𝑟 𝑟 a Gaussian prior distribution on the coefficients 𝛽 of the model 𝑠𝑖 = 𝜙⊤ 𝛽 + 𝜀𝑖 , where 𝜙𝑖 = ( 𝑠𝑖 1 , . . . , 𝑠𝑖 𝐾𝑟 ), 𝑖 and 𝜀𝑖 ∼ N (0, 𝛼−1 ) with 𝛼 ∼ Gamma( 𝛼1 , 𝛼2 ). By conditioning on {( 𝜙𝑖 , 𝑠𝑖 )} 𝑖 ∈ C , we obtain a posterior distribution over 𝛽 with mean 𝜇 𝛽| C and covariance Σ |𝛽C . As 𝑅ˆ, we use the AIPW estimator: ! ∑︁ 1 1 1 ∑︁ ˆ= 𝑅 𝑠𝑖 + 𝜇𝑖| C − 𝜇𝑖| C , 𝑛 1 + 𝑁 𝑛− 𝑛 𝑁 − 𝑛 𝑛 𝑖∈ C 𝑖∈ C 1 ∑︁
𝑖 ∈ 𝐸\C
𝛽 ˆ recovers the true performance 𝑅★ when all items have been evaluated. where 𝜇 𝑖 | C = 𝜙⊤ 𝜇 | C . Note that 𝑅 𝑖 Posterior uncertainty is obtained by propagating the posterior distribution of 𝛽 through the AIPW correction, yielding: ( 𝜙¯𝐸\C − 𝜙¯ C ) ⊤ Σ |𝛽C ( 𝜙¯𝐸\C − 𝜙¯ C ) + 𝛼 ˆ −1 1𝑛 + 𝑁 1− 𝑛 Var𝑅 = , 2 1 + 𝑁 𝑛− 𝑛 Í Í where 𝜙¯ C = 1𝑛 𝑖 ∈ C 𝜙𝑖 , 𝜙¯𝐸\C = 𝑁 1− 𝑛 𝑖 ∈ 𝐸\C 𝜙𝑖 , and 𝛼 ˆ is the posterior mean of 𝛼. C is augmented by selecting items at random. Figure 6 in Appendix C.3 shows that this Bayesian extension matches or exceeds the performance of the original AIPW method. 5. Bayes-RS-Learn: A Bayesian ridge regression extension of Random-Sampling-Learn (Zhang et al., 2025), obtained by placing a Gaussian prior on the coefficients 𝛽 of the model 𝑅 = 𝑠⊤C 𝛽 + 𝜀, where 𝑟 𝐾𝑟 𝜀 ∼ N (0, 𝛼 −1 ), 𝛼 ∼ Gamma( 𝛼1 , 𝛼2 ). We perform inference on 𝛽 by conditioning on {( 𝑠 C𝑘 , 𝑅★𝑟𝑘 )} 𝑘=1 . This 𝛽 𝛽 yields a posterior distribution over 𝛽 with mean 𝜇 | C and covariance Σ | C . We then obtain 𝑅ˆ and Var𝑅 as:
ˆ = 𝑠⊤C 𝜇 , 𝑅 |C 𝛽
Var𝑅 = 𝑠⊤C Σ |𝛽C 𝑠 C + 𝛼 ˆ −1 ,
where 𝛼 ˆ −1 is the posterior mean of 𝛼−1 . 6. ProEval: The method in Huang et al. (2026) for the setting with no prior performance data for the target model (called New Model setting with Score Features in the original paper), without abstention Í rule. The method defines the target model’s performance as 𝑅 = 𝑁1 𝑖𝑁=1 𝜃𝑖 , where 𝜃 = ( 𝜃1 , . . . , 𝜃𝑁 ) is a Gaussian random vector whose prior covariance matrix is given by the empirical covariance of the reference model scores. 𝑅ˆ = 𝔼[ 𝑅 | 𝑠 C ] and Var𝑅 = Var( 𝑅 | 𝑠 C ). Because 𝑅 is defined as the average of the latent abilities, the estimator 𝑅ˆ can deviate significantly from the true performance 𝑅★ even for large coreset sizes, as discussed in Section 4. The confidence interval induced by Var𝑅 can be interpreted as a Bayesian credible interval. 7. IRT: Two variants of IRT where C is selected randomly, depending on the score type: one for binary scores based on the 2-PL model (Polo et al., 2024) (p-IRT) and one for continuous scores based on the Beta distribution, specifically the extension of Noel and Dauvier (2007) provided in Chen et al. (2019). The item parameters are learned using the reference models’ scores, and the scores 𝑠 C are used to fit the target model-specific parameters. As the performance estimator, we deploy Eq. (5) with 𝑤 = 1/ 𝑁 . Because we fit the model parameters separately for each benchmark, treating them as scalars rather than vector-valued as in Polo et al. (2024), applying the generalized variant (gp-IRT) does not improve performance. We fit the parameters using MCMC. This provides samples from the posterior distribution of the parameters, allowing us to directly obtain the posterior distribution of 𝑅 and its variance, which we use as 𝑅ˆ and Var𝑅 . Due to the computational cost of MCMC, we update the parameters only every 50 evaluations (resulting in piecewise-constant curves in the plots).
Metrics We evaluate our framework using four metrics: RMSE Log, RMSE Gain, Evaluation-Weighted Integrated RMSE, and Cost-Accuracy Trade-Off. By computing the Root Mean Square Error (RMSE) over 200 initialization seeds, all metrics intrinsically account for the empirical variance of the performance estimates across different samples, thereby providing a rigorous evaluation of each method’s precision. The first three metrics evaluate the methods when stopping at fixed coreset sizes 𝑛 = 1, . . . , 𝑁 (thereby not using the 𝜖1 , 𝜖2 stopping criterion). Evaluating across the entire range of coreset sizes as input stopping points enables us to determine whether any method dominates the others regardless of the stopping point. In contrast, the Cost-Accuracy Trade-Off measures estimation error at the specific coreset sizes reached when the 𝜖1 , 𝜖2 9
BayesAME: Bayesian Active Model Evaluation
stopping criterion is met, demonstrating which method achieves the lowest estimation error at the smallest automatically determined coreset size. 1. RMSE Log: The RMSE on a logarithmic scale, multiplied by 100. 2. RMSE Gain: The relative improvement of a given method over the RS-Mean baseline, calculated as: RMSE(RS-Mean) . RMSE Gain(Method) = 20 log10 RMSE(Method) A positive gain indicates a relative reduction in estimation error compared to ∫the baseline. 𝑁 3. Evaluation-Weighted Integrated RMSE: The weighted integral 100 0 𝑛′ RMSE( 𝑛′ ) 𝑑𝑛′ , which ensures that late-stage errors are penalized more heavily than early-stage ones. 4. Cost-Accuracy Trade-Off: The RMSE (multiplied by 100) computed at the coreset sizes dynamically determined by the 𝜖1 , 𝜖2 stopping criterion, plotted against the average coreset size 𝑛¯ across seeds. We vary 𝜖2 ∈ {0.005, 0.0075, 0.01, 0.015, 0.02, 0.025, 0.03}, and fix 𝜖1 = 0.005 since empirical results indicate that 𝜖2 is the primary factor governing the stopping decision. We also evaluate our method using Spearman’s rank correlation to assess how well the estimated performance preserves the true relative ordering of the models. The results, presented in Appendix C.3, demonstrate that our method works well for tasks such as model comparison and ranking.
Interpolation and Extrapolation Reference Setup To investigate the interpolation and extrapolation regimes with granularity, we partition the available models into a pool of potential reference models and a fixed set of 10 target models, following an approach similar to Zhang et al. (2025). In the interpolation regime, the target models are selected uniformly at random, with all remaining models assigned to the reference pool. In the extrapolation regime, we induce a distribution shift by designating the 10 highest-performing models as the target set, discarding the next 20% of models to create a performance gap, and using the remaining lower-performing models as potential reference models. To assess the impact of reference model quantity, we vary the size of the set of reference models by sampling 10%, 50%, or 90% of the reference pool uniformly at random. 5.1.1. Results Figure 1 shows the RMSE Log (top three rows) and RMSE Gain (bottom three rows) for GPQA with 𝑠𝑖 ∈ {0, 1} (left two columns) and MMLU-Pro with 𝑠𝑖 ∈ [0, 1] (right two columns) across different percentages of reference models (10%, 50%, and 90%) for the interpolation and extrapolation regimes. Overall, when considering the entire range of coreset sizes, BayesAME consistently achieves the lowest RMSE Log, demonstrating superior estimation accuracy and precision. The relative gains of BayesAME and the two other most competitive methods, BayesAME-RS and Bayes-AIPW, over the RS-Mean baseline can be fully appreciated in the RMSE Gain plots (bottom three rows), which exclude the less competitive methods (Seq-APW, Bayes-RS-Learn, ProEval, and IRT) for clarity. The plots also illustrate that the gain scales positively with the number of reference models, and that this trend is more pronounced for BayesAME. This confirms that a richer reference pool establishes a stronger prior, which BayesAME successfully exploits to isolate and select the most critical benchmark items. In the two cases where 𝐵 < 𝑁 − 𝑁0 , which occur for GPQA with 10% of reference models, BayesAME and BayesAME-RS coincide by construction and perform comparably to Bayes-AIPW. Figure 2 displays the Cost-Accuracy Trade-Off for GPQA with 𝑠𝑖 ∈ {0, 1} (left two columns) and MMLU-Pro with 𝑠𝑖 ∈ [0, 1] (right two columns). By providing the most accurate performance estimates for any threshold value at the lowest cost across both the interpolation and extrapolation regimes, BayesAME establishes a strictly superior Pareto frontier. Plots corresponding to Figures 1 and 2 for the remaining benchmarks, which show similar conclusions, are provided in Appendix C.3. We summarize these results for 50% of reference models via the EvaluationWeighted Integrated RMSE in Table 1. As the table shows, BayesAME outperforms the other methods in 10
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
100
10 2 101
0
20
40
60
80
100
10 2 101
100
100
10 1
10 1
10 1
10 1
0
20
40
60
80
100
10 2
0
20
Interpolation 5 4 3
10%
80
100
6
5 4 3 2
1
1
0 1 6
20
40
60
80
2 6
2
1
1
0 1 80
100
10 2
Interpolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
20
40
60
80
10
40
60
80
10
6
4
4
0
2
2
1
0
2 6 5 4 3
2
2
1
1
0 1
20
40
60
80
10
40
60
80
10 8
6
6
4
4
0
2
2
1
0 20
40
60
80
20
40
60
80
20
40
60
80
20
40
60
80
0 20
8
2
BayesAME BayesAME-RS BayesAIPW RS-Mean
0 20
8
3
80
80
6
4
60
60
8
5
40
40
2
2
20
20
0
3
2
0
1
4
60
10 2
4
3
40
100
2
4
20
80
0
5
6
60
4
5
2
40
Extrapolation 6
BayesAME BayesAME-RS BayesAIPW RS-Mean
2
2
50%
60
100
10 2
90%
40
0 20
40
60
80
Figure 1 | Single-target Setting. GPQA with binary scores (left two columns) and MMLU-Pro with continuous scores (right two columns). RMSE Log (top three rows) and RMSE Gain (bottom three rows) across varying proportions of reference models (10%, 50%, and 90%) for the interpolation and extrapolation regimes. The 𝑥 -axis denotes the coreset size 𝑛 as a percentage of the benchmark size 𝑁 . nearly all cases. For binary scores, we have 𝐵 < 𝑁 − 𝑁0 for MMLU-Pro, ARC-Challenge, and MuSR in both the interpolation and extrapolation regimes, and for IFEVAL in the extrapolation regime. For continuous scores, we have that 𝐵 < 𝑁 − 𝑁0 up to some tolerance for ARC-Challenge in the extrapolation regime. In these cases, BayesAME-RS coincides with BayesAME by construction. The table clearly quantifies the overall benefit of using richer continuous scores over binary scores. In Appendix C.3, we demonstrate this advantage via the other three metrics. Our results demonstrate that active coreset selection becomes significantly advantageous over random coreset selection whenever reference models provide sufficiently unique item representations (e.g., via larger reference pools or continuous scores), and show that it is possible to strongly outperform RS-Mean. While non-unique representations cause BayesAME to default to random selection (matching BayesAME-RS and Bayes-AIPW), richer item signals enable active selection to effectively target informative items and strongly outperform RS-Mean, addressing the concerns raised by Zhang et al. (2025).
11
BayesAME: Bayesian Active Model Evaluation
10%
Interpolation 0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 0.1
50%
0.0
90%
Extrapolation
0.7
80
85
100
0.0
0.8
0.8
0.7
0.7
0.6
0.6
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.1 95
Extrapolation 0.9
0.5
0.2
90
Interpolation 0.9
80
85
0.3 0.2 90
95
100
0.1
0.5 0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean 20
40
0.3 0.2 60
80
100
0.1
0.9
0.9
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.4
0.4
0.2
0.2
0.3
0.3
0.1
0.1
0.2
0.0
0.0
0.1
80
85
90
95
100
60
80
100
0.1
0.7 0.6
0.5
0.5
0.4
0.4
0.2
0.3
0.3
0.1
0.2
0.4
0.3
0.3
0.2 0.1 100
40
0.8
0.4
0.0
80
85
90
95
100
0.1
20
40
60
80
100
20
40
60
80
100
20
40
60
80
100
0.2 20
0.6
0.5
95
100
0.7
0.6
0.5
90
95
0.8
0.6
85
90
0.9
0.7
80
85
0.9
0.7
0.0
80
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 20
40
60
80
100
0.1
Figure 2 | Single-target Setting. GPQA with binary scores (left two columns) and MMLU-Pro with continuous scores (right two columns). Cost-Accuracy Trade-Off across varying proportions of reference models (10%, 50%, and 90%) for the interpolation and extrapolation regimes. The 𝑥 -axis denotes the average coreset size 𝑛¯ as a percentage of the benchmark size 𝑁 . 5.2. Multi-Target Setting We compare the multi-target version of BayesAME (Section 3.2) against its single-target counterpart (Section 3.1) applied independently to each target model. To evaluate the multi-target formulation under varying degrees of correlation, we construct sets of 10 target models exhibiting either low or high correlation. Specifically, we form the low-correlation set using a random selection of 10 models, and the high-correlation set using the 10 most correlated models overall. We construct the set of reference models by sampling 50% of the reference model pool uniformly at random. Figure 3 shows RMSE Log, RMSE Gain, and Cost-Accuracy Trade-Off for IFEval with 𝑠𝑖 ∈ {0, 1} (left two columns) and GPQA with 𝑠𝑖 ∈ [0, 1] (right two columns). Additional results for GPQA with 𝑠𝑖 ∈ {0, 1} are given in Fig. 14 of Appendix C.4. The primary conclusion drawn from these plots is that jointly evaluating multiple target models provides a significant efficiency advantage when the models exhibit correlated behaviors. Specifically, in the high-correlation regime, BayesAME Multi-target demonstrates a steeper decline in RMSE Log and reach a higher RMSE Gain much earlier compared to BayesAME Single-target. This is further reflected in the Cost-Accuracy Trade-Off plots, where BayesAME Multi-target establishes a strictly superior Pareto frontier, achieving the same level of estimation accuracy at a notably lower coreset size across all threshold values. However, the low-correlation regime plots reveal a practical nuance of the multi-target framework. When target models do not share meaningful behavioral similarities, attempting to jointly model their latent abilities introduces a degree of negative transfer. As observed in the Cost-Accuracy Trade-Off plots, the multi-target approach slightly underperforms BayesAME Single-target, exhibiting higher estimation errors for a given evaluation budget. Crucially, however, this degradation is not catastrophic, the performance of the multi-target method degrades gracefully to match that of the competitive baseline Bayes-AIPW. This suggests that while the active selection process expends some resources attempting to learn nonexistent cross-model correlations, the underlying framework remains highly robust. Nonetheless, to maximize evaluation efficiency, practitioners should apply the multi-target extension primarily when there is a reasonable prior expectation of correlated performance.
12
BayesAME: Bayesian Active Model Evaluation
Benchmarks
BayesAME
BayesAME-RS
RS-Mean
Seq-APW
Bayes-AIPW
Bayes-RS-Learn
ProEval
IRT
Binary GPQA
Interpolation Extrapolation
61.54 68.91
66.38 72.65
76.48 81.73
866.61 751.08
67.21 73.41
99.06 403.29
96.29 418.11
89.61 91.40
MMLU-Pro
Interpolation Extrapolation
30.18 37.20
30.18 37.20
40.20 44.89
1141.06 821.40
30.68 37.69
90.59 670.20
95.56 701.84
46.38 156.33
BBH
Interpolation Extrapolation
31.70 38.35
36.35 42.02
50.42 52.59
495.01 448.04
35.54 41.34
50.63 448.53
57.82 600.98
50.95 158.41
ARC-Challenge
Interpolation Extrapolation
36.53 36.56
36.53 36.56
56.09 53.21
1391.22 768.83
36.68 38.03
70.84 327.45
74.06 364.21
44.19 88.58
MuSR
Interpolation Extrapolation
47.86 54.19
47.86 54.19
69.27 71.28
949.90 719.86
49.18 56.72
69.12 289.76
68.54 353.44
84.51 104.41
IFEval
Interpolation Extrapolation
49.53 58.08
56.14 58.08
72.45 64.54
1464.91 488.54
56.71 60.52
76.22 466.45
78.53 1017.33
79.06 212.27
GPQA
Interpolation Extrapolation
30.58 40.40
34.42 40.75
45.25 56.64
50.22 62.52
35.21 46.77
49.08 262.65
54.03 313.65
81.44 142.56
MMLU-Pro
Interpolation Extrapolation
13.11 26.82
17.10 31.10
27.80 40.83
39.99 47.07
17.19 31.33
72.66 482.43
82.62 570.76
66.44 311.12
BBH
Interpolation Extrapolation
13.04 29.80
17.54 35.57
35.21 47.10
49.38 48.50
18.01 33.64
26.13 303.09
40.45 612.49
61.84 388.23
ARC-Challenge
Interpolation Extrapolation
26.04 33.16
26.47 33.16
47.67 49.07
93.82 43.52
26.19 33.05
46.84 292.45
53.57 354.75
46.10 144.55
MuSR
Interpolation Extrapolation
27.83 36.96
33.42 42.56
50.67 60.19
76.21 70.64
34.34 43.69
50.02 246.90
62.88 372.08
57.39 211.37
Natural QA Openbook
Interpolation Extrapolation
28.26 20.83
29.68 28.36
45.91 41.76
1185.87 1044.02
29.12 27.70
275.90 226.18
286.37 262.55
129.11 127.84
Continuous
Table 1 | Single-target Setting. Evaluation-Weighted Integrated RMSE for 50% of reference models.
6. Discussion and Conclusions A primary objective of this work was to provide an efficient model evaluation method that is capable of automatically determining a coreset size when reliable performance estimation takes priority over efficiency, while remaining flexible to accept a predefined coreset size. To this end, we introduced a simple, highperforming method, BayesAME. Furthermore, we proposed an extension to the multi-target setting that leverages performance correlations among target models to further reduce the coreset size. Another objective was to address skepticism in recent literature, which suggests that random coreset selection and simply using a random sample mean baseline gives competitive performance. Our study overturns this belief by showing that, when leveraging richer information (such as continuous scores and sufficiently varied sets of reference models), looking across the entire spectrum of coreset sizes, and accounting for the empirical variance of the performance estimates across different samples, non-random coreset selection and more robust performance estimation remain highly advantageous. Interestingly, while our results challenge the literature’s reliance on random selection, they do validate the underlying intuition that simple methods are difficult to outperform. We found that introducing highly sophisticated, principled modeling of the scores did not translate to better performance. For example, when dealing with binary or bounded continuous scores, a more principled model of the latent abilities using logit transformations (as described in Appendix B.3) failed to outperform the coarser joint Gaussian approach. Though relying on an improper prior for bounded data, the Gaussian approach proved far more effective because it allowed for highly efficient, closed-form evaluations of information gain. Similarly, our attempts to find richer, more complex representations of items (for example, including prompt embeddings), rather than simply relying on reference scores, did not yield tangible improvements. There are several limitations of BayesAME that provide promising avenues for future research. Currently, our framework assumes scalar scores. A natural step is to adapt it to handle more complex, multi-dimensional, or unstructured evaluation formats. In addition, future iterations could investigate ways to integrate architecture details, training hyperparameters, or training data distributions to further refine the prior covariance.
13
BayesAME: Bayesian Active Model Evaluation
Low correlation
High correlation
101
100
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
10 2
20
40
101
100
100
10 1
60
80
100
10 2
6
6
5
5
4
4
3
3
2
2
1
1
0
0
1 2 10
Low correlation
101
40
60
80
100
30
40
50
60
70
80
2 10
90
1.0
1.0
0.8
0.8
0.6
0.6
0.4
0.4 0.2
0.2
10 2 10 8
20
30
40
50
60
20
30
40
50
60
70
80
70
80
90
100
6
4
4
2
2
75
80
85
90
95
100
70
75
80
85
90
95
100
20
30
40
50
60
70
80
90
20
30
40
50
60
70
80
90
80
90
100
0 20
30
40
50
60
70
80
2 10
90
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0.2
0.1 70
10 2 10 8
6
2 10
90
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
0
1 20
100
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
20
High correlation 101
0.1 50
60
70
80
90
100
50
60
70
100
Figure 3 | Multi-target Setting. IFEval with binary scores (left two columns) and GPQA with continuous scores (right two columns). The rows represent RMSE Log, RMSE Gain, and Cost-Accuracy Trade-Off, respectively, for 50% of reference models for the low and high correlation regimes.
Acknowledgments The authors would like to thank Michalis Titsias for valuable discussions and providing feedback on the manuscript.
References M. A. Álvarez, L. Rosasco, and N. D. Lawrence. Kernels for vector-valued functions: A review. Foundations and Trends® in Machine Learning, 4:195–266, 2012. A. M. Bean, N. Seedat, S. Chen, and J. R. Schwarz. Scales++: Compute efficient evaluation subset selection with cognitive scales embeddings. arXiv preprint arXiv:2510.26384, 2025. G. Berrada, J. Kossen, M. Razzak, F. B. Smith, Y. Gal, and T. Rainforth. Scaling up active testing to large language models. In Advances in Neural Information Processing Systems, 2025. S. Bowyer, A. Locatelli, and K. Cao. Efficient benchmarking is just feature selection and multiple regression. arXiv preprint arXiv:2605.25773, 2026. R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu. A limited memory algorithm for bound constrained optimization. SIAM Journal on Scientific Computing, 16:1190–1208, 1995. Y. Chen, T. S. Filho, R. B. Prudencio, T. Diethe, and P. Flach. 𝛽 3 -IRT: A new item response model and its applications. In International Conference on Artificial Intelligence and Statistics, pages 1013–1021, 2019. P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. A. Fisch, D. Deutsch, J. Maynez, A. Agarwal, J. Berant, W. Cohen, A. Globerson, and J. Eisenstein. CollabEval: Statistically efficient collaborative model evaluation via matrix completion. arXiv preprint arXiv:2607.05046, 2026. C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, and T. Wolf. Open LLM leaderboard v2, 2024.
14
BayesAME: Bayesian Active Model Evaluation
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In International Conference on Learning Representations, 2021. C.-Y. Hsu and S. Shekhar. arXiv:2607.17409, 2026.
Efficient sequential evaluation of large language models.
arXiv preprint
Y. Huang, W. Zeng, A. Kumaresan, and Z. Wang. ProEval: Proactive failure discovery and efficient performance estimation for generative AI evaluation. In International Conference on Machine Learning, 2026. H. Jiang, S. Kwon, J. Luo, Z. Xiao, and S. Zhang. Can we trust item response theory for AI evaluation? arXiv preprint arXiv:2607.15190, 2026. A. G. Journel and C. J. Huijbregts. Mining Geostatistics. Academic Press, 1978. A. Kipnis, K. Voudouris, L. M. Schulze Buschoff, and E. Schulz. metabench - A sparse benchmark of reasoning and knowledge in large language models. In International Conference on Learning Representations, 2025. J. Kossen, S. Farquhar, Y. Gal, and T. Rainforth. Active testing: Sample-efficient model evaluation. In International Conference on Machine Learning, 2021. A. Kulesza and B. Taskar. k-DPPs: Fixed-size determinantal point processes. In International Conference on Machine Learning, pages 1193–1200, 2011. T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466, 2019. P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Re, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. WANG, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. S. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. A. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda. Holistic evaluation of language models. Transactions on Machine Learning Research, 2023. L. Liao, Q. Zhang, R. Wu, and G. Fang. Toward a unified framework for data-efficient evaluation of large language models. arXiv preprint arXiv:2510.04051, 2025. Z. Liu, J. Zhang, C. Liu, and Y. Zhu. Active testing of large language models via approximate Neyman allocation. arXiv preprint arXiv:2605.10075, 2026. I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. D. J. C. MacKay. The evidence framework applied to classification networks. Neural Computation, 4(5):720–736, 1992. Y. Noel and B. Dauvier. A Beta item response model for continuous bounded responses. Applied Psychological Measurement, 31(1):47–73, 2007. M. Osborne, R. Garnett, S. Roberts, C. Hart, S. Aigrain, and N. Gibson. Bayesian quadrature for ratios. In International Conference on Artificial Intelligence and Statistics, pages 832–840, 2012. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and É. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. Y. Perlitz, E. Bandel, A. Gera, O. Arviv, L. Ein-Dor, E. Shnarch, N. Slonim, M. Shmueli-Scheuer, and L. Choshen. Efficient benchmarking (of language models). In Conference of the North American Chapter of the Association for Computational Linguistics, pages 2519–2536, 2024.
15
BayesAME: Bayesian Active Model Evaluation
F. M. Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin. tinyBenchmarks: evaluating LLMs with fewer examples. In International Conference on Machine Learning, 2024. C. E. Rasmussen and C. K. I. Williams. Gaussian Processes for Machine Learning. The MIT Press, 2006. D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. GPQA: A graduate-level google-proof Q&A benchmark. In First Conference on Language Modeling, 2024. G. Saranathan, C. Xu, M. P. Alam, T. Kumar, M. Foltin, S. Y. Wong, and S. Bhattacharya. SubLIME: Subset selection via rank correlation prediction for data-efficient LLM evaluation. In Annual Meeting of the Association for Computational Linguistics, pages 30572–30593, 2025. Z. Sprague, X. Ye, K. Bostrom, S. Chaudhuri, and G. Durrett. MuSR: Testing the limits of chain-of-thought with multistep soft reasoning. In International Conference on Learning Representations, 2024. M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, and J. Wei. Challenging BIG-Bench tasks and whether chain-of-thought can solve them. In Findings of the Association for Computational Linguistics, pages 13003–13051, 2023. R. Vivek, K. Ethayarajh, D. Yang, and D. Kiela. Anchor points: Benchmarking models with much fewer examples. In Conference of the European Chapter of the Association for Computational Linguistics, pages 1576–1601, 2024. G. Wang, Z. Chen, B. Li, and H. Xu. Cer-Eval: Certifiable and cost-efficient evaluation framework for LLMs. arXiv preprint arXiv:2505.03814, 2025. Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. In Advances in Neural Information Processing Systems (Track on Datasets and Benchmarks), 2024. G. Zhang, F. E. Dorner, and M. Hardt. How benchmark prediction from fewer data misses the mark. In Advances in Neural Information Processing Systems, 2025. H. Zhao, M. Li, L. Sun, and T. Zhou. BenTo: Benchmark reduction with in-context transferability. In International Conference on Learning Representations, 2025. J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023.
16
BayesAME: Bayesian Active Model Evaluation
A. BayesAME Details A.1. Posterior Distribution Computation In this section, we show the equivalence between 𝑝 ( 𝜃 | 𝑠 C ) and 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ) stated in Section 3.1. Based on this, we derive the two equivalent expressions for 𝜇 𝜃| C and Σ𝜃| C given in Eq. (7) and Eq. (8). Recall that ¯𝑠 C1:𝐵 = (¯𝑠 C1 , . . . , ¯𝑠 C𝐵 ), where ¯𝑠 C𝑏 = 𝑛1𝑏 2 bucket 𝑏. As 𝑝 (¯𝑠 C𝑏 | 𝜃𝑏 ) = N 𝜃𝑏 , 𝜎 𝑛𝑁𝑏 𝑏 , we have: 1
𝑝 ( 𝑠 C𝑏 | 𝜃𝑏 ) ∝ exp −
Í
𝑖 ∈ C𝑏 𝑠𝑖 and C𝑏 denotes the subset of 𝑛𝑏 items in C that are in
! ∑︁
2𝜎2 𝑁𝑏 𝑖 ∈ C
( 𝑠𝑖 − 𝜃𝑏 )
2
𝑏
= exp −
1
! ∑︁
2𝜎2 𝑁𝑏 𝑖 ∈ C
(( 𝑠𝑖 − ¯𝑠 C𝑏 ) + (¯𝑠 C𝑏 − 𝜃𝑏 ))
2
𝑏
= exp −
= exp −
!!
1
∑︁
2𝜎2 𝑁𝑏 1
2
( 𝑠𝑖 − ¯𝑠 C𝑏 ) + 𝑛𝑏 (¯𝑠 C𝑏 − 𝜃𝑏 )
2
𝑖 ∈ C𝑏
! ∑︁
2𝜎2 𝑁𝑏 𝑖 ∈ C
( 𝑠𝑖 − ¯𝑠 C𝑏 )
2
exp −
𝑛𝑏
2𝜎2 𝑁𝑏
(¯𝑠 C𝑏 − 𝜃𝑏 )
2
𝑏
𝑛𝑏
∝ exp −
2𝜎2 𝑁𝑏 ∝ 𝑝 (¯𝑠 C𝑏 | 𝜃𝑏 ) ,
(¯𝑠 C𝑏 − 𝜃𝑏 )
2
where the proportionality is with respect to 𝜃𝑏 . This results into: 𝑝 ( 𝜃 | 𝑠 C ) ∝ 𝑝 ( 𝜃) 𝑝 ( 𝑠 C | 𝜃) = 𝑝 ( 𝜃)
𝐵 Ö
𝐵 Ö
𝑝 ( 𝑠 C𝑏 | 𝜃𝑏 ) ∝ 𝑝 ( 𝜃)
𝑏=1; 𝑛𝑏 >0
𝑝 (¯𝑠 C𝑏 | 𝜃𝑏 ) = 𝑝 ( 𝜃) 𝑝 (¯𝑠 C1:𝐵 | 𝜃) ∝ 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ) ,
𝑏=1; 𝑛𝑏 >0
which establishes that 𝑝 ( 𝜃 | 𝑠 C ) = 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ). Having shown this equivalence, we now derive the explicit form for 𝑝 ( 𝜃 | 𝑠 C ) = 𝑝 ( 𝜃 | ¯𝑠 C1:𝐵 ) through the Gaussian conditioning rule, which states that, for a joint distribution 𝑝 ( 𝑥1 , 𝑥2 ), the conditional distribution 𝑝 ( 𝑥1 | 𝑥2 ) has −1 ( 𝑥 − 𝜇 ) and covariance Σ − Σ Σ −1 Σ . mean 𝜇 1 + Σ12 Σ22 2 2 11 12 22 21 Recall that 𝐻 C and 𝐷 C denote the submatrices of 𝐻 and 𝐷 formed by the rows corresponding to C. The joint distribution of ( 𝜃, 𝑠 C ) can be written as: 𝜃 𝜃 Σ Σ𝜃 𝐻 ⊤ 𝜃 𝜇 C ∼N , . 𝑠C 𝐻 C 𝜇𝜃 𝐻 C Σ𝜃 𝐻 C Σ𝜃 𝐻 ⊤ C + 𝐷C Applying the Gaussian conditioning rule, we obtain: 𝜃 ⊤ −1 𝜃 𝜇 𝜃| C = 𝜇 𝜃 + Σ𝜃 𝐻 ⊤ C ( 𝐻C Σ 𝐻C + 𝐷C ) (𝑠C − 𝐻C 𝜇 ), 𝜃 𝜃 ⊤ −1 Σ𝜃| C = Σ𝜃 − Σ𝜃 𝐻 ⊤ C ( 𝐻C Σ 𝐻C + 𝐷C ) 𝐻C Σ ,
which corresponds to Eq. (7). Let Δ be a diagonal noise matrix with entries Δ𝑏,𝑏 = 𝜎2 𝑁𝑏 /𝑛𝑏 . The joint distribution of ( 𝜃, ¯𝑠 C1:𝐵 ) can be written as: 𝜃
¯𝑠 C1:𝐵
∼N
𝜇𝜃 Σ𝜃 𝜃 , 𝜇 Σ𝜃
Σ𝜃
Σ +Δ 𝜃
.
17
BayesAME: Bayesian Active Model Evaluation
Applying the Gaussian conditioning rule, we obtain: 𝜇 𝜃| C = 𝜇 𝜃 + Σ𝜃 ( Σ𝜃 + Δ) −1 (¯𝑠 C1:𝐵 − 𝜇 𝜃 ) , Σ𝜃| C = Σ𝜃 − Σ𝜃 ( Σ𝜃 + Δ) −1 Σ𝜃 ,
which corresponds to Eq. (8).
A.2. Performance Estimator In this section, we discuss the estimator 𝑅ˆ𝐶𝑉 (Eq. (4)) that we use in the non-unique reference regime 𝐵 < 𝑁 − 𝑁0 . To simplify the exposition, we adopt an item-level, rather than bucket-level, description. Specifically, for an item ( 𝑥 𝑖 , 𝑦𝑖 ) belonging to bucket 𝑏, we use 𝜇 𝑖 | C to denote 𝔼[ 𝜃𝑏 | 𝑠 C ]. We can rewrite Eq. (4) at the item-level as: ˆCV = 𝑅
1 ∑︁ 𝑛
𝑖∈ C
𝑁
𝑠𝑖 − 𝜇 𝑖 | C +
1 ∑︁ 𝑁
𝜇 𝑗| C .
𝑗=1
This estimator can be seen as an in-sample proxy of the following control variate estimator that uses out-ofsample posteriors: 𝑁 1 ∑︁ © 1 ∑︁ ª ˆCVO = 𝑅 𝜇 𝑗 | C\{ 𝑖 } ® . 𝑠𝑖 − 𝜇 𝑖 | C\{ 𝑖 } + 𝑛 𝑁 𝑖∈ C « 𝑗=1 ¬ Note that the in-sample proxy estimator avoids the computational bottleneck of recomputing the posterior for every held-out item. ˆCVO is unbiased with respect to the random selection of the coreset as shown below. Let C = ( 𝑖1 , . . . , 𝑖𝑛 ) be 𝑅 constructed by sampling items uniformly at random with replacement, such that 𝑖𝑘 ∼ 𝑈 (1, . . . , 𝑁 ). 𝑅ˆCVO can be expressed as: 𝑛 𝑁 1 ∑︁ © 1 ∑︁ ª 𝜇 𝑗 | C\{ 𝑖𝑘 } ® . 𝑠𝑖𝑘 − 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 } + 𝑛 𝑁 𝑘=1 « 𝑗=1 ¬ Taking the unconditional expectation with respect to the random selection of the coreset C yields:
ˆCVO = 𝑅
𝑛
𝔼[ 𝑅ˆCVO ] =
1 ∑︁ 𝑛
𝔼[ 𝑠𝑖𝑘 ] −
𝑘=1
𝑛 𝑁 1 ∑︁ ª 1 ∑︁ © 𝔼 𝜇 𝑗 | C\{ 𝑖𝑘 } ® . 𝔼 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 } − 𝑛 𝑁 𝑘=1 « 𝑗=1 ¬
(10)
We first focus on the term 𝔼 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 } . Using the Law of Total Expectation, we have: 𝔼 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 } = 𝔼 C\{ 𝑖𝑘 } 𝔼 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 } |C \ { 𝑖𝑘 } . Because 𝑖𝑘 is drawn independently from the remaining coreset items C \ { 𝑖𝑘 }, the inner expectation over the uniform draw of 𝑖𝑘 is exactly the population mean:
𝔼 𝜇 𝑖𝑘 | C\{ 𝑖𝑘 }
∑︁ 𝑁 1 𝑁 1 ∑︁ = 𝔼 C\{ 𝑖𝑘 } 𝜇 𝑗 | C\{ 𝑖𝑘 } = 𝔼 𝜇 𝑗 | C\{ 𝑖𝑘 } . 𝑁 𝑗=1 𝑁 𝑗=1
Plugging this into Eq. (10), the second term cancels out, resulting in: 𝑛
𝔼[ 𝑅ˆCVO ] =
∑︁ 𝑁 1 𝑁 1 ∑︁ 𝔼[ 𝑠𝑖𝑘 ] = 𝔼 C\{ 𝑖𝑘 } [𝔼[ 𝑠𝑖𝑘 | C \ { 𝑖𝑘 }]] = 𝔼 C\{ 𝑖𝑘 } 𝑠 𝑗 = 𝑠 𝑗 = 𝑅. 𝑛 𝑛 𝑛 𝑁 𝑗=1 𝑁 𝑗=1 𝑘=1 𝑘=1 𝑘=1
1 ∑︁
𝑛
1 ∑︁
𝑛
1 ∑︁
Remark: In practice, the coreset is sampled without replacement. This breaks strict independence, introducing a weak dependence between 𝑖𝑘 and C \ { 𝑖𝑘 } and yielding a residual bias. However, this value is negligible when 𝑁 ≫ 1. 18
BayesAME: Bayesian Active Model Evaluation
A.3. Stopping Criterion Recall that, to monitor posterior uncertainty about the random variable 𝑅, we consider the width of the 95% credible interval for 𝑅, which is governed by Var( 𝑅 | 𝑠 C ). Crucially, this quantity is not tied to a particular choice of point estimator. This distinction is important in the non-unique reference regime, 𝐵 < 𝑁 − 𝑁0 , where we use the control-variate estimator 𝑅ˆCV in Eq. (4). Our stopping criterion deliberately separates these two roles: the stability condition ( 𝑖) is applied to the reported estimator 𝑅ˆCV , whereas the credible-interval condition ( 𝑖𝑖) is applied to the target quantity 𝑅. Í Í A natural alternative would be to define a random variable 𝑅CV = 1𝑛 𝑖 ∈ C 𝑠𝑖 + 𝑏𝐵=1 𝑁𝑁𝑏 − 𝑛𝑛𝑏 𝜃𝑏 and use Var( 𝑅CV | 𝑠 C ) in the stopping criterion. However, this variance is not reliable as a measure of uncertainty about 𝑅 since it vanishes whenever 𝑁𝑏 / 𝑁 = 𝑛𝑏 /𝑛, which occurs, e.g., when there is only one bucket. Another possible estimator-specific diagnostic is the posterior MSE, 𝔼[( 𝑅 − 𝑅ˆCV ) 2 | 𝑠𝐶 ] = Var( 𝑅 | 𝑠 C ) + ( 𝑅ˆBayes − 𝑅ˆCV ) 2 . This is a valid posterior risk criterion under squared-error loss: the first term captures residual uncertainty about 𝑅, while the second captures the conditional bias of 𝑅CV relative to the posterior mean. We use Var( 𝑅 | 𝑠 C ) as the default uncertainty measure because it preserves a direct credible-interval interpretation for the benchmark performance and avoids conflating posterior uncertainty with the deliberate correction introduced by the control-variate estimator. The estimator-specific discrepancy is instead monitored through the stability condition ( 𝑖).
A.4. Hyperparameter Optimization Single-target Setting We tune the hyperparameters 𝛼 and 𝛽 , which govern the prior covariance Σ𝜃 , by minimizing the negative log-marginal likelihood L ( 𝛼, 𝛽 ) of the coreset scores. As established in Appendix A.1, the conditional likelihoods 𝑝 ( 𝑠 C | 𝜃) and 𝑝 (¯𝑠 C1:𝐵 | 𝜃) are proportional up to a constant independent of 𝜃. As a result, their negative log-marginal likelihoods differ only by a constant independent of 𝛼 and 𝛽 . Therefore, to evaluate the objective efficiently, we define L ( 𝛼, 𝛽 ) as: ( 1 1 𝜃 ⊤ −1 𝜃 ( 𝑠 C − 𝐻 C 𝜇 𝜃 ) ⊤ ( 𝐻 C Σ𝜃 𝐻 ⊤ C + 𝐷 C ) ( 𝑠 C − 𝐻 C 𝜇 ) + 2 log | 𝐻 C Σ 𝐻 C + 𝐷 C | if 𝑛 < 𝐵 L ( 𝛼, 𝛽 ) = 12 (11) 1 𝜃 ⊤ 𝜃 −1 𝜃 𝜃 𝑠 C1:𝐵 − 𝜇 ) ( Σ + Δ) (¯𝑠 C1:𝐵 − 𝜇 ) + 2 log | Σ + Δ | if 𝑛 ≥ 𝐵 2 (¯ For additional efficiency and to maintain numerical stability, we employ a Cholesky decomposition of the respective covariance matrix. We perform this minimization via an L-BFGS-B optimizer (Byrd et al., 1995), constraining the parameters with a small lower bound of 10−5 to ensure the covariance matrix remains positive definite.
Multi-target Setting We tune the hyperparameters 𝛼𝑙 and 𝛽𝑙 which govern the prior covariance Σ𝜃,𝑙 , as well 𝑡 as the model-specific weights {𝑤𝑙𝑡 }𝑡𝐾=1 for each 𝑙 = 1, . . . , 𝐿. Additionally, we learn the noise variance 𝜎2 instead of keeping it fixed as for the single-target setting. This turns minimizing the negative log-marginal likelihood into a highly non-convex, joint optimization problem. Consequently, instead of L-BFGS-B, we employ AdamW (Loshchilov and Hutter, 2019). AdamW leverages momentum to successfully navigate complex loss surfaces and provides adaptive, per-parameter learning rates, which are essential for handling parameters that operate on fundamentally different scales. Because the hyperparameters naturally stabilize as the coreset grows, we deploy an adaptive computational schedule that executes 500 gradient steps during the initial active learning iterations and progressively decreases this count in later iterations to significantly reduce computational overhead. To ensure numerical stability throughout the training process, we clip the values of 𝛼𝑙 , 𝛽𝑙 , and 𝜎2 after each update, enforcing a lower bound of 10−4 .
19
BayesAME: Bayesian Active Model Evaluation
B. BayesAME Single-target Variants B.1. Batch Selection Strategies We extend the approach from the main text by introducing selection strategies that identify an optimal batch of items 𝐴, rather than a single item. We focus on the nearly unique reference regime 𝐵 ≥ 𝑁 − 𝑁0 . To simplify the presentation, we denote Σ | C = 𝐻Σ𝜃| C 𝐻 ⊤ + 𝐷, which in the case 𝐵 = 𝑁 simplifies to Σ | C = Σ𝜃| C + 𝜎2 𝐼 𝑁 . The extension of the expected information gain Eq. (6) from one item to a batch of items 𝐴 is given by: 𝛼IG ( 𝐴) = H ( 𝑅 | 𝑠 C ) − 𝔼𝑠 𝐴 [H ( 𝑅 | 𝑠 C , 𝑠 𝐴 )]
1 1 = − log 1 − 2 ( Σ 𝐴,𝑈 | C 1 |𝑈 | ) ⊤ Σ −1 ( Σ 1 ) , 𝐴,𝑈 | C | 𝑈 | 𝐴,𝐴 | C 2 𝑁 Var( 𝑅 | 𝑠 C ) where Σ 𝐴,𝑈 | C denotes the submatrix of Σ | C of columns and rows corresponding to 𝐴 and the unevaluated set 𝑈 = 𝐸 \ C, respectively, and 1 | 𝑈 | is a |𝑈 |-dimensional vector of ones. A fundamental property of Gaussian conditioning where the noise is not a function of the score value implies that Σ | C depends only on the score indexes rather than their actual values. This means that batch selection strategies simplify to a purely combinatorial problem of identifying an optimal batch of indices. Once this optimal batch is identified, the specific sequence in which these items are evaluated is arbitrary. Because the expected information gain is strictly monotonically increasing with respect to the variance reduction, maximizing 𝛼IG ( 𝐴) is equivalent to maximizing the variance reduction term: 𝛽 ( 𝐴) =
1 𝑁2
( Σ 𝐴,𝑈 | C 1 |𝑈 | ) ⊤ Σ −1 𝐴,𝐴 | C ( Σ 𝐴,𝑈 | C 1 | 𝑈 | ) .
(12)
Therefore, for a batch 𝐴 of specified size | 𝐴 | = 𝑘, we obtain the following optimization problem: 𝐴★ = argmax 𝐴 ⊂𝑈, | 𝐴 |=𝑘 𝛽 ( 𝐴) .
(13)
Batch Size | 𝐴 | = 2 When | 𝐴 | = 2, we can derive a closed-form expression of 𝛽 ( 𝐴 = { 𝑖, 𝑗 }) that can be efficiently vectorized: 2 2 Σ𝑖,𝑖 | C Σ𝑖, 𝑗 | C −1 𝑇𝑖 1 𝑇𝑖 Σ 𝑗, 𝑗 | C + 𝑇 𝑗 Σ𝑖,𝑖 | C − 2𝑇𝑖𝑇 𝑗 Σ𝑖, 𝑗 | C 1 , 𝛽 ({ 𝑖, 𝑗 }) = 2 𝑇𝑖 𝑇 𝑗 = 2 𝑇𝑗 Σ𝑖, 𝑗 | C Σ 𝑗, 𝑗 | C 𝑁 𝑁 Σ𝑖,𝑖 | C Σ 𝑗, 𝑗 | C − Σ𝑖,2 𝑗 | C where 𝑇𝑖 =
Í
𝑘 ∈ 𝑈 Σ𝑖,𝑘 | C and similarly for 𝑇 𝑗 .
Batch Size | 𝐴 | > 2 As the batch size | 𝐴 | grows, an exhaustive search over all ||𝑈𝐴 || subsets requires O (|𝑈 | | 𝐴 | | 𝐴 | 3 ) operations, which quickly becomes computationally infeasible even for moderate |𝑈 |. We address this by replacing the exhaustive search with an approximation algorithm that scales linearly with |𝑈 |. Specifically, we use heuristic warm-start strategies to identify initial candidate batches, followed by a combination of local search and Sequential Monte Carlo (SMC). Coupling local search with SMC prevents the algorithm from getting stuck in local optima: we maintain a population of candidate batches, apply local search to each independently, and resample them according to their fitness 𝛽 ( 𝐴). We use the following warm-start strategies: • LARS: Solve an ℓ1 -penalized continuous relaxation of Eq. (13). In particular, for 𝑤 ∈ ℝ𝑁 define the following function: 1 𝑔 ( 𝑤) = 𝑤⊤ Σ | C 𝑤 − 𝑤⊤ 𝑣, 2 𝑐 where 𝑣 = Σ𝐸,𝑈 | C 1 |𝑈 | . Fix a subset 𝐴 and let 𝑤 𝐴 = 0. The gradient of 𝑔 with respect to 𝑤 𝐴 is ∇𝑤𝐴 𝑔 = Σ 𝐴,𝐴 | C 𝑤 𝐴 − 𝑣 𝐴 , 20
BayesAME: Bayesian Active Model Evaluation
where 𝑣 𝐴 = Σ 𝐴,𝑈 | C 1 |𝑈 | . Setting the gradient to zero yields the optimizer 𝑤★𝐴 = Σ −1 𝑣 . Substituting 𝑤★𝐴 𝐴,𝐴 | C 𝐴 into 𝑔 gives 𝑁2 1 −1 𝑔 ( 𝑤★𝐴 ) = − 𝑣⊤ 𝛽 ( 𝐴) . 𝐴 Σ 𝐴,𝐴 | C 𝑣 𝐴 = − 2 2 Therefore, maximizing 𝛽 ( 𝐴) over all subsets 𝐴 of size 𝑘 is equivalent to solving the ℓ0 -constrained quadratic program 1 min 𝑤⊤ Σ | C 𝑤 − 𝑤⊤ 𝑣 subject to ∥ 𝑤 ∥ 0 ≤ 𝑘. 𝑁 𝑤 ∈ℝ 2 Since this problem is NP-hard, we instead consider the standard ℓ1 relaxation min
𝑤 ∈ℝ 𝑁
1 ⊤ 𝑤 Σ | C 𝑤 − 𝑤⊤ 𝑣 + 𝜆 ∥ 𝑤 ∥ 1 , 2
𝜆 > 0.
To solve this, we apply the least angle regression (LARS) algorithm to efficiently extract an initial active set of | 𝐴 | items. • DPP: Sample multiple diverse initial batches using a 𝑘-determinantal point process (𝑘-DPP) (Kulesza and Taskar, 2011) restricted to size | 𝐴 |, using the posterior covariance matrix of the candidates as the likelihood kernel. • Clust: Perform 𝐾 -means clustering on the reference model score vectors 𝜙𝑖 to partition the candidate pool into | 𝐴 | distinct clusters. Generate multiple initial sets by uniformly sampling one item from each cluster. Crucially, during the local search phase, evaluating a neighbor where only a single item in the candidate batch is swapped can be computed highly efficiently using rank-2 updates. Suppose we wish to recompute the inverse in Eq. (12) after a swap. Let 𝑀 𝐴 = Σ 𝐴,𝐴 | C be the original | 𝐴 | × | 𝐴 | matrix and 𝑀 𝐴′ = Σ 𝐴, ˜ 𝐴˜| C be the updated matrix. If the 𝑚-th item is swapped from 𝐴 to 𝐴˜, then 𝑀 𝐴 and 𝑀 𝐴′ differ only by their 𝑚-th row and column. Let 𝑒𝑚 ∈ ℝ | 𝐴 | denote the 𝑚-th standard basis vector (a vector of zeros with a 1 at the 𝑚-th coordinate). Let 𝑐𝑚 be the 𝑚-th column of 𝑀 𝐴 and ˜𝑐𝑚 be the corresponding column in 𝑀 𝐴′ . Defining the difference vector 𝑑 = ˜𝑐𝑚 − 𝑐𝑚 , we can express the update as: ⊤ ⊤ 𝑀 𝐴′ = 𝑀 𝐴 + 𝑑𝑒⊤ 𝑚 + 𝑒𝑚 𝑑 − 𝑑 𝑚 𝑒𝑚 𝑒𝑚 , where 𝑑𝑚 denotes the 𝑚-th coordinate of 𝑑 . This can be factorized into 𝑀 𝐴′ = 𝑀 𝐴 + 𝑈𝑉 ⊤ , where 𝑈 = 𝑒𝑚 𝑑 − 21 𝑑𝑚 𝑒𝑚 , 𝑉 = 𝑑 − 12 𝑑𝑚 𝑒𝑚 𝑒𝑚 . This rank-2 factorization allows us to apply the Woodbury matrix identity: ( 𝑀 𝐴 + 𝑈𝑉 ⊤ ) −1 = 𝑀 𝐴−1 − 𝑀 𝐴−1𝑈 ( 𝐼2 + 𝑉 ⊤ 𝑀 𝐴−1𝑈 ) −1𝑉 ⊤ 𝑀 𝐴−1 . Because 𝑉 ⊤ 𝑀 𝐴−1𝑈 is only a 2 × 2 matrix, evaluating this update is extremely fast. We can plug this inverted matrix directly back into Eq. (12), reducing the computational complexity of evaluating a neighbor to just O (| 𝐴 | 2 ) instead of O (| 𝐴 | 3 ). Once the optimal batch 𝐴★ is found, we can proceed in two different ways: (i) Myo: We evaluate the target model on all items in 𝐴★ simultaneously. Batch evaluation allows us to maximize throughput when multiple evaluations can be processed in parallel; or (ii) Non-Myo: We select a single item from it at random. Because this item is selected based on an optimal future horizon rather than just the immediate next step, this renders the active selection non-myopic. Figure 4 presents the RMSE Log, RMSE Gain, and Cost-Accuracy Trade-Off for GPQA with binary (left two columns) and continuous (right two columns) scores. We evaluate the selection strategies outlined above using batch sizes of | 𝐴 | = 2 and | 𝐴 | = 5. We observe that the overall performance of these batch approaches is similar to that of the single-item selection strategy. Table 2 compares the Evaluation-Weighted Integrated RMSE achieved by BayesAME with singlecandidate selection, random selection, bath selection with | 𝐴 | = 2 using both myopic (Myo) and non-myopic (Non-Myo) criteria, and Bayes-AIPW across the different benchmarks. Consistent with the results of Figure 4, BayesAME with batch size | 𝐴 | = 2 exhibits performance similar to that of myopic single-candidate selection under both the myopic and non-myopic variants. 21
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean
50%
100
10 1
10 2
90%
101
0
20
40
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean
100
10 1
60
80
100
10 2 101
0
20
40
101
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean
100
10 1
60
80
100
10 2 101
0
20
40
10 1
60
80
10 2
100
101
100
100
10 1
10 1
10 1
10 1
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation
40
60
80
10 2
100
6
6
6
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
20
40
60
2
80
20
6
40
60
2
80
20
6
40
60
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
0
0
0
0
1
1
1
1
40
60
2
80
20
Interpolation
40
60
2
80
20
40
60
50%
0.4 0.2
90%
70
75
80
85
0.6 0.4 0.2 90
95
100
70
0.8
0.8
0.6
0.6
0.4
0.4
0.2
0.2
70
75
80
85
90
95
100
70
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean 80
85
0.6 0.4 0.2 90
80
85
90
40
60
80
100
20
95
100
20 1.0
95
100
20
60
80
40
60
80
0.8
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean 30
40
50
60
0.6 0.4 0.2 70
80
90
100
20 1.0
0.8
0.8
0.6
0.6
0.4
0.4
20
40
1.0
0.2 75
20
Extrapolation
0.8
75
0
Interpolation
0.8
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean
100
2
80
1.0
0.6
80
1
Extrapolation
0.8
60
6
5
20
40
2
80
5
2
20
Extrapolation
5
6
0
Interpolation
6
2
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean
100
100
Interpolation
50%
Extrapolation
101
100
10 2
90%
Interpolation
101
BayesAME BayesAME |A|=2 Myo BayesAME |A|=2 Non-Myo BayesAME |A|=5 Myo LARS BayesAME |A|=5 Myo DPP BayesAME |A|=5 Myo Clust BayesAME-RS RS-Mean 30
40
50
60
70
80
90
100
30
40
50
60
70
80
90
100
0.2 30
40
50
60
70
80
90
100
20
Figure 4 | Single-target Setting. GPQA with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top two rows) and RMSE Gain (central two rows), and Cost-Accuracy Trade-Off (bottom two rows) across 50%, and 90% of reference models (percentages with unique reference scores 𝐵 = 𝑁 ). B.2. Metadata-Weighted Prior Distribution When rich metadata (such as, model size, training data, or hyperparameter configurations) is available for both the reference and target models, we can formulate a weighted prior distribution. Specifically, we define a weight vector 𝑤 = ( 𝑤1 , . . . , 𝑤𝐾𝑟 ), where 𝑤𝑘 ≥ 0 captures the prior positive correlation between reference model Í 𝐾𝑟 𝑘 and the target model, with 𝑘=1 𝑤𝑘 = 1. We then consider the prior distribution 𝑝 ( 𝜃) = N ( 𝜇 𝜃 , Σ𝜃 ) with 𝜇 𝜃𝑏 = 𝜙𝑏⊤ 𝑤,
𝜃 Σ𝑏,𝑏 ′ = 𝛼 exp
𝛽 − ( 𝜙𝑏 − 𝜙𝑏′ ) ⊤𝑊 ( 𝜙𝑏 − 𝜙𝑏′ ) , 𝐾𝑟
where 𝑊 = diag( 𝑤). In practice, public leaderboards rarely provide sufficiently detailed metadata to calibrate these weights reliably, so we assign uniform weights across the reference models 𝑤𝑘 = 1/ 𝐾𝑟 .
22
BayesAME: Bayesian Active Model Evaluation
Benchmarks
BayesAME
BayesAME-RS
BayesAME |A|=2 Myo
BayesAME |A|=2 Non-Myo
Bayes-AIPW
Binary GPQA
Interpolation Extrapolation
61.54 68.91
66.38 72.65
62.08 69.43
61.51 69.50
67.21 73.41
MMLU-Pro
Interpolation Extrapolation
30.18 37.20
30.18 37.20
-
-
30.68 37.69
BBH
Interpolation Extrapolation
31.70 38.35
36.35 42.02
30.51 38.47
30.98 38.66
35.54 41.34
ARC-Challenge
Interpolation Extrapolation
36.53 36.56
36.53 36.56
-
-
36.68 38.03
MuSR
Interpolation Extrapolation
47.86 54.19
47.86 54.19
-
-
49.18 56.72
IFEval
Interpolation Extrapolation
49.53 58.08
56.14 58.08
49.99 -
50.00 -
56.71 60.52
GPQA
Interpolation Extrapolation
30.58 40.40
34.42 40.75
30.84 41.17
29.88 40.76
35.21 46.77
MMLU-Pro
Interpolation Extrapolation
13.11 26.82
17.10 31.10
13.14 26.24
13.32 26.60
17.19 31.33
BBH
Interpolation Extrapolation
13.04 29.80
17.54 35.57
12.81 30.62
13.13 30.49
18.01 33.64
ARC-Challenge
Interpolation Extrapolation
26.04 33.16
26.47 33.16
-
-
26.19 33.05
MuSR
Interpolation Extrapolation
27.83 36.96
33.42 42.56
27.94 35.70
27.43 36.57
34.34 43.69
Natural QA Openbook
Interpolation Extrapolation
28.26 20.83
29.68 28.36
28.41 20.43
28.89 20.79
29.12 27.70
Continuous
Table 2 | Evaluation-Weighted Integrated RMSE for 50% reference models. A dash - denotes that we are in a non-unique reference regime. Therefore, selection defaults to random and the entry is omitted. B.3. Logit-Normal Prior Distribution While achieving strong empirical performance, a Gaussian prior distribution on 𝜃 is not the best choice for binary and bounded continuous scores from a modeling viewpoint, as it assigns positive probability mass outside the true support. Below, we describe more appropriate alternative formulations that we considered.
Case 𝑠𝑖 ∈ [0, 1] We introduce a logit-normal model that respects the support of scores. Specifically, we map both the latent abilities and the observed scores to the unconstrained real line using the logit transformation. We define the latent variable vector 𝑓 = ( 𝑓1 , . . . , 𝑓 𝐵 ), where 𝑓𝑏 = logit( 𝜃𝑏 ) = ln 1−𝜃𝑏𝜃𝑏 , and assign a Gaussian prior distribution on it. The scores are similarly transformed, and we assume a Gaussian conditional distribution 𝑝 (logit( 𝑠𝑖 ) | 𝑓𝑏 ) = N ( 𝑓𝑏 , 𝑁𝑏 𝜎2 ) for an item 𝑖 belonging to bucket 𝑏. The posterior distribution of 𝑓 = ( 𝑓1 , . . . , 𝑓 𝐵 ) conditioned on 𝑠 C is given by: Ö 𝑝( 𝑓 | 𝑠C ) ∝ 𝑝( 𝑓 ) 𝑝 (logit( 𝑠𝑖 ) | 𝑓 ) . 𝑖∈ C
Under this specification, 𝑝 ( 𝑓 | 𝑠 C ) is itself Gaussian and inference over 𝑓 remains fully tractable.
Case 𝑠𝑖 ∈ {0, 1} Similarly to above, consider the latent variable vector 𝑓 = ( 𝑓1 , . . . , 𝑓𝑏 ), where 𝑓𝑏 = logit( 𝜃𝑏 ), and assign a Gaussian prior distribution on it. In the case of binary responses, a more appropriate model would use a Bernoulli likelihood of the form 𝑝 ( 𝑠𝑖 | 𝑓 ) = 𝜎 ( 𝑓𝑏 ) 𝑠𝑖 (1 − 𝜎 ( 𝑓𝑏 )) 1− 𝑠𝑖 ,
for an item 𝑖 belonging to bucket 𝑏. Then, the posterior distribution of 𝑓 = ( 𝑓1 , . . . , 𝑓 𝐵 ) conditioned on 𝑠 C is given by: Ö 𝑝( 𝑓 | 𝑠C ) ∝ 𝑝( 𝑓 ) 𝑝 ( 𝑠𝑖 | 𝑓 ) . 𝑖∈ C
23
BayesAME: Bayesian Active Model Evaluation
50%
Interpolation
Interpolation
Extrapolation
101
101
101
100
100
100
100
BayesAME BayesAME-RS BayesAME-Logit BayesAME-Logit-RS
10 1
10 2 101
90%
Extrapolation
101
0
20
40
BayesAME BayesAME-RS BayesAME-Logit BayesAME-Logit-RS
10 1
60
80
100
10 2 101
0
20
40
BayesAME BayesAME-RS BayesAME-Logit BayesAME-Logit-RS
10 1
60
80
100
10 2 101
0
20
40
60
80
100
10 2 101
100
100
100
100
10 1
10 1
10 1
10 1
10 2
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
40
BayesAME BayesAME-RS BayesAME-Logit BayesAME-Logit-RS
10 1
60
80
100
10 2
0
20
40
60
80
100
0
20
40
60
80
100
Figure 5 | Single-target Setting. GPQA with binary scores (left two columns) and continuous scores (right two columns). RMSE Log across 50% and 90% of reference models. In this case, we can resort to variational inference or a Laplace approximation (Rasmussen and Williams, 2006) to compute the posterior. However, we are interested in the posterior of 𝜃 = 𝜎 ( 𝑓 ) = (1 + 𝑒− 𝑓 ) −1 , conditioned on 𝑠 C . The posterior mean can be computed as: ∫ 𝔼[ 𝜃𝑏 | 𝑠 C ] =
𝜎 ( 𝑓𝑏 ) 𝑝 ( 𝑓 | 𝑠 C ) d 𝑓 ,
which does not admit a closed-form solution because of the nonlinearity of 𝜎. We consider three strategies for approximating this integral. • Monte Carlo estimation. The main drawback is that obtaining accurate estimates for all 𝑁 items requires a large number of samples, which is prohibitive inside the sequential active learning loop. • Other numerical integration methods. We can approximate 𝜎 ( 𝑓𝑏 ) in the integral with a rescaled probit function that has the same slope at the origin (MacKay, 1992), obtaining: √︁ 𝔼[ 𝜃𝑏 | 𝑠 C ] ≈ 𝜎 𝔼[ 𝑓𝑏 | 𝑠 C ] 1 + 𝜋 Var( 𝑓𝑏 | 𝑠 C )/8 . • Linear approximation. Following Osborne et al. (2012), we linearize the logistic term around a reference point 𝑓𝑏,0 , that is, 𝜎 ( 𝑓𝑏 ) = 𝜎 ( 𝑓𝑏,0 ) + 𝜎 ( 𝑓𝑏,0 ) (1 − 𝜎 ( 𝑓𝑏,0 )) ( 𝑓𝑏 − 𝑓𝑏,0 ) Under this approximation, the posterior mean of 𝜃 admits a closed-form expression. To select the reference point, we introduce an auxiliary multivariate Gaussian model in the original space and set 𝜎 ( 𝑓𝑏,0 ) to be its posterior mean given 𝑠 C , i.e., 𝜎 ( 𝑓𝑏,0 ) = 𝔼[ 𝜃𝑏 | 𝑠 C ], clipping the value to [0, 1] if necessary to remain within the valid range. Although this auxiliary Gaussian model is not appropriate as it does not have the correct range, it is used solely to inform the choice of 𝑓𝑏,0 . While these alternative formulations offer a more principled treatment of the score support, they require approximate inference at every step and introduce substantial computational overhead. In practice, we found that these approximations resulted in worse performance. For example, Figure 5 compares BayesAME and BayesAME-RS with their corresponding versions using logit-normal models (BayesAME-Logit and BayesAME-Logit-RS) on GPQA for both binary and continuous scores, using the linear approximation described above to derive the posterior for 𝜃. For both settings and regardless of the selection strategy, BayesAME and BayesAME-RS outperform the logit-normal formulation.
24
BayesAME: Bayesian Active Model Evaluation
C. Experimental Details and Additional Results C.1. Implementations Details BayesAME For the single-target setting, we use Algorithm 1 with 𝜎2 = 10−4 when 𝐵 ≥ 𝑁 − 𝑁0 and 𝜎2 = 10−1 otherwise, 𝑁0 = 10, 𝑊 = 30. As described in Appendix A.4, optimization is performed using L-BFGS-B with 𝛼0 = 𝛽0 = 1, and 𝐹 = 50. For the multi-target setting, we set 𝑁0 = 10, and 𝑊 = 30. Of the 𝐿 covariances, half are squared exponential (Eq. (1)) and the rest are Matérn-3/2. As described in Appendix A.4, optimization is performed using AdamW, with 𝛼𝑙 and 𝛽𝑙 initialized to 1, each 𝑤𝑖𝑡 to a sample from N (0, 0.01), and 𝜎2 to 0.1 for binary scores and to 0.01 for continuous scores. We use 𝐹 = 50, and learning rate 𝜂 = 0.01 for binary scores and 𝜂 = 0.001 for continuous scores. We set the margin 𝑁0 = 10 to safely handle regimes where reference scores are nearly unique across items. Empirically, we identified a specific failure mode for active selection: when almost all items have unique combinations of reference scores (effectively forming singleton buckets), but one single bucket contains a significantly larger number of items, the active selection biases the posterior estimation and degrades performance. Enforcing 𝑁0 = 10 prevents the algorithm from relying on active selection in this skewed regime, yielding stable results across all scenarios considered in this work. While fixing 𝑁0 = 10 serves as a safe value, because the total number of buckets and the exact distribution of items across them are fully known to the algorithm, a more sophisticated future approach could explicitly incorporate this distributional information to set more suitable, benchmark-specific values. We conjecture that larger threshold values may be admissible when bucket sizes are strictly balanced, as the observed estimation bias appears to be driven by the presence of a single, severely unbalanced bucket. When selecting items in batches, we scale the window size 𝑊 proportionally. Specifically, for a batch size | 𝐵 |, we set 𝑊 = 30/| 𝐵 |.
Baselines For Bayes-AIPW, Bayes-RS-Learn, and Seq-APW, we adapt the implementations of AIPW, Random-Sampling-Learn, and APW, respectively, provided by Zhang et al. (2025) in https://github.com /socialfoundations/benchmark-prediction. In Bayes-AIPW and Bayes-RS-Learn, we replace the linear ridge regression with its Bayesian counterpart as implemented in scikit-learn library using the default setting (Pedregosa et al., 2011). For ProEval, we use the implementation provided by Huang et al. (2026) in https://github.com/google-deepmind/proeval. For IRT, we implement the binary and continuous approaches based on the methodology outlined in Polo et al. (2024) and Chen et al. (2019), respectively.
C.2. Additional Benchmark Details For BBH, we only consider a subset of scenarios: temporal sequences, salient translation error detection, tracking shuffled objects (seven objects), geometric shapes, and reasoning about colored objects. For MMLU-Pro, we also restrict our evaluation to a subset of subjects, namely business and law, resulting in a total of 1890 items.
C.3. Additional Results Single Target Setting Figure 6 shows that our Bayesian extension of AIPW matches or outperforms the original method of Zhang et al. (2025). We present extended results for all benchmarks considered. To facilitate navigation of the plots, Table 3 summarizes the figures corresponding to each benchmark and evaluation metric.
25
BayesAME: Bayesian Active Model Evaluation
Plot Type
GPQA
MMLU-Pro
BBH
ARC-Challenge
MuSR
IFEval
Natural QA Openbook
RMSE Log, RMSE Gain & Cost-Accuracy Tradeoff
7
8
9
10
11
12
12
Spearman Correlation
13
–
13
–
–
–
–
Table 3 | Figure index for the extended experimental results across all benchmarks. Interpolation 6 5
10%
4
4
4 3
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
20
40
60
80
2 6
20
40
60
80
2 6
20
40
60
80
2 6
5
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2
2
2
2
20
40
60
80
6
20
40
60
80
6
20
40
60
80
6
5
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
20
40
60
80
2
20
Interpolation 6 5 4
10%
5
2
2
40
60
80
2
20
Extrapolation 6
BayesAIPW AIPW RS-Mean
5 4
40
60
80
2
6
BayesAIPW AIPW RS-Mean
5 4
6
BayesAIPW AIPW RS-Mean
5 4
3
3
3
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2
2
2
2
40
60
80
6
20
40
60
80
6
20
40
60
80
6
5
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2 6
20
40
60
80
2 6
20
40
60
80
2 6
20
40
60
80
2 6
5
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2
2
2
2
20
40
60
80
20
40
60
20
40
60
80
20
40
60
80
40
60
80
20
Extrapolation
2
20
BayesAIPW AIPW RS-Mean
Interpolation
3
6
50%
4 3
6
90%
5
Extrapolation 6
BayesAIPW AIPW RS-Mean
3
6
50%
5
Interpolation 6
BayesAIPW AIPW RS-Mean
3
2
90%
Extrapolation 6
BayesAIPW AIPW RS-Mean
80
20
40
60
80
BayesAIPW AIPW RS-Mean
20
40
60
80
20
40
60
80
20
40
60
80
Figure 6 | Single-target Setting. GPQA (top three rows) and MMLU-Pro (bottom three rows) with binary scores (left two columns) and continuous scores (right two columns). RMSE Log for the AIPW ablation study across varying proportions of reference models (10%, 50%, and 90%) ablation for AIPW. C.4. Multi-target Setting Figure 14 show the results comparing BayesAME Multi-target with BayesAME Single-target on GPQA using both binary and continuous scores. Figure 15 reports results for the alternative selection strategy introduced in Section 3.2, which selects items based on the average marginalized information gain across target models.
26
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation 6
BayesAME BayesAME-RS BayesAIPW RS-Mean
3
10%
101
10 1
4
4 3
60
80
100
10 2
Interpolation 6
BayesAME BayesAME-RS BayesAIPW RS-Mean
5
40
5 4 3
Extrapolation 6
BayesAME BayesAME-RS BayesAIPW RS-Mean
5 4 3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2
2
2
2
20
6
50%
10 2
100
5
40
60
80
20
6
40
60
80
20
6
40
60
80
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
20
6
40
60
2
80
20
6
40
60
2
80
20
6
40
60
2
80
6
5
5
5
5
4
4
4
4
3
3
3
3
2
2
2
2
1
1
1
1
0
0
0
0
1
1
1
1
2
2
2
2
20
40
60
80
20
Interpolation 0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 0.1 0.0
80
85
0.1 95
100
0.7
0.0
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0.2
0.1
0.1 80
85
90
95
100
0.0
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0.2
0.1 0.0
80
20
80
85
60
80
95
0.8
0.6
0.6
100
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean 50
60
0.2 70
80
60
80
20
40
60
80
40
60
80
Extrapolation
0.8
0.2
40
20
Interpolation
0.4
90
40
20
90
100
BayesAME BayesAME-RS BayesAIPW RS-Mean 50
60
70
80
90
100
50
60
70
80
90
100
50
60
70
80
90
100
0.7
0.6
0.0
60
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2
90
40
Extrapolation
BayesAME BayesAME-RS BayesAIPW RS-Mean
6
5
2
90%
100
100
Interpolation
10%
80
100
6
50%
60
100
10 2
90%
40
85
90
95
100
0.0
0.8
0.6
0.6
0.4
0.4
0.2 80
85
90
95
100
0.2 50
60
70
80
90
100
0.8
0.8
0.6
0.6
0.4
0.4
0.2
0.1 80
0.8
80
85
90
95
100
0.2 50
60
70
80
90
100
Figure 7 | Single-target Setting. GPQA with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 27
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
10%
101
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
10 1
6
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
40
60
80
100
10 2
Interpolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
4
4
4
4
2
2
2
2
0
0
0
20
10
50%
10 2
100
8
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
20
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
40
60
80
20
Interpolation 1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.0
BayesAME BayesAME-RS BayesAIPW RS-Mean 50
60
0.2 70
80
60
20
90
100
1.0
0.0
0.8
0.6
0.6
0.4
0.4
0.2
0.2
0.0
0.0
50
60
70
80
90
100
1.0
1.0
0.8
0.8
60
80
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.9
0.8
0.8
0.7
0.7
0.6
0.6
50
60
0.2 70
80
90
100
0.1
0.4
20
40
0.3 0.2 60
80
100
60
70
80
90
100
0.1 0.9
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2 50
40
60
80
100
0.1
0.9
0.9
0.8
0.8
0.7
0.7 0.6
0.6
0.6
0.5
0.5
0.4
0.4
0.4
0.4
0.3
0.3
0.0
0.2 50
60
70
80
90
100
0.0
0.2 50
60
70
80
90
100
0.1
40
60
80
40
60
80
BayesAME BayesAME-RS BayesAIPW RS-Mean 20
40
60
80
100
20
40
60
80
100
20
40
60
80
100
0.2 20
0.6
0.2
20
0.5
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.9
0.1
80
Extrapolation
0.9
0.3
60
20
0.5
1.0
0.8
40
Interpolation
0.4
40
0
80
Extrapolation
1.0
0.2
40
20
0 20
10
8
20
BayesAME BayesAME-RS BayesAIPW RS-Mean
0 20
10
8
10
90%
100
100
Interpolation
10%
80
100
10
50%
60
100
10 2
90%
40
0.2 20
40
60
80
100
0.1
Figure 8 | Single-target Setting. MMLU-Pro with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 28
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
10%
101
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
10 1
6
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
40
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
60
80
100
10 2
Interpolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
6
4
4
4
2
2
2
2
0
0
0
20
40
60
80
20
10
40
60
80
0 20
10
40
60
80
10
8
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
20
10
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
40
60
80
20
Interpolation 1.0
0.8
0.8
0.6
0.6
0.2
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean 50
60
0.2 70
80
90
100 1.0
0.8
0.8
0.6
0.6
0.4 0.2 60
70
80
90
90
100
20
30
40
0.2 50
60
70
80
90
100
0.4
0.2
0.2
0.4 0.2
60
70
80
90
100
30
40
50
60
70
80
90
100 1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.4
0.2
0.2 60
70
80
80
40
60
80
90
100
20
30
40
50
60
70
80
90
100
20
30
40
50
60
70
80
90
100
20
30
40
50
60
70
80
90
100
0.2 20
1.0
50
60
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.4
0.6
50
40
0.6
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.4
0.6
100
0.8
0.4
0.8
90
1.0
0.8
100
20
Extrapolation
0.8
0.2 80
80
20
1.0
0.4
70
60
Interpolation
0.6
0.6
80
80
0.8
0.8
70
60
1.0
1.0
60
40
1.0
1.0
50
20
BayesAME BayesAME-RS BayesAIPW RS-Mean 60
40
0
80
0.6
50
1.0
50
60
Extrapolation
1.0
0.4
40
20
0 20
10
8
20
BayesAME BayesAME-RS BayesAIPW RS-Mean
8
4
10
50%
10 2
100
8
90%
100
100
Interpolation
10%
80
100
10
50%
60
100
10 2
90%
40
0.2 20
30
40
50
60
70
80
90
100
Figure 9 | Single-target Setting. BBH with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 29
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
10%
101
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
10 1
6
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
40
60
80
100
10 2
Interpolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
6
4
4
4
2
2
2
2
0
0
0
20
40
60
80
20
10
40
60
80
0 20
10
40
60
80
10
8
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
20
10
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
40
60
80
20
Interpolation 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
60
60
70
70
70
80
80
80
60
90
90
90
100
100
100
0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
20
0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0
40
60
80
0.8 0.7
0.6
0.6
0.5
0.5
70
0.1
80
90
100
0.0
0.3
70
80
60
80
40
60
80
30
40
50
0.1 60
70
80
90
100
90
100
90
100
0.0
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0.2
0.0
30
40
50
60
70
80
90
100
30
40
50
60
70
80
90
100
30
40
50
60
70
80
90
100
0.1 30
40
50
60
70
80
90
100
0.0
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
0.2
0.0
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2
0.1 60
40
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.1 60
20
Extrapolation
0.7
0.2 80
80
20
0.8
0.3
BayesAME BayesAME-RS BayesAIPW RS-Mean 70
60
Interpolation
0.4
60
40
0
80
Extrapolation
BayesAME BayesAME-RS BayesAIPW RS-Mean 60
40
20
0 20
10
8
20
BayesAME BayesAME-RS BayesAIPW RS-Mean
8
4
10
50%
10 2
100
8
90%
100
100
Interpolation
10%
80
100
10
50%
60
100
10 2
90%
40
0.1 30
40
50
60
70
80
90
100
0.0
Figure 10 | Single-target Setting. ARC-Challenge with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 30
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
10%
101
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
10 1
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
6
6
60
80
100
10 2
Interpolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8
40
8 6
Extrapolation 10
BayesAME BayesAME-RS BayesAIPW RS-Mean
8 6
4
4
4
4
2
2
2
2
0
0
0
20
10
50%
10 2
100
8
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
20
40
60
80
20
10
40
60
80
40
60
80
10
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
40
60
80
20
40
Interpolation
60
20
Extrapolation
40
60
80
0.8
0.8
0.7
0.7
0.7
0.6
0.6
0.6
0.6
0.5
0.5
0.5
0.5
0.4
0.4
0.4
0.2 0.1 0.8
70
75
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 0.1 80
85
90
95
100
0.8
70
75
0.3 0.2 0.1 80
85
90
95
100
50
60
0.3 0.2 0.1 70
80
90
100
0.8
0.8
0.7
0.7
0.7
0.6
0.6
0.6
0.6
0.5
0.5
0.5
0.5
0.4
0.4
0.4
0.4
0.3
0.3
0.3
0.3
0.2
0.2
0.2
0.1
0.1
0.8
75
80
85
90
95
100
0.8
75
80
85
90
95
100
60
70
80
90
100
0.8
0.8
0.7
0.7
0.7
0.6
0.6
0.6
0.6
0.5
0.5
0.5
0.5
0.4
0.4
0.4
0.4
0.3
0.3
0.3
0.3
0.2
0.2
0.2
0.2
0.1
0.1
0.1
75
80
85
90
95
100
70
75
80
85
90
95
100
60
80
40
60
80
BayesAME BayesAME-RS BayesAIPW RS-Mean 50
60
70
80
90
100
50
60
70
80
90
100
50
60
70
80
90
100
0.1 50
0.7
70
40
0.2
0.1 70
20
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.7
70
80
Extrapolation
0.8
0.3
60
20
Interpolation
0.7
BayesAME BayesAME-RS BayesAIPW RS-Mean
40
0
80
0.8
0.3
20
0 20
10
8
20
BayesAME BayesAME-RS BayesAIPW RS-Mean
0 20
10
8
10
90%
100
100
Interpolation
10%
80
100
10
50%
60
100
10 2
90%
40
0.1 50
60
70
80
90
100
Figure 11 | Single-target Setting. MuSR with binary scores (left two columns) and continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 31
BayesAME: Bayesian Active Model Evaluation
Interpolation
Extrapolation
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
10%
100
10 1
10 2 101
Interpolation
101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
Extrapolation
101
0
20
101
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
0
20
BayesAME BayesAME-RS BayesAIPW ProEval Seq-APW Bayes-RS-Learn IRT RS-Mean
100
10 1
40
60
80
100
10 2 101
100
100
100
10 1
10 1
10 1
10 1
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
50%
100
0
10 2
90%
101
0
20
40
60
80
100
10 2 101
0
20
10%
101
0
20
40
60
80
100
10 2 101
10 1
10 1
10 1
10 1
8
0
20
40
60
80
100
10 2
0
20
40
60
80
100
10 2
0
20
Extrapolation 12
BayesAME BayesAME-RS BayesAIPW RS-Mean
10 8
40
60
80
100
10 2
Interpolation 12
BayesAME BayesAME-RS BayesAIPW RS-Mean
10 8
Extrapolation 12
BayesAME BayesAME-RS BayesAIPW RS-Mean
10 8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
2
20
12
50%
10 2
100
10
40
60
2
80
20
12
40
60
20
12
40
60
2
80
12
10
10
10
8
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
20
12
40
60
2
80
20
12
40
60
20
12
40
60
2
80
12
10
10
10
10
8
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
2
20
40
60
2
80
20
Interpolation 1.2
1.0
1.0
0.8
0.8
0.6
0.2
60
70
75
0.4 0.2 80
85
90
95
100 1.2
1.0
1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.2
0.2 75
80
85
90
95
100 1.2
1.0
1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.2 80
85
90
95
100
2
80
75
80
85
90
95
100
1.0 0.8
0.0 20
80
85
90
95
70
75
80
85
90
95
100
0.4
30
40
0.2 50
60
70
80
90
100
0.0 20
1.0
1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.0 20
20
40
60
80
40
60
80
20
BayesAME BayesAME-RS BayesAIPW RS-Mean 30
40
50
60
70
80
90
100
30
40
50
60
70
80
90
100
30
40
50
60
70
80
90
100
0.2 30
40
50
60
70
80
90
100
0.0 20
1.0
1.0
0.8
0.8
0.6
0.6
0.4
0.4
0.2
0.2
0.0 100 20
80
0.6
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 75
60
Extrapolation
0.8
0.2
0.2 75
60
1.0
0.4
BayesAME BayesAME-RS BayesAIPW RS-Mean
70
1.2
70
40
0.6
70
1.2
70
20
Interpolation
0.6
BayesAME BayesAME-RS BayesAIPW RS-Mean
40
0
2
80
Extrapolation
1.2
0.4
40
20
0
2
80
BayesAME BayesAME-RS BayesAIPW RS-Mean
0
2
80
10
2
90%
100
100
Interpolation
10%
80
100
12
50%
60
100
10 2
90%
40
30
40
50
60
70
80
90
100
0.0 20
Figure 12 | Single-target Setting. IFEval with binary scores (left two columns) and Natural QA Openbook with continuous scores (right two columns). RMSE Log (top three rows), RMSE Gain (middle three rows), and Cost-Accuracy Trade-Off (bottom three rows) across varying proportions of reference models. 32
BayesAME: Bayesian Active Model Evaluation
10%
Interpolation
1.0
0.8
0.8
0.8
0.8
0.6
0.6
0.6
0.6
0.4
0.4
0.4
50%
1.0
BayesAME BayesAME-RS 0.2 BayesAIPW RS-Mean 0
20
40
60
80
100
0.0 1.0
BayesAME BayesAME-RS BayesAIPW RS-Mean 0
20
60
80
100
0.0 1.0
0
20
40
60
80
100
0.0 1.0
0.8
0.8
0.8
0.6
0.6
0.6
0.6
0.4
0.4
0.4
0.4
0.2
0.2
0.2
0.2
0.0
0.0
0.0
0
20
40
60
80
100
1.0
0
20
40
60
80
100
1.0
0
20
40
60
80
100
0.0 1.0
0.8
0.8
0.8
0.8
0.6
0.6
0.6
0.6
0.4
0.4
0.4
0.4
0.2
0.2
0.2
0.0
0
20
40
60
80
100
0.0
0
20
Interpolation
10%
40
0.4
BayesAME BayesAME-RS 0.2 BayesAIPW RS-Mean
0.2
0.8
1.0
40
60
80
100
0.0
20
Extrapolation
40
60
80
100
0.0
1.0
0.8
0.8
0.8
0.8
0.6
0.6
0.6
0.6
0.4
0.4
0.4
1.0
0
20
40
60
80
100
0.0 1.0
BayesAME BayesAME-RS BayesAIPW RS-Mean 0
20
40
60
80
100
0.0 1.0
0
20
40
60
80
100
0.0 1.0
0.8
0.8
0.8
0.6
0.6
0.6
0.4
0.4
0.4
0.4
0.2
0.2
0.2
1.0
20
40
60
80
100
0.0 1.0
0
20
40
60
80
100
0.0 1.0
20
40
60
80
100
0.0 1.0
0.8
0.8
0.8
0.6
0.6
0.6
0.6
0.4
0.4
0.4
0.4
0.2
0.2
0.2
0.2
0.0
0.0
0.0
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
80
100
BayesAME BayesAME-RS BayesAIPW RS-Mean 0
20
40
60
80
100
0
20
40
60
80
100
0
20
40
60
80
100
0.2 0
0.8
0
60
0.4
0.6
0
40
BayesAME BayesAME-RS 0.2 BayesAIPW RS-Mean
0.2
0.8
0.0
20
Extrapolation
1.0
0.0
0
Interpolation
1.0
BayesAME BayesAME-RS 0.2 BayesAIPW RS-Mean
BayesAME BayesAME-RS BayesAIPW RS-Mean
0.2 0
1.0
0.2
50%
Extrapolation
1.0
0.0
90%
Interpolation
1.0
0.2
90%
Extrapolation
1.0
0
20
40
60
80
100
0.0
Figure 13 | Single-target Setting. GPQA (top three rows) and BBH (bottom three rows) with binary scores (left two columns) and continuous scores (right two columns). Spearman’s rank correlation across varying proportions of reference models.
33
BayesAME: Bayesian Active Model Evaluation
Low correlation
High correlation
Low correlation
High correlation
101
101
101
101
100
100
100
100
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
10 2 10 8
20
30
40
50
60
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
70
80
90
100
10 2 10 8
20
30
40
50
60
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
70
80
90
100
10 2 10 8
20
30
40
50
60
70
80
90
100
10 2 10 8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
0
2 10
2 10
2 10
2 10
20
30
40
50
60
70
80
90
0.8
0.8
0.6
0.6
0.4
20
30
40
50
60
70
80
90
0.4
0.2
0.2
20
30
40
50
60
70
80
90
0.8
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.3
0.3
0.2
85
90
95
100
80
85
90
95
20
30
40
50
60
70
80
90
20
30
40
50
60
70
80
90
80
90
100
0.2
0.1 80
BayesAME Single-target BayesAME Multi-target BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
0.1
100
50
60
70
80
90
100
50
60
70
100
Figure 14 | Multiple-target Setting. GPQA with binary scores (left two columns) and with continuous scores (right two columns). The rows represent RMSE Log, RMSE Gain, and Cost-Accuracy Trade-Off, respectively. We compare the multitarget method against methods that do not model correlations across targets.
Low correlation
High correlation
101
100
20
30
40
50
60
70
90
100
10 2 10 10
101
100
BayesAME Single-target BayesAME Multi-target BayesAME Multi-target Avg. Marg. IG BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
80
High correlation
101
100
BayesAME Single-target BayesAME Multi-target BayesAME Multi-target Avg. Marg. IG BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
10 2 10 10
Low correlation
101
20
30
40
50
60
70
100
BayesAME Single-target BayesAME Multi-target BayesAME Multi-target Avg. Marg. IG BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
80
90
100
10 2 10 10
20
30
40
50
60
70
80
90
100
10 2 10 10
8
8
8
8
6
6
6
6
4
4
4
4
2
2
2
2
0
0
0
10
20
30
40
50
60
70
80
90
10
20
30
40
50
60
70
80
90
10
30
40
50
60
70
80
90
10
0.8
0.7
0.7
0.6
0.6
0.5
0.5
0.4
0.4
0.4
0.3
0.3
0.2
0.2
0.2
0.8
0.6
0.6
0.4 0.2
0.1 0.0
80
85
90
95
100
0.0
80
85
90
95
100
20
30
40
50
60
70
80
90
20
30
40
50
60
70
80
90
100
0 20
0.8 0.8
BayesAME Single-target BayesAME Multi-target BayesAME Multi-target Avg. Marg. IG BayesAME-RS Single-target BayesAIPW RS-Mean
10 1
0.1 50
60
70
80
90
100
50
60
70
80
90
100
Figure 15 | Multiple-target Setting. GPQA with binary scores (left two columns) and with continuous scores (right two columns). The rows represent RMSE Log, RMSE Gain, and Cost-Accuracy Trade-Off, respectively. We compare the different selection strategies in the multitarget setting.
34