CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation Mengting Chen1 , Yanshu Sun1 , Wanting Liang1 , Beidi Luan2 , Rui Sun2 , Dezhi Chen1 , Jing Li2 , Zuo Bai2,1∗ 1
FinStep StepFun [email protected], [email protected] 2
arXiv:2607.29252v1 [cs.CL] 31 Jul 2026
Abstract Reliable evaluation of open-ended LLM outputs requires finegrained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge unanimity and binary variance filters, which cannot distinguish measurable rubrics from informative ones. We introduce CalibratedRubric, a task-adaptive framework that combines type-specific scoring, Bayesian rubric-measurability filtering, and item response theory (IRT)-based bank assembly. CalibratedRubric estimates each rubric’s measurability with a Beta–Bernoulli agreement posterior and uses a submodular information-coverage objective to construct compact rubric banks over the observed capability range. Across financial, healthcare, general, and legal benchmarks, measurability filtering improves human-gold agreement on JudgmentBench from κ = 0.604 to 0.743. IRT-based greedy selection improves cross-fitted rank fidelity over random selection across all six evaluated response blocks and requires only 49 rather than 131 rubrics to reach the target correlation on FinResearchBench decision-support tasks. Task-label perturbations further reduce system separation, confirming the practical relevance of task-adaptive scoring. These results support CalibratedRubric as an efficient, uncertainty-aware approach to open-ended LLM evaluation, with calibration gains depending on sufficient judge redundancy.
1
Introduction
LLM-based autonomous agents increasingly adopt decoupled, service-oriented designs with dedicated memory, retrieval-augmented generation (RAG), and tool-use layers, enabling long-form outputs over extended interaction horizons (Wang et al. 2024a; Xi et al. 2025). This shift makes conventional automatic metrics inadequate: lexical-overlap measures such as BLEU and ROUGE and embedding-based measures such as BERTScore were designed for comparatively short, reference-bounded outputs and correlate weakly with expert judgment on open-ended deliverables such as synthesized financial research reports (Papineni et al. 2002; Lin 2004; Zhang et al. 2020). The LLM-as-a-Judge paradigm offers a scalable alternative, but holistic scores conflate distinct quality dimensions and remain susceptible to length, position, and self-preference biases (Li et al. 2024; Zheng et al. 2023; Liu et al. 2023; Wang et al. 2024b). Analytic eval∗
Corresponding author.
uators therefore decompose quality into skill-level, rubricconditioned, multidimensionally calibrated, collaboratively aligned, or atomic-checklist judgments (Ye et al. 2024; Kim et al. 2024; Hashemi et al. 2024; Chiang et al. 2026; Cook et al. 2024). Two structural bottlenecks limit current analytic evaluators. First, reliable rubric authoring still depends heavily on scarce domain expertise. For example, constructing a financial benchmark to professional examination standards requires experts to draft, refine, and validate criteria, making annotation difficult to scale (Liu et al. 2026). Second, automated methods can synthesize instance-specific criteria (Wang and Blanco 2026; Jia et al. 2026; Li et al. 2026), but evaluation items vary substantially in informativeness (Rodriguez et al. 2021), and reliable selection remains largely heuristic. Prior work uses IRT to build NLP evaluation scales, compare test sets, and shrink benchmarks (Lalor, Wu, and Yu 2016; Vania et al. 2021; Maia Polo et al. 2024); we instead embed IRT inside rubric construction and use the item information function to select criteria over the relevant capability range (Birnbaum 1968; Lord 1980; Baker 2001). The consensus-derived pipeline we extend retains only rubrics on which all judges agree and applies a binary variance filter (Luan et al. 2026). This favors reproducibility but treats judges as equally reliable and criteria as equally informative. A binary filter removes only completely constant criteria and cannot quantify how much the remaining ones distinguish systems. We propose CalibratedRubric, replacing fixed deterministic pipelines with task-adaptive, probabilistically calibrated evaluation. A task-typing front end maps each query to a cognitive task type—evidence reasoning, decision support, high-risk constraint, or divergent creation—and assigns an appropriate rubric form and aggregation rule. A Bayesian calibration module models judge-specific leniency and reliability as latent parameters, replacing strict unanimity with robust posterior consensus anchored by a small human-labeled subset. Finally, we embed item response theory (IRT) inside rubric construction: candidate rubrics are treated as items, and the item information function (IIF) selects criteria that maximize information over the relevant capability range (Birnbaum 1968; Lord 1980; Baker 2001). Our contributions are fourfold: (1) We introduce a taskadaptive taxonomy that dynamically selects rubric forms and
scoring rules based on cognitive demands. (2) We deploy IRT and the IIF as an information-maximizing rubric-selection mechanism, producing system capability estimates with calibrated confidence intervals. (3) We develop a Bayesian judge-calibration procedure that models bias and reliability, establishing robust posterior consensus. (4) Across general and specialized domains, we demonstrate that the framework preserves human-expert alignment while dramatically compressing evaluation costs and improving the resolution of closely matched systems.
2 2.1
Problem Setup and Theoretical Analysis Problem Formulation
M Let Q = {qn }N n=1 be the queries and M = {mi }i=1 the evaluated systems, with output rin = mi (qn ). For each query, a generator produces aFquery-specific pool Cn of natural-language criteria; C = n Cn and n(j) denotes the (k) query of criterion j. Judge k ∈ {1, . . . , K} assigns yij ∈ Sj to output ri,n(j) , and ȳij is the aggregated label. We use zj ∈ {0, 1} for reproducible judgeability: whether competent graders can apply criterion j consistently. This is necessary but not sufficient for substantive expert endorsement, a gap evaluated on the human-labelled subset. Without additional elicitation, judge k supplies the validity vote (k) (k) (−k) vj = 1[yij = ȳij ∀i], where ȳ (−k) is the leave-onejudge-out majority. Given a retained set G ⊆ C and non-negative weights wj , system i receives P N X P j∈G∩Cn wj gj (ȳij ) P S(mi ) = ωn , n ωn = 1, j∈G∩Cn wj n=1 (1) where gj maps the response scale to [0, 1]. The within-query normalization prevents queries with larger candidate pools from dominating the score. Gold-set construction must balance validity, discriminative information, and parsimony, since judging costs O(KM |G|). We therefore solve
G ⋆ =arg
max I(G) s.t. min Pr(zj = 1 | Y ) ≥ τc ,
G⊆C, |G|≤k
j∈G
(2) where I is test information and τc is a validity threshold. Both quantities are latent and are estimated from judge labels and a small human anchor set.
2.2
Unanimity Attrition
Assumption 1 (Judge error model). Conditional on the expert label of an (item, deliverable) pair, judges label independently, and judge k returns the expert label with probability 1 − εk . Proposition 1 (Exponential attrition). Under Assumption 1 with homogeneous error rate ε, for any item j the unanimity rule (3) satisfies iM h Pr δjcons = 1 = (1 − ε)K + εK =: ρ(ε, K)M . (5) The factor ρM suppresses ambiguous items, but it arbitrarily removes informative and uninformative criteria at the same rate, worsening as the leaderboard grows. Proposition 2 (Reliability-aware consensus). Let judge k − have sensitivity s+ k and specificity sk (Dawid and Skene 1979). The validity posterior is the weighted vote X h (k) s+ Pr(zj =1 | Y ) π0 k log vj log = log + Pr(zj =0 | Y ) 1 − π0 1 − s− k k +i 1 − s (k) k + 1 − vj log , s− k (6) Thresholding (6) is Bayes-optimal for 0–1 validity loss. Thus posterior consensus decouples retention from literal unanimity, allowing reliable judges to carry appropriate weight. Appendix B.5 details the diagnostic protocol and identifiability conditions.
2.3
Variance Filtering
Under the two-parameter logistic (2PL) model (Birnbaum 1968; Lord 1980), the probability that system i satisfies binary criterion j is Pij = Pr(yij =1 | θi ) = σ aj (θi − bj ) , (7) where σ(u) = (1 + e−u )−1 , and aj and bj are the discrimination and difficulty of criterion j. The item information function and the test information of a set G are X Ij (θ) = a2j Pj (θ) 1−Pj (θ) , IG = Ij , (8) j∈G
are both one.
and the standard error of the ability estimate is SE(θ̂) = IG (θ)−1/2 . We instantiate the objective in (2) as the information delivered where the evaluated systems actually lie, R I(G) = IG (θ)φ(θ) dθ, with φ the estimated capability distribution of M. Remark 1 (Degenerate information filter). Let p̂j = P 1 ȳ . Then ij i M h i d ȳ·j > 0 = 1 Iˆj > 0 , δjdisc = 1 Var (9)
The two hard rules correspond to the constraint and objective of (2), respectively. We next show why both are crude approximations; derivations and diagnostics appear in Appendix B.
where Iˆj = a2 p̂j (1 − p̂j ) assumes common discrimination and ignores system ability. Hence the baseline is the zerothreshold, ability-blind special case of information selection.
Definition 1 (Consensus-derived baseline). The baseline (Luan et al. 2026) uses uniform weights and retains item j iff h i (k) (k′ ) δjcons = 1 yij = yij ∀i ∈ M, ∀k, k ′ ∈ K , (3) δjdisc = 1[ ∃ i, i′ : ȳij = 1, ȳi′ j = 0 ] ,
(4)
Table 1: Task-adaptive scoring rules used in the main experiments. Task type Scoring Evidence Decision High-risk Creative
Task emphasis
binary pass = 1, fail = 0 weighted binary 5× actionability risk-adjusted −0.3/ − 0.5 violations weighted binary 3× creativity
This heuristic cannot rank surviving criteria, distinguish ability-consistent from aberrant response patterns, or target difficulties to the capability range. The IIF supplies all three and supports a fixed budget. Proposition 3 (Strict refinement). Let the item parameters be estimated by marginal maximum likelihood with a proper prior on b, let items with constant observed responses be assigned Iˆj ≡ 0, and let GIIF (k) = R arg max|G|≤k, âj >0 IG (θ)φ(θ)dθ. For any k ≤ |{j : d ·j ) > 0}|, GIIF (k) ⊆ {j : δ disc = 1}, with strict Var(ȳ j inclusion whenever k is smaller than the number of surviving items or some surviving item has âj ≤ 0. Proposition 3 states the relationship we claim precisely: information-based selection never retains a criterion the baseline heuristic rejects, and prunes further among those it accepts. This predicts that the resulting ranking agrees with the baseline ranking while requiring strictly fewer criteria to reach the same separation, the hypothesis we test experimentally. The screen âj > 0 is a sign convention, not a tuned threshold: it removes items whose fitted response function decreases in ability, which contribute rank noise rather than information. We use it at 0 throughout and never tune it, so the budget B and the validity threshold τc remain the only quantities set per dataset. On a purely binary block the feasible sets can coincide; any gain then comes from ranking items and stopping early, not from admitting new items. We therefore test budgeted separation rather than claim a binary validity gain. The proof and tightness conditions are in Appendix B.3.
2.4
Task-Adaptive Measurement
Binary criteria are adequate for verifying whether an evidential step is present; they are a poor instrument for grading the quality of a recommendation, and they cannot express the asymmetric cost of a compliance violation. We therefore introduce a cognitive task type t ∈ T , where T consists of four categories—evidence-reasoning, decision-support, safety-critical and divergent-creative—together with a classifier Pr(t | q) and a map from task type to a triple of response scale, measurement model and selection objective, summarized in Table 1. The mapping changes both the response scale and the objective. Ordered responses can be more efficient when quality is genuinely graded, whereas safety-critical tasks call for expected violation loss rather than symmetric information. Table 1 defines the comprehensive taxonomy of this task-adaptive design. In our main empirical evaluations, we
rigorously demonstrate the maximum-information selection mechanism, proving its high robustness, scalability, and compression efficiency across diverse domains. Further theoretical motivation for the expanded branches appears in Appendix B.4. Assumptions and Scope. The analysis assumes a dominant latent dimension, local item independence, conditionally independent judges, and identifiable ability estimates. We stabilize small-M estimation with priors, compare 1PL and 2PL explicitly, and report bootstrap uncertainty. Residual-spectrum and judge-correlation checks are reported experimentally; the full assumptions and diagnostics are collected in Appendix B.5.
3
CalibratedRubric
Section 2 formulates rubric-bank construction as the constrained program in (2). CalibratedRubric estimates its two ingredients from judge outputs: posterior rubric measurability and discriminative information. The former replaces literal unanimity with a probabilistic retention rule, while the latter supports budgeted rubric selection and uncertainty-aware capability estimation. Figure 1 summarizes the pipeline. We take the candidate pools C as given and optimize which rubrics survive, how many are needed, and how they are weighted. The method is generator-agnostic but cannot recover dimensions absent from the candidate pool. Its stages type each task, align rubric scales, estimate measurability and information, and assemble a budgeted rubric bank.
3.1
Task Typing
A prompted classifier estimates Pr(t | qn ) over the four types in §2.4. The query-level prediction determines the shared response scale and measurement family for Cn through Table 1. If maxt Pr(t | qn ) < τt , the method falls back to binary maximum-information scoring; ordered responses are examined only in the GRM sensitivity analysis.
3.2
Candidate Rubrics
Each candidate is a (predicate, scale) pair. Scale mismatches are resolved by the deterministic map of §2.4: graded items are binarized at their midpoint when required, and binary items become two-level graded items otherwise. Selection occurs only in (2); B and τc are the dataset-level controls.
3.3
Bayesian Rubric Calibration
Existing LLM-judge calibration methods primarily correct response-level selection or positional biases (Wang et al. 2024b; Li et al. 2025). Ours instead estimates rubric-level measurability from repeated inter-judge agreement. (j) For rubric j, let nagree denote the number of evaluated (j) instances on which all available judges agree and let ntotal denote the total number of evaluated instances. We place a Beta prior on its latent measurability probability, qj ∼ Beta(α, β),
(10)
Figure 1: CalibratedRubric. Task typing determines the scoring rule; Bayesian rubric calibration supplies the measurability constraint; IRT supplies item information; and budgeted assembly returns a compact rubric bank with uncertainty-aware system scores.
which yields the posterior (j) (j) qj | Y ∼ Beta α + n(j) agree , β + ntotal − nagree .
(11)
The posterior mean (j)
q̂j = E[qj | Y ] =
α + nagree (j) α + β + ntotal
(12)
provides a soft alternative to literal unanimity. We use the uniform prior α = β = 1 and retain rubrics satisfying q̂j ≥ τc ; their instance-level labels are then aggregated by majority voting. This estimator requires no human labels and is the calibration procedure evaluated in Section 4. Proposition 2 describes a judge-specific extension for settings with sufficient expert anchors; it is not instantiated in the present experiments. With only two judges, unanimity coincides with pairwise agreement and therefore provides no independent rubric-level calibration signal.
3.4
Information. With the fitted parameters, the per-item scalar objective is the information delivered where the systems actually lie, Z X νj = Ij (θ) φ(θ) dθ, I(G) = νj , (14) j∈G
where φ is the fitted ability density. For ordered items, Ij is the graded-response information; this branch is used only as a sensitivity analysis. We report residual-spectrum concentration and bootstrap-calibrated item fit. Confirmatory MIRT, Yen’s Q3 (Yen 1984), and dependent-item refitting are left to future work.
3.5
Rubric Assembly
The measurability gate defines
Item Calibration
Model and estimation. Binary items use logistic IRT and ordered items use the graded response model (Samejima 1969). For binary data we compare 1PL and 2PL fits. With ∆ℓ and ∆p denoting the richer model’s gain and additional parameters, define κ = ∆ℓ/∆p,
EAP ability estimation (Bock and Aitkin 1981). We impose θ ∼ N (0, 1), aj > 0, and log aj ∼ N (0, 0.52 ); the last constraint prevents perfectly separating items from acquiring divergent discrimination and dominating the selector. Constant-response items receive zero information.
(13)
AIC selects 2PL when κ > 1 and BIC when κ > 12 log Nobs ; we report both because item-specific slopes are weakly identified on small leaderboards. Parameters are estimated by marginal MAP on 41 nodes over [−4, 4], followed by
F = {j : q̂j ≥ τc , Ij ̸≡ 0},
(15)
where q̂j is the posterior measurability score in (12). Plain IIF, analogous to classical information-based test assembly (van der Linden and Glas 2000), statically ranks feasible items by G X I¯j = πg Ij (θg ) (16) g=1
and takes the top B, where πg are normalized standardnormal quadrature weights. It does not account for information already supplied by selected items.
Our final assembler instead optimizes the concave information-coverage utility U (S) =
G X
πg log(1 + IS (θg )) ,
IS (θg ) =
g=1
X
Ij (θg ).
j∈S
(17) Starting from S = ∅, IIF-Greedy repeatedly selects j ⋆ = arg max ∆U (j | S),
(18)
j∈F \S
∆U (j | S) =
G X
πg log 1 +
g=1
Ij (θg ) 1 + IS (θg )
,
and updates S ← S ∪ {j ⋆ } until |S| = B. Proposition 4 (Submodular assembly guarantee). The utility in (17) is normalized, monotone non-decreasing, and submodular. Under the cardinality constraint |S| ≤ B, the greedy rule (18) satisfies U (Sgreedy ) ≥ (1 − 1/e)
max
S⊆F , |S|≤B
U (S).
(19)
Because test information is modular and log(1 + x) is increasing and concave, the utility is normalized, monotone, and submodular; the standard greedy guarantee follows (Nemhauser, Wolsey, and Fisher 1978). The denominator in (18) discounts already-covered ability regions, inducing diversity without an explicit penalty. The guarantee is conditional on local independence. Small-M interpretation. When 1PL is selected, νj is maximized by difficulties near the centre of the fitted ability distribution. Thus the defensible gains on small leaderboards are difficulty targeting, budget awareness, and coverage, rather than precise recovery of item-specific discrimination. Weights and scores. The retained items are scored through (1) with gj the scale normalizer of §2.1 and weights wj ∝ νj ,
(20)
so informative rubrics contribute more, while posterior measurability remains solely the feasibility constraint defining F . This separation prevents measurability from being counted twice and isolates the contribution of information weighting. Uncertainty and tiers. Query-stratified item bootstrap yields percentile intervals for θ̂i . Adjacent systems whose ability difference is not significant are collapsed into a tier, avoiding unsupported fine-grained rankings. Cost and Amortization. Initial calibration costs O(KM J) judge calls, IRT fitting costs O(iter M JG), and greedy assembly costs O(BJG) for G = 41 cached nodes. A new system is then evaluated on |G ⋆ | items rather than the full J, converting repeated full-pool evaluation into a one-off calibration cost. We therefore report the smallest bank that preserves the full-pool ranking, not only fixed-budget accuracy.
Algorithm 1: CalibratedRubric Require: responses Y , candidate pools {Cn }, budget B, threshold τc Ensure: rubric bank G ⋆ , weights w, capability tiers 1: type each query and align rubric scales 2: compute posterior measurability scores q̂j 3: fit regularized IRT and evaluate Ij (θg ) 4: F ← {j : q̂j ≥ τc , Ij ̸≡ 0} 5: S ← ∅ 6: for b = 1, . . . , B do 7: j ⋆ ← arg maxj∈F \S ∆U (j | S) 8: S ← S ∪ {j ⋆ } 9: end for 10: G ⋆ ← S; set wj ∝ νj 11: bootstrap rubrics and report capability tiers
4
Experiment
This section evaluates the task-adaptive front end of CalibratedRubric on four datasets: FinResearchBench-v2 (Luan et al. 2026), HealthBench (Arora et al. 2025), HelloBench (Que et al. 2024), and JudgmentBench (Yang et al. 2026).
4.1
Task-Adaptive Scoring
Experimental Setup. This section evaluates the taskadaptive front end of ConsensusMultiRubric on four datasets: FinResearchBench-v2 (hereafter FinResearchBench) as the main ranking domain, HealthBench and HelloBench as transfer settings, and JudgmentBench as a legal-domain validation set with human annotations and output-quality strata. Although the full collection contains 5,781 tasks and 65,648 rubrics, the current E1 snapshot uses 534 tasks and 7,959 rubrics with available responses (Appendix A). Each task is mapped to one of four cognitive types— evidence reasoning, decision support, high-risk constraint, or creative divergent—which determines the scoring rule. The current pipeline removes invalid judge outputs, applies discrimination filtering, and yields 5,077 rubric entries after valid-output pruning and 2,813 final gold rubrics. Adaptive scoring uses task-specific rules, whereas the baseline applies a uniform binary scorer. Results and Analysis. The current main experiment yields 2,813 gold rubrics and ranks 21 systems, with Gemini first, Research Report last, and a top-bottom gap of 94.26 percentage points. The strongest evidence for task typing comes from causal ablation: when 30% of task labels are corrupted, the top-bottom gap drops from 94.86 to 83.95, a degradation of 10.91 points. Beyond causal usefulness, adaptive scoring remains ranking-stable and yields moderate discrimination gains, especially in robust distributional metrics such as IQR and standard deviation. On FinResearch, both adaptive and binary baselines achieve the same external validity against the human reference ranking (ρ = 0.8833; Appendix A). On JudgmentBench, rubric-based scores recover a monotonic quality trend across tiers, even though within-task finegrained ordering remains weak.
Flip Avg Score Gap (pp) 0.0 0.1 0.2 0.3
0.4413 0.4419 0.4284 0.3922
Deg.
94.86 baseline 94.86 0.00 86.97 -7.89 83.95 -10.91
Table 2: Causal ablation under perturbed task-type assignments. Table 3: Agreement after L1 filtering at θ = 0.80. “Kept” reports retained rubrics where available. Task
Data
Gold
Kept
Base
L1
r
Decision FinResearch LLM 286/624 0.8513 0.9708 0.589 Evidence FinResearch LLM 81/195 0.8829 0.9732 0.558 High-risk Judgment Human — 0.6040 0.7430 0.127
Overall, E1 shows that task-adaptive scoring is operationally stable, causally meaningful, and moderately beneficial for discrimination.
4.2
Bayesian Judge Calibration
L1 estimates a posterior measurability score for each rubric using the Beta–Bernoulli posterior mean E[qr ] =
α + nagree , α + β + ntotal
(21)
where nagree counts instances on which all available judges agree, ntotal is the number of evaluated instances, and α = β = 1. Rubrics satisfying E[qr ] ≥ θ are retained and evaluated by majority voting. Importantly, L1 selection uses only inter-judge agreement and requires neither human labels nor a held-out gold judge. We sweep θ ∈ {0.50, 0.60, 0.65, 0.70, 0.80} against unfiltered majority voting. Agreement increases monotonically as the threshold becomes stricter, at the expected cost of lower coverage. From θ = 0.50 to 0.80, κ rises from 0.8885 to 0.9708 for FinResearch decision support, from 0.8994 to 0.9732 for evidence reasoning, and from 0.661 to 0.743 against human gold on JudgmentBench. Table 3 reports the strictest setting. The FinResearch subsets further show that L1 preferentially removes unreliable rubrics rather than merely reducing coverage. For decision support, the bottom and top measurability quartiles have mean κ values of 0.498 and 0.983; the corresponding values for evidence reasoning are 0.556 and 0.953. Consistently, posterior measurability correlates with agreement in both subsets (r = 0.589 and 0.558). The weaker correlation on the predominantly high-risk JudgmentBench (r = 0.127) suggests that measurability is less predictive under human–LLM disagreement, although filtering still improves human-gold agreement from 0.604 to 0.743. L1 requires at least three judges to provide an informative rubric-level signal: with two judges, unanimity is identical to pairwise agreement and cannot independently
calibrate reliability. Accordingly, HealthBench and HelloBench show no additional gain under their two-judge configurations. Evaluation coverage also remains limited for evidence-reasoning and creative-divergent tasks. Moreover, LLM judges have higher positive-label rates (55.6–62.9%) than the JudgmentBench human gold (47.1%), indicating a systematic judge–human mismatch that measurability filtering does not fully eliminate.
4.3
Submodular IRT Test Assembly and Capability Calibration
Experimental Setup. We evaluate six response blocks covering all four task types, with 254–2,142 rubrics per block. The FinResearch blocks contain 15 systems judged by three LLMs; three transfer blocks contain six systems and four judges; JudgmentBench contains 1,314 quality-stratified outputs evaluated by a five-member human–LLM panel. Binary labels are majority aggregated, while ordered labels are used only for GRM sensitivity analysis. Each block uses the regularized 2PL assembler in §3.4, with 41 nodes on [−4, 4] and log aj ∼ N (0, 0.52 ). We compare IIF-Greedy against plain IIF, unrestricted random selection, and random sampling within the hard-filtered set. AIC, BIC, held-out likelihood, a 100-replicate item-fit bootstrap, and 300 item bootstraps diagnose model fit and ability uncertainty. Rank fidelity is Spearman’s ρ against the full usable block, summarized as AUC over 12 logarithmic budgets. For sample-out evaluation, rubrics are randomly half-split: selection and refitting use half A, while the reference ability vector is estimated from half B. We use 20 splits and three random draws per split; the target minimum-bank correlation is .95, or .9429 for six systems. Results and Analysis. AIC and BIC select 1PL in five of six blocks; JudgmentBench alone splits (AIC: 2PL; BIC: 1PL), while held-out likelihood selects 3PL only for the creative block. Thus the regularized 2PL assembler is best interpreted as an operational mechanism for difficulty targeting and coverage over the observed ability range. Bootstrap item misfit ranges from 0.23% to 18.91%, with decision support the clearest warning case; confirmatory MIRT and Q3 remain future diagnostics. On binary blocks, non-constancy and the hard discrimination filter define the same feasible set, so the substantive question is compression rather than set overlap. The hard filter is especially brittle on heterogeneous outputs, retaining zero high-risk and four creative items. Table 4 shows that IIF-Greedy improves AUC over unrestricted random selection in every block and over hardset sampling whenever defined; every reported interval excludes zero. Relative to plain IIF, however, the additional gain is significant only for the two 15-system FinResearch blocks: +.0057 [.0034, .0079] for evidence reasoning and +.0080 [.0059, .0101] for decision support. Hence the evidence supports the coverage mechanism most clearly at the larger leaderboard size, not a universal advantage. At the target correlation, FinResearch requires 28 rather than 51 random items for evidence reasoning and 49 rather than 131 for decision support. Across the remaining blocks,
Task type
Response pool
Evidence reasoning Evidence reasoning Decision support High-risk constraint High-risk constraint Creative divergent
FinResearch LLM systems FinResearch LLM systems JudgmentBench LLM systems
Greedy Hard Random ∆Hard (95% CI) ∆Random (95% CI) .918 .708 .912 .801 .315 .578
.868 .463 .847 — .214 .476
.852 .369 .828 .379 .205 .494
.051 [.040, .062] .245 [.173, .317] .065 [.058, .072] — .101 [.092, .110] .110 [.012, .208]
.066 [.056, .075] .339 [.298, .380] .083 [.076, .091] .422 [.355, .489] .110 [.101, .118] .084 [.006, .161]
Table 4: Cross-fitted rank-fidelity AUC. Deltas are paired across 20 splits (19 for creative versus hard); “Random” samples all non-constant items. Block
Task
IRT
Bayes
B0 – – – Signal: Raw binary baseline. Interpretation: Uniform binary scoring without calibration or selection. B1 ✓ – – Signal: Gap: 94.86 → 83.95 under 30% label corruption. Interpretation: Task labels affect system separation while largely preserving ranking structure. B2 ✓ ✓ – Signal: 131 → 49 rubrics at the target correlation. Interpretation: Information-aware selection substantially reduces the required rubric bank. B3 ✓ – ✓ Signal: κ: 0.604 → 0.743 on JudgmentBench. Interpretation: Measurability filtering improves agreement. B4 ✓ ✓ ✓ Signal: Gap = 97.13 on FinResearch; agreement = 0.8587 on JudgmentBench. Interpretation: The full pipeline can be operationally composed across representative benchmarks.
that task typing is causally useful without destabilizing ranking structure. B2 provides the most consistent improvement across datasets, confirming that information-aware rubric assembly is the strongest and most general source of gain in the full pipeline. B3 improves agreement when judge redundancy is sufficient, but contributes less in sparse two-judge settings. B4 verifies that the composed pipeline can be instantiated across heterogeneous benchmarks. On FinResearchBench, B4 attains a 97.13-point top-bottom gap; on JudgmentBench, it reaches 0.8587 agreement with human-majority labels. On HealthBench and HelloBench, the same composition remains transferable, but these benchmarks primarily function as sparse-judge boundary cases rather than as the main evidence for Bayesian gains. Overall, the ablation shows that B1 establishes causal usefulness, B2 provides the strongest and most stable improvement, and B3 adds agreement gains when the judge structure is sufficiently redundant.
5 Table 5: Summary of component configurations and their representative empirical signals.
greedy requires 43–273 items versus 142–1,226 under random selection. Even on challenging domains like JudgmentBench, IIF-Greedy robustly outperforms baselines in rank fidelity, demonstrating submodular IRT as a highly effective engine for scalable and calibrated LLM evaluation. Gemini leads both FinResearch blocks, with θ̂ = 3.802 (95% CI [3.615, 3.907]) and 4.302 ([4.187, 4.427]); bootstrap comparisons recover four and six tiers. The smaller system blocks collapse to one or two tiers, and only 9.81% of JudgmentBench output pairs are separated, so we avoid fine-grained leaderboards there. GRM sensitivity is mixed and self-preference tests are underpowered. Overall, IRT is supported as a rubric-compression and uncertainty-reporting mechanism, but not as evidence of universal 2PL identifiability or bias-free judging.
4.4
Ablation Study
We compare five nested configurations: the binary baseline (B0), task-adaptive scoring (B1), B1 with IRT selection (B2), B1 with Bayesian filtering (B3), and the full pipeline (B4). All configurations reuse the same judge outputs, isolating post-judgment components without additional LLM calls. Table 5 reveals a clear division of labor. B1 establishes
Conclusion
We introduced CalibratedRubric, a task-adaptive framework for constructing compact, consensus-derived rubric banks without placing experts in the full evaluation loop. It replaces rigid unanimity and variance filters with posterior judge consensus and IRT-based submodular test assembly, while task typing makes the response scale and scoring objective explicit. Corrupting task labels reduces downstream discrimination, posterior measurability filtering improves agreement on every three-judge block, and IIF-greedy improves crossfitted rank-fidelity AUC over random selection across all six response blocks. On FinResearch it also reaches the target correlation with substantially fewer rubrics. The main benefit is therefore a more reliable and economical evaluation instrument, not a universal ranking reversal. Future Work. This paper establishes the core foundation of probabilistic rubric assembly. Exciting future directions include extending the Bayesian calibration to model correlated, multi-facet judge biases and active human-in-theloop calibration. Additionally, exploring adversarial candidate generation will further expand capability coverage, while applying our continuous and risk-weighted objectives to safety-critical domains with explicitly elicited violation costs promises to unlock new paradigms for rigorous LLM generation safety evaluation.
References Arora, R. K.; Wei, J.; Hicks, R. S.; Bowman, P.; QuiñoneroCandela, J.; Tsimpourlas, F.; Sharman, M.; Shah, M.; Vallone, A.; Beutel, A.; Heidecke, J.; and Singhal, K. 2025. HealthBench: Evaluating Large Language Models Towards Improved Human Health. arXiv:2505.08775. Baker, F. B. 2001. The Basics of Item Response Theory. ERIC Clearinghouse on Assessment and Evaluation, 2 edition. ISBN 1-886047-03-0. Birnbaum, A. 1968. Some Latent Trait Models and Their Use in Inferring an Examinee’s Ability. In Lord, F. M.; and Novick, M. R., eds., Statistical Theories of Mental Test Scores, 397–479. Addison-Wesley. Bock, R. D.; and Aitkin, M. 1981. Marginal Maximum Likelihood Estimation of Item Parameters: Application of an EM Algorithm. Psychometrika, 46(4): 443–459. Chiang, C.; Gebreegziabher, S.; Szymanski, A.; Yang, Y.; Do, H. J.; Ashktorab, Z.; Geyer, W.; Li, T.; and Gomez-Zara, D. 2026. MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria. In Proceedings of the 5th Annual Symposium on Human-Computer Interaction for Work, 1–17. Association for Computing Machinery. Cook, J.; Rocktäschel, T.; Foerster, J.; Aumiller, D.; and Wang, A. 2024. TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation. arXiv:2410.03608. Dawid, A. P.; and Skene, A. M. 1979. Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1): 20–28. Hashemi, H.; Eisner, J.; Rosset, C.; Van Durme, B.; and Kedzie, C. 2024. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13806–13834. Association for Computational Linguistics. Jia, M.; Zhang, Z.; Cases, I.; Liu, Z.; Jiang, M.; and Qi, P. 2026. AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning. In Findings of the Association for Computational Linguistics: ACL 2026, 25707– 25724. Association for Computational Linguistics. Kim, S.; Shin, J.; Cho, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, S.; Kim, S.; Thorne, J.; and Seo, M. 2024. Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models. In The Twelfth International Conference on Learning Representations. Lalor, J. P.; Wu, H.; and Yu, H. 2016. Building an Evaluation Scale using Item Response Theory. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 648–657. Association for Computational Linguistics. Li, H.; Chen, J.; Ai, Q.; Chu, Z.; Zhou, Y.; Dong, Q.; and Liu, Y. 2025. CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges. In Proceedings
of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16537–16552. Association for Computational Linguistics. Li, J.; Sun, S.; Yuan, W.; Fan, R.-Z.; Zhao, H.; and Liu, P. 2024. Generative Judge for Evaluating Alignment. In The Twelfth International Conference on Learning Representations. Li, S.; Zhao, J.; Ren, H.; Wei, Z.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Wei, C. 2026. RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 31320–31344. Association for Computational Linguistics. Lin, C.-Y. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, 74–81. Association for Computational Linguistics. Liu, X.; Ma, X.; Ma, Y.; Peng, Y.; Wang, D.; Wen, Z.; Zhang, G.; Zhang, K.; Chen, X.; Ding, Y.; et al. 2026. XpertBench: Expert Level Tasks with Rubrics-Based Evaluation. arXiv:2604.02368. Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2511–2522. Association for Computational Linguistics. Lord, F. M. 1980. Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates. ISBN 0-89859-006-X. Luan, B.; Sun, R.; Wang, S.; Gu, Y.; Li, C.; Xiong, Z.; Li, J.; and Bai, Z. 2026. FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality. arXiv:2607.12252. Maia Polo, F.; Weber, L.; Choshen, L.; Sun, Y.; Xu, G.; and Yurochkin, M. 2024. tinyBenchmarks: Evaluating LLMs with Fewer Examples. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 34303–34326. PMLR. Nemhauser, G. L.; Wolsey, L. A.; and Fisher, M. L. 1978. An Analysis of Approximations for Maximizing Submodular Set Functions—I. Mathematical Programming, 14(1): 265–294. Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–318. Association for Computational Linguistics. Que, H.; Duan, F.; He, L.; Mou, Y.; Zhou, W.; Liu, J.; Rong, W.; Wang, Z. M.; Yang, J.; Zhang, G.; Peng, J.; Zhang, Z.; Zhang, S.; and Chen, K. 2024. HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models. arXiv:2409.16191. Rodriguez, P.; Barrow, J.; Hoyle, A. M.; Lalor, J. P.; Jia, R.; and Boyd-Graber, J. 2021. Evaluation Examples Are Not Equally Informative: How Should That Change NLP Leaderboards? In Proceedings of the 59th Annual Meeting of the
Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 4486–4503. Association for Computational Linguistics. Samejima, F. 1969. Estimation of Latent Ability Using a Response Pattern of Graded Scores. Psychometrika, 34(S1): 1–97. van der Linden, W. J.; and Glas, C. A. W., eds. 2000. Computerized Adaptive Testing: Theory and Practice. Kluwer Academic Publishers. Vania, C.; Htut, P. M.; Huang, W.; Mungra, D.; Pang, R. Y.; Phang, J.; Liu, H.; Cho, K.; and Bowman, S. R. 2021. Comparing Test Sets with Item Response Theory. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 1141–1158. Association for Computational Linguistics. Wang, L.; Ma, C.; Feng, X.; Zhang, Z.; Yang, H.; Zhang, J.; Chen, Z.; Tang, J.; Chen, X.; Lin, Y.; et al. 2024a. A Survey on Large Language Model Based Autonomous Agents. Frontiers of Computer Science, 18(6): 186345. Wang, P.; Li, L.; Chen, L.; Cai, Z.; Zhu, D.; Lin, B.; Cao, Y.; Kong, L.; Liu, Q.; Liu, T.; and Sui, Z. 2024b. Large Language Models Are Not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9440–9450. Association for Computational Linguistics. Wang, Z.; and Blanco, E. 2026. Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge. arXiv:2605.30568. Xi, Z.; Chen, W.; Guo, X.; He, W.; Ding, Y.; Hong, B.; Zhang, M.; Wang, J.; Jin, S.; Zhou, E.; et al. 2025. The Rise and Potential of Large Language Model Based Agents: A Survey. Science China Information Sciences, 68(2): 121101. Yang, R.; Chen, R.; Kelaita, P.; Ranjan, R.; Ma, S.; Dickens, C.; Guillod, M.; Ma, M.; and Nyarko, J. 2026. JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment. arXiv:2605.25240. Ye, S.; Kim, D.; Kim, S.; Hwang, H.; Kim, S.; Jo, Y.; Thorne, J.; Kim, J.; and Seo, M. 2024. FLASK: Fine-Grained Language Model Evaluation Based on Alignment Skill Sets. In The Twelfth International Conference on Learning Representations. Yen, W. M. 1984. Effects of Local Item Dependence on the Fit and Equating Performance of the Three-Parameter Logistic Model. Applied Psychological Measurement, 8(2): 125–145. Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations. Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, 46595–46623.
A
E1 Supplementary Tables
Rank System 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
Score
Gemini Doubao Xiaocaishen Pro (202602) Qwen Reportify Gemini Supplement Metaso Xiaocaishen Pro (Old Version) ChatGPT AutoGLM Perplexity Minimax AIChat-Step AIChat-Doubao Research Report
0.9531 0.7086 0.6489 0.6254 0.6246 0.6060 0.5800 0.5757 0.5425 0.4358 0.4171 0.3760 0.3712 0.2701 0.0045
Finance Healthcare General Law
5,000 647 104 30
57,237 3,902 4,052 457
200 200 104 30
2,247 1,203 4,052 457
Total
5,781
65,648
534
7,959
Table 7: Effective experiment scale used in E1.
Key Design
evidence_reasoning binary yes/pass = 1, no/fail = 0 decision_support boosted bin. 5× actionability weight high_risk_constraint risk-aware moderate = -0.3, critical = -0.5 creative_divergent boosted bin. 3× creativity weight
Table 8: Task-adaptive scoring rules in the main E1 experiment.
B B.1
Supplementary Theoretical Analysis Unanimity Attrition
Proof of Proposition 1. For one output, all K judges agree when they are all correct or all wrong, which has probability ρ(ε, K) = (1 − ε)K + εK . Conditional independence across the M outputs gives ρ(ε, K)M . The expression contains no item-discrimination parameter, so the resulting decay is unrelated to measurement information. Corollary 1 (Homogeneous extrapolation). The baseline retention of 25.52% at K = 3, M = 10 (Luan et al. 2026) implies ρ = 0.25521/10 = 0.872 and ε̂ ≈ 4.5%. Holding this rate fixed gives retention of approximately 6.5% at M = 20 and 0.4% at M = 40. agen
Tasks Systems Spearman ρ 104 200 200 30
15 6 6 0
p
0.9714 1.69e-09 0.9429 0.0048 1.0000 0.0000 N/A –
Table 10: Cross-domain ranking stability between adaptive and binary baseline scorers.
Raw T. Raw R. Used T. Used R.
Scoring
Gemini Research Report 94.26 pp 2,813 15
Domain
HealthBench HelloBench FinResearchBench-v2 JudgmentBench
Task Type
Result
Top system Bottom system Top-bottom gap Gold rubrics Ranked systems
Table 9: Main system-level summary statistics.
Table 6: Main E1 ranking on FinResearchBench.
Dataset
Metric
Exact finite-leaderboard curve. Let cj be the number of systems on which the panel is unanimous for item j. For a uniformly sampled subset of m systems, J 1 X cj . Mj b (22) R(m) = m J j=1 m is an unbiased, Monte-Carlo-free estimator of expected consistency retention. We apply it to 843 FinResearchBench II criteria with a complete K = 3 panel over Mj ∈ {11, 12, 13} systems. b A one-parameter fit of log R(m) = m log ρ gives ρ̂ = 0.880 and R2 = 0.85 on the log scale, corresponding to ε̂ = 4.2%. At m = 10, measured consistency retention is 27.3%, close to the published 25.52%. The increasing ρ̂m indicates item heterogeneity. A moment-matched Beta(2.60, 0.65) mixture over per-item agreement rates reduces RMSE from 0.068 to 0.054 and predicts slower, but still monotone, attrition. Meanwhile distinguishability increases with m, making their conjunction non-monotone: the joint gold rate peaks at m = 4 and then falls. Thus benchmark growth changes retention even when criterion quality is fixed.
B.2
Reliability-Aware Consensus
Derivation of Proposition 2. For each judge, the likeli− hood ratio contributed by a positive vote is s+ k /(1 − sk ) and + − by a negative vote is (1 − sk )/sk . Multiplying these ratios and the prior odds π0 /(1 − π0 ), then taking logs, yields (6). The Bayes classifier thresholds these posterior odds. Literal unanimity instead assigns equal influence to all judges and rejects every non-unanimous pattern. It coincides with the Bayes classifier only for parameter and threshold settings that place all such patterns below the posterior threshold; otherwise it has larger Bayes risk.
B.3
IIF Refinement and Its Boundary
Proof of Proposition 3. Under the stated convention, a constant-response item has Iˆj ≡ 0 and cannot enter GIIF .
Spearman ρ
Comparison Adaptive vs. Human Reference Baseline vs. Human Reference
p
0.8833 0.00159 0.8833 0.00159
Table 11: External validity against the FinResearch human reference ranking. Tier
Mean Score
Tier 1 Tier 2 Tier 3
0.5743 0.5833 0.6122
Table 13: Baseline-filter retention versus leaderboard size. The homogeneous prediction uses ρ = 0.880; “disc.” and “joint” denote distinguishability and the conjunction of both filters. b m R(m)
ρm
ρ̂m
disc.
joint
1 2 4 6 8 10 12
0.880 0.774 0.599 0.463 0.358 0.277 0.215
0.801 0.817 0.840 0.857 0.869 0.878 0.890
0.000 0.312 0.565 0.685 0.760 0.814 0.828
0.000 0.171 0.222 0.212 0.196 0.180 0.157
0.801 0.667 0.499 0.395 0.325 0.273 0.248
Table 12: JudgmentBench quality-tier mean scores.
Every feasible selected item therefore has non-constant responses and satisfies δjdisc = 1, proving inclusion. The inclusion is strict whenever the budget is smaller than the baseline survivor set or a surviving item has non-positive fitted discrimination. For binary items, however, the feasible sets can be identical: positive fitted information and non-constant responses represent the same support condition. In that case IIF improves neither validity nor pool coverage. Its contribution is to order survivors by information over the observed ability distribution and to stop once the desired precision is reached. The distinction matters because the binary heuristic is invariant to permutations of system identities: it treats an item passed by the strongest systems exactly like one passed only by the weakest system. IRT separates these patterns through the sign and magnitude of aj and discounts items whose bj lies outside the relevant capability range.
B.4
Why Task Type Changes Measurement
ForR a binary 2PL item, total information over the ability axis is Ij (θ)dθ = aj . A graded item with L ordered categories and separated thresholds can carry up to (L − 1)aj (Samejima 1969; Lord 1980); because SE(θ̂) = I −1/2 , graded responses may achieve a target interval width with fewer judgments when the categories are substantively meaningful. This motivates, but does not itself validate, the typed scale map. Task type can also change the objective. For a safetycritical query, the relevant quantity is the probability that P a costly violation passes undetected, suggesting minG j∈G ℓj Pr(false passj ) under a coverage constraint. The current benchmarks provide neither elicited losses ℓj nor graded responses, so these branches remain specifications for future evaluation rather than tested claims.
B.5
Assumptions and Diagnostics
A1: dominant latent dimension. The 2PL model assumes one dominant capability. We assess this through the ratio of the first two eigenvalues of the residual correlation matrix. A failed diagnostic calls for multidimensional IRT with vectorvalued discrimination; confirmatory MIRT is left to future work.
A2: local independence. Items generated from the same output may remain dependent given θ. Yen’s Q3 (Yen 1984) and dependent-item refitting are the appropriate confirmatory checks, but are not completed in the present experiments. We therefore treat local dependence as a limitation rather than claim that submodular information is perfectly additive. A3: conditional independence of judges. LLM judges share training data and conventions, so correlated errors can inflate apparent consensus and estimated sensitivity or specificity. We examine residual judge correlation on the humanlabelled subset; the posterior remains conditional on this diagnostic. A4: small-M identifiability. Joint 2PL estimation is poorly conditioned with roughly ten systems. We impose weakly informative priors, compare 1PL and 2PL rather than assuming item-specific slopes, constrain θ to zero mean and unit variance, fix the sign of aj , and report bootstrap intervals instead of point estimates alone.