Racer: Role-Aligned Competence Estimation for Human-AI Routing Joshua Strong
1* , Emma Sun 1 , Alexander Capstick 1 , Pramit Saha 1 , Cheng
Ouyang1 , and J. Alison Noble
1
1 Department of Engineering Science, University of Oxford, UK
arXiv:2609.21953v1 [cs.LG] 18 Sep 2026
* Correspondence: [email protected]
Preprint. Work in Progress — September 21, 2026 Work in progress. This preprint reports the current method and preliminary evaluation. Conclusions are restricted to the protocols and baselines reported in this version. Abstract Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be querydependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose Racer—Role-Aligned Competence Estimation for Routing—a role-relative framework for estimating an unseen expert’s competence from context. Racer estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, Racer benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human–AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the Racer family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.
Keywords: learning to defer; human-AI collaboration; calibration; unseen experts; invariance; expert competence
1
Introduction
In high-stakes prediction, the relevant question is often not only what an automated model should predict, but who should make the final decision. Learning to defer (L2D) formalizes this question by jointly training a classifier and a rejector1 : on each input, the system either predicts 1 Jointly training the classifier and rejector is the one-stage L2D setup. In the two-stage setup, the classifier is frozen
and only the rejector is trained [10–13].
1
autonomously or defers to a human expert2 [9, 14, 18]. Under the standard 0–1 system loss, the Bayes rule compares the model’s probability of being correct with the expert’s probability of being correct [14, Eqn. 5]. L2D therefore targets the accuracy of the combined human-AI system rather than model accuracy alone, deferring cases when the expert is expected to be more likely to predict correctly. In healthcare, such decision allocation is part of broader human–AI collaboration [19], with applications including guided deferral for medical-report classification [17]. Classical L2D assumes a fixed expert, but it is plausible to assume that practical deployments may involve rotating clinicians, annotators, or specialists who may be unavailable, newly hired, or drawn from a different site. Population-adaptive L2D methods address this by conditioning the rejector on a small context set of expert behavior for the currently available expert so that the rejector may adapt at test-time [20, 22]. The first approach by Tailor et al. [22] involved a query-conditioned context encoder which compared the current case (query) with similar examples in the expert’s context and could, in principle, route instance by instance. Yet this flexibility comes with a weak inductive bias. Standard population encoders often expose labels and expert predictions in fixed class coordinates, so a rejector can learn policies tied to label identities rather than to transferable expert roles. Here, a role is a candidate true class 𝑦 for which we assess the expert’s competence. Describing competence relative to this role means using relations such as whether a context label matches 𝑦, without tying the estimator to a fixed class name or index. Unseen experts can be in-distribution (ID), drawn from the same population distribution as the training experts, or out-of-distribution (OOD), drawn from a different population distribution. Identity-Free Deferral (IFD) shows that this matters for unseen and OOD experts: by replacing latent expert embeddings with role-indexed classwise competence profiles, IFD enforces invariance to coherent class relabellings [20]. However, IFD’s limitation is the opposite one: a classwise profile is robust, but it is constant within each class and cannot express that an expert may be reliable on some cases of a class and unreliable on others. This paper asks whether one can obtain both desirable properties simultaneously: querydependent adaptation and identity-free transfer. We propose Racer—Role-Aligned Competence Estimation for Routing—a framework for estimating the competence profile of an unseen expert from context using only role-relative information. The central object is a posterior-predictive competence functional Γ(𝑥, 𝑦, 𝐶 𝑒 ) ∈ [0, 1], which estimates how likely the expert represented by context 𝐶 𝑒 is to be correct on query 𝑥 under candidate class role 𝑦. The expert’s overall correctness is then obtained by combining this rolewise competence with the model posterior. Thus Racer treats expert competence as a probability-scale object to be estimated, and not as a hidden feature inside a routing score as in L2D-Pop [22]. At the same time, unlike IFD [20], which estimates role-aligned but classwise competence profiles that are constant within each class, Racer makes competence query-dependent while preserving the same identity-free, role-relative inductive bias. This distinction is important because the Bayes deferral rule compares expert correctness and model correctness on a shared probability scale. A rejector should therefore know not only that the expert looks better than the model on the training distribution, but how likely the expert is to be correct for this query and this role. Racer estimates this quantity with a strictly proper expertcorrectness loss, such as binary cross-entropy, whose population minimizer is the conditional probability of expert correctness. Thus the competence head is trained toward a calibrated probability target, while the final deferral decision can still be optimized with a one-stage L2D surrogate loss. In this sense, Racer preserves the classifier–rejector co-adaptation benefits 2 In this work, we define an expert as a skilled yet potentially biased and fallible decision-maker.
2
of L2D training, but changes the expert interface: the system routes through a role-aligned competence estimate rather than through an unconstrained expert embedding. Empirically, this design yields consistent routing gains in controlled synthetic settings and competitive or best budget-swept performance on the completed real-expert chest-radiography benchmarks. Our contributions are: • We formulate instance-dependent identity-free deferral for unseen experts, combining querydependent adaptation with coherent class-relabelling invariance. • We introduce Racer, a role-aligned competence-estimation framework with nonparametric and neural kernel-pooling instantiations for estimating expert correctness from context. • We establish the probability target theoretically: Racer estimates the posterior-predictive competence needed by the Bayes L2D comparison, is trained with a strictly proper expertcorrectness loss, and yields coherently invariant deferral decisions. • We evaluate on controlled synthetic-expert benchmarks and real-expert chest-radiography benchmarks. Racer gives the strongest aggregate routing on the nominal OOD split in the CIFAR-100 experiments, a strong context-scaling trend in the synthetic-expert PathMNIST histopathology study, and competitive or best budget-swept performance on VinDr-CXR and CheXpert.
2
Background
This section reviews the theory needed to motivate Racer. We first recall the Bayes comparison and augmented-softmax surrogate for standard L2D. We then separate two notions of decision consistency and calibrated expert-correctness estimation. Next, we move to populationadaptive L2D, where the expert-correctness term must be inferred from context, and discuss the representation problem caused by absolute class coordinates. We then review coherent class relabelling and IFD, which gives an identity-free but classwise solution. Finally, we identify the missing probability object: a query-conditioned, role-relative posterior-predictive competence function for unseen experts. Appendix E discusses the broader literature, including clinical collaboration and hierarchical deferral.
2.1
Learning to defer
Let 𝒳 be an input space and 𝒴 B {1, . . . , 𝐾} a finite label space. A random example (𝑋 , 𝑌) is drawn from a distribution 𝑃 over 𝒳 × 𝒴 , and an expert produces a prediction 𝑀 ∈ 𝒴 . A deferral system consists of a classifier ℎ : 𝒳 → 𝒴 and a rejector 𝑟 : 𝒳 → {0, 1}, where 𝑟(𝑥) = 0 means the classifier predicts and 𝑟(𝑥) = 1 means the system defers. Under 0–1 system loss, the risk is ℛ(ℎ, 𝑟) B E [(1 − 𝑟(𝑋))1{ℎ(𝑋) ≠ 𝑌} + 𝑟(𝑋)1{𝑀 ≠ 𝑌}] . Writing 𝜂 𝑦 (𝑥) B P(𝑌 = 𝑦 | 𝑋 = 𝑥), the Bayes classifier and rejector are [14] ℎ★(𝑥) = arg max 𝑦∈𝒴 𝜂 𝑦 (𝑥),
𝑟 ★(𝑥) = 1{P(𝑀 = 𝑌 | 𝑋 = 𝑥) ≥ max 𝑦∈𝒴 𝜂 𝑦 (𝑥)}.
(1)
Thus L2D is not selective classification with a fixed rejection cost. It is a comparison between two decision makers: the classifier and the expert. A standard optimization strategy introduces the augmented action space 𝒴 ⊥ B 𝒴 ∪ {⊥}, where ⊥ denotes deferral. Let 𝑔 𝑦 (𝑥) be class logits and 𝑔⊥ (𝑥) be a deferral logit, and define Π𝑎 (𝑥) B
exp(𝑔𝑎 (𝑥)) exp(𝑔⊥ (𝑥)) +
Í𝐾
𝑘=1 exp(𝑔 𝑘 (𝑥))
3
,
𝑎 ∈ 𝒴 ⊥.
The cross-entropy L2D surrogate of Mozannar and Sontag [14] is ℒCE (𝑥, 𝑦, 𝑚) B − log Π 𝑦 (𝑥) − 1{𝑚 = 𝑦} log Π⊥ (𝑥).
(2)
The first term rewards the correct class, while the second rewards deferral when the training expert is correct. Through the shared softmax normalization, this objective jointly trains the classifier and rejector and can encourage a capacity-limited classifier to complement the expert observed during training [14, Sec. 5.1 and Fig. 1]. This training-time specialization does not imply that the classifier adapts to an unseen expert at deployment. The consistency of this surrogate can be seen from its pointwise conditional risk. The full population risk is the expectation of (2) over (𝑋 , 𝑌, 𝑀), but by the tower property it decomposes as an expectation over 𝑋 of conditional losses. For fixed 𝑥, the relevant conditional population loss is 𝒮(Π | 𝑥) B −
Õ 𝑦
𝜂 𝑦 (𝑥) log Π 𝑦 (𝑥) − 𝑞(𝑥) log Π⊥ (𝑥),
𝑞(𝑥) B P(𝑀 = 𝑌 | 𝑋 = 𝑥).
Here Π(· | 𝑥) ranges over distributions on 𝒴 ⊥ . Let Π★(· | 𝑥) ∈ arg min𝜌∈Δ(𝒴 ⊥ ) 𝒮(𝜌 | 𝑥) denote a conditional population minimizer. A weighted cross-entropy calculation gives Π★𝑦 (𝑥) = Consequently,
𝜂 𝑦 (𝑥) 1 + 𝑞(𝑥)
Π★⊥ (𝑥) ≥ max Π★𝑦 (𝑥) 𝑦
Π★⊥ (𝑥) =
,
⇐⇒
𝑞(𝑥) . 1 + 𝑞(𝑥)
(3)
𝑞(𝑥) ≥ max 𝜂 𝑦 (𝑥), 𝑦
which is precisely the Bayes L2D comparison in (1). The augmented softmax surrogate is therefore decision-consistent, or Bayes consistent in the sense of recovering the optimal decision boundary between prediction and deferral [1].
2.2
From decision consistency to expert-correctness estimation
Decision consistency does not by itself make the augmented deferral coordinate an expertcorrectness probability. Eq. 3 shows that, at the conditional population optimum, Π★⊥ (𝑥) =
𝑞(𝑥) , 1 + 𝑞(𝑥)
𝑞(𝑥) = P(𝑀 = 𝑌 | 𝑋 = 𝑥).
Thus the augmented deferral coordinate is a normalized decision variable, but not the expertcorrectness probability itself. Recovering 𝑞(𝑥) from the ideal optimum requires the odds transform Π★⊥ (𝑥) 𝑞(𝑥) = . 1 − Π★⊥ (𝑥) This distinction is harmless if the goal is only the Bayes defer/predict decision, but it matters whenever the system needs probabilities for reasons such as calibrated triage, threshold selection, budgeted routing, or comparison between multiple experts. Recent calibrated L2D work studies this issue in fixed-expert and fixed multi-expert settings. Verma and Nalisnick [23] show that the standard augmented-softmax coordinate should not be read directly as a calibrated probability of expert correctness, and propose a one-vs-all formulation that parameterizes expert correctness more explicitly. Verma et al. [24] extend related ideas to fixed multi-expert L2D, where expert-specific probabilities must be comparable across experts. Cao et al. [2] further clarify that the issue is not the use of softmax as such, but 4
the probability structure as class probabilities live on a simplex, whereas expert correctness is a separate scalar in [0, 1]. For the present paper, the lesson is that the expert side of the L2D comparison should be treated as a probability-scale object, and not just as an augmented action coordinate. In the population-adaptive setting, this probability must additionally be inferred from the context set of an unseen expert.
2.3
Population-adaptive L2D and L2D-Pop
The preceding discussion assumes a fixed expert. Population-adaptive L2D extends this setting by assuming that experts are drawn from a population and that the expert available at test time may be unseen. The rejector must therefore depend on both the query input and the competence of the available expert, inferred from context. At deployment the system observes a small context set 𝐶 𝑒 B {(𝑥 𝑖𝐶 , 𝑦 𝑖𝐶 , 𝑚 𝑖𝐶 )}𝐵𝑖=1 , where 𝐵 is the number of context examples and 𝑚 𝑖𝐶 is expert 𝑒’s prediction on context input 𝑥 𝑖𝐶 . The population risk in this setup is ℛpop (ℎ, 𝑟) B E [(1 − 𝑟(𝑋 , 𝐸))1{ℎ(𝑋) ≠ 𝑌} + 𝑟(𝑋 , 𝐸)1{𝑀𝐸 ≠ 𝑌}] , where 𝑟 : 𝒳 × ℰ → {0, 1} is an expert-conditioned rejector. The corresponding identityconditioned Bayes rule is 𝑟 ★(𝑥, 𝑒) = 1{P(𝑀 𝑒 = 𝑌 | 𝑋 = 𝑥, 𝐸 = 𝑒) ≥ max 𝑦 𝜂 𝑦 (𝑥)}.
(4)
Thus the model posterior is still query-dependent through 𝑥, but the expert-correctness term is now expert-dependent and must be inferred from the information available for the test-time expert. In the main framework developed in Section 3, the image-only classifier is fixed at deployment, while context adapts competence estimates and routing. We also consider an extension that conditions the class posterior on context. Neither approach requires test-time parameter updates. L2D-Pop addresses this problem by learning a context encoder [22]. Let 𝜔 denote the parameters of the image encoder and classifier head. Write 𝜑 𝜔 : 𝒳 → R𝑑 for the image encoder, 𝜓 for the context encoder, and 𝑑𝜗 for the deferral head that maps query features and an expert representation to a scalar deferral logit. In the query-independent (QI) and query-conditioned (QC) variants, respectively, L2D-Pop forms 𝜓 𝑒 B 𝜓(𝐶 𝑒 ) or 𝜓 𝑒 (𝑥) B 𝜓(𝑥, 𝐶 𝑒 ), and implements the expert-conditioned deferral logit as 𝑔⊥ (𝑥, 𝑒) B 𝑑𝜗 (𝜑 𝜔 (𝑥), 𝜓 𝑒 ) or 𝑔⊥ (𝑥, 𝑒) B 𝑑𝜗 (𝜑 𝜔 (𝑥), 𝜓 𝑒 (𝑥)). The query-conditioned version is especially appealing because it can attend to context examples that are similar to the current query. L2D-Pop is therefore a strong utility baseline because it extends one-stage L2D training to unseen experts while preserving the classifier–rejector co-adaptation provided by the augmented softmax objective. Our comparisons use this softmax variant. The original L2D-Pop paper also develops a one-vs-all surrogate and reports experiments with it [22, Apps. A.3 and D.1]; that variant is not evaluated here. With sufficient data, model capacity, and successful optimization, a context-conditioned rejector can recover the Bayes comparison in Eq. (4). Its learned routing score, however, need not provide an explicit estimate of the expert’s probability of being correct. The softmax L2D-Pop variant considered here learns a latent routing margin, ΔPop (𝑥, 𝐶 𝑒 ) B 𝑔⊥ (𝑥, 𝑒) − max 𝑦 𝑔 𝑦 (𝑥), whose purpose is to compare expert and model reliability. This margin is an object for a defer/predict decision. The softmax variant does not directly parameterize expert correctness as a bounded probability, although the ideal augmented-softmax output recovers it through 5
the odds transform above. It also does not explicitly estimate the rolewise competence law P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ). Consequently, a good empirical routing margin can be supported by several signals that are correlated on the training population: expert competence, model uncertainty, class frequency, context-set composition, acquisition or site artifacts, or shortcuts tied to absolute class coordinates. These signals may improve ID routing performance, but they need not remain aligned under expert shift. This issue is amplified by the way standard population encoders represent class-indexed context information. A typical context token has the form [𝜑(𝑥 𝑖𝐶 ), 𝑒lab (𝑦 𝑖𝐶 ), 𝑒lab (𝑚 𝑖𝐶 )], where 𝑒lab (·) is a one-hot or learned label embedding. Such tokens expose absolute label identities to the expert encoder and rejector. They allow the model to learn coordinate-specific rules, such as deferring when an expert appears strong on a particular class index, rather than forcing the transferable rule: defer when the expert is strong on the relevant role. This is the gap addressed by identity-free approaches [20].
2.4
Coherent class relabelling
The identity issue can be formalized through a symmetry. Class names are arbitrary: renaming them coherently should rename the classifier output, but should not change the binary defer/predict decision. Let 𝔖𝐾 be the symmetric group on 𝒴 . For a permutation 𝜋 ∈ 𝔖𝐾 , define the relabelled posterior 𝜂𝜋𝑦 (𝑥) B 𝜂𝜋−1 (𝑦) (𝑥) and the relabelled context 𝜋𝐶 𝑒 B {(𝑥 𝑖𝐶 , 𝜋(𝑦 𝑖𝐶 ), 𝜋(𝑚 𝑖𝐶 ))}𝐵𝑖=1 . Definition 1 (Coherent class relabelling). A coherent class relabelling applies the same permutation 𝜋 ∈ 𝔖𝐾 to every class-indexed object: labels, expert predictions, model posteriors, and candidate roles. A classifier is coherently equivariant if ℎ 𝜋 (𝑥) = 𝜋(ℎ(𝑥)). A rejector is coherently invariant if 𝑟 𝜋 (𝑥, 𝜋𝐶 𝑒 ) = 𝑟(𝑥, 𝐶 𝑒 ). The Bayes classifier is equivariant under coherent class relabelling, and the Bayes rejector is invariant [20, Propn. 3.1]. Architectures that expose absolute class coordinates need not satisfy these symmetry conditions. This is a distribution-free architectural vulnerability: even if such models perform well on ID experts, their hypothesis class permits identity-conditioned shortcuts that are not stable under expert shifts.
2.5
Identity-Free Deferral
IFD addresses this vulnerability by replacing unstructured expert embeddings with explicit Bayesian competence profiles [20]. For expert 𝑒, class 𝑦, and context counts 𝑛 𝑒 ,𝑦 B
Õ 𝑖
1{𝑦 𝑖𝐶 = 𝑦},
𝑡 𝑒 ,𝑦 B
Õ 𝑖
1{𝑦 𝑖𝐶 = 𝑦, 𝑚 𝑖𝐶 = 𝑦 𝑖𝐶 },
IFD estimates the classwise correctness probability 𝜃𝑒 ,𝑦 ≈ P(𝑀 𝑒 = 𝑌 | 𝑌 = 𝑦, 𝐶 𝑒 ). The rejector does not receive this profile as a labelled vector in fixed coordinates. Instead, it receives role-indexed quantities, such as the posterior mean and uncertainty at the model’s top-ranked class, the expert’s best estimated class, and symmetric summaries across classes. This is the key strength of IFD: it is data-efficient, interpretable, and coherently invariant by construction. In this sense, IFD is closer to probability-scale competence estimation than a latent deferralmargin method: its profile entries are explicit posterior estimates of classwise expert correctness. 6
Its limitation is therefore not that competence is hidden inside an unconstrained routing score, but that the estimated probability is classwise rather than query-conditioned. IFD approximates the expert’s rolewise competence by ΓIFD (𝑥, 𝑦, 𝐶 𝑒 ) B 𝜃𝑒 ,𝑦 ,
(5)
which removes the dependence on query 𝑥. This is appropriate when expert ability is mostly classwise. It is biased when performance varies within class, for example when an expert is reliable on prototypical cases but not atypical ones, or when image quality, acquisition site, subtype, or local morphology changes the expert’s reliability without changing the label.
2.6
The missing probability object
Table 1: Main conceptual comparison. L2D-Pop is expressive but learns a latent routing margin; IFD is role-aligned but classwise. Racer estimates the query-conditioned role competence object whose marginalization gives expert correctness. Method
Context interface
Learned or estimated object
L2D-Pop QI/QC (softmax) [22]
Latent embedding 𝜓(𝐶 𝑒 ) or 𝜓(𝑥, 𝐶 𝑒 )
Routing margin 𝑔⊥ (𝑥, 𝜓) − max 𝑦 𝑔 𝑦 (𝑥)
IFD [20]
Bayesian classwise profile read through role-indexed summaries Role-relative query–context competence estimator
Racer (Ours)
Strength and limitation
Strong one-stage routing and, for QC, instance dependence; not explicitly calibrated as expert correctness and not generally identity-free ΓIFD (𝑥, 𝑦, 𝐶 𝑒 ) = 𝜃𝑒,𝑦 Explicit probability-scale competence and coherent invariance; query-independent within class ΓRacer (𝑥, 𝑦, 𝐶 𝑒 ) ≈ P(𝑀𝐸 = Instance-dependent, 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = probability-scale, and 𝐶𝑒 ) identity-free; requires informative context geometry and sufficient support
The preceding subsections identify two separate requirements for population-adaptive deferral. From calibrated L2D, the expert side of the Bayes comparison should be a probabilityscale correctness estimate. From coherent relabelling and IFD, this estimate should be expressed in role-relative rather than absolute class coordinates. What remains missing is the instancedependent version of this object for an unseen expert represented only by context. For an unseen expert, the Bayes population rejector with the information available at deployment requires the posterior-predictive correctness probability 𝑞★(𝑥, 𝐶 𝑒 ) B P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ). Conditioning on the unknown true class gives via the law of total probability 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ 𝑦∈𝒴
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 )P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ).
Under the context-posterior invariance condition formalized in Assumption 1, the expert context does not alter our label posterior once 𝑥 is known, i.e. 𝑌 ⊥⊥ 𝐶𝐸 | 𝑋. This should be understood as a context-acquisition condition rather than as a property of every possible historical expert 7
record. Appendix B describes a more general context-conditioned posterior variant. Under this condition, the decomposition becomes 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ 𝑦∈𝒴
𝜂 𝑦 (𝑥) P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) .
|
{z
Γ★ (𝑥,𝑦,𝐶 𝑒 )
}
Thus the central population-adaptive competence object is Γ★(𝑥, 𝑦, 𝐶 𝑒 ) B P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ). This posterior-predictive quantity averages over uncertainty about the expert given the labelled context examples in 𝐶 𝑒 . When the expert identity 𝐸 = 𝑒 is known, the corresponding oracle target is P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐸 = 𝑒). This view clarifies the relationship between existing methods. L2D-Pop can learn the Bayes routing margin implicitly. Its softmax variant does not directly parameterize 𝑞★(𝑥, 𝐶 𝑒 ) as a bounded probability or explicitly estimate the rolewise competence law Γ★(𝑥, 𝑦, 𝐶 𝑒 ). IFD estimates competence explicitly and identity-free, but uses the query-independent approximation ΓIFD (𝑥, 𝑦, 𝐶 𝑒 ) = 𝜃𝑒 ,𝑦 . Racer is designed to estimate the missing object directly: a queryconditioned, role-relative, probability-scale competence function whose marginalization against the model posterior gives expert correctness. Table 1 summarizes our positioning: L2D-Pop provides expressive context-conditioned routing but leaves expert competence implicit; IFD estimates competence explicitly and identity-free but only classwise; Racer targets the missing query-conditioned, role-relative competence probability.
3
Methodology
Racer separates classification, expert-competence estimation, and routing. The classifier produces an estimate 𝑝 𝜔 (𝑦 | 𝑥) of the label posterior 𝜂 𝑦 (𝑥), the competence module estimates how likely the expert is to be correct on a candidate role, and the final deferral rule compares the induced expert-correctness probability with the classifier’s confidence in its predicted class, max 𝑦 𝑝 𝜔 (𝑦 | 𝑥). We study two competence estimators. Racer-KNN uses a nonparametric 𝑘-nearest-neighbour estimate of expert correctness within each candidate class. Racer-kernel, our default implementation, applies a learned competence network to kernel-weighted context summaries and classifier statistics. Both use context to estimate expert competence while retaining an image-only class posterior. Section 3.4 gives their full specifications. This section defines the competence target, describes the role-relative estimators, and proves the two properties needed for transfer: probability-scale expert-correctness estimation and coherent-relabelling invariance.
3.1
Problem setup and competence target
We consider a population of experts. An expert identity 𝐸 ∈ ℰ is drawn from a population distribution. At deployment, expert 𝑒 may be unseen. The system observes a context set of 𝐵 labelled examples of that expert’s predictions, 𝐶𝐸 = 𝐶 𝑒 = {(𝑥 𝑖𝐶 , 𝑦 𝑖𝐶 , 𝑚 𝑖𝐶 )}𝐵𝑖=1 , and then receives a query 𝑥. We write 𝑎 𝑖𝐶 B 1{𝑚 𝑖𝐶 = 𝑦 𝑖𝐶 } for context correctness. The system must either output a class or defer to the expert represented by 𝐶 𝑒 . 8
Recall that 𝜂 𝑦 (𝑥) denotes the true label posterior. Let 𝑝 𝜔 (𝑦 | 𝑥) be the classifier’s estimated posterior, and 𝑝 𝜔,max (𝑥) B max 𝑦 𝑝 𝜔 (𝑦 | 𝑥). If the expert identity were known, the oracle expert-correctness term would be 𝑞 or 𝑒 (𝑥) B P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐸 = 𝑒). Racer does not assume access to this identity-level oracle. Its deployable target is the posteriorpredictive correctness probability conditioned on the context available at test time, 𝑞★(𝑥, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ). The corresponding rolewise competence target is Γ★(𝑥, 𝑦, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ). Assumption 1 (Context-posterior invariance). For the deployment contexts considered in this paper,
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) = P(𝑌 = 𝑦 | 𝑋 = 𝑥) = 𝜂 𝑦 (𝑥),
∀𝑥, 𝑦 and feasible contexts 𝐶 𝑒 .
Assumption 1 should be read as a context-acquisition condition rather than a property of arbitrary historical logs. It is appropriate for randomized or stratified warm-up contexts sampled independently of the future query stream, where 𝐶 𝑒 reveals expert behavior but not additional case-mix information. Historical contexts tied to site, scanner, ward, service line, referral pathway, patient cohort, or other prevalence-shifting workflow variables may violate it. In such deployments, the Bayes expert-correctness term becomes 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ 𝑦∈𝒴
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 )Γ★(𝑥, 𝑦, 𝐶 𝑒 ),
and we consider an optional extension, Racer-C, where C denotes context-conditioned classification. It retains Racer-kernel’s competence-estimation architecture and adds a role-equivariant posterior adapter with parameters 𝜌, replacing 𝑝 𝜔 (𝑦 | 𝑥) with 𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ). The adapted posterior is used both to combine rolewise competence estimates and to determine the classifier’s prediction and confidence. Appendix B gives the construction, and Appendix D.3 evaluates it under controlled shifts in case mix. Under Assumption 1, we retain Racer-kernel as the default, avoiding an additional posterior-estimation task that can increase variance when context is limited. Under Assumption 1, the law of total probability gives 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ).
𝑦∈𝒴
Racer estimates Γ★ with a role-aligned competence functional Γ(𝑥, 𝑦, 𝐶 𝑒 ) ∈ [0, 1]. The induced expert-correctness estimate and plug-in rejector are
b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) B
Õ 𝑦∈𝒴
𝑝 𝜔 (𝑦 | 𝑥)Γ(𝑥, 𝑦, 𝐶 𝑒 ),
b 𝑟𝜏 (𝑥, 𝐶 𝑒 ) B 1{b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) − 𝑝 𝜔,max (𝑥) ≥ 𝜏}.
(6)
Here 𝜏 controls the deferral budget or cost. The rest of the methodology answers two questions: (i) how should Γ be parameterized so that it is expressive but identity-free?, and (ii) how should it be trained so that it remains a probability-scale competence estimate rather than only a routing feature?
9
3.2
Role-relative information
Let 𝜑 𝜔 : 𝒳 → R𝑑 be an image encoder, with 𝑧 B 𝜑 𝜔 (𝑥) and 𝑧 𝑖𝐶 B 𝜑 𝜔 (𝑥 𝑖𝐶 ). For a candidate role Í 𝑦, define the posterior rank 𝑅 𝜔 (𝑦; 𝑥) B 1 + 𝑘∈𝒴 1{𝑝 𝜔 (𝑘 | 𝑥) > 𝑝 𝜔 (𝑦 | 𝑥)}. Racer allows class information to enter the competence estimator only through relations to the candidate role. For query 𝑥, role 𝑦, and context item 𝑖, admissible quantities include 𝑧, 𝑧 𝑖𝐶 , 𝑠(𝑧, 𝑧 𝑖𝐶 ) ,
|
{z
}
query–context geometry
𝑎 𝑖𝐶
1{𝑦 𝑖𝐶 = 𝑦}, 1{𝑚 𝑖𝐶 = 𝑦},
,
|{z}
|
{z
}
𝑝 𝜔 (𝑦 | 𝑥), 𝑅 𝜔 (𝑦; 𝑥),
context correctness candidate-role relations 𝑝 𝜔 (𝑦 𝑖𝐶 | 𝑥 𝑖𝐶 ), 𝑅 𝜔 (𝑦 𝑖𝐶 ; 𝑥 𝑖𝐶 ) , H(𝑝 𝜔 (· | 𝑥))
|
|
{z
query role state
}
{z
context role state
}
|
{z
.
}
query uncertainty
The precise feature list is less important than the invariance rule. Labels may appear through equality relations, ranks, posterior values at roles, or correctness indicators, but not through learned label embeddings, untied class-specific channels, or fixed class-coordinate vectors. Image features are allowed as non-class-indexed geometry. Any class-indexed object derived from the classifier must transform equivariantly under coherent relabelling (Definition 1). Formally, let 𝑇 denote a summary map used by a competence estimator. We call 𝑇 roleadmissible if, for every class permutation 𝜋 ∈ 𝔖𝐾 and every posterior satisfying 𝑝 𝜋𝜔 (𝜋(𝑦) | 𝑥) = 𝑝 𝜔 (𝑦 | 𝑥), 𝑇(𝑥, 𝜋(𝑦), 𝜋𝐶 𝑒 , 𝑝 𝜋𝜔 ) = 𝑇(𝑥, 𝑦, 𝐶 𝑒 , 𝑝 𝜔 ). Racer instantiations use only role-admissible summaries and apply the same functions across roles. This is the architectural difference between a role-aligned competence estimator and a generic population encoder.
3.3
Same-role kernel pooling
Both Racer-KNN and Racer-kernel use a query-conditioned same-role pooling primitive. Let 𝑧˜ B 𝜑 𝜔 (𝑥)/∥𝜑 𝜔 (𝑥)∥2 and 𝑧˜ 𝑖𝐶 B 𝜑 𝜔 (𝑥 𝑖𝐶 )/∥𝜑 𝜔 (𝑥 𝑖𝐶 )∥2 , define the cosine similarity 𝑠 𝑖 (𝑥) B 𝑧˜ ⊤ 𝑧˜ 𝑖𝐶 , and set 𝐾 𝑖 (𝑥) B exp(𝑠 𝑖 (𝑥)/𝜏ker ). For candidate role 𝑦, the same-role similarity mass is 𝑆 𝑒 ,𝑦 (𝑥) B
Õ𝐵 𝑖=1
1{𝑦 𝑖𝐶 = 𝑦}𝐾 𝑖 (𝑥),
𝑁𝑒 ,𝑦 B
Õ𝐵 𝑖=1
1{𝑦 𝑖𝐶 = 𝑦}.
(7)
These quantities compare the query only to context examples whose true label matches the candidate role 𝑦. A normalized local average is undefined without same-role support. We define a local competence statistic with an explicit fallback and optional prior smoothing. Let 𝜇𝑒 ,0 ∈ [0, 1] be a global-context prior, for example 𝜇𝑒 ,0 B
𝛼pop 𝜇pop +
Í𝐵
𝐶 𝑖=1 𝑎 𝑖
𝛼pop + 𝐵
,
where 𝜇pop is a population-level expert-accuracy prior and 𝛼 pop ≥ 0 is its strength, with 𝛼pop + 𝐵 > 0. For local prior strength 𝛼0 ≥ 0, define
Í 𝛼 𝜇 + 𝐵𝑖=1 1{𝑦 𝑖𝐶 = 𝑦}𝐾 𝑖 (𝑥)𝑎 𝑖𝐶 0 𝑒 ,0 , 𝛼0 + 𝑆 𝑒 ,𝑦 (𝑥) > 0, 𝑎¯ 𝑒 ,𝑦 (𝑥) B 𝛼0 + 𝑆 𝑒 ,𝑦 (𝑥) 𝜇𝑒 ,0 , 𝛼 0 + 𝑆 𝑒 ,𝑦 (𝑥) = 0.
(8)
Thus 𝑎¯ 𝑒 ,𝑦 (𝑥) is a similarity-weighted empirical correctness estimate for expert 𝑒 on context examples that are both similar to 𝑥 and have candidate role 𝑦, shrunk toward a global context 10
prior when same-role support is weak and 𝛼0 > 0. If 𝑆 𝑒 ,𝑦 (𝑥) = 0, the estimate equals 𝜇𝑒 ,0 . It is query-dependent through 𝑠 𝑖 (𝑥) and identity-free because the candidate class appears only through the equality mask 1{𝑦 𝑖𝐶 = 𝑦}. In the supplied implementation, 𝛼0 = 0: supported roles use a normalized local mean, while unsupported roles fall back to global context accuracy. The global prior is the empirical context accuracy for 𝐵 > 0, with 1/2 used for empty context. Kernel weights are computed with a stable softmax; the auxiliary similarity-mass feature caps exponential arguments at 30 for numerical stability.
3.4
Two competence estimators
Nonparametric estimator. Racer-KNN uses the local same-role statistic with uniform weights on the nearest same-role neighbors (and zero weight on the remaining items): ΓKNN (𝑥, 𝑦, 𝐶 𝑒 ) B 𝑎¯ 𝑒 ,𝑦 (𝑥). This estimator is transparent in a case-based sense: its competence estimate is the empirical correctness of expert 𝑒 on nearby same-role context examples. The implemented KNN variant uses a uniform average over the min(𝑘, 𝑁𝑒 ,𝑦 ) nearest same-role items, rather than exponential weights, and uses global context accuracy when no same-role item is available. Learned kernel estimator. Racer-kernel adds a role-shared competence network on top of the normalized kernel statistic. Its learned mapping supplies smoothing beyond the unsupportedrole fallback. Let margin𝜔 (𝑦; 𝑥) B 𝑝 𝜔 (𝑦 | 𝑥) − max 𝑝 𝜔 (𝑘 | 𝑥) 𝑘≠𝑦
be the posterior margin of role 𝑦. The role-relative summary is
𝑢𝑒 ,𝑦 (𝑥) B 𝑎¯ 𝑒 ,𝑦 (𝑥), log(1 + 𝑁𝑒 ,𝑦 ), log(1 + 𝑆 𝑒 ,𝑦 (𝑥)),
𝑝 𝜔 (𝑦 | 𝑥), 𝑅 𝜔 (𝑦; 𝑥), margin𝜔 (𝑦; 𝑥), H(𝑝 𝜔 (· | 𝑥)), 𝜇𝑒 ,0 . A small MLP 𝑔𝜃 is shared across all roles and defines Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ) B sigm 𝑔𝜃 (𝑢𝑒 ,𝑦 (𝑥)) .
Because the same network is applied to every role and all inputs are role-relative, Racer-kernel can learn nonlinear smoothing and calibration without learning class-identity-specific rules. In our probability-scale interpretation, 𝑢𝑒 ,𝑦 (𝑥) is a summary of the full information (𝑥, 𝑦, 𝐶 𝑒 ); the proper loss below targets conditional correctness given that summary at its population optimum.
3.5
Proper competence training and plug-in calibration
For a target tuple (𝑥, 𝑦, 𝑚) from expert 𝑒 and context 𝐶 𝑒 , define 𝑎 B 1{𝑚 = 𝑦}. The neural competence head is trained with the binary proper scoring rule ℒcomp (𝜃) B E −𝑎 log Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ) − (1 − 𝑎) log(1 − Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 )) ,
(9)
where the expectation is over sampled experts, their contexts, and target examples from the same expert. When the classifier posterior is intended to remain a calibrated estimate of 𝜂, the representation 𝜑 𝜔 and posterior 𝑝 𝜔 used inside 𝑢𝑒 ,𝑦 (𝑥) are best treated as fixed or stop-gradient inputs to the competence loss. If end-to-end feature sharing is used instead, classifier calibration should be rechecked after joint training. 11
Proposition 1 (Proper target and summary projection). Let 𝐴 B 1{𝑀𝐸 = 𝑌} and let 𝑈 B 𝑇(𝑋 , 𝑌, 𝐶𝐸 , 𝑝 𝜔 ) be any summary made available to a scalar predictor 𝑔(𝑈) ∈ [0, 1], with BCE extended to the boundary by continuity. The conditional minimizer of the binary cross-entropy over all measurable functions of 𝑈 is 𝑔★(𝑈) = P(𝐴 = 1 | 𝑈). In particular, if 𝑈 = (𝑋 , 𝑌, 𝐶𝐸 ), then 𝑔★(𝑥, 𝑦, 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ). More generally, the unrestricted population BCE optimum given the role-relative summary 𝑢𝑒 ,𝑦 (𝑥) equals its conditional correctness probability; it recovers the full competence functional when that summary is sufficient. This does not certify finite-network optimization or calibration under distribution shift. Proof. Condition on 𝑈 = 𝑢 and write 𝛼 B P(𝐴 = 1 | 𝑈 = 𝑢). The conditional risk for a scalar prediction 𝑔 ∈ (0, 1) is −𝛼 log 𝑔 − (1 − 𝛼) log(1 − 𝑔). Its derivative is −𝛼/𝑔 + (1 − 𝛼)/(1 − 𝑔), whose unique zero is 𝑔 = 𝛼. Strict convexity gives the unique minimizer for 0 < 𝛼 < 1; for 𝛼 ∈ {0, 1} the corresponding boundary value minimizes the extended loss. Taking 𝑈 = (𝑋 , 𝑌, 𝐶𝐸 ) yields P(𝐴 = 1 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ). Full proof in Appendix A.2. Under context-posterior invariance, the plug-in estimate recovers expert correctness when the classifier posterior and rolewise competence estimates recover their respective conditional probability targets. Marginal calibration of the two predictors alone is insufficient. Proposition 2 (Plug-in calibration). Under Assumption 1, if 𝑝 𝜔 (𝑦 | 𝑥) = 𝜂 𝑦 (𝑥) and Γ(𝑥, 𝑦, 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ) for all 𝑦, then
b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) =
Õ 𝑦
𝑝 𝜔 (𝑦 | 𝑥)Γ(𝑥, 𝑦, 𝐶 𝑒 ) = 𝑞★(𝑥, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ).
Proof. Apply the decomposition in (3.1) and substitute the assumed conditional probability targets for 𝑝 𝜔 and Γ. The next bound states the corresponding plug-in decision guarantee. It is intentionally a decision-regret statement, not a claim that finite neural training is automatically Bayes-consistent. The bound concerns ordinary 0–1 system loss without an additional deferral cost. It does not establish optimality for macro AURSBAC, nor does zero-threshold Bayes alignment alone guarantee optimal ranking across deferral budgets. Proposition 3 (Plug-in routing regret). Let ℎ 𝜔 (𝑥) B arg max 𝑦 𝑝 𝜔 (𝑦 | 𝑥) and let b 𝑟0 (𝑥, 𝐶 𝑒 ) = ★ ★ ★ 1{b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) ≥ 𝑝 𝜔,max (𝑥)}. Let ℎ (𝑥) = arg max 𝑦 𝜂 𝑦 (𝑥) and 𝑟 (𝑥, 𝐶 𝑒 ) = 1{𝑞 (𝑥, 𝐶 𝑒 ) ≥ 𝜂max (𝑥)}, where 𝜂max (𝑥) B max 𝑦 𝜂 𝑦 (𝑥). Under the 0–1 system loss with no additional deferral cost, ℛ(ℎ 𝜔 , b 𝑟0 ) − ℛ(ℎ★ , 𝑟 ★)
" ≤ E 3 𝑝 𝜔 (· | 𝑋) − 𝜂(· | 𝑋) 1 +
# Õ
𝜂 𝑦 (𝑋) Γ(𝑋 , 𝑦, 𝐶𝐸 ) − Γ★(𝑋 , 𝑦, 𝐶𝐸 ) .
𝑦∈𝒴
Proof. Condition on (𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) and abbreviate 𝑞 B 𝑞★(𝑥, 𝐶 𝑒 ), 𝑞ˆ B b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ), 𝜂max B max 𝑦 𝜂 𝑦 (𝑥), and 𝑝ˆ max B 𝑝 𝜔,max (𝑥). The Bayes conditional correctness is max{𝜂max , 𝑞}, while the plug-in system has conditional correctness (1 − b 𝑟0 )𝜂 ℎ 𝜔 (𝑥) + b 𝑟0 𝑞. Add and subtract the correctness that would be obtained by using the Bayes top class whenever the plug-in system predicts. The resulting excess is at most 𝜂max − 𝜂 ℎ 𝜔 (𝑥) + |𝑞 − 𝜂max | 1{sign(𝑞 − 𝜂max ) ≠ sign( 𝑞ˆ − 𝑝ˆ max )}.
12
The first term is at most ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 . On the sign-disagreement event, |𝑞 − 𝜂max | ≤ | 𝑞ˆ − 𝑞| + |𝑝ˆ max − 𝜂max |. Moreover, | 𝑝ˆ max − 𝜂max | ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 , and | 𝑞ˆ − 𝑞| ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 +
Õ 𝑦
𝜂 𝑦 (𝑥)|Γ(𝑥, 𝑦, 𝐶 𝑒 ) − Γ★(𝑥, 𝑦, 𝐶 𝑒 )|.
Combining these inequalities and taking expectation gives the claim. Full proof in App. A.4. This is the sense in which Racer differs from a pure routing model. The final rejector may be trained with a decision surrogate, but the expert side of the comparison has a separate probability target and an explicit path from estimation error to deferral regret.
3.6
Invariance and instance dependence
Under a coherent relabelling 𝜋 ∈ 𝔖𝐾 , assume the classifier posterior transforms equivariantly: 𝑝 𝜋𝜔 (𝜋(𝑦) | 𝑥) = 𝑝 𝜔 (𝑦 | 𝑥). Then the quantities used by Racer are preserved at corresponding roles: same-role masks, correctness indicators, similarities, ranks, entropies, posterior values, support counts, similarity masses, the global prior 𝜇𝑒 ,0 , and the smoothed statistic 𝑎¯ 𝑒 ,𝑦 (𝑥). Theorem 1 (Coherent-relabelling invariance). Under posterior equivariance, both Racer-KNN and Racer-kernel satisfy Γ𝜋 (𝑥, 𝜋(𝑦), 𝜋𝐶 𝑒 ) = Γ(𝑥, 𝑦, 𝐶 𝑒 ), where Γ denotes the corresponding estimator ΓKNN or Γ𝜃 . Consequently, b 𝑞 𝜋𝜔 (𝑥, 𝜋𝐶 𝑒 ) = b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ), 𝜋 𝑝 𝜔,max (𝑥) = 𝑝 𝜔,max (𝑥), and the plug-in rejector in Eq. (6) is coherently invariant. Proof. For every context item 𝑖, 1{𝑦 𝑖𝐶 = 𝑦} = 1{𝜋(𝑦 𝑖𝐶 ) = 𝜋(𝑦)} and 1{𝑚 𝑖𝐶 = 𝑦 𝑖𝐶 } = 1{𝜋(𝑚 𝑖𝐶 ) = 𝜋(𝑦 𝑖𝐶 )}. Image features and query–context similarities are unchanged, while posterior equivariance preserves posterior values, ranks, margins, and entropies at corresponding roles. Therefore 𝑁𝑒 ,𝑦 , 𝑆 𝑒 ,𝑦 (𝑥), 𝜇𝑒 ,0 , 𝑎¯ 𝑒 ,𝑦 (𝑥), and the summary vector 𝑢𝑒 ,𝑦 (𝑥) are unchanged after replacing (𝑦, 𝐶 𝑒 ) by (𝜋(𝑦), 𝜋𝐶 𝑒 ). The result for Racer-KNN follows immediately. The result for Racer-kernel follows because the same function 𝑔𝜃 is shared across all roles. Summing over relabelled roles gives invariance of b 𝑞 𝜔 , and the maximum posterior is unchanged, so the binary rejector is invariant. Full proof in App. A.5. This theorem concerns coherent relabelling of the task representation: all labels, expert predictions, candidate roles, and class-indexed posteriors are renamed together. It does not claim invariance to an expert-only competence shift relative to fixed class names; in that case the Bayes deferral action itself may change, and the context should cause the competence estimate to change. Identity-freeness does not imply query-independence. If two same-class queries 𝑥 and 𝑥 ′ have different similarities to correct and incorrect same-role context examples, then the kernel masses in (7) and the smoothed statistic in (8) can differ. Hence Γ(𝑥, 𝑦, 𝐶 𝑒 ) ≠ Γ(𝑥 ′ , 𝑦, 𝐶 𝑒 ) even for the same candidate role 𝑦. This is precisely the capacity missing from classwise profiles.
3.7
Bayes-aligned deferral training
The competence loss trains Γ𝜃 as a probability-scale expert correctness estimate. To train an explicit deferral head while preserving one-stage L2D co-adaptation, let 𝑓𝜔,𝑦 (𝑥) be class logits and let 𝑑𝜗 (𝑥, 𝐶 𝑒 ) be an invariant deferral logit, for example a function of b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ), 𝑝 𝜔,max (𝑥), their difference, and H(𝑝 𝜔 (· | 𝑥)). In the implementation, all these inputs are detached: 𝑑𝜗 (𝑥, 𝐶 𝑒 ) B 𝐷𝜗 sg b 𝑞 𝜔 , 𝑝 𝜔,max , b 𝑞 𝜔 − 𝑝 𝜔,max , H(𝑝 𝜔 )
13
.
Define Π 𝑦 (𝑥, 𝐶 𝑒 ) B
Π⊥ (𝑥, 𝐶 𝑒 ) B
exp( 𝑓𝜔,𝑦 (𝑥)) exp(𝑑𝜗 (𝑥, 𝐶 𝑒 )) +
Í𝐾
𝑘=1 exp( 𝑓𝜔,𝑘 (𝑥))
exp(𝑑𝜗 (𝑥, 𝐶 𝑒 )) exp(𝑑𝜗 (𝑥, 𝐶 𝑒 )) +
Í𝐾
𝑘=1 exp( 𝑓𝜔,𝑘 (𝑥))
,
.
For a labelled example (𝑥, 𝑦), the context-only deferral loss is ℒRACER (𝑥, 𝑦, 𝐶 𝑒 ) B − log Π 𝑦 (𝑥, 𝐶 𝑒 ) − sg[Γ(𝑥, 𝑦, 𝐶 𝑒 )] log Π⊥ (𝑥, 𝐶 𝑒 ), where sg[·] denotes stop-gradient. Detaching both the loss weight and the deferral-head inputs blocks routing-loss gradients to the competence-head parameters 𝜃. The shared encoder and classifier still receive routing gradients through the class logits, and competence BCE can update the query representation. Thus this is parameter-level isolation of the competence head, not a guarantee that joint training preserves its predictions or calibration. The kernel training objective adds 𝜆comp ℒcomp to this routing loss. Theorem 2 (Bayes alignment). Let Γ★(𝑥, 𝑦, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) and Í 𝑞★(𝑥, 𝐶 𝑒 ) = 𝑦 𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ). For fixed (𝑥, 𝐶 𝑒 ), the conditional minimizer of the idealized loss − log Π𝑌 − Γ★(𝑥, 𝑌, 𝐶 𝑒 ) log Π⊥ over distributions on 𝒴 ⊥ satisfies Π★𝑦 (𝑥, 𝐶 𝑒 ) =
𝜂 𝑦 (𝑥) 1 + 𝑞★(𝑥, 𝐶 𝑒 )
Π★⊥ (𝑥, 𝐶 𝑒 ) =
,
𝑞★(𝑥, 𝐶 𝑒 ) . 1 + 𝑞★(𝑥, 𝐶 𝑒 )
Therefore Π★⊥ (𝑥, 𝐶 𝑒 ) ≥ max 𝑦 Π★𝑦 (𝑥, 𝐶 𝑒 ) if and only if 𝑞★(𝑥, 𝐶 𝑒 ) ≥ max 𝑦 𝜂 𝑦 (𝑥). Proof. Condition on (𝑥, 𝐶 𝑒 ). Taking expectation over 𝑌 gives the conditional risk −
Õ 𝑦
𝜂 𝑦 (𝑥) log Π 𝑦 −
Õ
𝑦
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ) log Π⊥ = −
Õ 𝑦
𝜂 𝑦 (𝑥) log Π 𝑦 − 𝑞★(𝑥, 𝐶 𝑒 ) log Π⊥ .
This is cross-entropy with nonnegative weights {𝜂 𝑦 (𝑥)} 𝑦∈𝒴 and 𝑞★(𝑥, 𝐶 𝑒 ), whose total mass is 1 + 𝑞★(𝑥, 𝐶 𝑒 ). Normalizing these weights gives the stated minimizer. The decision equivalence follows by multiplying Π★⊥ ≥ max 𝑦 Π★𝑦 by 1 + 𝑞★(𝑥, 𝐶 𝑒 ). Full proof in A.6. Thus Racer does not reject the Mozannar-style augmented surrogate. It uses it at the decision layer, where it is appropriate, while the expert-competence channel is separately constrained to estimate a probability. The theorem is deliberately conditional: full statistical consistency of an implemented system requires the classifier posterior, competence estimator, model class, and optimization to approach their population targets. Proposition 3 makes the plug-in dependence explicit.
3.8
Relation to existing methods
Racer can be viewed as interpolating between the strengths of the two main baselines, L2D-Pop [22] and IFD [20]. If the query-dependent kernel is removed and Γ(𝑥, 𝑦, 𝐶 𝑒 ) is replaced by a classwise posterior mean 𝜃𝑒 ,𝑦 , the modelling assumption reduces to that of IFD. This has low variance and exact invariance, but cannot adapt within class. If instead the role-relative competence channel is replaced by a generic latent context encoder trained only through the augmented softmax margin, one obtains the L2D-Pop interface. This has high expressivity and strong one-stage routing, but rolewise competence is not explicitly estimated and the architecture need not be role-aligned. 14
Racer keeps the useful part of both. Like L2D-Pop (QC), it is query-conditioned. Like IFD, it removes absolute class-identity channels. Unlike both, it exposes an explicit estimate of Γ★(𝑥, 𝑦, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ), so calibration, budget ranking, and multi-expert comparison can be based on a probability-scale estimate of expert correctness rather than on a latent routing score.
3.9
Multi-expert routing 𝐽
If multiple experts are available with contexts {𝐶 𝑒 𝑗 } 𝑗=1 , compute
b 𝑞 𝜔,𝑗 (𝑥) B b 𝑞 𝜔 (𝑥, 𝐶 𝑒 𝑗 ) =
Õ 𝑦
𝑝 𝜔 (𝑦 | 𝑥)Γ(𝑥, 𝑦, 𝐶 𝑒 𝑗 )
for each expert. With optional workload or cost penalty 𝜆 𝑗 (𝑥) ≥ 0, choose 𝑗★(𝑥) B arg max{b 𝑞 𝜔,𝑗 (𝑥) − 𝜆 𝑗 (𝑥)} 𝑗
and defer only if
b 𝑞 𝜔,𝑗★ (𝑥) − 𝜆 𝑗★ (𝑥) − 𝑝 𝜔,max (𝑥) ≥ 𝜏. Calibration is essential in this extension because the quantities b 𝑞 𝜔,𝑗 (𝑥) must be comparable across experts and contexts.
4
Experiments
We evaluate whether a role-relative competence interface improves the two quantities needed for population-adaptive deferral: budget-swept routing utility and probability-scale expertcorrectness estimation. The main experiments are organized around three research questions: 1. Can a method convert additional context for an unseen expert into better query-specific routing? 2. Does the method remain useful when the label space is large and expert competence varies across roles? 3. Do the resulting competence estimates behave like probabilities rather than only ranking scores? A controlled context-posterior stress test for the optional Racer-C extension is reported in Section D.3. Full split details, subtype construction, expert-generation parameters, and run counts are given in Appendix D.
4.1
Evaluation setup
The synthetic benchmark has two main parts. The PathMNIST [6, 25, 26] context-scaling study fixes a locally complementary expert setting and varies the context size, isolating whether a router can exploit additional evidence when competence varies within class. CIFAR-100 [7] stresses class cardinality and synthetic expert variation by varying 𝐾, expert strength, hidden subtype dependence, and expert-profile permutations. The context-posterior stress test is kept in the appendix because it is diagnostic rather than part of the main method comparison: it asks when the more flexible Racer-C extension is worth estimating. Throughout the synthetic experiments, 𝐾 denotes the number of observable classes and 𝐵 denotes the number of context examples supplied for the queried expert. Expert competence can depend on a hidden within-class subtype. The parameter 𝜌 controls this dependence: 𝜌 = 0 gives classwise competence, where all subtypes of a class share the same accuracy, while 15
𝜌 = 1 gives full hidden-subtype specialization. The nominal OOD permutation parameter 𝜆id controls an expert-only permutation of the class/subtype competence profile while image labels and classifier coordinates remain fixed. The generator samples exchangeable class effects and independent random subtype orderings for each class. Therefore a full permutation at 𝜆id = 1 preserves the expert-profile population law, just as 𝜆id = 0 does. We retain “OOD” as the recorded split label for comparability with the saved results; these synthetic comparisons establish performance on unseen simulated experts, not robustness to an established expertpopulation distribution shift. The normalized context support 𝐵/𝐾 is therefore the expected number of context examples per observed role under a balanced context draw. The real-expert benchmarks are VinDr-CXR [16] and CheXpert [5], each of which contains persistent multi-rater annotations for medical images. These datasets are less controlled than the synthetic benchmark, but they test whether the same role-relative interface remains useful when expert variability, label noise, and context support are governed by real annotation processes rather than by a simulator. We follow Strong et al. [20] and use AURSAC, the area under the system-accuracy curve obtained by sweeping the deferral threshold, as the primary routing metric for multiclass tasks. Higher AURSAC means better routing over the full range of deferral budgets. For multi-label chest radiography, we report macro AURSBAC, the macro-average of per-pathology area under the system balanced accuracy curve. For methods that expose a bounded expertcorrectness probability b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ), we also report Brier score [3] and 15-bin ECE [4, 15] against the realized event 1{𝑀𝐸 = 𝑌}. L2D-Pop and classifier-confidence routers are included in routing comparisons, but are omitted from calibration comparisons because they do not directly output a finite expert-correctness probability. RACER budget sweeps rank queries by estimated expert correctness minus model confidence; the auxiliary deferral-head logit is used during training. Appendix G.1 gives the methodspecific losses, weights, inference scores, and component updates, together with context sources, context/query separation, and expert-selection protocols (Tables 14 and 15).
4.2
Synthetic context scaling on PathMNIST
The PathMNIST context-scaling experiment is the most direct test of the RACER mechanism. We fix 𝐾 = 9, full subtype dependence (𝜌 = 1), full expert-profile permutation (𝜆id = 1), and good/mid/bad subtype accuracies (0.98, 0.70, 0.30). The context size varies over 𝐵 ∈ {9, 25, 50, 100, 200, 500, 1000}, shown as 𝐵/𝐾. At each context size, the same context/query episodes are reused across methods and expert groups. Each method is trained separately for each 𝐵; further sampling details are in Appendix D.1. Figure 1 reports AURSAC gain over the classifier-confidence router. At 𝐵/𝐾 = 1, all adaptive methods are below the classifier baseline, consistent with the fact that same-role context is too sparse for reliable local competence estimation. As context grows, the methods separate. Racer-kernel crosses above the classifier baseline around 𝐵/𝐾 = 5.6 and shows increasing overall gain at the subsequent reported context sizes. At 𝐵/𝐾 = 111, it reaches an overall gain of +0.0273 AURSAC, with +0.0260 on nominal unseen-OOD experts and +0.0280 on unseen-ID experts. Racer-KNN also improves at large context, but reaches roughly half the final gain of the learned kernel head. IFD-score and IFD-MLP improve mildly as their classwise profiles become less noisy, while L2D-Pop QI/QC remain below the classifier baseline throughout the sweep.
16
AURSAC gain over classifier
Nominal unseen OOD
Unseen ID
0.02
0.03
0.02
0.02
0.01
0.01
0.00
0.00
−0.01
−0.01
0.00
−0.02
Overall
0.03
−0.02
−0.02 1
2.8
5.6
11
22
56
111
1
2.8
Context size B/K
5.6
11
22
56
111
1
2.8
Context size B/K
IFD-score
IFD-MLP
L2D-Pop QI
5.6
11
22
56
111
Context size B/K
L2D-Pop QC
RACER-KNN
RACER-Kernel
Figure 1: PathMNIST context-scaling utility in the hardest locally complementary synthetic setting (𝜌 = 1, 𝜆id = 1). The y-axis shows AURSAC gain over the classifier-confidence router; positive values indicate better budget-swept routing than classifier uncertainty alone. The x-axis is normalized context support 𝐵/𝐾 with 𝐾 = 9. Shaded bands for RACER methods show one standard deviation over three seeds. We partition the results into (i) unseen OOD experts, (ii) unseen ID experts, and (iii) overall experts. Figure 2 shows the corresponding calibration curves. The Brier score follows the utility pattern: Racer-kernel is best across the context range and improves from 0.2251 at 𝐵/𝐾 = 1 to 0.2029 at 𝐵/𝐾 = 111. IFD-score and IFD-MLP approach Brier scores around 0.222, while Racer-KNN starts poorly at small context before converging to a similar range. ECE is more nuanced. IFD-score and IFD-MLP have the lowest final ECE, while Racer-kernel has slightly higher final ECE but substantially better Brier and routing utility. In these results, smoother classwise probabilities and query-conditioned estimates behave differently: a classwise profile can be well bin-calibrated while remaining insufficiently discriminative for routing when competence varies within class.
ECE
Nominal unseen OOD
Unseen ID 0.4
0.4
0.3
0.3
0.3
0.2
0.2
0.2
0.1
0.1
0.1
0.0
0.0
0.0 1
2.8
5.6
11
22
56
111
1
2.8
Context size B/K
Nominal unseen OOD
Brier score
Overall
0.4
0.45
5.6
11
22
56
111
1
0.35
0.35
0.35
0.30
0.30
0.30
0.25
0.25
0.25
0.20 5.6
11
22
56
111
56
111
56
111
0.20 1
2.8
Context size B/K
5.6
11
22
56
111
Context size B/K IFD-score
22
Overall 0.40
2.8
11
Unseen ID
0.40
1
5.6
Context size B/K
0.40
0.20
2.8
Context size B/K
IFD-MLP
RACER-KNN
1
2.8
5.6
11
22
Context size B/K RACER-Kernel
Figure 2: PathMNIST context-scaling calibration. Mean ECE and Brier score are reported for the plotted probability-output methods. RACER bands show one standard deviation over three seeds. Lower is better. The learned Racer-kernel head has the best Brier score among these plotted methods, while classwise IFD variants have slightly lower final ECE but weaker routing utility. 17
This experiment explains why context size affects the interfaces differently. Classwise profiles reduce variance as 𝐵 grows, but cannot remove within-class bias. Latent population encoders observe the larger context, but in this stress test do not convert it into a better routing margin. Racer-kernel uses the additional evidence at the right level of structure: role-matched and query-near context is smoothed by a role-shared competence head, yielding both lower Brier score and larger routing gains. A representation-geometry ablation in Section D.2 checks that this gain is not an artifact of defining hidden subtypes in the same geometry used for local pooling: Racer-kernel remains strongest when subtypes are generated from classifier/auxiliary feature mixtures, while a random within-class subtype negative control drives same-role neighbor purity to chance and removes the Racer gain. To separate Racer’s role-relative competence interface from the specific kernel-pooling estimator, we additionally evaluate Racer-DeepSets, a permutation-invariant neural competence estimator trained with the same BCE target and the same role-admissible features as RACERkernel. Racer-DeepSetsis a strong non-kernel baseline, but RACER-kernel remains best on PathMNIST context scaling, with monotonic context gains, higher final overall AURSAC gain at B=1000 (+0.0273 vs. +0.0194), and lower final Brier score (0.2029 vs. 0.2197). Full results are in App. H.1.
4.3
Class-scale synthetic routing on CIFAR-100
The CIFAR-100 experiments complement the PathMNIST sweep by stressing label-space size and superclass-structured confusions. Table 2 reports the reference model, expert, and oracle accuracies. In the 𝐾 = 20 setting, the classifier is already much stronger than the average expert, so the possible improvement over classifier-confidence routing is modest. The 𝐾 = 100 setting is more diagnostic: model and expert accuracy are closer, hidden subtype specialization matters more, and the oracle gap leaves more room for better routing. 𝐾
Profile
Model acc.
Expert acc.
Oracle
20 20 100 100
Weak Strong Weak Strong
0.8609 ± 0.0085 0.8629 ± 0.0124 0.6560 ± 0.0043 0.6548 ± 0.0067
0.6465 ± 0.0113 0.6696 ± 0.0164 0.6438 ± 0.0072 0.6709 ± 0.0171
0.9005 ± 0.0049 0.9095 ± 0.0078 0.8255 ± 0.0025 0.8362 ± 0.0070
Table 2: OOD reference quantities for the synthetic CIFAR-100 experiments. The oracle is an upper reference for routing and is not a learned method. Variability is the standard deviation over four settings and three seeds.
18
Nominal OOD AURSAC gain over classifier
K=20, weak K=20, strong
K=100, weak K=100, strong
0.005 0.000 0.005 0.010 0.015 RACER-kernel
L2D-Pop (QI)
L2D-Pop (QC)
RACER-KNN
IFD-MLP
IFD-score
Figure 3: OOD AURSAC gain over the classifier-confidence router. Across both label-space sizes and both expert-strength profiles, Racer-kernel is the only method with a positive aggregate OOD gain over the classifier-confidence baseline. Bars show means over four settings and three seeds, without inferential error bars. “OOD” denotes the recorded synthetic split.
𝐾
Method
Weak OOD
Strong OOD
ΔOOD
Weak overall
Strong overall
Racer-kernel (Ours) Classifier-conf. (Baseline) L2D-Pop (QI) [22] L2D-Pop (QC) [22] Racer-KNN (Ours) IFD-MLP [20] IFD-score [20]
0.8087 ± 0.0086 0.8007 ± 0.0063 0.7986 ± 0.0071 0.7994 ± 0.0083 0.7995 ± 0.0075 0.7936 ± 0.0093 0.7932 ± 0.0078
0.8199 ± 0.0110 0.8130 ± 0.0094 0.8096 ± 0.0091 0.8090 ± 0.0105 0.8082 ± 0.0133 0.8063 ± 0.0099 0.8043 ± 0.0126
+0.0112 +0.0122 +0.0109 +0.0096 +0.0087 +0.0127 +0.0111
0.8072 ± 0.0032 0.7996 ± 0.0025 0.7973 ± 0.0021 0.7983 ± 0.0032 0.7969 ± 0.0032 0.7922 ± 0.0055 0.7907 ± 0.0051
0.8222 ± 0.0105 0.8151 ± 0.0085 0.8124 ± 0.0089 0.8114 ± 0.0094 0.8104 ± 0.0132 0.8078 ± 0.0090 0.8068 ± 0.0120
20
Racer-kernel (Ours) Classifier-conf. (Baseline) L2D-Pop (QI) [22] 100 L2D-Pop (QC) [22] IFD-score [20] IFD-MLP [20] Racer-KNN (Ours)
0.7299 ± 0.0050 0.7253 ± 0.0031 0.7229 ± 0.0043 0.7215 ± 0.0041 0.7186 ± 0.0040 0.7169 ± 0.0036 0.7082 ± 0.0037
0.7461 ± 0.0091 0.7388 ± 0.0084 0.7361 ± 0.0089 0.7339 ± 0.0099 0.7310 ± 0.0095 0.7296 ± 0.0097 0.7227 ± 0.0093
+0.0162 +0.0135 +0.0132 +0.0124 +0.0124 +0.0127 +0.0146
0.7277 ± 0.0041 0.7233 ± 0.0034 0.7207 ± 0.0036 0.7194 ± 0.0045 0.7165 ± 0.0038 0.7147 ± 0.0030 0.7060 ± 0.0032
0.7461 ± 0.0068 0.7391 ± 0.0060 0.7364 ± 0.0071 0.7342 ± 0.0075 0.7312 ± 0.0083 0.7297 ± 0.0072 0.7227 ± 0.0082
Table 3: Synthetic CIFAR-100 routing utility. Entries are mean ± standard deviation AURSAC over 4 stress-test settings and 3 seeds. The “Classifier-conf.” baseline sweeps a classifierconfidence deferral rule and uses no expert context. ΔOOD is strong-minus-weak OOD AURSAC. Table 4 reports descriptive AURSAC differences between Racer-kernel and classifierconfidence routing, matched by (𝜌, 𝜆id , seed) within each 𝐾/profile block. Racer-kernel is higher in 11/12, 11/12, 11/12, and 12/12 matched cells. These cells share seeds and experimental factors; they are not asserted to be independent replicates. We report no inferential tests or confidence intervals for these comparisons in this checkpoint.
19
𝐾
Expert profile
20 20 100 100
Weak Strong Weak Strong
Cells
Mean ΔOOD
12 12 12 12
+0.0080 +0.0070 +0.0046 +0.0072
Table 4: Descriptive paired CIFAR-100 nominal OOD AURSAC differences for Racer-kernel versus the classifier-confidence router. ΔOOD is AURSAC(Racer-kernel) minus AURSAC(classifier) on unseen-OOD experts. Each block contains four settings and three seeds. The cell count describes the experimental grid, not a count of independent replications. Figure 3 and Tables 3 and 4 support the intended bias–variance role of the learned kernel head. Racer-KNN is transparent and local, but its same-role neighborhoods become noisy when 𝐾 is large and finite context support is sparse. Racer-kernel keeps the same identity-free local evidence while learning how to smooth it using support, similarity mass, posterior confidence, entropy, and global context accuracy. The remaining oracle gap shows that finite-context instance-level competence estimation is still the limiting factor, especially in the full 𝐾 = 100 label space.
4.4
CIFAR-100 expert-correctness calibration
Calibration on the CIFAR-100 stress tests asks whether the competence score can be used as a probability, not merely as a ranking feature. This is central to Racer because the Bayes L2D decision compares model correctness and expert correctness on a shared probability scale. 𝐾 20 20 100 100
Profile Weak Strong Weak Strong
IFD-score
Racer-KNN
Racer-kernel
Brier
ECE
Brier
ECE
Brier
ECE
0.2472 0.2397 0.2504 0.2476
0.1109 0.1106 0.1126 0.1357
0.2788 0.2727 0.3274 0.3137
0.1725 0.1753 0.2500 0.2382
0.2281 0.2206 0.2293 0.2205
0.0209 0.0191 0.0135 0.0133
Table 5: OOD calibration of expert-correctness probabilities. Lower is better for both Brier score and 15-bin ECE. Values are averaged over all (𝜌, 𝜆id ) settings and seeds. Racer-kernel has the lowest Brier score and ECE in every 𝐾/profile setting, with ECE between 0.013 and 0.021. Its mean predicted expert correctness is also close to realized expert accuracy: for strong OOD experts, b 𝑞 𝜔 = 0.6712 versus expert accuracy 0.6696 at 𝐾 = 20, and b 𝑞 𝜔 = 0.6745 versus 0.6709 at 𝐾 = 100. IFD-score is systematically under-confident in the high-cardinality strong-expert setting, predicting 0.5352 mean correctness against realized accuracy 0.6709. Racer-KNN is closer to unbiased on average but has poor bin-wise calibration, which explains its high ECE and weak routing in Table 3. Overall, local role-relative evidence is useful, but high-cardinality finite-context deferral requires learned smoothing and calibration to turn that evidence into a reliable probability-scale competence estimate.
4.5
VinDr-CXR: real radiologist deferral
VinDr-CXR is a real-radiologist, multi-label deferral benchmark. We treat each finding as a binary-relevance task and report macro AURSBAC. Because the official VinDr-CXR test labels are 20
Figure 4: Real-expert chest-radiography results. Panels (i)–(ii) show macro balanced accuracy as the deferral budget increases; curves are means and shaded bands denote standard errors over matched runs. Panels (iii)–(iv) summarize expert-correctness calibration for methods with probability-scale outputs. consensus-only and have no radiologist identity, all L2D evaluation uses the multi-radiologist training annotations: for expert 𝑒, the deferred prediction is 𝑒’s own label and the target is the leave-one-radiologist-out consensus of the other two readers. We use the co-annotation-cluster holdout, which treats the dominant R8/R9/R10 cluster as the OOD expert group. No hospital shift is verified. Outcomes measure agreement with the other two readers, with a pathology masked when those readers disagree; they do not measure independent clinical ground truth. Protocol details, masking rules, and the co-annotation counts are in Section G. Table 6 gives the five-seed summary. The pooled “Overall” metric sweeps one global deferral budget over all valid rows; the group columns sweep the budget within each expert group. Racer-C has the best pooled/global-budget AURSBAC, the lowest Brier score and ECE, and a positive context-posterior NLL gain, indicating that its posterior adapter learns useful case-mix information on average. However, Racer-kernel is stronger in the group-conditional seen and unseen-OOD columns, and its unreported group-weighted AURSBAC is higher (0.7012 versus 0.6950 for Racer-C). Thus Racer-C is useful when a single global budget may be allocated unevenly across the pooled population, while Racer-kernel remains the more robust default under group-conditional expert-shift evaluation. The classifier-confidence baseline is also intentionally strong on VinDr-CXR. This is not surprising: the target radiologist is very accurate in this leave-one-radiologist-out construction, so simply deferring uncertain classifier cases to that radiologist is already a competitive policy. In the five-seed evaluation, the deferred expert has 97.00±0.03% raw accuracy overall, 98.62±0.41% on unseen-ID cases, and 96.28 ± 0.04% on unseen-OOD cases, whereas the image-only classifier is 90.26 ± 0.04% overall and 87.80 ± 0.06% on unseen-OOD cases. The classifier baseline’s strong OOD performance should therefore be read as a property of the benchmark: when the available expert is already highly reliable, the marginal room for learned expert-specific routing is small. Detailed expert-accuracy diagnostics are in Table 16.
4.6
CheXpert: multi-expert human–AI deferral
CheXpert evaluates a different real-world regime: multiple deferral targets are available at once, including human readers and external AI systems. The AI experts are CARZero, KAD, and CheXZero. Each run chooses one candidate human reader, includes two external AI systems as seen experts, and holds out the third AI system until test time. The primary metric is seen-plus-unseen macro AURSBAC. We use three candidate human readers, three held-out AI choices, and three seeds, yielding 27 matched cells per method. Details of the consensus construction, expert pools, and held-out AI protocol are in Section G. Table 7 reports the compact primary table. The two learned Racer variants are the strongest 21
Method
Pooled
Seen
ID
OOD
Brier ↓
ECE ↓
Ctx.
Racer-C 0.7019 ± 0.0016 0.7373 ± 0.0180 0.6359 ± 0.0363 0.6841 ± 0.0040 0.0297 ± 0.0004 0.0177 ± 0.0106 0.0542 ± 0.0250 Racer-kernel 0.6973 ± 0.0041 0.7436 ± 0.0155 0.6811 ± 0.0413 0.6901 ± 0.0057 0.0301 ± 0.0003 0.0243 ± 0.0055 0.0000 ± 0.0000 Classifier-conf. 0.6878 ± 0.0078 0.7057 ± 0.0152 0.7083 ± 0.0466 0.6898 ± 0.0077 – – 0.0000 ± 0.0000 L2D-QC [22] 0.6839 ± 0.0128 0.7005 ± 0.0605 0.6806 ± 0.0330 0.6787 ± 0.0110 – – 0.0000 ± 0.0000 IFD-MLP [20] 0.6761 ± 0.0098 0.6749 ± 0.0516 0.6868 ± 0.0629 0.6832 ± 0.0082 0.0324 ± 0.0010 0.0378 ± 0.0186 0.0000 ± 0.0000 L2D-QI [22] 0.6685 ± 0.0174 0.6788 ± 0.0580 0.6942 ± 0.0478 0.6805 ± 0.0054 – – 0.0000 ± 0.0000 Racer-KNN 0.6680 ± 0.0028 0.6779 ± 0.0234 0.6499 ± 0.0191 0.6768 ± 0.0024 0.0311 ± 0.0014 0.0207 ± 0.0017 0.0000 ± 0.0000 IFD-score [20] 0.6606 ± 0.0100 0.6487 ± 0.0853 0.7055 ± 0.0847 0.6813 ± 0.0059 0.0318 ± 0.0012 0.0303 ± 0.0117 0.0000 ± 0.0000
Table 6: VinDr-CXR real-radiologist deferral. Overall AURSBAC is pooled over all valid rows and therefore corresponds to one global deferral budget. Seen, unseen-ID, and unseen-OOD columns compute group-conditional AURSBAC. Brier and ECE evaluate expert-correctness probabilities and are omitted for methods that only output routing scores. Ctx. gain is the label-NLL improvement of 𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ) over the image-only posterior. Method Racer-C Racer-kernel IFD-MLP [20] IFD-score [20] L2D-QC [22] L2D-QI [22] Racer-KNN Classifier-conf.
Seen+unseen AURSBAC
Seen AURSBAC
Brier ↓
ECE ↓
0.7243 ± 0.0341 0.7241 ± 0.0259 0.7105 ± 0.0405 0.7013 ± 0.0444 0.7007 ± 0.0219 0.6986 ± 0.0265 0.6956 ± 0.0360 0.6931 ± 0.0360
0.7244 ± 0.0328 0.7276 ± 0.0224 0.7139 ± 0.0349 0.7092 ± 0.0431 0.7026 ± 0.0202 0.7029 ± 0.0224 0.6987 ± 0.0348 0.7001 ± 0.0340
0.0947 ± 0.0071 0.0916 ± 0.0068 0.0861 ± 0.0061 0.0849 ± 0.0052 – – 0.0844 ± 0.0061 –
0.0418 ± 0.0134 0.0366 ± 0.0098 0.0247 ± 0.0075 0.0293 ± 0.0067 – – 0.0534 ± 0.0088 –
Table 7: CheXpert multi-expert human–AI deferral. Values are mean ± standard deviation over 27 matched cells. Seen+unseen AURSBAC is the primary metric; Brier and ECE are reported only for probability-output methods and concern each method’s selected expert, not common expert–query pairs. by seen-plus-unseen AURSBAC, with nearly identical means. Racer-kernel is slightly stronger on seen experts and remains our main method because it improves balanced-accuracy routing without estimating the additional context-conditioned label posterior. Racer-C has a similar pooled mean. Both learned variants have higher mean AURSBAC than the IFD baselines in this table. These are descriptive comparisons; this version makes no significance claim. The extended metrics are in Section G.4.
4.7
Experimental summary, limitations, and future directions
The experiments support three main conclusions. First, role-relative competence estimation is most useful when expert skill has local structure that classifier confidence alone cannot capture. This is clearest in the PathMNIST context sweep, where Racer-kernel converts increasing context into positive routing gain under hidden subtype dependence. Second, identity-free local evidence must be smoothed carefully in high-cardinality settings: Racer-KNN becomes noisy as same-role support shrinks, while Racer-kernel remains the strongest aggregate router on the nominal OOD split of CIFAR-100 for both 𝐾 = 20 and 𝐾 = 100. Third, probability-scale competence estimation is not the same as maximizing a routing score. The learned kernel head gives the best Brier and ECE on the CIFAR-100 calibration stress tests, whereas smoother classwise profiles can have slightly lower ECE but weaker routing utility in the PathMNIST
22
large-context regime. The context-posterior stress test in Section D.3 confirms this choice: when 𝜆ctx = 0, Racer-C is essentially tied with Racer-kernel (AURSAC 0.8157 versus 0.8149), while at 𝜆ctx = 1, where context is maximally informative about case mix, Racer-C improves AURSAC by 0.0131 and reduces label NLL by 0.1154. Several limitations remain. The method relies on an informative query–context geometry; if the image representation does not organize examples by the factors that drive expert competence, local role-relative pooling cannot recover the right competence law. Finite context is also a fundamental constraint, especially when 𝐾 is large and the expected same-role support 𝐵/𝐾 is small. The real-expert benchmarks are valuable because they contain persistent raters, but they cover a limited range of clinical tasks and annotation protocols, and classifier confidence remains competitive. The higher reported mean for Racer-kernel does not by itself establish a statistically reliable advantage; this checkpoint reports descriptive comparisons without significance tests. Finally, calibration remains metric-dependent: Brier score rewards sharp useful probabilities, while ECE can favor smoother but less discriminative estimates. Code availability.
The code will be released upon publication.
References [1] Peter L Bartlett, Michael I Jordan, and Jon D McAuliffe. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101(473):138–156, 2006. [2] Yuzhou Cao, Hussein Mozannar, Lei Feng, Hongxin Wei, and Bo An. In defense of softmax parametrization for calibrated and consistent learning to defer. Advances in Neural Information Processing Systems, 36:38485–38503, 2023. [3] W Brier Glenn et al. Verification of forecasts expressed in terms of probability. Monthly weather review, 78(1):1–3, 1950. [4] Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR, 2017. [5] Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 590–597, 2019. [6] Jakob Nikolas Kather, Johannes Krisam, et al. Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study. PLOS Medicine, 16 (1):1–22, 01 2019. [7] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. [8] Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In International conference on machine learning, pages 3744–3753. PMLR, 2019. [9] David Madras, Toni Pitassi, and Richard Zemel. Predict responsibly: improving fairness and accuracy by learning to defer. Advances in neural information processing systems, 31, 2018. 23
[10] Anqi Mao, Christopher Mohri, Mehryar Mohri, and Yutao Zhong. Two-stage learning to defer with multiple experts. Advances in neural information processing systems, 36:3578–3606, 2023. [11] Yannis Montreuil, Shu Heng Yeo, Axel Carlier, Lai Xing Ng, and Wei Tsang Ooi. Two-stage learning-to-defer for multi-task learning. arXiv preprint arXiv:2410.15729, 2024. [12] Yannis Montreuil, Axel Carlier, Lai Xing Ng, and Wei Tsang Ooi. Adversarial robustness in two-stage learning-to-defer: Algorithms and guarantees. arXiv preprint arXiv:2502.01027, 2025. [13] Yannis Montreuil, Axel Carlier, Lai Xing Ng, and Wei Tsang Ooi. Why ask one when you can ask 𝑘? two-stage learning-to-defer to the top-𝑘 experts. arXiv preprint arXiv:2504.12988, 2025. [14] Hussein Mozannar and David Sontag. Consistent estimators for learning to defer to an expert. In International Conference on Machine Learning, pages 7076–7087. PMLR, 2020. [15] Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In Proceedings of the AAAI conference on artificial intelligence, volume 29, 2015. [16] Ha Quy Nguyen, Hieu Huy Pham, le tuan linh, Minh Dao, and lam khanh. VinDr-CXR: An open dataset of chest X-rays with radiologist annotations. PhysioNet, June 2021. doi: 10.13026/3akn-b287. URL https://doi.org/10.13026/3akn-b287. Version 1.0.0. [17] Joshua Strong, Qianhui Men, and J. Alison Noble. Trustworthy and practical AI for healthcare: A guided deferral system with large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 28413–28421, 2025. doi: 10.1609/aaai.v39i27.35063. URL https://doi.org/10.1609/aaai.v39i27.35063. [18] Joshua Strong, Emma Sun, Harry Rogers, Helen Higham, and Alison Noble. Learning to defer: A survey, December 2025. URL https://doi.org/10.5281/zenodo.17843044. [19] Joshua Strong, Harry Rogers, Emma Sun, Anna Louise Todsen, Jody Ede, Cherry Lumley, Nick Yeung, Helen Higham, and J. Alison Noble. Human-AI collaboration in healthcare: A scoping review. npj Digital Medicine, 2026. doi: 10.1038/s41746-026-02918-6. URL https://doi.org/10.1038/s41746-026-02918-6. [20] Joshua Strong, Pramit Saha, Yasin Ibrahim, Cheng Ouyang, and Alison Noble. Identityfree deferral for unseen experts. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=4YG9ufFg58. [21] Joshua Strong, Pramit Saha, Emma Sun, Helen Higham, and J. Alison Noble. Coherent hierarchical multi-label learning to defer for medical imaging. arXiv preprint arXiv:2605.02734, 2026. doi: 10.48550/arXiv.2605.02734. URL https://arxiv.org/abs/2605.02734. [22] Dharmesh Tailor, Aditya Patra, Rajeev Verma, Putra Manggala, and Eric Nalisnick. Learning to defer to a population: A meta-learning approach. In International Conference on Artificial Intelligence and Statistics, pages 3475–3483. PMLR, 2024. [23] Rajeev Verma and Eric Nalisnick. Calibrated learning to defer with one-vs-all classifiers. In International Conference on Machine Learning, pages 22184–22202. PMLR, 2022.
24
[24] Rajeev Verma, Daniel Barrejón, and Eric Nalisnick. Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles. In International Conference on Artificial Intelligence and Statistics, pages 11415–11434. PMLR, 2023. [25] Jiancheng Yang, Rui Shi, and Bingbing Ni. Medmnist classification decathlon: A lightweight automl benchmark for medical image analysis. In IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 191–195, 2021. [26] Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification. Scientific Data, 10(1):41, 2023. [27] Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. Deep sets. Advances in neural information processing systems, 30, 2017.
25
Contents of Appendix A Detailed Proof Derivations A.1 Conditional minimizer of the augmented L2D surrogate . . . . . . . . . . . . . . A.2 Expanded proof of Proposition 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.3 Expanded proof of Proposition 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.4 Expanded proof of Proposition 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.5 Expanded proof of Theorem 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . A.6 Expanded proof of Theorem 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
27 27 28 28 29 30 31
B Relaxing Context-Posterior Invariance
32
C Computational Complexity
34
D Synthetic Protocol Details D.1 PathMNIST context-scaling protocol . . . . . . . . . . . . . . . . . . . . . . . . . . D.2 Representation-geometry ablation . . . . . . . . . . . . . . . . . . . . . . . . . . . D.3 Context-posterior stress-test protocol . . . . . . . . . . . . . . . . . . . . . . . . . . D.4 CIFAR-100 class-scale protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
34 34 35 36 36
E Extended Literature Review E.1 Learning with rejection and learning to defer . . . . . . . . . . . . . . . . . . . . . E.2 Clinical collaboration and hierarchical deferral . . . . . . . . . . . . . . . . . . . . E.3 Calibration in fixed-expert and fixed multi-expert L2D . . . . . . . . . . . . . . . . E.4 Population-adaptive deferral and unseen experts . . . . . . . . . . . . . . . . . . . E.5 Decision consistency, probability estimation, and representation . . . . . . . . . . E.6 Identity-free and equivariant architectures . . . . . . . . . . . . . . . . . . . . . . . E.7 Real-human expert benchmarks . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
38 38 38 38 39 39 42 42
F Descriptive comparisons and variability
42
G Additional protocol notes G.1 Implementation and evaluation protocol . . . . . . . . . . . . . . . . . . . . . . . . G.2 VinDr-CXR protocol details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . G.3 CheXpert multi-expert protocol details . . . . . . . . . . . . . . . . . . . . . . . . . G.4 Extended CheXpert metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
42 42 44 44 45
H Ablations 45 H.1 Ablation: Role-relative DeepSets competence estimator . . . . . . . . . . . . . . . 45
26
A
Detailed Proof Derivations
This appendix expands the short proof sketches in the main text. The statements and assumptions are the same as in the corresponding main-body propositions and theorems; the goal here is only to show the intermediate algebra.
A.1
Conditional minimizer of the augmented L2D surrogate
For completeness, we first derive the conditional minimizer used in Section 2.1. Fix 𝑥 and write 𝑞(𝑥) = P(𝑀 = 𝑌 | 𝑋 = 𝑥). The conditional augmented-softmax risk is 𝒮(Π | 𝑥) = −
Õ
𝜂 𝑦 (𝑥) log Π 𝑦 − 𝑞(𝑥) log Π⊥ ,
𝑦∈𝒴
where Π ∈ Δ(𝒴 ⊥ ). Define nonnegative weights 𝑤 𝑦 B 𝜂 𝑦 (𝑥), with total mass 𝑊 B
Õ
𝑤⊥ B 𝑞(𝑥),
𝑤 𝑦 + 𝑤 ⊥ = 1 + 𝑞(𝑥).
𝑦
The optimization problem is min Í
Π 𝑎 ≥0,
Õ
−
𝑎 Π 𝑎 =1
𝑤 𝑎 log Π𝑎 .
𝑎∈𝒴 ⊥
For coordinates with 𝑤 𝑎 > 0, the Lagrangian is
! 𝒥 (Π, 𝜆) B −
Õ
𝑤 𝑎 log Π𝑎 + 𝜆
Õ
𝑎
Π𝑎 − 1 .
𝑎
The stationarity condition gives 𝜕𝒥 𝑤𝑎 =− + 𝜆 = 0, Π𝑎 𝜕Π𝑎 Summing over 𝑎 and using 1=
Í
Π𝑎 =
𝑤𝑎 . 𝜆
𝑎 Π 𝑎 = 1 yields
Õ𝑤 𝑎
so
𝑎
𝜆
=
Therefore Π★𝑦 (𝑥) =
𝑊 , 𝜆 𝜂 𝑦 (𝑥) 1 + 𝑞(𝑥)
hence
𝜆 = 𝑊 = 1 + 𝑞(𝑥).
Π★⊥ (𝑥) =
,
𝑞(𝑥) . 1 + 𝑞(𝑥)
Coordinates with zero weight receive zero probability in the closure of the simplex; the same formula applies by continuity. The induced decision is Bayes-aligned because Π★⊥ (𝑥) ≥ max Π★𝑦 (𝑥) 𝑦
⇐⇒
max 𝑦 𝜂 𝑦 (𝑥) 𝑞(𝑥) ≥ 1 + 𝑞(𝑥) 1 + 𝑞(𝑥)
27
⇐⇒
𝑞(𝑥) ≥ max 𝜂 𝑦 (𝑥). 𝑦
A.2
Expanded proof of Proposition 1
Let 𝐴 = 1{𝑀𝐸 = 𝑌} and let 𝑈 = 𝑇(𝑋 , 𝑌, 𝐶𝐸 , 𝑝 𝜔 ) be the information exposed to a scalar predictor 𝑔(𝑈) ∈ (0, 1). The binary cross-entropy risk is ℒ(𝑔) B E −𝐴 log 𝑔(𝑈) − (1 − 𝐴) log(1 − 𝑔(𝑈)) .
Using the tower property, this risk decomposes pointwise in 𝑈: ℒ(𝑔) = E𝑈 E −𝐴 log 𝑔(𝑈) − (1 − 𝐴) log(1 − 𝑔(𝑈)) | 𝑈
.
Fix a value 𝑈 = 𝑢 and define 𝛼(𝑢) B P(𝐴 = 1 | 𝑈 = 𝑢). For a scalar prediction 𝑠 ∈ (0, 1), the conditional risk is ℓ 𝑢 (𝑠) B −𝛼(𝑢) log 𝑠 − (1 − 𝛼(𝑢)) log(1 − 𝑠). Its first derivative is
𝛼(𝑢) 1 − 𝛼(𝑢) + , 𝑠 1−𝑠 and the stationarity condition ℓ 𝑢′ (𝑠) = 0 gives ℓ 𝑢′ (𝑠) = −
𝛼(𝑢) 1 − 𝛼(𝑢) = 𝑠 1−𝑠
⇐⇒
𝛼(𝑢)(1 − 𝑠) = (1 − 𝛼(𝑢))𝑠
The second derivative is ℓ 𝑢′′ (𝑠) =
⇐⇒
𝑠 = 𝛼(𝑢).
𝛼(𝑢) 1 − 𝛼(𝑢) + >0 𝑠2 (1 − 𝑠)2
for 𝑠 ∈ (0, 1) whenever 0 < 𝛼(𝑢) < 1, so this stationary point is the unique minimizer. If 𝛼(𝑢) ∈ {0, 1}, the infimum is attained at the corresponding boundary value by continuity. Hence the population minimizer over measurable functions of 𝑈 is 𝑔★(𝑈) = P(𝐴 = 1 | 𝑈). When 𝑈 = (𝑋 , 𝑌, 𝐶𝐸 ), this becomes 𝑔★(𝑥, 𝑦, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ). On the event 𝑌 = 𝑦, expert correctness 𝑀𝐸 = 𝑌 is the event 𝑀𝐸 = 𝑦, so
P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ). Thus binary cross-entropy is a strictly proper loss for the rolewise competence probability, and a restricted summary 𝑈 = 𝑢𝑒 ,𝑦 (𝑥) targets the corresponding conditional projection P(𝐴 = 1 | 𝑢𝑒 ,𝑦 (𝑥)).
A.3
Expanded proof of Proposition 2
By definition,
𝑞★(𝑥, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ).
Condition on the unknown true class and apply the law of total probability: 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 )P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ).
𝑦∈𝒴
28
On the event 𝑌 = 𝑦, the event 𝑀𝐸 = 𝑌 is the same as 𝑀𝐸 = 𝑦, so
P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ). Under Assumption 1,
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) = P(𝑌 = 𝑦 | 𝑋 = 𝑥) = 𝜂 𝑦 (𝑥). Therefore
𝑞★(𝑥, 𝐶 𝑒 ) =
Õ
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ).
𝑦
If 𝑝 𝜔 (𝑦 | 𝑥) = 𝜂 𝑦 (𝑥) and Γ(𝑥, 𝑦, 𝐶 𝑒 ) = Γ★(𝑥, 𝑦, 𝐶 𝑒 ) for all 𝑦, then the plug-in estimate satisfies
b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) =
Õ
𝑝 𝜔 (𝑦 | 𝑥)Γ(𝑥, 𝑦, 𝐶 𝑒 ) =
𝑦
Õ
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ) = 𝑞★(𝑥, 𝐶 𝑒 ).
𝑦
This proves that the induced expert-correctness estimate is calibrated with respect to the information (𝑥, 𝐶 𝑒 ) when both the classifier posterior and the rolewise competence functional equal their respective conditional probability targets.
A.4
Expanded proof of Proposition 3
Fix (𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) and abbreviate 𝑞 = 𝑞★(𝑥, 𝐶 𝑒 ),
𝑞ˆ = b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ),
𝜂max = max 𝜂 𝑦 (𝑥), 𝑦
𝑝ˆ max = 𝑝 𝜔,max (𝑥).
Let 𝑦★ ∈ arg max 𝑦 𝜂 𝑦 (𝑥) and 𝑦ˆ B ℎ 𝜔 (𝑥) ∈ arg max 𝑦 𝑝 𝜔 (𝑦 | 𝑥). The Bayes conditional correctness is max{𝜂max , 𝑞}. The plug-in system predicts class 𝑦ˆ when 𝑞ˆ < 𝑝ˆ max and defers otherwise, so its conditional correctness is (1 − b 𝑟0 )𝜂 𝑦ˆ + b 𝑟0 𝑞. We first separate classifier error from routing error. If the plug-in routing decision agrees with the Bayes comparison between 𝑞 and 𝜂max , then the only possible loss relative to Bayes occurs when both systems predict, in which case the plug-in classifier receives 𝜂 𝑦ˆ instead of 𝜂max . If the routing signs disagree, the additional loss is at most |𝑞 − 𝜂max |. Hence the conditional excess risk is bounded by 𝜂max − 𝜂 𝑦ˆ + |𝑞 − 𝜂max |1{sign(𝑞 − 𝜂max ) ≠ sign( 𝑞ˆ − 𝑝ˆ max )}. The classifier term is controlled by posterior error. Since 𝑝 𝜔 ( 𝑦ˆ | 𝑥) ≥ 𝑝 𝜔 (𝑦★ | 𝑥), 𝜂max − 𝜂 𝑦ˆ = (𝜂 𝑦★ − 𝑝 𝑦★ ) + (𝑝 𝑦★ − 𝑝 𝑦ˆ ) + (𝑝 𝑦ˆ − 𝜂 𝑦ˆ ) ≤ |𝜂 𝑦★ − 𝑝 𝑦★ | + |𝑝 𝑦ˆ − 𝜂 𝑦ˆ | ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 . On the sign-disagreement event, the two real numbers 𝑞 − 𝜂max and 𝑞ˆ − 𝑝ˆ max have opposite signs or one is zero. Therefore the distance from 𝑞 − 𝜂max to zero is no larger than its distance to 𝑞ˆ − 𝑝ˆ max : |𝑞 − 𝜂max | ≤ |(𝑞 − 𝜂max ) − ( 𝑞ˆ − 𝑝ˆ max )| ≤ |𝑞 − 𝑞ˆ | + |𝜂max − 𝑝ˆ max |. 29
The maximum posterior map is 1-Lipschitz in ℓ∞ , hence also in ℓ1 : |𝜂max − 𝑝ˆ max | ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥∞ ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 . It remains to bound | 𝑞ˆ − 𝑞|. Write 𝑞ˆ − 𝑞 =
Õ
𝑝 𝑦 Γ𝑦 −
Õ
𝑦
𝜂 𝑦 Γ★𝑦 =
Õ
Õ
𝑦
𝑦
𝑦
(𝑝 𝑦 − 𝜂 𝑦 )Γ 𝑦 +
𝜂 𝑦 (Γ 𝑦 − Γ★𝑦 ),
where 𝑝 𝑦 B 𝑝 𝜔 (𝑦 | 𝑥), Γ 𝑦 B Γ(𝑥, 𝑦, 𝐶 𝑒 ), and Γ★𝑦 B Γ★(𝑥, 𝑦, 𝐶 𝑒 ). Since 0 ≤ Γ 𝑦 ≤ 1 and | 𝑞ˆ − 𝑞| ≤
Õ
|𝑝 𝑦 − 𝜂 𝑦 | |Γ 𝑦 | +
Õ
𝑦
𝜂 𝑦 |Γ 𝑦 − Γ★𝑦 | ≤ ∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 +
𝑦
Õ
Í
𝑦 𝜂 𝑦 = 1,
𝜂 𝑦 |Γ 𝑦 − Γ★𝑦 |.
𝑦
Combining the three bounds gives the conditional excess risk ≤ 3∥𝑝 𝜔 (· | 𝑥) − 𝜂(· | 𝑥)∥1 +
Õ
𝜂 𝑦 (𝑥)|Γ(𝑥, 𝑦, 𝐶 𝑒 ) − Γ★(𝑥, 𝑦, 𝐶 𝑒 )|.
𝑦
Taking expectation over (𝑋 , 𝐶𝐸 ) proves Proposition 3.
A.5
Expanded proof of Theorem 1
Fix a coherent class relabelling 𝜋 ∈ 𝔖𝐾 . The relabelled context is 𝜋𝐶 𝑒 = {(𝑥 𝑖𝐶 , 𝜋(𝑦 𝑖𝐶 ), 𝜋(𝑚 𝑖𝐶 ))}𝐵𝑖=1 , and posterior equivariance means 𝑝 𝜋𝜔 (𝜋(𝑦) | 𝑥) = 𝑝 𝜔 (𝑦 | 𝑥) for every role 𝑦. We verify each ingredient of the RACER summaries. First, same-role masks are preserved: 1{𝜋(𝑦 𝑖𝐶 ) = 𝜋(𝑦)} = 1{𝑦 𝑖𝐶 = 𝑦}. Context correctness is also preserved: 1{𝜋(𝑚 𝑖𝐶 ) = 𝜋(𝑦 𝑖𝐶 )} = 1{𝑚 𝑖𝐶 = 𝑦 𝑖𝐶 }. Image features and query–context similarities do not involve class names, so 𝐾 𝑖 (𝑥) is unchanged. It follows that Õ Õ 𝑁𝑒𝜋,𝜋(𝑦) = 1{𝜋(𝑦 𝑖𝐶 ) = 𝜋(𝑦)} = 1{𝑦 𝑖𝐶 = 𝑦} = 𝑁𝑒 ,𝑦 , 𝑖
and
𝑆 𝜋𝑒 ,𝜋(𝑦) (𝑥) =
Õ
𝑖
1{𝜋(𝑦 𝑖𝐶 ) = 𝜋(𝑦)}𝐾 𝑖 (𝑥) =
𝑖
Õ
1{𝑦 𝑖𝐶 = 𝑦}𝐾 𝑖 (𝑥) = 𝑆 𝑒 ,𝑦 (𝑥).
𝑖
The global context prior 𝜇𝑒 ,0 depends only on the preserved correctness indicators and the context size, so it is unchanged. Therefore the smoothed local statistic obeys 𝑎¯ 𝜋𝑒 ,𝜋(𝑦) (𝑥) =
𝛼0 𝜇𝑒 ,0 +
Í
𝐶 𝐶 𝐶 𝑖 1{𝜋(𝑦 𝑖 ) = 𝜋(𝑦)}𝐾 𝑖 (𝑥)1{𝜋(𝑚 𝑖 ) = 𝜋(𝑦 𝑖 )} = 𝑎¯ 𝑒 ,𝑦 (𝑥). 𝛼0 + 𝑆 𝜋𝑒 ,𝜋(𝑦) (𝑥)
30
Posterior equivariance preserves the role posterior value, 𝑝 𝜋𝜔 (𝜋(𝑦) | 𝑥) = 𝑝 𝜔 (𝑦 | 𝑥), the rank of the role,
𝑅 𝜋𝜔 (𝜋(𝑦); 𝑥) = 𝑅 𝜔 (𝑦; 𝑥),
the role margin, and the entropy of the posterior vector. Hence the full summary satisfies 𝑢𝑒𝜋,𝜋(𝑦) (𝑥) = 𝑢𝑒 ,𝑦 (𝑥). For Racer-KNN, this immediately gives Γ𝜋KNN (𝑥, 𝜋(𝑦), 𝜋𝐶 𝑒 ) = ΓKNN (𝑥, 𝑦, 𝐶 𝑒 ). For Racer-kernel, the same equality holds because the MLP 𝑔𝜃 is shared across roles: Γ𝜋𝜃 (𝑥, 𝜋(𝑦), 𝜋𝐶 𝑒 ) = sigm(𝑔𝜃 (𝑢𝑒𝜋,𝜋(𝑦) (𝑥))) = sigm(𝑔𝜃 (𝑢𝑒 ,𝑦 (𝑥))) = Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ). Finally, change variables 𝑦 ′ = 𝜋(𝑦) in the expert-correctness marginal:
b 𝑞 𝜋𝜔 (𝑥, 𝜋𝐶 𝑒 ) = =
Õ
𝑝 𝜋𝜔 (𝑦 ′ | 𝑥) Γ𝜋 (𝑥, 𝑦 ′ , 𝜋𝐶 𝑒 )
𝑦 ′ ∈𝒴
Õ
𝑝 𝜋𝜔 (𝜋(𝑦) | 𝑥) Γ𝜋 (𝑥, 𝜋(𝑦), 𝜋𝐶 𝑒 )
𝑦∈𝒴
=
Õ
𝑝 𝜔 (𝑦 | 𝑥) Γ(𝑥, 𝑦, 𝐶 𝑒 )
𝑦∈𝒴
=b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ). The maximum posterior is also invariant because relabelling only permutes the coordinates: 𝑝 𝜋𝜔,max (𝑥) = 𝑝 𝜔,max (𝑥). Thus any plug-in threshold rule depending on b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) − 𝑝 𝜔,max (𝑥) is coherently invariant.
A.6
Expanded proof of Theorem 2
Fix (𝑥, 𝐶 𝑒 ) and consider the idealized conditional loss ℓ (Π | 𝑥, 𝐶 𝑒 ) B E − log Π𝑌 − Γ★(𝑥, 𝑌, 𝐶 𝑒 ) log Π⊥ | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 .
Expanding the expectation over 𝑌 gives ℓ (Π | 𝑥, 𝐶 𝑒 ) = −
Õ
𝜂 𝑦 (𝑥) log Π 𝑦 −
𝑦
By definition,
Õ
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ) log Π⊥ .
𝑦
𝑞★(𝑥, 𝐶 𝑒 ) =
Õ
𝜂 𝑦 (𝑥)Γ★(𝑥, 𝑦, 𝐶 𝑒 ),
𝑦
so the conditional loss is ℓ (Π | 𝑥, 𝐶 𝑒 ) = −
Õ
𝜂 𝑦 (𝑥) log Π 𝑦 − 𝑞★(𝑥, 𝐶 𝑒 ) log Π⊥ .
𝑦
31
This is a weighted cross-entropy over the augmented action space. The weights are 𝑤⊥ = 𝑞★(𝑥, 𝐶 𝑒 ),
𝑤 𝑦 = 𝜂 𝑦 (𝑥), with total mass 𝑊=
Õ
𝜂 𝑦 (𝑥) + 𝑞★(𝑥, 𝐶 𝑒 ) = 1 + 𝑞★(𝑥, 𝐶 𝑒 ).
𝑦
The same Lagrange multiplier calculation as in Section A.1 gives the minimizer Π★𝑦 (𝑥, 𝐶 𝑒 ) =
𝜂 𝑦 (𝑥) 1 + 𝑞★(𝑥, 𝐶 𝑒 )
Π★⊥ (𝑥, 𝐶 𝑒 ) =
,
𝑞★(𝑥, 𝐶 𝑒 ) . 1 + 𝑞★(𝑥, 𝐶 𝑒 )
The denominator is positive, so the induced defer/predict comparison is Π★⊥ (𝑥, 𝐶 𝑒 ) ≥ max Π★𝑦 (𝑥, 𝐶 𝑒 ) 𝑦
⇐⇒
max 𝑦 𝜂 𝑦 (𝑥) 𝑞★(𝑥, 𝐶 𝑒 ) ≥ ★ 1 + 𝑞 (𝑥, 𝐶 𝑒 ) 1 + 𝑞★(𝑥, 𝐶 𝑒 )
⇐⇒
𝑞★(𝑥, 𝐶 𝑒 ) ≥ max 𝜂 𝑦 (𝑥). 𝑦
This is exactly the Bayes L2D comparison with the posterior-predictive expert correctness probability conditioned on the available expert context.
B
Relaxing Context-Posterior Invariance
The main method assumes Assumption 1, so that the expert context is used to infer expert competence but not the label posterior of the query. This appendix details Racer-C, the context-conditioned classification extension of Racer-kernel, for deployments in which 𝐶 𝑒 may also reveal case mix, site, workflow, or acquisition information that changes P(𝑌 | 𝑋). Without Assumption 1, the Bayes-relevant expert correctness probability is 𝑞★(𝑥, 𝐶 𝑒 ) =
Õ 𝑦∈𝒴
P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) . | {z }| {z }
(10)
Γ★ (𝑥,𝑦,𝐶 𝑒 )
𝜂 𝐶𝑦 (𝑥,𝐶 𝑒 )
The competence term is unchanged: Racer still estimates Γ★(𝑥, 𝑦, 𝐶 𝑒 ). The additional object is the context-conditioned label posterior 𝜂 𝐶𝑦 (𝑥, 𝐶 𝑒 ) B P(𝑌 = 𝑦 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ), which can be approximated by a role-equivariant posterior adapter 𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ). A role-equivariant posterior adapter. Let 𝑝 𝜔 (𝑦 | 𝑥) be the image-only classifier posterior. A minimal implementation keeps this posterior as a strong prior and learns a shared role-relative logit adjustment, 𝑏 𝜌,𝑦 (𝑥, 𝐶 𝑒 ) B log 𝑝 𝜔 (𝑦 | 𝑥) + 𝑎 𝜌 (𝑥, 𝑦, 𝐶 𝑒 ),
𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ) B softmax 𝑦 𝑏 𝜌,𝑦 (𝑥, 𝐶 𝑒 ).
(11)
The code implements the equivalent construction 𝑏 𝜌,𝑦 = 𝑓𝜔,𝑦 + 𝑎 𝜌 using the original pre-softmax class logits. The two forms differ only by a common softmax-normalization constant; in particular, 𝑎 𝜌 = 0 recovers 𝑝 𝜔 . This describes the computation used by the supplied implementation. Here 𝑎 𝜌 is produced by the same network for every candidate role. Its inputs should be role-admissible summaries, for example the image-only posterior value and rank at role 𝑦, the same-role support 32
𝑁𝑒 ,𝑦 , the same-role similarity mass 𝑆 𝑒 ,𝑦 (𝑥), local context label density, and symmetric context summaries. To preserve coherent relabelling equivariance, the adapter should not use learned class embeddings, untied class-specific heads, or fixed class-coordinate channels. It is useful to separate two context channels conceptually. The posterior adapter should use 𝐶 𝑒𝑌 B {(𝑥 𝑖𝐶 , 𝑦 𝑖𝐶 )}𝐵𝑖=1 to estimate case-mix information, while the competence head uses 𝐶 𝑒𝑀 B {(𝑥 𝑖𝐶 , 𝑦 𝑖𝐶 , 𝑚 𝑖𝐶 )}𝐵𝑖=1 to estimate expert behavior. The expert predictions 𝑚 𝑖𝐶 are therefore part of the competence signal, not the default label-posterior signal. The resulting Racer-C expert-correctness estimate is
b 𝑞 𝜌,𝜃 (𝑥, 𝐶 𝑒 ) B
Õ
𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 )Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ).
(12)
𝑦∈𝒴
The autonomous prediction and plug-in rejector should use the same context-conditioned posterior, ℎ 𝜌 (𝑥, 𝐶 𝑒 ) B arg max 𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ), 𝑦∈𝒴
b 𝑟𝜏𝐶 (𝑥, 𝐶 𝑒 ) B 1{b 𝑞 𝜌,𝜃 (𝑥, 𝐶 𝑒 ) − max 𝜋𝜌 (𝑦 | 𝑥, 𝐶 𝑒 ) ≥ 𝜏}. (13) 𝑦
Using 𝜋𝜌 in (12) but comparing against 𝑝 𝜔,max (𝑥) would be internally inconsistent: if the context changes the label posterior, it changes both the expert-correctness marginalization and the model-correctness side of the Bayes comparison. Training. The posterior adapter can be trained episodically with a standard multiclass proper loss, ℒpost (𝜌) B −E log 𝜋𝜌 (𝑌 | 𝑋 , 𝐶𝐸𝑌 ), (14) where the target query is held out from its context. The competence head retains Section 3.1 and Proposition 1’s binary proper loss for Γ𝜃 . A plug-in implementation can therefore minimize ℒpost + 𝜆comp ℒcomp , with an optional Bayes-aligned deferral surrogate obtained by replacing the class logits in Section 3.7 with the context-conditioned logits 𝑏 𝜌,𝑦 (𝑥, 𝐶 𝑒 ) and replacing b 𝑞 𝜔 with b 𝑞 𝜌,𝜃 . Why this is not the default method. Racer-C must estimate both the rolewise competence law Γ★ and the context-conditioned label posterior 𝜂 𝐶 . Errors in this additional posterior affect both expert-correctness estimates and classifier confidence. Estimating it can increase variance when contexts are small or 𝐵/𝐾 is low. It may also learn associations with context composition that fail to transfer across deployments. Under Assumption 1, the target satisfies 𝜂 𝐶𝑦 (𝑥, 𝐶 𝑒 ) = 𝜂 𝑦 (𝑥), so context provides no additional information about the query label given 𝑥. We therefore use Racer-kernel as the default under this assumption and evaluate Racer-C as an optional extension when context is informative about query case mix.
33
C
Computational Complexity
For one query and one candidate expert, a direct implementation of the rolewise sums in (7)–(8) would cost 𝑂(𝐾𝐵) after image features are available. In practice the context embeddings, labels, correctness indicators, 𝑁𝑒 ,𝑦 , and 𝜇𝑒 ,0 are cached for each expert. The query embedding is computed once, the 𝐵 query–context similarities cost 𝑂(𝐵𝑑) for feature dimension 𝑑, and the rolewise numerator and denominator are obtained by one scatter-add over the 𝐵 context items, costing 𝑂(𝐵 + 𝐾). Applying the shared calibration MLP to all roles costs 𝑂(𝐾𝑐 𝑔 ) for a small per-role network cost 𝑐 𝑔 , and the final plug-in sum is 𝑂(𝐾). Thus the optimized per-query, per-expert cost is 𝑂(𝐵𝑑 + 𝐵 + 𝐾𝑐 𝑔 ) plus the classifier forward pass, rather than 𝑂(𝐾𝐵) pooling. Routing among 𝐽 candidate experts scales linearly in 𝐽 but is batchable over experts and roles. For very large 𝐾 or 𝐵, the same form admits standard engineering approximations, such as evaluating only high-posterior roles or using approximate nearest-neighbor context retrieval.
D
Synthetic Protocol Details
This appendix collects protocol details omitted from the main experiments section. The synthetic experiments have four parts: a PathMNIST context-scaling study that fixes the hardest locally complementary expert setting and varies context size, a representation-geometry ablation that tests whether subtype recoverability rather than simulator leakage explains the PathMNIST gains, a CIFAR-100 class-scale study that varies class cardinality, expert strength, hidden subtype dependence, and expert-profile permutation, and a context-posterior stress test that controls the degree to which the expert context reveals query case mix.
D.1
PathMNIST context-scaling protocol
The PathMNIST context-scaling experiment uses the MedMNIST PathMNIST dataset at 64px resolution with all nine classes. The run contains 107,180 total images, split into 89,996 training, 10,004 validation, and 7,180 test images. The context size 𝐵 is varied while the class count is fixed at 𝐾 = 9, and the main figures plot the normalized support ratio 𝐵/𝐾: 𝐵 𝐵/𝐾
9 1.0
25 2.8
50 5.6
100 11
200 22
500 56
1000 111
Table 8: PathMNIST context sizes used in the context-scaling ablation. We use the strongest locally complementary synthetic expert profile: full within-class subtype dependence (𝜌 = 1), full expert-profile permutation (𝜆id = 1), and good/mid/bad subtype accuracies (0.98, 0.70, 0.30) with three hidden subtypes per class. Experts are simulated by extracting image features, clustering training images into class-conditional hidden subtypes, assigning each expert a class/subtype competence tensor, and sampling expert predictions for every image. ID experts preserve the class-role profile; nominal OOD experts receive an expert-only permuted competence profile across observed classes. Each run uses 32 seen training experts, 8 unseen-ID validation experts, 8 unseen-OOD validation experts, 32 seen evaluation experts, 8 unseen-ID test experts, and 8 unseen-OOD test experts. Early stopping uses unseen-ID validation experts. At evaluation time, fixed context/query episodes are sampled and reused across expert groups and methods; context sets are class-balanced when possible. Each point uses three seeds.
34
Each method is trained separately for each 𝐵; this is not an inference-only context sweep of a fixed checkpoint. Larger context also leaves fewer query images within each evaluation split, so the query set need not be identical across context sizes.
D.2
Representation-geometry ablation
The main PathMNIST study uses class-conditional image-feature clusters to create hidden expert subtypes, and Racer also uses query–context geometry to recover local competence. We therefore test whether the gains depend on using the same representation to generate and retrieve hidden subtypes. Holding the PathMNIST setting fixed at 𝐾 = 9, 𝐵 = 500, 𝜌 = 1, 𝜆id = 1, and subtype accuracies (0.98, 0.70, 0.30), we generate hidden subtypes from mixtures of classifier features and auxiliary image-quality features, √
𝜓 𝛼 (𝑥) B normalize 𝛼𝜙cls (𝑥) +
1 − 𝛼2 𝜓
aux (𝑥)
,
(15)
with 𝛼 ∈ {1.0, 0.75, 0.5, 0.25, 0.0}. Here 𝜙cls is the default classifier/deep representation and 𝜓aux contains low-level color, contrast, texture, sharpness, and stain-quality summaries. For this ablation only, we also evaluate an auxiliary-geometry Racer-kernel variant that keeps the same competence interface but uses 𝜓aux for same-role similarity. Finally, we include a random-subtype negative control in which hidden subtypes are assigned independently within class. This destroys query–context recoverability while preserving the same class/subtype expert-competence tensor. Method
𝛼 = 1.00
Classifier-conf. 0.8134 ± 0.0099 IFD-score [20] 0.8295 ± 0.0034 IFD-MLP [20] 0.8114 ± 0.0048 L2D-Pop (QC) [22] 0.8124 ± 0.0038 Racer-KNN (Ours) 0.8313 ± 0.0052 Racer-kernel (Ours) 0.8432 ± 0.0030 Racer-kernel-aux (Ours) 0.8312 ± 0.0026
𝛼 = 0.75
𝛼 = 0.50
𝛼 = 0.25
0.8193 ± 0.0032 0.8357 ± 0.0017 0.8179 ± 0.0053 0.8104 ± 0.0201 0.8342 ± 0.0064 0.8583 ± 0.0045 0.8356 ± 0.0027
0.8137 ± 0.0117 0.8345 ± 0.0067 0.8151 ± 0.0074 0.8132 ± 0.0036 0.8453 ± 0.0076 0.8563 ± 0.0052 0.8367 ± 0.0074
0.8172 ± 0.0035 0.8324 ± 0.0028 0.8156 ± 0.0056 0.8106 ± 0.0230 0.8330 ± 0.0019 0.8573 ± 0.0025 0.8342 ± 0.0122
𝛼 = 0.00
Random
0.8167 ± 0.0044 0.8231 ± 0.0011 0.8342 ± 0.0054 0.8081 ± 0.0024 0.8193 ± 0.0091 0.8089 ± 0.0076 0.8173 ± 0.0040 0.8122 ± 0.0021 0.8397 ± 0.0101 0.8121 ± 0.0009 0.8593 ± 0.0068 0.8136 ± 0.0010 0.8363 ± 0.0055 0.8194 ± 0.0009
Table 9: PathMNIST representation-geometry ablation. Entries are overall AURSAC mean ± standard deviation over three seeds at 𝐵 = 500, 𝐾 = 9, 𝜌 = 1, and 𝜆id = 1. The 𝛼 columns generate hidden subtypes from classifier/auxiliary feature mixtures using (15). The random column assigns hidden subtypes independently within class. Bold denotes the best method in each column.
Subtype source 𝛼 = 1.00 𝛼 = 0.75 𝛼 = 0.50 𝛼 = 0.25 𝛼 = 0.00 Random
Racer-kernel gain
Same-role purity
Brier ↓
ECE ↓
+0.0298 +0.0390 +0.0426 +0.0401 +0.0425 −0.0095
0.6013 0.7609 0.7888 0.7819 0.7853 0.3338
0.2017 0.1850 0.1800 0.1838 0.1787 0.2247
0.0234 0.0207 0.0184 0.0259 0.0142 0.0122
Table 10: Mechanism diagnostics for Racer-kernel in the representation-geometry ablation. Gain is AURSAC improvement over the classifier-confidence router. Same-role purity is the fraction of the top-𝑘 same-true-role context neighbors that share the query’s hidden subtype under the geometry used by Racer-kernel. With three hidden subtypes, chance purity is 1/3. 35
Racer-kernel remains the strongest method across all structured classifier/auxiliary subtype geometries. Its gain over the classifier increases from +0.0298 AURSAC at 𝛼 = 1.0 to roughly +0.039–+0.043 for 𝛼 < 1.0. The same-role purity diagnostic explains why this is not a representation-mismatch failure case: auxiliary-defined subtypes are still highly recoverable by the default Racer-kernel geometry, with purity around 0.76–0.79 for 𝛼 < 1.0. In contrast, random within-class subtypes drive purity to chance and remove the routing gain, with Racer-kernel falling 0.0095 AURSAC below the classifier. These results support the intended mechanism: role-relative local competence estimation helps when expert-relevant within-class structure is recoverable from context geometry, and it does not spuriously improve when that structure is absent.
D.3
Context-posterior stress-test protocol
The context-posterior stress test reuses the PathMNIST label space and synthetic expert simulator, but adds a latent case-mix domain 𝑆 ∈ {1, . . . , 𝐿}. For each domain 𝑠, we sample a domainspecific class prior 𝛼 𝑠 ∈ Δ𝐾−1 and define a global prior 𝛼 0 . Context labels are drawn from 𝛼 𝑠 , while query labels are drawn from 𝛼𝜆 (𝑠) B (1 − 𝜆ctx )𝛼0 + 𝜆ctx 𝛼 𝑠 .
(16)
Thus 𝜆ctx = 0 satisfies the intended context-posterior invariance regime, while 𝜆ctx = 1 makes the context maximally informative about query case mix. Expert competence generation is held fixed across the sweep, so the experiment isolates context-posterior shift rather than changing the expert-reliability task. Racer-C uses the posterior adapter from Section B; all other components are matched to Racer-kernel. 𝜆
Kernel AURSAC
Racer-C AURSAC
Kernel Brier
Racer-C Brier
Δ NLL
0.00 0.25 0.50 1.00
0.8149 ± 0.0034 0.8173 ± 0.0007 0.8187 ± 0.0058 0.8211 ± 0.0071
0.8157 ± 0.0043 0.8160 ± 0.0033 0.8229 ± 0.0039 0.8342 ± 0.0040
0.2254 ± 0.0014 0.2252 ± 0.0003 0.2239 ± 0.0005 0.2222 ± 0.0018
0.2253 ± 0.0015 0.2251 ± 0.0007 0.2244 ± 0.0012 0.2213 ± 0.0008
0.0050 ± 0.0056 −0.0032 ± 0.0170 −0.0212 ± 0.0058 −0.1154 ± 0.0018
Table 11: Context-posterior stress test. Δ label NLL is NLL(𝜋𝜌 ) − NLL(𝑝 𝜔 ), so negative values indicate that the context-conditioned posterior improves held-out label-posterior fit. At small violation strengths Racer-C is tied with Racer-kernel; when the context strongly carries case-mix information, Racer-C improves both label fit and routing. The results support the modelling tradeoff. Under the intended regime (𝜆ctx = 0), the posterior adapter provides essentially no routing benefit. As the violation strengthens, the adapter becomes useful: at 𝜆ctx = 1, Racer-C improves AURSAC by 0.0131 and reduces label NLL by 0.1154. We therefore keep Racer-kernel as the default estimator and use Racer-C only as an extension for deployments where the context is allowed to reveal stable case-mix information.
D.4
CIFAR-100 class-scale protocol
The 𝐾 = 100 setting uses all CIFAR-100 fine labels. The 𝐾 = 20 setting uses a seeded superclassbalanced subset: CIFAR-100 has twenty superclasses, and we select exactly one fine class from each superclass using seed 0. Thus 𝐾 = 20 is a balanced low-cardinality stress test rather than
36
the first twenty CIFAR classes, while 𝐾 = 100 retains the full fine-label and superclass-confusion structure. Setting
Fine classes
Train
Val.
Test
𝐾 = 20 𝐾 = 100
20 100
9,000 45,000
1,000 5,000
2,000 10,000
Table 12: Synthetic CIFAR-100 label-space splits. The canonical training split is stratified into train/validation subsets.
Superclass
Selected fine class
Superclass
Selected fine class
aquatic_mammals flowers fruit_and_vegetables household_furniture large_carnivores large_natural_outdoor medium_mammals people small_mammals vehicles_1
beaver poppy mushroom chair wolf forest raccoon boy mouse train
fish food_containers household_electrical insects large_manmade_outdoor large_omnivores_herbivores noninsect_invertebrates reptiles trees vehicles_2
shark bottle clock beetle bridge elephant crab lizard oak_tree streetcar
Table 13: The seeded superclass-balanced 𝐾 = 20 subset. Exactly one fine class is selected from each CIFAR-100 superclass. For every image, we construct an unobserved subtype. We extract frozen ResNet-18 image features, cluster the training images within each true class into three subtypes, and assign validation and test images to the nearest within-class subtype centroid. Expert correctness then depends on the pair (𝑌, subtype). When an expert is incorrect, the wrong label is sampled from a structured confusion prior that places most of its mass on fine classes from the same CIFAR-100 superclass when such alternatives are available. We use two expert-strength profiles. The weak profile is the original setting with base expert accuracy 0.65, class-skill standard deviation 0.55, global-skill standard deviation 0.20, and three subtypes. When within-class specialization is active, its good/mid/bad subtype accuracies are approximately 0.90, the sampled classwise baseline around 0.65, and 0.40. The strong profile uses a cleaner locally complementary structure: good, mid, and bad subtypes have accuracies 0.98, 0.70, and 0.30, respectively. Two binary factors define the stress-test grid. The subtype factor 𝜌 ∈ {0, 1} controls whether competence is classwise or instance-dependent: 𝜌 = 0 makes all subtypes within a class share the same skill, while 𝜌 = 1 activates full hidden-subtype specialization. The permutation factor 𝜆id ∈ {0, 1} controls whether the nominal OOD group receives an expert-only permutation of its classwise competence profile. Because class effects and subtype assignments are exchangeable across classes, this full permutation preserves the population law. It changes a realized expert’s mapping to fixed roles, but does not create a population shift. The synthetic “OOD” results therefore concern a separately sampled unseen expert group, not verified out-of-distribution transfer. Each run uses 32 seen-training experts, 8 unseen-ID validation experts, 8 unseen-OOD validation experts, 32 seen-evaluation experts, 8 unseen-ID test experts, and 8 unseen-OOD test experts. Early stopping uses the unseen-ID validation experts. For each 𝐾 and expert-strength profile, the grid contains seven methods, four (𝜌, 𝜆id ) settings, and three seeds, for 84 runs; all four grids completed 84/84 runs. 37
E
Extended Literature Review
E.1
Learning with rejection and learning to defer
Learning to defer is closely related to selective classification and learning with a rejection option. Classical rejection methods abstain when the model is uncertain, often comparing model confidence to a fixed cost of rejection. L2D replaces the fixed rejection cost with a human or external expert whose correctness depends on the input. This changes the Bayes rule from “reject when the model is uncertain” to “defer when the expert is more likely to be correct than the model.” Early L2D formulations introduced this human–AI risk objective, while later work developed consistent surrogates and practical neural training objectives [9, 14]. The distinction matters for our setting. A selective classifier only needs to estimate its own reliability. A deferral system must compare two agents: the model and the expert. For unseen experts, the expert side of this comparison must be inferred from context. This is why our method focuses on estimating expert competence rather than only learning a rejector.
E.2
Clinical collaboration and hierarchical deferral
The scoping review of Strong et al. [19] examines human–AI collaboration across healthcare tasks, highlighting the roles of workflow, interaction design, and appropriate reliance alongside predictive performance. A concrete application is the guided deferral system of Strong et al. [17], which classifies medical reports with large language models and supplies guidance to humans on deferred cases. Racer focuses on estimating the correctness of an available expert from context; it does not generate guidance for that expert. Strong et al. [21] study coherent deferral for hierarchical multi-label medical imaging, constraining prediction and delegation actions to respect the label taxonomy and handoff semantics. This concerns the compatibility of actions across related labels. The coherent relabelling studied here instead concerns invariance to class names in a multiclass decision. Extending Racer’s competence estimates to hierarchical deferral would additionally require an appropriate structured decision rule.
E.3
Calibration in fixed-expert and fixed multi-expert L2D
Several recent papers study calibration in L2D. Verma and Nalisnick [23] show that the standard augmented-softmax surrogate of Mozannar and Sontag [14] is decision-consistent but does not directly yield calibrated estimates of expert correctness. The relevant expert-correctness estimator is an odds transform of the deferral coordinate and can exceed 1. Their one-vs-all alternative directly parameterizes expert correctness with a bounded sigmoid and improves calibration. Verma et al. [24] extend the calibration analysis to fixed multi-expert L2D. They show that symmetric softmax formulations can couple expert-correctness estimates across experts, so the estimate for one expert depends on the deferral coordinates of other experts. This is undesirable when one wants calibrated expert-specific probabilities. Cao et al. [2] further show that the problem is not the use of softmax itself, but the symmetric treatment of class probabilities and expert correctness. The correct probability object has product structure: class probabilities live in a simplex, while expert correctness is a separate scalar. Our work builds on this lesson but changes the problem setting. Prior calibrated L2D work asks whether a fixed-expert or fixed multi-expert system can estimate P(𝑀 = 𝑌 | 𝑋 = 𝑥) or P(𝑀 𝑗 = 𝑌 | 𝑋 = 𝑥). Racer asks whether a system can estimate P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ) for an unseen expert represented only by a small context. This introduces an additional statistical challenge: the
38
expert-correctness probability must be inferred from a set-valued context and should generalize across experts.
E.4
Population-adaptive deferral and unseen experts
Population-adaptive L2D studies deferral to experts not necessarily observed during training. L2D-Pop learns to condition a rejector on a small context set of expert behavior [22]. Queryindependent variants compress the context into a fixed expert representation; query-conditioned variants allow the current input to attend to relevant context examples. This makes them expressive enough to model instance-dependent expertise. However, high-capacity context encoders can exploit absolute class identity when they consume one-hot labels, class embeddings, or class-specific channels. This creates a failure mode under coherent relabellings or shifted expert populations. IFD addresses this by removing expert identity and absolute class-coordinate shortcuts, but it does so through classwise competence profiles [20]. Racer keeps the identity-free symmetry while restoring query dependence.
E.5
Decision consistency, probability estimation, and representation
A subtle point in comparing L2D-Pop, IFD, and Racer is that all three can be Bayes-aligned as defer/predict decision rules, yet they differ substantially as estimators of expert competence. This distinction is central to our empirical findings. The augmented softmax is a decision surrogate. For a fixed query 𝑥 and expert context 𝐶 𝑒 , let 𝑞★(𝑥, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑌 | 𝑋 = 𝑥, 𝐶𝐸 = 𝐶 𝑒 ).
𝜂 𝑦 (𝑥) denote the label posterior,
The conditional population form of the augmented softmax L2D loss is 𝒮(Π | 𝑥, 𝐶 𝑒 ) B −
Õ
𝜂 𝑦 (𝑥) log Π 𝑦 (𝑥, 𝐶 𝑒 ) − 𝑞★(𝑥, 𝐶 𝑒 ) log Π⊥ (𝑥, 𝐶 𝑒 ),
𝑦
where Π is a distribution on 𝒴 ∪{⊥}. By a Lagrange multiplier calculation, the unique minimizer over the simplex is Π★𝑦 (𝑥, 𝐶 𝑒 ) = Therefore
𝜂 𝑦 (𝑥) 1 + 𝑞★(𝑥, 𝐶 𝑒 )
Π★⊥ (𝑥, 𝐶 𝑒 ) ≥ max Π★𝑦 (𝑥, 𝐶 𝑒 ) 𝑦
Π★⊥ (𝑥, 𝐶 𝑒 ) =
,
⇐⇒
𝑞★(𝑥, 𝐶 𝑒 ) . 1 + 𝑞★(𝑥, 𝐶 𝑒 )
𝑞★(𝑥, 𝐶 𝑒 ) ≥ max 𝜂 𝑦 (𝑥). 𝑦
Equivalently, at the logit level, ★ ★ 𝑔⊥ (𝑥, 𝐶 𝑒 ) − 𝑔★ 𝑦 (𝑥) = log 𝑞 (𝑥, 𝐶 𝑒 ) − log 𝜂 𝑦 (𝑥),
up to an additive constant shared by all logits. The augmented softmax loss therefore learns the correct Bayes comparison. This is why one-stage L2D training is powerful: the classifier and rejector can co-adapt to optimize system performance rather than being trained as isolated modules.
39
Decision consistency is not the same as calibrated competence estimation. The previous calculation shows that the augmented softmax is a good surrogate for the decision boundary. It does not imply that the deferral coordinate is itself a calibrated estimate of expert correctness. At the optimum, 𝑞★(𝑥, 𝐶 𝑒 ) , Π★⊥ (𝑥, 𝐶 𝑒 ) = 1 + 𝑞★(𝑥, 𝐶 𝑒 ) so recovering 𝑞★ requires the odds transform 𝑞★(𝑥, 𝐶 𝑒 ) =
Π★⊥ (𝑥, 𝐶 𝑒 ) . 1 − Π★⊥ (𝑥, 𝐶 𝑒 )
Away from the ideal optimum, this transform is not constrained to lie in [0, 1]. Thus the augmented action probability is best viewed as a normalized decision variable, not as the expert-correctness probability itself. This distinction is harmless if the goal is only asymptotic defer/predict consistency, but it matters when the output is used for calibration, budget ranking, multi-expert comparison, or transfer to unseen experts. L2D-Pop learns a latent routing margin. L2D-Pop extends the augmented softmax surrogate to population-adaptive deferral by conditioning the deferral logit on a context representation: 𝑑Pop (𝑥, 𝐶 𝑒 ) B 𝑔⊥ (𝑥, 𝜓(𝐶 𝑒 ))
or
𝑑Pop (𝑥, 𝐶 𝑒 ) B 𝑔⊥ (𝑥, 𝜓(𝑥, 𝐶 𝑒 )).
At the population optimum, the resulting margin can recover the Bayes comparison. However, the method is not asked to estimate the components of that comparison separately. The context encoder and rejector learn the scalar quantity 𝑑Pop (𝑥, 𝐶 𝑒 ) − max 𝑔 𝑦 (𝑥) ≈ log 𝑞★(𝑥, 𝐶 𝑒 ) − log 𝑝 max (𝑥). 𝑦
This scalar can fit the training routing objective even when its internal explanation is not a transferable expert-competence law. For example, a high deferral margin may be caused by high expert competence, low model confidence, a frequent class, a site-specific artifact, context-set composition, or a fixed class-coordinate shortcut. In ID settings these explanations may be correlated and therefore empirically useful. Under expert shift they can come apart. This is especially important because standard population encoders often encode class labels and expert predictions in fixed coordinates: [𝜑(𝑥 𝑖𝐶 ), 𝑒(𝑦 𝑖𝐶 ), 𝑒(𝑚 𝑖𝐶 )]. Such encodings allow the rejector to learn identity-conditioned rules, e.g. “defer when coordinate 𝑗 is strong.” These rules are not ruled out by the augmented softmax objective, because they can reduce training risk whenever coordinate identity is predictive in the training population. But they do not express the invariant rule needed for transfer: “defer when the expert is strong on the relevant role.” IFD estimates competence explicitly, but classwise. IFD makes the opposite tradeoff. From context counts 𝑛 𝑒 ,𝑦 =
Õ
1{𝑦 𝑖𝐶 = 𝑦},
𝑡 𝑒 ,𝑦 =
𝑖
Õ
1{𝑦 𝑖𝐶 = 𝑦, 𝑚 𝑖𝐶 = 𝑦 𝑖𝐶 },
𝑖
it constructs a Bayesian profile 𝜃𝑒 ,𝑦 | 𝐶 𝑒 ∼ Beta(𝛼 𝑦 + 𝑡 𝑒 ,𝑦 , 𝛽 𝑦 + 𝑛 𝑒 ,𝑦 − 𝑡 𝑒 ,𝑦 ) 40
with mean 𝜇𝑒 ,𝑦 and variance 𝜎𝑒2,𝑦 . In the notation of our paper, IFD uses the restricted approximation ΓIFD (𝑥, 𝑦, 𝐶 𝑒 ) = 𝜇𝑒 ,𝑦 . This explicit profile improves interpretability, data efficiency, and invariance because the rejector receives only role-indexed summaries such as competence at the model-top class or expert-best class. The limitation is the missing 𝑥 dependence: all examples with the same candidate class share the same competence estimate. Racer separates competence estimation from system-level routing. Racer keeps the onestage L2D philosophy but changes the expert interface. The competence channel estimates the posterior-predictive role law Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ) ≈ Γ★(𝑥, 𝑦, 𝐶 𝑒 ) = P(𝑀𝐸 = 𝑦 | 𝑋 = 𝑥, 𝑌 = 𝑦, 𝐶𝐸 = 𝐶 𝑒 ) using a proper expert-correctness loss. The model posterior and competence estimate are then combined as Õ b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) = 𝑝 𝜔 (𝑦 | 𝑥)Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 ). 𝑦
The final deferral decision compares b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ) with 𝑝 𝜔,max (𝑥), either directly through the plug-in threshold or through an invariant decision head. This design preserves the benefit of joint training. The classifier, competence estimator, and deferral head can still be optimized together for system performance. But the competence estimate is not identified only through the augmented action coordinate. It has its own probability target. Detaching both the loss weight and the deferral-head inputs blocks routing gradients to its dedicated parameters, while shared representations still change through joint training: − log Π 𝑦 (𝑥, 𝐶 𝑒 ) − sg[Γ𝜃 (𝑥, 𝑦, 𝐶 𝑒 )] log Π⊥ (𝑥, 𝐶 𝑒 ). This does not establish calibration of the trained system. The quantity reported and compared as expert correctness is the probability-scale, role-relative estimate b 𝑞 𝜔 (𝑥, 𝐶 𝑒 ). Why this matters for ID and OOD experts. For ID experts, query conditioning can reduce the approximation bias of classwise IFD when competence varies within a class. Direct competence supervision supplies an additional modelling constraint relative to a latent routing interface; a universal variance reduction is not established here. Under population shift, coherent invariance excludes absolute class-coordinate shortcuts, but does not by itself guarantee better routing. L2D-Pop can learn a correct ID margin through absolute class-coordinate or context-composition shortcuts. IFD removes those shortcuts but cannot adapt within class. Racer combines the two desirable properties: it is role-aligned like IFD and query-conditioned like L2D-Pop. Therefore, when an unseen expert has the same structural competence pattern under different roles, or when competence depends on query-local subtypes, Racer has the appropriate inductive bias: learn competence relative to the candidate role, then compare it to model correctness. When the distinction disappears. In the infinite-data, realizable, perfectly optimized setting, L2D-Pop’s augmented softmax margin can recover the Bayes defer/predict comparison. If one only cares about the binary decision and the test experts are drawn from the same distribution, this may be enough. The distinction becomes important when we ask for more: probabilityscale expert-correctness estimates, budget-swept rankings, multi-expert comparability, and robust transfer to experts whose classwise or instancewise competence differs from the training population. 41
E.6
Identity-free and equivariant architectures
The coherent-relabelling requirement in L2D is an instance of a broader principle: the architecture should respect the symmetries of the task. If class names are arbitrary, then renaming classes coherently should not change a binary defer/predict decision. IFD enforces this by passing only role-indexed and symmetric summaries to the rejector. Racer extends this idea to instancedependent competence estimation. Candidate labels enter only through role-relative relations such as 1{𝑦 𝑖𝐶 = 𝑦}, ranks, posterior values at roles, and correctness indicators. No learned class embedding or untied class-specific channel is used. This perspective is related to permutationinvariant and equivariant neural architectures, including DeepSets and Set Transformers [8, 27]. However, our invariance is not merely permutation invariance over context examples; it is coherent relabelling invariance over class names. The context set may be reordered, and the class labels may be renamed, yet the predicted expert correctness and deferral decision should remain semantically unchanged.
E.7
Real-human expert benchmarks
A practical difficulty in evaluating L2D is the scarcity of datasets with persistent expert annotations. Many common classification datasets provide only a single label, aggregate votes, or sparse crowd annotations. Clean deferral evaluation requires either a fixed expert with labels across examples or enough repeated annotator behavior to construct expert contexts. For this reason, the completed experiments in this manuscript focus on radiologist and human–AI benchmarks with persistent annotations. These datasets test the deployment-relevant problem of estimating expert competence from finite context under realistic expert variability and annotation noise.
F
Descriptive comparisons and variability
The real-benchmark summaries are descriptive. CheXpert standard deviations range over 27 cells (three candidate humans, three held-out AI systems, and three seeds); VinDr-CXR standard deviations in the main summary range over five seeds. Cells sharing seeds, images, experts, or pathologies need not be independent. We omit the earlier cell-level inferential analysis pending a justified treatment of that dependence. No significance conclusion is drawn from these summaries.
G
Additional protocol notes
The completed result bundle includes full group-level and per-pathology breakdowns. The main text reports compact aggregate tables because the fine-grained real-expert tables are large.
G.1
Implementation and evaluation protocol
Training objectives and inference scores. Let 𝐿aug (𝑤) B − log Π 𝑦 − 𝑤 log Π⊥ , and let 𝐿CE denote class cross-entropy. The main implementations use 𝜆comp = 1; chest-radiography Racer-C additionally uses 𝜆post = 1 and image-only CE weight 0.1.
42
Method
Training objective
Classifier-conf. IFD-score IFD-MLP L2D-Pop QI/QC Racer-KNN Racer-kernel / DeepSets Racer-C (CXR)
𝐿CE 𝐿CE 𝐿aug (LCB 𝑦 1{𝑦 = 𝑘best }) 𝐿aug (1{𝑚 = 𝑦}) 𝐿aug (sg[Γ 𝑦 ]) 𝐿aug (sg[Γ 𝑦 ]) + 𝐿comp
Score swept at inference −𝑝max
b 𝑞 − 𝑝 max
𝑑 − max 𝑦 𝑓 𝑦 𝑑 − max 𝑦 𝑓 𝑦 b 𝑞 − 𝑝 max b 𝑞 − 𝑝 max
Kernel objective with adapted logits, plus 𝐿post + 0.1𝐿CE,base
b 𝑞 − max 𝑦 𝜋 𝑦
Table 14: Implemented objectives and budget-ranking scores. The auxiliary RACER routing head is trained but its logit is not the primary RACER budget score. Multi-label losses are averaged over pathologies with valid targets. The image encoder and classifier are updated during training; context image features are evaluated without gradients, while query features and posterior summaries remain differentiable for competence training. Dedicated RACER competence parameters receive BCE gradients; routing gradients are stopped at both the competence loss weight and the deferral-head inputs. The separate context-posterior stress test also adds image-only auxiliary CE when its configured weight is positive.
Protocol
Context source and separation
Selection and evaluation
Synthetic
Training contexts exclude the current minibatch. Evaluation samples labelled context within the validation/test split and removes those record indices from its query set. Disjoint image splits; each expert’s evaluation context and queries are disjoint subsets of images they annotated. Up to 64 context images are used, leaving at least one query. Disjoint record splits; training contexts contain 64 examples. Validation/test contexts occupy half their respective splits, with disjoint query remainders.
Separate training for each 𝐵; checkpoint selection uses unseen-ID validation AURSAC. Test queries are scored using their images and expert contexts.
VinDr-CXR
CheXpert
Leave-one-reader-out agreement target; disagreement between the other readers is masked. R8/R9/R10 define a co-annotation-cluster holdout. One candidate human plus AI experts; the held-out expert is an AI system. Calibration uses the expert selected by each method.
Table 15: Information and split protocol in the supplied implementations. Contexts contain historical image labels and expert responses. Query outcomes are used to score predictions, not as competence-head inputs at evaluation. Context sources, separation, and expert selection. In CheXpert’s multi-expert evaluation, IFD-score and RACER choose the available expert with highest b 𝑞 ; IFD-MLP and L2D-Pop choose by their learned routing scores. Classifier confidence uses reproducible random expert selection and ranks queries by classifier uncertainty. Expert availability is determined by the annotation mask. Brier and ECE therefore compare different selected expert–query pairs and jointly reflect expert selection and correctness estimation, rather than a fixed-pair calibration comparison. Reported best budgets and best accuracies summarize the test curves retrospectively; they are not validation-selected operating points. 43
G.2
VinDr-CXR protocol details
The most important VinDr-CXR caveat is that the official test labels are consensus-only and have no rad_id. All L2D evaluation therefore uses the multi-radiologist train annotations with image-level and expert-level splits. For radiologist 𝑒, the expert prediction is 𝑒’s own binary annotation and the target is the leave-one-radiologist-out consensus of the other two readers. If those two readers disagree, the pathology is masked. The all-rater majority label is used only for classifier-only training and image-level stratification. VinDr-CXR does not provide a hospital identifier in the annotation file, so the co-annotationcluster holdout is defined from co-annotation frequencies. The dominant co-annotation cluster is R8/R9/R10; we use this as an inferred OOD expert group. This is a radiologist/co-annotationcluster shift proxy, not a verified hospital label. Group
Classifier acc.
Expert acc.
Overall Seen Unseen-ID Unseen-OOD
0.9026 ± 0.0004 0.9700 ± 0.0003 0.9860 ± 0.0015 0.9942 ± 0.0010 0.9628 ± 0.0060 0.9862 ± 0.0041 0.8780 ± 0.0006 0.9628 ± 0.0004
Expert bal. acc.
Classifier AURSBAC
0.8365 ± 0.0032 0.8314 ± 0.0108 0.8323 ± 0.0739 0.8414 ± 0.0014
0.6878 ± 0.0078 0.7057 ± 0.0152 0.7083 ± 0.0466 0.6898 ± 0.0077
Table 16: VinDr-CXR expert-strength diagnostics for the classifier-confidence baseline. Values are mean ± standard deviation over the five final seeds. The deferred radiologist is substantially more accurate than the image-only classifier, especially on unseen-OOD cases, which explains why a simple classifier-confidence deferral rule is competitive. Expert balanced accuracy is reported because raw multi-label accuracy is inflated by frequent negative pathology labels.
Top triples Experts R8/R9/R10 R2/R3/R5 R1/R2/R3 R2/R3/R6 R12/R13/R16
Top pairs
Count
Experts
Count
5501 115 104 100 99
R8/R10 R8/R9 R9/R10 R2/R3 R1/R2
5931 5621 5529 1042 911
Table 17: VinDr-CXR co-annotation structure used to define the inferred R8/R9/R10 OOD expert cluster.
G.3
CheXpert multi-expert protocol details
CheXpert is evaluated as a multi-label binary-relevance task. Each run treats one human reader as the candidate human expert, includes two of the three external AI systems as seen experts during training and validation, and holds out the remaining AI system until test time. The primary candidate-reader set is {bc1, bc5, bc7}; bc2 and bc3 are retained for consensus construction. For a deferred human reader, the target is formed from the remaining human readers, so the deferred reader’s own annotation is not used as its target. The same reference target is used for every candidate AI system in that run: a positive label requires at least three positive votes among the four remaining human readers, with all four votes available. The full grid has three candidate human experts, three held-out AI choices, and three seeds, for 27 matched cells per method. All adaptive methods train with 𝐵 = 64; the multi-expert 44
evaluator uses half of each validation/test image split as context (67/66 images for the saved 134/133-image splits), with the remainder as queries, rather than enforcing the training context size. The three external AI systems are CARZero, KAD, and CheXZero. They were chosen because they are pretrained chest-radiograph systems, rather than small models trained only inside our experiment, and they represent strong contemporary CXR recognition paradigms: zero-shot or language-aligned CXR prediction and knowledge-guided chest-radiograph diagnosis. This makes the AI deferral targets plausible high-performing clinical-imaging experts. In each run, two of these AI systems are available during training and validation, while the third is held out until test time; rotating the held-out system tests whether a router trained with human and AI context can generalize to an unseen pretrained CXR model.
G.4
Extended CheXpert metrics
H
Ablations
H.1
Ablation: Role-relative DeepSets competence estimator
To isolate the contribution of explicit same-role kernel pooling, we implemented RACERDeepSets, a permutation-invariant neural competence estimator trained with the same BCE competence target as RACER-kernel. RACER-DeepSets receives only role-admissible query– context features: candidate-role equality indicators, context correctness, query–context similarity, posterior values and ranks, and symmetric context summaries. It uses no class embeddings, class-specific heads, or absolute class-coordinate channels, preserving the coherent-relabelling constraint used by Racer. Table 19 reports AURSAC gain over classifier confidence for all methods and context sizes. RACER-DeepSets is the strongest non-kernel neural ablation, but RACER-kernel is consistently better at moderate and large context sizes. In particular, at 𝐵 = 1000, RACER-kernel reaches +0.0273 overall AURSAC gain, compared with +0.0194 for RACER-DeepSets. Table 20 reports calibration metrics. RACER-kernel also gives lower Brier score at all nontrivial context sizes and remains substantially better at large context sizes, reaching 0.2029 at 𝐵 = 1000, compared with 0.2197 for RACER-DeepSets. These results suggest that the main conceptual contribution is the role-relative probability-scale competence interface, while explicit same-role kernel pooling provides a useful finite-context smoothing bias. Method IFD-score IFD-MLP L2D-Pop (QI) L2D-Pop (QC) Racer-KNN Racer-kernel RACER-DeepSets
𝐵=9
𝐵 = 25
𝐵 = 50
𝐵 = 100
𝐵 = 200
𝐵 = 500
𝐵 = 1000
−0.0084 ± 0.0023 −0.0119 ± 0.0072 +0.0000 ± 0.0081 −0.0055 ± 0.0078 −0.0012 ± 0.0047 +0.0025 ± 0.0045 +0.0044 ± 0.0052 −0.0143 ± 0.0056 −0.0113 ± 0.0036 −0.0045 ± 0.0037 −0.0068 ± 0.0020 −0.0043 ± 0.0096 −0.0020 ± 0.0049 +0.0003 ± 0.0062 −0.0099 ± 0.0045 −0.0191 ± 0.0116 −0.0100 ± 0.0023 −0.0192 ± 0.0162 −0.0112 ± 0.0117 −0.0049 ± 0.0042 −0.0061 ± 0.0040 −0.0060 ± 0.0013 −0.0146 ± 0.0088 −0.0076 ± 0.0056 −0.0061 ± 0.0023 −0.0050 ± 0.0013 −0.0042 ± 0.0038 −0.0092 ± 0.0070 −0.0086 ± 0.0034 −0.0128 ± 0.0081 −0.0049 ± 0.0083 −0.0039 ± 0.0065 −0.0008 ± 0.0102 +0.0058 ± 0.0124 +0.0147 ± 0.0064 −0.0080 ± 0.0034 −0.0006 ± 0.0036 +0.0098 ± 0.0144 +0.0129 ± 0.0025 +0.0186 ± 0.0037 +0.0251 ± 0.0068 +0.0273 ± 0.0051 −0.0057 ± 0.0058 −0.0041 ± 0.0072 +0.0074 ± 0.0082 +0.0019 ± 0.0091 +0.0156 ± 0.0085 +0.0204 ± 0.0017 +0.0194 ± 0.0029
Table 19: PathMNIST context-scaling AURSAC gain over the classifier-confidence router. Entries are mean ± standard deviation over three seeds, paired by seed against classifier confidence at the same context size. The classifier baseline is omitted because its gain is zero by definition. Bold denotes the best non-classifier method at each context size.
45
Method
Metric
𝐵=9
𝐵 = 25
𝐵 = 50
𝐵 = 100
𝐵 = 200
𝐵 = 500
𝐵 = 1000
Brier ↓ 0.2560 ± 0.0016 0.2519 ± 0.0013 0.2416 ± 0.0004 0.2339 ± 0.0010 0.2284 ± 0.0004 0.2237 ± 0.0007 0.2224 ± 0.0006 IFD-score ECE ↓ 0.1054 ± 0.0046 0.1358 ± 0.0027 0.1064 ± 0.0003 0.0807 ± 0.0059 0.0567 ± 0.0015 0.0328 ± 0.0003 0.0225 ± 0.0018 Brier ↓ 0.2563 ± 0.0016 0.2516 ± 0.0019 0.2417 ± 0.0007 0.2339 ± 0.0012 0.2283 ± 0.0005 0.2236 ± 0.0008 0.2222 ± 0.0004 IFD-MLP ECE ↓ 0.1055 ± 0.0047 0.1351 ± 0.0040 0.1067 ± 0.0009 0.0808 ± 0.0054 0.0567 ± 0.0029 0.0320 ± 0.0007 0.0214 ± 0.0026 Brier ↓ 0.4302 ± 0.0062 0.3013 ± 0.0035 0.2572 ± 0.0002 0.2387 ± 0.0009 0.2306 ± 0.0005 0.2253 ± 0.0019 0.2221 ± 0.0013 Racer-KNN ECE ↓ 0.4252 ± 0.0068 0.2171 ± 0.0059 0.1441 ± 0.0023 0.0978 ± 0.0051 0.0732 ± 0.0015 0.0633 ± 0.0053 0.0572 ± 0.0025 Brier ↓ 0.2251 ± 0.0007 0.2233 ± 0.0013 0.2200 ± 0.0010 0.2168 ± 0.0006 0.2102 ± 0.0018 0.2054 ± 0.0010 0.2029 ± 0.0006 Racer-kernel ECE ↓ 0.0122 ± 0.0094 0.0131 ± 0.0035 0.0220 ± 0.0076 0.0365 ± 0.0027 0.0302 ± 0.0047 0.0249 ± 0.0061 0.0255 ± 0.0070 Brier ↓ 0.2249 ± 0.0007 0.2250 ± 0.0012 0.2219 ± 0.0017 0.2191 ± 0.0032 0.2188 ± 0.0031 0.2209 ± 0.0047 0.2197 ± 0.0022 RACER-DeepSets ECE ↓ 0.0107 ± 0.0010 0.0144 ± 0.0146 0.0363 ± 0.0098 0.0398 ± 0.0226 0.0445 ± 0.0162 0.0534 ± 0.0320 0.0466 ± 0.0176
Table 20: PathMNIST context-scaling expert-correctness calibration. Entries are mean ± standard deviation over three seeds. Brier score and 15-bin ECE are reported only for methods that output probability-scale expert-correctness estimates. Bold denotes the best value at each context size for the corresponding metric.
46