ConceptioArchivearXiv CS
arXiv CSopen access

Uncertainty quantification for trustworthy deep learning: Methods and measures

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Uncertainty quantification for trustworthy deep learning: Methods and measures H. Martin Gillisa,∗, Thomas Trappenberga a Faculty of Computer Science, Dalhousie University, Halifax, Nova Scotia, Canada

Abstract

arXiv:2607.28248v1 [stat.ML] 30 Jul 2026

The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures. Keywords: uncertainty quantification; epistemic uncertainty; ensemble diversity; calibration; out-of-distribution detection; large language models

Abbreviations AU AUPR AUROC BALD BE BMA BNN BvSB CKA CLUE CreDE DDU DUE DUQ ECE ELBO EPCE EPJS EPKL EU GP HMC JS KFAC KL LLE LLL LLM

Aleatoric Uncertainty Area Under the Precision-Recall Curve Area Under the ROC Curve Bayesian Active Learning by Disagreement BatchEnsemble Bayesian Model Averaging Bayesian Neural Network Best-versus-Second-Best Centered Kernel Alignment Counterfactual Latent Uncertainty Explanation Credal Deep Ensemble Deep Deterministic Uncertainty Deterministic Uncertainty Estimation Deterministic Uncertainty Quantification Expected Calibration Error Evidence Lower Bound Objective Expected Pairwise Cross-Entropy Expected Pairwise Jensen-Shannon Expected Pairwise Kullback-Leibler Epistemic Uncertainty Gaussian Process Hamiltonian Monte Carlo Jensen-Shannon Kronecker-factored approximate curvature Kullback-Leibler Last-Layer Ensemble Last-Layer Laplace Large Language Model

∗ Corresponding author.

E-mail address: [email protected] (H. Martin Gillis).

MAP Maximum a Posteriori MCD Monte Carlo Dropout MCMC Markov Chain Monte Carlo MHML Multi-Head Multi-Loss MI Mutual Information MIMO Multi-Input Multi-Output MOD Maximized Overall Diversity MSP Maximum Softmax Probability MultiSWAG Multiple Stochastic Weight Averaging Gaussian NLL Negative Log-Likelihood NNGP Neural Network Gaussian Process ODIN Out-of-Distribution Detector OOD Out-of-Distribution RBF Radial Basis Function ROC Receiver Operating Characteristic SE Snapshot Ensemble SGD Stochastic Gradient Descent SGLD Stochastic Gradient Langevin Dynamics SNGP Spectral Normalized Gaussian Process SWAG Stochastic Weight Averaging Gaussian TU Total Uncertainty UQ Uncertainty Quantification VBLL Variational Bayesian Last Layer VGE Variance-Gated Ensemble VGMU Variance-Gated Margin Uncertainty VGN Variance-Gated Normalization VI Variational Inference

1. Introduction The deployment of deep neural networks in safety-critical domains such as autonomous driving, medical diagnostics, environmental monitoring, and scientific modeling is now

widespread. However, conventional architectures lack reliable uncertainty estimates accompanying their predictions. A model that is 95% confident in an incorrect class provides false assurance and may go unquestioned, whereas a model that reports high uncertainty allows a human to intervene before a decision is made. Consequently, a substantial body of research addresses Uncertainty Quantification (UQ) in deep learning, seeking to enable deterministic models with calibrated measures of predictive confidence.

What this survey provides. Rather than attempt broader coverage than the works above, we go deeper on a coherent area of the field and on the measurement question those surveys treat only briefly: 1. Method depth where the field has moved fastest: extended treatment of efficient ensemble approximations such as BatchEnsemble (BE), Snapshot Ensemble (SE), Stochastic Weight Averaging Gaussian (SWAG), Multiple Stochastic Weight Averaging Gaussian (MultiSWAG), and Credal Deep Ensemble (CreDE), and last-layer/singlepass approaches including Last-Layer Laplace (LLL), Spectral Normalized Gaussian Process (SNGP), LastLayer Ensemble (LLE), and deterministic uncertainty methods. 2. A unified treatment of uncertainty measures. We separate the method that produces a predictive ensemble from the measure that summarizes its uncertainty, and compare the standard mutual-information decomposition against pairwise divergence measures. 3. An explicit evaluation section (Section 11) that grounds the qualitative ratings in the comparison tables in a stated evidentiary basis.

1.1. Aleatoric and epistemic uncertainty It is useful to recall the distinction between aleatoric and epistemic uncertainty (Kendall and Gal, 2017; Hüllermeier and Waegeman, 2021). Aleatoric uncertainty arises from inherent noise in the data-generating process and is irreducible given the observation model. Epistemic uncertainty reflects ignorance about the true model parameters and, in principle, is reducible with additional data. From a Bayesian perspective, epistemic uncertainty is captured by the posterior distribution over model parameters p(w | D). Methods differ in how they approximate this posterior (or bypass it altogether) and this distinction organizes the sections that follow. We stress at the outset that the decomposition of predictive entropy into aleatoric/epistemic components is a modeling convention rather than a property intrinsic to the data. The same variability may be labeled aleatoric or epistemic depending on the chosen model class and feature set, a subtlety we return to in Section 14.

Scope and exclusions. The survey centers on ensemblebased and approximate Bayesian methods for classification, with regression treated where the decomposition differs. Evidential/prior-network methods (Section 8), conformal prediction (Section 9), and post-hoc calibration (Section 10) are summarized and related to the core families, and uncertainty in large language models is treated as an emerging adjacent area (Section 16). This treatment is deliberately uneven. The evidential/credal family is discussed more fully because it bears directly on the aleatoric/epistemic decomposition and the measure question (Section 14.6), whereas conformal prediction and post-hoc calibration are included as complementary tools rather than for exhaustive coverage.

1.2. Relation to prior surveys UQ in deep learning is already served by several broad surveys, and a new survey must justify itself against them. Gawlikowski et al. (2023) give a comprehensive overview organized primarily by model architecture: Bayesian neural networks, ensembles, single deterministic networks, and test-time augmentation, together with an extensive applications section; their taxonomy is broad but largely orthogonal to the sources of uncertainty a method targets. Abdar et al. (2021) review UQ techniques, applications, and challenges across both Bayesian and frequentist methods with a very large bibliography, but provide limited head-to-head comparison of method advantages and disadvantages. Hüllermeier and Waegeman (2021) focus conceptually on the aleatoric and epistemic distinction and its formalization rather than architectures. Mena et al. (2021) review UQ for classification primarily from a Bayesian standpoint, under-weighting frequentist and ensemble methods. Most recently, He et al. (2026) organizes UQ methods by the uncertainty sources they address, explicitly critiquing architecturefirst and Bayesian-first taxonomies. A recent general review of predictive uncertainty estimation with machine learning offers complementary breadth (Tyralis and Papacharalampous, 2024), as do a framework-oriented overview (Zhang et al., 2020), a tutorial pitched at engineering and health-prognostics audiences (Nemani et al., 2023), and domain-focused reviews such as Bayesian UQ for image segmentation (Valiuddin et al., 2025).

1.3. Organization Section 2 introduces the generative-mechanism and networkscope taxonomy that organizes the method families. Section 3– 7 review the five core method families. Section 8–Section 10 situate adjacent families. Section 11 consolidates evaluation methodology (metrics and distribution-shift benchmarks), establishing the evidentiary basis for the comparisons that follow. Section 12 consolidates all method families in comparison tables. Section 13 reviews ensemble diversity theory and metrics, and Section 14 treats uncertainty measures and their decompositions, including recent critiques and a sensitivity analysis. Section 15 treats the decision-time tasks of out-of-distribution detection and selective prediction. Section 16 orients the reader to the emerging area of uncertainty in large language models. Section 17 synthesizes the findings (Table 3–Table 8) and identifies open problems. Figure 1 summarizes this structure as a pipeline that deliberately separates the method producing a predictive ensemble from the measure that summarizes its uncertainty. A chronological overview of the surveyed methods is given in Table 1. 2

Uncertainty Quantification

(a)

Measures summarize uncertainty

Foundations & Evaluations diversity & metrics

Frontier & Open Directions emerging areas

Bayesian inference VI · MCMC · Laplace · deep GPs

Total predictive entropy · EPCE

Ensemble diversity BVD · disagreement · CKA

LLM uncertainty semantic entropy · verbalized

MC Dropout test-time dropout sampling

Aleatoric expected entropy

Evaluation ECE / NLL · AUROC · OpenOOD

Open directions hybrids · efficient epistemic

Deep ensembles gold standard

Epistemic mutual information · EPKL

Efficient ensembles BatchEnsemble · SWAG · CreDE

Efficient measures variance-gated · VGE/VGMU

Single-pass & adjacent SNGP · evidential · conformal

Decision-time tasks OOD · selection · active learning

Closed-form Sampling

Mechanism

(b)

Explicit

Methods produce an ensemble

Last-layer multi-head ensembles MIMO, MHML

Deep ensembles ( ), MultiSWAG, CreDE

high

Last-layer MCD/Laplace, SNGP (sampled)

BNN (VI), MCMC, MCD, SWAG

Evidential/ prior/posterior, DUQ/DDU/DUE

VBLL, Last-layer Laplace (MF), SNGP (MF)

Linearized/KFAC Laplace, NNGP

Last-layer

Full

Uncertainty Quality

low

Single

Network scope

Measure colour: information-theoretic vs. divergence / moment-based

Figure 1: Overview of uncertainty quantification (UQ) in deep learning. Panel (a). Taxonomy used throughout the survey. UQ methods (Section 3–Section 12) produce a predictive ensemble; uncertainty measures (Section 14) summarize that ensemble into total, aleatoric, epistemic, and efficient quantities, and support decision-time tasks such as OOD detection, selective prediction, and active learning (Section 15). Methods and measures are largely decoupled. Among methods that produce a predictive ensemble (explicit members or posterior samples), any such method can be paired with any sample-based measure; closed-form single-pass methods instead expose only their intrinsic measure unless sampled. The remaining branches cover foundations and evaluation-ensemble diversity (Section 13) and calibration/OOD metrics (Section 11); frontier topics, including LLM uncertainty (Section 16) and open directions (Section 17). Measure labels are colored by approach: information-theoretic (green) vs. divergence/moment-based (orange). Panel (b). Method families positioned by network scope (horizontal, increasing left to right: single deterministic network, shared backbone/last layer, full network) and mechanism, how the predictive ensemble is produced (vertical: explicit members, posterior sampling, closed-form predictive). Fill color encodes uncertainty quality, graded from the OOD and calibration ratings of Table 3 (light = near the maximum-softmax-probability baseline, dark = top-tier); deep ensembles are the gold standard (⋆). Quality and computational cost are decoupled: cost rises from left to right and toward the explicit row, so full-network, explicit-ensemble methods are the most expensive, yet high quality also appears at low cost in the bottom-left, where distance-aware single-pass methods (DDU/DUQ/DUE) reach top-tier OOD detection. A closed-form predictive spans all three scopes (single-pass deterministic networks; the last-layer mean-field methods, VBLL and mean-field Laplace/SNGP; and full-network linearized/KFAC Laplace with the infinite-width NNGP limit). A single deterministic network admits neither multiple trained members nor a posterior to sample.

2. Generative mechanisms and network scope

in these axes as it is introduced.

We organize the method families that follow along two axes, summarized in Table 2 and visualized in Figure 1. The first axis is the generative mechanism, how a method produces the predictive ensemble from which uncertainty is read: by explicit members (independently parameterized predictors, deterministic at inference), by posterior sampling (Monte Carlo draws from an approximate weight posterior at inference), or by a closed-form predictive (an analytic predictive distribution from a single deterministic pass). The second axis is the network scope over which that mechanism operates: the full network, a shared backbone with an independent last layer, or a single deterministic network. Computational cost generally rises toward explicit members and full-network scope, and the strongest uncertainty (between-mode diversity) arises only in the explicitmember, full-network cell. We flag the coordinate each family

3. Bayesian neural networks Mechanism: Posterior sampling (VI, MCMC) or closed-form predictive (linearized/Laplace, NNGP), full-network scope; last-layer variants in Section 7.1. 3.1. Framework and Bayesian model averaging The Bayesian treatment of neural networks, first developed by MacKay (1992) and Neal (1995), provides the most principled framework for UQ (see Jospin et al. (2022) for a practitioner-oriented tutorial and Arbel et al. (2023) for a critical review of the discussions). Rather than learning a single point estimate of the weight vector w, a Bayesian Neural Network (BNN) places a prior p(w) over the weights; a choice 3

Table 1: Timeline of uncertainty quantification methods, grouped by family and ordered by year within each group. References point to the originating work. Year

Method

Key contribution

Reference

Bayesian inference 1992 Bayesian framework 1995 BNN/HMC 2011 Variational inference 2011 SGLD 2015 Bayes by Backprop 2016 MC Dropout 2018 NNGP

Practical Bayesian backpropagation Bayesian learning for neural networks Practical VI for networks Scalable MCMC via gradient noise Weight-uncertainty VI Dropout as approximate inference Infinite-width networks as Gaussian processes

MacKay (1992) Neal (1995) Graves (2011) Welling and Teh (2011) Blundell et al. (2015) Gal and Ghahramani (2016) Lee et al. (2018a)

Ensembles 2015 TreeNets 2017 Deep Ensembles 2020 Anchored Ensembles 2024 CreDE

Shared-trunk branching ensemble Simple, scalable predictive UQ Prior-anchored approximate Bayesian ensembling Credal (interval-valued) deep ensembles

Lee et al. (2015) Lakshminarayanan et al. (2017) Pearce et al. (2020) Wang et al. (2024)

Efficient approximations 2017 Snapshot Ensembles 2019 SWAG 2019 Sub-Ensembles 2020 BatchEnsemble 2020 MultiSWAG 2021 Masksembles 2022 Layer Ensembles

Cyclical learning-rate checkpoints as an ensemble Gaussian from the SGD trajectory Trunk-shared partial ensembles Rank-one member perturbations Multi-basin Bayesian model averaging Fixed-mask MCD–ensemble interpolation Single-pass per-layer ensembling

Huang et al. (2017) Maddox et al. (2019) Valdenegro-Toro (2019) Wen et al. (2020) Wilson and Izmailov (2020) Durasov et al. (2021) Kushibar et al. (2022)

Last-layer and single-pass 2015 LLE (TreeNets) 2020 Last-layer Bayes 2020 DUQ 2021 Laplace redux 2021 MIMO 2021 DUE 2022 SNGP 2023 DDU 2023 Multi-Head Multi-Loss 2023 Epinet 2024 Repulsive LLE 2024 VBLL

Branching at final classification layer Fixes ReLU overconfidence far from data RBF distance to class centroids Scalable post-hoc Laplace approximation Multi-input multi-output subnetworks Deep-kernel distance-aware UQ Distance-aware GP output layer Feature-density deterministic UQ Per-head loss functions Auxiliary epistemic network Function-space repulsion for diversity Variational Bayesian last layers

Lee et al. (2015) Kristiadi et al. (2020) van Amersfoort et al. (2020) Daxberger et al. (2021) Havasi et al. (2021) van Amersfoort et al. (2021) Liu et al. (2022) Mukhoti et al. (2023) Galdran et al. (2023) Osband et al. (2023) Steger et al. (2024) Harrison et al. (2024)

Evidential, conformal, and calibration 2017 Temperature scaling 2018 Evidential deep learning 2018 Prior Networks 2019 Conformalized quantile regression 2020 Posterior Networks 2023 Conformal prediction 2025 Flexible evidential deep learning

Post-hoc calibration baseline Dirichlet evidence outputs Dirichlet distributional uncertainty Adaptive distribution-free intervals Density-aware Dirichlet, without OOD samples Finite-sample distribution-free coverage Flexible-Dirichlet parameterization

Guo et al. (2017) Sensoy et al. (2018) Malinin and Gales (2018) Romano et al. (2019) Charpentier et al. (2020) Angelopoulos and Bates (2023) Yoon and Kim (2025)

Table 2: Classification of UQ methods by how the predictive ensemble is generated (mechanism, rows) and network scope (columns, increasing left to right to match Figure 1). Bold marks methods that explore multiple loss-landscape modes (between-mode diversity); these consistently give the strongest uncertainty and all occupy the explicit-member, full-network. Computational cost generally rises toward the upper right. A single deterministic network admits neither multiple trained members nor a posterior to sample. Last-Layer Laplace and SNGP each appear under both the sampling and closed-form rows because the same last-layer Gaussian posterior may be consumed either way (by Monte Carlo sampling or by a closed-form mean-field approximation), with the closed-form mode the default in each case. The closed-form, full-network (linearized/KFAC Laplace and the NNGP limit) is listed for taxonomic completeness and treated only briefly (Section 3); it is not separately rated in Table 3. Mechanism

Shared backbone/last-layer

Full network

Explicit members (trained, deterministic at inference)

Multi-head/Last-Layer Ensemble (MIMO, MHML, LLCM, repulsive LLE, TreeNet, epinet), Layer Ensemble

Deep Ensemble, Anchored Ensemble, CreDE, MultiSWAG, Snapshot Ensemble, BatchEnsemble, Masksembles, Sub-Ensemble

Posterior sampling (Monte Carlo at inference)

Last-Layer MCD, Last-Layer Laplace (sampled), SNGP (sampled)

BNN (VI), BNN (MCMC), MCD, SWAG

Evidential/Prior/Posterior Network, DUQ, DDU, DUE

VBLL, Last-Layer Laplace (mean-field), SNGP (mean-field)

Linearized/KFAC Laplace, NNGP (infinite-width limit)

Closed-form predictive (analytic, single-pass)

Single deterministic network

4

that affects the posterior and remains under-examined (Fortuin et al., 2022; Fortuin, 2022) and seeks the posterior p(w | D) via Bayes’ rule. Predictions follow by marginalizing over this posterior through the posterior predictive distribution Z p(y | x, D) = p(y | x, w) p(w | D) dw. (1)

most practical Laplace variants are last-layer, we treat them in Section 7.1. We note here only that the linearized-Laplace view connects BNNs to Gaussian-process inference and underpins several single-pass methods. The neural-network/Gaussian Process (GP) connection is made explicit by the infinite-width Neural Network Gaussian Process (NNGP) limit (Lee et al., 2018a), with deep Gaussian processes a related hierarchical construction (Damianou and Lawrence, 2013). Applied across the full weight space, these linearized and KFAC Laplace approximations return a closed-form predictive in a single pass, as does the NNGP limit; together they occupy the closed-form, full-network cell of Table 2, the full-network analogue of the last-layer mean-field methods of Section 7.1. Hybrid schemes that jointly model structural and parametric uncertainty have also been explored (Hubin and Storvik, 2019).

w

This integral, Bayesian Model Averaging (BMA), captures epistemic uncertainty. A posterior with substantial parameter uncertainty results in a correspondingly wide predictive distribution. In the limit of infinite data, the posterior concentrates and epistemic uncertainty vanishes. The challenge is that the posterior is intractable for all but the simplest networks, motivating the extensive research on approximate inference. Beyond standard classification and regression, the same Bayesian treatment extends to structured prediction tasks such as deep survival analysis (Monod et al., 2025).

3.5. Practical limitations Despite their theoretical correctness, BNNs have consistently underperformed deep ensembles in empirical benchmarks (Gustafsson et al., 2020; Ovadia et al., 2019). Fort et al. (2020) offered an explanation. Commonly used variational approximations concentrate around a single posterior mode, capturing only local uncertainty. The true posterior of overparameterized networks is highly multi-modal, with distinct modes (basins) separated by high-loss barriers. By constraining the approximate posterior to be unimodal, variational BNNs systematically underestimate the epistemic uncertainty arising from multiple qualitatively different solutions. This “mode collapse” is a recurring theme in comparisons of Bayesian and ensemble methods.

3.2. Variational inference Graves (2011) and Blundell et al. (2015) developed practical Variational Inference (VI) schemes for BNNs; see Blei et al. (2017) for a general treatment and Hoffman (2013) for black-box and stochastic variants that scale VI to large datasets. Bayes by Backprop (Blundell et al., 2015) parameterizes a factorized (mean-field) Gaussian approximate posterior qθ (w) and   optimizes θ by minimizing DKL qθ (w) ∥ p(w | D) , equivalently maximizing the Evidence Lower Bound Objective (ELBO). At inference, weight samples are drawn from the approximate posterior and predictions are averaged to approximate the BMA integral. While principled and compatible with backpropagation, this doubles the number of learnable parameters and adds significant overhead in both training and inference.

4. Monte Carlo Dropout Mechanism: Posterior sampling, full-network scope.

3.3. Markov chain Monte Carlo An alternative family draws asymptotically exact posterior samples via Markov Chain Monte Carlo (MCMC). Neal (1995) applied Hamiltonian Monte Carlo (HMC) to neural networks; Welling and Teh (2011) proposed Stochastic Gradient Langevin Dynamics (SGLD) to scale MCMC to large datasets by injecting calibrated noise into Stochastic Gradient Descent (SGD) updates. MCMC is more faithful to the true posterior than VI but is computationally expensive, sensitive to sampling hyperparameters, and yields no closed-form posterior. Recent full-batch HMC studies (Izmailov et al., 2021) provide a goldstandard reference posterior and confirm that inexpensive approximations often differ substantially.

4.1. Dropout as approximate inference Dropout (Srivastava et al., 2014) was introduced as regularization that stochastically zeroes a fraction of hidden units during training, reducing feature co-adaptation. Gal and Ghahramani (2016) re-interpreted it: retaining dropout at test time and performing stochastic forward passes yields predictions whose empirical mean approximates the predictive mean and whose empirical variance estimates predictive uncertainty. This corresponds to VI with a specific approximate posterior, a product of Bernoulli distributions over binary masks multiplied by pointestimated weight matrices. The appeal of Monte Carlo Dropout (MCD) lies in its simplicity. It requires no change to training beyond standard dropout layers and costs only forward passes at inference.

3.4. Linearized and Laplace approximations Mechanism: Closed-form predictive, full-network scope; lastlayer variants in Section 7.1.

4.2. Practical limitations

A complementary route is the Laplace approximation, which fits a Gaussian at a Maximum a Posteriori (MAP) estimate using (an approximation to) the loss Hessian. Modern, scalable variants use Kronecker-factored or last-layer Hessians and can be applied post-hoc (Daxberger et al., 2021). Since the

The Bernoulli variational family constrains the posterior to a narrow region of weight space around the point estimate. As Fort et al. (2020) showed, this concentrates around a single mode of the loss landscape, failing to capture multi-modal 5

structure. Consequently MCD tends to underestimate epistemic uncertainty, particularly far from the training distribution. Additional limitations include sensitivity to the dropout rate (which simultaneously controls regularization strength and uncertainty magnitude) and high correlation among sub-networks. A line of work addresses these by making the dropout distribution learnable or otherwise more expressive (Boluki et al., 2020; Xie et al., 2022), by adopting alternative variational families such as α-divergences (Li and Gal, 2017), or by dropping connections rather than units (Mobiny et al., 2021). Ovadia et al. (2019) confirmed empirically that MCD generally underperforms deep ensembles under dataset shift, especially on Outof-Distribution (OOD) detection. MCD nonetheless remains a widely used baseline and a lightweight option in resourceconstrained settings.

Pearce et al. (2020) showed that regularizing each member toward an independent draw from the prior (anchored ensembling) makes the ensemble approximate posterior inference. 5.3. Practical limitations The principal limitation is linear cost. Training M independent networks requires M× the compute, memory, and storage of a single model. In practice M = 5 is common, with diminishing returns beyond that point (Fort et al., 2020). Although Wilson and Izmailov (2020) argued for a Bayesian interpretation, the method lacks the formal variational or sampling justification of traditional Bayesian approaches. For large models where even one training run is expensive, the factor-of-five overhead can be prohibitive, motivating the efficient approximations. Two cautions modulate the gold-standard narrative: (i) Abe et al. (2022) question whether the full ensemble is necessary, finding much of its benefit recoverable at lower cost, and (ii) Ashukha et al. (2020) show that common in-distribution uncertainty metrics overstate ensemble gains unless test-time calibration is controlled.

5. Deep Ensembles Mechanism: Explicit members, full-network scope. 5.1. Method and empirical performance Lakshminarayanan et al. (2017) proposed deep ensembles as a non-Bayesian approach to predictive UQ, building on a long tradition of ensemble learning (Dietterich, 2000). The method trains M networks independently from different random initializations, each minimizing the Negative Log-Likelihood (NLL), and averages their predictions. For regression, at a single input x each member m outputs a mean µm (x) and variance σ2m (x), enabling a decomposition of predictive uncertainty into an aleatoric component (the mean of the individual variances) and an epistemic component (the variance of the individual means). The spread is taken across members at that fixed input; the explicit law-of-total-variance form is given in Section 14.8 (Equation 11). Despite its simplicity, the method is remarkably effective. Multiple independent benchmarks (Gustafsson et al., 2020; Ovadia et al., 2019) establish deep ensembles as the de facto gold standard for UQ in deep learning, consistently outperforming approximate Bayesian methods on calibration, accuracy, and OOD detection, although their apparent advantage narrows once individual members are properly calibrated (Rahaman and Thiery, 2021).

6. Efficient ensemble approximations Mechanism: Explicit members (SWAG: posterior sampling), full-network scope.

6.1. BatchEnsemble Wen et al. (2020) introduced BatchEnsemble, in which each member m modifies a shared weight matrix W via elementwise multiplication with a member-specific rank-one matrix, Wm = W ⊙ (rm s⊤m ), where rm and sm are trainable memberspecific vectors and ⊙ denotes the element-wise (Hadamard) product. The parameter overhead is minimal (two vectors per member per layer), and ensemble inference can be vectorized within a single batch. However, Zamyatin et al. (2026) recently showed that BatchEnsemble members are near-identical in function space, behaving more like a single model than a true ensemble. The rank-one perturbations are small relative to the shared weights and are overwhelmed by the shared gradient signal during training, resulting in negligible functional diversity and poor OOD detection.

5.2. Loss-landscape perspective Fort et al. (2020) explained this success via the loss landscape. Networks from different seeds converge to distinct modes separated by high-loss barriers, each member sampling a different basin of attraction. This multi-modal exploration is precisely what unimodal approximate posteriors fail to achieve, accounting for the systematic performance differences. Wilson and Izmailov (2020) formalized the view that deep ensembles can be understood as approximate Bayesian model averaging over multiple posterior modes, with their success deriving from marginalizing across distinct solutions. Wild et al. (2023) subsequently established a more rigorous link between deep ensembles and (variational) Bayesian inference. Meanwhile,

6.2. Snapshot ensembles Huang et al. (2017) collect checkpoints along a single training trajectory using a cyclical learning-rate schedule. Each snapshot captures a different point in weight space, forming an ensemble without independent training runs. Since all snapshots arise from one trajectory, they typically remain within the same basin of attraction, limiting functional diversity relative to independently trained ensembles. 6

forward pass already encodes uncertainty. This is motivated by evidence that networks need not be fully stochastic to produce useful uncertainty (Sharma et al., 2023) and by analyses that decouple the representation from the uncertainty head (Brosse et al., 2020). This shared cost regime, however, is not a single mechanism: it spans all three of the generative strategies of Table 2. Within it, uncertainty may be produced by explicit members (multi-head and last-layer ensembles), by posterior sampling restricted to the last layer (last-layer dropout, or a sampled last-layer Laplace or SNGP posterior), or by a closedform predictive read analytically from a single pass (Variational Bayesian Last Layers (VBLLs), the mean-field Laplace and SNGP approximations, and the deterministic uncertainty methods). We flag the governing mechanism as each method is introduced.

6.3. Stochastic weight averaging–Gaussian Maddox et al. (2019) proposed SWAG, fitting a Gaussian to the SGD trajectory in the final training phase; at inference, weight samples approximate BMA. SWAG achieves good indistribution calibration with a single training run, but the posterior is confined to a single mode, limiting OOD detection. Wilson and Izmailov (2020) showed that combining SWAG posteriors from multiple independently trained networks, MultiSWAG, recovers much of the gap with deep ensembles, reinforcing that multi-modal exploration matters for high-quality uncertainty. 6.4. Credal deep ensembles Wang et al. (2024) introduced CreDEs, which extend standard deep ensembles by producing interval-valued outputs rather than point predictions. Using distributionally robust optimization-inspired training, CreDEs learn credal sets, convex sets of probability distributions that quantify both aleatoric and epistemic uncertainty. Empirically, CreDEs improve OOD detection over standard ensembles while maintaining comparable in-distribution performance, offering a promising route to representing epistemic uncertainty via imprecise probabilities (Walley, 1991). CreDE connects to the credal or evidential viewpoint of Section 8 and to the axiomatic critique of Section 14.6.

7.1. Bayesian and stochastic last-layers Mechanism: Posterior sampling or closed-form predictive, lastlayer scope. The Laplace approximation builds a Gaussian posterior centered at the MAP estimate from the loss Hessian. Applied to the last layer only, it is tractable and can be applied post-hoc to any trained network without retraining (Daxberger et al., 2021). Kristiadi et al. (2020) showed that even a minimal Bayesian treatment of the last layer measurably reduces the overconfidence of ReLU networks. The approach inherits two limitations: (i) it assumes a unimodal Gaussian posterior, and (ii) by restricting uncertainty to the last layer, it ignores epistemic uncertainty in the feature representation itself. A variational alternative, the VBLL (Harrison et al., 2024), trains the last-layer posterior deterministically and marginalizes it analytically, returning a closed-form predictive distribution in a single-pass rather than by sampling; it is thus the closed-form member of this last-layer family (Table 2). Restricting dropout to the final layer follows the same logic. Sampling several passes through a last-layer dropout mask yields a cheap uncertainty estimate that Brosse et al. (2020) found clearly improves on a point-estimate softmax, yet recovers only part of the benefit of full-network MCD or a deep ensemble, again because the shared representation is treated deterministically.

6.5. Further efficient approaches Several additional methods trade exactness for cost. Deep sub-ensembles share a common trunk and ensemble only a subset of layers (Valdenegro-Toro, 2019, 2023); Masksembles interpolate between MCD and deep ensembles using a fixed set of overlapping masks (Durasov et al., 2021); and layer ensembles treat per-layer weight distributions as a single-pass source of ensemble members (Oleksiienko and Iosifidis, 2023; Kushibar et al., 2022). As with the methods above, the recurring challenge is that parameter sharing limits functional diversity. Kirsch (2025) sharpens this concern for large models, showing that implicit weight sharing can collapse epistemic uncertainty across nominally independent members, an “ensemble of ensembles” that behaves like far fewer effective models. Complementary directions automate ensemble construction through joint neural-architecture and hyperparameter search (Egele et al., 2022; Herron et al., 2020), amortize members through a shared hypernetwork (Chauhan et al., 2024), or share a backbone with multiple lightweight heads.

7.2. Spectral-normalized neural Gaussian process Mechanism: Closed-form predictive (mean-field) or posterior sampling, last-layer scope. Liu et al. (2022) proposed SNGP, replacing the final dense layer with a random-feature approximation to a Gaussian process and enforcing spectral normalization throughout the backbone. Spectral normalization makes the learned representation approximately distance-preserving (inputs far from training data in input space remain far in feature space), so the GP output layer can assign appropriate uncertainty. SNGP matches the calibration and OOD-detection quality of a deep ensemble at near single-forward-pass cost, despite being a single deterministic model, making it one of the strongest single-model methods. The GP output layer returns a Gaussian over the logits rather than a point, so the class probabilities are the expected

7. Last-layer and single-pass approaches Mechanism: Spans all three strategies: explicit members, posterior sampling, and closed-form predictive, across sharedbackbone/last-layer and single-network scope; flagged per method below. A practically important class of methods seeks ensemblelevel uncertainty at (near) single-model cost, either by confining the source of uncertainty to the output layer over a shared backbone or by designing a single deterministic network whose 7

softmax under that Gaussian. This is precisely the BMA integral of Equation 1 specialized to the approximate last-layer posterior of SNGP, an intractable integral approximated either by Monte Carlo sampling the output logits or, as Liu et al. (2022) do by default, by a closed-form mean-field approximation that rescales the logits using the predictive variance and avoids sampling at inference. Its combination of distance-aware representations and a probabilistic output layer is a compelling inductive bias, although it does not capture between-mode epistemic uncertainty.

last-layer ensemble members with Monte Carlo dropout passes so that diversity is drawn from separately-instantiated heads and stochastic passes together (Lee et al., 2015; Brosse et al., 2020; Schweighofer et al., 2023a; Gillis et al., 2026a,b). This approach has a classical antecedent: the Bayesian committee machine (Tresp, 2000), where predictions are combined from separately-trained experts within a Bayesian framework, and last-layer ensembles (committee machines) can be read as its modern, shared-backbone realization.

7.3. Deterministic uncertainty methods

8. Evidential and prior-network approaches

Mechanism: Closed-form predictive, single deterministic network.

Mechanism: Closed-form predictive, single deterministic network.

A related line of methods avoids sampling entirely. Deterministic Uncertainty Quantification (DUQ) (van Amersfoort et al., 2020) uses Radial Basis Function (RBF) distances to class centroids with a gradient penalty enforcing sensitivity to input changes; Deep Deterministic Uncertainty (DDU) (Mukhoti et al., 2023) fits a Gaussian mixture in a (spectrally regularized) feature space and separates epistemic (feature-density) from aleatoric (softmax entropy) signals; and Deterministic Uncertainty Estimation (DUE) extends this line with deep-kernel learning over a distance-aware feature space (van Amersfoort et al., 2021). Orthonormal certificates (Tagasovska and LopezPaz, 2019) take a different single-model route, learning a set of diverse functions trained to vanish on the training data so that non-zero responses on new inputs signal epistemic, out-ofdistribution uncertainty. These deterministic uncertainty methods share the reliance on distance-aware representations of SNGP and similarly target OOD detection with a single-pass. They are attractive when latency is critical but, like SNGP, do not represent multi-modal epistemic uncertainty. Postels et al. (2022) further caution that the reliability of deterministic epistemic estimates is sensitive to architecture and training choices, influencing their single-pass appeal.

A distinct family of methods predicts the parameters of a distribution over predictions in a single forward pass, so uncertainty is produced directly rather than derived post hoc from the predictions. For classification, evidential deep learning (Sensoy et al., 2018) outputs the parameters of a Dirichlet over the class simplex, interpreting predictions as “evidence” and reading epistemic uncertainty from the total evidence mass. Prior networks (Malinin and Gales, 2018) similarly parameterize a Dirichlet but explicitly train for distributional uncertainty by exposing the model to OOD data during training, enabling a clean separation of in-distribution aleatoric uncertainty from distributional (epistemic) uncertainty; a later reverse-KullbackLeibler (KL) training scheme improved their uncertainty and adversarial robustness (Malinin and Gales, 2019). Posterior networks (Charpentier et al., 2020) replace the OOD-exposure requirement with a normalizing-flow density over the latent space, resulting in density-aware Dirichlet parameters without auxiliary OOD data. More recent variants include flexibleDirichlet parameterizations for greater expressiveness (Yoon and Kim, 2025), post-hoc Dirichlet meta-models that equip a pretrained network with evidential outputs (Shen et al., 2023), and losses that explicitly widen the in-/out-of-distribution representation (Nandy et al., 2021). These methods are appealing for their single-pass cost and explicit uncertainty semantics, and they connect directly to the credal/imprecise-probability viewpoint that runs as a thread through this survey: the credal deep ensembles of Section 6, the evidential and prior/posterior networks here, and the axiomatic credal-set foundations of Section 14.6. Their known weaknesses are a sensitivity to the choice of evidential regularizer (the loss term controlling how much the Dirichlet concentrates), difficulty calibrating the evidence scale, and a tendency to be overconfident on far-OOD inputs. This far-OOD overconfidence is a property of the base Dirichlet model, which has no built-in mechanism to recognize inputs far from the training data. Prior networks remove it by training against auxiliary OOD data, whereas posterior networks remove it with a latentspace density model that requires no OOD data at all. The effectiveness of prior networks therefore relies on how well the chosen auxiliary outliers represent the OOD encountered at test time, a dataset-dependent assumption that posterior networks avoid.

7.4. Last-layer ensembles and the branching-depth continuum Mechanism: Explicit members, last-layer scope. Lee et al. (2015) introduced TreeNets, where an ensemble shares a common trunk and branches at a chosen depth into independent sub-networks. Branching depth trades parameter efficiency against ensemble diversity. Last-layer ensembles are the extreme case of a TreeNet, where branching occurs at the final classification layer. Members share the entire backbone and differ only in output heads. Several architectures explore this space: Havasi et al. (2021) proposed Multi-Input Multi-Output (MIMO); Galdran et al. (2023) proposed Multi-Head MultiLoss (MHML) with per-head calibration; Steger et al. (2024) proposed repulsive last-layer ensembles that encourage functional diversity via a repulsive loss; and Osband et al. (2023) augment a base network with an auxiliary “epinet” that models epistemic uncertainty at low additional cost. Each navigates the same question: How much functional diversity is achievable when all members share a feature representation? One practical route combines both diversity sources at the head, pairing 8

9. Conformal prediction and distribution-free UQ

under shift (Ovadia et al., 2019). Interactions with ensembling are subtle. Calibrating members can mitigate accuracy– calibration trade-offs under shift (Kumar et al., 2022), whereas naively combining ensembles with data augmentation can harm calibration (Wen et al., 2021). Calibration methods and UQ methods are therefore complementary. The former corrects the scale of confidence, the latter aims to produce confidence that moves appropriately with epistemic state.

Mechanism: Post-hoc and model-agnostic; produces no predictive ensemble of its own and applies to any base predictor at any scope. Conformal prediction (Angelopoulos and Bates, 2023) provides finite-sample, distribution-free coverage guarantees. Given an exchangeable calibration set and a target level 1 − α, it returns prediction sets (classification) or intervals (regression) that contain the true label with probability at least 1 − α, regardless of the underlying model. Complementary to the methods surveyed here, it applies on top of any base predictor, including an ensemble or BNN, and converts a heuristic uncertainty score into a calibrated set. For regression, conformalized quantile regression (Romano et al., 2019) combines quantile regression (Koenker and Bassett, 1978; Tagasovska and Lopez-Paz, 2019) with conformal calibration to obtain adaptive intervals with coverage guarantees. The principal caveats are the exchangeability assumption and the fact that marginal coverage does not imply conditional coverage. Exchangeability requires that the calibration and test points be statistically interchangeable (their joint distribution is unchanged by reordering), which fails under distribution shift when test inputs are drawn differently from the calibration set. Marginal coverage means the 1 − α guarantee holds on average over all inputs, but not necessarily within any specific subgroup; conditional coverage, the stronger property of holding for every input or class, is not guaranteed.

11. Evaluation: Metrics and benchmarks Before comparing the method families head to head, we fix the basis on which they are judged. The quality of an uncertainty estimate is multifaceted. A method may be well calibrated yet uninformative, or accurate in-distribution yet unreliable under shift. Evaluating UQ therefore requires complementary metrics together with benchmarks that probe behavior across conditions. We review calibration, OOD, and selectiveprediction metrics, the shift and corruption suites used to stress them, and the reporting practices that make comparisons meaningful. Calibration metrics. ECE (Naeini et al., 2015; Guo et al., 2017) bins predictions by confidence and averages the difference between confidence and accuracy; it is widely used but sensitive to bin size. The Brier score (Brier, 1950) and negative log-likelihood are proper scoring rules (Gneiting and Raftery, 2007) that jointly reward calibration and sharpness. OOD and selective-prediction metrics. OOD detection is typically scored by Area Under the ROC Curve (AUROC) and Area Under the Precision-Recall Curve (AUPR) using an uncertainty score to separate in- from out-of-distribution inputs (the Maximum Softmax Probability (MSP) baseline (Hendrycks and Gimpel, 2017) is the standard reference point). Standardized suites such as OpenOOD (Yang et al., 2022) and Uncertainty Baselines (Nado et al., 2022) consolidate these protocols, datasets, and reference implementations, permitting reproducible comparison. Selective prediction is summarized by risk–coverage curves.

10. Post-hoc calibration Mechanism: Post-hoc and model-agnostic; produces no predictive ensemble of its own and rescales the confidence of any base predictor. Calibration (the agreement between predicted confidence and empirical accuracy) is used throughout this survey as an evaluation, so we briefly note the methods that target it directly. Temperature scaling (Guo et al., 2017) rescales logits by a single learned temperature on a held-out set and is a strong, nearfree baseline that preserves accuracy; Platt scaling and isotonic regression are classical alternatives (Platt, 1999; Zadrozny and Elkan, 2002). Training-time alternatives target calibration directly. Focal loss reduces overconfidence and has been analyzed as an implicit calibration regularizer (Mukhoti et al., 2020), while label smoothing (Zhang et al., 2021) and logit normalization (Wei et al., 2022) similarly temper confidence. Measuring calibration is itself delicate. Expected Calibration Error (ECE) is sensitive to binning and estimator choice (Vaicenavicius et al., 2019), and overconfidence is not always the dominant failure mode it is assumed to be (Wang et al., 2021). Calibration is also architecture-dependent. Modern networks are often better calibrated out of the box than the older models that motivated much of this literature (Minderer et al., 2021). Crucially, post-hoc calibration improves in-distribution calibration but does not by itself confer epistemic awareness or OOD robustness. Calibrated confidence can remain confidently wrong

Shift and corruption benchmarks. The most informative evaluations stress models under distribution shift. CIFAR-10C/CIFAR-100-C and ImageNet-C (Hendrycks and Dietterich, 2019) apply parameterized corruptions at increasing severity; the shift suite of Ovadia et al. (2019) tracks calibration and accuracy as a function of severity and is the canonical demonstration that ensembles degrade most gracefully. For OOD detection, common pairs include CIFAR-10 vs. SVHN/CIFAR-100 and ImageNet vs. iNaturalist/Places. Reporting practice. Following Ovadia et al. (2019), we recommend that UQ methods be evaluated across multiple metrics, since in-distribution calibration is a weak predictor of behavior under shift. Where possible, methods should be compared at matched compute, because much of the apparent advantage of ensembles is acquired with M× training cost. 9

12. Comparative summary of methods

where pm is the m-th member predictive distribution and the ensemble prediction p̄ is the centroid combiner of Wood et al. (2023), here the normalized geometric mean of the member distributions. The term p̊m is the centroid of member m prediction with respect to the training-set distribution (its expected prediction over random training draws D), the reference against which the bias and variance terms are measured. The diversity term is the average KL divergence from p̄ to each member pm . Since this term subtracts from the loss, greater diversity always benefits the ensemble for a given bias and variance. This generalizes the classical regression decomposition of Ueda and Nakano (1996), which partitions expected squared-error loss into bias, variance scaled by 1/M, and pairwise covariance scaled by (1 − 1/M) (Brown et al., 2005). Diversity is thus a formally defined component of ensemble loss. A subtlety worth emphasizing: the “diversity always helps” statement holds for fixed bias and variance; in practice, interventions that increase diversity often change bias and variance simultaneously, so the net effect on loss is not guaranteed.

Having surveyed the core and adjacent method families (Section 3–Section 10) and fixed the evaluation basis (Section 11), we now consolidate the methods here. Table 3 compares methods on computational cost and uncertainty quality, and Table 4 lists their advantages and disadvantages; both are read against the generative-mechanism taxonomy of Table 2 (Section 2). The OOD and calibration entries are qualitative three-level ratings (L/M/H), not reproduced metrics. Each is anchored to the published benchmarks summarized in Section 11 and defined in note (a) of Table 3 relative to two fixed reference points, the MSP baseline (Hendrycks and Gimpel, 2017) and deep ensembles on the standard OOD pairs and shift suite of Ovadia et al. (2019). The Evidence column names the principal source behind each rating. Where direct OOD-detection or shift benchmarking was unavailable, the rating is inferred from method structure and identified with a dagger (†). These are therefore relative, evidence-anchored placements rather than precise scores, and should be interpreted against Section 11. However, two main qualifications apply. First, the supporting results are drawn from studies that differ in architecture, dataset, and shift type, and since uncertainty rankings are themselves dataset- and shift-dependent, the ratings are indicative placements rather than directly comparable measurements. Second, ratings marked with a dagger (†) are conservative inferences from method structure rather than direct benchmarks and should not be over-interpreted; where a rating is primarily from the originating work of a method, the standardized suites of Section 11 (OpenOOD (Yang et al., 2022), Uncertainty Baselines (Nado et al., 2022)) offer an independent point of comparison.

13.2. Diversity metrics Kuncheva and Whitaker (2003) systematically studied ten diversity measures for classifier ensembles, categorized as pairwise and non-pairwise. The most widely used are summarized below. Pairwise disagreement. For classifiers hi , h j , the disagreement  is Di j = P hi (x) , h j (x) , averaged over pairs. Simple and interpretable, but it treats all disagreements equally regardless of whether they improve the ensemble. Cohen’s κ. Agreement corrected for chance; values near 1 indicate low diversity, near 0 indicate chance-level agreement (high diversity), and negative values indicate worse-thanchance agreement.

13. Ensemble diversity

Prediction (vote) entropy. A non-pairwise measure 13.1. Definition and the bias–variance–diversity decomposition

C

H=−

The comparison in Section 12 repeatedly traced uncertainty quality back to how much functional diversity a method achieves; we now make that notion precise. The theoretical motivation for diversity rests on the bias–variance–diversity decomposition. For an ensemble of M members each producing a predictive distribution pm over classes, Wood et al. (2023) proved an exact decomposition of the expected ensemble crossentropy loss

+

M 1 X

  ED DKL p̊m ∥ pm M m=1   M  1 X  − ED  DKL (p̄ ∥ pm ) M m=1

(3)

where nc is the number of members predicting class c. It is the hard-vote analogue of the predictive entropy of Equation 4; the entropy/mutual-information decomposition of Section 14 formalizes how its epistemic part is separated from the aleatoric. Variance of predictions. For classification, the variance of predicted probability vectors across members at a given input; for regression, Var[pi (x)], which is exactly the epistemic component of the variance decomposition (Equation 11).

M

  1 X y · log p̊m −ED y · log p̄ = − M m=1

nc 1 X nc log , M c=1 M

(bias) Prediction cosine similarity. The cosine similarity between members’ predicted probability (or logit) vectors at a given input; low similarity (equivalently, high cosine distance) indicates that members distribute mass differently and serves as a scaleinvariant proxy for functional diversity. Fort et al. (2020) used prediction cosine similarity together with functional distance to show that independently initialized networks occupy distinct loss-landscape modes.

(variance)

(diversity), (2) 10

Table 3: Comparative summary of UQ methods, grouped by family. Cost in units of single-model training/inference; M = ensemble size, S = stochastic passes, SN = spectral normalization, GP = Gaussian process. OOD/Calibration ratings (L/M/H) are defined in note (a) and summarize the benchmarks of Section 11; the Evidence column gives the principal source supporting each rating. Method

Training

Inference

Uncertainty

OODa

Calibrationa

Bayesian inference BNN (VI) BNN (MCMC) MC Dropout Last-Layer Laplace

2× high 1× post-hoc

S× S× S× 1× (+Hessian)

AU & EU AU & EU AU & EU EU

M M M L

M M M M

Ovadia et al. (2019) Izmailov et al. (2021) Ovadia et al. (2019) Kristiadi et al. (2020)

Ensembles and efficient approximations Deep Ensembleb M×

AU & EU

H

H

Anchored Ensemble BatchEnsemble Snapshot Ensemblec (Multi)SWAG Sub-Ensemble Masksembles Layer Ensemble

M× 1× 1× (M×) 1× 1×+ 1× 1×

M× 1× (batched) M× S× M× (partial) S× 1×

AU & EU EU EU EU EU EU EU

M–H L M M M M M

M–H L M M M M M

CreDEd

AU & EU

M+

M+

Ovadia et al. (2019); Gustafsson et al. (2020) Pearce et al. (2020)† Zamyatin et al. (2026) Huang et al. (2017)† Wilson and Izmailov (2020) Valdenegro-Toro (2019)† Durasov et al. (2021) Oleksiienko and Iosifidis (2023)† Wang et al. (2024)

Single-pass and last-layer SNGP DUQ/DDU/DUE TreeNets/LLE

1× (+SN) 1× (+regression) 1×

1× 1× 1× (batched)

AU & EU EU (feature-density) EU

M–H H L–M

M–H M–H M

Liu et al. (2022) Mukhoti et al. (2023) Lee et al. (2015); Havasi et al. (2021)†

Evidence

Distributional and post-hoc Evidential/Prior/Posterior Network Conformal wrapper

AU & EU (Dirichlet)

M–He

M

Charpentier et al. (2020)

+calibration set

negligible

set/interval coverage

n/af

guaranteedg

Post-hoc calibration

+calibration set

rescaled confidence

n/a

in-distribution onlyh

Angelopoulos and Bates (2023) Guo et al. (2017)

H = consistently top-tier, comparable to deep ensembles on standard OOD pairs (CIFAR-10 vs. SVHN/CIFAR-100, ImageNet vs. iNaturalist/Places) and the shift suite of Ovadia et al. (2019); M = clearly above the maximum softmax probability baseline (Hendrycks and Gimpel, 2017) but below deep ensembles; L = near or below that baseline. Intermediate marks (L–M, M–H, M+) denote placement between adjacent levels. b Explores multiple modes of the loss landscape. c Limited multi-modal exploration (single basin). d Improvement over the standard deep ensemble baseline; rating from the proposing work, with limited independent benchmarking to date. e Higher with OOD exposure (prior networks) or density modeling (posterior networks). f Conformal targets coverage, not an OOD score per se. g Marginal coverage guarantee under exchangeability. h Improves in-distribution calibration only; no epistemic awareness. † Rating inferred from method structure; limited direct OOD-detection or shift benchmarking in the cited work. a OOD/calibration quality relative to two reference points:

Centered kernel alignment. A representation-level measure, Centered Kernel Alignment (CKA), comparing internal features across members. Two networks with low CKA similarity have learned qualitatively different representations even when their training-set predictions agree, which is relevant for OOD detection, where representational diversity can yield disagreement on novel inputs.

sity because independently initialized networks reach different loss-landscape modes; Fort et al. (2020) quantified this via prediction cosine similarity and functional distance. MCD and SWAG produce members confined to a single basin with high mutual similarity. BatchEnsemble members, despite distinct parameterization, are near-identical in function space (Zamyatin et al., 2026). Snapshot ensembles achieve intermediate diversity. For last-layer approaches, diversity can only arise in the output mapping; with a shared backbone, independent random initialization can break symmetry among heads, but how strongly shared gradient pressure converges them remains open. Strategies for enhancing multi-head diversity Maximized Overall Diversity (MOD) (Jain et al., 2020), adversarial diversity (Ramé and Cord, 2021), repulsive function-space inference (Steger et al., 2024), and per-head weighted losses (Galdran et al., 2023)) remain an active area of research.

Vendi score. A recent information-theoretic diversity metric, defined as the exponential of the Shannon entropy of the eigenvalues of a member-similarity matrix, that quantifies effective ensemble size without reference labels (Friedman and Dieng, 2023). 13.3. Diversity across UQ methods Functional diversity varies sharply across methods and predicts uncertainty quality. Deep ensembles achieve high diver11

Table 4: Advantages and disadvantages of uncertainty quantification methods, grouped by family. Method Bayesian inference BNN (VI) BNN (MCMC) MC Dropout Last-Layer Laplace

Advantages

Disadvantages

Principled Bayesian framework; captures weight uncertainty; naturally regularizes; compatible with backpropagation. Asymptotically exact posterior samples; gold-standard reference posterior; more faithful than variational approximations. Simple; no architecture changes; reuses existing dropout layers; cheap training. Post-hoc; no retraining; tractable Hessian; closed-form posterior.

Doubles parameters; expensive; concentrates on a single mode; underperforms ensembles under shift. Very expensive; sensitive to sampling hyperparameters; no closed-form posterior; hard to scale. Restricted to a single mode; correlated sub-networks; sensitive to dropout rate; underestimates OOD uncertainty. Ignores feature uncertainty; over-confident without representation constraints; limited to the last layer.

Ensembles and efficient approximations Deep Ensemble Multi-modal posterior exploration; approximate BMA over modes; strong calibration and OOD detection; no special architecture. Anchored Ensemble Prior-anchored regularization makes the ensemble approximate Bayesian posterior inference. BatchEnsemble Near single-model cost; vectorized inference; minimal parameter overhead. Snapshot Ensemble Ensemble for the cost of one training run; simple cyclical learning-rate schedule. (Multi)SWAG Cheap posterior from the SGD trajectory; closed-form Gaussian; combinable across runs for multi-modal coverage. Sub-Ensemble Shares a trunk; cheaper than full ensembles; diversity tunable via branch depth. Masksembles Interpolates between MC Dropout and deep ensembles using fixed complementary masks; tunable diversity/cost. Layer Ensemble Single-pass uncertainty from per-layer weight distributions; low overhead. CreDE Interval-valued (credal) outputs; improved OOD detection over standard deep ensembles. Single-pass and last-layer SNGP Single forward pass; distance-aware; strong calibration; composable with ensembles. DUQ/DDU/DUE Single forward pass; distance/density-aware OOD detection; conceptually simple. LLE/LLCM Near single-model cost; ensemble diversity from output heads; composable with any backbone. Distributional and post-hoc Evidential/Prior/Posterior Single pass; explicit uncertainty semantics; clean Network aleatoric/epistemic separation (with OOD exposure or density). Conformal Model-agnostic; finite-sample, distribution-free coverage guarantees; wraps any base predictor. Post-hoc calibration Near-free; preserves accuracy; strong in-distribution calibration baseline.

Linear cost scaling in M; high memory footprint; diminishing returns beyond M = 5. Linear cost in M; exact posterior recovery only under restrictive conditions. Near-identical members; behaves like a single model; poor OOD detection. Limited functional diversity; stays within a single basin; inferior to independent ensembles. Standard SWAG confined to one mode; modified learning-rate schedule; MultiSWAG costs M× training. Diversity limited by the shared trunk; branch passes still scale with M. Single-basin like MC Dropout at low mask diversity; requires mask-count/overlap tuning. Limited functional diversity (within one trained model). Same cost as deep ensembles; relatively new with limited cross-domain benchmarking. Requires spectral normalization; quality depends on the number of random features; single-mode only. Single-mode; sensitive to feature regularization; gradient-penalty/Gaussian-mixture fitting overhead. Diversity limited by the shared representation and shared training signal (common batches/ordering); may collapse under shared gradients; achievable diversity is an open question. Regularizer-sensitive; evidence-scale calibration is hard; far-OOD overconfidence without exposure/density. Exchangeability assumption (violated under shift); marginal, not conditional, coverage. Improves in-distribution calibration only; no epistemic awareness; degrades under shift.

x, wm ) (noting that the aggregation rule matters, as averaging logits rather than probabilities shifts the resulting confidence and calibration (Tassi et al., 2022))

14. Uncertainty measures and decompositions The methods above produce sets of predictions (from ensemble members, stochastic passes, or posterior samples), but quantifying the uncertainty in those predictions is itself nontrivial. This section reviews the principal measures, emphasizing information-theoretic decompositions and divergencebased measures for classification; regression is treated in Section 14.8. Table 5 places these measures on a timeline for an overview.

H[p(y | x, D)] = −

C X

p(y = c | x) log p(y = c | x).

(4)

c=1

Predictive entropy is high whether the input is inherently ambiguous (aleatoric) or the model lacks knowledge (epistemic); used alone it cannot separate the two. This is a critical limitation when only the epistemic part is actionable. In active learning, selecting examples by total entropy directs the labeling effort on irreducibly noisy inputs that also score high in aleatoric uncertainty, whereas the goal is to query inputs the model could still learn from. The same conflation can make predictive entropy actively misleading as an OOD score (Kirsch

14.1. Total predictive entropy The most direct measure of total predictive uncertainty for classification is the entropy of the averaged predictive distribuP tion. With M members (or S passes), p(y | x, D) = M1 m p(y | 12

Table 5: Timeline of uncertainty measures for ensemble-based deep learning, grouped by measure category and ordered by year within each group. Year

Measure

Total uncertainty and decompositions 1965 Predictive entropy 2017 Variance decomposition 2018 Entropy/MI decomposition 2023 2026

EPKL/EPCE Variance-gated (VGMU, VGN)

Epistemic measures 2005 JS-divergence acquisition 2011 BALD 2023 MI critique 2023 Pairwise-distance estimators 2023 Axiomatic epistemic measure 2025 Integral imprecise metrics Aleatoric measures 2022 β-NLL Ensemble diversity 1996 Bias–variance (regression) 2003 Classifier diversity measures 2023 Bias–variance–diversity

Key idea

Reference

Information-theoretic roots; total predictive uncertainty Regression aleatoric/epistemic split Additive total/aleatoric/epistemic decomposition

Ash (1965) Lakshminarayanan et al. (2017) Depeweg et al. (2018); Smith and Gal (2018) Schweighofer et al. (2023a) Gillis et al. (2026b)

Pairwise divergence measures; epistemic (EPKL) and total (EPCE) Variance-gated total/aleatoric/epistemic decomposition; epistemic margin score JS divergence for active learning Mutual information MI underestimates epistemic uncertainty under finite ensembles Bound entropy via pairwise divergences Credal-set desiderata for epistemic uncertainty Credal-set imprecise-probability metrics

Melville et al. (2005) Houlsby et al. (2011) Wimmer et al. (2023) Berry and Meger (2023) Sale et al. (2023) Chau et al. (2025)

Training objective for the aleatoric variance head; corrects heteroscedastic NLL

Seitzer et al. (2022)

Ensemble error decomposition; diversity as the covariance term Ten pairwise/non-pairwise metrics Exact ensemble cross-entropy decomposition

Ueda and Nakano (1996) Kuncheva and Whitaker (2003) Wood et al. (2023)

and Mukhoti, 2021).

members disagree but each is confident in a different class, average entropy (AU) is low while mixture entropy (TU) is high, giving high MI; but when members produce diffuse but distinct distributions, both are high, so MI stays low despite genuine disagreement. The decomposition also requires access to individual member predictions. These issues motivated Schweighofer et al. (2023a) to propose pairwise divergence measures that bypass the standard entropy decomposition.

14.2. The entropy/mutual-information decomposition The information-theoretic decomposition (Ash, 1965), formalized by Depeweg et al. (2018) and popularized by Smith and Gal (2018), disentangles the two sources   H[p(y | x, D)] = Ew∼p(w|D) H[p(y | x, w)] | {z } | {z } TU

AU

+ MI[y; w | x, D] . | {z }

(5)

14.4. Pairwise divergence measures

EU

Schweighofer et al. (2023a) argued that the standard additive decomposition implicitly assumes the BMA distribution equals the true posterior predictive (an assumption that breaks under finite ensembles) and proposed pairwise measures that avoid it. For total predictive uncertainty they introduced the Expected Pairwise Cross-Entropy (EPCE)

The expected conditional entropy captures Aleatoric Uncertainty (AU); the Mutual Information (MI) between prediction and parameters captures Epistemic Uncertainty (EU); together they constitute the Total Uncertainty (TU). In the ensemble approximation M

MI[y; w | x, D] ≈ H[p(y | x, D)] −

M

1 X H[p(y | x, wm )], (6) M m=1

EPCE =

the difference between mixture entropy and average member entropy. This additive decomposition is standard for active learning (Houlsby et al., 2011; Gal et al., 2017), selective prediction (Geifman and El-Yaniv, 2017), and OOD detection (Smith and Gal, 2018). Decomposition-aware estimators tailored to ensembles refine the finite-M estimates (Liu et al., 2019), and supervised variants train directly for the aleatoric/epistemic split (Li et al., 2025).

M

 1 XX  H p(y | x, wi ), p(y | x, w j ) , 2 M i=1 j=1

(7)

P with H[p, q] = − c p(c) log q(c). For epistemic uncertainty they introduced the Expected Pairwise Kullback-Leibler (EPKL) divergence, M

EPKL =

M

  1 XX DKL p(y | x, wi ) ∥ p(y | x, w j ) , M 2 i=1 j=1

(8)

which decomposes as the sum of MI and the reverse MI and upper-bounds MI by Jensen’s inequality (Schweighofer et al., 2023a); the pairwise-KL form also appears in Malinin and Gales (2018). EPKL is zero when all members agree and grows with disagreement. Its weaknesses are that KL is asymmetric,

14.3. Critiques of the additive decomposition The decomposition has drawn recent criticism for its finiteensemble behavior. Wimmer et al. (2023) showed MI can underestimate epistemic uncertainty in multi-class settings. When 13

DKL [pi ∥p j ] , DKL [p j ∥pi ], and unbounded, causing instability when a member assigns near-zero probability to a class. A remedy replaces KL with the symmetric, bounded Jensen– Shannon divergence

Interpretation: Sensitivity and stability trade-off. The unboundedness of EPKL is precisely what makes it both most responsive to member disagreement and most prone to instability, whereas EPJS and MI sacrifice some responsiveness for numerical stability. The defensible, model-independent claim is structural. The symmetry and boundedness of EPJS remove a known failure mode of EPKL, while its pairwise form avoids the mixture behavior that limits MI. We do not claim that EPJS is the best OOD detector; which measure performs best in practice depends on the ensemble and the nature of the shift, and must be established empirically on standard suites (Section 11). We present this discussion and comparison as a synthesis of the properties of the measures rather than as a new empirical result.

DJS [pi ∥p j ] = 12 DKL [pi ∥m] + 12 DKL [p j ∥m], m = 21 (pi + p j ) (9) bounded in [0, 1] (using base-2 logarithms throughout this discussion), giving the Expected Pairwise Jensen-Shannon (EPJS) divergence M

EPJS =

M

  1 XX DJS p(y | x, wi ) ∥ p(y | x, w j ) . M 2 i=1 j=1

(10)

Pairwise-distance epistemic estimators (including JensenShannon (JS)-divergence variants) have appeared in the literature. Berry and Meger (2023) develop pairwise-distance estimators that bound entropy via pairwise divergences for regression ensembles and connect them to Bayesian Active Learning by Disagreement (BALD), building on a longer line of ensemble-disagreement estimators that quantify epistemic uncertainty through divergences between member predictive distributions, such as the expected pairwise KL divergence of Malinin and Gales (2018); Malinin et al. (2020). The JS divergence in particular has a long history as an uncertainty signal, e.g. for active learning (Melville et al., 2005), and continues to be used as an epistemic (variance) proxy and full-distribution disagreement score in recent work (Schirmer et al., 2023). We note, however, that information-theoretic decompositions of this kind have been shown to behave inconsistently as epistemic measures under posterior mismatch and finite ensembles (Wimmer et al., 2023).

14.6. Axiomatic perspectives Beyond empirical comparison, an axiomatic line asks what a proper epistemic measure should satisfy. Sale et al. (2023), drawing on proper scoring rules and the geometry of credal sets, argued that a valid epistemic measure should be zero when the posterior concentrates on a single model and increase monotonically with posterior spread, establishing formal desiderata that existing measures meet to varying degrees. This links the measure question to the credal/evidential methods of Section 6 and Section 8. Related formalizations include integral impreciseprobability metrics over credal sets (Chau et al., 2025), a rate– distortion view of uncertainty quantification (Apostolopoulou et al., 2024), and epistemic estimation via explicitly adversarial models (Schweighofer et al., 2023b). 14.7. Efficient variance-gated measures A practical drawback of the pairwise measures above is their O(M 2C) cost, which becomes prohibitive for large ensembles and many classes. To reduce this cost, VarianceGated Ensembles (VGEs) (Gillis et al., 2025, 2026b) gate member probabilities by a signal-to-noise factor derived from the ensemble mean and per-class predictive spread, yielding a total/aleatoric/epistemic decomposition at O(MC), a marginbased epistemic score, Variance-Gated Margin Uncertainty (VGMU) extends the Best-versus-Second-Best (BvSB) criterion (Joshi et al., 2009) at O(C), and a differentiable trainingtime calibration layer, Variance-Gated Normalization (VGN). Like the other efficient approximations surveyed here, variancegating trades the explicit information-theoretic semantics of the entropy decomposition for lower cost, and its behavior under distribution shift has not yet been benchmarked against MI and the pairwise measures on the standard suites of Section 11. It nonetheless illustrates the broader search for efficient, nonentropy-based epistemic measures motivated by the critiques of Wimmer et al. (2023).

14.5. Comparative stability of divergence-based measures The three epistemic measures above (MI, EPKL, EPJS) differ in structural properties that reflect directly on their numerical behavior (Table 6). Table 6 describes the symmetry and boundedness of just these three epistemic divergence measures. The broader description of all uncertainty measures and their classification by type and computational approach is consolidated separately in Table 7 and Table 8 (Section 14.9). EPKL is asymmetric and unbounded. When member pi assigns nearzero mass to a class for which p j assigns non-negligible mass, DKL [pi ∥p j ] diverges, disproportionately amplifying the aggregate. This makes EPKL the most sensitive of the three to distributional differences but also the most fragile, since a single tailclass asymmetry can dominate the score and compress the in/out-of-distribution margin that governs detection. This effect is exacerbated under shift, where members lose a shared reference and extrapolate independently. EPJS replaces KL with the symmetric Jensen–Shannon divergence, bounded in [0, 1], so any single class pair’s contribution is restricted and the nearzero-probability pathology is removed. MI sidesteps unboundedness altogether by operating on the mixture rather than on pairs, but this is known to underestimate epistemic uncertainty for diffuse-but-distinct members (Wimmer et al., 2023).

14.8. Variance-based measures for regression For regression the decomposition takes an algebraic form via the law of total variance. When each member outputs a mean 14

Table 6: Structural properties of the three epistemic measures compared in Section 14.5. Symmetry and boundedness are the properties that govern numerical stability; sensitivity trade-offs against them. Measure

Basis

Symmetric

Bounded

Principal limitation

MI

Mixture-based

n/a

Yes (≤ log2 C bits)

EPKL EPJS

Pairwise (KL) Pairwise (JS)

No Yes

No Yes ([0, 1] bit)

Underestimates EU for diffuse-but-distinct members (Wimmer et al., 2023) Unbounded; unstable near zero-probability classes Less sensitive than EPKL; O(M 2 C) cost

µm (x) and variance σ2m (x) (Lakshminarayanan et al., 2017) M M  1 X 2 1 X Var[y | x] = σm (x) + µm (x) − µ̄(x) 2 . | {z } M M TU | m=1 {z } | m=1 {z } AU

15.1. Score-based and post-hoc detection A series of post-hoc refinements improve separability over the MSP baseline (Section 11) without retraining: (i) temperature scaling with input perturbation, Out-of-Distribution Detector (ODIN) (Liang et al., 2020); (ii) Mahalanobis distance to class-conditional Gaussians in feature space (Lee et al., 2018c), with a relative variant for near-OOD (Ren et al., 2021); (iii) the free energy of the logits (Liu et al., 2020); and, more recently, (iv) scores derived from the gap between the top logit and the remainder (Liang et al., 2025). Training-time approaches instead learn an explicit confidence or detection signal, via a confidence branch (DeVries and Taylor, 2018) or by exposing the model to auxiliary outliers (Lee et al., 2018b). Fort et al. (2021) show that strong pretrained representations can dramatically improve on what these scores can achieve, particularly for near-OOD.

(11)

EU

The mean of individual variances captures aleatoric uncertainty; the variance of individual means captures epistemic uncertainty. This is exact under a mixture-of-Gaussians assumption. For distribution-free intervals, quantile regression (Koenker and Bassett, 1978; Tagasovska and Lopez-Paz, 2019) and its conformalized variant (Romano et al., 2019) offer alternatives. Training the aleatoric head is itself delicate. The Gaussian negative-log-likelihood objective scales the gradient of each point by its predicted variance, systematically under-sampling poorly-fit regions, which β-weighted variants correct (Seitzer et al., 2022). Hybrid models that pair a flexible (e.g. normalizing-flow) aleatoric density with a separate epistemic estimator further decouple the two components (Van Katwyk and Bergen, 2025).

15.2. Bayesian and ensemble approaches Since OOD inputs are precisely where epistemic uncertainty should be highest, the method families of this survey are themselves OOD detectors. Bayesian neural networks, dropoutbased posteriors (Nguyen et al., 2022), and ensembles supply epistemic scores directly; combining a Bayesian treatment with outlier exposure (Wang and Aitchison, 2021) or with training on synthesized OOD pseudo-inputs (Segonne et al., 2022) further sharpens detection. The standardized suites of Section 11 consolidate these methods and protocols. A recurring caution (Section 14) is that predictive entropy alone can be a poor OOD score (Kirsch and Mukhoti, 2021).

14.9. Summary of uncertainty measures Whereas Table 6 focused only on the three epistemic divergence measures, the two tables here cover the full set: Table 7 compares all the uncertainty measures introduced above on task, cost, strengths, and limitations, and Table 8 classifies them by uncertainty type (total, aleatoric, epistemic) and computational approach (information-theoretic vs. moment/divergencebased).

15.3. Selective prediction and abstention 15. Out-of-distribution detection and selective prediction

Selective prediction equips a model with a reject option, trading coverage for accuracy along a risk–coverage curve (Geifman and El-Yaniv, 2017; Hendrickx et al., 2024). Practical instantiations range from trust scores that flag likely misclassifications (Jiang et al., 2018) to calibrated selective classifiers with coverage guarantees (Fisch et al., 2022). A complementary response to high uncertainty is to abstain. The quality of the abstention decision inherits directly from the quality of the underlying uncertainty estimate, relating this task back to the method and measure choices of the preceding sections. A related decision-time capability is to explain the uncertainty, identifying the input features responsible for it via counterfactual approaches such as Counterfactual Latent Uncertainty Explanation (CLUE) (Antorán et al., 2021); we return to this largely open direction in Section 17.

With the methods that produce predictive ensembles (Section 3–Section 12) and the measures that summarize their uncertainty (Section 14) in place, we turn to the two decisiontime tasks that ultimately judge them: flagging inputs that fall outside the training distribution, and abstaining when a prediction cannot be trusted. Both reduce to thresholding an uncertainty score, yet how that score is constructed is itself a substantial research question; one that is largely independent of which method produced the predictive distribution (Table 9). A third, closely related decision-time use of the same uncertainty scores is active learning, where high epistemic uncertainty selects the most informative points to label (Houlsby et al., 2011; Gal et al., 2017) via the BALD criterion of Section 14. 15

Table 7: Comparative summary of uncertainty measures, grouped by the uncertainty they capture. All measures are for classification except where noted. The Key reference column gives the introducing or defining reference for each measure. Measure

Costa

Strengths

Limitations

Key reference

Total uncertainty Predictive Entropy EPCE

O(C) O(M 2 C)

Single scalar; info-theoretic Pairwise total uncertainty; avoids BMA-as-posterior assumption

Does not separate aleatoric/epistemic Quadratic in M

Ash (1965) Schweighofer et al. (2023a)

Aleatoric uncertainty Expected entropy

O(MC)

Mean of member entropies; aleatoric component of the decomposition

Conflated with total uncertainty if used alone

Depeweg et al. (2018)

Epistemic uncertainty MI

O(MC)

Principled; expected info gain

Houlsby et al. (2011)

EPKL

O(M 2 C)

EPJS

O(M 2 C)

VGMU

O(C)

Improved epistemic sensitivity; robust to posterior mismatch Bounded; symmetric; numerically stable Margin-based; decision-focused; epistemic-aware

Assumes BMA ≡ true posterior; underestimates under finite ensembles Quadratic in M; unbounded Quadratic in M; less sensitive than EPKL Uses top-2 classes; insensitive to tail-classes

Schweighofer et al. (2023a); Melville et al. (2005) Gillis et al. (2026b)

Behavior gating-parameter dependent

Gillis et al. (2026b)

Gaussian assumption

Lakshminarayanan et al. (2017)

Full decompositions (multiple types) Variance-gated decomposition O(MC) Variance decompositionb

O(M)

Linear-time gated decomposition (TU/AU/EU); differentiable Exact TU/AU/EU split under Gaussian mixture; tractable

Schweighofer et al. (2023a)

a Costs denote the additional cost of computing each measure given the M member predictions and their aggregate p̄; C is the number of classes. b Defined for regression; all other measures listed are for classification.

Table 8: Classification of uncertainty measures by type and computational approach. Uncertainty

Information-theoretic

Moment/divergence-based

Total Aleatoric Epistemic

Predictive entropy H[ p̄] Expected entropy E[H[pm ]] Mutual information I(y; w | x, D)

EPCE; total predictive variance (regression); variance-gated decomposition 1 P 2 Mean variance M σm (regression); variance-gated decomposition EPKL, EPJS, VGMU, variance of means (regression); variance-gated decomposition

Table 9: Out-of-distribution detection and selective-prediction approaches (Section 15). Post-hoc scores operate on a fixed trained classifier; training-time and Bayesian/ensemble approaches modify training or inference. Standardized evaluation: OpenOOD (Yang et al., 2022); strong pretrained representations raise the achievable ceiling (Fort et al., 2021). Approach

Signal/mechanism

Requirements

Representative work

None Tune T, ϵ Fit Gaussians Fit Gaussians None None

Hendrycks and Gimpel (2017) Liang et al. (2020) Lee et al. (2018c) Ren et al. (2021) Liu et al. (2020) Liang et al. (2025)

Auxiliary confidence head Train against auxiliary outliers Train against synthesized OOD inputs Diverse functions trained to vanish on in-distribution data; non-zero response flags OOD

Retraining Auxiliary OOD data Generate inputs Train certificates (no OOD data)

DeVries and Taylor (2018) Lee et al. (2018b) Segonne et al. (2022) Tagasovska and Lopez-Paz (2019)

Bayesian/ensemble BNN/dropout posterior Bayesian + outlier exposure

Posterior disagreement (epistemic) Posterior with outlier exposure

Sampling Auxiliary OOD data

Nguyen et al. (2022) Wang and Aitchison (2021)

Selective prediction/abstention Softmax response Trust score Calibrated selective classification Reject option (survey)

Confidence threshold; risk–coverage Agreement with k-NN class structure Coverage-guaranteed abstention Taxonomy of abstaining classifiers

None Reference set Calibrated set None

Geifman and El-Yaniv (2017) Jiang et al. (2018) Fisch et al. (2022) Hendrickx et al. (2024)

Post-hoc scores (training-free; operate on a fixed model) MSP Maximum class probability (baseline) ODIN Temperature scaling + input perturbation Mahalanobis Class-conditional feature distance Relative Mahalanobis Near-OOD-corrected distance Energy Free energy of the logits LogitGap Gap between top and remaining logits Training-time Learned confidence Outlier exposure OOD pseudo-inputs Orthonormal certificates

16

16. Uncertainty in large language models

16.4. Downstream use: Selective generation and conformal methods As in Section 15, the measure should support decisionmaking. Conformal and selective-generation methods provide coverage-style guarantees for conditional language models (those that generate an output conditioned on an input, as in summarization or translation) (Ren et al., 2023) and for LLMas-judge pipelines (Badshah et al., 2026), while self-certainty scores guide Best-of-N response selection (Kang et al., 2025). Here “coverage” is the generative analogue of the classification guarantee. Rather than a prediction set that contains the true label with probability 1 − α (e.g., 90% when α = 0.1), the method returns a set of sampled generations (or an abstention) that contains an acceptable response at that same level, under the caveats on exchangeability noted in Section 9. As in the classification setting of Section 14, most current scores are heuristic. They are validated by correlation with downstream errors rather than derived as estimators of a defined quantity. Semantic entropy (Kuhn et al., 2023), response-consistency checks (Manakul et al., 2023), and verbalized confidence (Tian et al., 2023) each supply a usable signal, but none targets a posterior, a proper score, or a coverage level, and so none carries a guarantee that survives a change of model, prompt, or domain. A principled measure, by contrast, would estimate a well-specified object and retain its calibration or coverage property under those shifts; conformal generation (Ren et al., 2023) is the clearest step in this direction but is still confined to narrow output settings; although, non-exchangeable variants relax the exchangeability assumption via nearest-neighbor calibration (Ulmer et al., 2024). Building measures that are principled, calibration-stable, and semantically aware, remains the central open problem in this area.

The methods surveyed so far target a categorical (or realvalued) predictive distribution over a fixed output space. Autoregressive Large Language Models (LLMs) break both assumptions; a prediction is a variable-length token sequence, and many forms express the same meaning, so a token-level probability conflates linguistic variation with genuine semantic uncertainty. This has caused a fast-growing and still largely heuristic research literature, which we present briefly here (Table 10) and which two recent surveys cover in depth (Shorinwa et al., 2024; Huang et al., 2024). The area is an emerging adjacent field rather than the main focus of the present survey. 16.1. Likelihood- and consistency-based measures The simplest signals read off the token distribution model. That is, the (length-normalized) sequence log-likelihood or pertoken entropy serves as a confidence proxy, but it is lengthbiased and, as in classification (Section 14), cannot by itself separate aleatoric from epistemic uncertainty. Semantic approaches lift the ensemble-disagreement view of Section 14 from a fixed label set to a space of meanings. Kuhn et al. (2023) sample multiple generations, cluster them by bidirectional entailment (two answers share a meaning if and only if each entails the other under a natural language inference model), and compute semantic entropy over the clusters. This scales to reliable hallucination detection (Farquhar et al., 2024). Black-box variants quantify uncertainty purely from sampledresponse consistency, such as SelfCheckGPT (Manakul et al., 2023). 16.2. Verbalized and elicited confidence A distinctively LLM capability is asking the model to state its confidence. Kadavath et al. (2022) showed that models can be trained to predict whether their own answers are correct (“knowing what they know”), and prompting strategies elicit confidence scores whose calibration varies widely (Tian et al., 2023; Xiong et al., 2024). These self-reports are cheap but fragile, and their calibration degrades under distribution shift and adversarial prompting.

17. Summary and open research directions The organizing contribution of this survey is a separation we have applied throughout. The method that produces a predictive ensemble is distinct from the measure that summarizes its uncertainty and the measures to evaluate the quality of the uncertainty estimates (Figure 1). The two stages communicate only through the set of member predictions, so in principle any method can be paired with any measure, yet the literature has largely studied them apart. This interchangeability is strongest among methods that expose an explicit set of member predictions (the explicit-member and posterior-sampling rows of Table 2), for which the sample-based measures (predictive entropy, the mutual-information decomposition, and the pairwise divergences) are freely substitutable; closed-form singlepass methods return only a marginal predictive plus a methodspecific epistemic signal (Dirichlet precision, feature density, or GP variance), so the sample-based decompositions apply to them only once their predictive is sampled, which moves them up a row in Table 2. Reading the field through this lens clarifies that progress on generating diverse predictions does not automatically transfer to reading uncertainty off them, and that comparisons are meaningful only against a stated evidentiary basis (Section 11).

16.3. Bayesian, ensemble, and decomposition views The Bayesian, ensemble, and decomposition views developed above (Section 14) transfer only partially to the generative setting. Epistemic uncertainty can be captured through (implicit) ensembles, but Kirsch (2025) show that scale induces an epistemic-uncertainty collapse across nominally independent members (Section 6). The token- and sequence-level entropy and MI decomposition was first extended from a fixed label set to autoregressive structured prediction using ensembles of autoregressive models (Malinin and Gales, 2021). Decomposition ideas have since been adapted to the generative setting for natural-language generation and for in-context learning (Jayasekera et al., 2025), and uncertainty has been argued to be decision-critical for LLM agents (Felicioni et al., 2024). 17

Table 10: Families of uncertainty quantification for large language models. Access indicates whether a method needs model internals/logits (white-box), only sampled output text (black-box), or a held-out calibration set. Cost is in forward passes/samples per query; K = number of sampled generations, M = ensemble size. Family

Accessa

Uncertainty signal

Cost

Representative work

Token likelihood

White

1

Malinin and Gales (2021)

Semantic entropy

Whiteb

K (+entailment)

Kuhn et al. (2023); Farquhar et al. (2024)

Sampled consistency

Black

K

Manakul et al. (2023)

Verbalized confidence

Black

1 (+prompt)

Ensemble/Bayesian

White/Grey

Kadavath et al. (2022); Tian et al. (2023); Xiong et al. (2024) Kirsch (2025); Jayasekera et al. (2025)

Conformal/selective

Calibration set

Sequence log-likelihood/per-token entropy (length-biased) Entropy over meaning-equivalence clusters of sampled generations Agreement/similarity across independently sampled responses Model-elicited (self-reported) confidence score Disagreement across members, seeds, or prompts; decomposed AU/EU Coverage-calibrated output sets or abstention

M wrapper

Ren et al. (2023); Badshah et al. (2026); Kang et al. (2025)

a White = requires logits/internal states; Black = output text only; Grey = requires multiple stochastic passes. b Black-box if an external entailment/clustering model supplies the semantic grouping.

identities at sub-quadratic cost. A concrete near-term step is to validate bounded pairwise measures (EPJS) against MI and EPKL on the OOD/shift suites of Section 11, closing the gap between the structural argument of Section 14.5 and empirical performance.

17.1. Key findings The survey is consolidated in two sets of comparison tables placed with their subject matter: (i) the method tables (Table 3 and Table 4, in Section 12; the taxonomy of Table 2 in Section 2); and (ii) the measure tables (Table 7–Table 8, in Section 14), their qualitative ratings should be read against Section 11. From these, several patterns emerge. First, there is a consistent trade-off between computational cost and uncertainty quality. Methods that explore multiple posterior modes (deep ensembles, CreDE, MultiSWAG) produce superior uncertainty but at the cost of training multiple models. Efficient approximations via weight sharing, trajectory collection, or stochastic masking invariably sacrifice diversity and yield less informative uncertainty under shift. Second, closedform predictive methods (those that read uncertainty analytically from a single deterministic pass) can nonetheless achieve strong uncertainty with appropriate inductive biases: SNGP, DDU, and prior/posterior networks pursue this by different means. Third, the classification in Table 2 explains why deep ensembles and their extensions lead benchmarks; they are the explicit-member methods that operate over the full network, the only cell in which between-mode (multi-modal) diversity is achieved. Fourth, for measures, information-theoretic estimators exist for all three uncertainty types in classification, whereas moment- and divergence-based alternatives are either regression-only or quadratic in M; efficient, non-entropy-based epistemic measures for classification remain under-explored.

Diversity and calibration under shift. The relationship between ensemble diversity and calibration under covariate shift needs deeper theory; current diversity metrics are largely empirical. Hybrid architectures. Since a method that produces predictions and a measure that summarizes them are separable (Figure 1), single-pass techniques can in principle be combined with ensembling. For example, a small deep ensemble of SNGP or DDU backbones, or a single distance-aware backbone carrying several last-layer heads, would add between-mode diversity to a distance-aware single-pass model at a fraction of cost of full ensembles. Whether such hybrids recover most of the uncertainty quality of a deep ensemble for that added cost is an open question; systematic evaluation is lacking. Input-dependent uncertainty structure. The classification analogue of heteroscedastic regression (that ensemble reliability varies across inputs and classes) suggests per-class uncertainty structure that current aggregation largely ignores. Explaining uncertainty estimates. Beyond producing and measuring uncertainty, explaining why a model is uncertain about a given input (Section 15) remains under-explored. Counterfactual approaches such as CLUE are increasingly relevant for trust and deployment but remain largely absent from the ensemble-and-measure pipeline surveyed here.

17.2. Open research questions Last-layer diversity and its limits. Whether last-layer diversity alone can recover most of the benefit of full deep ensembles impacts directly on efficient architecture design. The extent to which shared gradients collapse classification heads, and what mitigates it, is open (Kirsch, 2025).

Uncertainty in large language models. The most consequential open frontier is generative UQ (Section 16). The ensemble-andmeasure framework transfers only partially to this setting. Semantic entropy (Kuhn et al., 2023; Farquhar et al., 2024) is the generative analogue of the predictive-entropy decomposition, and sampled-response consistency (Manakul et al., 2023) mirrors the pairwise-divergence measures of Section 14.4. Three

Efficient epistemic measures for classification. The demonstration that additive entropy decompositions can fail under finite ensembles and posterior mismatch (Schweighofer et al., 2023a; Wimmer et al., 2023) motivates measures that bypass entropy 18

questions stand out: (i) whether a principled aleatoric/epistemic split exists for free-form generation (Jayasekera et al., 2025); (ii) whether verbalized confidence can be made calibrationstable under distribution shift and adversarial prompting (Tian et al., 2023; Xiong et al., 2024; Kadavath et al., 2022); and (iii) whether the epistemic-collapse of large models (Kirsch, 2025) can be counteracted without the prohibitive cost of full ensembles.

URL: https://proceedings.mlr.press/v119/ van-amersfoort20a/van-amersfoort20a.pdf. Angelopoulos, A.N., Bates, S., 2023. Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning 16, 494–591. doi:10.1561/2200000101. Antorán, J., Bhatt, U., Adel, T., Weller, A., Hernández-Lobato, J.M., 2021. Getting a CLUE: A method for explaining uncertainty estimates, in: International Conference on Learning Representations (ICLR), pp. 1–34. URL: https:// openreview.net/pdf?id=XSLF1XFq5h.

Conclusion. On the method side the evidence is consistent. Full-network, multi-modal exploration, deep ensemble, and their multi-modal extensions continue to set the standard for uncertainty quality, while efficient single-pass and last-layer approaches recover much of it when equipped with the right inductive bias, but do not capture the between-mode epistemic component. On the measure side the picture is less settled. The additive entropy/mutual-information decomposition remains the default despite known failures under finite ensembles and posterior mismatch, and bounded or efficient alternatives such as EPJS and variance-gated scores are structurally attractive but still await systematic empirical validation. The practitioner’s choice therefore remains a deliberate cost–quality trade-off rather than a single dominant framework. The same framework indicates where the field is heading. The most consequential frontier is generative UQ (Section 16), which requires lifting the diversity, decomposition, and calibration tools surveyed here from a fixed label set into a space of semantic equivalences, where even the aleatoric/epistemic split must be redefined. Whether the method–measure separation that organized this survey survives that move is, in our view, the most important question it leaves open.

Apostolopoulou, I., Eysenbach, B., Nielsen, F., Dubrawski, A., 2024. A rate-distortion view of uncertainty quantification, in: International Conference on Machine Learning (ICML), pp. 1–24. URL: https://openreview.net/pdf? id=zMGUDsPopK. Arbel, J., Pitas, K., Vladimirova, M., Fortuin, V., 2023. A primer on Bayesian neural networks: Review and debates. arXiv preprint. doi:10.48550/arXiv.2309.16314. Ash, R.B., 1965. Information theory. Dover Publications. Ashukha, A., Lyzhov, A., Molchanov, D., Vetrov, D., 2020. Pitfalls of in-domain uncertainty estimation and ensembling in deep learning, in: International Conference on Learning Representations (ICLR), pp. 1–30. URL: https:// openreview.net/pdf?id=BJxI5gHKDr. Badshah, S., Emami, A., Sajjad, H., 2026. Scope: Selective conformal optimized pairwise LLM judging, in: International Conference on Machine Learning (ICML), pp. 1–23. doi:10.48550/arXiv.2602.13110.

References Abdar, M., Pourpanah, F., Hussain, S., Rezazadegan, D., Liu, L., Ghavamzadeh, M., Fieguth, P., Cao, X., Khosravi, A., Acharya, U.R., Makarenkov, V., Nahavandi, S., 2021. A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion 76, 243–297. doi:10.1016/j.inffus.2021.05.008.

Berry, L., Meger, D., 2023. Efficient epistemic uncertainty estimation in regression ensemble models using pairwisedistance estimators, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–34. URL: https: //openreview.net/pdf?id=PsVp5Gxm0s. Blei, D.M., Kucukelbir, A., McAuliffe, J.D., 2017. Variational inference: A review for statisticians. Journal of the American Statistical Association 112, 859–877. doi:10.1080/ 01621459.2017.1285773.

Abe, T., Buchanan, E.K., Pleiss, G., Zemel, R., Cunningham, J.P., 2022. Deep ensembles work, but are they necessary?, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–15. URL: https://openreview.net/ pdf?id=Wl1ZIgMqLlq.

Blundell, C., Cornebise, J., Kavukcuoglu, K., Wierstra, D., 2015. Weight uncertainty in neural networks, in: International Conference on Machine Learning (ICML), pp. 1613– 1622. URL: https://proceedings.mlr.press/v37/ blundell15.pdf.

van Amersfoort, J., Smith, L., Jesson, A., Key, O., Gal, Y., 2021. On feature collapse and deep kernel learning for single forward pass uncertainty, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1– 16. URL: https://bayesiandeeplearning.org/2021/ papers/28.pdf. Workshop: Bayesian deep learning.

Boluki, S., Ardywibowo, R., Dadaneh, S.Z., Zhou, M., Qian, X., 2020. Learnable bernoulli dropout for Bayesian deep learning, in: International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 3905– 3916. URL: https://proceedings.mlr.press/v108/ boluki20a/boluki20a.pdf.

van Amersfoort, J., Smith, L., Teh, Y.W., Gal, Y., 2020. Uncertainty estimation using a single deep deterministic neural network, in: International Conference on Machine Learning (ICML), pp. 9690–9700. 19

Brier, G.W., 1950. Verification of forecasts expressed in terms of probability. Monthly Weather Review 78, 1–3. doi:10. 1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2.

Egele, R., Maulik, R., Raghavan, K., Lusch, B., Guyon, I., Balaprakash, P., 2022. AutoDEUQ: Automated deep ensemble with uncertainty quantification, in: International Conference on Pattern Recognition (ICPR), pp. 1908– 1914. URL: https://ieeexplore.ieee.org/stamp/ stamp.jsp?arnumber=9956231.

Brosse, N., Riquelme, C., Martin, A., Gelly, S., Moulines, E., 2020. On last-layer algorithms for classification: Decoupling representation from uncertainty estimation. arXiv preprint. doi:10.48550/arXiv.2001.08049.

Farquhar, S., Kossen, J., Kuhn, L., Gal, Y., 2024. Detecting hallucinations in large language models using semantic entropy. Nature 630, 625–630. doi:10.1038/ s41586-024-07421-0.

Brown, G., Wyatt, J., Harris, R., Yao, X., 2005. Diversity creation methods: A survey and categorisation. Information Fusion 6, 5–20. doi:10.1016/j.inffus.2004.04.004.

Felicioni, N., Maystre, L., Ghiassian, S., Ciosek, K., 2024. On the importance of uncertainty in decision-making with large language models. Transactions on Machine Learning Research , 1–27URL: https://openreview.net/pdf?id= YfPzUX6DdO.

Charpentier, B., Zügner, D., Günnemann, S., 2020. Posterior network: Uncertainty estimation without OOD samples via density-based pseudo-counts, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1356– 1367. URL: https://dl.acm.org/doi/pdf/10.5555/ 3495724.3495839.

Fisch, A., Jaakkola, T., Barzilay, R., 2022. Calibrated selective classification. Transactions on Machine Learning Research , 1–25URL: https://openreview.net/pdf?id= zFhNBs8GaV.

Chau, S.L., Caprio, M., Muandet, K., 2025. Integral imprecise probability metrics, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–42. URL: https:// openreview.net/pdf?id=KM2XzHq2Rm.

Fort, S., Hu, H., Lakshminarayanan, B., 2020. Deep ensembles: A loss landscape perspective. arXiv preprint. doi:10.48550/ arXiv.1912.02757.

Chauhan, V.K., Zhou, J., Lu, P., Molaei, S., Clifton, D.A., 2024. A brief review of hypernetworks in deep learning. Artificial Intelligence Review 57. doi:10.1007/ s10462-024-10862-8.

Fort, S., Ren, J., Lakshminarayanan, B., 2021. Exploring the limits of out-of-distribution detection, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–14. URL: https://openreview.net/pdf?id=j5NrN8ffXC.

Damianou, A.C., Lawrence, N.D., 2013. Deep gaussian process, in: International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 1–9. URL: https: //proceedings.mlr.press/v31/damianou13a.pdf.

Fortuin, V., 2022. Priors in Bayesian deep learning: A review. International Statistical Review 90, 563–591. doi:10.1111/ insr.12502.

Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., Hennig, P., 2021. Laplace redux – effortless Bayesian deep learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–15. URL: https://openreview.net/pdf?id=gDcaUj4Myhn.

Fortuin, V., Garriga-Alonso, A., Ober, S.W., Wenzel, F., Ratsch, G., Turner, R.E., van der Wilk, M., Aitchison, L., 2022. Bayesian neural network priors revisited, in: International Conference on Learning Representations (ICLR), pp. 1–35. URL: https://openreview.net/pdf?id=xkjqJYqRJy.

Depeweg, S., Hernández-Lobato, J.M., Doshi-Velez, F., Udluft, S., 2018. Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning, in: International Conference on Machine Learning (ICML), pp. 1–10. URL: https://proceedings.mlr.press/v80/ depeweg18a/depeweg18a.pdf.

Friedman, D., Dieng, A.B., 2023. The Vendi score: A diversity evaluation metric for machine learning. Transactions on Machine Learning Research , 1–26URL: https: //openreview.net/pdf?id=g97OHbQyk1. Gal, Y., Ghahramani, Z., 2016. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning, in: International Conference on Machine Learning (ICML), pp. 1050–1059. URL: https://proceedings.mlr.press/ v48/gal16.pdf.

DeVries, T., Taylor, G.W., 2018. Learning confidence for outof-distribution detection in neural networks. arXiv preprint. doi:10.48550/arXiv.1802.04865. Dietterich, T.G., 2000. Ensemble methods in machine learning, in: Multiple Classifier Systems, Springer. pp. 1–15. doi:10. 1007/3-540-45014-9_1.

Gal, Y., Islam, R., Ghahramani, Z., 2017. Deep Bayesian active learning with image data, in: International Conference on Machine Learning (ICML), pp. 1183–1192. URL: https: //proceedings.mlr.press/v70/gal17a/gal17a.pdf.

Durasov, N., Bagautdinov, T., Baque, P., Fua, P., 2021. Masksembles for uncertainty estimation, in: Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13539–13548. doi:10.1109/cvpr46437.2021.01333.

Galdran, A., Verjans, J., Carneiro, G., González Ballester, M.A., 2023. Multi-head multi-loss model calibration, in: International Conference on Medical Image Computing and 20

Computer-Assisted Intervention, pp. 108–117. 1007/978-3-031-43898-1_11.

doi:10.

Harrison, J., Willes, J., Snoek, J., 2024. Variational Bayesian last layers, in: International Conference on Learning Representations (ICLR), pp. 1–30. URL: https://openreview. net/pdf?id=Sx7BIiPzys.

Gawlikowski, J., Tassi, C.R.N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R., Zhu, X.X., 2023. A survey of uncertainty in deep neural networks. Artificial Intelligence Review 56, 1513–1589. doi:10.1007/ s10462-023-10562-9.

Havasi, M., Jenatton, R., Fort, S., Liu, J.Z., Snoek, J., Lakshminarayanan, B., Dai, A.M., Tran, D., 2021. Training independent subnetworks for robust prediction, in: International Conference on Learning Representations (ICLR), pp. 1–13. URL: https://openreview.net/pdf?id= OGg9XnKxFAH.

Geifman, Y., El-Yaniv, R., 2017. Selective classification for deep neural networks, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 4885–4894. URL: https://proceedings. neurips.cc/paper_files/paper/2017/file/ 4a8423d5e91fda00bb7e46540e2b0cf1-Paper.pdf.

He, W., Jiang, Z., Xiao, T., Xu, Z., Li, Y., 2026. A survey on uncertainty quantification methods for deep learning. ACM Computing Surveys 58, 179. doi:10.1145/3786319.

Gillis, H.M., Xu, I., Misiuk, B., Brown, C.J., Trappenberg, T., 2026a. Last-layer committee machines for uncertainty estimations of benthic imagery. Neural Networks 200. doi:10. 1016/j.neunet.2026.108819.

Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., Davis, J., 2024. Machine learning with a reject option: A survey. Machine Learning 113, 3073–3110. doi:10.1007/ s10994-024-06534-x.

Gillis, H.M., Xu, I., Trappenberg, T., 2025. Uncertainty estimation using variance-gated distributions, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–21. URL: https://openreview.net/pdf?id=vTDvrwtaG0. Workshop: Mathematical Foundations and Operational Integration of Machine Learning for Uncertainty-Aware Decision-Making Workshop.

Hendrycks, D., Dietterich, T., 2019. Benchmarking neural network robustness to common corruptions and perturbations, in: International Conference on Learning Representations (ICLR), pp. 1–16. URL: https://openreview.net/pdf? id=HJz6tiCqYm. Hendrycks, D., Gimpel, K., 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks, in: International Conference on Learning Representations (ICLR), pp. 1–12. URL: https://openreview.net/ pdf?id=Hkg4TI9xl.

Gillis, H.M., Xu, I., Trappenberg, T., 2026b. Variancegated ensembles: An epistemic-aware framework for uncertainty estimation. Transactions on Machine Learning Research , 1–55URL: https://openreview.net/pdf?id= fNMZjV1gje.

Herron, E.J., Young, S.R., Potok, T.E., 2020. Ensembles of networks produced from neural architecture search, in: High Performance Computing. Springer. volume 12321 of Lecture Notes in Computer Science, pp. 223–234. doi:10.1007/ 978-3-030-59851-8_14.

Gneiting, T., Raftery, A.E., 2007. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association 102, 359–378. doi:10.1198/ 016214506000001437. Graves, A., 2011. Practical variational inference for neural networks, in: Conference on Neural Information Processing Systems (NIPS), pp. 2348–2356. URL: https: //papers.nips.cc/paper_files/paper/2011/file/ 7eb3c8be3d411e8ebfab08eba5f49632-Paper.pdf.

Hoffman, M.D., 2013. Stochastic variational inference. Journal of Machine Learning Research 14, 1303–1347. URL: https://www.jmlr.org/papers/ volume14/hoffman13a/hoffman13a.pdf.

Guo, C., Pleiss, G., Sun, Y., Weinberger, K.Q., 2017. On calibration of modern neural networks, in: International Conference on Machine Learning (ICML), pp. 1321–1330. URL: https://proceedings.mlr.press/v70/guo17a/ guo17a.pdf.

Houlsby, N., Huszár, F., Ghahramani, Z., Lengyel, M., 2011. Bayesian active learning for classification and preference learning. arXiv preprint. doi:10.48550/arXiv.1112. 5745. Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J.E., Weinberger, K.Q., 2017. Snapshot ensembles: Train 1, get M for free, in: International Conference on Learning Representations (ICLR), pp. 1–14. URL: https://openreview.net/ pdf?id=BJYwwY9ll.

Gustafsson, F.K., Danelljan, M., Schön, T.B., 2020. Evaluating scalable Bayesian deep learning methods for robust computer vision, in: Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1289–1298. URL: https://openaccess.thecvf.com/content_ CVPRW_2020/papers/w20/Gustafsson_Evaluating_ Scalable_Bayesian_Deep_Learning_Methods_for_ Robust_Computer_Vision_CVPRW_2020_paper.pdf.

Huang, H.Y., Yang, Y., Zhang, Z., Lee, S., Wu, Y., 2024. A survey of uncertainty estimation in LLMs: Theory meets practice. arXiv preprint. doi:10.48550/arXiv.2410.15326. 21

Hubin, A., Storvik, G., 2019. Combining model and parameter uncertainty in Bayesian neural networks. arXiv preprint. doi:10.48550/arXiv.1903.07594.

Kendall, A., Gal, Y., 2017. What uncertainties do we need in Bayesian deep learning for computer vision?, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 5580–5590. URL: https://proceedings. neurips.cc/paper_files/paper/2017/file/ 2650d6089a6d640c5e85b2b88265dc2b-Paper.pdf.

Hüllermeier, E., Waegeman, W., 2021. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Machine Learning 110, 457–506. doi:10.1007/s10994-021-05946-3.

Kirsch, A., 2025. (implicit) ensembles of ensembles: Epistemic uncertainty collapse in large models. Transactions on Machine Learning Research , 1–28URL: https:// openreview.net/pdf?id=ON7dtdEHVQ.

Izmailov, P., Vikram, S., Hoffman, M.D., Wilson, A.G., 2021. What are Bayesian neural network posteriors really like?, in: International Conference on Machine Learning (ICML), pp. 1–12. URL: https://proceedings.mlr.press/v139/ izmailov21a/izmailov21a.pdf.

Kirsch, A., Mukhoti, J., 2021. On pitfalls in OoD detection: Predictive entropy considered harmful, in: International Conference on Machine Learning (ICML), pp. 1–12. URL: https://www.gatsby.ucl.ac.uk/~balaji/ udl2021/accepted-papers/UDL2021-paper-092.pdf. Workshop: Uncertainty and robustness in deep learning.

Jain, S., Liu, G., Mueller, J., Gifford, D., 2020. Maximizing overall diversity for improved uncertainty estimates in deep ensembles, in: AAAI Conference on Artificial Intelligence, pp. 4264–4271. doi:10.1609/aaai.v34i04.5849.

Koenker, R., Bassett, G., 1978. Regression quantiles. Econometrica 46, 33–50. doi:10.2307/1913643.

Jayasekera, I.S., Si, J., Valdettaro, F., Chen, W., Faisal, A.A., Li, Y., 2025. Variational uncertainty decomposition for in-context learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–62. URL: https: //openreview.net/pdf?id=MqGZIJxZ1z.

Kristiadi, A., Hein, M., Hennig, P., 2020. Being Bayesian, even just a bit, fixes overconfidence in ReLU networks, in: International Conference on Machine Learning (ICML), pp. 1–11. URL: https://proceedings.mlr.press/v119/ kristiadi20a/kristiadi20a.pdf.

Jiang, H., Kim, B., Guan, M.Y., Gupta, M., 2018. To trust or not to trust a classifier, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 5546–5557. URL: https://proceedings. neurips.cc/paper_files/paper/2018/file/ 7180cffd6a8e829dacfc2a31b3f72ece-Paper.pdf.

Kuhn, L., Gal, Y., Farquhar, S., 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, in: International Conference on Learning Representations (ICLR), pp. 1–19. URL: https:// openreview.net/pdf?id=VD-AYtP0dve.

Joshi, A.J., Porikli, F., Papanikolopoulos, N., 2009. Multiclass active learning for image classification, in: Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2372–2379. doi:10.1109/CVPR.2009.5206627.

Kumar, A., Ma, T., Liang, P., Raghunathan, A., 2022. Calibrated ensembles can mitigate accuracy tradeoffs under distribution shift, in: Conference on Uncertainty in Artificial Intelligence (UAI), pp. 1–11. URL: https://openreview. net/pdf?id=S2EuSP8sqgc.

Jospin, L.V., Laga, H., Boussaid, F., Buntine, W., Bennamoun, M., 2022. Hands-on Bayesian neural networks—a tutorial for deep learning users. IEEE Computational Intelligence Magazine 17, 29–48. doi:10.1109/MCI.2022.3155327.

Kuncheva, L.I., Whitaker, C.J., 2003. Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Machine Learning 51, 181–207. doi:10.1023/A: 1022859003006.

Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., Kaplan, J., 2022. Language models (mostly) know what they know. arXiv preprint. doi:10.48550/arXiv.2207.05221.

Kushibar, K., Campello, V.M., Moras, L.G., Linardos, A., Radeva, P., Lekadir, K., 2022. Layer ensembles: A singlepass uncertainty estimation in deep learning for segmentation, in: International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), pp. 514–524. doi:10.1007/978-3-031-16452-1_49.

Lakshminarayanan, B., Pritzel, A., Blundell, C., 2017. Simple Kang, Z., Zhao, X., Song, D., 2025. Scalable best-of-N and scalable predictive uncertainty estimation using deep enselection for large language models via self-certainty, in: sembles, in: Conference on Neural Information Processing Conference on Neural Information Processing Systems Systems (NIPS), pp. 6405–6416. URL: https://dl.acm. (NeurIPS), pp. 1–26. URL: https://proceedings. org/doi/pdf/10.5555/3295222.3295387. neurips.cc/paper_files/paper/2025/file/ 1c7eff166a8e345f664f0faa8f4e4d2e-Paper-Conference.Lee, J., Bahri, Y., Novak, R., Schoenholz, S.S., Pennington, J., pdf. Sohl-Dickstein, J., 2018a. Deep neural networks as Gaussian 22

processes, in: International Conference on Learning Representations (ICLR), pp. 1–17. URL: https://openreview. net/pdf?id=B1EA-M-0Z.

via distance-awareness. Journal of Machine Learning Research 24, 1–63. URL: https://www.jmlr.org/papers/ volume24/22-0479/22-0479.pdf.

Lee, K., Lee, H., Lee, K., Shin, J., 2018b. Training confidencecalibrated classifiers for detecting out-of-distribution samples, in: International Conference on Learning Representations (ICLR), arXiv. pp. 1–16. URL: https:// openreview.net/pdf?id=ryiAv2xAZ.

Liu, W., Wang, X., Owens, J.D., Li, Y., 2020. Energy-based out-of-distribution detection, in: Conference on Neural Information Processing Systems (NeurIPS), Curran Associates Inc.. pp. 1–12. URL: https://proceedings. neurips.cc/paper_files/paper/2020/file/ f5496252609c43eb8a3d147ab9b9c006-Paper.pdf.

Lee, K., Lee, K., Lee, H., Shin, J., 2018c. A simple unified framework for detecting out-of-distribution samples and adversarial attacks, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 7167–7177. URL: https://proceedings. neurips.cc/paper_files/paper/2018/file/ abdeb6f575ac5c6676b747bca8d09cc2-Paper.pdf, doi:10.48550/arXiv.1807.03888.

MacKay, D.J.C., 1992. A practical Bayesian framework for backpropagation networks. Neural Computation 4, 448–472. doi:10.1162/neco.1992.4.3.448. Maddox, W., Garipov, T., Izmailov, P., Vetrov, D., Wilson, A.G., 2019. A simple baseline for Bayesian uncertainty in deep learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https://dl. acm.org/doi/pdf/10.5555/3454287.3455466.

Lee, S., Purushwalkam, S., Cogswell, M., Crandall, D., Batra, D., 2015. Why M heads are better than one: Training a diverse ensemble of deep networks. arXiv preprint. doi:10.48550/arXiv.1511.06314.

Malinin, A., Gales, M., 2018. Predictive uncertainty estimation via prior networks, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 7047– 7058. URL: https://dl.acm.org/doi/pdf/10.5555/ 3327757.3327808.

Li, L., Chen, Y., Yue, X., 2025. Vicinal label supervision for reliable aleatoric and epistemic uncertainty estimation, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–30. URL: https://openreview.net/ pdf?id=hPfICQIDOm.

Malinin, A., Gales, M., 2019. Reverse KL-divergence training of prior networks: Improved uncertainty and adversarial robustness, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https:// proceedings.neurips.cc/paper_files/paper/2019/ file/7dd2ae7db7d18ee7c9425e38df1af5e2-Paper. pdf.

Li, Y., Gal, Y., 2017. Dropout inference in Bayesian neural networks with alpha-divergences, in: International Conference on Machine Learning (ICML), pp. 2052– 2061. URL: https://proceedings.mlr.press/v70/ li17a/li17a.pdf.

Malinin, A., Gales, M., 2021. Uncertainty estimation in autoregressive structured prediction, in: International Conference on Learning Representations (ICLR), pp. 1–31. URL: https://openreview.net/pdf?id=jN5y-zb5Q7m.

Liang, J., Hou, R., Hu, M., Chang, H., Shan, S., Chen, X., 2025. Revisiting logit distributions for reliable out-ofdistribution detection, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–33. URL: https: //openreview.net/pdf?id=FLdLPUqnsP.

Malinin, A., Mlodozeniec, B., Gales, M., 2020. Ensemble distribution distillation, in: International Conference on Learning Representations (ICLR), pp. 1–22. URL: https: //openreview.net/pdf?id=BygSP6Vtvr.

Liang, S., Li, Y., Srikant, R., 2020. Enhancing the reliability of out-of-distribution image detection in neural networks, in: International Conference on Learning Representations (ICLR), pp. 1–27. URL: https://openreview.net/pdf? id=H1VGkIxRZ.

Manakul, P., Liusie, A., Gales, M.J.F., 2023. SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models, in: Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9004–9017. doi:10.18653/v1/2023.emnlp-main.557.

Liu, J., Paisley, J., Kioumourtzoglou, M.A., Coull, B., 2019. Accurate uncertainty estimation and decomposition in ensemble learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https://proceedings. neurips.cc/paper_files/paper/2019/file/ 1cc8a8ea51cd0adddf5dab504a285915-Paper.pdf.

Melville, P., Yang, S.M., Saar-Tsechansky, M., Mooney, R.J., 2005. Active learning for probability estimation using Jensen-Shannon divergence, in: European Conference on Machine Learning, pp. 268–279. doi:10.1007/11564096_ 28. Mena, J., Pujol, O., Vitrià, J., 2021. A survey on uncertainty estimation in deep learning classification systems from a Bayesian perspective. ACM Computing Surveys 54, 1–35. doi:10.1145/3477140.

Liu, J.Z., Padhy, S., Ren, J., Lin, Z., Wen, Y., Jerfel, G., Nado, Z., Snoek, J., Tran, D., Lakshminarayanan, B., 2022. A simple approach to improve single-model deep uncertainty 23

Minderer, M., Djolonga, J., Romijnders, R., Hubis, F., Zhai, X., Houlsby, N., Tran, D., Lucic, M., 2021. Revisiting the calibration of modern neural networks, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–13. URL: https://openreview.net/pdf?id=QRBvLayFXI.

prognostics: A tutorial. Mechanical Systems and Signal Processing 205. doi:10.1016/j.ymssp.2023.110796. Nguyen, A.T., Lu, F., Munoz, G.L., Raff, E., Nicholas, C., Holt, J., 2022. Out of distribution data detection using dropout Bayesian neural networks, in: Conference on Artificial Intelligence, pp. 7877–7885. doi:10.1609/aaai.v36i7.20757.

Mobiny, A., Yuan, P., Moulik, S.K., Garg, N., Wu, C.C., Van Nguyen, H., 2021. DropConnect is effective in modeling uncertainty of Bayesian deep networks. Scientific Reports 11. doi:10.1038/s41598-021-84854-x.

Oleksiienko, I., Iosifidis, A., 2023. Layer ensembles, in: International Workshop on Machine Learning for Signal Processing (MLSP), pp. 1–6. doi:10.1109/MLSP55844.2023. 10286005.

Monod, M., Micheli, A., Bhatt, S., 2025. NeuralSurv: Deep survival analysis with Bayesian uncertainty quantification, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–53. URL: https://openreview.net/ pdf?id=c768Z1FwDL.

Osband, I., Wen, Z., Asghari, S.M., Dwaracherla, V., IBRAHIMI, M., Lu, X., Van Roy, B., 2023. Epistemic neural networks, in: Advances in Neural Information Processing Systems, pp. 2795–2823. URL: https://openreview. net/pdf?id=dZqcC1qCmB.

Mukhoti, J., Kirsch, A., van Amersfoort, J., Torr, P.H.S., Gal, Y., 2023. Deep deterministic uncertainty: A new simple baseline, in: Conference on Computer Vision and Pattern Recognition (CVPR), pp. 24384–24394. doi:10.1109/ CVPR52729.2023.02336.

Ovadia, Y., Fertig, E., Ren, J., Nado, Z., Sculley, D., Nowozin, S., Dillon, J.V., Lakshminarayanan, B., Snoek, J., 2019. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https://dl.acm.org/doi/pdf/10.5555/ 3454287.3455541.

Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P.H.S., Dokania, P.K., 2020. Calibrating deep neural networks using focal loss, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 15288–15299. URL: https: //proceedings.neurips.cc/paper/2020/file/ aeb7b30ef1d024a76f21a1d40e30c302-Paper.pdf.

Pearce, T., Leibfried, F., Brintrup, A., Zaki, M., Neely, A., 2020. Uncertainty in neural networks: Approximately Bayesian ensembling, in: International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 1–10. URL: https://proceedings.mlr.press/v108/ pearce20a/pearce20a.pdf.

Nado, Z., Band, N., Collier, M., Djolonga, J., Dusenberry, M.W., Farquhar, S., Feng, Q., Filos, A., Havasi, M., Jenatton, R., Jerfel, G., Liu, J., Mariet, Z., Nixon, J., Padhy, S., Ren, J., Rudner, T.G.J., Sbahi, F., Wen, Y., Wenzel, F., Murphy, K., Sculley, D., Lakshminarayanan, B., Snoek, J., Gal, Y., Tran, D., 2022. Uncertainty baselines: Benchmarks for uncertainty & robustness in deep learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1– 12. URL: https://bayesiandeeplearning.org/2021/ papers/21.pdf. Workshop: Bayesian Deep Learning.

Platt, J.C., 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods, in: Advances in Large Margin Classifiers. volume 10, pp. 61–74. Postels, J., Segu, M., Sun, T., Sieber, L., Van Gool, L., Yu, F., Tombari, F., 2022. On the practicality of deterministic epistemic uncertainty, in: International Conference on Machine Learning (ICML), pp. 1– 40. URL: https://proceedings.mlr.press/v162/ postels22a/postels22a.pdf.

Naeini, M.P., Cooper, G.F., Hauskrecht, M., 2015. Obtaining well-calibrated probabilities using Bayesian binning, in: AAAI Conference on Artificial Intelligence, pp. 2901–2907. doi:10.1609/aaai.v29i1.9602.

Rahaman, R., Thiery, A.H., 2021. Uncertainty quantification and deep ensembles, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–13. URL: https:// openreview.net/pdf?id=wg_kD_nyAF.

Nandy, J., Hsu, W., Lee, M.L., 2021. Towards maximizing the representation gap between in-domain & out-of-distribution examples, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https: //proceedings.neurips.cc/paper/2020/file/ 68d3743587f71fbaa5062152985aff40-Paper.pdf.

Ramé, A., Cord, M., 2021. DICE: Diversity in deep ensembles via conditional redundancy adversarial estimation, in: International Conference on Learning Representations (ICLR), pp. 1–31. URL: https://openreview.net/pdf? id=R2ZlTVPx0Gk.

Neal, R.M., 1995. Bayesian learning for neural networks. volume 118 of Lecture Notes in Statistics. Springer. doi:10. 1007/978-1-4612-0745-0.

Ren, J., Fort, S., Liu, J., Roy, A.G., Padhy, S., Lakshminarayanan, B., 2021. A simple fix to Mahalanobis distance for improving near-OOD detection, in: International Conference on Machine Learning (ICML), pp.

Nemani, V., Biggio, L., Huan, X., Hu, Z., Fink, O., Tran, A., Wang, Y., Zhang, X., Hu, C., 2023. Uncertainty quantification in machine learning for engineering design and health 24

1–8. URL: https://www.gatsby.ucl.ac.uk/~balaji/ udl2021/accepted-papers/UDL2021-paper-007.pdf. Workshop: Uncertainty and robustness in deep learning.

pp. 3183–3193. URL: https://dl.acm.org/doi/pdf/ 10.5555/3327144.3327239. Sharma, M., Farquhar, S., Nalisnick, E., Rainforth, T., 2023. Do Bayesian neural networks need to be fully stochastic?, in: International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 1–29. URL: https://proceedings. mlr.press/v206/sharma23a/sharma23a.pdf.

Ren, J., Luo, J., Zhao, Y., Krishna, K., Saleh, M., Lakshminarayanan, B., Liu, P.J., 2023. Out-of-distribution detection and selective generation for conditional language models, in: International Conference on Learning Representations (ICLR), arXiv. pp. 1–32. URL: https://openreview. net/pdf?id=kJUS5nD0vPB.

Shen, M., Bu, Y., Sattigeri, P., Ghosh, S., Das, S., Wornell, G., 2023. Post-hoc uncertainty learning using a Dirichlet metamodel, in: Conference on Artificial Intelligence, pp. 9772– 9781. doi:10.1609/aaai.v37i8.26167.

Romano, Y., Patterson, E., Candès, E., 2019. Conformalized quantile regression, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–11. URL: https:// proceedings.neurips.cc/paper_files/paper/2019/ file/5103c3584b063c431bd1268e9b5e76fb-Paper. pdf.

Shorinwa, O., Mei, Z., Lidard, J., Ren, A.Z., Majumdar, A., 2024. A survey on uncertainty quantification of large language models: Taxonomy, open research challenges, and future directions. ACM Computing Surveys 58, 1–38. doi:10. 1145/3744238.

Sale, Y., Caprio, M., Hüllermeier, E., 2023. Is the volume of a credal set a good measure for epistemic uncertainty?, in: Conference on Uncertainty in Artificial Intelligence (UAI), pp. 1795–1804. URL: https://proceedings. mlr.press/v216/sale23a/sale23a.pdf.

Smith, L., Gal, Y., 2018. Understanding measures of uncertainty for adversarial example detection. arXiv preprint. doi:10.48550/arXiv.1803.08533. Srivastava, N., Hinton, G.E., Krizhevsky, A., Sutskever, I., Salakhutdinov, R., 2014. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research 15, 1929–1958. URL: https://www.jmlr.org/papers/volume15/ srivastava14a/srivastava14a.pdf?utm_content= buffer79b43&utm_medium=social&utm_source= twitter.com&utm_campaign=buffer,.

Schirmer, M., Zhang, D., Nalisnick, E., 2023. Beyond topclass agreement: Using divergences to forecast performance under distribution shift, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–9. URL: https: //openreview.net/pdf?id=DwEENEkkLi. workshop on Distribution Shifts. Schweighofer, K., Aichberger, L., Ielanskyi, M., Hochreiter, S., 2023a. Introducing an improved information-theoretic measure of predictive uncertainty, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 3605– 3640. URL: https://raw.githubusercontent.com/ mlresearch/v286/main/assets/schweighofer25a/ schweighofer25a.pdf. Workshop: Mathematics of modern machine learning.

Steger, S., Knoll, C., Klein, B., Fröning, H., Pernkopf, F., 2024. Function space diversity for uncertainty prediction via repulsive last-layer ensembles, in: International Conference on Machine Learning (ICML), pp. 1–16. URL: https://openreview.net/pdf?id=FbMN9HjgHI. Workshop: Structured Probabilistic Inference & Generative Modeling.

Schweighofer, K., Aichberger, L., Ielanskyi, M., Klambauer, G., Hochreiter, S., 2023b. Quantification of uncertainty with adversarial models, in: Neural Information Processing Systems (NeurIPS), pp. 1–39. URL: https://openreview. net/pdf?id=5eu00pcLWa.

Tagasovska, N., Lopez-Paz, D., 2019. Single-model uncertainties for deep learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https: //dl.acm.org/doi/pdf/10.5555/3454287.3454863. Tassi, C.R.N., Gawlikowski, J., Fitri, A.U., Triebel, R., 2022. The impact of averaging logits over probabilities on ensembles of neural networks, in: International Joint Conferences on Artificial Intelligence (IJCAI), pp. 132–141. URL: https://ceur-ws.org/Vol-3215/19.pdf. workshop: Artificial intelligence safety.

Segonne, P., Zainchkovskyy, Y., Hauberg, S., 2022. Robust uncertainty estimates with out-of-distribution pseudo-inputs training. arXiv preprint. doi:10.48550/arXiv.2201. 05890. Seitzer, M., Tavakoli, A., Antic, D., Martius, G., 2022. On the pitfalls of heteroscedastic uncertainty estimation with probabilistic neural networks, in: International Conference on Learning Representations (ICLR), pp. 1–24. URL: https: //openreview.net/pdf?id=aPOpXlnV1T.

Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., Manning, C.D., 2023. Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback, in: Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5433–5442. doi:10.18653/v1/2023. emnlp-main.330.

Sensoy, M., Kaplan, L., Kandemir, M., 2018. Evidential deep learning to quantify classification uncertainty, in: Conference on Neural Information Processing Systems (NeurIPS), 25

Tresp, V., 2000. A Bayesian committee machine. Neural Computation 12, 2719–2741. doi:10.1162/ 089976600300014908.

Wang, K., Cuzzolin, F., Manchingal, S.K., Shariatmadar, K., Moens, D., Hallez, H., 2024. Credal deep ensembles for uncertainty quantification, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–33. URL: https://openreview.net/pdf?id=PCgnTiGC9K.

Tyralis, H., Papacharalampous, G., 2024. A review of predictive uncertainty estimation with machine learning. Artificial Intelligence Review 57, 1–65. doi:10.1007/ s10462-023-10698-8.

Wang, X., Aitchison, L., 2021. Bayesian OOD detection with aleatoric uncertainty and outlier exposure, in: Symposium on Advances in Approximate Bayesian Inference, pp. 1–10. URL: https://openreview.net/pdf?id= Otz7hLZkLfs.

Ueda, N., Nakano, R., 1996. Generalization error of ensemble estimators, in: International Conference on Neural Networks (ICNN), pp. 90–95. doi:10.1109/icnn.1996.548872.

Wei, H., Xie, R., Cheng, H., Feng, L., An, B., Li, Y., 2022. Mitigating neural network overconfidence with logit normalization, in: International Conference on Machine Learning (ICML), pp. 1–14. URL: https://proceedings.mlr. press/v162/wei22d/wei22d.pdf.

Ulmer, D., Zerva, C., Martins, A.F.T., 2024. Non-exchangeable conformal language generation with nearest neighbors, in: Findings of the Association for Computational Linguistics, pp. 1909–1929. URL: https://aclanthology.org/ 2024.findings-eacl.129.pdf.

Welling, M., Teh, Y.W., 2011. Bayesian learning via stochastic gradient langevin dynamics, in: International Conference on Machine Learning (ICML), pp. 681–688. URL: http:// www.icml-2011.org/papers/398_icmlpaper.pdf.

Vaicenavicius, J., Widmann, D., Andersson, C., Lindsten, F., Roll, J., Schön, T., 2019. Evaluating model calibration in classification, in: International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 3459– 3467. URL: https://proceedings.mlr.press/v89/ vaicenavicius19a/vaicenavicius19a.pdf.

Wen, Y., Jerfel, G., Muller, R., Dusenberry, M.W., Snoek, J., Lakshminarayanan, B., Tran, D., 2021. Combining ensembles and data augmentation can harm your calibration, in: International Conference on Learning Representations (ICLR), arXiv. pp. 1–21. URL: https://openreview.net/pdf? id=g11CZSghXyY.

Valdenegro-Toro, M., 2019. Deep sub-ensembles for fast uncertainty estimation in image classification, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1– 7. URL: https://bayesiandeeplearning.org/2019/ papers/39.pdf. Workshop: Bayesian deep learning.

Wen, Y., Tran, D., Ba, J., 2020. BatchEnsemble: An alternative approach to efficient ensemble and lifelong learning, in: International Conference on Learning Representations (ICLR), pp. 1–19. URL: https://openreview.net/pdf? id=Sklf1yrYDr.

Valdenegro-Toro, M., 2023. Sub-ensembles for fast uncertainty estimation in neural networks, in: International Conference on Computer Vision Workshops (ICCVW), pp. 4119–4127. doi:10.1109/ICCVW60793.2023.00445.

Wild, V.D., Ghalebikesabi, S., Sejdinovic, D., Knoblauch, J., 2023. A rigorous link between deep ensembles and (variational) Bayesian methods, in: Conference on Neural Information Processing Systems (NeurIPS), arXiv. pp. 1–30. URL: https://openreview.net/pdf?id=eTHawKFT4h& noteId=UBYXOwJMPl.

Valiuddin, A., van Sloun, R., Viviers, C., de With, P., van der Sommen, F., 2025. A review of Bayesian uncertainty quantification in deep a review of Bayesian uncertainty quantification in deep probabilistic image segmentation. Transactions on Machine Learning Research , 1–59URL: https: //openreview.net/pdf?id=Yzf4anYwao.

Wilson, A.G., Izmailov, P., 2020. Bayesian deep learning and a probabilistic perspective of generalization, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 4697–4708. URL: https://dl.acm.org/doi/pdf/10. 5555/3495724.3496118.

Van Katwyk, P., Bergen, K., 2025. Hybridflow: Quantification of aleatoric and epistemic uncertainty with a single hybrid model. Transactions on Machine Learning Research , 1–28URL: https://openreview.net/pdf?id= xRiEdSyVjY.

Wimmer, L., Sale, Y., Hofman, P., Bischl, B., Hüllermeier, E., 2023. Quantifying aleatoric and epistemic uncertainty in machine learning: Are conditional entropy and mutual information appropriate measures?, in: Conference on Uncertainty in Artificial Intelligence (UAI), pp. 2282– 2292. URL: https://proceedings.mlr.press/v216/ wimmer23a/wimmer23a.pdf.

Walley, P., 1991. Statistical Reasoning with Imprecise Probabilities. Monographs on Statistics and Applied Probability, Chapman and Hall. Wang, D.B., Feng, L., Zhang, M.L., 2021. Rethinking calibration of deep neural networks: Do not be afraid of overconfidence, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–12. URL: https:// openreview.net/pdf?id=NJS8kp15zzH.

Wood, D., Mu, T., Webb, A.M., Reeve, H.W.J., Luján, M., Brown, G., 2023. A unified theory of diversity in ensemble learning. Journal of Machine Learning Research 24, 1– 26

49. URL: https://www.jmlr.org/papers/volume24/ 23-0041/23-0041.pdf. Xie, J., Ma, Z., Lei, J., Zhang, G., Xue, J.H., Tan, Z.H., Guo, J., 2022. Advanced dropout: A model-free methodology for Bayesian dropout optimization. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 4605–4625. doi:10.1109/TPAMI.2021.3083089. Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., Hooi, B., 2024. Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs, in: International Conference on Learning Representations (ICLR), pp. 1–29. URL: https://openreview.net/pdf?id=gjeQKFxFpZ. Yang, J., Wang, P., Zou, D., Zhou, Z., Ding, K., Peng, W., Wang, H., Chen, G., Li, B., Sun, Y., Du, X., Zhou, K., Zhang, W., Hendrycks, D., Li, Y., Liu, Z., 2022. OpenOOD: Benchmarking generalized out-of-distribution detection, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–14. URL: https:// openreview.net/pdf?id=gT6j4_tskUt. Yoon, T., Kim, H., 2025. Uncertainty estimation by flexible evidential deep learning, in: Conference on Neural Information Processing Systems (NeurIPS), pp. 1–41. URL: https://openreview.net/pdf?id=N6ujq5Yfwa. Zadrozny, B., Elkan, C., 2002. Transforming classifier scores into accurate multiclass probability estimates, in: Conference on Knowledge Discovery and Data Mining (KDD), pp. 694–699. doi:10.1145/775047.775151. Zamyatin, A., Indri, P., Malhotra, S., Gärtner, T., 2026. Is BatchEnsemble a single model? on calibration and diversity of efficient ensembles, in: Conference on Neural Information Processing Systems (EurIPS), pp. 1–10. doi:10.48550/ arXiv.2601.16936. workshop on Epistemic Intelligence in Machine Learning. Zhang, C.B., Jiang, P.T., Hou, Q., Wei, Y., Han, Q., Li, Z., Cheng, M.M., 2021. Delving deep into label smoothing. IEEE Transactions on Image Processing 30, 5984–5996. doi:10.1109/TIP.2021.3089942. Zhang, J., Yin, J., Wang, R., 2020. Basic framework and main methods of uncertainty quantification. Mathematical Problems in Engineering 2020, 1–18. doi:10.1155/2020/ 6068203.

27

Record · ID 414125 · SHA-256 ac049a4577a0b761
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.