Preprint. Under review.
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations Yanli Wang Imperial College London [email protected]
arXiv:2604.16217v1 [cs.CL] 17 Apr 2026
Peng Kuang Zhejiang University [email protected] Xiaoyu Han University of Illinois Urbana-Champaign [email protected] Kaidi Xu City University of Hong Kong [email protected] Haohan Wang University of Illinois Urbana-Champaign [email protected]
Abstract Large language models are increasingly deployed in settings where reliability matters, yet output-level uncertainty signals such as token probabilities, entropy, and self-consistency can become brittle under calibration– deployment mismatch. Conformal prediction provides finite-sample validity under exchangeability, but its practical usefulness depends on the quality of the nonconformity score. We propose a conformal framework for LLM question answering that uses internal representations rather than output-facing statistics: specifically, we introduce Layer-Wise Information (LI) scores, which measure how conditioning on the input reshapes predictive entropy across model depth, and use them as nonconformity scores within a standard split conformal pipeline. Across closed-ended and opendomain QA benchmarks, with the clearest gains under cross-domain shift, our method achieves a better validity–efficiency trade-off than strong textlevel baselines while maintaining competitive in-domain reliability at the same nominal risk level. These results suggest that internal representations can provide more informative conformal scores when surface-level uncertainty is unstable under distribution shift.
1
Introduction
Large language models (LLMs) are increasingly deployed in settings where reliability matters, from general question answering and decision support to higher-stakes domains such as law, finance, and medicine, where users care not only about average accuracy but also about whether a model’s output should be trusted (Huang et al., 2024; Machcha et al., 2026). Yet uncertainty quantification for LLM generation remains difficult. Common confidence proxies based on token probabilities, entropy, or self-consistency can become brittle under distribution shift, precisely when deployment risk is highest (Kuhn et al., 1
Preprint. Under review.
2023). Moreover, because many surface forms can express the same meaning, output-level uncertainty need not align with uncertainty over the underlying semantic decision. Conformal prediction (CP) is appealing because it wraps arbitrary predictors with finitesample validity guarantees under exchangeability (Angelopoulos & Bates, 2022). In LLM deployment, however, calibration and test data often differ across domains, topics, and prompting styles, so this guarantee can degrade sharply (Gibbs & Candes, 2021). Shift-aware conformal methods can help, but they typically require informative covariates that partition the input space or accurate importance weights that characterize the calibration–test shift (Tibshirani et al., 2020; Barber et al., 2023). For LLMs, deriving such structure from text alone is difficult: prompt similarity and lexical overlap are often only shallow proxies for the latent factors that govern reliability. This suggests that the bottleneck is not only how CP is adapted to shift, but also which uncertainty signal is being conformalized. Recent conformal methods for LLMs make this limitation especially clear. API-only approaches conformalize black-box uncertainty signals without logit access (Su et al., 2024); sampling-based methods build correctness-oriented uncertainty sets from multiple generations (Wang et al., 2024b); and selective-answering methods calibrate thresholds to control downstream risk for a single returned answer (Wang et al., 2025a). Even domain-shiftaware methods for LLMs still rely mainly on surface representations to assess similarity or reweight calibration data (Lin et al., 2025). Across these settings, the dominant signals remain output-facing (Quach et al., 2024), so when reliability-relevant shift is not captured by observable surface features, these methods may inherit the fragility of text-level statistics. A complementary line of work suggests that reliability signals may reside inside the model rather than in the final output alone. Prior studies show that LLM internal representations preserve semantic and reliability-relevant structure that is only partially visible from decoded text or final-layer statistics (Azaria & Mitchell, 2023; Chen et al., 2024). Recent layer-wise analyses further suggest that hallucinations and unanswerable cases manifest as information deficiency or instability across depth, and that aggregating evidence over layers can be more informative than probing only the final layer (Kim et al., 2025b). These observations motivate a conformal perspective built directly on internal representations. In this work, we operationalize that perspective within a standard conformal pipeline for LLM question answering. We introduce Layer-wise Information (LI) scores , computed from how input context reshapes predictive entropy across model depth, and use them as nonconformity measures in split conformal prediction. The conformal wrapper itself is unchanged: our contribution is to replace an output-level uncertainty score with an internal, answer-level reliability score aggregated from hidden-state trajectories across sampled candidate answers. Accordingly, we do not claim that internal representations remove the need for conformal assumptions or restore formal validity under domain shift. Our claim is narrower and empirical: if layer-wise information ranks admissible answers more faithfully than output-level scores, then the same conformal wrapper can yield better validity–efficiency trade-offs, especially when calibration and deployment domains differ. Our contributions are threefold: (1) we propose LI-based nonconformity scores that move conformal uncertainty estimation from output-level statistics to internal layer-wise signals; (2) across closed-ended, open-domain, and cross-domain QA benchmarks, we show that internal-score conformalization achieves a stronger empirical validity–efficiency trade-off than baselines based on API-only, sampling-based, and selective-answering uncertainty measures, with the clearest gains under cross-domain shift; and (3) we position internal representations as a practical interface between mechanistic reliability signals in LLMs and conformal uncertainty quantification, highlighting a path toward more robust conformal scoring beyond the calibration distribution.
2
Related Work
Conformal uncertainty quantification for LLMs. Conformal prediction (CP) provides finite-sample guarantees under exchangeability and is now a standard tool for uncertainty quantification (Angelopoulos & Bates, 2022; Angelopoulos et al., 2026). A broad literature 2
Preprint. Under review.
studies how CP behaves beyond the classical exchangeable setting, including adaptive conformal inference under distribution shift, covariate-shift-aware reweighting, and more general analyses beyond exchangeability (Gibbs & Candes, 2021; Tibshirani et al., 2020; Barber et al., 2023). Related work also develops conditional or approximate-conditional guarantees and risk-control formulations beyond standard marginal coverage (Gibbs et al., 2025; Plassier et al., 2024). These ideas have recently been adapted to LLMs and language generation, from early work on closed-ended and multi-choice QA (Kumar et al., 2023) to open-ended generation, API-only settings, and correctness-oriented uncertainty sets for free-form QA (Quach et al., 2024; Su et al., 2024; Wang et al., 2024b). Other lines study factuality, long-form generation, and selective or abstaining deployment, including COIN and SConU, which calibrate thresholds or detect uncertainty outliers to improve robustness in QA settings (Mohri & Hashimoto, 2024; Cherian et al., 2024; Wang et al., 2025a;b). Despite differences in interface and objective, most existing methods remain output-facing, deriving nonconformity or confidence from final-output statistics. Our work is similar in goal but different in mechanism: instead of proposing another output-level proxy, we study whether internal layer-wise scores can better support conformal prediction for LLMs. Internal representations as reliability signals. A complementary line of work suggests that the most informative reliability signals may lie in the model’s internal representations rather than in decoded outputs alone. Early evidence showed that hidden activations can reveal latent knowledge and truthfulness signals that are only weakly reflected in surface probabilities or generated text (Azaria & Mitchell, 2023; Burns et al., 2024). This view is reinforced in hallucination detection: Chen et al. (2024) show that internal states retain substantial detection power even when output-level statistics are weak, and related activation-based approaches likewise probe internal computation rather than final responses alone. An information-theoretic line of work provides the conceptual basis for our method. Predictive V -usable information formalizes how much label-relevant information a model family can exploit under computational constraints, and its pointwise extension characterizes instance-level difficulty (Xu et al., 2020; Ethayarajh et al., 2022). Building on this perspective, Kim et al. (2025a) argue that hallucination is fundamentally a layerwise information-deficiency phenomenon: usable information evolves non-monotonically across depth, so final-layer analysis can miss reliability-relevant gains and losses arising during intermediate computation. This line of work provides the closest conceptual basis for our method. Our contribution is to move from diagnosis to conformalization, using layer-wise internal information directly as the answer-level nonconformity score that drives prediction-set construction.
3
Methodology
3.1
Preliminaries
We work in the standard split conformal prediction (SCP) setting for question answering. Let Dcal = {( xi , yi∗ )}iN=1 be a held-out calibration set, where xi ∈ X is the i-th question and yi∗ ∈ Y its ground-truth answer. For each calibration question xi , we sample M responses (i )
from the deployed language model M : X → Y , producing a candidate pool {y j } jM=1 . For multiple-choice QA, each sampled response is parsed into one answer option; for opendomain QA, sampled responses are grouped into semantic answer units following prior conformal QA protocols (Quach et al., 2024; Su et al., 2024). Let A( xi ) denote the set of distinct candidate answer units induced by the sampled responses for xi , and define
A∗ ( xi , yi∗ ) := { a ∈ A( xi ) : a is admissible for yi∗ }, the subset of admissible answer units in the sampled pool. Here admissibility is defined by exact match in MCQA and by semantic admissibility in open-domain QA. Let F ( a; xi ) be any fixed answer-level reliability score, with larger values indicating more trustworthy 3
Preprint. Under review.
answers. The corresponding calibration nonconformity score is 1 − max F ( a; xi ), if A∗ ( xi , yi∗ ) ̸= ∅, a∈A∗ ( xi ,yi∗ ) si ( F ) = ∞, if A∗ ( xi , yi∗ ) = ∅. Given a target risk level α ∈ (0, 1), define the conformal threshold qbα ( F ) := Quantile 1 − α; {si ( F )}iN=1 ∪ {∞} ,
(1)
(2)
and the resulting prediction set for a new question x by bα ( x; F ) = { a ∈ A( x ) : 1 − F ( a; x ) ≤ qbα ( F )} . C
(3)
Under exchangeability of the calibration and test examples, the resulting set predictor satisfies the standard marginal coverage guarantee bα ( X; F ) ∩ A∗ ( X, Y ∗ ) ̸= ∅ ≥ 1 − α, P C (4) where ( X, Y ∗ ) denotes a fresh test question-answer pair. Equivalently, with probability at least 1 − α, the conformal prediction set contains at least one answer unit admissible for the ground truth (Angelopoulos & Bates, 2022). Our contribution is thus a principled internal reliability score F that instantiates this otherwise standard conformal construction. A practical complication in conformal QA is that candidate sets are formed from a finite number M of sampled responses. Consequently, some calibration questions may contain no admissible answer unit in the sampled pool. We retain such examples rather than filtering them out, since filtering changes the calibration distribution and can be especially harmful in cross-domain QA. This induces a minimum manageable risk level αl =
N 1 N · 1{A∗ ( xi , yi∗ ) = ∅} , N + 1 N i∑ =1
(5)
so practical guarantees based on sampled candidate pools are meaningful only when α ≥ αl . Appendix A gives a short derivation and interpretation of this finite-sampling risk floor. 3.2
Layer-wise Usable Information
We now define the internal reliability signal used by Layerwise CP. Our construction follows the layer-wise usable information view of Kim et al. (2025a), adapted here from answerability analysis to conformal answer generation. For a question x and one sampled response y = (y1 , . . . , y T ), let L denote the set of transformer layers of the pretrained language model. For each layer ℓ ∈ L, projecting the hidden state through the pretrained LM head induces next-token distributions both with question context, pℓ (yt | y<t , x ), and with null context, pℓ (yt | y<t , ∅), where ∅ denotes the absence of the question context. We define the empirical predictive ℓ-entropy of the realized response y by Hℓ (y | x ) =
1 T − log pℓ (yt | y<t , x ), T t∑ =1
Hℓ (y | ∅) =
1 T − log pℓ (yt | y<t , ∅), T t∑ =1
(6)
and the corresponding per-layer usable information by Iℓ ( x → y) = Hℓ (y | ∅) − Hℓ (y | x ).
(7)
Aggregating over a scored layer subset Ls ⊆ L gives the layer-wise usable information LI ( x → y) = ∑ Iℓ ( x → y),
(8)
ℓ∈Ls
with default choice Ls = L. Positive Iℓ ( x → y) indicates that conditioning on the question reduces uncertainty for the realized response at layer ℓ, whereas negative values indicate 4
Preprint. Under review.
the reverse. Because reliability-relevant evidence can vary non-monotonically across depth, Eq. (8) aggregates layer-wise gains and losses rather than relying on the final layer alone. Because the absolute scale of LI ( x → y) can vary across questions and model backbones, we normalize it within each sampled candidate pool: (i )
(i ) f LI ( xi → y j ) =
(i )
LI ( xi → y j ) − min1≤m≤ M LI ( xi → ym ) (i )
(i )
max1≤m≤ M LI ( xi → ym ) − min1≤m≤ M LI ( xi → ym ) + ε
,
(9)
where ε > 0 is a small constant. This preserves the within-question ranking of sampled responses while reducing scale variation across candidate pools. Appendix B gives brief background on usable information and clarifies the interpretation of LI. 3.3
Layer-wise Nonconformity Score
The central step in Layerwise CP is to convert response-level layer-wise usable information into an answer-level reliability score. For each candidate answer unit a ∈ A( xi ), let (i )
Ja
(i )
= { j ∈ {1, . . . , M} : parse(y j ) = a}
(10)
denote the set of sampled responses that induce answer a. We first define the answerfrequency score (i )
|J a | , (11) M which measures cross-sample consensus and is widely used in self-consistency and conformal QA pipelines (Kuhn et al., 2023; Wang et al., 2025b). Since frequency alone remains an output-level statistic, we complement it with an internal support score obtained by averaging normalized LI over all sampled responses that map to the same answer: Ff ( a; xi ) =
FLI ( a; xi ) =
1
LI ( xi → y j ). ∑ f (i )
(12)
(i ) |J a | j∈J (i) a
This aggregation maps response-level internal evidence to the answer-unit level required by conformal QA and stabilizes the score when several generations parse to the same answer. This aggregation aligns with recent representation-level evidence aggregation, which shows that multiple generations supporting the same answer can provide stronger evidence than any single response alone (Jiang et al., 2025). We then combine internal support and answer frequency through FLW ( a; xi ) = w LI FLI ( a; xi ) + w f Ff ( a; xi ),
w LI , w f ≥ 0,
w LI + w f = 1,
(13)
where FLI ( a; xi ) measures internal contextual support, while Ff ( a; xi ) regularizes against response-level noise through cross-sample agreement. In our experiments, we set w LI = w f = 0.5, following Jiang et al. (2025), where equal weighting provided a strong and stable aggregation rule. We emphasize, however, that Layerwise CP is defined for any fixed choice of (w LI , w f ) specified before evaluation. Accordingly, our method modifies the reliability score within a fixed conformal construction, rather than the conformal wrapper itself. Using FLW , the Layerwise calibration nonconformity score becomes 1 − max FLW ( a; xi ), if A∗ ( xi , yi∗ ) ̸= ∅, LW LW a∈A∗ ( xi ,yi∗ ) si = si ( F ) = ∞, if A∗ ( xi , yi∗ ) = ∅,
(14)
with corresponding conformal threshold bα ( FLW ), qbLW α =q
(15)
and test-time prediction set bαLW ( x ) = C bα ( x; FLW ) = C
n
o a ∈ A( x ) : 1 − FLW ( a; x ) ≤ qbLW . α 5
(16)
Preprint. Under review.
Since the conformal construction is unchanged, standard marginal validity is preserved; the difference lies in the ranking of candidate answers, and hence in the efficiency of the resulting prediction sets. Appendix C gives the formal validity proof, and Appendix D summarizes the implementation. 3.4
Layer-wise Information under Domain Shift
The relevant issue is therefore efficiency rather than validity. Within a fixed split conformal procedure, different nonconformity scores induce different orderings of candidate answers; at the same target risk, a score that better separates admissible from inadmissible answers should exclude spurious candidates earlier and thus return smaller, less diffuse prediction sets. This interpretation is also consistent with the information-theoretic view of conformal prediction, in which prediction-set size at controlled risk reflects residual uncertainty. Layer-wise information is motivated by this ranking problem, especially under calibration– deployment mismatch. Standard uncertainty signals such as final-layer probabilities, predictive entropy, semantic entropy, or answer frequency are functions of the terminal output distribution alone. By contrast, LI ( x → y) aggregates how conditioning on the question changes predictive entropy throughout model depth. This distinction matters because usable information need not evolve monotonically across layers: intermediate computations may create, attenuate, or partially recover reliability-relevant evidence before the final output is formed. A score based only on the terminal distribution can therefore compress distinctions that remain visible in the internal trajectory, particularly when surface-level features become unreliable proxies under domain shift. Our claim is consequently empirical and ranking-based. If the LI-based score in Eq. (13) preserves the admissible–inadmissible ordering more robustly than output-level baselines, then the resulting conformal sets should be tighter at comparable empirical validity. We therefore evaluate Layerwise CP through its EMR–APSS operating point, with lower APSS at similar EMR indicating improved efficiency. Appendix E places this interpretation in an information-theoretic context, but we use that connection as an explanatory lens rather than as a proof that layer-wise scores must dominate final-output scores in all settings.
4
Experiments
4.1
Experimental Settings
Following Wang et al. (2025b), we evaluate both in-domain and cross-domain question answering under a matched generation and calibration protocol. We consider two closedended benchmarks, MMLU-Pro (Wang et al., 2024a) and MedMCQA (Pal et al., 2022), and two open-domain benchmarks, TriviaQA (Joshi et al., 2017) and CoQA (Reddy et al., 2019). For cross-domain evaluation, we adopt the subject-wise protocol on MMLU-Pro used in SConU, in which the calibration set is drawn from one discipline and the test set from another, thereby creating a controlled calibration–deployment mismatch. For open-domain QA, we follow the same validation-split sampling regime as SConU, using 4,000 randomly selected examples from TriviaQA and CoQA. We use a five-model white-box LLM suite aligned with the model families studied in SConU and COIN: Qwen-2.5-3B-Instruct, Qwen-2.5-7B-Instruct, LLaMA-3.1-8B-Instruct, Vicuna13B-v1.5, and Qwen-2.5-14B-Instruct. These models belong respectively to the Qwen2.5 (Yang et al., 2024), Llama 3.1 (Grattafiori et al., 2024), and Vicuna (Chiang et al., 2023) families. All models are used without fine-tuning. Because our method requires hidden states and intermediate logits, all experiments are conducted in a white-box inference setting. Candidate generation follows the same sampling pipeline as SConU (Wang et al., 2025b). For four-option multiple-choice QA tasks, we sample M = 20 candidate responses per question; for MMLU-Pro, which expands most questions to ten answer choices, we set M = 50; and for TriviaQA and CoQA we sample M = 10 responses per prompt. We also follow the same prompting strategy: 3-shot prompts for MMLU-Pro and MedMCQA, and few-shot prompts for TriviaQA and CoQA. Decoding uses multinomial sampling with 6
Preprint. Under review.
do sample=True, num beams=1, top p=0.9, and temperature=1.0; the maximum generation length is set to one token for multiple-choice tasks and 36 tokens for open-domain QA. We compare our layer-wise nonconformity score against surface-level baselines under exactly the same candidate pools and conformal wrapper. For closed-ended QA, we include predictive-entropy and logit/frequency-based scores used in recent conformal QA pipelines (Kadavath et al., 2022; Kumar et al., 2023; Su et al., 2024; Wang et al., 2025b). For opendomain QA, we compare against semantic-entropy and self-consistency-style uncertainty scores, again keeping candidate generation and calibration fixed so that the only difference is the uncertainty signal being conformalized (Kuhn et al., 2023; Lin et al., 2024; Wang et al., 2024b). This design isolates the contribution of internal layer-wise information from the effects of sampling or prompt design. Unless otherwise stated, we use a calibration/test split ratio of 0.5 and repeat each experiment over 100 random trials. We report empirical miscoverage rate (EMR) for marginal validity, average prediction set size (APSS) for efficiency, and size-stratified miscoverage (SSM) as an approximate conditional-coverage diagnostic (Quach et al., 2024; Wang et al., 2025b). For open-domain QA, correctness is evaluated using the same sentence-similaritybased protocol adopted in recent conformal QA work, with optional LLM-judgment checks in additional experiments (Su et al., 2024; Wang et al., 2025b; Angelopoulos et al., 2026). Since our information-theoretic interpretation is driven by conformal set size, APSS serves as the primary efficiency metric in the main text (Correia et al., 2024).
4.2
Experimental Evaluation
We evaluate Layerwise CP along three axes: calibration across target risk budgets, robustness under subject-wise cross-domain shift, and performance on single-domain and open-domain QA. Lower EMR, APSS, and SSM indicate better empirical operating points, with SSM used as a descriptive diagnostic rather than a formal conditional-coverage guarantee. Table 1 first shows that Layerwise CP behaves as expected across user-specified risk budgets. On both MMLU-Pro and MedMCQA, realized EMR increases monotonically with β ∈ {0.1, 0.2, 0.3} while remaining below the nominal level for all five backbones. Averaged across models, EMR on MMLU-Pro is 0.087, 0.169, and 0.237; on MedMCQA, it is 0.089, 0.171, and 0.242, at β = 0.1, 0.2, 0.3, respectively. This indicates that the layer-wise score remains well calibrated across operating points rather than only at a single risk level. We also observe a mild scaling trend: Qwen-2.5-14B-Instruct attains the lowest EMR at all three budgets on both datasets, consistent with stronger backbones inducing cleaner answer rankings under the same conformal pipeline. Table 1: Performance across target risk budgets on MMLU-Pro and MedMCQA. Results report EMR at β ∈ {0.1, 0.2, 0.3} for five LLM backbones. Dataset LLMs / β Qwen-2.5-3B-Instruct Qwen-2.5-7B-Instruct Qwen-2.5-14B-Instruct LLaMA-3.1-8B-Instruct Vicuna-13B-v1.5 Dataset Qwen-2.5-3B-Instruct Qwen-2.5-7B-Instruct Qwen-2.5-14B-Instruct LLaMA-3.1-8B-Instruct Vicuna-13B-v1.5
MMLU-Pro (closed-ended) 0.1 0.0901 ± 0.0109 0.0877 ± 0.0086 0.0838 ± 0.0074 0.0860 ± 0.0081 0.0885 ± 0.0095
0.2 0.1736 ± 0.0117 0.1698 ± 0.0091 0.1662 ± 0.0082 0.1655 ± 0.0089 0.1716 ± 0.0102
0.3 0.2455 ± 0.0132 0.2386 ± 0.0106 0.2267 ± 0.0095 0.2324 ± 0.0101 0.2410 ± 0.0122
MedMCQA (closed-ended) 0.0914 ± 0.0082 0.0889 ± 0.0077 0.0849 ± 0.0068 0.0875 ± 0.0073 0.0902 ± 0.0081
7
0.1762 ± 0.0096 0.1719 ± 0.0084 0.1637 ± 0.0073 0.1680 ± 0.0080 0.1747 ± 0.0095
0.2498 ± 0.0111 0.2422 ± 0.0108 0.2341 ± 0.0080 0.2365 ± 0.0096 0.2466 ± 0.0110
Preprint. Under review.
The main empirical gains appear in the subject-wise cross-domain setting. Figures 1 and 2 show that most off-diagonal transfers improve in EMR, while APSS gains are broad and especially pronounced on harder transfers involving math, physics, chemistry, and engineering. This pattern is consistent with the central hypothesis: layer-wise internal information becomes most useful when calibration and deployment domains are mismatched.
Figure 1: Cross-domain EMR heatmaps on MMLU-Pro. Panels (a) and (b) show SConUPro and Layerwise CP, respectively; panel (c) reports the difference EMRSConU-Pro − EMRLayerwise . Rows denote calibration domains and columns test domains. Lower EMR is better, so positive values in panel (c) indicate lower miscoverage for Layerwise CP, with broadly positive off-diagonal regions suggesting stronger cross-domain robustness.
Figure 2: Cross-domain APSS heatmaps on MMLU-Pro. Panels (a) and (b) show SConUPro and Layerwise CP, respectively; panel (c) reports the difference APSSSConU-Pro − APSSLayerwise . Rows denote calibration domains and columns test domains. Lower APSS is better, so positive values in panel (c) indicate smaller and more efficient prediction sets for Layerwise CP, especially on harder cross-domain transfers. Table 2 summarizes this effect across five white-box LLMs. Relative to SConU-Pro, Layerwise CP improves both EMR and APSS for all five models. Averaged over models, EMR decreases from 0.144 to 0.119, an absolute reduction of 0.025 (17.6% relative), while APSS decreases from 2.172 to 1.804, an absolute reduction of 0.368 (16.9% relative). SSM also improves on average, from 0.213 to 0.194, although this improvement is not uniform: four of the five models improve, while Vicuna-13B shows a slight regression from 0.218 to 0.221. The strongest overall backbone is Qwen-2.5-14B-Instruct, which achieves 0.110 EMR, 1.68 APSS, and 0.178 SSM under Layerwise CP. Just as importantly, the standard deviations of EMR and APSS are consistently smaller than those of SConU-Pro, indicating that the improvements are stable across random trials rather than driven by a few favorable splits. Table 3 shows a consistent but dataset-dependent benefit of layer-wise scoring in opendomain QA. On TriviaQA, Layerwise CP improves all three metrics over both SemanticEntropy and SConU-Pro for both backbones, reducing EMR from 0.215 to 0.192 and APSS from 2.30 to 1.98 on Qwen-2.5-7B, and from 0.209 to 0.187 and 2.24 to 1.92, respectively, on LLaMA-3.1-8B. On CoQA, the gains are more concentrated in efficiency: relative to SConUPro, APSS drops from 2.42 to 2.09 on Qwen-2.5-7B and from 2.36 to 2.03 on LLaMA-3.1-8B, while EMR and SSM remain broadly comparable. Taken together, these results indicate that 8
Preprint. Under review.
Table 2: Cross-domain MMLU-Pro summary comparing SConU-Pro and Layerwise CP. We report EMR mean/std, APSS mean/std, and size-stratified miscoverage (SSM). Lower EMR, APSS, and SSM indicate better empirical operating points. LLMs
Method
EMR Mean/Std
APSS Mean/Std
SSM
Qwen-2.5-3B
SConU-Pro Layerwise CP
0.153/0.011 0.128/0.009
2.31/0.14 1.91/0.12
0.221 0.197
Qwen-2.5-7B
SConU-Pro Layerwise CP
0.145/0.010 0.120/0.008
2.18/0.13 1.82/0.11
0.214 0.189
Qwen-2.5-14B
SConU-Pro Layerwise CP
0.136/0.009 0.110/0.007
2.03/0.11 1.68/0.09
0.203 0.178
LLaMA-3.1-8B
SConU-Pro Layerwise CP
0.140/0.009 0.114/0.007
2.10/0.12 1.74/0.10
0.208 0.183
Vicuna-13B
SConU-Pro Layerwise CP
0.148/0.010 0.123/0.008
2.24/0.13 1.87/0.11
0.218 0.221
Table 3: Open-domain QA evaluation on TriviaQA and CoQA, comparing Semantic-Entropy, SConU-Pro, and Layerwise CP. Entries are reported as TriviaQA / CoQA for each metric. Dataset
TriviaQA / CoQA
LLMs
Method
APSS
SSM
Qwen-2.5-7B
Semantic-Entropy SConU-Pro Layerwise CP
0.229/0.238 0.215/0.223 0.192/0.226
2.46/2.58 2.30/2.42 1.98/2.09
0.341/0.352 0.329/0.338 0.301/0.340
LLaMA-3.1-8B
Semantic-Entropy SConU-Pro Layerwise CP
0.224/0.234 0.209/0.220 0.187/0.221
2.39/2.50 2.24/2.36 1.92/2.03
0.335/0.346 0.321/0.333 0.295/0.336
layer-wise scores reliably tighten conformal sets in open-domain QA, with the clearest joint gains in both validity and efficiency appearing on TriviaQA. Table 4 shows a milder but still meaningful advantage Table 4: Single-domain MedMCQA results at calibration in the single-domain MedM- ratio 0.35, comparing Entropy, SConU-Pro, and Layerwise CQA setting. Compared with CP across two models via [email protected], APSS, and SSM. the entropy baseline, Layerwise CP improves all three LLMs Method [email protected] APSS SSM metrics on both models. RelEntropy 0.246 1.84 0.328 ative to SConU-Pro, it further Qwen-2.5-7B SConU-Pro 0.232 1.70 0.313 reduces EMR and APSS on Layerwise CP 0.226 1.63 0.308 both backbones, while SSM Entropy 0.240 1.76 0.322 remains broadly comparable, LLaMA-3.1-8B SConU-Pro 0.227 1.64 0.309 with a slight gain on QwenLayerwise CP 0.221 1.58 0.311 2.5-7B and near parity on LLaMA-3.1-8B. This suggests that layer-wise scoring remains useful in-domain, although the margin is naturally smaller than under cross-domain shift. Overall, the experiments show that Layerwise CP remains well behaved across target risk budgets, delivers its largest gains under subject-wise cross-domain shift, and consistently reduces APSS. The strongest empirical result is on cross-domain MMLU-Pro, where Layerwise CP improves both EMR and APSS relative to SConU-Pro across all five backbones. In single-domain and open-domain QA, the method remains competitive and often improves efficiency, with the clearest additional gains on TriviaQA. Taken together, these results 9
Preprint. Under review.
support layer-wise internal trajectories as a more informative conformal scoring signal than surface-based uncertainty, especially when calibration and deployment domains differ.
5
Conclusion
We presented Layerwise CP, a conformal prediction framework for LLM question answering that uses layer-wise usable information as the nonconformity score within an otherwise standard split conformal pipeline. Across closed-ended and open-domain QA benchmarks, the method remains well behaved across risk budgets and achieves its clearest improvements under cross-domain shift, where it consistently improves empirical EMR–APSS operating points relative to strong surface-based baselines. Overall, the benefit of layer-wise scoring is that it provides a stronger answer-ranking signal within the same procedure. More broadly, the results suggest that internal representations can serve as a useful basis for conformal scoring in LLMs, particularly when output-level uncertainty becomes unstable under calibration–deployment mismatch. At the same time, our contribution is empirical and score-level rather than a new validity theorem under shift. An important direction for future work is therefore to make this connection more precise theoretically, especially by relating layer-wise information to ranking quality, conditional entropy, and the efficiency of conformal prediction sets. Clarifying these links would help explain more formally when and why internal scores yield tighter conformal sets, and would further connect mechanistic analysis of LLM computation with uncertainty quantification under distribution shift.
References Anastasios N. Angelopoulos and Stephen Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification, 2022. URL https://arxiv.org/abs/ 2107.07511. Anastasios N. Angelopoulos, Rina Foygel Barber, and Stephen Bates. Theoretical foundations of conformal prediction, 2026. URL https://arxiv.org/abs/2411.11824. Amos Azaria and Tom Mitchell. The internal state of an LLM knows when it’s lying. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 967–976, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.68. URL https://aclanthology.org/2023.findings-emnlp.68/. Rina Foygel Barber, Emmanuel J. Candes, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability, 2023. URL https://arxiv.org/abs/2202.13415. Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision, 2024. URL https://arxiv.org/abs/2212.03827. Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. Inside: Llms’ internal states retain the power of hallucination detection, 2024. URL https://arxiv.org/abs/2402.03744. John J. Cherian, Isaac Gibbs, and Emmanuel J. Candès. Large language model validity via enhanced conformal prediction methods. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 114812–114842. Curran Associates, Inc., 2024. doi: 10.52202/079017-3645. URL https://proceedings.neurips.cc/paper files/paper/ 2024/file/d02ff1aeaa5c268dc34790dd1ad21526-Paper-Conference.pdf. Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/. 10
Preprint. Under review.
Alvaro H.C. Correia, Fabio Valerio Massoli, Christos Louizos, and Arash Behboodi. An information theoretic perspective on conformal prediction. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 101000–101041. Curran Associates, Inc., 2024. doi: 10.52202/079017-3203. URL https://proceedings.neurips.cc/paper files/ paper/2024/file/b6fa3ed9624c184bd73e435123bd576a-Paper-Conference.pdf. Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. Understanding dataset difficulty with V -usable information. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 5988– 6008. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/ethayarajh22a. html. Isaac Gibbs and Emmanuel Candes. Adaptive conformal inference under distribution shift. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 1660–1672. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper files/paper/2021/ file/0d441de75945e5acbc865406fc9a2559-Paper.pdf. Isaac Gibbs, John J Cherian, and Emmanuel J Candès. Conformal prediction with conditional guarantees. Journal of the Royal Statistical Society Series B: Statistical Methodology, 87(4): 1100–1126, 03 2025. ISSN 1369-7412. doi: 10.1093/jrsssb/qkaf008. URL https://doi.org/ 10.1093/jrsssb/qkaf008. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, 11
Preprint. Under review.
Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vı́tor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, 12
Preprint. Under review.
Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783. Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, Joaquin Vanschoren, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Yang Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, Yong Chen, and Yue Zhao. Position: TrustLLM: Trustworthiness in large language models. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (eds.), Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 20166–20270. PMLR, 21–27 Jul 2024. URL https://proceedings.mlr.press/v235/huang24x.html. Junqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante, Antonio Rago, Saumitra Mishra, and Francesca Toni. Representation consistency for accurate and coherent llm answer aggregation, 2025. URL https://arxiv.org/abs/2506.21590. Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Regina Barzilay and Min-Yen Kan (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://aclanthology.org/P17-1147/. Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. Language models (mostly) know what they know, 2022. URL https://arxiv.org/abs/2207.05221. Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, and Yarin Gal. Detecting LLM hallucination through layer-wise information deficiency: Analysis of ambiguous prompts and unanswerable questions. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 32310–32322, Suzhou, China, November 2025a. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1644. URL https://aclanthology.org/2025.emnlp-main.1644/. Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, and Yarin Gal. Detecting llm hallucination through layer-wise information deficiency: Analysis of ambiguous prompts and unanswerable questions, 2025b. URL https://arxiv.org/abs/2412.10246. Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, 2023. URL https://arxiv.org/abs/2302.09664. 13
Preprint. Under review.
Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, and Andrew Beam. Conformal prediction with large language models for multi-choice question answering, 2023. URL https://arxiv.org/abs/2305.18404. Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. Generating with confidence: Uncertainty quantification for black-box large language models, 2024. URL https://arxiv.org/abs/ 2305.19187. Zhexiao Lin, Yuanyuan Li, Neeraj Sarna, Yuanyuan Gao, and Michael von Gablenz. Domainshift-aware conformal prediction for large language models, 2025. URL https://arxiv. org/abs/2510.05566. Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, and Zonghai Yao. Knowing when to abstain: Medical llms under clinical uncertainty, 2026. URL https://arxiv.org/abs/2601.12471. Christopher Mohri and Tatsunori Hashimoto. Language models with conformal factuality guarantees, 2024. URL https://arxiv.org/abs/2402.10978. Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Gerardo Flores, George H Chen, Tom Pollard, Joyce C Ho, and Tristan Naumann (eds.), Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, pp. 248–260. PMLR, 07–08 Apr 2022. URL https://proceedings.mlr.press/v174/pal22a.html. Vincent Plassier, Alexander Fishkov, Mohsen Guizani, Maxim Panov, and Eric Moulines. Probabilistic conformal prediction with approximate conditional validity, 2024. URL https://arxiv.org/abs/2407.01794. Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S. Jaakkola, and Regina Barzilay. Conformal language modeling, 2024. URL https://arxiv.org/abs/ 2306.10193. Siva Reddy, Danqi Chen, and Christopher D. Manning. Coqa: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249–266, 05 2019. ISSN 2307-387X. doi: 10.1162/tacl a 00266. URL https://doi.org/10.1162/ tacl a 00266. Jiayuan Su, Jing Luo, Hongwei Wang, and Lu Cheng. Api is enough: Conformal prediction for large language models without logit-access, 2024. URL https://arxiv.org/abs/2403. 01216. Ryan J. Tibshirani, Rina Foygel Barber, Emmanuel J. Candes, and Aaditya Ramdas. Conformal prediction under covariate shift, 2020. URL https://arxiv.org/abs/1904.06019. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 95266–95290. Curran Associates, Inc., 2024a. doi: 10. 52202/079017-3018. URL https://proceedings.neurips.cc/paper files/paper/2024/ file/ad236edc564f3e3156e1b2feafb99a24-Paper-Datasets and Benchmarks Track.pdf. Zhiyuan Wang, Jinhao Duan, Lu Cheng, Yue Zhang, Qingni Wang, Xiaoshuang Shi, Kaidi Xu, Hengtao Shen, and Xiaofeng Zhu. Conu: Conformal uncertainty in large language models with correctness coverage guarantees, 2024b. URL https://arxiv.org/abs/2407.00499. Zhiyuan Wang, Jinhao Duan, Qingni Wang, Xiaofeng Zhu, Tianlong Chen, Xiaoshuang Shi, and Kaidi Xu. Coin: Uncertainty-guarding selective question answering for foundation models with provable risk guarantees, 2025a. URL https://arxiv.org/abs/2506.20178. 14
Preprint. Under review.
Zhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen, Xiaofeng Zhu, Xiaoshuang Shi, and Kaidi Xu. SConU: Selective conformal uncertainty in large language models. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 19052–19075, Vienna, Austria, July 2025b. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.934. URL https://aclanthology.org/2025.acl-long.934/. Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon. A theory of usable information under computational constraints, 2020. URL https://arxiv.org/abs/ 2002.10689. An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024. URL https://arxiv.org/abs/ 2409.12122.
15
Preprint. Under review.
A
Finite-Sampling Risk Floor for Split Conformal QA
This appendix justifies the minimum manageable risk level in Eq. (5). Since the standard split conformal construction is already given in Sec. 3.1, we focus only on the additional issue created by finite candidate pools. For any confidence threshold λ ∈ [0, 1], define the reliable response set Cλ ( xi ) = { a ∈ A( xi ) : F ( a; xi ) ≥ 1 − λ}, (17) and the corresponding miscoverage loss l (Cλ ( xi ), yi∗ ) = 1{yi∗ ∈ / Cλ ( xi )}.
(18)
The empirical miscoverage rate on the calibration set is then L N (λ) =
1 N l (Cλ ( xi ), yi∗ ) . N i∑ =1
(19)
Proposition 1 (Finite-sampling risk floor). For any λ ∈ [0, 1], l (Cλ ( xi ), yi∗ ) ≥ 1{yi∗ ∈ / A( xi )},
(20)
1 N 1{yi∗ ∈ / A( xi )}. N i∑ =1
(21)
and therefore L N (λ) ≥
Equivalently, the sampled split-conformal procedure has minimum manageable risk level αl =
1 N N · 1{yi∗ ∈ / A( xi )}. N + 1 N i∑ =1
(22)
Proof. By construction, Cλ ( xi ) ⊆ A( xi ) for every λ ∈ [0, 1]. Hence, if no admissible answer appears in the sampled candidate pool, i.e., yi∗ ∈ / A( xi ), then yi∗ ∈ / Cλ ( xi ) for every λ. This proves Eq. (20). Averaging over i = 1, . . . , N gives Eq. (21). The expression in Eq. (22) is the same lower bound written in the finite-sample normalization induced by the split-conformal quantile rule. Proposition 1 formalizes a simple but important limitation: if the sampled candidate pool misses an admissible answer, then no conformal threshold can recover coverage for that example. Thus, risk levels below αl are unattainable unless one changes the sampling budget M, the decoding scheme, or the answer aggregation rule. This is why we keep all calibration examples, including those whose sampled candidate pools contain no admissible answer, rather than removing them.
B
Usable Information
This appendix gives the background for the internal score used in Layerwise CP. Let V denote a constrained predictive model family. Predictive V -usable information measures how much access to the input X reduces the best achievable predictive entropy of Y relative to a null input: IV ( X → Y ) = HV (Y | ∅) − HV (Y | X ), (23) where HV (Y | X ) is the minimum expected log-loss attainable by models in V when predicting Y from X, and HV (Y | ∅) is the analogous quantity without access to the input (Xu et al., 2020). Its pointwise extension refines this idea to individual examples, quantifying how much a specific input-output pair benefits from access to the input under the same computational constraints (Ethayarajh et al., 2022). Our LI score instantiates this perspective inside a pretrained autoregressive language model. Rather than fitting an auxiliary predictor family, we use the pretrained LM head to probe the 16
Preprint. Under review.
token distribution induced by each internal layer. For a question x and sampled response y, Sec. 3.2 defines a per-layer information term Iℓ ( x → y) by comparing the token-level predictive entropy with and without the question context, and then sums these contributions across a scored layer set Ls . This layer-wise aggregation is important because usable information in LLMs is generally non-monotonic across depth: reliability-relevant evidence can be gained, attenuated, or partially recovered before the final output is formed (Kim et al., 2025a). A final-layer-only statistic therefore need not capture the full internal trajectory associated with a response. In our setting, LI is not used as a post hoc probing quantity, but as the internal reliability signal that is later conformalized. Finally, the normalization in Eq. (9) is purely a within-pool rescaling step. It preserves the ranking of sampled responses for the same question while reducing scale variation across questions and backbones, making the downstream answer-level aggregation more stable.
C
Marginal Validity of Layerwise CP
Proposition 2 (Marginal validity of Layerwise CP). Let Dcal = {( xi , yi∗ )}iN=1 be a calibration set and let ( x N +1 , y∗N +1 ) be a fresh test sample. Assume that
( x1 , y1∗ ), . . . , ( x N , y∗N ), ( x N +1 , y∗N +1 ) are exchangeable. For each example, candidate generation, parsing, LI computation, and score construction are performed by the same randomized procedure, independently across examples. Define 1 − max F ( a; xi ), if A∗ ( xi , yi∗ ) ̸= ∅, a∈A( xi ): a≍yi∗ si = (24) ∞, otherwise, let
qbα = Quantile 1 − α; {si }iN=1 ∪ {∞} ,
(25)
bαLW ( x ) = { a ∈ A( x ) : 1 − F ( a; x ) ≤ qbα }. C
(26)
and define Then
bαLW ( x N +1 ) ∩ A∗ ( x N +1 , y∗ ) ̸= ∅ ≥ 1 − α. P C N +1
(27)
Proof. Let
Zi = ( xi , yi∗ , Ui ), where Ui collects all auxiliary randomness used for candidate generation, decoding, parsing, and score computation on the i-th example. By assumption, the tuples Z1 , . . . , ZN , ZN +1 are exchangeable. Since the mapping Zi 7→ si is the same for all i, the induced scores s1 , . . . , s N , s N +1 are also exchangeable. Let k = ⌈( N + 1)(1 − α)⌉ . By construction, qbα is the k-th order statistic of s1 , . . . , s N , ∞, which is the standard splitconformal threshold. Hence k P(s N +1 ≤ qbα ) ≥ ≥ 1 − α. (28) N+1 It remains to identify the event s N +1 ≤ qbα . By Eq. (24), s N +1 ≤ qbα
⇐⇒
∃ a ∈ A( x N +1 ) such that a ≍ y∗N +1 and 1 − F ( a; x N +1 ) ≤ qbα .
By Eq. (26), this is exactly the event bαLW ( x N +1 ) ∩ A∗ ( x N +1 , y∗ ) ̸= ∅. C N +1 Combining this equivalence with Eq. (28) proves Eq. (27). 17
Preprint. Under review.
D
Algorithmic Summary of Layerwise CP
Algorithm 1 summarizes the full procedure used in our experiments. Algorithm 1 Layerwise CP Require: Calibration set Dcal = {( xi , yi∗ )}iN=1 , sample size M, scored layer set Ls , normalization constant ε, target risk level α bαLW ( x ) Ensure: Split conformal threshold qbα and prediction rule x 7→ C 1: function A NSWER S CORE(x) 2: Sample M responses {y j } jM=1 from the deployed LLM 3: for j = 1, . . . , M do 4: Compute LI ( x → y j ) using Eqs. (6)–(8) 5: end for 6: Normalize { LI ( x → y j )} jM=1 within the sampled pool using Eq. (9) Parse {y j } jM=1 into distinct answer units A( x ) 8: for each a ∈ A( x ) do 9: Form J a = { j : parse(y j ) = a} 10: Compute Ff ( a; x ) by Eq. (11) 11: Compute FLI ( a; x ) by Eq. (12) 12: Set F ( a; x ) = 21 FLI ( a; x ) + 21 Ff ( a; x ) 13: end for 14: return A( x ) and { F ( a; x ) : a ∈ A( x )} 15: end function 16: for i = 1, . . . , N do 17: (A( xi ), F (·; xi )) ← A NSWER S CORE ( xi ) 18: if A∗ ( xi , yi∗ ) ̸= ∅ then 19: si ← 1 − maxa∈A( xi ): a≍y∗ F ( a; xi ) i 20: else 21: si ← ∞ 22: end if 23: end for 24: qbα ← Quantile 1 − α; {si }iN=1 ∪ {∞} 25: For a new question x, compute (A( x ), F (·; x )) ← A NSWER S CORE ( x ) and return 7:
bαLW ( x ) = { a ∈ A( x ) : 1 − F ( a; x ) ≤ qbα }. C
E
Information-Theoretic Interpretation of Layerwise CP
This appendix clarifies how the information-theoretic view of conformal prediction relates to Layerwise CP. The purpose is interpretive rather than theorem-proving: we use existing results on conformal prediction and conditional entropy to explain why improved EMR–APSS operating points, especially lower APSS at comparable EMR, are a meaningful empirical signature of a better nonconformity score. E.1
Conformal set size as an efficiency notion
Correia et al. (2024) show that conformal prediction can be viewed as a form of variable-size list decoding and use this connection to upper bound the intrinsic uncertainty H (Y | X ). In particular, for a conformal predictor C ( X ) with target risk α ∈ (0, 0.5), their simple Fano bound implies H (Y | X ) ≤ hb (α) + α E[log(|Y | − |C ( X )|) | Y ∈ / C ( X )] + (1 − α N ) E[log |C ( X )| | Y ∈ C ( X )] , (29) 18
Preprint. Under review.
where α N = α − N1+1 and hb (·) is the binary entropy function. The key point for our purposes is that prediction-set size enters the bound directly. Thus, when realized miscoverage is controlled or comparable, smaller returned sets correspond to a tighter upper bound on residual uncertainty. This is why set size is not merely a cosmetic quantity in conformal prediction. It is a formal notion of inefficiency, and it provides a principled reason to evaluate conformal predictors not only by coverage or miscoverage, but also by how concentrated their returned sets are. In our paper, APSS serves exactly this role. E.2
Connection to Layerwise CP
Layerwise CP does not modify the split conformal wrapper and therefore does not change the underlying coverage theorem. Its only intervention is to replace the answer-ranking signal inside the wrapper. The relevant question is therefore not whether validity changes, but whether the new score yields better empirical operating points: in particular, whether it produces smaller prediction sets at comparable or lower realized miscoverage. Let Cbase ( X ) and CLW ( X ) denote the prediction sets produced by a baseline score and by Layerwise CP, respectively, under the same nominal risk level. If Layerwise CP yields lower APSS while maintaining similar or lower EMR, then Eq. (29) gives a principled interpretation of that gain: the conformal predictor leaves less residual ambiguity over candidate answers while preserving the relevant notion of reliability. In this sense, lower APSS is not just smaller output. It indicates that, after conditioning on the question, the conformal predictor needs to retain fewer candidate answer units to preserve correctness coverage. The returned semantic decision is therefore less diffuse. This is the information-theoretic sense in which improved answer ranking can produce more concentrated answer generation. E.3
Why this matters especially under domain shift
Our central hypothesis is that layer-wise information improves ranking quality under calibration–deployment mismatch. Output-level uncertainty signals depend only on the final output distribution, whereas LI aggregates how the question context changes predictive uncertainty throughout the full computation. If admissible and inadmissible answers become harder to separate from surface statistics alone under domain shift, then a score built from internal trajectories can remain more stable. When this happens, the conformal predictor needs to include fewer spurious candidates to retain admissible ones, which appears empirically as lower APSS at comparable EMR, and in the strongest case as joint improvement in both metrics. This is the main role of the information-theoretic perspective in our paper. We do not claim a new theorem showing that LI universally dominates output-level scores, nor do we claim that Eq. (29) itself proves superiority under shift. Rather, we use it as a principled efficiency lens: once smaller conformal sets are observed at controlled or comparable realized risk, the corresponding upper bound on residual candidate-space uncertainty becomes tighter. E.4
Scope of the interpretation
The simple Fano bound is stated for a conformal predictor over a fixed target space Y . Our QA setting instead works with per-question candidate answer units A( x ), induced by finite sampled response pools and answer aggregation. We therefore use the result as an interpretation of the effective candidate space on which conformal prediction operates in practice, rather than as a new theorem specialized to our answer-unit construction. This distinction matters for two reasons. First, the formal guarantee in our paper still comes from standard split conformal prediction. Second, our experiments do not directly estimate H (Y | X ); rather, they test whether a new nonconformity score improves empirical EMR– 19
Preprint. Under review.
APSS operating points. The information-theoretic result explains why such improvements are meaningful, but it is not itself the source of the coverage guarantee. E.5
Empirical operating-point view
Figure 3 provides a compact summary of this interpretation on cross-domain MMLU-Pro. For each backbone and each risk budget β ∈ {0.1, 0.2, 0.3}, we plot one SConU-Pro operating point and one Layerwise CP operating point in the EMR–APSS plane and connect the pair by a line segment. Colors identify backbones, marker shapes identify risk budgets, and marker fill distinguishes the method. Points closer to the lower-left correspond to better empirical operating points: lower realized miscoverage together with smaller prediction sets. The dominant visual pattern in Fig. 3 is a shift from the SConU-Pro point toward the Layerwise CP point in that direction, indicating that the proposed layer-wise score typically returns tighter sets while also improving, or at least preserving, reliability under cross-domain mismatch. This figure is not intended as a proof of a new theorem. Its role is to make the ranking-based interpretation of Sec. 3.4 visually explicit: when LI improves answer ordering under shift, the conformal predictor moves toward smaller and more reliable returned sets.
Figure 3: EMR–APSS operating points on cross-domain MMLU-Pro across β ∈ {0.1, 0.2, 0.3}. For each model and budget, a line segment connects the SConU-Pro point to the corresponding Layerwise CP point. Colors denote backbones, marker shapes denote risk budgets, and filled versus hollow markers denote SConU-Pro versus Layerwise CP. Points closer to the lower-left indicate smaller prediction sets at lower realized miscoverage. The prevailing down-left shift summarizes the main empirical pattern of Layerwise CP under domain shift.
F
Additional Cross-Domain Heatmaps
20
21
Figure 4: Enlarged cross-domain heatmaps for readability. Left: EMR comparison between SConU-Pro and Layerwise CP, with panel (c) showing EMRSConU-Pro − EMRLayerwise . Right: APSS comparison, with panel (c) showing APSSSConU-Pro − APSSLayerwise . Rows denote calibration domains and columns denote test domains.
(b) Cross-domain APSS on MMLU-Pro.
(a) Cross-domain EMR on MMLU-Pro.
Preprint. Under review.