Conceptio › Archive › arXiv CS
arXiv CSopen access

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention Andre Bacellar [email protected]

arXiv:2609.22056v1 [cs.IR] 18 Sep 2026

Abstract

1

Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility, Theorem 1): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success — a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUCAC gap between regimes. Second (Feature Regime Complementarity, Proposition 1): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer), and a constructive witness pair shows each is necessary in one regime and non-contributory in the other. We instantiate these principles in R EGIME A BSTAIN, which computes a Retrieval Confidence Score (RCS) — a logistic function of up to nine queryANN structural features, all available without any additional LLM call — and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLMjudge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE = 0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only −0.5pp AUC loss, confirming the domain-agnostic structure of regime features.

tem returns its best-matching top-k passages and feeds them to a language model. This design is sensible when the goal is to maximize recall across a broad query distribution. It is inappropriate when retrieval quality is heterogeneous and highconfidence errors are costly. Multi-hop QA is precisely such a setting: a query requires retrieving both a bridge passage and an answer passage, and the system frequently succeeds on one while missing the other. On MuSiQue with our LLM-judge pipeline (Bacellar, 2026a), 39.5% of test queries fail to retrieve all gold passages into the top-5. These are not low-confidence retrievals—the system returns a ranked list regardless, with no signal to downstream components that the evidence is incomplete. We argue that a retrieval system operating on heterogeneous queries should support a third action: abstention (⊥). When the system predicts low retrieval quality, it can flag the query rather than silently returning an incomplete context. Abstracting this into a calibrated score opens the door to query routing, graceful degradation, and human-in-the-loop escalation—all without modifying the underlying retrieval pipeline. We study what makes a query hard at retrieval time, without using gold labels. The insight comes from the regime-conditional framework of Bacellar (2026b): queries that sit near the transition between Q-dominant and B-dominant retrieval regimes are structurally ambiguous. Lightweight features from the ANN score distribution prove particularly predictive: hop-1 lift (how strongly the top ANN passage dominates the rest), hop-2 max score (quality of the expanded query search), and query length (longer questions correlate with harder multi-hop chains).

Introduction

Contributions.

Retrieval-augmented generation (RAG) pipelines always produce an answer: given a query, the sys-

1. We prove that CWAR is reducible if and only if retrieval features carry mutual information

about retrieval success, formally characterizing when abstention can help (Theorem 1). 2. We show empirically that no single ANN score feature achieves best predictive performance across all failure regimes (Proposition 1): a constructive witness pair exhibits opposite necessity/redundancy patterns across MuSiQue and HoVer, motivating multi-feature aggregation. 3. We define the Confident-Wrong-Answer Rate (CWAR) metric and introduce RCS, a logistic function of nine query-time features that is calibrated (ECE = 0.035) and requires no additional LLM inference. 4. We evaluate RCS against eight confidence baselines across five failure regimes spanning three datasets and two retrieval architectures, establishing consistent best or co-best AUCAC in all conditions. 5. We show cross-dataset transfer (−0.5pp AUC) and identify regime-specific dominant features: query length on MuSiQue, hop-1 concentration on HoVer.

2

Background and Related Work

Selective prediction. The study of abstaining classifiers dates to Chow (1970). Geifman and El-Yaniv (2017) formalize the accuracy-coverage trade-off: at coverage c, the system answers only the fraction c of queries it is most confident about, and accuracy among answered queries should increase as c decreases. Kamath et al. (2020) apply this framework to open-domain QA, training a confidence predictor on top of a QA model’s internal states. Our setup differs: we predict retrieval success (do the retrieved passages contain the gold answers?), not answer generation accuracy, using only pre-LLM features.

Conformal prediction for NLP. Conformal prediction (Angelopoulos et al., 2022) provides coverage guarantees under exchangeability. We include a split-conformal baseline (Section 6) and show that RCS outperforms it on the AUCAC metric while offering a continuous score for operating-point selection. Score-distributional signals for retrieval abstention. Concurrent work by Holdcroft et al. (2026) shows, on single-hop semantic, logical, and temporal benchmarks, that thresholding raw similarity magnitude is a poor abstention signal and that zero-cost distributional statistics of the retrieved score list (score gap, magnitude-andvariance) improve abstention AUROC by up to 0.16. We reach the same conclusion from a different direction and extend it in four ways: the setting is multi-hop, where failure concentrates in structurally identifiable regimes; the signals are combined by a learned model whose output is a calibrated probability (ECE = 0.035) rather than used as raw rankers; the analysis covers an LLM-judge pipeline as well as dense retrieval; and Theorem 1 gives the condition under which any such signal can reduce confident failure at all. Regime-conditional retrieval. Bacellar (2026b) show that multi-hop queries divide into two regimes: Q-dominant (bridge entity contained in the question) and B-dominant (bridge entity must be discovered from the hop-1 result). R EGIME A BSTAIN uses related features for a different task: predicting retrieval success rather than choosing a pipeline, providing a complementary abstention gate for the same regime-aware stack.

3

Problem Formulation

3.1

Confident-Wrong-Answer Rate (CWAR)

Let Q be a query set, R : q → {p1 , . . . , pk } a retrieval system, and G(q) the set of gold passages for query q. Define full success: y(q) = 1[G(q) ⊆ {p1 , . . . , pk }]

LLM uncertainty. Post-generation uncertainty is well-studied: Kadavath et al. (2022) show that large LLMs can self-predict answer accuracy; Kuhn et al. (2023) propose semantic entropy over sampled outputs. These require a full generation pass per query. RCS is computed in < 1ms from ANN scores available at retrieval time— no LLM call is needed for the abstention decision.

i.e., 1 if all gold passages appear in the top-k result, 0 otherwise.1 Given a confidence score s(q) ∈ [0, 1] and threshold τ , define the confident set Cτ = {q : 1 For MuSiQue and 2Wiki each query has two supporting passages. Full success therefore requires both to be in the top-5. HoVer queries have 2–4 supporting passages.

s(q) ≥ τ }. The CWAR at threshold τ is: |{q ∈ Cτ : y(q) = 0}| |Cτ | = 1 − Pr[y(q) = 1 | s(q) ≥ τ ] .

CWAR(τ ) =

When τ = 0 (answer every query), CWAR(0) = 1 − ȳ equals the base failure rate. As τ increases, coverage Cov(τ ) = |Cτ |/|Q| decreases; a good confidence score makes CWAR(τ ) decrease faster than Cov(τ ). The accuracy-coverage curve plots Pr[y = 1|s ≥ τ ] against Cov(τ ) as τ varies, and its area (AUC-AC) summarises performance across all operating points. 3.2

Retrieval Confidence Score (RCS)

We model s(q) = σ(w⊤ ϕ(q) + b) where σ is the logistic function and ϕ(q) ∈ Rd is a feature vector computed from ANN scores available after retrieval. d = 9 for pipelines with hop-2 expansion; d = 6 for hop-1-only pipelines (Section 5). 3.3

Calibration

RCS is calibrated if E[y(q) | RCS(q) = p] = p for all p. We measure calibration by Expected Calibration Error (ECE) with equal-frequency bins: ECE =

B X |Bb | b=1

n

|ȳBb − s̄Bb |

where Bb is the b-th bin.

4

Theoretical Properties of CWAR

4.1

CWAR Reducibility

The central deployment question is: when can confident-failure reduction be achieved at all? We show it is governed by the mutual information between retrieval features and retrieval success. Theorem 1 (CWAR Reducibility). Let y(q) ∈ {0, 1} be the full-success indicator and ϕ(q) the feature vector. Define the Bayes-optimal score s∗ (q) = P [y = 1 | ϕ(q)]. Then AUC-AC(s∗ ) > AUC-ACrandom if and only if I(y; ϕ(q)) > 0. Moreover, AUC-AC(s∗ ) ≥ AUC-AC(s) for any score s based on ϕ. Proof. (⇒, contrapositive) If I(y; ϕ) = 0 then y ⊥ ϕ, so s∗ (q) = ȳ for all q. Every threshold produces the same coverage and accuracy ȳ; the accuracy-coverage curve is flat and AUCAC(s∗ ) = ȳ = AUC-ACrandom .

(⇐) If I(y; ϕ) > 0 then Var(s∗ ) > 0. By the law of iterated expectation, E[s∗ ] = ȳ, so there exists a positive-measure set of queries with s∗ (q) > ȳ. Thresholding at any θ ∈ (ȳ, 1) concentrates the confident set on above-average queries: P [y = 1 | s∗ ≥ θ] > ȳ. The accuracy-coverage curve is strictly above ȳ for some coverage, so AUC-AC(s∗ ) > ȳ = AUC-ACrandom . The optimality of s∗ follows from the NeymanPearson lemma: any monotone threshold rule achieves maximum accuracy at a given coverage when ranking by P [y = 1 | ϕ], so no other function of ϕ can produce a higher accuracy-coverage curve. Empirical consequence. The theorem predicts the AUC-AC ordering across the five conditions. Dense-pipeline failures are nearly uniformly distributed across feature space (MuSiQue Dense: AUC-AC = 0.556 vs. random = 0.366, small gap); the LLM judge concentrates failures in structurally complex queries, increasing I(y; ϕ) and making failures more predictable (MuSiQue PropH: AUC-AC = 0.790 vs. random = 0.576, large gap). High CWAR alone does not guarantee reducibility: the dense conditions have the highest base CWAR (62.1%, 58.3%) yet the weakest AUC-AC. 4.2

Feature Regime Complementarity

The feature importance divergence in Table 4 reflects a structural property of multi-hop failure, not a dataset quirk. Proposition 1 (Feature Regime Complementarity). No single feature in {ϕ1 , . . . , ϕ9 } achieves best AUC-AC across all five failure regimes. Furthermore, there exist features ϕi ̸= ϕj such that ϕi is necessary in regime Ra and non-contributory in Rb , while ϕj exhibits the opposite pattern — the dominant predictive feature is regime-specific. Proof. First claim. Table 1 establishes this by exhaustion: no single baseline achieves best AUCAC in more than two of the five conditions. Querylen-inv achieves 0.947 on 2Wiki PropH but only 0.426 on MuSiQue Dense; Max-score achieves 0.839 on HoVer Dense but only 0.480 on 2Wiki Dense. Second claim (constructive). Leave-one-out ablation in Table 4 gives: QUERY- LEN is necessary on MuSiQue PropH (∆AUC = −0.012) and noncontributory on HoVer Dense (∆AUC = −0.002,

within noise); HOP 1- TOP 3 is necessary on HoVer Dense (∆AUC = −0.026) and non-contributory on MuSiQue PropH (∆AUC = −0.001). This pair forms the constructive witness. Note that entropy features (HOP 1-H, HOP 2-H) are non-contributory in both conditions (∆AUC ≥ 0), consistent with the Entropy baseline’s weak performance across regimes; the proposition does not claim universality for all nine features, only that the complementary structure exists. Consequence for confidence model design. Proposition 1 provides a principled motivation for multi-feature aggregation: any confidence model that uses only a single feature incurs avoidable AUC-AC loss on whichever regime that feature does not dominate. A logistic model trained jointly across regimes learns the complementary weighting — QUERY- LEN dominates on MuSiQue while HOP 1- TOP 3 dominates on HoVer — achieving regime-universal coverage that no single baseline can match.

• ϕ7 (HOP 2- MARGIN): gap between the top-2 hop-2 scores. (2)

• ϕ8 (HOP 2-H): normalised entropy of {sj }. • ϕ9 (QUERY- LEN): word count of the question. 5.2

Training

Given a held-out validation set V with gold labels y(q), we fit w, b by minimising binary cross-entropy with gradient descent until convergence (tolerance 10−7 ). Features are z-score normalised. Predictions on the test set use s(q) = σ(w⊤ ϕ̂(q) + b) where ϕ̂ uses validation-set normalisation statistics. 5.3

Threshold Selection

5

R EGIME A BSTAIN

Given a user-specified CWAR target γ ∈ (0, 1), we sweep τ over the validation set and select τ ∗ = min{τ : CWAR(τ ) ≤ γ}. Because CWAR(τ ) is monotone non-increasing in τ and equals 1 − ȳ at τ = 0 and 0 at τ = 1 (empty confident set), the desired τ ∗ always exists for any γ ≥ 0.

5.1

Feature Extraction

6

All features are computed from ANN similarity scores available after retrieval— no additional model inference is required. (1) Let {(ui , si )}K i=1 be the top-K hop-1 passage scores sorted in decreasing order, and (2) {(uj , sj )}j the hop-2 scores obtained by embedding N = 3 SVO-extracted bridge queries and taking the maximum score per passage. When hop-2 scores are unavailable (hop-1-only architectures), ϕ6 , ϕ7 , ϕ8 are dropped, giving d = 6. (1)

• ϕ1 (HOP 1- MAX): s1 , the maximum hop-1 score. (1)

(1)

• ϕ2 (HOP 1- MARGIN): s1 − s2 , the gap between the top-2 hop-1 passages. • ϕ3 (HOP 1- TOP 3): mean of the top-3 hop-1 scores. • ϕ4 (HOP 1-H): normalised Shannon entropy (1) of {si }. (1)

• ϕ5 (HOP 1- LIFT): s1 /s(1) 50 , peak prominence relative to the top-50 mean. (2)

• ϕ6 (HOP 2- MAX): maxj sj .

Experimental Setup

Datasets. We evaluate on three multi-hop benchmarks. MuSiQue (Trivedi et al., 2022) contains compositional 2–4-hop questions over Wikipedia; we use 1000 queries (486 tune, 514 test). 2WikiMultiHopQA (Ho et al., 2020) contains bridge and comparison questions; we use 1000 queries (509 tune, 491 test). HoVer (Shi et al., 2020) contains 4000 multi-hop factverification claims over Wikipedia; we use the 2000 SUPPORTED claims (1007 tune, 993 test). All splits use a deterministic MD5-hash 50/50 partition on query id. Retrieval pipelines. tectures:

We evaluate on two archi-

• Proposal H (LLM-judge) (Bacellar, 2026a): NV-Embed-v2 hop-1 ANN, N = 3 SVOexpanded hop-2 queries, 20-candidate pool, 3-way LLM judge, α = 0.10. Achieves R@5 = 0.8138 on MuSiQue and 0.9527 on 2Wiki. Full-success rate: MuSiQue 60.5%, 2Wiki 85.5%. Uses 9-feature RCS (d = 9). • Dense-only: hop-1 top-5 directly (no LLM judge, no SVO re-ranking). Applied to all three datasets; HoVer uses d = 6

1. Random: uniform Uniform[0, 1]. (1)

2. Entropy: 1 − H({si }). (1)

3. Max-score: s1 . (1)

(1)

4. Margin: s1 − s2 . (1)

MuSiQue (LLM-judge)

2Wiki (LLM-judge)

1/(1 + |q|/20) (inverse

7. Temp-scaled: {0.5, 1, 2, 5, 10} split.

s1 with T ∈ selected on the tune

(1) 1/T

8. MLP: two-layer network (32-32 hidden units) trained on the same 9 (or 6) features. We additionally evaluate a split-conformal baseline using hop-1 max-score as non-conformity; results at discrete coverage targets are reported separately since AUC-AC is not directly applicable. Metrics. Primary: AUC of the accuracycoverage curve (AUC-AC) on the test split, with 95% bootstrap CIs (2000 samples). Secondary: CWAR and accuracy at fixed coverage targets (70%, 50%); ECE (10 equal-frequency bins) and Brier score for calibration (MuSiQue only).

7

Results

7.1

Main Results: AUC-AC Across Conditions

Table 1 reports AUC-AC for all nine methods across five failure-regime conditions. RCS achieves best or co-best AUC-AC in all five conditions. On MuSiQue (LLM-judge), RCS achieves AUC-AC = 0.790, surpassing the next-best nonlearned method Lift (0.776, +1.4pp) and Tempscaled (0.774, +1.6pp); entropy-only 0.743 places behind Max-score 0.751. On 2WikiMultiHopQA, RCS and Query-len-inv are tied (0.947), both far above entropy (0.849) and random (0.841)— entropy adds only +0.8pp over random, while

MuSiQue (Dense)

2Wiki (Dense)

AUC-AC

0.6

Ran do Entr m Max opy -sco re Marg in QLe Lift Tem n-inv p-sca led MLP RCS

Ran do Entr m Max opy -sco re Marg in QLe Lift Tem n-inv p-sca led MLP RCS

Ran do Entr m Max opy -sco re Marg in QLe Lift Tem n-inv p-sca led MLP RCS

0.4

Figure 1: AUC-AC (test split) across five failureregime conditions for nine confidence methods. RCS (rightmost, dark) is best or co-best in all five columns. Random baseline (leftmost, gray) varies because base full-success rates differ by regime. Table 1: AUC-AC (test split) across five failureregime conditions. Bold = best in column; † = within bootstrap-CI overlap of best. Base CWAR = 1− fullsuccess rate.

5. Lift: s1 /s(1) 50 (single-feature). 6. Query-len-inv: length).

HoVer (Dense)

0.8

Ran do Entr m Max opy -sco re Marg in QLe Lift Tem n-inv p-sca led MLP RCS

Baselines. We compare eight confidence signals against RCS:

RCS achieves best or co-best AUC-AC across all five failure regimes 1.0

Ran do Entr m Max opy -sco re Marg in QLe Lift Tem n-inv p-sca led MLP RCS

(no hop-2 scores available). Full-success rates: MuSiQue 37.9%, 2Wiki 41.7%, HoVer 68.3%.

LLM-judge pipeline Method

MuSiQue

Dense pipeline

2Wiki HoVer MuSiQue 2Wiki

Base CWAR

39.5%

14.5% 31.7%

62.1% 58.3%

Random Entropy Max-score Margin Lift Query-len-inv Temp-scaled MLP RCS (ours)

0.576 0.743 0.751 0.724 0.776 0.751 0.774 0.742 0.790

0.841 0.661 0.849 0.741 0.912 0.839 0.893 0.730 0.902 0.825 0.947† 0.742 0.912 0.839 0.928 0.878† 0.947 0.873†

0.366 0.535 0.510 0.412 0.525 0.426 0.510 0.557† 0.556†

0.402 0.490 0.480 0.420 0.571 0.526 0.480 0.643 0.649

RCS gains +10.6pp. On HoVer (dense), MLP (0.878) and RCS (0.873) are within bootstrapCI overlap; both dominate all non-learned methods by at least +3.4pp. On the dense MuSiQue and 2Wiki conditions, RCS and MLP are again within CI overlap, with RCS leading 2Wiki dense (0.649 vs 0.643). No single non-learned heuristic is consistently competitive: entropy underperforms max-score and lift across all five conditions. 7.2

Operating Points (MuSiQue LLM-Judge)

Table 2 shows selected operating points on MuSiQue test. At 50% coverage, RCS achieves 79.4% accuracy: CWAR drops from 39.5% (baseline) to 20.6% (47.8% relative reduction). At 70% coverage, accuracy is 75.1% (+14.6pp over base). 7.3

Architecture Robustness

A key question for deployment is whether RCS generalizes across retrieval architectures. In the dense-only setting, the absence of an LLM judge substantially increases failure rates (MuSiQue: 62.1% CWAR vs. 39.5% for Proposal H; 2Wiki: 58.3% vs. 14.5%). Despite this regime shift, RCS retains its relative advantage: on both dense conditions it achieves best or co-best AUC-AC. Abso-

Accuracy

1.0

MuSiQue (LLM-judge): Accuracy vs Coverage

Table 3: Calibration on MuSiQue test (n = 514).

0.8

Method

ECE (↓)

Brier (↓)

0.6

Majority-class baseline RCS (ours)

— 0.035

0.239 0.183

0.4

Table 4: Leave-one-out feature ablation: ∆AUC-AC when feature removed. Negative = feature contributes positively.

0.2 0.0 0.0

Base accuracy (0.605)

0.2

0.4

0.6 Coverage

0.8

1.0

Feature removed

Figure 2: Accuracy-coverage curves on MuSiQue (LLM-judge pipeline, test split). RCS dominates baselines at every coverage level. The dashed line shows the unguarded base accuracy (60.5%); points above it are net gains from abstention. Table 2: Operating points on MuSiQue test (base accuracy = 60.5%). Method

Cov.

Acc.

CWAR

∆Acc

RCS (ours)

70% 50%

75.1% 79.4%

24.9% 20.6%

+14.6pp +18.9pp

Lift

70% 50%

73.2% 76.7%

26.8% 23.3%

+12.7pp +16.2pp

Entropy

70% 50%

68.6% 73.1%

31.4% 26.9%

+8.1pp +12.6pp

lute AUC-AC values are lower across all methods on dense MuSiQue (best: 0.557 vs. 0.790 on Proposal H), reflecting that dense-pipeline failures are more uniformly distributed and thus harder to predict from pre-retrieval features alone. The dense 2Wiki condition is more tractable (best: 0.649), suggesting that the structured nature of 2Wiki questions preserves predictive signal even without judge re-ranking. 7.4

Calibration

Table 3 shows calibration metrics on the MuSiQue test split. ECE = 0.035 indicates good calibration: among queries where RCS ≈ 0.7, approximately 70% are indeed fully successful. The Brier score improvement over the majority-class baseline (0.183 vs. 0.239) shows that the probabilistic predictions carry meaningful information beyond the prior. 7.5

Cross-Dataset Generalization

We train RCS on MuSiQue tune and evaluate on 2Wiki test (zero-shot transfer). AUC-AC

MuSiQue PropH HoVer Dense

None (full model)

0.790

0.873

QUERY- LEN HOP 1- LIFT HOP 2- MAX HOP 1- TOP 3 HOP 1- MAX HOP 2- MARGIN HOP 1-H HOP 2-H HOP 1- MARGIN

−0.012 −0.003 −0.002 −0.001 −0.000 +0.000 +0.000 +0.000 +0.001

−0.002 −0.015 n/a −0.026 −0.015 n/a +0.000 n/a −0.002

= 0.942, within 0.5pp of the 2Wiki in-domain model (0.947). This near-perfect transfer holds despite MuSiQue’s failure rate being 2.8× higher than 2Wiki’s (39.5% vs. 14.5%). The result confirms that regime features capture domainagnostic structure. 7.6

Feature Importance

Table 4 reports leave-one-out AUC drop on MuSiQue (9-feat) and HoVer (6-feat). The dominant feature differs by dataset: QUERY- LEN is strongest on MuSiQue (−1.2pp), while HOP 1TOP 3 dominates on HoVer (−2.6pp). This divergence reflects architectural differences: MuSiQue Proposal H failures are driven by compositional question complexity (longer = more bridging steps), whereas HoVer dense failures are driven by hop-1 retrieval concentration (multi-hop claims require multiple supporting passages, and s(1) 1:3 directly measures whether the top passages form a strong candidate pool). On MuSiQue, RCS extracts most of its signal from QUERY- LEN, HOP 1- LIFT, and HOP 2MAX; entropy features ( HOP 1-H, HOP 2-H) add zero marginal information, consistent with the weak standalone performance of the Entropy baseline. On HoVer, RCS’s gain over max-score (+3.4pp) arises primarily from HOP 1- TOP 3 and HOP 1- LIFT: the average quality of the top candidate pool is a better predictor of multi-passage claim coverage than the single peak score alone.

8

Discussion

Why entropy fails on 2Wiki. On 2Wiki, entropy-only barely exceeds random (0.849 vs. 0.841). Many 2Wiki failures occur in comparison subtypes: the bridge passage is retrieved with high confidence (peaked hop-1 distribution = low entropy), but the second gold passage is missed because SVO expansion from the bridge produces a weak query. Entropy captures only hop-1 uncertainty; RCS includes HOP 2- MAX, which captures hop-2 expansion quality and is decisive on 2Wiki. RCS vs. MLP. On HoVer dense and MuSiQue dense, MLP (0.878, 0.557) and RCS (0.873, 0.556) are within bootstrap-CI overlap. MLP’s competitive performance is expected: given the same features, a two-layer network can fit mild non-linearities that logistic regression misses. We prefer RCS for deployment: it is interpretable (one scalar weight per feature), always calibrated by construction, and avoids overfitting risk on the small training sets available in production (486– 1007 calibration queries). Query length as the dominant feature on MuSiQue. The strong negative effect of query length is consistent with the B RIDGE RAG analysis (Bacellar, 2026a): longer questions contain more bridging clauses, increasing the chance that SVO extraction produces a malformed hop2 query. At query time, length is perfectly observable and costs nothing to compute. A simple rule—“flag long questions for manual review”— would capture a substantial fraction of RCS’s benefit on MuSiQue, while on 2Wiki (where querylen-inv ties RCS) it is the dominant signal. On HoVer, however, length alone ranks 4th—hop1 concentration features become primary when multi-passage coverage (not query structure) is the binding constraint. Deployment framing: the cost of confident failures. CWAR has a concrete product interpretation. In a deployed multi-hop fact-verification system (HoVer dense pipeline, CWAR = 31.7%), over one third of verified claims are incorrect with no downstream signal of unreliability — the system returns a ranked list with the same interface confidence as correct results. In a multi-hop QA system (MuSiQue PropH, CWAR = 39.5%), nearly two-fifths of answers are built on incomplete evidence, delivered confidently to downstream components or users. The operating points

in Table 2 quantify the reduction: at 50% coverage, CWAR drops from 39.5% to 20.6% — the system answers half its queries and reduces its confident-failure rate by 47.8% relative. At 70% coverage the system still answers 70% of queries while reducing confident errors by 37% relative (39.5% → 24.9%). All of this is achieved at <1 ms per query from ANN scores already available at retrieval time, with no additional model inference. Theorem 1 guarantees these gains are maximal for LLM-judge pipelines: I(y; ϕ) is high for PropH failures, so the feature information is fully exploitable. The same theorem explains the smaller absolute gains on dense pipelines: their failures are closer to uniform in feature space, and no confidence model — however complex — can substantially reduce CWAR when I(y; ϕ) ≈ 0. Abstention vs. re-routing. R EGIME A BSTAIN abstains entirely when RCS < τ . An alternative is to re-route to a more expensive pipeline. Our experiments confirm that the rescue-judge pass from Bacellar (2026a) recovers 27 of the 203 MuSiQue test failures; RCS could trigger this fallback rather than outright abstention, limiting the cost overhead to the ∼20% of queries flagged at 80% coverage.

9

Conclusion

We presented R EGIME A BSTAIN, a calibrated retrieval abstention framework based on a logistic Retrieval Confidence Score derived from query-time regime features. Across three multihop datasets and two retrieval architectures (five failure regimes, CWAR 14.5%–62.1%), RCS achieves best or co-best AUC-AC against eight baselines in all conditions. On MuSiQue with the LLM-judge pipeline, RCS reduces CWAR from 39.5% to 20.6% at 50% coverage, with ECE = 0.035; on 2WikiMultiHopQA it outperforms entropy-only by +9.8pp AUC even though entropy is near-random on that dataset. Nearperfect cross-dataset transfer (−0.5pp AUC) confirms domain-agnostic structure in the regime features. Feature importance varies by dataset: query length dominates on MuSiQue, hop-1 concentration on HoVer. This finding has direct implications for production deployments: a calibrated combination of a few scalar features available at retrieval time is sufficient to gate retrieval quality across diverse architectures, with interpretable

failure-mode diagnostics built in. R EGIME A BSTAIN is fully composable with R EGIME ROUTER (Bacellar, 2026b): run RCS first, abstain or escalate if below τ , otherwise route to the appropriate retrieval pipeline. This adds a well-calibrated abstention gate to any pipeline based on iterative ANN search at negligible inference cost.

References Anastasios N. Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. 2022. Conformal risk control. In Proceedings of ICLR. Andre Bacellar. 2026a. BridgeRAG: Bridgeconditioned retrieval for multi-hop QA. arXiv preprint arXiv:2604.03384. Andre Bacellar. 2026b. Regime-conditional retrieval augmentation without LLMs at query time. arXiv preprint arXiv:2604.09019. C.K. Chow. 1970. On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory, 16(1):41–46. Yonatan Geifman and Ran El-Yaniv. 2017. Selective prediction in deep neural networks. In Proceedings of NeurIPS. Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of COLING, pages 6609– 6625. Jamie Holdcroft, Abdelrahman Abdallah, and Adam Jatowt. 2026. The magnitude mirage: Rethinking confidence for reasoning-intensive retrieval. arXiv preprint arXiv:2609.15578. Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, and 1 others. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Amita Kamath, Robin Jia, and Percy Liang. 2020. Selective question answering under domain shift. In Proceedings of ACL, pages 5684–5696. Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. arXiv preprint arXiv:2302.09664. Yichen Shi, Sravana Subramanian, Nikhita Bhutani, Eduard Hruschka, and Xin Luna Dong. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of EMNLP, pages 3441–3450.

Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10:539–554.

Record · ID 1006832 · SHA-256 8688e2f17643161c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.