T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Muhan Gao 1 Zih-Ching Chen 2 Kuan-Hao Huang 1
Abstract
extensive documents into a single context. Deep research pipelines (OpenAI, 2025), for instance, autonomously retrieve and synthesize information from numerous sources, while long-document analysis systems enable users to query entire books or legal documents (Ke et al., 2026; Chang et al., 2024; Guha et al., 2023; Su et al., 2025). These applications accumulate large volumes of text, often exceeding 100K tokens, before generating a final response. However, as models ingest more documents, they inevitably encounter information that is topically relevant yet ultimately misleading.
arXiv:2605.10828v1 [cs.AI] 11 May 2026
As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becomes critical. Prior work shows that semantically relevant yet misleading documents degrade performance, but the quantitative relationship between the proportion of distractors and performance remains unstudied. In this work, we systematically vary the hard-distractor proportion in fixed-length contexts, revealing a striking nonlinear pattern: as the proportion of hard distractors increases, performance drops sharply within the first small fraction, while the remainder of the range yields only marginal additional decline. We term this “ T HE F IRST D ROP OF I NK” effect, analogous to how a single drop of ink contaminates water. Our theoretical and empirical analyses grounded in attention mechanics show that hard distractors capture disproportionate attention even at small proportions, with diminishing marginal impact as their proportion grows. Controlled experiments further show that filtering gains mainly come from context-length reduction rather than distractor removal; substantial recovery requires reducing the hard-distractor proportion to near zero, highlighting the importance of upstream retrieval precision.
Prior work on long-context language models has primarily focused on how the position (Liu et al., 2024) and length (Bianchi et al., 2025; Levy et al., 2025; 2024) of relevant information affect performance. Less attention has been paid to the surrounding context itself. While research on short-context reasoning tasks (Shi et al., 2023; Yang et al., 2025a) and retrieval-augmented generation (RAG) systems (Lee et al., 2026; Jin et al., 2025) demonstrates that distractors can cause non-negligible performance drops, and Hong et al. (2025) reveals that this effect amplifies as context length grows, how performance degrades in long contexts as the proportion of misleading documents increases remains unexplored. A natural question arises: how does performance change as the proportion of distractors grows in long-context reasoning? In this work, we systematically vary the proportion of hard distractors within fixed-length contexts and identify T HE F IRST D ROP OF I NK effect as in Figure 1: as hard distractor proportion increases, performance drops sharply within the first small fraction, then plateaus with only marginal further decline. We provide a theoretical analysis grounded in the softmax attention mechanism, showing that attention on the gold document is a convex function of hard distractor proportion, with empirical validation on retrieval heads (Wu et al., 2025; Zhang et al., 2025b). This explains the observed nonlinearity: hard distractors dominate the softmax denominator even at small proportions, implying that partially removing them yields negligible recovery and only near-complete removal restores performance.
1. Introduction Recent advances in long-context language models (Anthropic, 2025) have given rise to applications that aggregate 1
Department of Computer Science and Engineering, Texas A&M University, College Station, TX, USA 2 NVIDIA AI Technology Center, NVIDIA Corporation, Santa Clara, CA, USA. Correspondence to: Muhan Gao <[email protected]>, Kuan-Hao Huang <[email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
These findings challenge the prevailing assumption in long1
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning Removing distractors yields ❌ proportional gains.
Attention Before Softmax:
With 100 distractor docs: Query:... Answer: _
8 ≈ 9 >> 1
Attention of gold doc after Softmax:
0% Hard:
10% Hard:
≈ 97% ≈ 21%!
Small fraction of hard ✅ distractors ruin the accuracy!
Figure 1. T HE F IRST D ROP OF I NK effect. Left: Conventional linear assumption (top, red dashed line) versus empirically observed nonlinear degradation (bottom, blue curve): a small fraction of hard distractors is sufficient to severely degrade accuracy. Middle: Hard distractors receive similar attention logits as gold documents (8 ≈ 9 ≫ 1), dominating the softmax competition even at low proportions. Right: With 100 distractor documents, attention on gold drops 76% by adding only 10% hard distractors. This convex relationship explains T HE F IRST D ROP OF I NK.
context applications that accumulating more documents improves performance. As long as even a small fraction of hard distractors remains in the context, performance is severely degraded; consequently, post-hoc filtering in most cases yields only marginal recovery. This suggests that preventing hard distractors from entering the context in the first place is more critical than filtering them afterward.
in-the-Middle” phenomenon: LLMs prioritize information at the start and end of contexts while neglecting middle portions. This limitation persists across various scenarios (Lee et al., 2025; Gao et al., 2024), motivating followup studies to develop methods for mitigating positional bias (Hsieh et al., 2024a; Zhang et al., 2024b; Wang et al., 2025). Beyond position, needle length also affects retrieval accuracy (Bianchi et al., 2025; Levy et al., 2025).
Contribution. (1) We identify T HE F IRST D ROP OF I NK effect across multiple models and datasets: as the proportion of hard distractors increases, performance degrades sharply within the first small fraction, then plateaus. (2) We provide a theoretical explanation showing that attention on the gold document is a strictly convex function of hard distractor proportion, and validate this empirically through attention logit measurements on retrieval heads. (3) We design controlled experiments to disentangle the effects of context length and distractor composition, showing that conventional filtering yields gains primarily from context reduction, and removing hard distractors only provides substantial benefit when their proportion is reduced to near zero.
Recent work has also examined the haystack itself. In original settings (Kamradt, 2023; Hsieh et al., 2024b), the haystack consists of irrelevant documents, posing no semantic confusion with the target needle. Yang et al. (2025b) use synthetically generated biographies to improve coherence between needles and haystack, better approximating realistic retrieval conditions. Further studies introduce semantically related distractors into the haystack and observe non-negligible performance degradation (Lee et al., 2026; Hong et al., 2025). However, under the long context setting, how performance varies with the proportion of distractors in the haystack remains unexplored. Information aggregation in agentic systems. The rise of agentic AI has fundamentally transformed how LLMs interact with external information. Rather than responding to a single query with a fixed context, modern systems such as deep research pipelines (OpenAI, 2025; Zhang et al., 2025a), multi-agent collaboration frameworks (Wu et al., 2023; Hong et al., 2024; Li et al., 2023), and tool-augmented agents (Schick et al., 2023; Qin et al., 2024; Yao et al., 2023) autonomously gather, aggregate, and synthesize information across multiple retrieval rounds. These systems routinely accumulate contexts exceeding 100K tokens before producing a final response (Singh et al., 2025). Recent work has shown that even context length alone can degrade performance (Du et al., 2025), further underscoring the challenges of information aggregation at scale.
2. Related Work Long-context understanding and evaluation. The ability to process long context has emerged as a critical capability for large language models (LLMs), with the context window extended from 4K to over 1M tokens (Anthropic, 2025; Team, 2024; Xiao et al., 2024; Ding et al., 2024; Peng et al., 2024). This expansion has motivated efforts to more effectively understand and evaluate long-context ability. Among various evaluation approaches, the “Needle-in-aHaystack” (NIAH) paradigm (Kamradt, 2023) is preferred due to its controllability and ease of construction (Hsieh et al., 2024b; Yen et al., 2025; Bai et al., 2024; Zhang et al., 2024a). A “needle” (a fact or short passage required to answer a query) is inserted into a “haystack” of unrelated filler text, and the model must locate and use the needle while ignoring the surrounding context.
This information aggregation process introduces an unavoidable challenge: while gathering relevant information, these systems inevitably accumulate unhelpful or misleading documents along the way (Jin et al., 2025; Shi et al., 2023;
Under this paradigm, Liu et al. (2024) identify the ”Lost2
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning 100
Llama-3.1-8B-Instruct - 128K tokens
Accuracy (%)
90 80 70 60 50 40
Qwen2.5-7B-Instruct - 128K tokens
100 90 80 70 60 50 40 30 20
100 90 80 70 60
0 10 20 30 40 50 60 70 80 90 100
0 10 20 30 40 50 60 70 80 90 100
nq_easy
triviaqa_random
Hard Proportion (%) nq_random
triviaqa_easy
Qwen3-Next-80B-Instruct - 128K tokens
0 10 20 30 40 50 60 70 80 90 100
Hard Proportion (%) popqa_easy
Hard Proportion (%)
popqa_random
hotpotqa_easy
hotpotqa_random
Figure 2. Accuracy as a function of hard distractor proportion at 128K context length across three models (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Qwen3-Next-80B-Instruct) on Natural Questions, TriviaQA, PopQA and HotpotQA. Across all configurations, introducing the first 10% of hard distractors (shaded region) causes steep performance degradation, while further increases yield only marginal decline. Despite substantial variation in absolute accuracy across datasets (e.g., HotpotQA shows the lowest baseline due to multi-hop reasoning), the nonlinear pattern persists, illustrating the T HE F IRST D ROP OF I NK effect.
Distractors. To control distractor difficulty, we use three categories of passages with varying degrees of relevance to the query q: (1) Easy (E): repetitions of a single filler sentence ”The grass is green. The sky is blue. The sun is yellow...”; (2) Random (R): arbitrary passages sampled from the Wikipedia 2019-08-01 dump from KILT (Petroni et al., 2021); (3) Hard (H): semantically related passages retrieved from Wikipedia using BM25, which are topically relevant to q but do not contain the answer. We use gpt-4o-mini to examine each hard distractor and filter out those that contain the answer in any form (including paraphrases or alternative expressions of a, prompt in §C), ensuring that hard distractors are genuinely misleading rather than inadvertently providing correct information. All three categories of distractors are normalized to approximately 100–150 tokens to avoid length bias.
Yang et al., 2025a). Prior work demonstrates that such noisy retrieval can significantly degrade LLM performance (Cuconasu et al., 2024; Yoran et al., 2024). In response, filtering and reranking have become standard techniques for improving RAG performance, operating under the assumption that removing distractors yields substantial gains (Glass et al., 2022; Yoran et al., 2024). However, these findings are derived from relatively short contexts of only a few thousand tokens. Whether the same assumptions hold, and whether existing mitigation strategies remain effective, as context windows scale to 128K tokens and beyond, remains underexplored.
3. Nonlinearity in Distractor Effects We study how the proportion of hard distractors affects model performance in a multi-document question answering setting, where a language model must locate relevant information among retrieved passages. Formally, given a query q, a gold passage J ∗ containing the answer, and a set of N distractor passages {P1 , . . . , PN }, the model must attend to J ∗ to produce the correct answer. We categorize distractors into three types based on their semantic relevance to q: easy (E), random (R), and hard (H), and systematically vary their proportions to study the resulting performance degradation. We begin by describing our experimental setup (§3.1) and then present our findings (§3.2).
Input. Given a target context length T and a hard distractor proportion p ∈ [0, 1] (see §A for specific values), we construct the input by concatenating the gold passage J ∗ with distractors sampled to fill the context. For each dataset, we consider two mixing strategies: (1) easy-hard mixing, where proportion p of distractors are from H and (1 − p) are from E (e.g., nq easy), and (2) random-hard mixing, where proportion p of distractors are from H and (1 − p) are from R (e.g., nq random). All passages are randomly shuffled before concatenation to avoid positional bias. For each setting (dataset × context length × hard proportion), we sample 200 examples for evaluation.
3.1. Experimental Setup
Evaluation. Hsieh et al. (2024b) employs string containment matching to evaluate QA tasks:
Dataset. We use Natural Questions (Kwiatkowski et al., 2019), TriviaQA (Joshi et al., 2017), PopQA (Mallen et al., 2023), and HotpotQA (Yang et al., 2018), covering both single-hop and multi-hop reasoning. Each sample is a tuple (q, a, J ∗ ), where q is the question, a is the gold answer, and J ∗ is the gold passage from which a can be derived.
N
Accuracy =
1 X max ⊮[lower(r) ⊆ lower(pi )] N i=1 r∈Ri
where pi denotes the model prediction, Ri is the set of 3
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
llama-3.2-1b-nim
nq_easy 0.26 0.77 0.62 0.76 0.59 0.79 nq_random 0.36 0.49 0.00 0.17 0.55 0.80 triviaqa_easy 0.55 0.93 0.66 0.73 0.56 0.74 triviaqa_random 0.29 0.16 0.62 0.53 0.46 0.53 popqa_easy 0.29 0.69 0.49 0.51 0.51 0.62 popqa_random 0.25 0.47 0.51 1.36 0.78 0.39 hotpotqa_easy 0.48 0.62 0.54 0.73 0.71 0.78 hotpotqa_random 0.18 0.42 -0.06 0.50 0.14 1.40 4K 8K 16K 32K 64K 128K
llama3.1-8b-chat
qwen2.5-7b-ins
0.13
0.53
0.67
0.64
0.23
0.53
0.44
0.40
0.58
0.27
0.78
0.46
0.73
0.53
0.61
-0.11 0.21
0.24
0.84
0.47
0.81
0.87
0.70 -0.20 0.34
0.50
1.04
0.43
0.89
0.39
0.67
0.61
0.66
0.75
0.21
0.67
0.48
0.48
0.69
0.55
0.61
0.62
0.84
0.96
0.82
0.75
0.73
0.73
0.71
0.64
0.35
0.36
0.46
0.54
0.29
0.55
0.57
0.18
0.83
0.36
0.28
0.04
0.58
0.75
0.55
1.00
0.54
0.67
0.35
0.00
0.33
0.21
0.16
0.58
0.66
0.31
0.74
0.31
0.46
0.48
0.67
2.00
0.40
0.20
0.47
0.25
0.24
0.12
0.10
0.25
0.39
0.47
0.57
0.36
0.59
0.40
0.44
0.45
0.80
0.60
0.14
0.31
0.23
0.36
0.42
0.67
0.55
0.67
0.85
0.80
0.92
0.22
0.58
0.44
0.55
0.84
0.72
0.25
0.52
0.63
0.60
0.53
0.65
0.39
0.87
0.44
0.70
0.40
1.10
1.29
0.32
0.18
0.38
0.56
1.11
0.00
0.12
0.47
0.56
0.46
0.40
4K 8K 16K 32K 64K 128K
0.62
qwen3-next-80b-ins
-0.12 0.44
4K 8K 16K 32K 64K 128K
4K 8K 16K 32K 64K 128K
Figure 3. Drop ratio of accuracy degradation across different context lengths, models, and datasets. The drop ratio measures the fraction of total performance loss that occurs in the first 10% of hard distractors. A linear degradation would yield 0.1. Negative values indicate the first 10% of hard distractors does not further degrade performance. Darker green indicates higher drop ratios (stronger nonlinearity), while orange/red indicates values near or below the linear baseline. The prevalence of green across the table confirms that the T HE F IRST D ROP OF I NK effect is consistent across models and datasets.
reference answers, and ⊮[·] is the indicator function. A prediction is considered correct if any reference answer appears as a substring within it (case-insensitive). However, we observe that string matching suffers from false negatives (e.g., “3” vs. “three”, “Bill Clinton” vs. “William Jefferson Clinton”, “1986-2013” vs. “from 1986 to 2013”). Taking HotpotQA as an example, we identify 36 out of 200 samples (18%) where the model’s response is semantically correct but marked incorrect by string matching.
which means 58% of the total degradation occurs in the first 10% of hard distractors. These results show the T HE F IRST D ROP OF I NK effect, which is contrary to the linear degradation assumption where each additional hard distractor contributes equally and the expected ratio would be 0.1.
4. The First Drop Matters Most In this section we theoretically analyze the mechanistic reason of T HE F IRST D ROP OF I NK effect based on the transformer’s attention mechanism.
Therefore, we follow Yen et al. (2025) and use gpt-4o-mini as an LLM judge to verify correctness. The judge receives only the gold document J ∗ , question q, correct answer a, and model output, and determines whether the output is semantically correct (prompt in §C). We manually check 100 samples on each dataset and find the judge produces only 17 false negatives in total (4.25%), consistent with prior findings that LLM judges achieve Cohen’s κ of 0.72–0.91 with human judgment (Yen et al., 2025).
4.1. Preliminaries and Notations Attention mechanism. For a sequence of T tokens with hidden representations {hi }Ti=1 ∈ Rd , the attention mechanism computes query, key, and value projections: qi = WQ hi ,
vj = WV hj
The attention logits, weights, and output are:
3.2. T HE F IRST D ROP OF I NK Effect
T X exp(zi,j ) qi⊤ kj , αi,j = PT zi,j = √ αi,j vj , oi = dk ℓ=1 exp(zi,ℓ ) j=1
We demonstrate the results for 3 models on the length of 128K tokens in Figure 2, more detailed results for all the models can be found in §A.
(1)
An autoregressive model predicts the next token based on the last position’s attention over all preceding tokens.
Accuracy shows a nonlinear relationship with hard distractor proportion. As shown in Figure 2, the initial increase in hard distractor proportion (0–10%, shaded region) causes disproportionately large performance drops compared to subsequent increases (10–100%). To quantify this asymmetry, we compute the ratio of accuracy drop in the 0– 10% region versus the total drop from 0–100% in Figure 3: Drop Ratio =
kj = WK hj ,
Retrieval task. When predicting the answer to query q, the relevant information lies in the gold passage J ∗ , a span of tokens within the context. Retrieval succeeds if the model attends sufficiently to J ∗ when generating the answer. Prior work on retrieval heads (Wu et al., 2025; Zhang et Pal., 2025b) has shown that the attention weight αi,J ∗ := j∈J ∗ αi,j on the target passage strongly correlates with downstream accuracy: higher attention mass on the gold passage leads to higher probability of correct answer generation.
Acc(0%) − Acc(10%) Acc(0%) − Acc(100%)
A linear degradation would yield a ratio of 0.1; significantly higher values indicate front-loaded degradation. For example, on nq easy at 128K context, Qwen2.5-7B-Instruct exhibits a drop ratio of 0.58,
Logit margin. Let i denote the position of the last token, from which the model generates the answer. For a passage P spanning multiple tokens, we define its aggregate logit as 4
(a) Fixed gap e h = 4 Same ba = e4 same shape
0.02
e = 5, e = 6, e = 7, e = 8,
h=1 h=2 h=3 h=4
0.01 0.00
2
0
2 4 6 Hard Distractor Proportion (%)
8
Normalized Attention
Attention i, *
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
10
(b) Fixed h = 2, varying gap
1.00
Larger ba
0.75
gap = 1 gap = 2 gap = 3 gap = 4 gap = 5
stronger First Drop
0.50 0.25 0.00
0
10
20 30 40 Hard Distractor Proportion (%)
50
Figure 4. Two controlling factors of the theoretical attention curve (Remark 4.3). (a) When the margin gap ∆e − ∆h is fixed at 4, all curves share identical shape (same b/a = e4 ) but differ in vertical position, controlled by 1/a. Faded lines extend into p < 0 to illustrate the shape equivalence. (b) When ∆h is fixed at 2, increasing ∆e enlarges the ratio b/a, producing more convex curves and amplifying the “First Drop of Ink” effect (shaded region, 0–10%).
P 1 zi,P := |P| j∈P zi,j , representing how strongly the last token attends to passage P. The margin between the target passage J ∗ and a distractor passage P is:
The denominator decomposes as: X j∈J ∗
∆P := zi,J ∗ − zi,P
|
where zi,j is the attention logit defined in Eq. (1).
X
exp(zi,j ) +
j∈E
{z
target passage J ∗
}
|
X
exp(zi,j ) +
j∈H
{z
}
weaker distractors E
|
X
exp(zi,j )
j∈O
{z
}
hard distractors H
|
{z
other tokens O
}
By the definition of logit margin, for tokens in weaker distractors we have zi,j = zi,J ∗ − ∆e , and for tokens in hard distractors we have zi,j = zi,J ∗ − ∆h . Thus: X exp(zi,j ) = (1 − p) · Td · exp(zi,J ∗ ) · e−∆e
In our two mixing strategies (§3.1), we always mix hard distractors (H) with a weaker distractor type (either easy or random). For notational simplicity in the following analysis, we use E to denote the weaker distractor set and ∆e to denote its characteristic margin: 1 X 1 X ∆e := ∆P , ∆h := ∆P |E| |H| P∈E
exp(zi,j ) +
j∈E
X
exp(zi,j ) = p · Td · exp(zi,J ∗ ) · e−∆h
j∈H
X
P∈H
exp(zi,j ) = To · exp(zi,J ∗ ) · e−∆o
j∈O
Since hard distractors are semantically more similar to the query and compete more strongly for attention, we have ∆h ≪ ∆e . We empirically validate this in §5.
Substituting and factoring out exp(zi,J ∗ ) from numerator and denominator, and noting that Tg ≪ Td (the target passage is small relative to distractors): αi,J ∗ (p) 1 = 1 + (1 − p) · Td · e−∆e + p · Td · e−∆h + To · e−∆o 1 = 1 + (1 − p)a + pb + c
4.2. Theoretical Explanation of Nonlinear Degradation Lemma 4.1 (Attention Weight with Mixed Distractors). Consider a context of total length T tokens, consisting of: (1) A target passage J ∗ with Tg tokens; (2) Distractor passages with Td tokens, where proportion p ∈ [0, 1] are from H and (1 − p) are from the weaker category; (3) Other tokens (query, instructions) with To tokens, where T = Tg + Td + To . The aggregate attention weight on the target passage is: 1 αi,J ∗ (p) = 1 + (1 − p) · a + p · b + c
where a := Td · e−∆e , b := Td · e−∆h , and c := To · e−∆o . Lemma 4.2 (Monotonicity and Convexity). Let f (p) = 1 αi,J ∗ (p) = 1+(1−p)a+pb+c . Then f ′ (p) < 0 (strictly decreasing) and f ′′ (p) > 0 (strictly convex) for all p ∈ [0, 1].
where a := Td · e−∆e and b := Td · e−∆h represent the aggregate contributions from weaker and hard distractors respectively, and c := To · e−∆o denotes the contribution from other tokens.
Proof. Let D(p) := 1+(1−p)a+pb+c = 1+a+c+p(b− a). Since ∆h ≪ ∆e (hard distractors have smaller margins), we have e−∆h > e−∆e , and thus b = Td · e−∆h > Td · e−∆e = a. Let γ := b−a > 0. Then D(p) = 1+b+c+pγ.
Proof. From Eq. (1), the aggregate attention weight on the target passage is: P j∈J ∗ exp(zi,j ) αi,J ∗ = PT j=1 exp(zi,j )
First derivative: f ′ (p) = − 5
γ D(p)2
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Since γ > 0 and D(p) > 0, we have f ′ (p) < 0.
5. Validation of Theoretical Explanation
Second derivative:
One natural question is: does the model truly exhibit a clear gap between attention on semantically similar versus dissimilar distractors (i.e., ∆h ≪ ∆e )? To answer this, we measure ∆e and ∆h by computing the attention logit difference between the target passage and distractor passages. Rather than averaging across all attention heads or selecting a specific layer, we follow Zhang et al. (2025b) to identify the sparse subset of heads (approximately 1–2%) responsible for retrieving relevant information from context, as their attention mass directly correlates with retrieval success.
f ′′ (p) =
2γ 2 D(p)3
Since γ 2 > 0 and D(p) > 0, we have f ′′ (p) > 0. Remark 4.3 (Simplified Form for Large Context). When a, b ≫ 1 (i.e., Td is sufficiently large), the constant term in the denominator becomes negligible:
Specifically, given a query q and context containing gold passage J ∗ among distractors, we score each attention head h by the attention mass it allocates from query tokens to the gold passage. While Zhang et al. (2025b) use post-softmax attention weights, we observe numerical underflow in long contexts (128K tokens) where attention weights become vanishingly small. We therefore use pre-softmax logits instead: 1 X 1 X tq →td Scoreh (q) = zh |q| t ∈q |J ∗ | ∗
1 1 α(p) = ≈ 1 + (1 − p)a + pb (1 − p)a + pb 1 1 = · a 1 + p ab − 1 This reveals two key insights, as illustrated by Figure 4: • Vertical position is controlled by 1/a = e∆e /Td : larger ∆e shifts the curve upward.
q
• Curve shape is controlled solely by b/a = e∆e −∆h : only the margin gap matters, not the absolute values.
td ∈J
t →t where zhq d is the attention logit from query token tq to document token td in head h. For each setting of dataset
Takeaway: The monotonicity confirms that increasing the hard distractor proportion always hurts attention on the target. The strict convexity further implies that this degradation is front-loaded: the first few percent of hard distractors cause disproportionately large drops, while subsequent increases have diminishing impact. Together, these properties provide the theoretical basis for T HE F IRST D ROP OF I NK effect.
and hard proportion, we use 50 samples to identify the topscoring heads as retrieval heads, and then measure ∆e and ∆h on the remaining 150 samples. The identified heads are highly stable: Pearson correlation between train and test scores on the top-16 heads is 0.96±0.01, and Spearman rank correlation across all heads is 0.99 ± 0.00. We report results for Llama-3.1-8B-Instruct in Figure 5; results for Llama-3.2-1B-Instruct and additional details are provided in §B.
Llama-3.1-8B-Instruct: easy
Margin separation confirms theoretical assumption. Figure 5 shows the measured margins for Llama-3.1-8B-Instruct across different hard proportions. We observe a clear and consistent separation: ∆e ≈ 7–10 while ∆h ≈ 2–3, yielding an average gap of 5.83. This confirms our theoretical assumption that ∆h ≪ ∆e . To understand the practical implication, consider a 128K context with Td = 128000 distractor tokens. The ratio b/a = e∆e −∆h ≈ e5.83 ≈ 340 means that each hard distractor token contributes 340× more to the softmax denominator than an easy distractor token. Even at just 10% hard proportion, hard distractors account for 0.1×340 0.1×340+0.9×1 ≈ 97% of the total distractor contribution, completely dominating the attention competition.
hard
easy hard
Delta Value (attention logits)
10 8 6
8.0
6.9
Avg Margin: 5.83
6.1
5.2
4.6
4.1
40
60
80
90
4 2 1
20
Hard Proportion (%)
Figure 5. Empirical measurement of logit margins ∆e and ∆h on retrieval heads for Llama-3.1-8B-Instruct. Green bars show ∆e (margin to easy distractors) and brown bars show ∆h (margin to hard distractors). The gap ∆e − ∆h (annotated values) remains substantial across all hard proportions, with an average of 5.83. This validates the theoretical assumption ∆h ≪ ∆e .
Shrinking gap reinforces T HE F IRST D ROP OF I NK effect. One might argue that the margin gap (∆e − ∆h ) decreases as hard proportion increases: from 8.0 at 1% to 4.1 at 90%, and wonder whether this undermines our theory. In fact, the opposite is true: this observation reinforces T HE 6
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
F IRST D ROP OF I NK effect. The largest margin gap occurs precisely when the hard proportion is lowest, meaning the first few hard distractors enjoy the maximum competitive advantage (b/a ≈ e8.0 ≈ 2980) over easy distractors. As more hard distractors are added, the gap shrinks and so does their marginal impact (b/a ≈ e4.1 ≈ 60 at 90%). This is exactly the pattern our theory predicts: first drops sharply and then plateaus.
dynamics, degrading performance even when the attention distribution appears more favorable. Indeed, effective temperature scaling typically requires adjustment during training or fine-tuning (Ryan, 2024; Ram et al., 2025), as the model must learn to adapt its representations to the modified softmax behavior. Our results suggest that T HE F IRST D ROP OF I NK effect cannot be mitigated through simple inference-time interventions.
Llama-3.1-8B-Instruct - 128K tokens 90
Accuracy (%)
Attention logit
T=0.9 T=1.0
Attention (T = 1) small
Attention (T < 1)
large
80
Figure 7. Effect of temperature scaling on attention distribution. From left to right: pre-softmax logits, attention weights at τ = 1, and attention weights at τ < 1. Colors denote easy distractors, hard distractors, and target passage. Lower temperature sharpens the softmax, suppressing hard distractors while maintaining attention on the target.
70
60 0
20
40
60
80
100
Hard Proportion (%) Figure 6. Effect of softmax temperature scaling on accuracy across hard proportions (nq easy, Llama-3.1-8B-Instruct). Lower temperature (τ = 0.9) consistently degrades performance despite theoretically sharpening attention toward the target, indicating that inference-time temperature adjustments cannot mitigate T HE F IRST D ROP OF I NK effect.
6.2. Incremental Filtering of Hard Distractors Our main experiments vary the hard proportion while fixing the context length. In practice, however, filtering removes unwanted passages entirely, reducing the overall context length. This creates a confound: when filtering improves performance, is the gain due to removing hard distractors, or simply due to shorter context? We design two experiments to disentangle these factors: (1) We compare Filter Hard versus Filter Random with symmetric starting compositions, removing only hard or weaker distractors respectively, isolating the effect of filtering strategy. (2) We perform Proportional Reduction, shrinking context length while holding the hard distractor ratio fixed by removing both types proportionally, isolating the pure effect of context length. Both experiments are conducted on Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct for all four datasets.
6. Implications for Mitigation Strategies 6.1. Inference-Time Temperature Scaling Our theoretical analysis in §4 indicates that the softmax function’s exponential nature causes hard distractors to dominate the attention competition despite their lower logits than the target. A natural hypothesis is that decreasing the softmax temperature τ during inference could “sharpen” the attention distribution, amplifying the target passage’s advantage as the highest-logit tokens (Figure 7). Specifically, we modify the attention computation as: exp(zi,j /τ ) αi,j = PN ℓ=1 exp(zi,ℓ /τ )
Filter Hard vs. Filter Random. Filter Hard begins with 80% hard distractors (≈ 102K) and 20% random distractors (≈ 26K), progressively removing hard distractors and reducing context length by approximately 20K tokens at each step until 27K tokens remain. Filter Random begins with the reversed composition (20% hard, 80% random) and removes random distractors at the same pace. Both experiments end at 27K tokens but with opposite compositions: Filter Hard ends with nearly all random distractors, while Filter Random ends with nearly all hard distractors. Table 1 summarizes the composition at each step.
where τ < 1 produces a sharper distribution that concentrates more attention on the target passage. Results. Figure 6 shows the effect of temperature scaling on Llama-3.1-8B-Instruct across different hard proportions on nq easy. Contrary to our hypothesis, decreasing temperature consistently degrades performance across all hard proportions. Why does this fail? Although lower temperature theoretically sharpens attention toward the target, the model was trained with τ = 1 and its learned dynamics are calibrated to this setting. Modifying τ at inference time disrupts these
Figure 8 shows the results. From 131K to 47K tokens, both filtering strategies yield nearly identical performance gains regardless of whether hard or random distractors are 7
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
NQ Accuracy (%)
POPQA
69
HOTPOTQA 69
90
66
87
63
84
72
60
81
60
68
78
57
57
80 76
1K
13
0K
11
9K
8
9K
6
7K
4
7K
2
13
0K
11
9K
8
9K
6
7K
4
7K
2
60
45
50
40
40
35
30
30
13
11
8
9K
6
9K
4
7K
2
7K
11
9K
8
9K
6
7K
4
7K
2
13
1K
11
0K
8
9K
9K
6
7K
4
7K
K
27
K
1K 10K 89K 69K 1
47
K
27
56 48 40 32 0K
9K
9K
8 6 13 11 Filter Random
Filter Hard
47
64
1K
2
1K 10K 89K 69K 1
13
Qwen2.5-7B-Inst
50
0K
0K
50 45
1K
54
13
70
35
63
1K
55
40
66
54 1K
Llama-3.1-8B-Inst
84
Accuracy (%)
TRIVIAQA
93
4
7K
7K
2
13
K
Figure 8. Filter Hard vs. Filter Random. Both strategies yield similar gains from removing the first 80K tokens, indicating that performance recovery comes from context length reduction rather than filtering strategy. The two strategies begin to diverge below 47K tokens (shaded region), where Filter Hard has a near-zero hard distractor proportion. This suggests that the gains from partial filtering are largely attributable to context reduction rather than the removal of hard distractors themselves.
removed. This indicates that the performance improvement has little to do with the filtering strategy itself, and comes almost entirely from reducing context length. However, the two curves diverge between 47K and 27K tokens (shaded area). At 27K, Filter Hard has reduced the hard proportion to near zero, consistently outperforming Filter Random, which ends with nearly all hard distractors. This asymmetry confirms that the benefit of filtering hard distractors emerges only when their proportion is reduced to near zero; above this threshold, the filtering strategy’s benefit is marginal.
This section shows that changing the hard proportion from moderate to high levels has limited marginal effect. In this regime, shortening the context can dominate the observed recovery. The divergence between Filter Hard and Filter Random at the shortest context lengths should therefore be viewed as an idealized boundary case: a clear strategyspecific gain appears only when the hard proportion is pushed close to zero, which is difficult to achieve in realistic retrieval pipelines and is not the regime targeted by most filtering methods. This distinction reconciles the two findings: hard distractor composition has its largest marginal effect near the initial contamination boundary, whereas context length dominates after the context has already been substantially contaminated.
Proportional reduction. To isolate the pure effect of context length, we shrink context from 131K to 27K tokens while maintaining a fixed hard distractor ratio (20%, 50%, or 80%) throughout. These ratios are chosen to lie beyond the initial first-drop region, where performance has largely entered the saturated regime. At each step, we remove both hard and easy distractors proportionally, for example, we remove 4K tokens of hard distractors and 16K tokens of easy distractors in the 20% setting. In this way, we ensure the composition remains constant as the length decreases.
Table 1. Experimental design for incremental filtering. Both experiments start at 128K tokens and progressively reduce context length. The symmetric design allows us to separate the effects of context length reduction from distractor composition.
Figure 9 shows the results. The three curves largely overlap despite varying hard distractor ratios, indicating that performance scales with context length rather than composition. Together with the Filter Hard vs. Filter Random results, this suggests that filtering benefits observed in practice may be primarily a byproduct of context shortening. 8
Context Length
Filter Hard Hard Random
Filter Random Hard Random
131K (start) 110K 89K 69K 47K 27K (end)
80% 76% 71% 62% 44% 3%
20% 24% 29% 38% 56% 97%
20% 24% 29% 38% 56% 97%
80% 76% 71% 62% 44% 3%
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
78
90
75
87
72
84
69
81
POPQA
78
66 1K 10K 89K 69K 47K 27K 1
1K 10K 89K 69K 47K 27K 1
13
13
55
64
50
56
45
48
40
40
35
32
13
24 1K 0K 9K 9K 47K 27K 13 11 8 6
1K 10K 89K 69K 47K 27K 1
80% Hard
HOTPOTQA
64 62 60 58 56 54
65.0
1K 0K 9K 9K 47K 27K 13 11 8 6 48
1K 0K 9K 9K 47K 27K 13 11 8 6 60
44
54
40
48
36
42
32
36
28
1K 0K 9K 9K 47K 27K 13 11 8 6 50% Hard 20% Hard
62.5 60.0 57.5 55.0
30
Qwen2.5-7B-Inst
Accuracy (%)
TRIVIAQA
Llama-3.1-8B-Inst
Accuracy (%)
NQ
1K 0K 9K 9K 47K 27K 13 11 8 6
Figure 9. Proportional Reduction. Context is reduced from 131K to 27K while maintaining a fixed hard distractor ratio (20%, 50%, or 80% hard) by removing documents proportionally from each distractor category. Across both models and all datasets, the three curves follow similar trajectories: reducing the context length consistently improves performance, while varying the fixed hard ratio within this moderate-to-high range has only a limited marginal effect. This suggests that, once the context already contains a non-negligible fraction of hard distractors, the observed recovery is driven primarily by context length reduction rather than by the exact hard-distractor ratio.
7. Limitations and Implications
8. Conclusion
Limitations. We use multi-document QA as the experimental setting throughout this paper due to its controllability and ease of evaluation. However, we acknowledge that generalizing our findings to other long-context scenarios (e.g., summarization, code understanding, or multi-turn dialogue) remains an important direction for future work. While we provide both empirical characterization and mechanistic understanding of T HE F IRST D ROP OF I NK effect, we have not yet identified an effective mitigation strategy.
In this work, we identify the T HE F IRST D ROP OF I NK effect: in long-context settings, a small fraction of hard distractors causes disproportionately severe performance degradation, while subsequent additions have diminishing impact. We provide a theoretical explanation grounded in attention mechanics and validate this theory by measuring logit margins on retrieval heads. Our findings challenge the assumption that filtering yields proportional gains and suggest that retrieval precision is far more critical than incremental filtering in long context settings.
Implications. Our work identifies the T HE F IRST D ROP OF I NK effect, implying that removing 90% of hard distractors may recover only a fraction of the lost performance, while the remaining 10% continues to dominate attention. For practitioners, this suggests prioritizing retrieval precision over recall to prevent hard distractors from entering the context in the first place. Additionally, our findings imply that disentangling the effects of context length and distractor composition matters, which prior work often conflates when evaluating long-context models.
Impact Statement We do not foresee any direct negative societal consequences of this work. However, we note that improved understanding of attention mechanisms could potentially be misused to craft adversarial inputs; we encourage the community to develop robust defenses alongside mechanistic insights.
Acknowledgements We sincerely thank Daniel Khashabi, Taiming Lu, and the Texas A&M NLP community for their helpful comments and feedback. 9
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
References
Chakraborty, T., Rose, C., and Peng, V. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pp. 23281– 23298. Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.fi ndings-emnlp.1264/.
Anthropic. Claude Sonnet 4 now supports 1M tokens of context, August 2025. URL https://claude.com /blog/1m-context. Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 3119– 3137. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL- LONG.172. URL https://doi.org/10.18653/v1/2024.acl -long.172.
Gao, M., Lu, T., Yu, K., Byerly, A., and Khashabi, D. Insights into LLM Long-context Failures: When Transformers Know but Don’t Tell. In Al-Onaizan, Y., Bansal, M., and Chen, Y. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Findings of ACL, pp. 7611– 7625. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.FINDINGS-EMNLP.447. URL https://doi.org/10.18653/v1/2024.fin dings-emnlp.447. Glass, M. R., Rossiello, G., Chowdhury, M. F. M., Naik, A., Cai, P., and Gliozzo, A. Re2G: Retrieve, Rerank, Generate. In Carpuat, M., de Marneffe, M., and Ruı́z, I. V. M. (eds.), Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pp. 2701–2715. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.NAACL-MAIN.194. URL https://doi.org/10.18653/v1/2022.naa cl-main.194.
Bianchi, O., Koretsky, M. J., Willey, M., Alvarado, C. X., Nayak, T., Asija, A., Kuznetsov, N., Nalls, M. A., Faghri, F., and Khashabi, D. Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find, 2025. URL https://arxiv.org/abs/2505.18148. Chang, Y., Lo, K., Goyal, T., and Iyyer, M. BooookScore: A systematic exploration of book-length summarization in the era of LLMs. In The Twelfth International Conference on Learning Representations, 2024. URL https://op enreview.net/forum?id=7Ttk3RzDeu. Cuconasu, F., Trappolini, G., Siciliano, F., Filice, S., Campagnano, C., Maarek, Y., Tonellotto, N., and Silvestri, F. The Power of Noise: Redefining Retrieval for RAG Systems. In Yang, G. H., Wang, H., Han, S., Hauff, C., Zuccon, G., and Zhang, Y. (eds.), Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024, pp. 719–729. ACM, 2024. doi: 10.1145/3626772.3657834. URL https: //doi.org/10.1145/3626772.3657834. Ding, Y., Zhang, L. L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M. LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens. In Salakhutdinov, R., Kolter, Z., Heller, K. A., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, Proceedings of Machine Learning Research, pp. 11091–11104. PMLR / OpenReview.net, 2024. URL https://proceeding s.mlr.press/v235/ding24i.html.
Guha, N., Nyarko, J., Ho, D. E., Ré, C., Chilton, A., Aditya, K., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D. N., Zambrano, D., Talisman, D., Hoque, E., Surani, F., Fagan, F., Sarfaty, G., Dickinson, G. M., Porat, H., Hegland, J., Wu, J., Nudell, J., Niklaus, J., Nay, J. J., Choi, J. H., Tobia, K., Hagan, M., Ma, M., Livermore, M. A., Rasumov-Rahe, N., Holzenberger, N., Kolt, N., Henderson, P., Rehaag, S., Goel, S., Gao, S., Williams, S., Gandhi, S., Zur, T., Iyer, V., and Li, Z. LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/pap er/2023/hash/89e44582fd28ddfea1ea4dc b0ebbf4b0-Abstract-Datasets_and_Benc hmarks.html.
Du, Y., Tian, M., Ronanki, S., Rongali, S., Bodapati, S. B., Galstyan, A., Wells, A., Schwartz, R., Huerta, E. A., and Peng, H. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. In Christodoulopoulos, C.,
Hong, K., Troynikov, A., and Huber, J. Context Rot: How Increasing Input Tokens Impacts LLM Performance. Technical report, Chroma, July 2025. URL https: //research.trychroma.com/context-rot. 10
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., and Schmidhuber, J. MetaGPT: Meta Programming for A Multi-agent Collaborative Framework. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https: //openreview.net/forum?id=VtmBAGCN7o.
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A. P., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. Natural Questions: a Benchmark for Question Answering Research. Trans. Assoc. Comput. Linguistics, 7:452–466, 2019. doi: 10.1162/TACL A 00276. URL https://doi.org/10.1162/tacl_a_00276.
Hsieh, C., Chuang, Y., Li, C., Wang, Z., Le, L. T., Kumar, A., Glass, J. R., Ratner, A., Lee, C., Krishna, R., and Pfister, T. Found in the middle: Calibrating Positional Attention Bias Improves Long Context Utilization. In Ku, L., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, Findings of ACL, pp. 14982–14995. Association for Computational Linguistics, 2024a. doi: 10.18653/V1/ 2024.FINDINGS-ACL.890. URL https://doi.or g/10.18653/v1/2024.findings-acl.890.
Lee, J., Chen, A., Dai, Z., Dua, D., Sachan, D. S., Boratko, M., Luan, Y., Arnold, S., Perot, V., Dalmia, S., Hu, H., Lin, X., Pasupat, P., Amini, A., Cole, J. R., Riedel, S., Naim, I., Chang, M.-W., and Guu, K. LOFT: Scalable and more realistic long-context evaluation. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, pp. 6713–6738, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-889176-195-7. doi: 10.18653/v1/2025.findings-naacl.374. URL https://aclanthology.org/2025.fi ndings-naacl.374/.
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B. RULER: What’s the real context size of your long-context language models? In First Conference on Language Modeling, 2024b. URL https: //openreview.net/forum?id=kIoBbc76Sy.
Lee, S., Jo, Y., Seo, M., Lee, M., and Seo, M. Lost in the Noise: How Reasoning Models Fail with Contextual Distractors. CoRR, abs/2601.07226, 2026. doi: 10.48550 /ARXIV.2601.07226. URL https://doi.org/10 .48550/arXiv.2601.07226.
Jin, B., Yoon, J., Han, J., and Arik, S. Ö. Long-context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https: //openreview.net/forum?id=oU3tpaR8fm.
Levy, M., Jacoby, A., and Goldberg, Y. Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models. In Ku, L., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pp. 15339– 15353. Association for Computational Linguistics, 2024. doi: 10.18653/V1/2024.ACL- LONG.818. URL https://doi.org/10.18653/v1/2024.acl -long.818.
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Barzilay, R. and Kan, M.-Y. (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://aclantho logy.org/P17-1147/. Kamradt, G. Needle In A Haystack - pressure testing LLMs, 2023. URL https://github.com/gkamradt/ LLMTest_NeedleInAHaystack. GitHub repository.
Levy, S., Mazor, N., Shalmon, L., Hassid, M., and Stanovsky, G. More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, pp. 19539–19547. Association for Computational Linguistics, 2025. URL https://aclantho logy.org/2025.findings-emnlp.1064/.
Ke, W., Zheng, Y., Li, Y., Xu, H., Nie, D., Wang, P., and He, Y. Large Language Models in Document Intelligence: A Comprehensive Survey, Recent Advances, Challenges, and Future Trends. ACM Trans. Inf. Syst., 44(1):18:1– 18:64, 2026. doi: 10.1145/3768156. URL https: //doi.org/10.1145/3768156.
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B. CAMEL: Communicative Agents for ”Mind” Exploration of Large Language Model Society. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural In11
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
formation Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/pap er/2023/hash/a3621ee907def47c1b952ad e25c67698-Abstract-Conference.html.
Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id= dHng2O0Jjr. Ram, D., Xia, W., and Soatto, S. Learning to Focus: Focal Attention for Selective and Scalable Transformers. CoRR, abs/2511.06818, 2025. doi: 10.48550/ARXIV.2511.0681 8. URL https://doi.org/10.48550/arXiv.2 511.06818.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguistics, 12:157–173, 2024. doi: 10.1162/ TACL A 00638. URL https://doi.org/10.116 2/tacl_a_00638.
Ryan, N. Introducing a learnable temperature value into the softmax self-attention scores, 8 2024.
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H. When Not to Trust Language Models: Investigating Effectiveness of Parametric and Nonparametric Memories. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 9802–9822. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL -LONG.546. URL https://doi.org/10.18653 /v1/2023.acl-long.546.
Schick, T., Dwivedi-Yu, J., Dessı̀, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023. URL http://papers.nips.cc/paper_files/pap er/2023/hash/d842425e4bf79ba039352da 0f658a906-Abstract-Conference.html.
OpenAI. Introducing Deep Research. https://openai .com/index/introducing-deep-research/, 2025. Accessed: 2025-01-20.
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, Proceedings of Machine Learning Research, pp. 31210–31227. PMLR, 2023. URL https://proceedings.mlr.press/v202/s hi23a.html.
Peng, B., Quesnelle, J., Fan, H., and Shippole, E. YaRN: Efficient Context Window Extension of Large Language Models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https: //openreview.net/forum?id=wHBfxhZu1u. Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., Cao, N. D., Thorne, J., Jernite, Y., Karpukhin, V., Maillard, J., Plachouras, V., Rocktäschel, T., and Riedel, S. KILT: a Benchmark for Knowledge Intensive Language Tasks. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tür, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pp. 2523–2544. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.NAA CL-MAIN.200. URL https://doi.org/10.186 53/v1/2021.naacl-main.200.
Singh, A., Ehtesham, A., Kumar, S., and Khoei, T. T. Agentic Retrieval-augmented Generation: A Survey on Agentic RAG. CoRR, abs/2501.09136, 2025. doi: 10.48550/ARXIV.2501.09136. URL https://doi. org/10.48550/arXiv.2501.09136. Su, X., Yu, Z., Cui, Y., Liu, A., Lin, X., Wang, Y., Liang, H., Li, W., Shen, L., and Cao, X. Dynamic Analysis and Adaptive Discriminator for Fake News Detection. In Gurrin, C., Schoeffmann, K., Zhang, M., Rossetto, L., Rudinac, S., Dang-Nguyen, D., Cheng, W., Chen, P., and Benois-Pineau, J. (eds.), Proceedings of the 33rd ACM International Conference on Multimedia, MM 2025, Dublin, Ireland, October 27-31, 2025, pp. 8164–8173. ACM, 2025. doi: 10.1145/3746027.3755337. URL ht tps://doi.org/10.1145/3746027.3755337.
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., and Sun, M. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. In The Twelfth International 12
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Team, G. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024. URL https: //arxiv.org/abs/2403.05530.
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 2369–2380. Association for Computational Linguistics, 2018. doi: 10.18653/V1/D18-1259. URL https://doi.org/10.18653/v1/d18-1259.
Wang, Z., Zhang, H., Li, X., Huang, K., Han, C., Ji, S., Kakade, S. M., Peng, H., and Ji, H. Eliminating Position Bias of Language Models: A Mechanistic Approach. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https: //openreview.net/forum?id=fvkElsJOsN. Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., and Wang, C. AutoGen: Enabling Next-gen LLM Applications via Multi-agent Conversation, 2023. URL https://arxiv.org/abs/2308 .08155.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/forum ?id=WE_vluYUL-X.
Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y. Retrieval Head Mechanistically Explains Long-context Factuality. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https: //openreview.net/forum?id=EytBpUGB1Z.
Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D. HELMET: How to evaluate long-context models effectively and thoroughly. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview .net/forum?id=293V3bJbmE.
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M. InfLLM: Training-free Longcontext Extrapolation for LLMs with an Efficient Context Memory. In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J. M., and Zhang, C. (eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URL http://papers .nips.cc/paper_files/paper/2024/hash /d842425e4bf79ba039352da0f658a906-Abs tract-Conference.html.
Yoran, O., Wolfson, T., Ram, O., and Berant, J. Making Retrieval-augmented Language Models Robust to Irrelevant Context. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https: //openreview.net/forum?id=ZS4m74kZpH. Zhang, W., Li, Y., Bei, Y., Luo, J., Wan, G., Yang, L., Xie, C., Yang, Y., Huang, W., Miao, C., Zou, H. P., Luo, X., Zhao, Y., Chen, Y., Chan, C., Zhou, P., Zhang, X., Zhang, C., Shang, J., Zhang, M., Song, Y., King, I., and Yu, P. S. From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents. CoRR, abs/2506.18959, 2025a. doi: 10.48550/ARXIV.250 6.18959. URL https://doi.org/10.48550/a rXiv.2506.18959.
Yang, M., Huang, E., Zhang, L., Surdeanu, M., Wang, W. Y., and Pan, L. How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pp. 13329–13347. Association for Computational Linguistics, 2025a. doi: 10.18653/V1/2025.EMNLP-M AIN.674. URL https://doi.org/10.18653/v 1/2025.emnlp-main.674.
Zhang, W., Yin, F., Yen, H., Chen, D., and Ye, X. Queryfocused Retrieval Heads Improve Long-context Reasoning and Re-ranking. In Christodoulopoulos, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pp. 23791–23805. Association for Computational Linguistics, 2025b. doi: 10.18653/V1/2025.EMNLP- MAIN.1214. URL https://doi.org/10.18653/v1/2025.emn lp-main.1214.
Yang, Y., Huang, Z., Zhu, W., Qiu, Z., Yuan, F., Pan, J. Z., and Titov, I. A Controllable Examination for Long-context Language Models. CoRR, abs/2506.02921, 2025b. doi: 10.48550/ARXIV.2506.02921. URL http s://doi.org/10.48550/arXiv.2506.02921.
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M. K., Han, X., Thai, Z. L., Wang, S., Liu, Z., and Sun, M. 13
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
∞Bench: Extending Long Context Evaluation Beyond 100K Tokens, 2024a. URL https://arxiv.org/ abs/2402.13718. Zhang, Z., Chen, R., Liu, S., Yao, Z., Ruwase, O., Chen, B., Wu, X., and Wang, Z. Found in the Middle: How Language Models Use Long Contexts Better via Plugand-play Positional Encoding. In Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J. M., and Zhang, C. (eds.), Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024b. URL http://papers.nips.cc/paper_files /paper/2024/hash/6ffdbbe354893979367 f93e2121e37dd-Abstract-Conference.ht ml.
14
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
A. Detailed Experiment Results As mentioned in §3.2, below are results for all the models across different settings. Table 2. Accuracy (%) of Llama-3.2-1B-Instruct across different hard distractor proportions and context lengths. Hard % indicates the proportion of hard distractors, with the remaining being easy distractors (Easy) or random Wikipedia passages (Random). Easy Hard %
4K
Random
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
76.0 75.0 75.0 75.0 71.0 70.0 65.5 61.5 62.0 60.0 52.5 57.5
68.5 58.5 50.5 50.5 48.0 38.0 38.5 36.0 34.0 26.0 29.0 31.0
61.0 43.5 43.0 43.0 42.5 37.0 35.0 30.0 24.5 22.0 22.5 26.5
66.0 51.0 46.5 40.0 40.5 30.5 31.0 25.0 27.0 21.0 19.0 23.5
53.0 43.5 37.0 36.5 34.5 37.0 35.5 33.0 28.0 26.0 26.0 24.0
55.5 35.0 31.0 29.5 30.0 29.0 29.0 29.5 25.0 24.0 22.0 23.0
75.0 69.5 70.0 69.0 70.5 68.5 67.0 65.0 54.0 58.5 57.0 57.5
50.5 37.0 47.0 46.5 40.0 39.0 34.0 30.0 31.5 30.0 27.0 31.0
30.5 32.5 32.5 34.0 30.5 30.5 31.0 24.5 28.0 21.5 20.5 27.0
24.0 28.0 32.0 30.0 22.5 23.5 22.0 21.0 23.5 24.0 21.0 24.0
38.0 33.5 29.5 31.0 28.5 32.5 26.0 27.0 23.0 22.0 28.0 24.5
30.5 27.0 28.5 23.5 29.0 26.5 25.5 24.5 22.5 24.0 25.5 23.0
84.5 77.5 77.5 77.5 74.0 71.5 66.5 66.5 62.0 58.5 61.0 63.5
78.0 64.5 56.5 56.5 52.0 43.0 50.5 41.5 37.5 34.0 40.5 40.5
72.5 56.0 51.5 50.5 45.5 46.0 43.0 40.0 34.0 32.5 32.5 33.0
78.5 55.5 51.0 48.0 49.5 43.0 41.0 35.5 31.0 29.0 30.0 31.0
72.0 57.0 51.0 54.0 53.5 46.5 43.0 40.5 36.0 27.0 26.5 25.0
68.5 44.5 43.5 46.5 45.0 35.0 35.5 31.0 26.0 21.5 23.5 25.0
74.5 74.0 73.0 76.0 75.5 70.5 69.5 70.0 63.5 63.0 60.5 63.5
58.5 58.0 56.5 57.5 57.5 56.0 49.5 50.0 41.0 42.5 43.0 40.5
56.0 46.5 47.5 50.0 47.0 41.5 39.0 35.0 36.5 29.5 32.5 33.0
42.0 44.0 41.5 43.5 37.5 34.0 33.0 31.5 32.0 31.5 27.0 31.0
54.5 52.0 49.5 49.5 44.5 42.0 35.0 29.5 29.5 29.0 27.5 25.5
41.5 43.5 39.0 42.5 36.5 31.0 26.5 28.0 26.0 20.5 21.5 25.0
88.5 85.0 85.0 85.0 83.5 81.0 77.5 66.0 64.5 62.5 63.0 63.5
81.0 76.0 68.0 68.0 62.5 54.0 46.5 50.5 41.0 37.0 42.0 42.0
74.0 63.0 60.0 60.5 55.0 55.0 47.5 39.5 34.0 32.5 35.0 43.0
72.0 61.5 57.0 59.0 55.0 50.0 41.0 35.0 32.0 27.5 28.5 27.5
64.0 53.5 52.5 47.0 45.5 46.0 48.0 39.5 37.0 33.5 28.5 29.5
61.5 48.0 46.5 43.0 38.0 43.5 39.0 32.0 33.5 30.5 32.5 31.5
84.0 84.0 86.5 86.5 77.5 78.5 77.5 73.5 71.0 66.5 62.0 64.0
72.5 59.5 61.5 56.5 52.0 61.0 60.0 42.0 42.5 37.0 48.0 42.0
37.5 46.0 44.0 41.5 42.5 45.0 39.0 39.5 35.5 32.5 35.0 43.0
30.5 31.0 30.0 33.5 30.0 28.5 32.0 32.5 36.0 31.0 33.5 27.5
41.0 46.0 42.5 39.0 39.0 37.5 41.5 38.0 34.5 29.0 36.5 29.5
29.5 28.5 30.0 27.5 34.5 34.5 35.5 33.0 33.5 31.0 26.5 31.5
64.5 59.0 59.5 59.0 54.5 54.5 51.0 49.5 42.0 44.0 43.5 43.0
56.5 48.5 48.0 48.0 43.5 43.5 45.0 39.0 37.0 35.5 35.5 33.5
53.0 45.0 45.5 42.0 34.5 37.5 34.5 29.5 29.5 27.0 24.5 27.5
58.5 44.0 46.5 45.5 41.5 33.0 32.0 25.0 26.0 23.0 23.5 23.0
54.5 44.0 46.0 41.0 41.5 33.5 31.5 30.5 22.0 23.0 25.0 25.5
51.0 32.0 28.5 31.5 26.0 26.0 20.0 15.5 22.0 19.0 19.0 18.0
53.5 52.0 52.0 52.0 49.5 51.5 48.5 46.0 46.0 43.5 42.5 43.0
48.5 45.0 46.0 46.5 43.0 43.0 45.0 38.5 39.0 35.5 35.5 33.5
37.0 37.5 36.0 32.5 36.0 37.5 28.0 33.0 31.0 26.0 29.0 27.5
29.0 31.5 26.5 31.5 28.5 25.5 23.5 22.5 20.0 23.5 22.0 22.5
35.0 37.5 35.0 36.5 30.5 33.5 28.0 25.0 23.5 24.0 24.5 26.0
26.0 23.0 23.0 22.0 20.5 19.0 21.5 19.0 17.0 17.0 21.0 18.0
Natural Questions 0 1 2 3 5 10 20 40 60 80 90 100 TriviaQA 0 1 2 3 5 10 20 40 60 80 90 100 PopQA 0 1 2 3 5 10 20 40 60 80 90 100 HotpotQA 0 1 2 3 5 10 20 40 60 80 90 100
15
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning Table 3. Accuracy (%) of Llama-3.1-8B-Instruct across different hard distractor proportions and context lengths. Hard % indicates the proportion of hard distractors, with the remaining being easy distractors (Easy) or random Wikipedia passages (Random). Easy Hard %
4K
Random
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
89.0 89.0 89.0 89.0 90.0 89.5 91.0 85.0 85.5 84.0 85.0 82.5
91.5 89.0 89.5 89.5 87.5 86.0 88.0 86.5 82.0 80.5 79.0 83.0
90.0 87.5 89.5 88.5 88.5 88.0 85.0 79.5 79.5 78.0 75.0 75.5
88.0 86.5 88.0 83.5 85.0 80.0 79.5 75.0 75.0 71.5 73.0 73.5
87.5 87.5 86.5 83.5 82.5 77.5 76.5 73.5 74.0 73.0 72.5 71.0
87.0 85.5 82.0 78.0 76.0 72.5 70.5 66.0 63.5 62.5 64.5 62.0
88.5 89.5 88.5 88.5 90.0 89.0 87.0 87.0 84.0 83.0 84.0 85.0
90.0 91.0 90.5 90.0 86.0 88.5 85.5 83.0 84.0 83.0 83.0 84.0
90.0 88.5 90.5 86.0 89.0 86.5 83.5 82.5 79.0 79.0 75.5 77.0
90.0 83.5 83.5 83.5 77.5 79.5 81.0 77.0 72.5 71.5 77.5 74.0
86.0 86.0 81.5 81.0 79.0 78.5 77.5 72.5 71.5 70.5 70.0 71.0
82.5 76.5 76.0 75.5 71.0 66.0 64.0 62.5 63.0 65.5 62.0 62.5
97.0 95.0 95.0 95.0 95.0 94.0 94.0 93.5 92.5 93.0 93.0 90.0
96.5 94.5 93.5 94.0 94.0 95.0 93.5 89.5 91.0 91.5 89.5 90.0
97.0 95.0 94.5 95.0 94.0 93.0 93.0 90.5 91.0 90.0 91.0 88.0
96.0 93.0 94.0 94.5 90.0 90.5 89.0 88.0 88.0 85.5 84.5 84.5
96.5 93.0 93.5 92.0 90.0 90.5 89.0 86.5 85.5 81.5 84.0 83.5
96.0 92.5 89.0 90.5 89.5 83.5 82.5 82.0 80.0 78.0 78.0 77.5
97.5 97.0 95.5 95.0 97.0 95.5 95.5 93.5 91.0 93.0 92.0 91.0
95.5 95.0 95.0 95.5 94.0 93.0 93.5 94.0 91.5 90.5 90.0 91.0
95.0 95.5 93.5 94.0 93.0 91.5 92.5 90.0 89.5 88.0 88.5 88.0
94.0 94.5 93.0 93.0 90.5 91.0 88.5 87.5 87.5 85.0 83.5 84.0
93.5 92.5 92.0 92.0 87.5 88.0 88.0 85.0 81.5 84.0 83.5 83.5
93.5 87.0 87.5 86.0 86.0 83.5 80.0 80.0 76.0 76.0 76.0 77.0
96.0 97.0 96.5 97.5 93.5 96.0 96.0 93.0 94.0 92.5 90.0 89.0
96.0 97.0 95.5 95.5 95.5 94.5 94.5 92.5 93.0 91.0 91.5 89.0
94.5 94.5 94.5 94.5 93.0 93.0 91.5 91.0 91.0 89.5 87.5 92.0
94.0 94.5 93.0 94.0 94.0 92.0 87.0 84.5 84.0 85.5 81.5 89.0
94.5 92.5 92.0 92.0 88.0 84.0 80.5 83.0 76.0 75.5 76.5 78.0
92.0 92.5 87.5 86.0 84.0 78.5 77.5 75.5 71.0 70.5 71.5 75.0
96.0 97.0 97.5 96.0 95.0 95.5 95.5 95.0 95.0 92.0 92.0 89.0
96.5 97.5 98.0 97.5 96.5 96.0 94.0 93.5 93.0 90.5 91.5 89.0
98.0 96.0 96.5 97.0 96.5 96.0 95.0 91.5 91.5 88.5 90.0 92.0
96.0 95.0 93.0 95.5 94.0 90.5 86.5 85.5 84.5 84.0 82.0 87.0
95.0 95.5 95.5 90.0 90.0 86.0 85.0 76.0 74.0 72.0 76.0 78.0
92.5 90.0 90.0 88.5 83.5 78.5 76.0 73.0 72.5 70.5 68.0 75.0
86.5 85.0 85.0 85.0 81.5 77.5 75.0 70.5 69.0 70.5 73.0 67.0
85.5 82.5 76.0 76.5 74.5 74.5 70.0 66.0 65.5 65.5 65.5 64.0
83.5 76.5 72.0 72.0 69.0 68.5 64.5 61.0 61.5 58.0 61.0 57.0
80.0 71.5 69.5 68.0 65.5 60.5 58.0 61.0 56.0 58.0 57.0 58.0
81.0 71.5 71.0 66.5 66.5 63.5 60.5 62.0 60.0 60.5 59.0 62.0
70.0 60.0 57.5 57.5 51.0 47.5 47.5 46.5 44.0 45.5 45.5 48.0
80.0 74.5 77.5 73.5 74.0 75.5 75.0 74.5 70.5 71.5 68.5 67.0
71.0 67.5 71.5 71.5 68.0 64.5 65.0 65.0 64.5 63.5 63.5 64.0
67.0 67.5 65.5 65.5 64.5 63.0 60.5 58.0 59.0 54.5 58.0 57.0
60.5 61.0 60.5 61.0 56.0 57.0 55.0 54.0 57.0 54.0 55.5 58.0
61.0 61.5 62.5 61.0 62.0 59.0 58.0 60.0 54.5 55.5 56.0 63.0
49.0 51.0 49.0 44.5 45.0 43.5 38.0 41.5 45.0 41.5 44.0 48.0
Natural Questions 0 1 2 3 5 10 20 40 60 80 90 100 TriviaQA 0 1 2 3 5 10 20 40 60 80 90 100 PopQA 0 1 2 3 5 10 20 40 60 80 90 100 HotpotQA 0 1 2 3 5 10 20 40 60 80 90 100
16
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Table 4. Accuracy (%) of Qwen2.5-7B-Instruct across different hard distractor proportions and context lengths. Hard % indicates the proportion of hard distractors, with the remaining being easy distractors (Easy) or random Wikipedia passages (Random). Easy Hard %
4K
Random
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
84.5 82.0 81.5 83.5 80.0 82.0 77.0 75.0 76.0 73.0 73.5 74.0
83.5 78.0 76.0 77.0 79.5 75.0 71.0 69.0 67.0 66.5 67.5 63.0
75.5 75.5 74.5 73.0 72.0 69.0 69.5 67.5 61.5 62.0 65.0 62.0
75.0 73.0 70.5 72.5 72.0 65.5 58.5 53.0 48.5 57.0 53.5 55.5
75.5 71.0 71.5 68.0 67.5 65.0 62.0 54.0 57.0 51.0 49.0 54.5
68.0 69.0 52.5 51.5 42.5 45.5 34.0 41.0 40.0 30.5 29.5 32.5
78.5 72.0 80.5 74.0 70.5 72.0 74.5 75.0 76.5 73.0 71.0 73.5
71.5 71.5 73.5 69.0 72.0 68.0 72.0 69.0 67.0 63.5 66.5 64.5
70.5 69.5 70.5 68.5 72.0 72.5 68.0 66.0 66.5 62.5 60.5 61.0
74.5 64.0 65.0 61.5 65.5 64.5 54.5 54.5 53.5 47.0 45.5 55.0
70.0 60.5 58.0 62.0 63.0 58.0 55.5 51.5 50.0 56.0 46.0 54.0
44.5 35.0 36.5 32.0 36.5 32.5 35.5 33.0 33.5 32.5 33.0 32.0
89.0 86.5 86.5 86.5 85.5 81.0 83.0 76.0 77.0 75.0 74.5 74.0
87.5 84.0 84.5 83.5 81.5 79.0 76.5 74.0 70.5 74.0 73.5 70.5
89.5 79.0 82.5 79.5 76.0 76.5 76.5 69.5 72.0 70.5 68.5 70.0
83.0 75.0 74.5 71.5 67.5 64.0 65.5 58.5 56.0 63.5 60.5 59.5
82.0 73.0 69.5 66.5 64.0 59.0 58.5 57.5 53.5 52.0 58.0 57.0
81.0 63.0 56.5 53.0 44.5 41.0 44.0 34.5 34.0 33.5 32.5 32.0
88.5 86.0 85.5 87.0 87.0 86.0 81.5 80.5 77.0 72.5 74.5 74.5
84.0 79.5 81.5 80.0 80.5 74.5 74.0 74.0 71.0 70.5 72.5 70.0
78.5 78.0 80.0 76.5 78.5 74.5 75.0 74.0 73.0 69.0 67.5 70.0
74.5 74.0 72.0 73.0 68.0 70.5 59.5 64.5 55.5 62.5 60.0 59.0
67.0 65.0 65.0 61.0 66.5 66.5 58.0 55.0 52.0 48.5 54.5 56.5
42.5 39.0 31.5 34.5 35.0 34.0 27.5 29.5 30.5 30.0 28.0 32.5
93.0 90.0 89.0 89.5 88.0 89.0 89.0 84.5 82.0 83.5 80.0 84.0
93.0 92.5 88.0 87.5 88.5 81.5 83.0 83.0 78.0 76.5 77.5 70.0
88.5 83.0 85.0 85.0 87.0 84.5 80.0 80.0 71.5 74.5 75.5 73.0
88.5 87.0 83.5 81.0 77.5 76.0 72.0 71.0 63.5 60.0 61.0 58.5
89.5 83.0 75.0 70.5 72.0 70.5 63.5 62.0 53.0 54.5 50.0 50.5
86.0 68.0 60.0 61.5 48.0 47.0 45.0 37.0 38.0 34.5 27.5 28.0
91.0 91.5 90.0 91.5 89.5 89.0 88.5 87.5 87.0 84.5 85.5 84.0
89.5 87.5 88.0 87.5 86.0 81.5 79.5 81.5 76.5 71.0 76.0 70.5
82.0 86.0 85.5 83.5 83.5 80.0 79.5 75.5 72.5 76.0 77.0 72.5
82.0 77.0 77.5 77.0 76.5 72.5 73.0 65.0 63.0 60.0 60.5 57.5
74.5 71.5 71.0 67.0 63.5 62.0 64.0 55.5 51.5 49.5 46.5 53.0
91.0 41.0 39.5 34.0 36.0 42.0 32.0 32.5 33.5 27.0 29.5 29.0
75.5 73.5 75.0 74.5 74.0 73.0 63.5 68.5 65.0 69.5 64.0 66.5
76.0 73.0 75.5 73.5 69.0 64.0 63.5 55.0 58.0 53.5 55.5 56.5
77.0 72.5 65.5 67.5 68.0 65.0 61.0 52.0 53.5 48.5 49.5 48.0
73.5 67.5 60.5 61.5 59.0 55.5 54.5 48.5 44.0 44.5 41.0 40.0
70.0 56.0 56.5 55.5 50.0 49.0 44.5 43.0 38.0 42.5 45.0 39.5
68.5 52.0 44.0 39.0 39.5 35.0 27.5 24.0 24.5 24.5 22.0 18.5
71.0 68.5 68.5 67.5 67.0 62.0 67.0 64.5 61.5 65.5 64.0 67.0
67.0 62.0 61.5 65.5 63.0 62.5 57.0 54.5 55.0 49.5 53.0 58.0
63.5 65.0 63.0 56.0 58.0 60.0 48.0 53.0 48.5 47.0 44.5 48.0
62.5 63.0 56.5 60.5 56.5 53.5 45.0 46.5 44.5 42.5 38.5 41.0
50.5 53.5 51.5 53.0 45.5 45.5 44.0 37.5 41.0 38.5 41.5 38.5
24.5 26.5 29.5 25.0 26.5 21.0 22.0 22.0 23.5 23.5 21.0 19.0
Natural Questions 0 1 2 3 5 10 20 40 60 80 90 100 TriviaQA 0 1 2 3 5 10 20 40 60 80 90 100 PopQA 0 1 2 3 5 10 20 40 60 80 90 100 HotpotQA 0 1 2 3 5 10 20 40 60 80 90 100
17
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Table 5. Accuracy (%) of Qwen3-Next-80B-Instruct across different hard distractor proportions and context lengths. Hard % indicates the proportion of hard distractors, with the remaining being easy distractors (Easy) or random Wikipedia passages (Random). Easy Hard %
4K
Random
8K
16K
32K
64K
128K
4K
8K
16K
32K
64K
128K
93.0 92.0 91.5 92.0 92.0 91.5 91.0 91.0 88.5 88.0 87.5 88.0
92.0 92.0 92.0 92.0 90.5 88.5 90.5 89.5 88.0 87.0 87.5 85.0
93.5 91.0 90.5 90.0 85.5 88.5 88.5 86.5 84.5 86.5 82.5 84.0
92.0 90.5 87.5 88.0 89.5 86.5 88.0 83.5 84.5 83.0 84.5 83.5
91.0 87.0 87.0 88.5 86.0 87.0 84.5 83.5 83.5 80.0 83.5 82.5
91.5 88.0 89.5 83.0 87.0 84.5 81.5 82.0 79.5 79.0 80.0 77.5
94.0 91.5 90.0 92.5 90.0 91.0 89.0 87.0 86.5 87.5 87.0 88.0
90.5 91.5 90.0 88.5 89.5 86.5 88.5 87.5 88.0 88.5 86.0 84.5
90.0 90.0 90.0 90.5 90.5 87.5 86.5 89.5 85.0 81.0 83.5 85.5
89.5 89.0 90.5 88.0 88.5 85.5 85.5 84.0 84.0 82.5 83.5 84.0
90.5 90.5 90.5 89.5 87.0 85.0 85.5 81.5 82.0 81.5 81.5 80.5
91.5 87.0 85.0 85.5 85.5 82.0 82.0 79.5 78.5 78.0 77.0 76.5
99.5 98.0 98.5 98.5 97.5 96.5 96.5 95.0 95.5 96.0 95.5 93.5
99.5 98.5 97.0 97.5 98.0 95.5 94.5 95.5 95.5 96.0 94.0 94.5
99.5 98.0 97.5 96.5 96.5 94.0 94.5 95.5 93.0 95.0 92.0 94.0
99.5 96.5 95.5 94.0 94.0 94.5 94.0 93.0 94.5 92.0 92.5 90.5
99.5 96.0 96.5 95.0 97.0 95.0 94.5 94.0 92.0 93.5 92.5 90.0
99.5 95.5 94.0 95.0 96.0 95.5 93.0 92.0 89.5 89.0 88.0 87.0
98.0 95.5 97.0 95.5 95.0 96.5 95.0 94.0 96.5 94.5 96.0 95.0
97.5 97.5 97.0 97.0 96.5 94.5 95.0 95.5 94.0 95.0 92.0 93.0
98.0 97.0 97.0 96.0 96.0 94.5 96.0 94.5 94.5 93.5 94.5 95.0
98.5 98.0 97.0 97.5 96.5 95.0 94.5 94.0 93.0 92.5 92.0 91.5
98.5 97.0 96.5 96.0 96.0 93.5 96.5 93.0 92.5 90.0 91.0 90.5
98.0 96.5 95.0 94.0 93.5 94.0 93.0 92.5 89.0 85.5 86.5 87.0
97.5 98.0 97.5 97.5 96.5 96.5 97.5 98.5 97.0 96.0 97.0 96.0
97.5 98.0 97.5 97.5 97.5 95.5 96.5 95.5 94.5 93.5 92.5 92.5
98.0 97.0 97.0 97.0 97.5 97.0 95.0 91.0 91.0 93.5 93.0 91.0
98.0 96.5 97.5 97.5 97.5 93.5 94.5 89.5 92.0 86.0 88.5 87.0
98.0 97.5 97.0 97.0 97.5 95.0 93.0 88.0 86.5 86.0 86.0 84.0
98.0 97.0 96.0 94.5 93.0 94.0 87.5 86.0 84.5 82.5 81.5 80.0
98.5 98.5 98.0 98.0 97.0 97.0 97.5 97.0 98.0 97.0 96.0 95.0
98.5 98.5 97.5 98.5 97.5 97.5 97.0 96.0 94.0 92.0 91.5 92.0
99.0 98.0 97.0 98.5 98.5 97.0 96.5 90.5 91.5 92.5 92.5 91.5
98.5 97.0 97.5 97.0 96.0 95.5 94.5 91.0 88.0 87.5 85.5 87.5
98.5 97.5 96.5 96.5 96.0 93.5 91.0 88.5 87.0 85.0 84.5 84.0
97.5 95.0 94.0 92.0 92.0 89.5 89.0 84.5 81.5 78.5 78.5 81.5
88.5 85.0 85.5 85.5 88.0 86.5 83.5 83.0 82.5 81.0 80.5 81.5
90.5 87.5 88.0 88.0 84.0 85.0 83.0 80.5 77.0 79.5 80.0 79.0
88.0 86.0 81.5 82.0 79.5 78.5 75.5 73.0 72.0 73.5 73.0 71.5
89.0 85.0 83.5 84.5 81.0 80.0 75.0 74.0 74.0 73.5 74.0 73.0
90.5 88.5 84.5 84.5 82.5 80.5 79.0 77.0 74.5 76.0 71.5 71.5
90.0 84.5 82.0 78.5 75.0 76.0 68.5 70.0 67.5 67.0 68.5 69.0
85.0 85.0 83.0 85.0 85.0 85.0 81.0 83.0 82.5 82.0 79.0 81.5
84.5 84.5 85.0 85.5 84.0 84.0 82.0 81.5 79.0 81.5 80.5 79.5
82.0 79.5 80.0 79.5 78.0 77.5 75.5 75.0 73.5 73.0 72.5 72.5
81.0 81.0 80.5 82.0 79.0 76.0 74.0 71.0 70.5 71.0 72.0 74.5
82.0 80.0 77.5 81.0 79.5 76.5 74.0 72.5 67.0 69.5 70.0 72.0
69.0 69.0 69.5 68.5 67.0 67.0 64.0 62.0 64.5 67.0 64.0 69.5
Natural Questions 0 1 2 3 5 10 20 40 60 80 90 100 TriviaQA 0 1 2 3 5 10 20 40 60 80 90 100 PopQA 0 1 2 3 5 10 20 40 60 80 90 100 HotpotQA 0 1 2 3 5 10 20 40 60 80 90 100
18
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
B. Margin Computation (∆e and ∆h ) As mentioned in §5, we calculate the margin for Llama-3.1-8b-Instruct and Llama-3.2-1b-Instruct. Below are results for Llama-3.2-1b-Instruct and the correlation results for both models.
Llama-3.2-1B-Instruct: easy
hard
Delta Value (attention logits)
10
easy hard
8 6
Avg Margin: 6.52 8.0
7.2
1
20
6.9
6.3
5.7
5.0
40
60
80
90
4 2
Hard Proportion (%)
Figure 10. Empirical measurement of logit margins ∆e and ∆h on retrieval heads for Llama-3.2-1B-Instruct. Green bars show ∆e (margin to easy distractors) and brown bars show ∆h (margin to hard distractors). The gap ∆e − ∆h remains substantial across all hard proportions, with an average of 6.52, which is more significant than the gap of the 8B model. This validates the theoretical assumption ∆h ≪ ∆ e . Table 6. Per-file train–test correlations for selected hard proportions on nq easy.
Hard Proportion 1% 20% 40% 60% 80% 90%
8B Pearson
8B Spearman
1B Pearson
1B Spearman
0.9498 0.9619 0.9753 0.9725 0.9508 0.9491
0.9838 0.9854 0.9915 0.9894 0.9832 0.9850
0.9599 0.9592 0.9546 0.9522 0.9522 0.9278
0.9953 0.9962 0.9970 0.9964 0.9922 0.9868
19
T HE F IRST D ROP OF I NK: Nonlinear Impact of Misleading Information in Long-Context Reasoning
C. Prompts Demonstration In this section, we demonstrate the prompts used for: (1)evaluating model’s output and (2) verifying the distractors not containing the answers. Evaluation Prompt You are an expert evaluator for question-answering systems. Your task is to determine if a model’s answer to a question is correct based on the provided documents and reference answer. # Documents {context} # Question {question} # Reference Answer {reference answer} # Model’s Answer {model answer} # Evaluation Task Based on the information in the provided documents and the reference answer, evaluate whether the model’s answer is correct. The answer is CORRECT if: 1. It matches or is semantically equivalent to the reference answer 2. It accurately answers the question using information from the documents 3. It does not contain extra hallucinated or incorrect information The answer should be considered CORRECT even if: • It uses slightly different wording but conveys the same meaning • It uses synonyms or alternative names for the same entity • It is a shorter or longer form of the reference answer (e.g., “Steven Weber” vs “Steven Robert Weber”) Respond with ONLY one of these two options: • CORRECT • INCORRECT
Distractor Verification Prompt You are evaluating whether a set of documents contains the answer to a given question. # Question {question} # Accepted Answers {’, ’.join(answer)} # Documents {docs text} # Task Determine if ANY of the documents above EXPLICITLY contain information that would allow someone to answer the question with one of the accepted answers. Please respond in the following JSON format: { "has_answer": true/false, "explanation": "Brief explanation...", "confidence": "high/medium/low" } # Critical Rules • You must ONLY use information that is EXPLICITLY stated in the provided documents • Do NOT use any external knowledge or make inferences based on your own knowledge • Do NOT assume facts that are not directly written in the documents • Answer ”has answer”: true ONLY if the answer is EXPLICITLY written in the documents • Consider synonyms and paraphrases (e.g., ”USA” and ”United States” are equivalent) • If the documents mention the topic but don’t EXPLICITLY contain the specific answer, count it as false # Examples • INVALID reasoning: ”Green Day is known for punk rock” - this uses external knowledge • VALID reasoning: ”The document states ’punk rock band Green Day’” - this quotes the document
20