LENS: The Sum Is Worse Than the Parts for Set-Level Poisoning in Retrieval-Augmented Generation Kaisheng Fan1 , Yishu Gao1 , Xunzhu Tang2 , Tegawend’e F. Bissyand’e2 , Weizhe Zhang1,3∗
arXiv:2609.35155v1 [cs.CR] 28 Sep 2026
1
School of Cyber Science and Technology, Harbin Institute of Technology, Harbin, China 2 SnT, University of Luxembourg, Luxembourg City, Luxembourg 3 Department of New Networks, Peng Cheng Laboratory, Shenzhen, China {fankaisheng, gaoyishu}@stu.hit.edu.cn, [email protected] {xunzhu.tang, tegawende.bissyande}@uni.lu Abstract
or drawn from open sources, poisoned documents that enter the retrievable corpus may be selected at inference time and shape the final answer. Many RAG poisoning attacks expose locally sufficient or individually detectable cues by directly stating the target answer, adding answer-shaped scaffolding (Zou et al. 2025), optimizing adversarial strings (Ben-Tov and Sharif 2025; Wang et al. 2026b), or using trigger-style poisoned documents (Chaudhari et al. 2026). Several defenses screen passages through document-local abnormality, conflict, answer support, or isolation tests (Xiang et al. 2026; Shen et al. 2025; Si et al. 2025). This leaves a different failure mode less explored: individually plausible documents can remain insufficient in isolation while their composition redirects answer selection. Because generation conditions on the composed retrieved set, passage-wise inspection evaluates a different unit from the one that produces the answer. We show that this mismatch enables set-level compositional poisoning: individually plausible documents occupy complementary semantic roles, so their composition installs an attacker-selected criterion for answer selection while proper subsets remain insufficient. LENS separates this criterion from the facts that support it, placing control in the induced evidence-set decision rule instead of a passage-local payload. The attack operates at inference time against a frozen RAG pipeline, making evidence composition the security boundary. Figure 1 contrasts this setting with single-document poisoning, where one passage is locally sufficient. To construct such attacks, we introduce LENS (Local Evidence, Non-local Steering), a query-aware, generator-blackbox multi-agent pipeline with nested planning and repair loops. The outer loop mines a typed, query-conditioned interpretation lens. This lens specifies how evidence should be read according to role, scope, time, category, naming convention, or referent. It then fixes a document plan that assigns the lens, true or locally supported complementary facts, and claims to avoid for subset safety. The inner loop synthesizes lens-setting and fact-completion documents, evaluates full-set success and subset leakage with an attacker-controlled local surrogate reader, and uses typed counterexamples to repair candidates toward strong full-set steering with low subset leakage. This compositional attack surface creates a new defense objective: identify suspicious cross-document dependence
Retrieval-augmented generation (RAG) aggregates evidence from multiple external documents, yet this joint integration creates an underexamined vulnerability: attack effects absent in individual documents can emerge through set-level composition. Existing coordinated attacks do not explicitly enforce that every proper subset remains insufficient in frozen singleround RAG. We formalize set-level compositional poisoning, where documents designed to remain individually plausible jointly redirect RAG outputs to a target answer, while proper subsets fail to induce the target on their own. To construct such attacks, we propose LENS, a generator-black-box multi-agent framework that casts construction as constrained evidence composition. LENS factorizes target inference into a queryconditioned interpretation lens and complementary facts, then uses a nested dual-loop workflow to concentrate steering in the full set while suppressing subset leakage. The outer loop plans the interpretation lens and semantic roles; the inner loop synthesizes documents and applies counterexample-guided repair. Across four benchmarks and three generators, returned packets achieve 0.852 full-set ASR and 0.784 post-retrieval ASR@5, while their strongest proper subsets reach only 0.069. Against construction baselines evaluated on the same frozen manifest, LENS improves all-attempt E2E-Strict@5 from 0.244 to 0.363, a 48.8% relative gain. A blinded human audit finds that 68.3% of returned packets combine an incorrect target, a definite answer-criterion shift, and no target entailment under the original semantics. Across four published defenses, LENS attains the highest defended all-attempt ASR@5, exceeding the strongest baseline by 0.141 on average. Together, these results establish evidence composition as a distinct RAG security boundary and position LENS as a concrete stress test for defenses that reason over document sets.
Introduction Retrieval-augmented generation (RAG) improves large language models (LLMs) by conditioning generation on a small top-K set of passages retrieved from external corpora (Lewis et al. 2020; Ram et al. 2023; Asai et al. 2024). This externalmemory interface lets LLMs use fresh, domain-specific, or user-provided information without changing model parameters, but it also turns corpus content into a security-critical input. When documents are user-uploaded, weakly curated, ∗
Corresponding author.
1
Single-document Poisoning
Set-level Compositional Poisoning
(risk inside one document)
(risk emerges from a document set)
Local Sufficient Query User
Clean Doc
Query
Retriever
User Poisoned Doc
Jointly Sufficient (Risk Emerges)
LENS insert
LLM Wrong Generator Answer
Retriever
LLM Wrong Generator Answer
Retrieval Corpus
Clean Doc
Retrieval Corpus
LENS insert
Figure 1: Single-document versus set-level compositional poisoning. Left: one poisoned document is locally sufficient to steer generation. Right: plausible LENS documents rarely induce the target in proper subsets, but the full set shifts generation toward the preselected non-gold target.
while preserving the evidence integration that enables legitimate multi-hop reasoning (Yang et al. 2018; Ho et al. 2020; Trivedi et al. 2022). We therefore evaluate attack suppression and clean multi-hop utility together, using their trade-off to characterize composition-aware defenses. This paper makes three contributions. First, we formalize set-level compositional poisoning with an explicit all-propersubset constraint: the full inserted set induces a preselected non-gold target, while every proper subset is constrained to maintain a low target rate. Second, we introduce LENS, a generator-black-box construction method that mines interpretation lenses, synthesizes lens-setting and fact-completion documents, and uses counterexample-guided repair with active subset checking and exhaustive final verification. Third, construction baselines under a shared manifest and evaluation, semantic controls, cross-generator evaluation, and all-attempt defense experiments establish the distinctive value of full-set dependence. LENS improves end-to-end strict set success by 48.8% over the strongest construction baseline and retains the highest attack success under every tested defense. Effective interventions must distinguish suspicious dependence from legitimate multi-hop composition.
independently measures residual ambiguity or compatibility. We study targeted benchmark-answer displacement, where corpus insertion redirects a clean system that outputs g toward t. This defines an attacker-directed output-integrity violation: an untrusted contributor preselects t and attempts to replace the clean-system answer with it. For S ⊆ D, define TR(x) (S) = Pr[Gx (q, X ∪ S) = t],
(1)
where x = sur denotes the construction surrogate and x = (x) d20 (S) for its observed m a target generator. We write TR (sur)
d20 , frequency over 20 generations. Construction uses TR (m) d20 . while final evaluation uses TR Subset dependence. A packet is subset-safe for target generator m at tolerance ϵ when ∀S ⊊ D :
TR(m) (S) ≤ ϵ.
(2)
This is a population property. We report the corresponding empirical all-subset pass rate under the 20-generation protocol, not a probability certificate. We quantify full-set dependence by
Problem Setup and Formalization
CJL(m) (D) = TR(m) (D) − max TR(m) (S).
Attack model. Let C = {D1 , . . . , Dn } be a corpus and X = RK (q; C) its clean top-K context for query q. The adversary inserts a packet D = {d1 , . . . , dk }, yielding C ′ = C ∪ D, but cannot modify the query, original documents, retrieval pipeline, or target generators. Construction uses only an attacker-controlled local surrogate, and target generators provide no attacked-query feedback before packet freezing. We study index-admitted corpus poisoning in frozen, benchmarkscale RAG. Let g be the benchmark gold and t a type-compatible target fixed before construction and alias-distinct from g. Benchmark-intended query semantics are fixed before insertion; frozen filters admit only candidates classified as incorrect under those semantics, and a blinded human audit
S⊊D
(3)
Reported values replace each probability with its observed frequency. CJL measures a full-set dependence gap, while token- and slot-matched controls test alternative context-size explanations. Retrieval and validity. End-to-end evaluation reports JRR@K, the fraction of packets retrieved in full, and Ret. ASR@K, the target-output rate after retrieval from C ′ . Document validity is evaluated separately from attack success and subset safety. A verifier filters candidates judged fabricated, speculative, answer-shaped, or conflicting and favors documents that remain plausible and consistent with the surrounding evidence. 2
(a) Setup & Threat
(b) Multi-Agent Construction: Lens Search and Document Synthesis-Repair Outer Loop: Lens Search & Planning
Question
Retrieval Top-K Results LLM Planner Agents
What nationality is athlete A
...
Athlete [A] officially competes for [Country X]
...
...
...
LLMs
Search Lens Lenses Candidate Queue
!!
!" ... !#
Selected Policy *
(ℓ, %, &"#$ , '%&'() )
...
Drafter
Humanizer
Reviser
How to read 'nationality'
same question, different criterion -> different answer
Counter
... Sample
LLM Verifier Agents
...
...
Success Subset Checker Checker
...
&! : nationality = country of birth
JRR@K: full set co-retrieved Subset-Safety
Supported Doc [A] was born in Country [Y]. [A] represented Country [X] in international competition.
target ) = Country Y Birth nationality Y (preselected target)
...
Lens-setting Doc
Question What nationality is athlete A
"#$%: nationality = sporting/offical
...
Nationality may denote birth country, not the country represented in competition.
Select !! Sporting Nationality X (benchmark gold)
"! "" ... "#
LLM Proposer Agents
Evidence Len Len Planner Analyzer Enumerator Scorer Composer
Black-box RAG clean passage
(c) Set-Level Compositional Failure
...
Alignment Validity Checker Checker
TR(%) ≥ ( )*+ TR(,) ≤ . '⊂)
Reflection
{E+ } {E+ ,E, } {E+ ,...,E, } |⋅|< 0
no proper subset induces = TR(,) ≤ > for all , ⊆ %
Full Set Success
Align(ℓ,6,%)= 1
{E!,...,E1 }
Valid(<, %) = 1
only the full set induces = TR(%) ≥ (
Inner Loop: Synthesis & Repair
Figure 2: Overview of LENS. The Planner defines the document plan, the Proposer writes lens-setting and fact-completion documents, and the Verifier returns typed counterexamples from full-set, subset, alignment, and validity checks. Construction uses a local surrogate, while attacked target-side runs occur only after packet freezing. Exhausted plans return to the Planner.
Method: LENS
non-default reading criterion available, while fact-completion documents provide the missing support needed under that criterion. At the role level, the Planner seeks D = Dlens ∪ Dfact such that
Overview: Constraint-Maintaining Construction LENS treats constraint violations as counterexamples for local document repair or, when a plan is exhausted, for outer-loop replanning. Figure 2 summarizes three functional roles. The Planner selects an interpretation lens and assigns complementary document roles, required facts, and claims to avoid. The Proposer realizes this plan as lens-setting and fact-completion documents. The Verifier evaluates full-set success, subset leakage, interface alignment, and document validity using the local surrogate, then returns a typed repair signal.
TR(sur) (Dlens ) ≤ ϵ, TR(sur) (Dfact ) ≤ ϵ, TR(sur) (D) ≥ τ. (5) This is a population-level planning objective. Construction instead uses 20-generation empirical rates over an active set A of proper subsets, initialized with all singletons and the subsets specified by Rrisk , then expanded when new leakage is found.
Interpretation-Lens Planning Document Synthesis and Diagnostic Verification
An interpretation lens is a typed, query-conditioned reading criterion, such as role, time, category, naming, or scope, rather than the inserted conclusion. For example, a role lens may map a creator query to a local creator-entry rule; separate documents supply the role and date facts needed under it. Using the clean evidence path, answer dimension, and locally supportable target-relevant facts, the Planner ranks lens and role decompositions by semantic fit, closure, validity, and leakage risk. This structured search rejects plans that require one document to state a decisive target relation. The selected plan is π = (ℓ, ρ, Freq , Cavoid , Rrisk ), (4)
The Proposer writes lens-setting documents that make the selected criterion available without instantiating the target relation. Fact-completion documents supply the remaining support without restating the lens, comparing candidate answers, or stating the target. Both roles must resemble ordinary reference prose and avoid answer-shaped scaffolding. The Verifier distinguishes two common failures. Subset leakage occurs when a proper subset makes the target locally decisive. Interface mismatch occurs when the lens and complementary facts operate over different answer dimensions. Accordingly, Align(ℓ, π, D) checks whether the query, clean support, lens, and target-relevant facts share the same selection criterion. Diagnostic contexts include clean-only, lens-only, fact-only, clean-plus-single-role, and full-set inputs, allowing the Verifier to separate genuine complementarity from a locally sufficient insert or target behavior already present in the clean reader. For a candidate packet D, plan π,
where ℓ is the lens, ρ assigns document roles, Freq lists closure facts, Cavoid lists leakage-prone claims, and Rrisk lists proper subsets of planned document slots predicted to have high leakage risk. Inner-loop repair preserves ℓ and ρ; changing either starts a new outer-loop plan. The key planning decision is to separate interpretation from factual completion. Lens-setting documents make a 3
Diagnostic view
Test / comparison
Full-set closure
d(sur) TR 20 (D) ≥ τ d(sur) maxS∈A TR 20 (S) ≤ ϵ
Tracked subset safety Interface alignment Document validity Tracked-set update
Failure signal
Repair objective
Full-set rate is too low
Supply the missing relation
Align(ℓ, π, D) Valid(X , D)
A tracked subset leaks Lens and facts use different criteria An insert is unsupported or conflicting
Remove the decisive cue Align the document roles Revoice, qualify, or remove
d(sur) Find Sv ⊊ D with TR 20 (Sv ) > ϵ
A new subset leaks
Add Sv to A
Table 1: Verifier diagnostic matrix. Each failed check returns a typed counterexample and a repair target. Newly discovered leaking subsets are added to the tracked constraints.
Experiments
lens ℓ, and active subset set A, the current diagnostic gate is (sur) d20 (D) ≥ τ, TR (sur) d Pass(D; ℓ, π, A) ⇐⇒ maxS∈A TR20 (S) ≤ ϵ, Align(ℓ, π, D) = 1, Valid(X , D) = 1. (6) Table 1 summarizes the resulting counterexamples and repair targets.
Experimental Setup Benchmarks, models, and pipeline. We evaluate HotpotQA, 2WikiMultihopQA, MuSiQue, and Natural Questions (Yang et al. 2018; Ho et al. 2020; Trivedi et al. 2022; Kwiatkowski et al. 2019) on a fixed manifest of 1,000 query–target attempts per benchmark. Each retained query is answered correctly by all three clean target RAGs; targets are type-compatible and fixed before construction, and failures are never replaced. LENS uses Qwen3.6-27B for construction and surrogate reading, while Qwen3.6-35B-A3B, Llama-3.1-70B-Instruct, and DeepSeek-V4-Pro serve as target generators (Qwen Team 2026; Grattafiori et al. 2024; DeepSeek-AI 2026). LENS and the LLM-based baselines follow the assigned k ∈ {2, 3} budget; Semantic Chameleon always uses its native twodocument sleeper–trigger pair. A LENS attempt succeeds only if its packet passes the exhaustive surrogate gate at τ = 0.50 and ϵ = 0.10. Frozen packets are evaluated on all three target generators without target-side selection. Conditional evaluation supplies X ∪ S; end-to-end evaluation indexes returned packets using BM25/BGE-M3 reciprocal-rank fusion and a Qwen3 reranker (Robertson and Zaragoza 2009; Chen et al. 2024; Cormack, Clarke, and Buettcher 2009; Zhang et al. 2025b). Metrics and evaluation. Each context uses 20 stochastic generations. Packet metrics use all returned packet–generator pairs, with JRR computed once per packet; all-attempt metrics retain the fixed manifest and score construction failures as zero. Full ASR is the conditional full-packet target rate, Max-sub. TR is the largest proper-subset rate, and CJL is their difference. Subset Pass@20 accepts a pair only when every proper subset produces the target at most twice in 20 generations. Ret. ASR measures post-retrieval target adoption. E2E-Strict@5 further requires complete packet retrieval, target rate at least 0.50 in the actual top-5 context, and target rate at most 0.10 for every controlled proper subset. Cond-Strict@5 applies the full-set threshold in the conditional interface while retaining the same controlled subset test. Only unhedged matches to t or g count. Clean Acc. uses a separate held-out utility set. Mechanism and mitigation analyses condition on returned artifacts; defense comparisons retain all attempts. Results are benchmark-macro averages with benchmark-stratified confidence intervals; method comparisons use paired resampling. Comparisons and checks. Construction baselines include direct multi-document synthesis, best-of-N synthesis, generic self-refinement, split evidence/payload synthesis, and
Counterexample-Guided Repair For a fixed lens, plan, and active subset set, the inner loop uses the tracked surrogate dependence gap (sur)
(sur)
(sur)
d20 (D) − max TR d20 (S). [ A,20 (D) = TR CJL S∈A
(7)
This score prioritizes repairs, while Eq. 6 remains the acceptance criterion. When A contains every proper subset, the score equals the 20-generation empirical surrogate-side dependence gap. During construction, it covers only the subsets currently tracked by the Verifier. Each counterexample identifies both the violated condition and the responsible document role. The Proposer then rewrites only the corresponding document while preserving the clean context, lens, and role assignment. Low full-set closure triggers the addition or clarification of missing support rather than direct insertion of the conclusion. Subset leakage triggers removal or qualification of the decisive cue. Alignment failures revise how a document realizes its assigned role, and validity failures revoice or remove unsupported content. This localized repair preserves the connection between each failure signal and the change intended to correct it. Whenever the Verifier discovers a violating Sv ⊊ D, it adds Sv to A for subsequent repairs. These tracked checks guide construction, but they do not define the final reported subset result. Before a packet is frozen, the Verifier exhaustively evaluates every proper subset under the same 20-generation surrogate protocol and reapplies the full-set, alignment, and validity checks. Target-side evaluation then independently repeats the exhaustive subset analysis for each target generator Gm . If repair repeatedly alternates between subset leakage and insufficient full-set closure, LENS exhausts the current plan and returns to the Planner rather than assuming that local improvements will compose. 4
Metric
Hotpot 2Wiki MuSiQue NQ
Avg.
All attempts Yield ↑ 0.655 Ret. ASR@5 ↑ 0.517 E2E-Strict@5 ↑ 0.401
0.640 0.490 0.364
0.705 0.611 0.477
0.500 0.625 0.357 0.494 0.209 0.363
Returned packets Full ASR ↑ 0.862 Max-sub. TR ↓ 0.036 JRR@5 ↑ 0.939 Ret. ASR@5 ↑ 0.790 E2E-Strict@5 ↑ 0.612
0.841 0.029 0.913 0.765 0.568
0.910 0.069 0.958 0.866 0.677
0.795 0.852 0.141 0.069 0.898 0.927 0.713 0.784 0.419 0.569
PoisonedRAG GASLITE BadRAG LENS
.507 .454 .473 .494
.219 .167 .194 .401
.286 .274 .259 .448
.157 .118 .141 .260
.274 .255 .291 .408
Human and joint validity outcome
Result
Document and target validity Locally supported inserts (n = 600) Plausible reference prose (n = 600) Incorrect frozen targets (n = 240)
552/600 (92.0%) 535/600 (89.2%) 211/240 (87.9%)
Joint semantic and behavioral validity SemValid returned packets (n = 240) 164/240 (68.3%) SemValid ∧ E2E-Strict pairs (n = 720) 331/720 (46.0%) E2E-Strict | SemValid 331/492 (67.3%) E2E-Strict | non-SemValid 78/228 (34.2%)
Value
Returned target-specific packets Intended Ret. ASR@5 0.735 Other-target Ret. ASR@5 0.042 Selectivity gap 0.693
Table 5: Blinded human validation. Document judgments use 600 inserts from 240 returned packets; target correctness uses a separate pre-construction sample of 240 frozen targets. SemValid is evaluated on the returned-packet audit pool and requires an incorrect target, a definite criterion shift, and no entailment under the original query semantics. Joint behavioral rates attach the three target-generator E2E-Strict outcomes to each of the 240 audited returned packets.
All 720 attempts Yield 0.588 Intended Ret. ASR@5 0.434 Other-target Ret. ASR@5 0.024 All 240 queries ≥ 2 selective targets All 3 selective targets
None SeCon Reliab. Robust Trust
Table 4: All-attempt defended ASR@5 on the shared 4,000attempt manifest. Construction failures and invalid outputs are retained and scored as zero. Each cell aggregates three target generators and 20 paired decoding seeds per attempt.
Table 2: Main results. The upper panel uses all 4,000 attempts and scores construction failures as zero. The lower uses all returned packets without target-side filtering. E2E-Strict@5 requires complete packet retrieval, target adoption in the actual top-5 context, and the controlled all-proper-subset check. Metric
Attack
0.538 0.158
Full ASR is 0.852, Max-sub. TR is 0.069, and 79.7% pass every proper-subset check. Moreover, 89.7% of Cond-Strict successes remain strict in the actual retrieved context, showing that the constructed dependence transfers from controlled evaluation to the deployed retrieval path. Human validation. Annotators are independent of the construction validators and blind to their acceptance decisions and model-side outcomes. Table 5 shows that inserted documents are predominantly locally supported and plausible, while 87.9% of targets sampled before construction are incorrect under the intended query semantics. More importantly, 68.3% of returned packets satisfy SemValid, and 46.0% of their generator evaluations are both SemValid and E2E-Strict. Strict success is nearly twice as frequent within SemValid packets as outside them, linking the behavioral effect to attacker-induced answer criteria. Construction baselines. Table 6 shows that LENS raises all-attempt E2E-Strict@5 from 0.244 for the strongest baseline to 0.363, a 48.8% relative gain. This advantage targets the capability studied here: isolating attack success to the complete document set. LENS lowers Max-sub. TR from 0.153 to 0.069, and the same conclusion holds under exact two-document matching, where it reaches 0.397 versus 0.247 for native Semantic Chameleon. Statistical robustness. The main conclusion is stable across
Table 3: Target controllability with three preselected targets per query. Packet rows condition on returned packets, attempt rows score failures as zero, and selective reachability requires a gap of at least 0.30.
protocol-faithful Semantic Chameleon (Thornton 2026). All share the fixed manifest, corpus, retrieval pipeline, targetagnostic admission validator, and target-side evaluation; candidate selection uses frozen construction-side signals only. LLM-based baselines additionally share the Qwen construction model and generation budget, while Semantic Chameleon uses its published native optimizer. We also compare PoisonedRAG, GASLITE, and BadRAG under RobustRAG, ReliabilityRAG, SeCon-RAG, and TrustRAG (Zou et al. 2025; Ben-Tov and Sharif 2025; Xue et al. 2024; Xiang et al. 2026; Shen et al. 2025; Si et al. 2025; Zhou et al. 2025). Additional checks cover three-target controllability, Llama-family construction transfer, and independent 100-draw confirmation.
Main Results Attack effectiveness. Table 2 shows that LENS returns packets for 62.5% of fixed attempts, yielding 0.494 all-attempt Ret. ASR@5 and 0.363 E2E-Strict@5 (95% CI: [0.348, 0.378]). Returned packets exhibit strong set dependence: conditional 5
Figure 4: Component ablations on the fixed 1,000-attempt ablation manifest. Natural runs use each variant’s executed budget; fixed-budget runs apply the same call, generation, and token caps. Points report the primary all-attempt E2EStrict@5 metric.
tains the highest all-attempt ASR@5 under all four defenses, exceeding the strongest baseline by 0.103–0.182 (0.141 on average); paired intervals exclude zero. Its advantage combines construction coverage with set-level effects that survive defenses built around document-local cues.
Mechanism and Ablation Component ablations. Figure 4 evaluates every component on an independent 1,000-attempt manifest. Full LENS reaches 0.403 all-attempt E2E-Strict@5. Under a shared cap of 15 proposal-or-repair calls, 800 diagnostic generations, and 1.05M tokens, it retains 0.327, twice the strongest component removal at 0.163. Removing subset updates preserves high full-set success but reduces fixed-budget E2E-Strict@5 to 0.134, identifying subset-aware repair as the source of set isolation. Mechanism controls. Figure 3 shows that matched packets achieve 0.785 Ret. ASR@K, compared with 0.265 after role scrambling and 0.127 after cross-lens substitution. JRR remains nearly unchanged across these conditions (0.863– 0.867), localizing the attack effect to the planned lens–fact correspondence. The larger residual under role scrambling is consistent with retaining all semantic atoms, whereas crosslens substitution replaces the criterion that makes the recipient facts jointly decisive.
Figure 3: Mechanism controls on 240 matched instances. Bars show Ret. ASR@K and lines show JRR@K. Each condition uses 720 realization–generator pairs, with JRR computed once per realization. We use K = 5 for k ∈ {2, 3} and K = 8 for the held-out k = 4 extension. construction thresholds: increasing τ trades coverage for packet strength, while the nested ϵ sweep changes subset acceptance without altering the full-set dependence pattern. Independent 100-draw evaluation confirms all-subset control for 68.6% of sampled pairs. All-attempt E2E-Strict remains stable across target generators (0.357–0.370) and transfers to Llama-guided construction. Target specificity and spillover. Table 3 reports a 0.693 intended-versus-alternate target gap, with at least two targets selectively reachable for 53.8% of queries. Across 1,200 unseen off-target query–packet pairs, full-packet retrieval is 1.7%, original-target spillover is 0.8% versus 0.3% in paired clean runs, and gold accuracy changes by 1.5 points. Even when same-entity queries retrieve individual documents more often (29.0%), the composed packet rarely transfers its behavioral effect. Robustness under defenses. Table 4 shows that LENS at-
Composition-Aware Mitigation Probes Local passage screening leaves LENS largely intact, whereas defenses that evaluate the retrieved set expose its compositional dependence. A set-level judge reduces Ret. ASR from 0.784 to 0.318, and combining judging with subset ablation reaches 0.164. Development-frozen selective auditing provides the strongest operating point: it reduces Ret. ASR to 0.284 while retaining 91.3% of baseline clean accuracy. These results establish a practical defense principle: allocate setlevel inspection to answers whose support is both distributed 6
Method
Attempt MaxPacket Attempt Yield Full sub. CJL Ret. Ret. E2E-Strict@5
Direct multi-doc Best-of-N direct Generic self-refine Split evidence/payload Semantic Chameleon LENS
.338 .447 .472 .503 .546 .625
.754 .823 .812 .842 .872 .852
.241 .198 .146 .172 .153 .069
.513 .625 .666 .670 .719 .783
.671 .741 .735 .766 .798 .784
.227 .331 .347 .385 .436 .494
.088 .154 .190 .216 .244 .363
Table 6: Construction baselines on the fixed 4,000-attempt manifest. All methods share query–target pairs, corpus, retrieval pipeline, target-agnostic admission, and target-side evaluation. LENS and the LLM-based baselines follow the assigned k = 2/3 budget; Semantic Chameleon retains its native two-document sleeper–trigger GCG protocol. Packet metrics condition on returned packets; attempt metrics score construction failures as zero.
proper subset remains insufficient. This all-proper-subset objective distinguishes set-level composition from coordinated retrieval, prompt injection, structured inference-chain poisoning, and attacker competition.
and unstable under document removal. Operational implications. The experiments identify joint retrieval as both the operational bottleneck and the defense lever. Incomplete packets rarely activate the target (0.028), while targeted hard negatives reduce Ret. ASR@5 to 0.519, showing that retrieval competition directly limits compositional steering. Attack construction and defense therefore meet at the same control point: which evidence relations survive retrieval and become jointly available to the generator. Extending this analysis across larger retrieval ecosystems and competitive corpus settings is a natural next step. Security implications. Together, these results redefine three units of RAG security analysis. Attack models should treat retrieved evidence sets as units of control; evaluation should test whether target adoption depends on the complete packet; and defenses should examine how passages jointly determine an answer. This makes evidence-set dependence measurable and frames defenses around distinguishing attacker-induced rules from legitimate multi-hop support.
Detection and defense. RAG defenses and transferable input-screening methods inspect or isolate passages and use filtering, answer aggregation, clean-evidence recovery, consistency graphs, clustering, or self-assessment (Qi et al. 2021; Robey et al. 2023; Xiang et al. 2026; Tan et al. 2025; Yao et al. 2025; Edemacu et al. 2025; Cheng et al. 2025; Kim, Lee, and Koo 2025; Chang et al. 2025a; Shen et al. 2025; Si et al. 2025; Zhou et al. 2025). These signals suit locally abnormal or conflicting poisons. LENS reduces both cues through plausible, locally insufficient documents. It therefore motivates defenses that evaluate how passages jointly support an answer, extending security inspection from document content to evidence composition.
Conclusion
Related Work
We introduced LENS, a set-level poisoning framework that shifts frozen RAG’s attack surface from individual documents to their composition. LENS combines interpretation-lens planning, role-separated synthesis, counterexample-guided repair, and empirical proper-subset verification to construct packets whose full set redirects answers while proper subsets maintain low target rates. This exposes a security blind spot: plausible, locally inconclusive documents can jointly induce an attacker-chosen decision rule. Mechanism controls and blinded audits localize this effect to the planned lens–fact interface and establish it as a distinct form of evidence-set control. Across target generators, LENS outperforms strong construction baselines on strict set-level success and survives published defenses in the retrieved context. LENS also provides a reusable stress test for evidence-set dependence, allowing RAG models, retrieval pipelines, and defenses to be evaluated against answer control that emerges only through cross-document composition. RAG security must therefore treat evidence composition as a first-class attack surface while preserving the cross-document reasoning that gives retrieval augmentation its value. Generative AI use disclosure. Generative AI tools were used for language editing, LaTeX assistance, and figure and code
RAG corpus poisoning. RAG corpus poisoning uses answer-bearing passages, triggers, gradient optimization, or black-box document construction to steer retrieved generation (Lewis et al. 2020; Ram et al. 2023; Izacard and Grave 2021; Asai et al. 2024). Representative methods span these threat models (Zou et al. 2025; Ben-Tov and Sharif 2025; Chaudhari et al. 2026; Wang et al. 2026b; Xian et al. 2025; Chang et al. 2025b; Zhang et al. 2025a; Nazary, Deldjoo, and di Noia 2025; Chen et al. 2025b). SilentRetrieval improves the fluency and retrieval transfer of such poisons (Qian 2026). Prior methods optimize attack success, retrievability, or stealth without constraining every proper subset. LENS makes this constraint its construction objective and places the target effect in evidence composition. Coordinated and composite attacks. Prior work studies sleeper–trigger pairs, agentic trajectories, prompt-injectionplus-database poisoning, adversarial KG inference chains, and competing attackers (Thornton 2026; Choi et al. 2026; Pan et al. 2026; Wang et al. 2026a; Zhao et al. 2025; Chen et al. 2025a; Huang et al. 2024). LENS isolates a complementary regime: frozen single-round text RAG, where one naturallanguage packet realizes an answer-level effect while every 7
drafting. The authors verified all scientific claims, experimental results, analyses, citations, and final text.
Cormack, G. V.; Clarke, C. L. A.; and Buettcher, S. 2009. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, 758–759. Association for Computing Machinery. DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. Edemacu, K.; Shashidhar, V. M.; Tuape, M.; Abudu, D.; Jang, B.; and Kim, J. W. 2025. Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation. arXiv preprint arXiv:2508.02835. Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783. Ho, X.; Nguyen, A.-K. D.; Sugawara, S.; and Aizawa, A. 2020. Constructing a Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps. In Proceedings of the 28th International Conference on Computational Linguistics, 6609– 6625. Barcelona, Spain (Online): International Committee on Computational Linguistics. Huang, H.; Zhao, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024. Composite Backdoor Attacks Against Large Language Models. In Findings of the Association for Computational Linguistics: NAACL 2024, 1459–1472. Mexico City, Mexico: Association for Computational Linguistics. Izacard, G.; and Grave, E. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 874–880. Online: Association for Computational Linguistics. Kim, M.; Lee, H.; and Koo, H. 2025. Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems. arXiv preprint arXiv:2511.01268. Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; Toutanova, K.; Jones, L.; Kelcey, M.; Chang, M.-W.; Dai, A. M.; Uszkoreit, J.; Le, Q.; and Petrov, S. 2019. Natural Questions: A Benchmark for Question Answering Research. Transactions of the Association for Computational Linguistics, 7: 453–466. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Kuttler, H.; Lewis, M.; Yih, W.-t.; Rocktaschel, T.; Riedel, S.; and Kiela, D. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, volume 33, 9459– 9474. Nazary, F.; Deldjoo, Y.; and di Noia, T. 2025. PoisonRAG: Adversarial Data Poisoning Attacks on RetrievalAugmented Generation in Recommender Systems. arXiv preprint arXiv:2501.11759. Pan, Y.; Zhang, Z.; Lei, J.; Jia, C.; Si, Q.; and Guo, H. 2026. FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents. arXiv preprint arXiv:2607.04718.
References Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; and Hajishirzi, H. 2024. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In International Conference on Learning Representations. Ben-Tov, M.; and Sharif, M. 2025. GASLITEing the Retrieval: Exploring Vulnerabilities in Dense Embedding-Based Search. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, 4364–4378. Association for Computing Machinery. Chang, C.-Y.; Jiang, Z.; Rakesh, V.; Pan, M.; Yeh, C.-C. M.; Wang, G.; Hu, M.; Xu, Z.; Zheng, Y.; Das, M.; and Zou, N. 2025a. MAIN-RAG: Multi-Agent Filtering RetrievalAugmented Generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2607–2622. Vienna, Austria: Association for Computational Linguistics. Chang, Z.; Li, M.; Jia, X.; Wang, J.; Huang, Y.; Jiang, Z.; Liu, Y.; and Wang, Q. 2025b. One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems. In Findings of the Association for Computational Linguistics: EMNLP 2025, 18811–18825. Suzhou, China: Association for Computational Linguistics. Chaudhari, H.; Severi, G.; Abascal, J.; Suri, A.; Jagielski, M.; Choquette-Choo, C. A.; Nasr, M.; Nita-Rotaru, C.; and Oprea, A. 2026. Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation. ACM Transactions on AI Security and Privacy. Chen, J.; Xiao, S.; Zhang, P.; Luo, K.; Lian, D.; and Liu, Z. 2024. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. In Findings of the Association for Computational Linguistics: ACL 2024, 2318–2335. Bangkok, Thailand: Association for Computational Linguistics. Chen, L.; Yang, X.; Lu, Y.; Zhang, J.; Sun, X.; Liu, Q.; Wu, S.; Dong, J.; and Wang, L. 2025a. PoisonArena: Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation. arXiv preprint arXiv:2505.12574. Chen, Z.; Gong, Y.; Liu, J.; Chen, M.; Liu, H.; Cheng, Q.; Zhang, F.; Lu, W.; and Liu, X. 2025b. FlippedRAG: Black-Box Opinion Manipulation Adversarial Attacks to Retrieval-Augmented Generation Models. arXiv preprint arXiv:2501.02968. Cheng, Z.; Sun, J.; Gao, A.; Quan, Y.; Liu, Z.; Hu, X.; and Fang, M. 2025. Secure Retrieval-Augmented Generation Against Poisoning Attacks. In 2025 IEEE International Conference on Big Data, 1799–1806. ArXiv:2510.25025. Choi, C.; Kim, E.; Lee, K.; Chun, Y.; Jeong, J.; Kim, E.; Oh, M.; Jang, J.; and Chang, B. 2026. KidnapRAG: A Black-Box Attack for Hijacking Reasoning in Agentic Retrieval-Augmented Generation Systems. arXiv preprint arXiv:2607.00422. 8
of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 68292–68315. PMLR. Xiang, C.; Wu, T.; Zhong, Z.; Wagner, D.; Chen, D.; and Mittal, P. 2026. Certifiably Robust RAG against Retrieval Corruption. In Conference on Secure and Trustworthy Machine Learning (SaTML). Xue, J.; Zheng, M.; Hu, Y.; Liu, F.; Chen, X.; and Lou, Q. 2024. BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models. arXiv preprint arXiv:2406.00083. Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2369–2380. Brussels, Belgium: Association for Computational Linguistics. Yao, R.; Zhang, Y.; Song, S.; Gao, N.; and Tu, C. 2025. EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, 4034–4050. Suzhou, China: Association for Computational Linguistics. Zhang, C.; Zhang, X.; Lou, J.; Wu, K.; Wang, Z.; and Chen, X. 2025a. PoisonedEye: Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large VisionLanguage Models. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 76811–76830. PMLR. Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; Huang, F.; and Zhou, J. 2025b. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv preprint arXiv:2506.05176. Zhao, T.; Chen, J.; Ru, Y.; Zhu, H.; Hu, N.; Liu, J.; and Lin, Q. 2025. RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation. arXiv preprint arXiv:2507.08862. Zhou, H.; Lee, K.-H.; Zhan, Z.; Chen, Y.; Li, Z.; Wang, Z.; Haddadi, H.; and Yilmaz, E. 2025. TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation. arXiv preprint arXiv:2501.00879. Zou, W.; Geng, R.; Wang, B.; and Jia, J. 2025. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. In 34th USENIX Security Symposium (USENIX Security 25), 3827–3844. Seattle, WA: USENIX Association.
Qi, F.; Chen, Y.; Li, M.; Yao, Y.; Liu, Z.; and Sun, M. 2021. ONION: A Simple and Effective Defense Against Textual Backdoor Attacks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 9558–9566. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics. Qian, J. 2026. SilentRetrieval: Hijacking RetrievalAugmented Generation via Semantically-Preserving Adversarial Data Poisoning. arXiv preprint arXiv:2605.28074. Qwen Team. 2026. Qwen3.6 Model Collection. Hugging Face model collection. Ram, O.; Levine, Y.; Dalmedigos, I.; Muhlgay, D.; Shashua, A.; Leyton-Brown, K.; and Shoham, Y. 2023. In-Context Retrieval-Augmented Language Models. Transactions of the Association for Computational Linguistics, 11: 1316–1331. Robertson, S.; and Zaragoza, H. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4): 333–389. Robey, A.; Wong, E.; Hassani, H.; and Pappas, G. J. 2023. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. arXiv preprint arXiv:2310.03684. Shen, Z.; Imana, B.; Wu, T.; Xiang, C.; Mittal, P.; and Korolova, A. 2025. ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search. Advances in Neural Information Processing Systems, 38: 45662–45702. Si, X.; Zhu, M.; Qin, S.; Yu, L.; Zhang, L.; Liu, S.; Li, X.; Duan, R.; Liu, Y.; and Jia, X. 2025. SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG. Advances in Neural Information Processing Systems, 38: 70652–70681. Tan, X.; Luan, H.; Luo, M.; Sun, X.; Chen, P.; and Dai, J. 2025. RevPRAG: Revealing Poisoning Attacks in RetrievalAugmented Generation through LLM Activation Analysis. In Findings of the Association for Computational Linguistics: EMNLP 2025, 12999–13011. Suzhou, China: Association for Computational Linguistics. Thornton, S. 2026. Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems. arXiv preprint arXiv:2603.18034. Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2022. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics, 10: 539–554. Wang, H.; Liu, H.; Zhu, J.; Wang, Z.; Guo, Y.; and Tang, X. 2026a. PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems. arXiv preprint arXiv:2603.25164. Wang, H.; Zhang, R.; Wang, J.; Li, M.; Huang, Y.; Wang, D.; and Wang, Q. 2026b. Joint-GCG: Unified GradientBased Poisoning Attacks on Retrieval-Augmented Generation Systems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 35793–35801. Xian, X.; Wang, G.; Bi, X.; Zhang, R.; Srinivasa, J.; Kundu, A.; Fleming, C.; Hong, M.; and Ding, J. 2025. On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains. In Proceedings 9