Preprint. Under review.
S TILL T HERE , N O L ONGER S EEN : E XPOSING C OMPRESSION -I NDUCED R ISK IN L ARGE V ISION L ANGUAGE M ODELS
arXiv:2609.35002v1 [cs.CV] 28 Sep 2026
Qiankun Li1∗ Yuechen Zhang2∗ Bowen Chen2 Shilinlu Yan2 Zhenhong Zhou1 Kun Wang1† Li Sun2† 1 Nanyang Technological University 2 Beijing University of Posts and Telecommunications {cs-qiankun.li,wang.kun}@ntu.edu.sg [email protected] [email protected] {cbcbw,lulu_land,lsun}@bupt.edu.cn
A BSTRACT Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial input that remains correct under full-token inference but fails after compression, casting compression-induced risk as a paired failure attribution problem. Within a controlled diagnostic cohort, counterfactuals show that retained-set allocation causally changes compressed correctness and reveal a negative association between recovery and representation drift in displaced evidence. Motivated by these findings, we propose CIRA, a Compression-Induced Risk Attack for Large Vision-Language Models. Under a vision-encoder white-box setting, CIRA optimizes image perturbations through encoder-side objectives that manipulate token priorities across candidate compression budgets while preserving displaced evidence. CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across 12 dataset–compressor settings evaluated at four budgets, CIRA achieves a mean CSFR of 20.35% while limiting full-token attack success to 6.92%, with similar behavior on additional LVLM families. A cross-view selection-stabilization defense substantially suppresses CIRA, although Adaptive CIRA partially restores its effectiveness. These results show that compression-specific failures persist under restricted access and support paired evaluation of full-token and compressed inference for attributing risk to visual-token compression. Code is provided in the Github.
1
I NTRODUCTION
Large Vision-Language Models (LVLMs) support broad visual understanding by mapping images to long visual-token sequences (Alayrac et al., 2022; Li et al., 2023a; Liu et al., 2024b). Visual-token compression reduces this inference cost through token selection, aggregation, and compact visual representations (Shang et al., 2025; Yang et al., 2025b; Li et al., 2025; Bulat et al., 2026). Decoderside methods further reduce visual-token computation within the language model (Chen et al., 2024; Zhang et al., 2024; Xing et al., 2024). Meanwhile, adversarial attacks show that failures can be induced in the underlying LVLM through image-space or joint-modal perturbations (Zhang et al., 2022; Schlarmann & Hein, 2023; Yin et al., 2023). This raises the question: what risks are induced by visual-token compression itself, beyond vulnerabilities already present in the underlying LVLM? When a full-token LVLM is used for robustness assessment while a compressed variant is deployed, failures confined to the compressed path are not captured by that assessment. We therefore evaluate each adversarial input with and without token compression. A compression-specific failure (CSF) ∗ †
Equal contribution. Kun Wang and Li Sun are the corresponding authors.
1
Preprint. Under review.
Empirical Evidence
CIRA Framework
Compression-Specific Failure
Attack Optimization Vision Encoder
Is there a pillow in the image? Clean Image
Answer: No
Adv Image
Output
Objective I: Global Selection Hijacking Clean Priority Scores
Inverse Clean Priority Ranking
···
···
Pillow evidence inaccessible
Adversarial Priority Alignment Before → After Optimization
···
𝒓𝒄 = 𝒓𝒂𝒏𝒌↓ 𝒔𝒄
𝓡 𝒔𝒄 𝒊 =
𝒓𝒊𝒄 − 𝟏 𝑵−𝟏
···
𝒔𝒄 → 𝒔𝒂
𝓛𝑮𝑺𝑯
Objective II: Hidden-Evidence Preservation Budget-Marginal Displacement Weighting
2. Unknown Compression Budget
K=? 𝑲𝒎𝒊𝒏
𝑲𝒎𝒂𝒙
concealed budgets
Unknown retention boundary.
3. No Downstream Access Image
Vision Encoder
Compression
Clean rank
Adv rank
Representation Preservation
LLM
Access: Prompt-free and downstream-agnostic.
𝑲𝒎𝒊𝒏
Adv tokens
Limited Drift Clean tokens
𝑲𝒎𝒂𝒙
𝒘𝒊 = 𝑷𝒓𝑲∼𝑼𝒏𝒊𝒇 𝓚 𝒓𝒄𝒊 ≤ 𝑲 < 𝒓𝒂𝒊
Representation Space:
Weighted Preservation
Do not corrupt the evidence.
𝓛𝑯𝑬𝑷
𝒅𝒊 =
Full-Token Preservation; CSF Exposure.
Case Study:
···
···
Key Challenge 1. Preserve Displaced Evidence
Compressor LLM ···
Unknown Token Features (Adv)
𝟏 − 𝒄𝒐𝒔 𝒉𝒂𝒊 , 𝒉𝒄𝒊 , 𝒉𝒂𝒊 ≈ 𝒉𝒄𝒊 𝟐
: Is there a pillow in the image? LLaVA-1.5-7B | PruMerge | K=64 Priority Rank: 9→112 | Cos. Sim.: 0.91
Priority Rank: 2→72 | Cos. Sim.: 0.97 Priority Rank: 8→97 | Cos. Sim.: 0.93
···
Priority Rank: 6→117 | Cos. Sim.: 0.98
Critical Pillow Tokens
Pillow evidence accessible Compressed inference
Unknown Task Prompt
VE
Standardized Alignment
Answer: Yes
Quantitative Results:
Objective
Vision Encoder
Priority Reversal
Our Goal Full-token inference
Token Features (Clean) CIRA
Victim Inference
Full-Token Inference
: Yes
Compressed Inference
: No
Figure 1: Overview of the compression-specific failure setting, the CIRA attack framework and the resulting behavior under visual-token compression.
occurs when full-token inference remains correct but compressed inference fails. This paired criterion distinguishes failures introduced by compression from those already present in the underlying LVLM. Recent studies increasingly examine robustness under visual-token compression. Evidence shows that compression can reshape robustness in either direction, depending on how visual evidence is selected and retained (Wang et al., 2026a; Gu et al., 2026). Compression-aware attacks further demonstrate that adversarial behavior can depend on the compression path itself (Zhang et al., 2026c;a). These developments shift the question from whether compression affects robustness to which adversarial failures can be attributed specifically to compression and whether they can be selectively induced under deployment uncertainty. Specifically, can a single perturbation optimized through the target vision encoder preserve full-token correctness while inducing compressed-path failures when the deployed compressor and budget are unknown during optimization? To understand what governs this selectivity, we use controlled counterfactuals to isolate two factors. Retained-set interventions causally change compressed correctness by altering which visual evidence remains accessible after compression. Meanwhile, recovery is negatively associated with representation drift in displaced evidence. These results directly inform the attack design: reallocate token priority while preserving displaced evidence. In this direction, we introduce CIRA (Compression-Induced Risk Attack), an attack framework for inducing compression-specific failures. As illustrated in Figure 1, CIRA combines Global Selection Hijacking (GSH), which globally reallocates encoder-side proxy priorities, with HiddenEvidence Preservation (HEP), which limits representation drift in displaced clean high-priority tokens. With white-box access restricted to the vision encoder, CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across four compressors, three datasets, four budgets, and multiple LVLMs, CIRA induces compression-specific failures while keeping full-token degradation limited. We also introduce Translation-Consensus Selection (TCS), a cross-view selection-stabilization defense, and evaluate it against both standard and adaptive CIRA. Our contributions are threefold: ❶ Compression Risk Attribution. We formulate compression-induced risk through a cleanconditioned paired CSF criterion and show that retained-set allocation causally affects compressed correctness while recovery decreases with representation drift. 2
Preprint. Under review.
❷ Compression-Induced Risk Attack. We introduce CIRA, a target-encoder-only attack that combines priority reallocation with evidence preservation and induces paired CSFs across compressor– budget settings without configuration-specific optimization. ❸ Comprehensive Evaluation and Analysis. We evaluate CIRA across four compressors, three datasets, four budgets, and multiple LVLM families, with mechanistic evidence of global priority reallocation and limited representation drift in displaced evidence. We also test TCS as a selectionstabilization defense against both standard and adaptive CIRA.
2
R ELATED W ORK
Visual token compression. Token-selection methods exploit visual cues, diversity, and salience– coverage objectives (Zhang et al., 2025b; Alvar et al., 2025; Xu et al., 2026), while learned pruning addresses limitations of attention-based importance estimates (Takezoe et al., 2026). Hybrid methods combine pruning with clustering and merging (Endo et al., 2025; Dhouib et al., 2025; Yang et al., 2025a). Query-conditioned methods use textual instructions for token scoring or aggregation (Yu et al., 2026; Gao et al., 2026; Sun et al., 2026). Adaptive and progressive methods vary token reduction across inputs and layers (Ye et al., 2025; Chen et al., 2026; Li et al., 2026a), while video and multi-turn approaches address temporal redundancy and evolving context (Wang et al., 2025a; Shen et al., 2024; Li et al., 2026b; Wang et al., 2026c). Adversarial robustness of vision-language models. Transfer-based attacks exploit set-level guidance, prompt-robust objectives, and iterative multimodal alignment (Lu et al., 2023; Luo et al., 2024; Liu et al., 2024a; Xie et al., 2025). Encoder-based and self-supervised attacks support cross-task and cross-model transfer (Zhang et al., 2025a; Hu et al., 2025; Zhang et al., 2026b). VEAttack disrupts visual representations (Mei et al., 2026b), while PA-Attack uses prototype-guided gray-box attacks (Mei et al., 2026a). Defenses include adversarial encoder fine-tuning (Schlarmann et al., 2024) and adversarial pre-training and instruction tuning (Wang et al., 2025c). Test-time defenses use prompt adaptation and augmented-view consistency (Sheng et al., 2025; Liu et al., 2025). Robustness under visual token compression. Visual-token compression introduces distinct robustness risks and defense opportunities. Safety-Aware Pruning (SAP) (Wang et al., 2026a) mitigates pruning-induced vulnerabilities, while robustness-oriented pruning (Gu et al., 2026) removes visually misaligned tokens. On the attack side, CAGE (Zhang et al., 2026c) targets tokens expected to survive unknown compression settings. CAA (Zhang et al., 2026a) directly manipulates token-selection rankings, with white-box attacks tailored to known compression configurations and transfer attacks using surrogate models. Our focus is on selectively inducing compression-specific failures under encoder-only access, without knowledge of the deployed compressor or budget.
3
D IAGNOSING C OMPRESSION -S PECIFIC FAILURES
In this section, we investigate why adversarial inputs can fail only after compression and use controlled diagnostics to inform the design of CIRA. 3.1
P ROBLEM S ETUP AND C OMPRESSION -S PECIFIC FAILURE
Let C denote a visual-token compressor and K its compression budget. For a dataset D = C,K {(xi , qi , yi )}M (xi , qi ) denote full-token and compressed inference, rei=1 , let fθ (xi , qi ) and fθ spectively. Given a task-specific binary evaluator χ(ŷ, y) ∈ {0, 1}, we define ci (x) = χ(fθ (x, qi ), yi ) , ci,C,K (x) = χ fθC,K (x, qi ), yi . For a fixed C, we write ci,K (x) ≡ ci,C,K (x) and omit the compressor subscript below. For each compression budget K, we restrict evaluation to samples for which both full-token and compressed inference produce correct answers before attack: SK = {i | ci (xi ) = 1, ci,K (xi ) = 1} . 3
(1)
Preprint. Under review.
A compression-specific failure (CSF) occurs when an adversarial input remains correct under fulltoken inference but fails after compression: CSF adv FK = i ∈ SK ci (xadv (2) i ) = 1, ci,K (xi ) = 0 . Aggregate accuracy can obscure instance-level prediction changes introduced by VLM acceleration (Sun et al., 2025). Our paired definition focuses on adversarial inputs that remain correct under full-token inference but fail after compression. 3.2
R ETAINED -S ET A LLOCATION AND C OMPRESSED C ORRECTNESS
Diagnostic setting. We test whether retained-set allocation can change compressed correctness while the adversarial encoder state and compression capacity remain fixed. LLaVA–VisionZip exposes direct-token membership, providing a controlled interface for retained-set intervention. From each downstream-agnostic VEAttack trajectory, we retain the earliest checkpoint satisfying the CSF criterion, termed a first-observed CSF. Counterfactual intervention. For each CSF, we exchange equal numbers of retained and omitted direct tokens, reconstruct compression, and rerun inference at the same encoder state and capacity. A post-hoc answer-aware oracle selects the exchanged token identities (Appendix C). Recovery criteria. Let ρ denote the fraction of exchanged direct-token slots. We compare guided reallocation with a size-matched random exchange. Cumulative recovery credits correction at any tested ρ′ ≤ ρ, whereas exact recovery uses only ρ. ✰ Observation 1: Controlled retained-set reallocation changes compressed correctness. Figure 2 shows that guided recovery exceeds matched-random recovery for every tested budget–fraction pair with ρ > 0, with an average advantage of 22.5–24.1 pp across the 12 dataset– budget cells. With encoder state, capacity, and exchange size fixed, the contrast shows that robustness depends on token identity, not only on retention count. 3.3 R ETAINED -S ET A LLOCATION C OMPONENTS AND R EPRESENTATION D RIFT Figure 2: Cumulative CSF recovery under Having established retained-set allocation sensi- guided and matched-random retained-set extivity, we next separate restoration from removal changes across direct-token exchange fractions. and ask when restored evidence remains useful. Observations 2–3 use ρ∗ = 7.5% as a shared intermediate diagnostic point, where guided and random recovery are already clearly separated. Representation drift is the restored tokens’ mean half-cosine distance, [1 − cos(hc , ha )]/2, between their clean and adversarial representations; Table 1(b) reports this distance multiplied by 100. ✰ Observation 2: Recovery depends on what is restored and removed. Table 1(a) decomposes guided reallocation into restoring omitted tokens with high retrospective answer support and removing retained tokens with low support; larger scores indicate stronger support for the reference answer. Exact recovery rises from 17.4% under matched random exchange to 42.0% under guided reallocation. Both restoration and removal contribute to recovery, with their coordinated use providing the largest gain. ✰ Observation 3: Restoration benefits are smaller under greater representation drift. Table 1(b) shows that evidence restoration is most effective for tokens with limited representation drift: the average effect falls from +20.5 pp in the lower-drift tertile to +10.3 pp in the upper. A within-cell continuous analysis shows the same negative association. 4
Preprint. Under review.
Table 1: Mechanistic diagnostics at ρ∗ = 7.5%: exact recovery under retained-set interventions (a) and evidence-restoration effects across representation-drift tertiles (b), with construction and aggregation details in Appendix C. (a) Retained-Set Counterfactuals Retained-Set Change Random + Evidence restoration + Low-support removal Guided reallocation
(b) Association with Representation Drift
Exact Recovery (%) Avg. K = 192 K = 128 K = 64 K = 32 17.9 31.1 34.5 46.4
17.2 33.5 31.9 44.4
18.4 34.0 25.4 38.5
16.1 34.9 20.1 38.6
17.4 33.4 28.0 42.0
Drift Tertile
Mean Drift (%)
Lower Middle Upper Lower–Upper Gap
31.4 39.5 45.2 –
Evidence Effect (pp) K = 192 K = 128 K = 64 K = 32
Avg.
+16.6 +15.3 +5.4 +11.2
+20.5 +14.0 +10.3 +10.3
+18.9 +12.7 +11.5 +7.4
+18.2 +7.0 +18.0 +0.2
+28.5 +20.8 +6.1 +22.3
Design implication. Together, these diagnostics motivate CIRA’s objective: make evidence inaccessible after compression while preserving its utility under full-token inference.
4
M ETHODOLOGY
4.1
ATTACK F ORMULATION
We consider an untargeted, image-specific attack with white-box access only to the vision encoder. The downstream question, reference answer, language model, compressor, and exact deployed compression budget are unavailable during optimization; the attacker knows only the admissible range K = {K ∈ N : Kmin ≤ K ≤ Kmax }. For an evaluation tuple (x, q, y), the ideal compression-specific outcome is to preserve full-token correctness while causing compressed inference to fail: max min L fθC,K (x + δ, q), y s.t. c(x + δ) = 1, (3) ∥δ∥∞ ≤ϵ K∈K
where c(·) denotes full-token correctness and L increases with compressed prediction error. Equation 3 defines the desired compression-specific behavior but is not optimized directly. CIRA instead replaces it with compressor-independent encoder-side objectives and evaluates cross-configuration transfer empirically. 4.2
G LOBAL S ELECTION H IJACKING
Let si (x) denote the encoder-side proxy priority score of visual token i, where larger values indicate higher retention priority. We collect these scores in the priority-score vector s(x) = [s1 (x), . . . , sN (x)] ∈ RN , whose descending order defines the priority ranking over tokens. We write sc = s(x) and sa = s(x + δ) for the clean and adversarial priority-score vectors, respectively. We first construct the clean rank, its inverse-priority encoding, and a stabilized standardization operator: rc − 1 u − ū1 . (4) rc = rank↓ (sc ), R(sc )i = i , Zϵs (u) = −1/2 N −1 max N ∥u − ū1∥2 , ϵs Here, rc ranks tokens from high to low clean priority, while R maps high-priority tokens near zero and low-priority tokens near one. Zϵs standardizes its input vector and prevents a degenerate denominator when its variance vanishes. To drive a global priority inversion, CIRA maximizes the standardized alignment between adversarial priorities and the inverse clean priority ranking: 1 LGSH (δ) = Alignϵs (sa , R(sc )) ≡ ⟨Zϵs (sa ), Zϵs (R(sc ))⟩ . (5) N This differentiable, scale-invariant objective encourages clean high-priority tokens to move downward while promoting clean low-priority tokens. 4.3
H IDDEN -E VIDENCE P RESERVATION
Priority reallocation may also perturb the representations of displaced evidence, undermining fulltoken correctness. HEP therefore focuses preservation on clean high-priority tokens that cross 5
Preprint. Under review.
candidate retention boundaries. CIRA computes the adversarial descending ranks ra = rank↓ (sa ) and defines HK (δ) = {i : ric ≤ K < ria }, the clean proxy Top-K tokens whose adversarial ranks move beyond the retention boundary at compression budget K. Under an unknown compression budget, their displacement weights are wi =
Pr K∼Unif(K)
[i ∈ HK (δ)] =
[min(Kmax , ria − 1) − max(Kmin , ric ) + 1]+ . Kmax − Kmin + 1
(6)
Here, [z]+ = max(z, 0), and wi is the fraction of admissible budgets under which token i becomes hidden. To limit representation drift in these tokens, let hci and hai denote their clean and adversarial visualtoken features. We normalize their directions and measure the resulting representation drift by c a 1 e c = hi , e a = hi , e c )⊤ h ea , h h di = 1 − (h (7) i i i i c a ∥hi ∥2 ∥hi ∥2 2 where the factor 1/2 normalizes cosine distance to [0, 1]. Hidden-Evidence Preservation aggregates these token-level distances using the budget-marginal weights: N X sg(wi ) nP o. LHEP (δ) = − w ei di , w ei = (8) N max i=1 j=1 sg(wj ), ϵh The stop-gradient freezes rank-derived weights within an update, while ϵh stabilizes normalization; weights are recomputed at the next update so evidence displaced across more of K receives greater protection. 4.4
J OINT O PTIMIZATION
CIRA combines the two objectives as max δ
LCIRA (δ) = LGSH (δ) + λ · LHEP (δ),
s.t.
∥δ∥∞ ≤ ϵ.
(9)
Here, λ controls the preservation strength. We maximize equation 9 using projected sign-gradient ascent over the valid-image domain. The complete optimization procedure is given in Appendix A.
5
E XPERIMENTS
5.1
E XPERIMENTAL S ETUP
Models and Benchmarks. We evaluate LLaVA-v1.5-7B (Liu et al., 2024b) on 1,000 randomly sampled image–question pairs from each of POPE (Li et al., 2023b), TextVQA (Singh et al., 2019), and MME (Fu et al., 2025), covering object hallucination, scene-text understanding, and general visual perception and reasoning, respectively. We additionally evaluate Qwen3-VL-8B-Instruct (Bai et al., 2025) and InternVL3.5-8B (Wang et al., 2025b). Compression Settings. We evaluate VisionZip (Yang et al., 2025b), VisPruner (Zhang et al., 2025b), PruMerge (Shang et al., 2025), and FastV (Chen et al., 2024) at Keval = {32, 64, 128, 192}. For each input, one adversarial image is reused across all compressor–budget settings. Baselines. We compare CIRA with the downstream-agnostic VEAttack (Mei et al., 2026b) and CAGE (Zhang et al., 2026c); CAA (Zhang et al., 2026a) is reported as a stronger-access reference. Evaluation Metrics. On the clean-eligible set SK defined in Eq. 1, we report the CompressionSpecific Failure Rate (CSFR) and full-token attack success rate: X 1 X 1 adv CSFRK = ci (xadv ASR = 1 − ci (xadv i ) 1 − ci,K (xi ) , i ) . |SK | |Sfull | i∈SK i∈Sfull (10) 6
Preprint. Under review.
Table 2: CSFR and Full ASR on LLaVA-v1.5-7B. For downstream-agnostic attacks, red and blue cells mark highest CSFR and lowest Full ASR, and underlining marks second best; CAA† uses the downstream question and language model during optimization and is excluded from the markings. Metric
Full ASR ↓
POPE TextVQA MME VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA 38.65
39.48
Full-token Inference (Compressor-independent) 2.72 4.96 60.95 71.35 3.82 10.61
38.26
39.77
3.41
5.18
CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
4.05 4.98 5.05 6.52 5.15
3.19 5.72 9.71 9.48 7.03
1.72 2.87 3.46 4.57 3.16
7.61 14.05 22.47 26.96 17.77
VisionZip 4.87 5.32 6.28 9.76 6.56
2.92 5.32 6.28 5.24 4.94
2.14 2.05 2.39 5.01 2.90
11.35 18.21 27.41 31.70 22.17
2.83 3.33 3.64 4.56 3.59
4.04 5.55 5.97 5.19 5.19
1.62 1.80 2.92 2.86 2.30
7.40 10.26 19.07 22.17 14.73
CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
3.42 4.60 6.27 7.59 5.47
3.79 4.35 7.45 10.46 6.51
1.23 2.97 4.68 8.29 4.29
5.87 9.07 15.29 22.21 13.11
VisPruner 2.92 3.50 3.80 3.40 5.89 5.47 8.07 7.33 5.17 4.93
1.75 2.18 4.08 4.63 3.16
10.50 13.18 19.13 25.57 17.09
2.56 2.88 3.33 4.13 3.22
3.50 4.79 5.36 5.66 4.83
1.22 2.33 4.47 6.73 3.69
7.14 11.23 18.70 21.87 14.73
CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
5.33 5.39 6.96 9.40 6.77
3.65 5.54 9.38 11.76 7.58
2.66 2.92 3.63 3.44 3.16
22.86 29.59 32.53 32.29 29.32
PruMerge 6.02 4.86 6.51 5.35 8.29 6.64 9.18 6.70 7.50 5.89
3.44 4.86 4.49 5.20 4.50
28.66 36.23 43.91 48.68 39.37
2.64 3.62 3.50 4.06 3.45
4.69 5.28 5.78 5.78 5.38
1.75 3.02 3.04 2.81 2.66
20.64 22.93 29.98 35.16 27.18
CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
4.61 6.85 9.90 12.32 8.42
4.74 6.59 10.47 10.40 8.05
2.30 2.90 3.44 4.48 3.28
7.04 11.46 22.24 25.44 16.55
FastV 4.29 7.10 9.11 7.80 7.08
1.43 2.15 4.08 3.76 2.85
9.61 13.55 23.74 32.95 19.96
2.77 3.17 3.65 4.24 3.46
2.90 3.89 5.18 5.55 4.38
1.38 1.44 2.89 6.20 2.98
5.81 7.64 15.07 20.39 12.23
4.09 5.38 7.67 8.38 6.38
Here, Sfull = {i : ci (xi ) = 1}, and Avg. CSFR is the arithmetic mean of CSFRK over K ∈ Keval for a fixed dataset–compressor setting. Broader summaries weight each reported dataset–compressor setting equally. CSFR counts post-attack failures confined to the compressed path, normalized over inputs answered correctly by both clean inference paths; ASR is normalized over inputs answered correctly by clean full-token inference. Clean and post-attack accuracy results are reported in Appendix G. Implementation Settings. We run all experiments on a single NVIDIA GeForce RTX 4090 GPU. All downstream-agnostic attacks use ϵ = 4/255 and 100 optimization steps; CIRA uses projected sign-gradient ascent with λ = 0.8. Further protocol details and sensitivity analyses appear in Appendix B and Appendix D. 5.2
M AIN R ESULTS
In this section, we evaluate CIRA’s compression selectivity and cross-configuration transfer, and compare it with stronger-access attacks. Compression-specific selectivity. Across the four budgets and 12 dataset–compressor settings in Table 2, CIRA achieves a mean CSFR of 20.35%, compared with 5.49% for VEAttack and 5.92% for CAGE. Its Full ASR is 6.92%, substantially below 45.95% and 50.20%, respectively. CIRA therefore induces more compression-specific failures while causing much less full-token degradation than either downstream-agnostic baseline. Cross-configuration transfer. The same adversarial image is reused without re-optimization across selection-, pruning-, and merging-based compression rules and all evaluated budgets. CIRA remains effective across these configurations, reaching 31.96% mean CSFR on PruMerge compared with 6.28% for CAGE. Averaged equally over all 12 dataset–compressor settings, its CSFR increases 7
Preprint. Under review.
Figure 3: CIRA-induced priority reallocation across the Figure 4: Priority reallocation and repreclean visual-token ranking at K = 128. sentation drift across token groups. from 12.04% at K = 192 to 28.78% at K = 32. Compression-selective behavior also extends to Qwen3-VL and InternVL (Table 10); qualitative cross-setting examples appear in Appendix I. Comparison with stronger-access CAA. CAA† optimizes with access to the downstream question and language model, whereas CIRA uses only the vision encoder. Despite its lower Full ASR (3.32% for CAA† versus 6.92% for CIRA), CAA† achieves only 3.24% mean CSFR, well below CIRA’s 20.35%. The contrast shows that preserving the full-token prediction is not sufficient to induce compression-specific failures; CIRA’s priority-reallocation objective addresses this distinct requirement under narrower access. ➪ Takeaway. CIRA induces compression-specific failures across heterogeneous compression rules and budgets while limiting full-token degradation. 5.3
M ECHANISTIC A NALYSIS
To connect CIRA’s behavior to its objectives, we test whether it globally reallocates token priority while limiting representation drift in displaced evidence. Priority reallocation. Figure 3 shows a near-monotonic priority reallocation across all three benchmarks: the highest-priority octile is demoted by 0.47–0.53 normalized-rank units, whereas the lowest-priority octile is promoted by 0.73–0.76, with the sign changing near the median. This cross-quantile pattern is not confined to a single retention boundary. Priority reallocation versus representation drift. Figure 4 shows that rank displacement and representation drift behave differently across clean-priority groups. Rank demotion peaks at 53.0% and 53.7% in the top 10% and 10–30% groups, where representation drift is lowest at 12.4% and 12.7%. The lowest-priority 60–100% group shows the reverse pattern, with 7.8% rank demotion and 26.1% drift. This separation indicates that CIRA reallocates clean high-priority tokens while limiting their representation drift, consistent with the HEP objective. Component ablation. On VisionZip, Table 3 isolates the two objectives. Removing GSH reduces CSFR averaged across the three benchmarks and four budgets from 18.22% to 3.68%. Removing HEP leaves this mean nearly unchanged (17.04%) but raises mean Full ASR from 6.92% to 23.10%; replacing it with global representation preservation yields 15.71% mean CSFR and 8.66% Full ASR. Global preservation yields lower CSFR and higher Full ASR than HEP, supporting preservation focused on displaced evidence rather than a uniform constraint. ➪ Takeaway. GSH drives compression-specific failure induction, while HEP limits representation drift in displaced evidence and substantially reduces full-token degradation. 8
Preprint. Under review.
Table 3: CIRA ablation on VisionZip; arrows Table 4: TCS evaluation on VisionZip; arrows denote show changes from full CIRA. changes from matched no-defense baselines. Metric
Full CIRA
Full ASR ↓ CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
4.96 7.61 14.05 22.47 26.96 17.77
Full ASR ↓ CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
10.61 11.35 18.21 27.41 31.70 22.17
Full ASR ↓ CSFR@192 ↑ CSFR@128 ↑ CSFR@64 ↑ CSFR@32 ↑ Avg. CSFR ↑
5.18 7.40 10.26 19.07 22.17 14.73
6
Selection Preservation w/o Sel. w/o Pres. Global Pres. POPE 2.01↓2.95 15.72↑10.76 4.73↓0.23 1.60↓6.01 9.94↑2.33 4.79↓2.82 1.99↓12.06 15.42↑1.37 9.95↓4.10 4.39↓18.08 22.07↓0.40 19.95↓2.52 5.63↓21.33 26.81↓0.15 26.37↓0.59 3.40↓14.37 18.56↑0.79 15.26↓2.51 TextVQA 6.91↓3.70 36.00↑25.39 15.82↑5.20 3.90↓7.45 9.94↓1.41 8.58↓2.77 2.05↓16.16 15.20↓3.02 12.53↓5.69 5.43↓21.98 23.70↓3.72 23.04↓4.37 6.44↓25.26 26.73↓4.97 36.28↑4.57 4.46↓17.71 18.89↓3.28 20.11↓2.06 MME 2.40↓2.77 17.57↑12.40 5.44↑0.26 1.62↓5.79 7.55↑0.14 5.80↓1.61 1.93↓8.33 11.33↑1.06 7.73↓2.53 4.09↓14.98 17.23↓1.84 14.45↓4.62 5.08↓17.09 18.57↓3.60 19.05↓3.12 3.18↓11.55 13.67↓1.06 11.76↓2.97
Metric Full ASR CSFR@192 ↓ CSFR@128 ↓ CSFR@64 ↓ CSFR@32 ↓ Avg. CSFR ↓ Full ASR CSFR@192 ↓ CSFR@128 ↓ CSFR@64 ↓ CSFR@32 ↓ Avg. CSFR ↓ Full ASR CSFR@192 ↓ CSFR@128 ↓ CSFR@64 ↓ CSFR@32 ↓ Avg. CSFR ↓
CAGE None TCS POPE 4.96 39.48 7.61 2.60↓5.01 3.19 4.34↑1.15 14.05 3.40↓10.65 5.72 4.79↓0.93 22.47 2.84↓19.63 9.71 5.68↓4.03 26.96 4.21↓22.75 9.48 5.52↓3.96 17.77 3.26↓14.51 7.03 5.08↓1.95 TextVQA 10.61 71.35 11.35 2.39↓8.96 2.92 1.59↓1.33 18.21 2.07↓16.14 5.32 2.48↓2.84 27.41 4.50↓22.91 6.28 3.43↓2.85 31.70 3.09↓28.61 5.24 4.99↓0.25 22.17 3.01↓19.16 4.94 3.12↓1.82 MME 5.18 39.77 7.40 2.72↓4.68 4.04 2.31↓1.73 10.26 2.64↓7.62 5.55 2.50↓3.05 19.07 4.53↓14.54 5.97 4.09↓1.88 22.17 5.75↓16.42 5.19 4.67↓0.52 14.73 3.91↓10.82 5.19 3.39↓1.80 None
CIRA TCS
Adaptive CIRA None TCS
6.99 10.07 20.08 28.00 16.29
3.78 6.57↓0.42 10.71↑0.64 13.65↓6.43 21.92↓6.08 13.21↓3.08
11.45 8.38 8.17↓0.21 11.09 9.30↓1.79 26.09 14.13↓11.95 38.90 25.65↓13.25 21.11 14.31↓6.80 6.57 3.64 5.03↑1.39 9.53 7.22↓2.31 18.54 11.68↓6.86 23.17 18.35↓4.82 13.72 10.57↓3.15
S ELECTION S TABILIZATION D EFENSE
To test whether stabilizing token selection can suppress CIRA, TCS exploits cross-view priority stability: clean high-priority evidence tends to remain stable under small translations, whereas attackinduced replacements are more view-sensitive. At the LLaVA–VisionZip token-priority interface, TCS uses four translated views V = {T0,0 , Td,0 , T0,d , Td,d } with d = 7 pixels. For view v, s(v) is its priority-score vector, Av maps the translated score grid back to the reference coordinates, and rank 1 denotes the highest priority. TCS selects the K tokens with highest aligned rank-quantile consensus: ri Av (s(v) ) − 1 1 X (v) (v) TCS qi = 1 − , q̄i = qi , IK = TopK(q̄, K). (11) N −1 |V| v∈V
Rank quantiles make priorities comparable across views despite differences in score scale. The selected indices are applied to the unshifted-view features, leaving aggregation and full-token inference unchanged. Construction and mechanism details appear in Appendix F; complementary utility results are reported in Appendix G. Defense effectiveness. Within the matched evaluation blocks in Table 4, TCS reduces standard CIRA CSFR averaged across the three datasets and four budgets from 18.22% to 3.39%, an 81.4% relative reduction, compared with a 32.5% reduction for CAGE. Suppression strengthens as the compression budget decreases, with the largest reductions under tighter compression. Adaptive stress test. Because TCS is deterministic and public, we also evaluate an adaptive attacker that optimizes the CIRA objectives over the same four views while sharing one image-space perturbation: i 1 X h (v) (v) Ladapt (δ) = LGSH (δ) + λLHEP (δ) . (12) |V| v∈V
Adaptive CIRA uses the same access assumptions and perturbation budget as standard CIRA. Under TCS, its mean CSFR rises from 3.39% to 12.70%, reaching 74.5% of Adaptive CIRA’s matched undefended value (17.04%). TCS therefore retains a smaller but nonzero effect against the adaptive attack. Additional optimization and mechanism diagnostics appear in Appendix F.3.
7
C ONCLUSION
Visual-token compression changes not only inference cost but also which visual evidence remains available after compression. By pairing full-token and compressed inference on the same adversarial 9
Preprint. Under review.
input, we attribute failures specifically to the compression path. CIRA induces such failures across unknown compressor and budget settings by reallocating token priorities while limiting full-token degradation. Controlled diagnostics further show that retained-set allocation affects compressed correctness and that recovery is negatively associated with representation drift in displaced evidence. Selection stabilization substantially suppresses CIRA, although adaptive optimization partially restores its effectiveness. Together, these results support treating the compression boundary as a security-relevant component of LVLM deployment and motivate paired robustness evaluation for compressed LVLMs.
AI U SE S TATEMENT Generative AI tools were used in a limited supporting role for this work, including language editing and polishing, literature retrieval and discovery, research ideation and technical execution support, and drafting and revising parts of the manuscript. In particular, these tools were used to improve the clarity and fluency of writing, help identify related literature, provide technical suggestions for coding and experimental workflows, and assist in refining sections of the paper during revision. All AI-assisted content was independently checked, validated, and, where necessary, modified by the authors. The authors retain full responsibility for the scientific content, methodological choices, experimental evidence, interpretations, and final form of the manuscript.
R EPRODUCIBILITY S TATEMENT We provide comprehensive details to facilitate the reproduction and verification of CIRA. The threat model, compression-specific failure formulation, and attack objectives are described in Sections 3 and 4, including the Global Selection Hijacking (GSH) and Hidden-Evidence Preservation (HEP) objectives and their joint optimization. The optimization procedure and experimental details are provided in Appendices A and B, respectively. Main results, mechanistic analyses, and defense evaluations are reported in Sections 5.2, 5.3, and 6. Additional analyses, including retained-set diagnostics, sensitivity studies, cross-model evaluation, task-utility results, limitations, and qualitative cases, are provided in Appendices C, D, E, G, H, and I. Code and scripts for reproducing the reported experiments are provided in the Github repository.
E THICS S TATEMENT This work investigates the adversarial robustness of LVLMs under visual-token compression in a controlled research setting. Our experiments use publicly available benchmark datasets (e.g., POPE, TextVQA, and MME) and publicly released pretrained models (e.g., LLaVA-v1.5-7B, Qwen3-VL-8BInstruct, and InternVL3.5-8B), involving no human subjects or personally identifiable information. We recognize the dual-use nature of adversarial robustness research. CIRA is developed to identify compression-specific vulnerabilities and support robustness evaluation and defense development. The proposed attack is evaluated under a restricted threat model with bounded image-space perturbations and vision-encoder-only access. We transparently report the attack assumptions and evaluation protocols to facilitate reproducibility and responsible security research.
R EFERENCES Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022. Saeed Ranjbar Alvar, Gursimran Singh, Mohammad Akbari, and Yong Zhang. Divprune: Diversitybased visual token pruning for large multimodal models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9392–9401. IEEE, 2025. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. 10
Preprint. Under review.
Adrian Bulat, Yassine Ouali, and Georgios Tzimiropoulos. Compress & cache: Vision token compression for efficient generation and retrieval. Advances in Neural Information Processing Systems, 38:31943–31968, 2026. Junjie Chen, Xuyang Liu, Zichen Wen, Yiyu Wang, Siteng Huang, and Honggang Chen. Variationaware vision token dropping for faster large vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3489–3499, 2026. Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large visionlanguage models. In European Conference on Computer Vision, pp. 19–35. Springer, 2024. Hyeonwoo Cho, Donghyeon Baek, Yewon Kim, and Bumsub Ham. Improving visual token reduction via rectifying distortions for efficient multimodal llm inference. arXiv preprint arXiv:2606.01711, 2026. Mohamed Dhouib, Davide Buscaldi, Sonia Vanier, and Aymen Shabou. Pact: Pruning and clusteringbased token reduction for faster visual language models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14582–14592. IEEE, 2025. Mark Endo, Xiaohan Wang, and Serena Yeung-Levy. Feather the throttle: Revisiting visual token pruning for vision-language model acceleration. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 22826–22835. IEEE, 2025. Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, Rongrong Ji, Caifeng Shan, and Ran He. Mme: A comprehensive evaluation benchmark for multimodal large language models. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (eds.), Advances in Neural Information Processing Systems, volume 38, Main Conference. Curran Associates, Inc., 2025. doi: 10.52202/085713-4899. URL https://proceedings.neurips.cc/paper_files/ paper/2025/file/d79a27cf2772fe00be7f341efc0eb517-Paper-Datasets_ and_Benchmarks_Track.pdf. Tianxiao Gao, Shanwei Zhao, Shuo Fang, Shiai Zhu, and Chenguang Ma. Quietprune: Query-guided early token pruning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3553–3562, 2026. Shishen Gu, Jiequan Cui, Wenbo Hu, Zenglin Shi, Zhenzhen Hu, and Richang Hong. Visual token compression enhances robustness of mllms. arXiv preprint arXiv:2607.22716, 2026. Kai Hu, Weichen Yu, Li Zhang, Alexander Robey, Andy Zou, Chengming Xu, Haoqi Hu, and Matt Fredrikson. Transferable adversarial attacks on black-box vision-language models. arXiv preprint arXiv:2505.01050, 2025. Yihong Huang, Fei Ma, Yihua Shao, Jingcai Guo, Zitong Yu, Laizhong Cui, and Qi Tian. N\" uwa: Mending the spatial integrity torn by vlm token pruning. arXiv preprint arXiv:2602.02951, 2026. Ao Li, Yuxiang Duan, Jinghui Zhang, Congbo Ma, Yutong Xie, Gustavo Carneiro, Mohammad Yaqub, and Hu Wang. Transprune: Token transition pruning for efficient large vision-language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 39529–39538, 2026a. Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PmLR, 2023a. Wentong Li, Yuqian Yuan, Jian Liu, Dongqi Tang, Song Wang, Jie Qin, Jianke Zhu, and Lei Zhang. Tokenpacker: Efficient visual projector for multimodal llm. International Journal of Computer Vision, 133(10):6794–6812, 2025. Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 conference on empirical methods in natural language processing, pp. 292–305, 2023b. 11
Preprint. Under review.
Zhenyu Li, Zuchao Li, Ping Wang, Lefei Zhang, and Haojun Ai. Vista-llm: Decoupled query-guided visual token pruning for efficient long-video large language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 13171–13187, 2026b. Daizong Liu, Mingyu Yang, Xiaoye Qu, Pan Zhou, Xiang Fang, Keke Tang, Yao Wan, and Lichao Sun. Pandora’s box: Towards building universal attackers against real-world large vision-language models. Advances in Neural Information Processing Systems, 37:52127–52158, 2024a. Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 26286–26296. IEEE, 2024b. Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari, and Mingkun Xu. Self-calibrated consistency can fight back for adversarial robustness in vision-language models. arXiv preprint arXiv:2510.22785, 2025. Dong Lu, Zhiqiang Wang, Teng Wang, Weili Guan, Hongchang Gao, and Feng Zheng. Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 102–111. IEEE, 2023. Haochen Luo, Jindong Gu, Fengyuan Liu, and Philip Torr. An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models. arXiv preprint arXiv:2403.09766, 2024. Jianxin Ma, Shibo Jin, and Lujuan Dang. Visual token compression via run-length pruning in multimodal large language models. Pattern Recognition Letters, 2026. Hefei Mei, Zirui Wang, Chang Xu, Jianyuan Guo, and Minjing Dong. Pa-attack: Guiding gray-box attacks on lvlm vision encoders with prototypes and attention. arXiv preprint arXiv:2602.19418, 2026a. Hefei Mei, Zirui Wang, Shen You, Minjing Dong, and Chang Xu. Veattack: Downstream-agnostic vision encoder attack against large vision language models. In International Conference on Learning Representations, volume 2026, pp. 18135–18161, 2026b. Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pp. 21527–21536, 2024. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PmLR, 2021. Christian Schlarmann and Matthias Hein. On the adversarial robustness of multi-modal foundation models. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 3679–3687. IEEE, 2023. Christian Schlarmann, Naman Deep Singh, Francesco Croce, and Matthias Hein. Robust clip: Unsupervised adversarial fine-tuning of vision embeddings for robust large vision-language models. arXiv preprint arXiv:2402.12336, 2024. Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 22857–22867. IEEE, 2025. Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models. In International conference on learning representations, volume 2024, pp. 30853–30885, 2024. Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, et al. Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434, 2024. 12
Preprint. Under review.
Lijun Sheng, Jian Liang, Zilei Wang, and Ran He. R-tpt: Improving adversarial robustness of visionlanguage models through test-time prompt tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 29958–29967, 2025. Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8309–8318. IEEE, 2019. Guohao Sun, Yufei Wang, Sizhuo Ma, Yuege Xie, Yuting Cheng, Zhiqiang Tao, and Jian Wang. Ifprune: Information-flow guided token pruning for efficient vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3522–3531, 2026. Yizheng Sun, Hao Li, Chang Xu, Hongpeng Zhou, Chenghua Lin, Riza Theresa Batista-Navarro, and Jingyuan Sun. Does acceleration cause hidden instability in vision language models? uncovering instance-level divergence through a large-scale empirical study. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 8453–8467, 2025. Rinyoichi Takezoe, Yaqian Li, Zi-Hao Bo, Anzhou Hou, Mo Guang, and Kaiwen Long. Learnpruner: Rethinking attention-based token pruning in vision language models. In International Conference on Learning Representations, volume 2026, pp. 66381–66400, 2026. Han Wang, Yuxiang Nie, Yongjie Ye, Yanjie Wang, Shuai Li, Haiyang Yu, Jinghui Lu, and Can Huang. Dynamic-vlm: Simple dynamic visual token compression for videollm. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 20812–20823. IEEE, 2025a. Shuailong Wang, Xinyu Lyu, Shengming Yuan, Jingkuan Song, Heng Tao Shen, and Lianli Gao. Understanding and mitigating token-pruning-induced vulnerabilities in VLMs. In Forty-third International Conference on Machine Learning, 2026a. URL https://openreview.net/ forum?id=D3OHVbePvz. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, Guanzhou Chen, Zichen Ding, Changyao Tian, Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Yuchen Duan, Xuehui Wang, Zhi Hou, Haoran Hao, Tianyi Zhang, Songze Li, Xiangyu Zhao, Haodong Duan, Nianchen Deng, Bin Fu, Yinan He, Yi Wang, Conghui He, Botian Shi, Junjun He, Yingtong Xiong, Han Lv, Lijun Wu, Wenqi Shao, Kaipeng Zhang, Huipeng Deng, Biqing Qi, Jiaye Ge, Qipeng Guo, Wenwei Zhang, Songyang Zhang, Maosong Cao, Junyao Lin, Kexian Tang, Jianfei Gao, Haian Huang, Yuzhe Gu, Chengqi Lyu, Huanze Tang, Rui Wang, Haijun Lv, Wanli Ouyang, Limin Wang, Min Dou, Xizhou Zhu, Tong Lu, Dahua Lin, Jifeng Dai, Weijie Su, Bowen Zhou, Kai Chen, Yu Qiao, Wenhai Wang, and Gen Luo. Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency, 2025b. URL https://arxiv.org/abs/2508.18265. Yahong Wang, Juncheng Wu, Zhangkai Ni, Longzhen Yang, Yihang Liu, Chengmei Yang, Ying Wen, Lianghua He, Xianfeng Tang, Hui Liu, et al. When token pruning is worse than random: Understanding visual token information in vllms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 31910–31919, 2026b. Yi Wang, Haofei Zhang, Qihan Huang, Anda Cao, Gongfan Fang, Wei Wang, Xuan Jin, Jie Song, Mingli Song, and Xinchao Wang. Rethinking token reduction for large vision-language models. arXiv preprint arXiv:2603.21701, 2026c. Zeyu Wang, Cihang Xie, Brian Bartoldson, and Bhavya Kailkhura. Double visual defense: Adversarial pre-training and instruction tuning for improving vision-language model robustness. arXiv preprint arXiv:2501.09446, 2025c. Peng Xie, Yequan Bie, Jianda Mao, Yangqiu Song, Yang Wang, Hao Chen, and Kani Chen. Chain of attack: On the robustness of vision-language models against transfer-based adversarial attacks. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14679– 14689. IEEE, 2025. 13
Preprint. Under review.
Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang, Feng Wu, et al. Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction. arXiv preprint arXiv:2410.17247, 2024. Tong Xu, Hailong Shi, and Xingyu Gao. Score: Salience-coverage reduction for vision token pruning in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 24686–24695, 2026. Cheng-Yu Yang, Shao-Yuan Lo, and Yu-Lun Liu. Reroute, don’t remove: Recoverable visual token routing for vision-language models. arXiv preprint arXiv:2606.12412, 2026. Longrong Yang, Dong Shen, Chaoxiang Cai, Kaibing Chen, Fan Yang, Tingting Gao, Di Zhang, and Xi Li. Libra-merging: Importance-redundancy and pruning-merging trade-off for acceleration plug-in in large vision-language model. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9402–9412. IEEE, 2025a. Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. Visionzip: Longer is better but not necessary in vision language models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19792–19802. IEEE, 2025b. Xubing Ye, Yukang Gan, Yixiao Ge, Xiao-Ping Zhang, and Yansong Tang. Atp-llava: Adaptive token pruning for large vision language models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 24972–24982. IEEE, 2025. Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, and Fenglong Ma. Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models. Advances in Neural Information Processing Systems, 36:52936–52956, 2023. Hanxun Yu, Wentong Li, Xuan Qu, Song Wang, Junbo Chen, and Jianke Zhu. Visiontrim: Unified vision token compression for training-free mllm acceleration. arXiv preprint arXiv:2601.22674, 2026. Jiaming Zhang, Qi Yi, and Jitao Sang. Towards adversarial attack on vision-language pre-training models. In Proceedings of the 30th ACM international conference on multimedia, pp. 5005–5013, 2022. Jiaming Zhang, Junhong Ye, Xingjun Ma, Yige Li, Yunfan Yang, Yunhao Chen, Jitao Sang, and DitYan Yeung. Anyattack: Towards large-scale self-supervised adversarial attacks on vision-language models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19900–19909. IEEE, 2025a. Qizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang, Zhiyong Zhuo, Jiajun Cao, Shaobo Guo, Qi She, and Shanghang Zhang. Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 20857–20867. IEEE, 2025b. Xiaomei Zhang, Zhaoxi Zhang, Leo Yu Zhang, Yanjun Zhang, Guanhong Tao, and Shirui Pan. Less is more–until it breaks: Security pitfalls of vision token compression in large vision-language models. arXiv preprint arXiv:2601.12042, 2026a. Xinwei Zhang, Li Bai, Tianwei Zhang, Youqian Zhang, Qingqing Ye, Yingnan Zhao, Ruochen Du, and Haibo Hu. Grounding-driven attack: Improving encoder-based adversarial transferability against large vision-language models. arXiv preprint arXiv:2602.09431, 2026b. Xinwei Zhang, Hangcheng Liu, Li Bai, Hao Wang, Qingqing Ye, Tianwei Zhang, and Haibo Hu. On the adversarial robustness of large vision-language models under visual token compression. arXiv preprint arXiv:2601.21531, 2026c. Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al. Sparsevlm: Visual token sparsification for efficient vision-language model inference. arXiv preprint arXiv:2410.04417, 2024.
14
Preprint. Under review.
A PPENDIX C ONTENTS Appendix A. CIRA Optimization Procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 Appendix B. Detailed Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 Appendix C. Controlled Retained-Set Allocation Diagnostics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 Appendix D. Sensitivity Analyses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 Appendix E. Cross-Model Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 Appendix F. Selection Stabilization Defense . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 Appendix G. Complementary Task-Utility Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 Appendix H. Limitation Discussion and Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 Appendix I. Qualitative Case Studies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27
A
CIRA O PTIMIZATION P ROCEDURE
Algorithm 1 gives the image-space optimization induced by CIRA’s two encoder-side objectives. Clean features and priority scores are cached once; adversarial scores, ranks, and displacement weights are refreshed at every step. The priority rule P maps encoder attention to the token-score vector used by GSH. Algorithm 1 Projected sign-gradient optimization of CIRA. Require: Image x; encoder E; priority rule P Require: Candidate interval [Kmin , Kmax ]; ϵ, ϵs , ϵh Require: Step size α; steps T ; preservation weight λ Ensure: Adversarial image xadv 1: K ← {Kmin , . . . , Kmax } 2: (Hc , sc ) ← (E(x), P(x)) 3: rc ← Rank↓ (sc ) 4: uc ← R(sc ) 5: Initialize δ (0) ← 0 6: for τ = 0, . . . , T − 1 do 7: x(τ ) ← x + δ (τ ) 8: (Ha , sa ) ← (E(x(τ ) ), P(x(τ ) )) 9: LGSH ← Alignϵs (sa , uc ) 10: ra ← Rank↓ (sa ) 11: wi ← sg PrK∼Unif(K) [ric ≤ K < ria ] c 12: di ← 21 (1 − cos(h , hai )) P i wi di i 13: LHEP ← − max(P wi ,ϵh ) i 14: LCIRA ← LGSH + λLHEP 15: g ← sign(∇δ LCIRA ) 16: δ (τ +1) ← Π∆ϵ (x) (δ (τ ) + αg) 17: end for 18: return xadv ← x + δ (T )
Sorting remains outside of the gradient path, sg(·) blocks gradients through the discrete weights, and ∆ϵ (x) = {δ : ∥δ∥∞ ≤ ϵ, x + δ ∈ [0, 1]d } is the feasible perturbation set. The update therefore uses only the vision encoder, priority-scoring rule, and candidate compression-budget interval.
B
D ETAILED E XPERIMENTAL S ETUP
B.1
M ODELS
LLaVA-v1.5-7B. Our primary model is LLaVA-v1.5-7B (Liu et al., 2024b) with a CLIP ViT-L/14336 vision encoder (Radford et al., 2021). The released model feeds 576 penultimate-layer patch tokens to its multimodal projector. We use LLaVA for the full benchmark–compressor matrix and for the mechanism, ablation, sensitivity, and defense studies. 15
Preprint. Under review.
Additional LVLM families. We additionally evaluate Qwen3-VL-8B-Instruct (Bai et al., 2025) and InternVL3.5-8B (Wang et al., 2025b). Qwen3-VL uses its native dynamic resolution and spatial merger, so its visual-token count varies by image. InternVL uses one 448 × 448 tile, producing 256 tokens after spatial downsampling. Accordingly, we use family-specific evaluation budgets of {96, 64, 32} for Qwen3-VL and {128, 64, 32} for InternVL. CIRA is optimized separately for each model using its native vision encoder. B.2
B ENCHMARKS
We use fixed, randomly sampled 1,000-pair subsets from POPE (Li et al., 2023b), TextVQA (Singh et al., 2019), and MME (Fu et al., 2025). POPE evaluates object hallucination through balanced object-presence questions; TextVQA tests reading of text in natural images and provides ten reference answers per question; MME covers perception and cognition with binary questions. Within each model–benchmark setting, all attacks are evaluated on the same image–question pairs using a common prompt, decoding protocol, and answer evaluator. POPE and MME outputs are lowercased, stripped of punctuation, and matched by the first normalized yes/no token. TextVQA uses its EvalAI-style normalizer and accepts a match to any reference answer. These task-specific rules provide the per-example binary correctness required by the paired CSF definition. For TextVQA, we use a match to any normalized reference rather than the benchmark-level soft agreement score; MME is evaluated per question rather than by its aggregate category score. The same evaluators are used for clean, full-token, and compressed inference. B.3
C OMPRESSION M ECHANISMS
VisionZip. VisionZip (Yang et al., 2025b) retains high-attention tokens and merges the remainder around uniformly sampled contextual tokens. Compression budgets K = 32, 64, 128, 192 use (dominant, contextual) counts (27, 5), (54, 10), (108, 20), and (162, 30), respectively. VisPruner. VisPruner (Zhang et al., 2025b) combines attention-based importance with feature diversity. Half of each budget is assigned to important tokens and the remainder to diverse tokens. PruMerge. PruMerge (Shang et al., 2025) selects attention-ranked representatives and merges nearby tokens by feature similarity. We use K −1 representatives together with one attention-weighted residual aggregate, yielding exactly K output tokens. FastV. FastV (Chen et al., 2024) ranks visual tokens using language-model attention and removes low-ranked tokens after layer 2. These configurations are used for the primary LLaVA evaluation. For Qwen3-VL and InternVL, the corresponding compression rules are instantiated on their native visual-token interfaces while preserving each method’s selection and aggregation principle. Throughout the evaluation, compression budget K denotes the number of post-compression visual tokens passed to subsequent computation, whether obtained through selection, merging, or both. B.4
I MPLEMENTATION
Inference. All experiments run on one NVIDIA GeForce RTX 4090 GPU. Decoding is deterministic (do_sample=False) with at most 64 new tokens, using each model’s native image processor and conversation template. Priority-score instantiation. CIRA derives token priorities from model-native late-layer visual attention. Let Lscore denote the encoder layers used to compute token priority scores. We use the final, third-to-last, and fifth-to-last encoder blocks, corresponding to relative layer indices (−1, −3, −5). For the CLIP encoder in the primary setting, the score from Section 4.2 is si (x) =
1 |Lscore |
H X X
ℓ,h α0,i (x),
ℓ∈Lscore h=1
16
i = 1, . . . , N,
(13)
Preprint. Under review.
ℓ,h where α0,i is attention from the class token to patch i at head h and visual layer ℓ, and H is the number of attention heads.
InternVL derives token priorities from class-to-patch visual self-attention and aggregates them over its spatial-downsampling groups. For Qwen3-VL, which lacks a class-token routing interface, we use the mean incoming attention over visual queries and aggregate scores within its native spatial-merger groups. For both models, priorities are averaged over the same three relative encoder layers to produce one score per downstream visual token. Preservation features. For LLaVA, HEP measures tokenwise cosine distance between final CLIP patch features after post-layer normalization. For Qwen3-VL, it averages the tokenwise distances of the merged main visual features and the DeepStack features. For InternVL, it uses the visual tokens returned by the model’s feature-extraction module after pixel shuffle and MLP projection. Optimization. All downstream-agnostic attacks use an ℓ∞ perturbation budget of 4/255 and 100 optimization steps. CIRA starts from the clean image and uses projected sign-gradient ascent with step size 1/255, λ = 0.8, and candidate budget range [Kmin , Kmax ] = [32, 192]. The remaining hyperparameters of VEAttack (Mei et al., 2026b) and CAGE (Zhang et al., 2026c) follow their released settings; CAGE uses its released budget interval [16, 192]. CAA† (Zhang et al., 2026a) optimizes a question-conditioned objective at language-model layer 2 with ϵ = 4/255 and 100 steps. One adversarial image is generated per image–question pair and then evaluated across compressors and budgets without re-optimization. Evaluation protocol. Each downstream-agnostic method generates one adversarial image per clean input. After optimization, the image is fixed and evaluated across all corresponding questions, compressors, and budgets. Clean eligibility is determined separately for each compressor–budget setting using SK in equation 1, while Full ASR is computed over Sfull defined in Section 5.1.
C
C ONTROLLED R ETAINED -S ET A LLOCATION D IAGNOSTICS
The diagnostic in Section 3 isolates retained-set allocation through counterfactual exchanges at fixed first-observed CSF states. Answer-aware scores are used only for post-hoc exchange selection after the failure state is fixed and are not part of CIRA optimization. C.1
D IAGNOSTIC C OHORT AND F IXED FAILURE S TATE
Diagnostic setting. We use LLaVA–VisionZip on POPE, TextVQA, and MME at K ∈ {32, 64, 128, 192}. VisionZip retains dominant patch tokens individually and aggregates the remainder into contextual tokens. We call the individually retained patches direct tokens; their capacity DK is smaller than the total compressed budget K. Common clean cohort. Let j index an image–question observation, comprising image xj , its associated question, and reference answer; index i is reserved for visual tokens. We retain only observations that are correct under full-token inference and under every evaluated compressed setting: cj (xj ) = 1, cj,K (xj ) = 1, ∀K ∈ {192, 128, 64, 32}, (14) where cj and cj,K denote full-token and compressed correctness for observation j. Every trajectory therefore begins from a common state without pre-existing compression errors in any evaluated path. Attack trajectory. We generate clean-initialized feature-objective VEAttack (Mei et al., 2026b) trajectories with ϵ = 4/255, step size 1/255, and at most 100 steps, evaluating predictions every 10 steps. For each observation and setting, we select the earliest evaluated checkpoint satisfying cj (xadv cj,K (xadv j ) = 1, j ) = 0. This first-observed CSF supplies the fixed state for all subsequent counterfactuals.
(15)
Balanced diagnostic cohort. Within each comparison, eligible images are ordered using a fixed sampling order and truncated to the smallest available image count across cells. All eligible questions associated with the retained images are then included. 17
Preprint. Under review.
C.2
C ONTROLLED C OUNTERFACTUAL R EALLOCATION
At the fixed adversarial state, exchange identities are selected using retrospective answer support computed from the adversarial representations. For adversarial token i, we apply a multiplicative gate e hi = gi hadv and define i ai = −
e adv , q) ∂LNLL (y | H ∂gi
.
(16)
gi =1
Larger ai indicates that token i provides stronger support for the reference answer. Let RK be the direct-token set and OK its complement. We rank OK in descending and RK in ascending order of ai . For exchange size r, guide RK,r = (RK \ Lr ) ∪ Hr ,
(17)
where Hr contains the r highest-support omitted tokens and Lr the r lowest-support retained tokens. We evaluate the nominal schedule ρ ∈ {0, 5%, 7.5%, 10%, 15%}, with rK (ρ) = round(ρDK ),
(18)
where DK is the number of direct slots. For ρ = (0, 5%, 7.5%, 10%, 15%), this gives r32 = (0, 1, 2, 3, 4), r64 = (0, 3, 4, 5, 8), r128 = (0, 5, 8, 11, 16), and r192 = (0, 8, 12, 16, 24). The matched-random control independently permutes the retained and omitted pools. Each observation uses 40 nested paths, with larger exchanges extending smaller ones. VisionZip merging and aggregation are recomputed after every exchange. C.3
R ETAINED -S ET A LLOCATION S ENSITIVITY AND C UMULATIVE R ECOVERY
Let YjKz,p (ρ) ∈ {0, 1} indicate correctness for observation j under condition z, path p, and ratio ρ. Guided reallocation has one path; the random control has P = 40 nested paths. Because recovery need not persist under a larger exchange, curves report cumulative recovery: CjKz,p (ρ) = max YjKz,p (ρ′ ). ′ ρ ≤ρ
(19)
Thus, CjKz,p (ρ) asks whether a CSF is corrected at any schedule point up to the displayed nominal ratio. Estimand and aggregation. aggregation:
Random paths are averaged within each observation before cell P
C jK,rand (ρ) =
1 X CjK,rand,p (ρ), P p=1
(20)
and we set C jK,guide (ρ) = CjK,guide,1 (ρ). Let G denote the 12 dataset–budget cells and JdK their analyzed image–question observations. The reported equal-cell average is X 1 X 1 µ bz (ρ) = C jK,z (ρ) . (21) |G| |JdK | j∈JdK
(d,K)∈G
The inner mean averages over image–question observations within each cell, while the outer mean b assigns equal weight to every dataset–budget cell. The paired contrast is ∆(ρ) = µ bguide (ρ) − µ brand (ρ). The average guided–random advantage ranges from 22.5 to 24.1 pp across all nonzero exchange ratios. 18
Preprint. Under review.
Table 5: Cumulative guided (G) and matched-random (R) recovery (%) by compression budget and exchange ratio, with ∆ = G − R reported in percentage points and Avg. denoting the equal-cell average. K 32 64 128 192 Avg.
C.4
G
ρ = 5% R ∆
29.1 39.1 35.8 44.4 37.1
12.3 16.5 13.8 15.6 14.5
+16.8 +22.6 +22.0 +28.8 +22.6
G
ρ = 7.5% R ∆
40.6 43.4 45.8 50.8 45.2
19.3 21.8 20.9 22.3 21.1
+21.4 +21.6 +24.9 +28.4 +24.1
G
ρ = 10% R ∆
45.5 44.1 49.4 53.6 48.2
23.6 25.9 25.4 27.0 25.5
+21.9 +18.3 +24.0 +26.6 +22.7
G
ρ = 15% R ∆
51.1 49.7 54.9 57.9 53.4
27.5 32.5 30.8 33.0 31.0
+23.6 +17.2 +24.1 +24.9 +22.5
FACTORIAL D ECOMPOSITION OF R ECOVERY
At ρ∗ = 7.5%, we separate which omitted tokens enter from which retained tokens leave. Incoming tokens are high-support evidence or random omissions; outgoing tokens are low-support retained tokens or random ones: Random in Evidence in
Random out RR ER
Low-support out RL EL
Let QjK,z ∈ [0, 1] denote exact recovery under arm z, averaged over random paths where applicable. We compute the two main effects and their interaction per observation before equal-cell aggregation: (QjK,ER − QjK,RR ) + (QjK,EL − QjK,RL ) , 2 (QjK,RL − QjK,RR ) + (QjK,EL − QjK,ER ) , ujK = 2 ηjK = QjK,EL − QjK,ER − QjK,RL + QjK,RR . ejK =
(22) (23) (24)
The four arm rates and observation-level effects are aggregated with equation 21; Table 6 reports the resulting point estimates by budget. Table 6: Exact recovery and factorial effects at ρ∗ = 7.5% by compression budget, with Avg. denoting the equal-cell average. K
Exact Recovery (%) RR ER RL EL
32 64 128 192 Avg.
16.1 18.4 17.2 17.9 17.4
34.9 34.0 33.5 31.1 33.4
20.1 25.4 31.9 34.5 28.0
38.6 38.5 44.4 46.4 42.0
Factorial Effect (pp) Evidence Restoration Low-Support Removal Interaction +18.6 +14.4 +14.4 +12.6 +15.0
+3.8 +5.7 +12.8 +15.9 +9.6
−0.3 −2.4 −3.9 −1.3 −2.0
Evidence restoration is positive at every budget, while the removal effect increases with K. Interaction estimates are negative and smaller in magnitude than either main effect at every budget. C.5
R EPRESENTATION -D RIFT M ODERATION
For a restored evidence token i, let hclean and hadv be its clean and adversarial encoder representations. i i We define 1 − cos hclean , hadv diag i i di = . (25) 2 Let EjK denote the high-support omitted tokens restored for observation j at budget K under the ρ∗ = 7.5% intervention. We define their mean representation drift as X diag 1 DjK = di . (26) |EjK | i∈EjK
19
Preprint. Under review.
Clean representations are used only in this post-hoc diagnostic. Let DdK and edK be within-cell means. We estimate the common within-cell association between drift and the evidence-restoration effect ejK from equation 24: P P DjK − DdK ejK − edK d,K j∈J dK βb = . (27) P P 2 d,K j∈JdK (DjK − D dK ) Thus, βb uses only within-cell variation. Drift tertiles provide a grouped summary, and the continuous slope summarizes the corresponding within-cell association. Table 7: Evidence-restoration effects across within-cell representation-drift tertiles and compression budgets, with mean drift defined as 100× half-cosine distance and Gap as the lower-minus-upper effect. K 32 64 128 192 Avg.
Lower Middle Upper Gap (pp) Drift (%) Effect (pp) Drift (%) Effect (pp) Drift (%) Effect (pp) 26.1 31.0 34.0 34.4 31.4
+28.5 +18.2 +18.9 +16.6 +20.5
38.0 39.3 40.4 40.2 39.5
+20.8 +7.0 +12.7 +15.3 +14.0
45.2 45.5 45.1 45.1 45.2
+6.1 +18.0 +11.5 +5.4 +10.3
+22.3 +0.2 +7.4 +11.2 +10.3
The estimated within-cell association is −8.1 pp per 0.1 increase in DjK . Budget-specific tertiles are not uniformly monotone and are therefore interpreted descriptively. The retained-set intervention establishes that changing token allocation can causally alter correctness within this fixed cohort, whereas the drift–recovery analysis supports an association rather than causal mediation.
D
S ENSITIVITY A NALYSES
We test whether CIRA’s selective operating point depends on individual design or optimization choices. Unless stated otherwise, each analysis varies one choice on LLaVA-v1.5-7B, POPE, and VisionZip while retaining the evaluation definitions in Section 5.1. D.1
P RESERVATION W EIGHT
The relative loss weight λ in equation 9 controls the tradeoff between failure induction and full-token preservation. We summarize this operating tradeoff by Selective Gap (Avg. CSFR minus Full ASR), used only as a configuration score because the two metrics have different conditioning sets. Table 8 shows that increasing λ reduces Full ASR while retaining substantial CSFR. Selective Gap remains stable for λ ∈ [0.6, 1.0], with λ = 0.8 attaining the highest observed value. We therefore use λ = 0.8 throughout the evaluation. Table 8: Sensitivity of Full ASR, CSFR, and Selective Gap to the preservation weight λ on POPE. Best values in each row are shown in bold. Metric Full ASR↓ CSFR@192↑ CSFR@128↑ CSFR@64↑ CSFR@32↑ Avg. CSFR↑ Selective Gap↑
D.2
λ = 0.2 9.57 8.59 15.67 22.07 27.56 18.47 8.90
λ = 0.4 6.03 9.08 15.30 22.07 27.26 18.43 12.40
λ = 0.6 5.79 9.20 15.55 22.47 25.93 18.29 12.50
λ = 0.8 4.96 7.61 14.05 22.47 26.96 17.77 12.81
λ = 1.0 4.61 6.50 13.68 21.94 27.41 17.38 12.77
C ANDIDATE C OMPRESSION -B UDGET I NTERVAL
The admissible interval [Kmin , Kmax ] determines the budget-marginal displacement weights in equation 6. Table 9 varies this interval while keeping the evaluation budgets fixed at K ∈ 20
Preprint. Under review.
{192, 128, 64, 32}. Selectivity varies modestly across the six candidate intervals. We use [32, 192] as the default because it matches the evaluation range; its Selective Gap is within 0.28 pp of the best observed value. Table 9: Sensitivity of Full ASR, CSFR, and Selective Gap to the candidate compression-budget interval on POPE. Best values in each row are shown in bold. Metric Full ASR↓ CSFR@192↑ CSFR@128↑ CSFR@64↑ CSFR@32↑ Avg. CSFR↑ Selective Gap↑
D.3
Vary Kmin [16, 192] [64, 192]
Default [32, 192]
4.62 7.38 14.86 21.14 27.43 17.70 13.09
4.96 7.61 14.05 22.47 26.96 17.77 12.81
5.56 6.40 12.86 21.68 28.17 17.28 11.71
5.68 8.86 13.23 23.01 29.20 18.57 12.89
4.97 7.63 14.73 19.95 26.55 17.21 12.24
5.33 6.52 12.86 22.61 25.22 16.80 11.48
P ERTURBATION B UDGET AND O PTIMIZATION S TEPS K = 192
Full
(a) 50 Attack Steps Attack Success Rate (%)
Vary Kmax [32, 128] [32, 384]
[32, 64]
29.7
30
25.5
25.3
25.1 22.9
21.6
20
17.0
15.7 9.4
9.1 6.0 4.3
2/255
6.9 4.7
4/255
5.2
6/255
Perturbation Budget ε
23.3
16.2
15.2 12.0
10
24.6
22.6
(d) 150 Attack Steps
28.3
28.0
28.7
27.9
6.06.7
8/255
7.8 4.1
2/255
9.8 7.7 5.2
4/255
4.3
6/255
Perturbation Budget ε
14.6
21.5
7.3 4.7
8/255
2/255
7.9 5.0
4/255
24.5
21.0
19.0
18.8
16.7 14.2
11.6 9.4 5.8
27.1
26.1 24.0
22.9
19.2
17.7
16.0
(c) 100 Attack Steps 29.2
28.9
28.7
28.5
K = 32
K = 64
(b) 80 Attack Steps 30.5
29.7
29.1 27.0
27.6
K = 128
7.7 5.4
6/255
Perturbation Budget ε
7.5 5.9
8/255
13.0 5.76.6
2/255
11.7 6.5 4.6
4/255
12.1 7.3 5.2
6/255
Perturbation Budget ε
11.5 8.5 6.4
8/255
Figure 5: Sensitivity of Full ASR and CSFR to the perturbation budget and number of optimization steps on POPE. Figure 5 varies ϵ ∈ {2, 4, 6, 8}/255 and the number of optimization steps in {50, 80, 100, 150}; the step size is set to ϵ/4. Across all 16 configurations, Full ASR remains between 4.14% and 6.38%, whereas Avg. CSFR ranges from 14.93% to 20.82%. Thus, substantial compression-specific failure rates persist while full-token degradation remains limited, and the default ϵ = 4/255, 100-step setting lies within a stable selective region rather than at an isolated optimum. D.4
S CORING -L AYER C ONFIGURATION S TUDY
The primary score in equation 13 averages class-to-patch attention over the scoring-layer configuration Lscore . We compare 14 single-layer, contiguous, and spaced multi-layer configurations using 100example design subsets from POPE, TextVQA, and MME, evaluated with VisionZip, PruMerge, and VisPruner. We select the configuration with the highest mean Selective Gap across these nine dataset–compressor environments and use it throughout the reported evaluation. As shown in Figure 6, the selected spaced late-layer triplet {−1, −3, −5} achieves 23.45% Avg. CSFR with 5.59% Full ASR, corresponding to a 17.86 pp Selective Gap. Its exclusion rate for clean high-priority tokens also varies less across observer layers than that of the single-layer {−2} configuration (range 0.055 versus 0.242). The negative association between cross-layer variation and selectivity (ρs = −0.70) indicates that configurations with more stable exclusion behavior across depth tend to exhibit higher selectivity. This association is descriptive rather than causal.
E
C ROSS -M ODEL S COPE
The main experiments establish CIRA across multiple compressors and budgets on LLaVA. We next test whether paired selectivity persists when the vision encoder, multimodal interface, and language 21
Preprint. Under review.
Figure 6: Selectivity and cross-layer exclusion stability across scoring-layer configurations. model change together, using Qwen3-VL-8B-Instruct (Bai et al., 2025) and InternVL3.5-8B (Wang et al., 2025b). Here K denotes the number of compressed visual tokens passed to subsequent computation in each model’s native interface. For compact reporting, we use K1 /64/32, where K1 = 96 for Qwen3-VL and K1 = 128 for InternVL. Avg. CSFR is the arithmetic mean over the three budgets within each model family. Table 10 reports the corresponding CSFR and Full ASR results, while clean accuracy is provided in Table 11. Table 10: CSFR and Full ASR on Qwen3-VL-8B-Instruct and InternVL3.5-8B. Compressor entries report CSFR at K1 /64/32 followed by their average. Bold and underlined values denote the best and second-best downstream-agnostic results, respectively; CAA† is shown as a stronger-access reference and excluded from these rankings. Dataset
POPE
TextVQA
MME
POPE
TextVQA
MME
Attack
Full ASR (%)↓
VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA
45.93 46.47 4.40 10.66 80.14 66.03 8.24 16.37 44.11 41.60 4.76 12.43
VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA VEAttack CAGE CAA† CIRA
51.46 45.01 9.85 11.68 56.92 55.24 5.03 12.03 40.63 37.25 5.53 5.76
VisionZip K1 /64/32/Avg.↑
VisPruner K1 /64/32/Avg.↑
Qwen3-VL-8B-Instruct (K1 = 96) 4.63 / 5.83 / 10.06 / 6.84 2.15 / 2.78 / 4.67 / 3.20 7.13 / 8.13 / 10.19 / 8.48 4.65 / 5.68 / 8.32 / 6.22 4.03 / 5.82 / 11.31 / 7.05 2.61 / 4.86 / 8.70 / 5.39 7.59 / 9.82 / 14.14 / 10.52 4.98 / 6.56 / 11.22 / 7.59 10.11 / 10.03 / 3.58 / 7.91 10.14 / 7.75 / 1.98 / 6.62 23.96 / 22.56 / 15.22 / 20.58 19.59 / 16.71 / 10.62 / 15.64 46.37 / 42.86 / 35.22 / 41.48 29.28 / 31.23 / 32.84 / 31.12 51.65 / 52.13 / 45.37 / 49.72 30.63 / 28.81 / 34.57 / 31.34 7.14 / 6.59 / 9.29 / 7.67 5.95 / 6.51 / 6.29 / 6.25 11.67 / 7.78 / 10.19 / 9.88 11.41 / 7.83 / 9.37 / 9.54 8.27 / 9.96 / 12.35 / 10.19 7.37 / 7.72 / 10.34 / 8.48 8.34 / 7.79 / 12.79 / 9.64 8.82 / 7.37 / 9.91 / 8.70 InternVL3.5-8B (K1 = 128) 6.01 / 8.17 / 10.35 / 8.18 4.20 / 7.44 / 7.91 / 6.52 7.67 / 13.25 / 17.45 / 12.79 3.82 / 5.61 / 8.19 / 5.87 4.35 / 7.50 / 8.78 / 6.88 1.90 / 4.14 / 6.66 / 4.23 4.87 / 7.90 / 11.65 / 8.14 3.04 / 4.01 / 8.42 / 5.16 7.50 / 12.77 / 14.81 / 11.69 3.83 / 7.17 / 11.67 / 7.56 11.00 / 16.23 / 16.24 / 14.49 6.20 / 10.99 / 11.41 / 9.53 10.58 / 22.73 / 23.58 / 18.96 13.64 / 15.66 / 23.68 / 17.66 33.17 / 36.80 / 36.93 / 35.63 10.55 / 17.00 / 21.58 / 16.38 2.57 / 4.44 / 4.91 / 3.98 2.27 / 4.05 / 4.44 / 3.58 7.13 / 9.14 / 9.16 / 8.48 4.07 / 4.93 / 6.32 / 5.10 2.92 / 3.58 / 8.11 / 4.87 2.41 / 6.19 / 8.31 / 5.64 6.88 / 10.63 / 12.50 / 10.00 3.98 / 7.45 / 8.72 / 6.72
PruMerge K1 /64/32/Avg.↑
FastV K1 /64/32/Avg.↑
3.91 / 5.01 / 7.49 / 5.47 2.32 / 3.95 / 4.60 / 3.62 7.23 / 8.23 / 8.95 / 8.14 4.65 / 6.25 / 11.49 / 7.46 2.47 / 3.33 / 7.06 / 4.29 8.44 / 11.61 / 14.80 / 11.61 3.65 / 4.39 / 7.94 / 5.33 8.56 / 10.08 / 13.07 / 10.57 10.22 / 7.87 / 3.34 / 7.14 3.54 / 3.58 / 3.65 / 3.59 21.54 / 18.12 / 16.56 / 18.74 12.38 / 11.00 / 12.41 / 11.93 31.69 / 30.19 / 31.91 / 31.27 43.81 / 42.20 / 38.69 / 41.57 34.88 / 32.13 / 35.20 / 34.07 38.31 / 36.57 / 34.31 / 36.40 7.60 / 8.31 / 9.69 / 8.54 2.58 / 3.84 / 6.58 / 4.33 12.47 / 11.87 / 15.09 / 13.14 4.42 / 6.64 / 8.53 / 6.53 9.71 / 9.72 / 12.08 / 10.51 9.16 / 11.71 / 15.66 / 12.18 7.14 / 5.26 / 7.05 / 6.49 5.28 / 6.77 / 10.21 / 7.42 5.28 / 7.56 / 9.03 / 7.29 7.13 / 11.29 / 12.99 / 10.47 3.48 / 5.33 / 8.49 / 5.77 9.68 / 13.31 / 16.35 / 13.11 0.77 / 2.21 / 2.65 / 1.88 7.26 / 20.03 / 28.32 / 18.54 2.83 / 3.26 / 4.75 / 3.61 3.82 / 8.60 / 12.26 / 8.23 7.52 / 10.50 / 11.86 / 9.96 5.24 / 7.37 / 12.68 / 8.43 6.64 / 10.37 / 13.50 / 10.17 4.94 / 9.89 / 14.83 / 9.89 5.25 / 9.51 / 10.53 / 8.43 10.18 / 33.63 / 45.69 / 29.84 17.35 / 21.20 / 15.79 / 18.11 10.03 / 25.36 / 38.28 / 24.56 3.28 / 3.93 / 5.71 / 4.31 1.72 / 3.27 / 5.35 / 3.45 4.38 / 5.73 / 7.82 / 5.98 4.70 / 7.02 / 8.43 / 6.72 2.81 / 2.81 / 5.04 / 3.55 3.44 / 13.68 / 18.88 / 12.00 5.37 / 5.61 / 7.17 / 6.05 4.01 / 6.54 / 11.38 / 7.31
Across both additional model families, CIRA continues to induce compression-specific failures while keeping Full ASR substantially below those of VEAttack and CAGE. The pattern is strongest on TextVQA, where CIRA achieves the highest Avg. CSFR among downstream-agnostic attacks for every compressor on both model families. Results on POPE and MME are more heterogeneous across compressors, but compression-specific failure induction remains observable under target-encoderonly access. Together, these results show that CIRA’s compression-selective behavior extends beyond LLaVA to distinct vision encoders and native visual-token interfaces. 22
Preprint. Under review.
F
S ELECTION S TABILIZATION D EFENSE
Translation-Consensus Selection (TCS) stabilizes priority rankings by aggregating aligned scores across spatially translated views. F.1
T RANSLATION -C ONSENSUS S ELECTION
For pixel displacement d, define V = {T0,0 , Td,0 , T0,d , Td,d }.
(28)
We set d = 7 pixels and construct each view by reflection-padding and cropping to the original size. The four views form one batched vision-encoder input. Score alignment and consensus. For view v ∈ V, let s(v) ∈ RN be its encoder-side priority-score vector. Operator Av inverse-aligns the score grid to T0,0 by bilinear sampling with reflection padding. With patch size p = 14, the offset is d/p = 0.5 patch: e s(v) = Av s(v) . (29) Equation 11 converts the aligned scores to descending rank quantiles, so each view contributes a priority ranking rather than a score scale. Selection interface. At compression budget K, TCS replaces the original priority ranking with the cross-view consensus ranking. The unshifted view supplies token features and the compressor’s keysimilarity metric, while translated views contribute aligned priority scores. Token counts, aggregation, and language-model input length remain unchanged. Matched evaluation and cost. We evaluate None and TCS on matched adversarial images, questions, references, eligibility sets, and compression budgets. Full-token inference is unchanged, so Full ASR is shared within each matched pair in Table 4. The four views are processed by the vision encoder in one batch, with no additional language-model inference. F.2
C ROSS -V IEW S UPPORT M ECHANISM
Cross-view support characterizes the contrast between stable clean evidence and view-specific adversarial replacements. At each compression budget K, the canonical clean Top-K set serves as the reference, while CIRA replacements are tokens that enter the canonical adversarial Top-K set from outside this reference. A candidate’s view support is the number of aligned views in which it remains within the Top-K set.
Figure 7: Cross-view support distributions of clean Top-K tokens and CIRA replacement tokens across compression budgets. 23
Preprint. Under review.
Figure 7 shows that CIRA replacements are predominantly view-specific. As K increases from 32 to 192, the one-view share decreases from 88.0% to 63.9%, while fewer than 2.5% are supported by all four views. Clean Top-K tokens show the opposite pattern: their four-view share increases from 20.0% to 44.2%. Averaging aligned rank quantiles therefore downweights isolated replacement spikes while favoring evidence supported across translations. F.3
A DAPTIVE E VALUATION
Standard CIRA is optimized on the unshifted view, with TCS applied only at evaluation. Adaptive CIRA instead optimizes equation 12 over all four public transformations, using the view-specific encoder objectives as differentiable surrogates for rank conversion and Top-K selection. For each view, we cache clean features and ranks and compute GSH and HEP with view-specific hiddenevidence weights over [Kmin , Kmax ]. The averaged gradient updates one shared perturbation using the original ϵ = 4/255, step size 1/255, 100 steps, and λ = 0.8. Adaptive CIRA retains the same target-encoder access as standard CIRA. Selection and rank response. For each condition, clean Top-K retention measures the fraction of tokens in the clean Top-K set that remain in the adversarial Top-K set under the corresponding ranking rule. Signed normalized rank change is (ria − ric )/(N − 1), where ric and ria denote clean and adversarial ranks under the same ranking rule, and positive values indicate demotion. Retention is averaged per image.
Figure 8: Clean Top-K retention and priority-reallocation profiles under CIRA, CIRA + TCS, and Adaptive CIRA + TCS. Across the four compression budgets, Figure 8(a) shows that TCS raises clean Top-K retention under CIRA from 0.3–12.3% to 62.9–70.8%. Adaptive CIRA reduces this retention under TCS to 6.5–34.7%. Figure 8(b) shows the corresponding priority reallocation: TCS attenuates both the demotion of clean high-priority tokens and the promotion of initially low-priority tokens, whereas Adaptive CIRA restores much of this signed reallocation. Together, the cross-view support patterns and the selection responses under TCS show that TCS suppresses view-fragile priority reallocation. Corresponding task-utility results are reported in Appendix G.
G
C OMPLEMENTARY TASK -U TILITY R ESULTS
CSFR is the primary clean-conditioned metric for compression-specific failure. We complement it with accuracy-based results that characterize clean utility under compression, post-attack performance, and task-utility recovery under TCS. 24
Preprint. Under review.
Table 11: Clean full-token and compressed accuracy (%) across model families, datasets, compressors, and family-specific compression budgets. The final value in each compressed cell reports the average across the listed budgets. Dataset
Full ACC
POPE TextVQA MME
84.6 60.3 79.2
POPE TextVQA MME
86.3 88.6 90.0
POPE TextVQA MME
82.2 71.5 88.6
G.1
VisionZip
VisPruner
PruMerge
LLaVA-v1.5-7B (K = 192/128/64/32/Avg.) 84.3/83.5/79.6/73.0/80.1 84.1/82.9/80.5/75.3/80.7 75.6/74.1/72.0/70.1/73.0 54.3/51.6/49.5/45.1/50.1 58.4/58.0/56.8/51.9/56.3 51.9/51.6/51.7/49.4/51.2 77.2/76.3/73.5/69.0/74.0 76.9/76.5/74.0/71.7/74.8 72.9/70.9/71.1/70.8/71.4 Qwen3-VL-8B-Instruct (K = 96/64/32/Avg.) 86.1/84.1/80.7/83.6 86.6/85.1/82.3/84.7 86.4/86.1/83.8/85.4 46.7/41.1/34.1/40.6 45.8/42.6/41.5/43.3 33.9/31.8/31.2/32.3 88.2/88.6/83.4/86.7 88.4/88.1/83.5/86.7 86.7/86.7/83.3/85.6 InternVL3.5-8B (K = 128/64/32/Avg.) 81.4/80.8/77.1/79.8 80.4/80.5/78.3/79.7 79.0/79.0/78.8/78.9 64.7/48.6/37.5/50.3 57.7/47.0/40.4/48.4 46.8/37.7/32.8/39.1 87.5/83.8/78.2/83.2 84.8/82.2/77.0/81.3 84.0/80.8/78.5/81.1
FastV 81.1/80.0/74.7/68.0/76.0 51.6/49.4/44.7/37.8/45.9 76.1/73.9/70.8/67.4/72.1 85.1/82.4/74.4/80.6 52.6/40.6/28.1/40.4 88.3/85.7/79.4/84.5 81.0/79.5/74.9/78.5 68.5/57.6/44.3/56.8 88.2/85.2/77.8/83.7
C LEAN U TILITY UNDER C OMPRESSION
Table 11 shows that high-budget settings preserve most POPE and MME accuracy across model families, while TextVQA generally exhibits larger losses under compression. Related analyses also find task-dependent visual-token requirements, with OCR tasks relying on visual information deeper into the decoder (Wang et al., 2026b). These values provide the clean reference for the post-attack comparisons below. G.2
P OST-ATTACK TASK U TILITY
Unlike CSFR, adversarial accuracy is computed over the complete evaluation set and therefore reflects both pre-existing compression errors and attack-induced failures. It provides a complementary view of overall task degradation under CIRA. Table 12: Full-token clean-to-adversarial accuracy and compressed adversarial accuracy (%) under CIRA across model families, datasets, compressors, and family-specific budgets. The final value in each compressed cell reports the budget average. Dataset
Full ACC
POPE TextVQA MME
84.6 → 82.6 60.3 → 57.1 79.2 → 77.2
POPE TextVQA MME
86.3 → 81.3 88.6 → 75.7 90.0 → 81.5
POPE TextVQA MME
82.2 → 76.1 71.5 → 65.9 88.6 → 85.3
VisionZip
VisPruner
PruMerge
LLaVA-v1.5-7B (compressed Adv. ACC: K = 192/128/64/32/Avg.) 78.4/74.0/64.6/58.0/68.8 80.0/77.7/71.7/64.1/73.4 63.3/55.9/51.2/50.3/55.2 46.7/40.0/32.2/26.1/36.3 51.3/49.5/44.6/38.6/46.0 36.9/32.2/27.4/25.0/30.4 73.2/70.7/65.5/59.5/67.2 73.2/70.9/64.1/60.5/67.2 62.2/58.9/55.4/51.5/57.0 Qwen3-VL-8B-Instruct (compressed Adv. ACC: K = 96/64/32/Avg.) 75.9/74.2/71.2/73.8 78.1/76.7/73.5/76.1 79.8/79.8/76.0/78.5 21.1/20.5/21.4/21.0 31.9/31.9/31.0/31.6 22.2/22.4/22.0/22.2 74.9/76.1/69.8/73.6 73.9/76.5/72.5/74.3 74.3/77.1/74.6/75.3 InternVL3.5-8B (compressed Adv. ACC: K = 128/64/32/Avg.) 75.5/75.4/71.8/74.2 75.8/75.9/73.4/75.0 75.7/75.0/75.4/75.4 42.3/32.3/27.1/33.9 53.9/43.6/36.1/44.5 39.5/31.8/29.5/33.6 81.4/75.7/71.5/76.2 81.3/77.8/74.9/78.0 80.2/78.3/74.7/77.7
FastV 76.8/71.4/61.0/53.4/65.7 45.5/41.9/33.5/25.7/36.7 73.2/71.6/63.7/58.8/66.8 76.0/73.9/67.7/72.5 30.7/25.6/20.6/25.6 79.9/77.5/70.9/76.1 77.3/73.1/70.2/73.5 58.4/42.7/29.6/43.6 83.8/79.8/72.9/78.8
Table 12 shows that CIRA reduces budget-averaged compressed accuracy by 3.3–20.8 pp across the evaluated model–dataset–compressor settings. After averaging compressors within each model– dataset pair and weighting the nine pairs equally, the mean compressed accuracy drop is 9.6 pp, compared with 5.4 pp under full-token inference. This aggregate view complements the paired CSFR analysis by quantifying end-task degradation under compressed inference. G.3
TASK -U TILITY R ECOVERY UNDER TCS
Table 13 reports matched task accuracy without and with TCS. Each comparison fixes the input image, question, and compression budget, with TCS applied only to compressed inference. Across the three benchmarks, TCS changes clean Avg. ACC by at most 0.52 pp in magnitude, while recovering 6.77–13.20 pp under CIRA, with larger gains at tighter budgets. Under Adaptive 25
Preprint. Under review.
Table 13: Matched compressed accuracy (%) without and with TCS on LLaVA-v1.5-7B with VisionZip under clean, CIRA, and Adaptive CIRA conditions across datasets and compression budgets. Entries show None → TCS, with annotations reporting the corresponding change. Evaluation
K = 192
Clean CIRA Adaptive CIRA
84.30 → 83.10 ↓1.20 78.40 → 81.30 ↑2.90 78.80 → 78.60 ↓0.20
Clean CIRA Adaptive CIRA
54.30 → 53.20 ↓1.10 46.70 → 52.40 ↑5.70 48.30 → 49.00 ↑0.70
Clean CIRA Adaptive CIRA
77.20 → 76.50 ↓0.70 73.20 → 75.00 ↑1.80 74.50 → 73.60 ↓0.90
K = 128
K = 64
POPE 83.50 → 82.50 ↓1.00 79.60 → 78.40 ↓1.20 74.00 → 80.90 ↑6.90 64.60 → 78.30 ↑13.70 76.40 → 76.00 ↓0.40 66.80 → 71.60 ↑4.80 TextVQA 51.60 → 51.30 ↓0.30 49.50 → 49.40 ↓0.10 40.00 → 50.90 ↑10.90 32.20 → 49.00 ↑16.80 45.10 → 46.90 ↑1.80 36.00 → 42.10 ↑6.10 MME 76.30 → 76.00 ↓0.30 73.50 → 74.30 ↑0.80 70.70 → 75.80 ↑5.10 65.50 → 74.00 ↑8.50 70.90 → 72.00 ↑1.10 62.90 → 68.40 ↑5.50
K = 32
Avg.
73.00 → 74.30 ↑1.30 58.00 → 75.40 ↑17.40 56.30 → 62.20 ↑5.90
80.10 → 79.58 ↓0.52 68.75 → 78.98 ↑10.23 69.58 → 72.10 ↑2.52
45.10 → 45.40 ↑0.30 26.10 → 45.50 ↑19.40 26.70 → 33.00 ↑6.30
50.13 → 49.83 ↓0.30 36.25 → 49.45 ↑13.20 39.03 → 42.75 ↑3.72
69.00 → 70.60 ↑1.60 59.50 → 71.20 ↑11.70 58.50 → 62.00 ↑3.50
74.00 → 74.35 ↑0.35 67.23 → 74.00 ↑6.77 66.70 → 69.00 ↑2.30
CIRA, TCS recovers less utility, consistent with the attack partially restoring the priority reallocation suppressed by TCS.
H
L IMITATIONS AND F UTURE W ORK
Temporal and contextual dependencies. The present evaluation is restricted to single-image inference. In video, multi-image, and multi-turn settings, evidence retention also depends on temporal redundancy and evolving context (Shen et al., 2024; Li et al., 2026b; Wang et al., 2026c). Run-Length Pruning, for instance, combines temporal redundancy removal with token distillation (Ma et al., 2026). Whether encoder-only priority manipulation remains compression-selective when evidence is distributed across frames or dialogue turns remains an open question. Compression mechanisms beyond token selection. Our diagnostics examine retained-set allocation and representation drift, but do not exhaust the mechanisms underlying compression-induced errors. Adaptive pruning changes allocation across inputs or layers (Ye et al., 2025; Chen et al., 2026; Li et al., 2026a); learned summarization transforms the token representation (Bulat et al., 2026); and recoverable routing permits deferred tokens to re-enter subsequent selection stages (Yang et al., 2026). These mechanisms complicate a description based on a single token ranking. Moreover, spatial disruption (Huang et al., 2026) and positional or attentional distortion (Cho et al., 2026) are not separately identified by our retained-set interventions. Paired CSF evaluation remains applicable, but attributing failures within these compression mechanisms requires additional diagnostics. Beyond task correctness. CSF is defined through task-level correctness rather than response safety or targeted attack success. Accordingly, preserving full-token correctness does not establish safety preservation, and compressed-path errors do not necessarily constitute safety-alignment failures. Targeted manipulation and multimodal jailbreaks (Zhang et al., 2025a; Qi et al., 2024; Shayegani et al., 2024) provide distinct settings for paired evaluation, requiring outcome criteria tailored to the corresponding security objective.
26
Preprint. Under review.
I
Q UALITATIVE C ASE S TUDIES
Figure 9: Qualitative examples of compression-specific failures on LLaVA-v1.5-7B across visual-token compressors and retention budgets.
27
Preprint. Under review.
Figure 10: Qualitative examples of compression-specific failures on Qwen3-VL-8B-Instruct across visual-token compressors and retention budgets.
28
Preprint. Under review.
Figure 11: Qualitative examples of compression-specific failures on InternVL3.5-8B across visual-token compressors and retention budgets.
29