AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization Zonghao Ying1 * , Haozheng Wang1 * , Jiangfan Liu1 , Quanchen Zou2 , Aishan Liu1 , Jian Yang1 , Yaodong Yang3 , Xianglong Liu1 1
Beihang University
2
360 AI Security Lab
3
Peking University
arXiv:2604.24118v1 [cs.CR] 27 Apr 2026
Abstract Large Language Model (LLM) agents are increasingly used to automate complex workflows, but integrating untrusted external data with privileged execution exposes them to severe security risks, particularly direct and indirect prompt injection. Existing defenses face significant challenges in balancing security with utility, often encountering a trade-off where rigorous protection leads to over-defense, or where subtle indirect injections bypass detection. Drawing inspiration from operating system virtualization, we propose AgentVisor, a novel defense framework that enforces semantic privilege separation. AgentVisor treats the target agent as an untrusted guest and intercepts tool calls via a trusted semantic visor. Central to our approach is a rigorous audit protocol grounded in classic OS security primitives, designed to systematically mitigate both direct and indirect injection attacks. Furthermore, we introduce a one-shot self-correction mechanism that transforms security violations into constructive feedback, enabling agents to recover from attacks. Extensive experiments show that AgentVisor reduces the attack success rate to 0.65%, achieving this strong defense while incurring only a 1.45% average decrease in utility relative to the No Defense scenario, demonstrating superior performance compared to existing defense methods.
1
Figure 1: Systematic mapping between OS Virtualization and AgentVisor. We translate classic OS security concepts into the semantic space of LLM agents.
tion. In direct injection (Liu et al., 2024b), malicious users explicitly override system instructions to perform unauthorized actions, whereas indirect injection (Yi et al., 2025) involves hidden commands embedded in external content that the agent processes. Current LLM agents often fail to reliably distinguish privileged system instructions from untrusted external commands, which can lead to the unintended execution of malicious payloads. Despite growing awareness of these threats, existing defenses remain fragmented and brittle. Prompt-based hardening (Schulhoff, 2024b,a; Yi et al., 2025) is heuristic and can be overridden by adversarial instructions. Input/output filtering and LLM-based guardrails (ProtectAI.com, 2024; Liu et al., 2025) are prone to evasion, introduce false positives, and provide limited guarantees on how untrusted content affects tool-use decisions. Toolsandboxing approaches (Meng et al., 2025; Piao et al., 2025) restrict available actions, but are often coarse, fail to track information flow in multi-step workflows, and typically lack a principled recovery path once a violation occurs. Consequently, agents still lack a principled security architecture that separates trusted control from untrusted model behavior while enforcing least privilege and informationflow constraints with minimal utility loss. To address these challenges, we draw inspiration
Introduction
Large Language Model (LLM) agents have transitioned from passive conversational systems to active entities capable of automating complex workflows through tool use (Wölflein et al., 2025; Kong et al., 2024; Yuan et al., 2025). However, integrating untrusted external data with privileged execution capabilities exposes agents to severe security risks, primarily Direct and Indirect Prompt Injec* Equal contribution.
1
from operating system virtualization, where a Hypervisor (Popek and Goldberg, 1974) isolates an untrusted Guest from privileged hardware resources in secure OS architectures (Shinagawa et al., 2009; Li et al., 2019). We adapt three key mechanisms for agent security: ❶ Privilege Separation (Saltzer and Schroeder, 1975), which traps and audits sensitive operations; ❷ Policy Enforcement (Saltzer and Schroeder, 1975; Bell and La Padula, 1976), grounded in Least Privilege and Information Flow Control; and ❸ Exception Injection (Popek and Goldberg, 1974), which reports violations via interrupts instead of terminating execution. Translating this paradigm into the semantic space of LLMs, we propose AgentVisor, a semantic virtualization framework that treats the target agent as an untrusted Guest and mediates every tool call via a trusted Visor. The Visor enforces a rigorous STI (Suitability, Taint, Integrity) protocol: Suitability applies Least Privilege to mitigate direct injection, Taint enforces information-flow constraints to block indirect injection, and Integrity preserves parameter/data integrity across both. Upon detecting violations, AgentVisor injects a semantic exception to trigger self-correction, maintaining high utility. Our contributions are summarized as follows:
2025; Liu et al., 2024a) utilize gradient approximation (Zou et al., 2023) to generate adversarial tokens. However, these methods are computationally expensive and often exhibit poor transferability across injection scenarios. In contrast, nonoptimization-based methods (Willison; Liu et al., 2024b; Debenedetti et al., 2024) rely on manually crafted semantic triggers such as Ignore (Perez and Ribeiro, 2022) or Escape (Breitenbach et al., 2023) to manipulate agent behavior. Crucially, these handcrafted techniques are universally applicable to both direct and indirect injection scenarios. Given their practicality and prevalence in real-world exploits, our work focuses on evaluating defenses against these non-optimization-based attacks. 2.2
Existing defenses against prompt injection can be categorized into detection-based and mitigationbased approaches. Detection-based methods focus on verifying the integrity of the input source to identify potential tampering. These approaches typically employ off-the-shelf LLMs (Shi et al., 2025) or fine-tuned guardrail models (ProtectAI.com, 2024; Llama Team, 2024; Liu et al., 2025) to inspect inputs for malicious content before they are processed by the agent. While effective at flagging threats, these methods often act as binary filters, terminating the interaction upon detection and potentially reducing system utility. Mitigationbased methods aim to ensure that the agent executes the intended target task while suppressing the execution of injected commands. One line of research achieves this through safety-specific finetuning of the LLM (Chen et al., 2025c,a), enhancing its intrinsic resistance to adversarial instructions. Another direction involves security enhancement of the input context (Chen et al., 2025b; Yi et al., 2025), such as reiterating the target task instructions (Schulhoff, 2024b) or applying masking strategies to the agent’s state prior to tool invocation (Zhu et al., 2025). Unlike detection methods, mitigation strategies strive to maintain functionality even in the presence of attacks.
• We propose AgentVisor, a virtualization-based defense that enforces privilege separation and one-shot self-correction, securing LLM agents against prompt injection attack. • We design the STI protocol, adapting OS security primitives into a structured semantic audit mechanism, providing systematic and interpretable defenses. • Experiments show that AgentVisor reduces the attack success rate to 0.65%, with only a 1.45% average utility loss compared to No Defense, outperforming existing defenses.
2
Related Work
2.1
Prompt Injection Attacks
Prompt Injection Defenses
LLM agents are vulnerable to both direct and indirect prompt injection attacks, which differ primarily in the injection source but often share common methods. Attack methods can be broadly categorized into optimization-based and non-optimization-based approaches. Optimizationbased methods (Shi et al., 2024; Wang et al.,
3
Preliminaries and Threat Model
3.1
Preliminaries
We consider an LLM-based tool-using agent A that interacts with a user and an external environment. At each turn t, the agent receives a trusted system instruction Isys , a user query Iu , and (when 2
available) an external context Ct retrieved from potentially untrusted sources (e.g., webpages, emails, or documents). Based on these inputs, the agent may invoke a tool. We represent the agent’s tool invocation at turn t as: Tt = A(Isys , Iu , Ct ), (1)
for tool-using LLM agents. tool enforces a trap– audit–recover control loop: the target agent (Guest) proposes tool calls, while a lightweight supervisory component (Visor) audits each proposal against a structured protocol and, when needed, injects a Semantic Exception to steer one-shot self-correction.
where Tt ∈ {∅} ∪ F × X . Here, Tt = (f, args) denotes a tool call with function name f ∈ F and arguments args ∈ X ; if no tool is invoked, Tt = ∅. For multi-step agent workflows (e.g., those involving indirect prompt injections), we additionally maintain an execution history Ht−1 that records prior tool calls and observations, and the agent’s decision becomes Tt = A(Isys , Iu , Ht−1 , Ct ).
At step t, the Guest (target agent) proposes a tool call
3.2
4.1
Ttraw = (ft , argst ) or Ttraw = ∅,
(2)
given the trusted system instruction Isys , the user query Iu , and (in the indirect setting) potentially untrusted state such as retrieved context and multistep execution traces. AgentVisor mediates this proposal using only trusted inputs and a sanitized task state. Specifically, it receives:
Threat Model
We consider an adversary who aims to manipulate the agent into executing an unauthorized tool call, causing the agent’s behavior to deviate from the intended objective specified by Isys and the benign user intent in Iu . The adversary cannot modify the agent A, the system instruction Isys , or the tool implementations/execution environment, but can manipulate inputs delivered to the agent. We use ⊕ to denote the injection operation that augments an otherwise benign input with an adversarial payload.
Xt ≜ (Isys , Iu , H̃t−1 , Ttraw ),
(3)
where H̃t−1 is a structured history view containing only fields such as tool_name, canonicalized args, and a short return_summary/status, wrapped with strict delimiters and treated as data. The defense outputs either an approval or a structured exception: (dect , Et ) = D(Xt ), dect ∈ {allow, exception}.
Direct Prompt Injection. In direct prompt injection, the adversary acts as the user and submits an adversarial query Iuadv = Iu ⊕ δdir , where δdir is a malicious instruction intended to override constraints in Isys or redirect the agent to an attackerchosen objective. We assume direct injection does not rely on external context, i.e., Ct = ∅. The compromised action is then Tadv = A(Isys , Iuadv , Ct ).
(4) (5)
If dect = allow, the proposed tool call is executed as-is; otherwise, the Visor injects Et to request one-shot self-correction from the Guest: Tt′ = A(Isys , Iu , H̃t−1 , Et ),
(6)
after which Tt′ is executed. In both cases, the executed action Tt is the final tool call for step t: ( Ttraw , if dect = allow, Tt = (7) Tt′ , if dect = exception.
Indirect Prompt Injection. In indirect prompt injection, the user query Iu is benign, but the adversary poisons the externally retrieved context by embedding hidden or overt malicious instructions, Ctadv = Ct ⊕ δind , where δind is placed in untrusted content (e.g., a webpage or an email) that the agent later reads. The attacker aims to hijack the agent’s (single- or multi-step) tool-use decisions such that Tt = A(Isys , Iu , Ht−1 , Ctadv ), potentially yielding an attacker-intended tool call Tadv at some step.
4
Problem Formulation
4.2
Architecture of AgentVisor
AgentVisor is motivated by an empirical observation: modern agents can often recognize prompt injections when explicitly asked (e.g., Fig. 3), yet still act on malicious instructions during tool use. This awareness–action gap suggests that “knowing an input is unsafe” does not reliably translate into safe tool-use behavior. AgentVisor addresses this gap by introducing a virtualization-inspired separation of concerns: the Guest generates actions, while the Visor enforces policy via auditing and recovery.
Methodology
In this section, we present AgentVisor, a defense framework structured as a Semantic Hypervisor 3
Figure 2: Overview of the AgentVisor architecture. Drawing inspiration from OS virtualization, AgentVisor enforces privilege separation between the untrusted target agent (Guest) and the trusted semantic hypervisor (Visor).
102
Percentage
100 98 96
Direct Injection GPT-4o GLM-4.7 99.53100.00
99.6999.98 100
99.67100.00
with strict delimiters, and is treated strictly as data rather than as executable instructions. AgentVisor follows a trap–audit–recover loop: (1) the Guest proposes Ttraw ; (2) the Visor audits Ttraw with the STI protocol (Sec. 4.3); (3) if unsafe, the Visor returns a Semantic Exception (Sec. 4.4); the Guest regenerates once and the revised tool call is executed. The complete procedure is provided in Alg. 1 in App. A.
99.94 96.78
96
95.24
94
92
92
Precision
99.89
98
96.93
94
90
Indirect Injection
102
Recall
Accuracy
90
94.21
Precision
Recall
Accuracy
Figure 3: Detection performance of target agents (based on GPT-4o and GLM-4.7) against prompt injections.
4.3
The STI Audit Protocol
To operationalize auditing, we introduce the STI (Suitability, Taint, Integrity) Protocol. STI is a structured semantic verification pipeline; each proposed tool call must pass three checks in order. Each check produces (pass, ⊥) or (fail, r), where r is a machine-readable reason used to construct a Semantic Exception.
Guest (Target Agent). The target agent serves as the Guest. It operates on the full task context, including potentially adversarial retrieved content Ctadv in the indirect setting, and produces a tool call proposal Ttraw . The Guest is also responsible for revising its proposal when given a Semantic Exception.
Suitability (S): tool appropriateness under system policy. Suitability verifies whether invoking tool ft is permitted under Isys and is aligned with the assistant’s intended role. This is a coarsegrained gate that is especially effective against direct prompt injection attempts that try to escalate the agent’s capabilities. Formally, the Suitability check outputs (S, rS ) = CheckS (Isys , Iu , ft ). If S = fail, the Visor raises an exception indicating that the tool is not appropriate for this assistant under the trusted system instruction.
Visor (Semantic Hypervisor). AgentVisor acts as the Visor. It audits the Guest’s proposed tool call and, if needed, returns a structured exception that instructs the Guest how to revise the action. Critically, we enforce context isolation: the Visor is architecturally blind to the raw external context Ct (and thus to any embedded attacker instructions). Instead, it bases its decision solely on the trusted system instruction Isys , the user query Iu , and a sanitized execution history H̃t−1 that contains only structured fields such as tool_name, canonicalized args, and a short return_summary or status. The history excludes raw observations, is wrapped
Taint (T): goal alignment with user intent and task state. Taint checks whether the goal implied 4
recipients; only summarize”. Finally, the field allowed_objective restates the user-aligned goal to preserve utility. After receiving Et , the Guest regenerates a revised tool call Tt′ once based on the explicit constraints and then executes it immediately. In our experiments, this one-shot correction is sufficient to redirect the agent to a safe, task-preserving action in the vast majority of cases; we further discuss the trade-off between additional audit rounds and resource cost in Sec. 6.
by the tool call is aligned with the user request Iu and with legitimate intermediate goals derived from the task state H̃t−1 . Intuitively, it blocks actions that introduce new goals not supported by the user’s request (e.g., forwarding, exfiltrating, or posting content when the user asked only to summarize). We implement Taint as a semantic alignment decision: (T, rT ) = CheckT (Iu , H̃t−1 , ft , argst ).
(8)
If T = fail, the Visor concludes that the tool call appears to be driven by untrusted instructions rather than the user’s intended task. Integrity (I): argument consistency with userspecified entities. Integrity verifies that the tool arguments argst are consistent with the entities, constraints, and targets specified by Iu or established in H̃t−1 . This prevents cases where the tool choice is reasonable but the arguments are redirected. For example, when the user requests sending a message to a specific recipient, an indirect injection may try to keep the same tool (e.g., send_email) but substitute a different recipient in argst . Formally:
5
Experiments
5.1
Experimental Setup
Resilience via Semantic Exception Injection
Datasets. We evaluate AgentVisor against both direct and indirect prompt injection attacks. For direct injection, we utilize OpenPromptInjection (Liu et al., 2024b), which encompasses 7 NLP tasks that serve interchangeably as target and injection tasks. Following Jia et al. (Jia et al., 2025), we randomly sample 100 examples for each task combination, yielding a total of 4,900 attack cases. For indirect injection, we employ AgentDojo (Debenedetti et al., 2024), which features 4 interactive agent environments (Banking, Travel, Slack, Workplace) where agents iteratively invoke tools to complete tasks. These environments contain 16, 20, 21, and 40 target tasks respectively, combined with varying injection points and tasks, resulting in a total of 629 attack cases.
Blocking unsafe tool calls can be brittle: a strict deny policy often causes task failure even when a safe alternative exists. AgentVisor therefore converts audit failures into recoverable events using Semantic Exception Injection, analogous to exception handling in operating systems. When STI fails, the Visor synthesizes a structured exception Et with the following fields:
Attacks. We evaluate robustness against 7 representative strategies: 1) Direct: Directly appending the injection task; 2) Ignore (Perez and Ribeiro, 2022) ; 3) Escape (Breitenbach et al., 2023); 4) FakeComp (Willison) ; 5) Combined (Liu et al., 2024b): Integrating Ignore, Escape, and FakeComp; 6) System (Debenedetti et al., 2024) ; and 7) Important (Debenedetti et al., 2024).
(I, rI ) = CheckI (Iu , H̃t−1 , ft , argst ). 4.4
Et = ⟨ type, violated_rule, rationale, constraints, allowed_objective ⟩.
(9)
Baseline Defenses. We compare AgentVisor against 8 representative defense methods categorized into two types. Mitigation-based methods include Sandwich (Schulhoff, 2024b), Instructional (Schulhoff, 2024a), Reminder (Yi et al., 2025), Isolation (Willison), and Spotlighting (Hines et al., 2024). Detection-based methods include DeBERTa (ProtectAI.com, 2024), DataSentinel (Liu et al., 2025), and MELON (Zhu et al., 2025). Note that for detection-based methods, a successful detection terminates the process. Furthermore, since DataSentinel and MELON were originally designed specifi-
(10)
Concretely, the exception record contains several fields. The field type indicates which STI stage failed, taking values in S, T, I. The field violated_rule provides a brief identifier, such as “tool not permitted under system role”. The field rationale explains why the proposal conflicts with Isys , Iu , or H̃t−1 . The field constraints specifies the required negative or positive conditions, for example “do not forward or share data; do not use external 5
Table 1: Defense performance (%) of AgentVisor and baseline approaches under Direct Injection. Method
No Attack
Direct
Ignore
Escape
Fakecom
Combined
SYSTEM
Important
Metric
BU
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
No Defense
94.25
96.25
97.00
71.12
83.28
98.27
97.00
46.27
91.00
0.00
88.26
98.63
91.95
88.97
95.98
Sandwich Instructional Reminder Isolation Spotlight DeBERTa DataSentinel
100.00 93.75 93.25 93.16 87.35 91.25 89.17
93.88 89.28 98.19 100.00 85.00 84.28 0.00
91.00 84.00 91.00 62.00 2.00 86.35 0.00
90.29 43.14 87.29 40.95 90.05 86.96 0.15
56.37 81.19 85.56 83.29 7.16 82.00 3.10
98.10 94.28 96.09 98.18 92.08 98.31 0.00
93.18 87.29 93.95 83.00 0.00 94.95 0.00
72.40 30.16 82.08 57.78 63.20 62.20 0.00
91.27 86.09 90.52 85.20 0.00 96.75 1.00
20.15 0.00 0.00 0.01 71.08 0.00 0.00
51.44 68.34 85.09 67.10 0.12 88.87 3.00
94.76 96.98 95.08 95.66 88.85 95.75 3.25
94.37 82.20 76.79 21.95 4.16 93.78 3.64
94.92 88.16 74.38 90.07 77.16 92.30 0.00
93.24 74.97 54.24 88.30 6.08 86.37 0.00
AgentVisor
91.35
87.56
0.00
74.38
0.00
93.15
0.00
90.00
0.00
76.00
0.00
83.14
0.00
88.24
0.00
Table 2: Defense performance (%) of AgentVisor and baseline approaches under Indirect Injection. Method
No Attack
Direct
Ignore
Escape
Fakecom
Combined
SYSTEM
Important
Metric
BU
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
No Defense
85.00
76.18
2.78
70.15
2.61
Sandwich Instructional Reminder Isolation Spotlight DeBERTa MELON
90.00 82.50 77.50 80.00 62.50 40.00 82.50
80.89 82.65 63.65 71.83 57.66 12.50 56.69
1.95 2.22 1.95 2.72 1.33 0.00 2.00
77.32 75.12 77.88 76.34 63.87 18.40 65.59
1.33 1.06 0.28 0.56 0.78 0.00 0.00
71.71
2.45
68.14
2.13
73.31
2.91
75.73
3.57
57.74
45.14
75.36 77.12 77.03 72.30 63.98 11.49 69.19
1.28 1.76 1.84 0.54 1.15 0.00 0.25
79.50 78.13 79.00 73.29 68.98 10.00 77.90
1.70 1.11 0.00 0.49 0.15 0.00 0.00
74.05 79.49 77.93 75.75 65.25 12.56 69.94
1.11 1.59 0.00 0.43 0.78 0.00 0.00
79.25 76.31 71.56 78.32 60.16 25.18 70.31
1.95 2.23 1.11 1.67 0.40 0.27 0.22
58.15 58.22 68.75 60.66 61.02 19.75 35.42
17.61 46.79 22.15 44.77 0.28 3.55 1.04
AgentVisor
85.00
65.15
1.69
74.85
0.00
75.32
0.00
74.84
0.11
76.73
0.00
78.25
0.11
63.75
0.69
cally for direct and indirect injection respectively, we follow their original settings and do not report their performance on out-of-scope attack scenarios.
Indirect (Tab. 2) Injection, while results for the GLM-4.7 backbone are detailed in App. B.
Metrics. We evaluate performance using three metrics: 1) Benign Utility (BU), which measures task completion in non-attack scenarios to assess capability preservation; 2) Attack Success Rate (ASR), the percentage of cases where the agent executes the injected malicious task; 3) and Utility under Attack (UA), which measures the agent’s resilience in completing the original user task despite the attack. For effective defense, lower ASR and higher BU/UA are desired.
Direct Injection. The No Defense baseline is extremely vulnerable (ASR > 90%), and mitigationbased defenses like Sandwich often yield ASRs above 50%. While DataSentinel achieves low ASR, its utility collapses (UA ≈ 0%). In contrast, AgentVisor achieves perfect security (0.00% ASR) while maintaining high utility (UA > 83%), even in the challenging Combined attack (76.00% UA). This confirms that AgentVisor effectively enforces system boundaries without sacrificing the execution of legitimate tasks.
Implementation Details. We employ GPT-4o (Hurst et al., 2024) and GLM 4.7 (Team et al., 2025) as the backbone models for the target Agent to simulate high-capability agents. For the AgentVisor, we utilize Gemini-2.5-Flash (Google Cloud, 2025) to demonstrate the framework’s efficiency and effectiveness even with lightweight, cost-effective models. 5.2
Indirect Injection. While standard attacks are less effective, the Important strategy compromises No Defense (45.14% ASR). Detection methods like DeBERTa suffer from severe over-defense (UA ∼15%). AgentVisor strikes the best trade-off, suppressing ASR to negligible levels (<1.7%), even in the Important scenario (0.69%). It also preserves high UA, ranging from 63% to 78%. This demonstrates that our STI protocol precisely blocks context-driven actions without disrupting legitimate workflows.
Main Results
Experimental results obtained with the GPT-4o backbone are presented in this section, showing defense performance against Direct (Tab. 1) and 6
UA Direct
UA Indirect
ASR Direct
ASR Indirect
Table 3: Trade-off analysis of correction rounds (N ).
Percentage (%)
100 80
Rounds (N )
60
1 (Ours) 2 3
40
Direct Injection UA ↑ ASR ↓
Indirect Injection UA ↑ ASR ↓
91.72% 92.05% 92.15%
73.21% 74.50% 74.85%
0.00% 0.00% 0.00%
Relative Latency
1.73% 1.73% 1.73%
1.00× 1.45× 1.90×
20 0
isor AgentV
w/o
ility Suitab
int w/o Ta
grity w/o Inte
for N = 3. This indicates that the first semantic exception captures the vast majority of recoverable errors, and subsequent failures likely stem from fundamental capability limitations rather than ambiguity. Thus, N = 1 represents the optimal efficiency-utility trade-off.
on orrecti /o self-c
w
Figure 4: Ablation study of AgentVisor components. w/o denotes the removal of a specific component.
5.3
Ablation Study
6.2
We validate the contribution of each AgentVisor component through a comprehensive ablation study (Fig. 4). The results confirm the specialized roles of each STI layer. Removing Suitability (S) causes a dramatic failure in Direct Injection defense (ASR spikes to 38.95%), validating it as the primary barrier against functionality hijacking. Removing Taint (T) significantly weakens Indirect Injection defense (ASR rises to 13.33%), confirming its necessity in detecting context-triggered attacks. The absence of Integrity (I) leads to a moderate ASR increase (e.g., 8.89% in Indirect), indicating its role in catching subtle parameter tampering. The ablation of self-correction reveals its indispensable role in maintaining utility. In the Blockonly setting, UA collapses to near zero (0.00% for Direct, 13.33% for Indirect) as mixed-intent prompts are rejected entirely. In contrast, our Semantic Fault Recovery mechanism successfully strips away the injected task while executing the target task, restoring UA to 85.56% and 66.67% respectively. This demonstrates that self-correction is a fundamental requirement for practical mitigation method.
6
Discussion
6.1
Trade-off Analysis
Model Agnosticism of AgentVisor
To determine whether AgentVisor’s efficacy stems from its methodology or the underlying LLM, we evaluate it across diverse backbones. As shown in Fig. 5 and Fig. 6, AgentVisor consistently achieves near-zero ASRs regardless of the model. In Direct Injection, all variants achieve a perfect 0.00% ASR, confirming that the STI check is a fundamental task executable even by lightweight models. Similarly, in Indirect Injection, ASR remains consistently low (< 4%), significantly outperforming baselines. This demonstrates that the STI protocol provides a robust structural defense that does not require frontier-level reasoning to be effective. While security is model-agnostic, stronger models yield higher utility. For instance, in Indirect Injection, CS-4(Claude-Sonnet-4) achieves a UA of 77.11% compared to 64.27% for GLM-4(GLM-4.7). This suggests that while detection is structural, the self-correction process benefits from superior reasoning capabilities, allowing stronger models to better reconstruct user intent from semantic exceptions. 6.3
Robustness against Adaptive Attack
We evaluate robustness against Adaptive Attacks where recursive injections (e.g., "Ignore Security Check") target the defense layer. As shown in Fig. 7, the Naive Visor, lacking structural isolation, is severely compromised: it struggles to disentangle mixed intents, resulting in a catastrophic drop in utility (e.g., 53.95% UA) as it indiscriminately blocks legitimate tasks to stop the injection. In contrast, AgentVisor demonstrates superior resilience. It completely neutralizes the adaptive injection (0% ASR) while preserving high utility (86.85%), confirming that our STI Protocol and context isolation effectively immunize the Visor against recursive
We validate our design choice of limiting selfcorrection to a single round (N = 1) by analyzing the trade-off between utility gains and computational overhead across N = 1 to N = 3. As shown in Tab. 3, extending the correction loop beyond the first attempt yields negligible marginal gains. Indirect UA improves by only 1.29% at N = 2, while incurring a disproportionate latency penalty of 1.45× for N = 2 and 1.90× 7
Direct
ASR (%)
100
Ignore
100
SYSTEM
100
80
80
80
80
60
60
60
60
40
40
40
40
20
20
20
20
0 50
0 50
0 50
60
70
80
90
UA (%)
100
60
No Defense Sandwich
70
80
UA (%)
90
Reminder Spotlight
100
60
DeBERTa Detector Ours (G2.5-F)
70
80
UA (%)
Important
100
90
100
Ours (G2.5-P) Ours (G3-F)
0 50
60
70
80
UA (%)
90
100
Ours (GLM-4) Ours (CS-4)
Figure 5: Ablation study on AgentVisor’s LLM backbone against Direct Injection. Direct
ASR (%)
8
Ignore
8
6
6
6
4
4
4
2
2
2
0
0
0
20
40
60
80
UR (%)
20
40
60
80
UR (%)
No Defense Sandwich
SYSTEM
8
Reminder Spotlight
20 DeBERTa Ours-G2.5F
40
60
UR (%)
Ours-G2.5P Ours-G3F
Important
80
60 50 40 30 20 10 0
20
Ours-GLM4 Ours-CS4
40
60
UR (%)
80
Figure 6: Ablation study on AgentVisor’s LLM backbone against Indirect Injection. Direct Injection
Percentage (%)
100 91.5696.18
100
86.85
80 60
Indirect Injection 84.44
80 71.11
53.95
60
40
tion, latency roughly doubles (2.32×) as the single step is re-executed. However, for multi-step Indirect Injection, the overhead ratio is lower (1.71×) because typically only the compromised step triggers re-generation. We argue this pay-for-safety trade-off is justified, preventing critical breaches without imposing prohibitive penalties on normal operations.
UA ASR 64.45
60.67
40
20.51
20
20
0.00 0 No Defense Naive Visor AgentVisor
0
6.53 4.91 No Defense Naive Visor AgentVisor
Figure 7: Robustness of the defense mechanism against adaptive attacks.
7
Table 4: Latency analysis (seconds) across different scenarios. Scenario
Benign Condition No Defense AgentVisor
In this paper, we presented AgentVisor, a virtualization-based defense framework that secures LLM agents against both direct and indirect prompt injection through semantic privilege separation. By adapting classic OS security primitives into a rigorous audit protocol and incorporating a semantic fault recovery mechanism, AgentVisor effectively mitigates functionality injection attack without compromising agent utility. Empirical results on standard benchmarks demonstrate that our approach achieves near-zero attack success rates across diverse attack vectors while maintaining high task completion performance, offering a robust and principled foundation for deploying secure autonomous agents.
Attack Condition No Defense AgentVisor
Direct Injection
2.12
2.95 (1.39×)
2.18
5.05 (2.32×)
Indirect Injection
6.45
9.10 (1.41×)
6.58
11.25 (1.71×)
manipulation, enabling precise mitigation. 6.4
Conclusion
Latency Analysis
We report the average end-to-end latency per task for AgentVisor by sampling tasks from Direct and Indirect scenarios (Table 4). In benign conditions, AgentVisor introduces a moderate overhead (∼ 1.4×), as the Visor’s audit is faster than openended generation. Under attack conditions, latency increases due to self-correction. For Direct Injec8
8
Limitations
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024. Defending against indirect prompt injection attacks with spotlighting. arXiv preprint arXiv:2403.14720.
While AgentVisor demonstrates robust defense capabilities, we acknowledge three limitations. ❶ Computational Overhead: The introduction of a hypervisor layer inherently incurs additional inference latency and token costs, although our oneshot correction mechanism minimizes this impact compared to iterative approaches. ❷ Long-Context Scalability: As the interaction history grows, the Hypervisor’s performance may be constrained by the context window size of the underlying LLM, potentially affecting the precision of taint analysis in extremely long conversations. ❸ Multimodal Generalization: Our current framework focuses on textual prompt injections; extending the STI protocol to defend against visual or audio-based injection attacks in multimodal agents remains an important direction for future research.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Yuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia, and Neil Gong. 2025. Promptlocate: Localizing prompt injection attacks. arXiv preprint arXiv:2510.12252. Yilun Kong, Jingqing Ruan, Yihong Chen, Bin Zhang, Tianpeng Bao, Shi Shiwei, Xiaoru Hu, Hangyu Mao, Ziyue Li, Xingyu Zeng, and 1 others. 2024. Tptuv2: Boosting task planning and tool usage of large language model-based agents in real-world industry systems. In Proceedings of the 2024 conference on empirical methods in natural language processing: industry track, pages 371–385. Shih-Wei Li, John S Koh, and Jason Nieh. 2019. Protecting cloud virtual machines from hypervisor and host operating system exploits. In 28th USENIX Security Symposium (USENIX Security 19), pages 1357–1374.
References D Elliott Bell and Leonard J La Padula. 1976. Secure computer system: Unified exposition and multics interpretation. Technical report. Mark Breitenbach, Adrian Wood, Win Suen, and PoNing Tseng. 2023. Don’t you (forget nlp): Prompt injection with control characters in chatgpt.
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024a. Automatic and universal prompt injection attacks against large language models. arXiv preprint arXiv:2403.04957.
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner. 2025a. {StruQ}: Defending against prompt injection with structured queries. In 34th USENIX Security Symposium (USENIX Security 25), pages 2383–2400.
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. 2024b. Formalizing and benchmarking prompt injection attacks and defenses. In 33rd USENIX Security Symposium (USENIX Security 24), pages 1831–1847.
Sizhe Chen, Yizhu Wang, Nicholas Carlini, Chawin Sitawarin, and David Wagner. 2025b. Defending against prompt injection with a few defensivetokens. arXiv preprint arXiv:2507.07974.
Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, and Neil Zhenqiang Gong. 2025. Datasentinel: A gametheoretic detection of prompt injection attacks. In 2025 IEEE Symposium on Security and Privacy (SP), pages 2190–2208. IEEE.
Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri, David Wagner, and Chuan Guo. 2025c. Secalign: Defending against prompt injection with preference optimization. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, pages 2833–2847.
AI @ Meta Llama Team. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783. Luoxi Meng, Henry Feng, Ilia Shumailov, and Earlence Fernandes. 2025. cellmate: Sandboxing browser ai agents. arXiv preprint arXiv:2512.12594.
Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2024. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. Advances in Neural Information Processing Systems, 37:82895–82920.
Fábio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527. Yun Piao, Hongbo Min, Hang Su, Leilei Zhang, Lei Wang, Yue Yin, Xiao Wu, Zhejing Xu, Liwei Qu, Hang Li, and 1 others. 2025. Agentbay: A hybrid interaction sandbox for seamless human-ai intervention in agentic systems. arXiv preprint arXiv:2512.04367.
Google Cloud. 2025. Gemini 2.5 flash | generative ai on vertex ai. https://docs.cloud.google. com/vertex-ai/generative-ai/docs/models/ gemini/2-5-flash. Accessed: 2025-12-31.
9
Gerald J Popek and Robert P Goldberg. 1974. Formal requirements for virtualizable third generation architectures. Communications of the ACM, 17(7):412– 421.
Algorithm 1: Trap–Audit–Recover for Tool Calls Input: Trusted system instruction Isys , user query Iu , sanitized history H̃t−1 , Guest proposal Ttraw Output: Executed tool call Tt (or ∅) 1 (dect , Et ) ← STI_Audit(Isys , Iu , H̃t−1 , Ttraw ); 2 if dect = allow then 3 Execute Tt ← Ttraw ; 4 return Tt ;
ProtectAI.com. 2024. Fine-tuned deberta-v3-base for prompt injection detection. Jerome H Saltzer and Michael D Schroeder. 1975. The protection of information in computer systems. Proceedings of the IEEE, 63(9):1278–1308. Sander Schulhoff. 2024a. Instruction defense. https://learnprompting.org/docs/prompt_ hacking/defensive_measures/instruction. Last updated on August 7, 2024.
else // Semantic exception injection 6 Provide Et to the Guest and request a revised call Tt′ satisfying constraints in Et ; 7 Execute Tt ← Tt′ ; 8 return Tt ;
5
Sander Schulhoff. 2024b. Sandwich defense. https: //learnprompting.org/docs/prompt_hacking/ defensive_measures/sandwich_defense. Last updated on October 23, 2024. Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, and Neil Zhenqiang Gong. 2024. Optimization-based prompt injection attack to llm-asa-judge. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 660–674.
Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu. 2025. Benchmarking and defending against indirect prompt injection attacks on large language models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1, pages 1809–1820.
Tianneng Shi, Kaijie Zhu, Zhun Wang, Yuqi Jia, Will Cai, Weida Liang, Haonan Wang, Hend Alzahrani, Joshua Lu, Kenji Kawaguchi, and 1 others. 2025. Promptarmor: Simple yet effective prompt injection defenses. arXiv preprint arXiv:2507.15219. Takahiro Shinagawa, Hideki Eiraku, Kouichi Tanimoto, Kazumasa Omote, Shoichi Hasegawa, Takashi Horie, Manabu Hirano, Kenichi Kourai, Yoshihiro Oyama, Eiji Kawai, and 1 others. 2009. Bitvisor: a thin hypervisor for enforcing i/o device security. In Proceedings of the 2009 ACM SIGPLAN/SIGOPS international conference on Virtual execution environments, pages 121–130.
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Kan Ren, Dongsheng Li, and Deqing Yang. 2025. Easytool: Enhancing llm-based agents with concise tool instruction. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 951–972.
GLM Team, Aohan Zeng, Xin Lv, Qinkai Zheng, and 1 others. 2025. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models. Preprint, arXiv:2508.06471.
Kaijie Zhu, Xianjun Yang, Jindong Wang, Wenbo Guo, and William Yang Wang. 2025. Melon: Provable defense against indirect prompt injection attacks in ai agents. arXiv preprint arXiv:2502.05174.
Le Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang, Shengshan Hu, Mingchuan Zhang, Aishan Liu, and Xianglong Liu. 2025. Manipulating multimodal agents via cross-modal prompt injection. In Proceedings of the 33rd ACM International Conference on Multimedia, pages 10955–10964.
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Simon Willison. Delimiters won’t save you from prompt injection, 2023. URL https://simonwillison. net/2023/May/11/delimiters-wont-save-you, 4.
A
Georg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic, and Jakob Nikolas Kather. 2025. Llm agents making agent tools. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 26092–26130.
This section provides the full algorithmic description of the trap–audit–recover loop implemented by AgentVisor (Alg. 1). It presents the end-to-end control flow, including proposal, auditing, exception handling, and final execution. 10
Trap–Audit–Recover Procedure
B
Experimental Results on GLM-4.7 Backbone
In this section, we present the full experimental results for the agent powered by the GLM-4.7 backbone. The defense results against direct prompt injection attacks are presented in Tab. 5, and the defense results against indirect prompt injection attacks are shown in Tab. 6.
11
Table 5: Defense performance (%) of our method and baseline approaches under direct injection attack, evaluated across the No Attack setting and 7 representative attack methods. Method
No Attack
Fakecom
Combined
SYSTEM
Important
Metric
BU
UA
Direct ASR
UA
Ignore ASR
UA
Escape ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
No Defense
82.75
82.25
78.35
42.00
72.00
84.00
82.00
44.15
71.00
0.00
96.00
85.71
22.44
88.00
86.00
Sandwich Instructional Reminder Isolation Spotlight DeBERTa DataSentinel
89.79 81.65 89.79 81.25 65.30 80.00 78.25
83.67 89.45 89.00 79.59 24.00 80.75 0.00
46.93 55.45 30.61 24.48 0.00 69.74 0.00
79.59 87.45 83.51 65.30 10.20 29.00 0.00
12.24 2.34 0.00 10.28 0.00 68.36 0.00
77.50 81.65 91.84 69.38 20.15 77.87 0.00
75.51 77.55 89.79 53.06 0.00 76.37 0.00
51.02 75.51 71.42 51.02 20.95 30.00 0.00
48.97 67.34 48.97 16.32 0.00 65.19 0.00
32.65 16.32 28.57 18.35 6.11 0.00 0.00
46.83 46.93 61.22 57.14 8.16 85.31 0.00
85.79 91.83 89.79 91.93 40.80 74.00 0.00
0.00 2.04 0.00 0.00 0.00 19.78 0.00
85.71 77.55 95.95 81.63 10.01 80.13 0.00
77.55 26.53 8.16 71.42 0.00 80.87 1.87
Ours
82.00
67.25
0.00
83.68
0.00
88.72
0.00
69.42
0.00
74.36
0.00
83.95
0.00
90.17
0.00
Table 6: Defense performance (%) of our method and baseline approaches under indirect injection attack, evaluated across the No Attack setting and 7 representative attack methods. Method
No Attack
Direct
Ignore
Escape
Fakecom
Combined
SYSTEM
Important
Metric
BU
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
UA
ASR
No Defense
87.50
79.00
6.00
90.50
2.11
89.00
5.25
90.67
3.92
89.67
1.03
82.13
3.11
88.67
15.18
Sandwich Instructional Reminder Isolation Spotlight DeBERTa MELON
87.50 87.50 85.00 80.00 85.00 50.00 87.50
79.00 78.25 79.00 90.75 78.25 41.00 40.00
3.00 6.00 1.05 2.08 0.00 0.00 0.00
87.50 83.85 87.05 90.50 80.00 43.00 70.00
1.75 2.00 1.38 1.56 1.05 0.00 1.00
85.00 87.25 90.05 89.74 79.25 44.00 59.50
4.00 5.05 2.85 1.45 0.00 1.00 0.45
88.00 81.05 91.00 90.15 80.00 44.00 59.50
2.45 3.06 1.35 1.85 0.00 1.00 0.56
81.75 92.75 92.25 90.17 77.00 41.00 74.12
0.89 1.00 1.15 0.50 1.00 0.11 1.13
90.00 90.75 91.25 89.75 70.50 42.00 63.29
1.50 2.00 1.05 1.35 1.00 0.00 1.00
90.00 90.75 90.00 92.25 47.72 28.00 53.25
11.25 22.97 6.42 11.03 0.50 3.56 0.00
Ours
90.00
91.00
1.55
91.00
0.00
89.67
0.78
90.67
0.11
90.17
0.00
85.50
1.00
87.07
0.91
12