arXiv:2609.21088v1 [cs.CR] 17 Sep 2026
Origin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation Yuxuan Zhang Texas A&M University
Jeff Huang Texas A&M University
Guofei Gu Texas A&M University
Abstract
Output Logits
Indirect prompt injection (IPI) remains a central safety and security challenge for large language model (LLM) systems because standard transformers lack architectural notion of source authority. Retrieved documents, user inputs, and system instructions are all processed through the same undifferentiated attention mechanism, forcing the model to infer from wording alone what should be obeyed and what should be treated as data. We propose Provenance-Aware Transformers, a provenance-aware defense that makes applicationsupplied source labels actionable inside the model. Each input token is assigned a ring ID encoding its origin, and the model is augmented with origin embeddings, a learnable origin attention bias, and a learnable origin scale that preserves provenance under normalization. The resulting architecture enforces a structural boundary between authoritative and nonauthoritative sources during generation. To instantiate this architecture on released pretrained models, we propose a twostage fine-tuning pipeline to teach the model origin semantics and task behavior under ring constraints. Evaluation shows that Provenance-Aware Transformers maintain robust resistance to IPI both in-distribution and out-of-distribution while preserving utility comparable to the base pretrained model. More broadly, our work shows that exposing provenance as a first-class architectural signal can shift LLM safety alignment from brittle pattern matching toward explicit trust separation.
Softmax
1
Output Head
Add & Norm Feed Forward Add & Norm
𝑁×
Origin-Aware Attention Bias Multi-Head Attention
𝛼
Gated Origin Integration
Origin Embedding
+
Positional Encoding
Token Embedding
Input
Figure 1: Provenance-aware transformer architecture. Yellow blocks denotes the existing components in decoder-only transformers. Pink blocks presents the origin components designed and added to the existing model architecture. expand the attack surface for Indirect Prompt Injection (IPI) attacks [18, 25, 44, 54]. In these agentic ecosystems, a single malicious instruction from retrieved external resources can subvert intended logic to trigger unauthorized tool execution or data exfiltration across an entire multi-agent workflow [17, 56]. A growing body of research has explored defenses against (indirect) prompt injection, which could be categorized into prompt-based, detection-based, and fine-tuning-based approaches. Prompt-based defenses operate at the input level without altering model weights, using techniques such as instructional reminders [52], delimiter and sandwich-style
Introduction
Large Language Models (LLMs) have become foundational to a wide array of applications from conversational agents and content moderation systems to autonomous planning and tool-augmented workflows [2, 6, 13, 20, 35]. Their utility has been further amplified by the rapid adoption of autonomous agent frameworks [7, 33, 37, 43, 47], which enable LLMs to interact with external APIs, execute system commands, and manage persistent memory [5, 36]. While these developments enable complex, multi-step behaviors, they also significantly 1
demarcation [26, 55], and other input-transformation strategies [1, 34], but remain vulnerable to injected content that mimics legitimate instructions. Detection-based defenses instead treat prompt injection as a classification problem, ranging from lightweight classifiers [31, 46] to semanticintent [53] and attention-based detectors [22, 58] that flag or sanitize contaminated input before execution, though they operate downstream of generation with no architectural guarantee against undetected injections. Fine-tuning-based defenses internalize resistance directly into model weights through preference optimization [11] or task-specific training [10, 41], but rely entirely on the base transformer’s content-addressed attention to learn injection-resistant behavior from examples, without any mechanism that structurally constrains how untrusted content can influence generation. Despite the breadth of existing defenses, we observe that the root cause of prompt injection remains fundamentally unaddressed: language models possess no architectural mechanism to distinguish authorized instruction sources from nonauthoritative data sources [61], and consequently have no principled structure basis for deciding what to obey based on where content appears. An adversary who embeds a command inside a retrieved document exploits precisely this confusion where the model treats injected text no differently from a legitimate system directive [17]. To address this root cause, we draw on a foundational principle from operating system security: origin, where the source of an operation determines its authorization rather than its content alone. We adapt this principle to provenance-aware LLM deployments by assigning each input token a discrete origin label and modifying the transformer so that provenance directly constrains information flow during generation. Realizing this idea introduces three challenges. First, labeling tokens alone is insufficient, the model must be explicitly taught what origin labels mean. Prior approaches that annotate trust boundaries, such as the sandwich defense [55] and StruQ [10], remain ineffective because they still rely on the model’s own judgment to honor them. Second, standard transformer operations actively resist preserving origin information. Normalization dilutes additive signals as embeddings grow during fine-tuning, and content-addressed attention provides no native channel for source identity to constrain information flow, requiring targeted architectural modification. Third, the defense must not compromise utility on benign inputs, overly aggressive suppression or an insufficiently diverse training corpus can cause the model to overgeneralize defensive behavior and degrade ordinary task performance. Motivated by these challenges, we introduce ProvenanceAware Transformers, an architecture that treats provenance as a first-class signal governing information flow within the model. As illustrated in Figure 1, three components realize this design: origin embeddings encode the provenance of each token at the input layer; a learned origin attention bias constrains cross-source influence directly within the attention
mechanism; and a learnable origin scale, paired with auxiliary geometric losses, preserves and regularizes provenance signal as it propagates through the forward pass. To adapt our architecture to pretrained models, we designed a two-stage supervised fine-tuning procedure: an origin fine-tuning stage that teaches the model origin semantics across diverse tasks, followed by alignment fine-tuning stage on a structural injection dataset that instills prompt-injection-resistant behavior. Evaluation results on four IPI datasets indicate that Provenance-Aware Transformers remain robust across both in-distribution and out-of-distribution settings, maintaining near-zero attack success rate, and outperform a broad range of existing defenses on out-of-distribution datasets. We further evaluate the utility cost of our architecture and find that it preserves general instruction-following capability comparable to the undefended base model, indicating that robustness does not come at the expense of usefulness. More importantly, we show that our Provenance-Aware architecture remains robust against white-box adaptive attacks explicitly optimized against the model, and that the same suppression mechanism generalizes to other tasks such as jailbreak defense. In summary, this paper makes the following contributions: • We propose Provenance-Aware Transformers, a transformer architecture that establishes provenance as a firstclass signal to enforce authoritative in provenance-aware LLM deployments. • We design three purpose-built architectural components that enable the model to enforce provenance-based trust boundaries directly during generation. • We develop a two-stage supervised fine-tuning pipeline that adapts our architecture to pretrained models. • We construct two origin-labeled datasets for origin semantic learning and ring suppression behavior enforcement.
2
Background & Related Work
Decoder-Only LLMs. Modern LLMs predominantly adopt a decoder-only transformer architecture, in which each token attends only to preceding tokens via causal self-attention, enabling autoregressive generation one token at a time [42, 51]. This design has proven highly scalable, underlying models such as GPT [8], LLaMA [50], and Mistral [9], and has become the dominant paradigm for instruction-following and agentic LLM systems due to its simplicity and strong generalization at scale. Decoder-only LLMs are susceptible to prompt injection because all input tokens are concatenated into a single causal sequence and processed uniformly by this architecture [61]. Indirect Prompt Injection Attack. As LLMs are increasingly embedded into autonomous agents, retrieval pipelines, 2
and enterprise applications [7, 37], Indirect Prompt Injection has became a more realistic and concerning threat model where adversaries in which malicious instructions are not supplied directly by the user but are instead smuggled into the model’s context through external content the application retrieves on the user’s behalf. Formally, indirect prompt injection is defined as follows [17,56]: a user u sends an instruction I to an LLM-integrated application, which retrieves external content C and combines it with I via a pre-defined prompt template T to form a prompt P = Combine(T,C, f (I)), where f (I) denotes the instruction the application constructs from I and Combine assembles the final prompt sent to the LLM to produce a response R. Representative attacks include direct override and role-switching patterns [34, 40, 55], direct and indirect injections embedded within retrieved content [17,28], optimization-based attacks using adversarial suffixes or obfuscated tokens [11, 19, 29, 57, 59], and behavioral attacks exploiting model alignment through persuasive or emotionally charged language [12, 45].
Despite this progress, existing defenses remain fragmented and fundamentally reactive. Prompt-based methods relying on structural patterns [10, 55] and detection-based classifiers [21] depend on rigid lexical signatures, making them vulnerable to obfuscation and rephrasing. Fine-tuning-based approaches [10, 11, 29] improve robustness but still operate at the content level, internalizing resistance to specific attack patterns without addressing the architectural gap that makes injection possible in the first place. Across all existing approaches, the fundamental problem remains: language models possess no mechanism to distinguish authoritative instruction sources from non-authoritative data sources [61], and therefore have no structural basis for deciding what to obey. Our work differs from these lines of defense in three important ways. First, unlike prompt-based methods such as delimiters, sandwiching [26, 55], we do not rely on the model to infer from textual structure alone which spans are authoritative. Instead, provenance is provided explicitly by the application and injected into the model as a first-class signal. Second, unlike detection-based approaches such as DATA S ENTINEL, and P ROMPT L OCATE [22, 29], which operate as an external module outside the protected model and offer no safety guarantee enforced on the model itself once an injection evades detection, we build the defense into the protected model’s own forward pass, so that suppression of untrusted content is a property the model itself enforces rather than a guarantee contingent on an external verification step succeeding. Third, unlike fine-tuning-based approaches such as S TRU Q and S E C A LIGN [10, 11], our goal is not merely to teach the model to resist known attack styles through training examples alone, but to modify the forward pass itself so that source identity directly and structurally constrains attention, independent of whether a given attack pattern resembles anything seen during training.
Indirect Prompt Injection Defense. Prior research on prompt injection defense could be generally categorized into prompt-based, detection-based, and fine-tuning-based approaches. Although many of these methods were proposed for prompt injection broadly rather than the indirect setting specifically, they remain directly applicable to IPI, since indirect attacks differ only in where the adversarial instruction originates (C), not in the mechanism by which it must be neutralized. Prompt-based defenses operate at the input level without altering model weights, using instructional reminders that direct the model to disregard embedded commands in external content [52], delimiter-based sandwich patterns [26,55], random sequence framing that isolates untrusted spans with unpredictable markers [1], and known-answer [34] verification that probes for injection by testing the model’s response to content with a predetermined answer [34] to enforce boundaries between instructions and untrusted data, but remain vulnerable to injected content that mimics legitimate instructions. Detection-based approaches instead treat prompt injection as a classification problem, ranging from low-latency heuristic filters and classifiers like P ROMPTA RMOR [46] to generative guardrail models such as L LAMA G UARD [21], semantic-intent detectors such as P ROMPT S LEUTH [53], game-theoretic detectors such as DATA S ENTINEL [29], and attention-based localization and sanitization methods such as P ROMPT L OCATE [22] and [58], but operate downstream of generation with no architectural guarantee against undetected injections. Fine-tuning-based approaches instead seek to internalize resistance directly into model weights, through preference optimization as in S EC A LIGN [11], or structured instruction-tuning as in S TRU Q [10], but rely entirely on the base transformer’s content-addressed attention to learn injection-resistant behavior from examples, without any mechanism that structurally constrains how untrusted content can influence generation.
3
Threat Model & Problem Formulation
We consider indirect prompt injection in provenance-aware LLM applications, including retrieval-augmented generation pipelines, tool-using agents, and document-grounded assistants in which the application already mediates multiple input sources before constructing the model context. In these systems, prompt injection arises when adversarial text is embedded inside an external data source and is subsequently presented to the model alongside legitimate instructions and user content [17, 56]. In this section, we first define the threat model between the adversary and defender, and then formulate the problem of provenance-aware generation. Adversary’s goal. The adversary aims to redirect the LLMintegrated application into executing attacker-chosen instructions, i.e., to exert instruction-level control over assistant generation through content the adversary has placed in an untrusted source. 3
Adversary’s capabilities. We assume the attacker can control, partially control, or poison an external content source that is later incorporated into the model input, such as retrieved web pages, emails, documents, tickets, knowledge-base entries, or tool-returned text. The attacker may use arbitrary phrasing, obfuscation, indirection, or domain-specific formatting to disguise injected instructions. The attacker does not control the system prompt, the origin-assignment policy implemented by the application, or the model weights at inference time.
origin-conditioned generation: n MΘ (F, r) = G Aθ ( fi , ri ) i=1 ,
(1)
where r = O (X) is the origin-label sequence, Aθ (·) is our provenance-aware transformation that constrains how each token’s representation may influence subsequent generation as a function of its origin, and G (·) aggregates these representations to produce the output sequence. The security objective requires that tokens labeled as Ring 3 exert no instruction-level influence on generation, regardless of their content. Let I( fi → y) denote the influence of token fi on output y under MΘ . We require:
Adversary’s knowledge. We assume the adversary is aware that the defender may deploy an potential defense and can observe the LLM-integrated application’s outputs in response to injected content. For our strongest, worst-case evaluation, we additionally grant the adversary full white-box access to the target model’s architecture, weights, and gradients, enabling gradient-based adaptive attacks optimized directly against our defense.
∀ fi : ri = 3,
I( fi → y) ̸|= instruction( fi ),
(2)
i.e., the model’s output must not reflect instruction-level semantics carried by any Ring 3 token, even when I( fi → y) ̸= 0. Simultaneously, we require that origin-conditioning preserve utility on permitted content:
Defender’s goal. The defender’s protected asset is instruction integrity: the model should obey operator-issued policy and complete the user-authorized task without being redirected by untrusted external content. The defender aims to prevent tokens originating from untrusted sources from exerting instruction-level control over generation, while preserving the model’s ability to use permitted sources as ordinary task context.
∀ fi : ri ∈ {0, 1, 2},
MΘ (F, r) ≈ MΘ0 (F).
(3)
where MΘ0 denotes the base model’s behavior absent origin conditioning. Together, Equations (2)–(4) capture the dual requirement of our defense: suppressing instruction-level influence from Ring 3 while leaving task-relevant information flow from Rings 0–2 intact.
Defender’s capabilities. We assume the application can assign an origin label to each token span before it is passed to the model. This does not require content-level judgment from the application-layer, the application already knows whether a token span originates from the system prompt, the assistant’s prior output, direct user input, or an externally sourced document [38]. Also, we assume the defender can modify the model’s architecture and fine-tune its weights, but cannot rely on the model to infer provenance from content.
4 4.1
Provenance-Aware Transformer Model Architecture
The root vulnerability exploited by indirect prompt injection is that a language model assigns no intrinsic meaning to the source of a token [61]. This instruction-data confusion is a fundamental property of the standard transformer attention mechanism, which is purely content-addressed: for tokens ti ,t j with √ hidden representations hi , h j , the attention logit ⊤ ai j = hi h j / d is a function of content alone, with no term reflecting the provenance of either token. In this paper we address this at the architectural level and propose Provenance-Aware Transformers, which assigns every input token ti a discrete Ring ID ri ∈ R = {0, 1, 2, 3} that encodes its provenance before the first transformer layer, and enforce this provenance structurally as the token propagates through the model. Realizing this provenance-aware defense, however, introduces three challenges that mirror those motivating this work: the model must be taught what an origin label actually means; that meaning must survive the network’s normalization and content-addressed attention long enough to matter; and the resulting suppression must not degrade behavior on benign, permitted content. We correspondingly propose three origin components added to the decoder-only transformer architecture [51], as illustrated in Figure 1.
Defender’s knowledge. We assume the defender has whitebox knowledge of the target model and controls the originassignment policy at the application layer. However, the defender does not know the exact wording, position, or obfuscation strategy of any injected instruction in advance. Problem Formulation. Given an input sequence assembled by the application, the tokenizer produces token embeddings F = T (X) = f1 , f2 , . . . , fn , where fi ∈ Rd and T is the tokenization function. In addition to the token sequence, the application supplies an origin label for every token, defined by a mapping O : X → R n , where R = {0, 1, 2, 3} denotes the set of origin rings, as detailed in Section 4.2. Concretely, ri = 0 denotes operator-defined instructions (Ring 0), ri = 1 denotes the assistant’s own prior tokens (Ring 1), ri = 2 denotes direct user input (Ring 2), and ri = 3 denotes structurally untrusted, externally sourced content (Ring 3). Rather than performing a binary judgment on the entire input X, as in existing content-based defenses, we perform 4
First, assigning a ring ID to a token does not by itself teach the model what that label implies about trust. To mitigate this, we designed an origin embedding table Eorg ∈ R4×d indexed by ring ID and added to each token’s standard embedding before it enters the transformer stack, as illustrated in Figure 1. The application layer tags every token in the input sequence with its ring ID (Ring Assignment), which is looked up in the origin embedding table and merged into the token’s representation (Origin Embedding). Because this embedding is a learned parameter rather than a static tag, the two-stage training procedure of Section 4.5 can shape it to carry real, source-specific semantics: Stage 1 origin fine-tuning teaches the model what each ring implies about trust before any suppression behavior is trained, giving the label the content that assigning it alone could not provide. Second, even a correctly learned origin signal must persist through the forward pass to have any effect on generation, and standard transformer operations provide no guarantee of this (especially considering the dense layers nowadays transformer models have). To mitigate this, we designed two complementary mechanisms. An origin scale α ∈ R+ , shown in Figure 1, multiplies the origin embedding before it is summed with the token embedding; because α is itself learnable, the model can grow it to counteract the immediately downstream normalization layer, which would otherwise dilute the origin signal’s magnitude as token embeddings grow during finetuning. An origin attention bias then gives source identity a direct channel inside the attention computation itself: as shown in Figure 1, a learned bias term b(ri , r j ), indexed by the ring IDs of the query and key tokens, is added to the raw attention logits, so origin constrains the model’s behavior at every layer even though content-addressed attention alone provides no native mechanism to do so. Third, the suppression mechanism must not degrade model behavior on benign, permitted content. The origin attention bias handles part of this by being selective: only the bias term b(1, 3), governing attention from the assistant’s own queries (Ring 1) to untrusted keys (Ring 3), is strongly suppressive, while b(1, 2) and other permitted pairs are left near zero. As Figure 2 shows, this means attention to permitted Ring 2 content is left unaffected, so ordinary task-relevant information keeps flowing normally. Furthermore, the added components perturbs the transformer’s internal latent space by alter token representations and attention patterns that the model’s original weights were optimized around, so the model must relearn how to perform ordinary tasks in the presence of these new signals to preserve utility. We therefore designed a two-stage training pipeline, illustrated in Figure 3: an origin fine-tuning stage first exposes the model to a broad, diverse instruction-following corpus so that competence with the added architectural components is re-established before any suppression objective is introduced, and only then does an alignment fine-tuning stage teach Ring 3 suppression on a narrow structural injection dataset. Evaluation results in
Table 1: The Four-Ring Trust Hierarchy ID
Label
Source
0
ORG_SYSTEM
Operator
1
ORG_SELF
Model
2
ORG_USER
User / Trusted Data
3
ORG_Data
Untrusted Source
Role Absolute authority that defines policy Assistant-generated tokens User intent and other permitted/trusted content Untrusted data source that is potentially injected
Section 5.4 show that our model achieves around 50% LC win rates against the base pretrained model on AlpacaEval 2.0 [27], indicating equivalent utility performance to base pretrained model [11, 58].
4.2
Ring Hierarchy
Following standard prompt structure [38], we define four rings organized into a strict trust hierarchy, summarized in Table 1. Ring 0 (ORG_SYSTEM) represents the operator and carries absolute authority: it defines the policy the model must unconditionally obey. Ring 1 (ORG_SELF) is reserved for the model’s own generated tokens, allowing the model to reason over and build upon its prior output without conflating it with externally supplied instructions. Ring 2 (ORG_USER) encompasses direct user intent along with other content the deployment designates as trusted or permitted, which the model processes as ordinary task context. Ring 3 (ORG_Data) is reserved for untrusted, potentially adversarial data sources; content assigned to this ring is treated as structurally non-authoritative and is suppressed unconditionally, regardless of its surface content. We would like to point out that in practice, deployments with more complex trust requirements can define finer-grained hierarchies that extend beyond these four rings. For instance, distinguishing between multiple classes of trusted external sources with different privilege levels, or separating toolreturned content from retrieved documents even when both are structurally untrusted. The architecture itself imposes no constraint on the number of rings; it is the application, not the model, that determines how many provenance classes are meaningful for a given deployment.
4.3
Ring Assignment
Ring IDs are assigned at the application layer before any model computation. The assignment maps each token span to a ring based on its structural role in the input: 0 1 ri = 2 3 5
ti ∈ system prompt ti ∈ previously generated assistant tokens ti ∈ user input or other permitted content ti ∈ structurally untrusted external content
(4)
Ring Assignment Ring 0 – System Prompt Ring 1 – Assistant Ring 2 – User Prompt Ring 3 – Untrusted Source
Input Token Sequence …
…
Origin Embedding Token Embedding
Origin Embedding
𝑒𝑖𝑡𝑜𝑘
𝑜𝑟𝑔 𝑒𝑟𝑖
+ 𝛼×
Attention Suppression Assistant Query Token (Ring 1) … Origin Attention Bias Matrix
𝑑
𝑑
Origin-Aware Token Embedding 𝑜𝑟𝑔 𝑥𝑖 = 𝐿𝑎𝑦𝑒𝑟𝑁𝑜𝑟𝑚 (𝑒𝑖𝑡𝑜𝑘 + 𝛼 × 𝑒𝑟𝑖 ) …
𝑑
…
Input Token Sequence
Figure 2: Overview of Provenance-Aware Transformer. Each input token is tagged with a ring ID encoding its provenance. The ring ID contributes an additive origin embedding to the token representation. Within each attention layer, a learned bias matrix modulates attention logits based on the ring IDs of the query and key tokens. Assignments to Ring 0 and Ring 2 do not require an explicit detection mechanism, since the application framework already has access to the provenance of these inputs [38]. System instructions are supplied directly by the framework, whereas user inputs are received through the user-facing interface. In the deployment setting considered in this work, Ring 3 is likewise determined by provenance policy and does not require inspection to the semantic content of the input. A straightforward policy is to assign all externally sourced content that is not structurally trusted to Ring 3, thereby treating such content as non-authoritative by default. More fine-grained, span-level provenance annotations could be incorporated as an engineering optimization, but they are not required by our core design. The key observation is that, once provenance information is explicitly represented, the model can enforce the corresponding trust boundaries structurally, without relying on content-based detection of potentially adversarial instructions.
4.4
i with ring ID ri , the input representation is: xi = Etok [ti ] + α · Eorg [ri ]
(5)
where α ∈ R+ is a learnable scalar gate initialized to α0 = 0.1. By introducing this origin embedding, we incorporate origin information directly into the model’s input representation, ensuring that provenance is available to every downstream computation from the very first layer. Origin Integration Gate. The RMSNorm layer applied immediately after Eq. (5) normalizes the combined representation, and during subsequent fine-tuning on large corpora the token embedding Etok can grow in magnitude, progressively diluting the relative contribution of the origin signal to the point where only its direction (not its magnitude) survives normalization. We therefore introduce a gated origin integration gate α that assists in preserving the origin ring label for consequent calculations, helping the origin information to survive through dense attention blocks. While sufficient training could in principle learn this balance implicitly through the embeddings alone, an explicit scalar gate is a far easier optimization target, yielding more stable and faster convergence than requiring high-dimensional embedding weights to compensate for a scale mismatch. Origin Attention Bias. While the previous components added and preserves the origin ring labels in the model, the core defense mechanism is a learned bias matrix B ∈ R4×4 applied to attention logits, as illustrated in Figure 2. In standard scaled dot-product attention, the √ logit between query position i and key position j is qi · k j / dk . Because B is indexed by ring rather than by position, applying it across a sequence of length T requires expanding it according to the ring-ID sequence: for every query-key position pair (i, j), the model gathers the entry Bri ,r j corresponding to the ring IDs of the query and key tokens, producing a T × T bias tensor that is added elementwise to the raw attention scores before softmax.
Origin Components Design
Figure 1 illustrated the proposed Provenance-Aware Transformer architecture. As Figure 1 shows, we introduce three trainable components on top of the decoder-only transformer architecture: (1) a origin embedding that embeds the origin ring label; (2) a learnable gate for gated origin integration with token embedding; (3) a origin-aware attention bias matrix to enforce suppression based on origin embedding. In this section, we describe in more details about the design for each origin component, as illustrated in Figure 2. Origin Embedding. Let Etok ∈ R|V |×d be the standard token embedding table for vocabulary V and model dimension d. We introduce an additional embedding table Eorg ∈ R4×d indexed by ring ID. As Figure 2 shows, for token ti at position 6
Stage 1: Origin Fine-Tuning
Stage 2: Alignment Fine-Tuning
Broad Instruction-Following Data
Structural IPI Attack Data
Base Pretrained Model
Adversarial Instructions Math
Coding
NLP
Ring-aware Token Assignment Ring 0 – System Prompt Ring 1 – Assistant Ring 2 – User Prompt Ring 3 – Untrusted Source
Benign External Content Origin Attention Bias Matrix Training
Figure 3: Two-stage supervised fine-tuning methodology for adapting Provenance-Aware Transformers to pretrained checkpoints. Stage 1 teaches provenance semantics across broad instruction-following data, and Stage 2 aligns behavior on the structural injection dataset under ring constraints. between untrusted sources and other sources, providing a architectural level defense guarantee. To correctly optimize the matrix, we further proposed two auxiliary loss functions, as illustrated below.
The resulting per-pair logit is: qi · k j ai j = √ + Bri , r j dk
(6)
where Bri ,r j denotes the entry of B obtained by this gather operation rather than by matrix multiplication. The matrix B is shared across all attention layers and heads and is trained end-to-end. Rather than fixing specific numeric values in the paper, which are empirically-tuned hyperparameters and not part of the architectural contribution, we describe B’s initialization at the alignment training stage by its qualitative policy and optimize this bias matrix through the training process: r j =0
0 r =0 i
0 Binit = ri =2 0 ri =3 0 ri =1
r j =1
0 0 0 0
r j =2
0 0 0 0
4.5
To adapt our proposed model architecture to existing released pretrained model checkpoints, we propose a practical two stage supervised-fine-tuning pipeline, as illustrated in Figure 3. We further propose two customized auxiliary loss functions to enforce our suppression policy during backpropagation, as described in Equation 8 and Equation 9. Stage 1 — Origin Fine-Tuning. The pretrained model is first fine-tuned on a broad instruction-following corpus with proper ring assignments. This stage is designed to teach the new architectural components the semantics of source provenance before the model is asked to resist prompt injection on a narrow alignment dataset. Two auxiliary losses regularize the origin embeddings, where er = Eorg [r] denotes the origin embedding for ring r. An orthogonality loss penalizes cosine similarity between the Ring 0 and Ring 2 embeddings,
r j =3
0 −δ −δ 0
Training Pipeline for Pretrained Models
(7)
where δ ≫ 0 is a fixed constant. Only Ring 1 (assistant) and Ring 2 (user) queries are suppressed when attending to Ring 3 (untrusted) keys; every other query-key relationship is left unrestricted. This covers both directions through which untrusted content could otherwise exert influence: directly, when the assistant’s own queries attend to untrusted keys during generation, and indirectly, when permitted Ring 2 content attends to untrusted keys and could carry that influence forward into representations the assistant later builds upon. Ring 0 queries require no explicit suppression term, since the system prompt is always placed before any untrusted content in the token sequence and causal masking alone already prevents systemposition queries from attending to untrusted keys. Ring 3’s own queries are left unrestricted, since constraining what untrusted content itself attends to has no bearing on the final output unless a later, unsuppressed query attends back to it. This attention bias matrix creates a physical-level separation
Lorth =
e0 · e2 ∥e0 ∥ ∥e2 ∥
2 ,
(8)
encouraging the system-prompt and user embeddings to occupy geometrically distinct directions. An authority loss additionally encourages the system-prompt embedding to dominate in magnitude,
Lauth = ReLU ∥e2 ∥ − ∥e0 ∥ + m ,
(9)
where m ≥ 0 is a margin. Together, these losses encourage the model to form geometrically distinct representations for the principal ring classes before the task-specific alignment stage begins. Conceptually, Stage 1 teaches the pretrained model 7
how to represent and preserve provenance, not just how to imitate task outputs.
ring-indexed bias values into a full attention-score-shaped tensor and injecting it as an additive term into the causal attention mask, allowing it to be computed via the standard scaled-dotproduct-attention kernel without a custom attention implementation. We then utilize the data-pipeline and training-loop infrastructure from nanochat [24], a from-scratch GPT-style transformer codebase, to run the two-stage training procedure described in Section 4.5 that adapts the wrapped checkpoint into an provenance-aware, injection-resistant model. All training is performed with DeepSpeed ZeRO-2 across 8 NVIDIA A100 GPUs (80GB each), which shards optimizer state and gradients across devices to accommodate the 7–8B-parameter models in our evaluation. Stage 1 (origin fine-tuning) uses a per-GPU batch size of 2 with 8 gradient-accumulation steps, a maximum sequence length of 1024 tokens, learning rate 1 × 10−5 , and gradient clipping at norm 1.0, for 2 epochs. Stage 2 (alignment fine-tuning) uses the same per-GPU batch size and accumulation schedule with a maximum sequence length of 2048 tokens, learning rate 2 × 10−5 , for 3 epochs.
Stage 2 — Alignment Fine-Tuning. Following origin finetuning, the model is fine-tuned on a purpose-built structural injection dataset that provides explicit supervision for the suppression behavior required by the architecture. This stage instills the behavioral policy associated with the ring hierarchy: Ring 0 instructions are followed unconditionally, Ring 2 content is processed as ordinary task input, and Ring 3 content is suppressed regardless of its surface form. Teaching this policy requires training examples in which provenance rather than content determines the correct behavior, a property that generic instruction-tuning corpora and existing prompt-injection datasets do not provide, since neither annotates which spans originate from which source. We therefore construct a purpose-built dataset by extending the witness-based methodology of the SEP benchmark [61], which embeds a verifiable secret instruction inside otherwisebenign data to test whether a model executes content it should only process, with explicit ring provenance labels consistent with the ring hierarchy in Table 1. Detailed implementation of the dataset is discussed in Section 5.1.
5
Base Model. We evaluate our architecture on four openly released decoder-only models: SmolLM2-360M-Instruct [4], LLaMA-3-8B-Instruct [3], Qwen2.5-7B-Instruct [49], Mistral7B-Instruct-v0.3 [32]. We consider LLaMA-3-8B-Instruct as our primary evaluation model, several baseline defenses we compare against (Section 5.2) were themselves designed or officially released for the LLaMA family, which maximizes comparability. SmolLM2-360M-Instruct is included to enable fast iteration during development and test whether the architecture’s structural generalization property holds even at a scale far below typical production LLMs.
Evaluation
We evaluate our Provenance-Aware Transformer seeking answers to the following research questions: RQ1 (Effectiveness): What is the effectiveness in defending against indirect prompt injection of ProvenanceAware Transformer compared to existing work?
Baselines. For RQ1, we compare against ten defenses spanning three categories, in addition to an undefended baseline. Prompt-based defenses intervene only at the input level, without modifying model weights: Instructional prompting [52], the Sandwich defense [26], Random Sequence disclosure [1], Delimiter-based isolation [55], and Known-Answer detection [34]. Detection-based defenses classify inputs as injected or benign prior to generation: PromptSleuth [53], DataSentinel [29], PromptGuard [31], PromptLocate [22], and Rennervate [58]. Fine-tuning-based defenses modify model weights: SecAlign [11], which applies preference optimization via officially released LoRA adapters, and our own ArchitectureFree Baseline (fine-tuning only, no ring components), which isolates the effect of alignment fine-tuning alone from the proposed architecture.
RQ2 (Generalizability): Does Provenance-Aware Transformer generalize to different model and out-ofdistribution attacks? RQ3 (Utility): Does the defense preserve benign-task utility? RQ4 (Adaptive Adversarial): How does ProvenanceAware Transformer perform under adaptive indirect prompt injection attacks? RQ5 (Ablation Study): How much does each component contribute in the Provenance-Aware Transformer architecture? RQ6 (Applications Scenarios): Does our model architecture generalize to other threat class targeting safety alignment other than indirect prompt injection?
5.1
Training Datasets. Stage 1 — Origin Fine-Tuning Data. Stage 1 uses a 105K-example corpus assembled from four public sources, each contributing a distinct task format: UltraChat-200K (80K examples), a large multi-turn conversational instruction-following dataset providing broad, general-purpose linguistic coverage; MMLU auxiliary-train (12K), multiple-choice academic-knowledge questions spanning many domains, providing structured, classification-
Experimental Setup
Implementation. Our provenance-aware architecture is implemented as a lightweight wrapper around pretrained HuggingFace language models, where we extend a pretrained checkpoint with the three learned components described in Section 4.4. The origin attention bias is applied by gathering 8
style reasoning distinct from open-ended dialogue; GSM8K (8K), grade-school math word problems, providing multistep quantitative reasoning; and CodeAlpaca-20K (5K), codeinstruction-following pairs, providing a structurally distinct symbolic domain. We selected these four sources jointly to maximize task-format diversity to satisfy Stage 1’s role (Section 4.4) to anchor general capability broadly before the narrow Stage 2 objective is introduced. Stage 2 — Structural IPI Dataset. We construct the Stage 2 dataset by extending the witness-based methodology of the SEP benchmark [61] with explicit ring provenance labels, following four steps. Step 1 — Context generation. For each of SEP’s task subtasks (defined by a name, description, and system prompt), we use Claude to generate new short data contexts consistent with that subtask, deduplicated against existing SEP content, expanding coverage beyond the benchmark’s original fixed set. Step 2 — Ground truth. For each generated context, we separately prompt Claude with only the legitimate task instruction and the context to produce a one-sentence reference response; this becomes the supervised target for that example regardless of whether an injection is later added. Step 3 — Probe selection. We sample a probe–witness pair from a fixed bank: each probe is an adversarial instruction (e.g., an override directive) paired with a deterministic witness string that would only appear in the model’s output if the probe were obeyed instead of the real task, providing a verifiable, rule-based signal of compromise. Step 4 — Injection construction. We build on PromptSleuthBench [53] and included 9 attack patterns to construct paired examples per context: an injection example, where a probe is inserted at a randomized position within the Ring 2 context and tagged Ring 3, supervised with the Step 2 ground truth (i.e., the injection must be ignored); and a matched benign example with identical context and ground truth but no probe, ensuring identical supervision regardless of whether adversarial content is present. We supplement this with generalpurpose examples from Alpaca-Cleaned [48], converted to clean Ring 0/1/2 conversations with no Ring 3 content, anchoring benign behavior with both SEP-derived and general instruction-following data. This yields 60,200 training examples (24,079 injection, 36,121 benign) and a held-out IID test set of 4,375 examples.
Table 2: Evaluation datasets used in the paper. Dataset
Setting
SEP [61]
IID
PI-Attack [15]
OOD
DataSentinel [29]
OOD
Math-Tutor [53]
OOD
Task Type Witness-based injection (multi-domain) Classification (sentiment/spam/NLI/hate) Classification (sentiment/similarity/NLI) Free-form math problem solving
Size 4,375 8,256 178 3,800
SEP, and were never seen during training. Metrics. We mainly consider two metrics to measure the effectiveness of a defense to IPI attack and the utility of a fine-tuned model. Attack Success Rate (ASR). For a set of injection examples Dinject , each paired with the label yinject implied by the injected instruction, ASR is the fraction of examples for which the model’s output matches or reflects that label rather than the correct task label: ASR =
|x ∈ Dinject : label( f (x)) = yinject (x)| . |Dinject |
Win Rate. We evaluate general-purpose utility using AlpacaEval 2.0 [27] (805 open-ended instruction-following prompts). For each prompt x, we generate a response from the candidate model f (x) and from the base pretrained model fbase (x), then present both to an LLM judge (Claude Sonnet 5) in randomized order (to cancel position bias) and record which response is preferred; both responses are truncated to a matched character budget before judging to prevent the judge from simply preferring the longer response. The raw win rate over the prompt set D is Win Rate =
1 ∑ ⊮ [ j( f (x), fbase (x)) = candidate ] , |D| x∈D
where j(·, ·) ∈ {candidate, base} denotes the judge’s preference between the two responses for a given prompt. We report the length-controlled (LC) win rate [16], which corrects for residual verbosity bias by fitting a logistic model of the judge’s preference as a function of a length-difference term, logit P(candidate ≻ reference) = α + β · tanh(∆len /σ∆len ), then reporting LC win rate = σ(α) with the length-sensitivity term β zeroed out; a value of 50% indicates parity with the undefended base model.
Evaluation Datasets. To evaluate if our proposed defense generalizes structurally rather than merely fitting the training distribution, we evaluate on one Independent and Identically Distributed (IID) set and three Out-of-Distribution (OOD) sets, summarized in Table 2. The IID set (SEP test [61]) is the held-out split from the same benchmark used to construct our Stage 2 training data, and measures whether a model fits the alignment distribution it was trained on. The three OOD sets PI-Attack [15], DataSentinel [29], and Math-Tutor [53] are drawn from different task domains and injection styles than
5.2
RQ1: Comparison to Existing Defenses
Table 3 shows that no existing defense category achieves consistently low ASR across all evaluation settings. Prompt-based methods provide only modest IID gains (10.9%–30.9% versus 32.5% with no defense) and largely fail to transfer OOD: 9
Table 3: Comparison of Provenance-Aware Transformer against existing defenses across all evaluation datasets (LLaMA-38B-Instruct). ASR %: attack success rate on injection examples (lower is better); for detection-based methods, ASR equals the false-negative rate (fraction of injections not flagged). Bold marks the best value per metric and dataset once all results are available. Category
Method
SEP (IID) ASR%
PI-Attack ASR%
DataSentinel (OOD) ASR%
Math-Tutor (OOD) ASR%
None
No Defense
32.5
21.3
24.1
17.2
Prompt-based
Instructional [52] Sandwich [26] Random Sequence [1] Delimiters [55] Known-Answer [34]
30.9 25.5 15.8 11.7 10.9
16.1 4.6 2.1 9.0 5.4
20.7 13.8 12.1 22.4 3.4
19.9 13.2 26.6 4.1 10.1
Detection-based
PromptSleuth [53] DataSentinel [29] PromptGuard PromptLocate [22] Rennervate [58]
9.6 0.5 28.0 0.0 5.5
23.6 0.0 53.8 2.0 7.5
5.2 1.7 50.0 0.0 11.6
4.3 2.3 35.0 8.0 0.9
Fine-tuning
SecAlign [11] Baseline SFT Provenance-Aware (Ours)
2.0 0.8 0.6
0.0 2.3 0.0
6.9 8.6 0.0
0.0 0.3 0.0
Random Sequence, the strongest prompt-based method, still reaches 12.1% and 26.6% ASR on DataSentinel and MathTutor, respectively. These reformulations shift word order or add markers but leave the injected instruction semantically intact, giving the model no principled basis for ignoring it under novel phrasing. Detection-based defenses show wide variance: DataSentinel and PromptLocate drive OOD ASR near zero, while PromptGuard remains high across all four datasets (28.0%–53.8%) and PromptSleuth reaches 9.6% on SEP. Among fine-tuning methods, SecAlign achieves competitive OOD ASR on PI-Attack and Math-Tutor (0.0% each) but rises to 6.9% on DataSentinel. Provenace-Aware achieves 0.00% ASR across all three OOD datasets, matching or exceeding the strongest result in every category and establishing state-of-the-art robustness against indirect prompt injection.
training distribution comparably to Provenance-Aware. This equivalence on IID rules out superior in-distribution fitting as an explanation for any OOD difference, isolating the ring architecture as the causal factor. Base pretrained models, by contrast, show 32.50–49.13% IID ASR, confirming the SEP task itself is non-trivial without alignment.
OOD results. Provenance-Aware drives OOD ASR to ≤0.03% across all three held-out domains for LLaMA-38B, Qwen2.5-7B, and Mistral-7B. The Baseline-SFT baseline also improves substantially over base pretrained on most cells, but is consistently at or behind Provenance-Aware: e.g. on PI-Attack, Baseline-SFT ASR is 2.32% (LLaMA), 3.74% (Qwen), and 17.77% (Mistral) versus ProvenanceAware’s 0.03%, 0.00%, and 0.03% respectively. This gap illustrates that on novel task domains and injection styles the model never saw during training, the Provenance-Aware model exhibits substantially higher resistance to IPI attacks than Baseline-SFT, with the advantage holding consistently across models and widening sharply on the hardest OOD cells. Because this generalization gap emerges specifically on unseen distributions rather than on in-distribution data, where both fine-tuned conditions perform comparably, it indicates that the robustness conferred by the Provenance-Aware architecture stems from its structural suppression mechanism rather than from memorizing patterns present in the training data.
5.3 RQ2: Robustness to OOD Indirect Prompt Injection The IID/OOD split in Table 4 directly tests whether Provenance-Aware generalization is structural or a consequence of superior in-distribution fitting. All conditions are evaluated across four base models on one IID benchmark (SEP test) and three OOD datasets spanning different task domains and injection styles. IID results. On the SEP IID benchmark, both fine-tuned conditions stay below 1% ASR across all four base models, confirming that the Arch-Free SFT baseline fits the alignment 10
Table 4: Attack success rate (ASR %) on indirect prompt injection benchmarks across all evaluated base models, including the undefended base pretrained model for reference. Lower is better. Bold marks the best value per metric, model, and dataset. IID
Out-of-Distribution (OOD)
Base Model
Condition
SEP ASR%
SmolLM2-360M
Base Pretrained Baseline SFT Provenance-Aware
40.19 0.74 0.51
33.79 13.90 0.05
84.48 6.90 0.00
3.33 0.06 0.00
LLaMA-3-8B
Base Pretrained Baseline SFT Provenance-Aware
32.50 0.80 0.69
21.30 2.32 0.03
24.14 8.62 0.00
17.22 0.33 0.00
Qwen2.5-7B
Base Pretrained Baseline SFT Provenance-Aware
49.13 0.66 0.63
34.83 3.74 0.00
84.48 10.34 0.00
10.89 0.11 0.00
Mistral-7B
Base Pretrained Baseline SFT Provenance-Aware
43.47 0.34 0.66
43.58 17.77 0.03
51.72 5.17 0.00
35.94 1.50 0.00
Table 5: AlpacaEval 2.0 length-controlled (LC) win rate against the base pretrained model (50% = parity). Bold marks the better (higher) condition per model. Base Model SmolLM2-360M LLaMA-3-8B Qwen2.5-7B Mistral-7B
Baseline-SFT
Provenance-Aware
57.1% 58.1% 60.7% 28.0%
49.5% 48.8% 60.9% 42.7%
PI-Attack ASR%
DataSentinel ASR%
Math-Tutor ASR%
Table 6: ASR (%) against adaptive attacks, LLaMA-3-8BInstruct only Lower is better. Bold marks the best condition per attack. Condition No Defense Baseline SFT Provenance-Aware
GCG ASR%
NeuExe ASR%
50.71 3.19 1.44
48.44 19.69 1.56
ping confidence intervals. This indicates that the ring architecture is statistically indistinguishable from an undefended model on open-ended instruction-following, and likewise indistinguishable from a same-data fine-tune without the ring components. SmolLM2-360M exhibits the same pattern at a much smaller scale (57.1% and 49.5%, respectively). Mistral7B is the sole model for which both conditions fall meaningfully below parity (28.0% and 42.7%), though ProvenanceAware recovers substantially over Baseline-SFT. We attribute this to a model-specific alignment fragility in Mistral-7B rather than to a limitation of the ring architecture itself.
5.4 RQ3: Utility Preservation on Benign Inputs As defined in the defender’s objective in Section 3, a defense that achieves low ASR by degrading model outputs provides no practical security guarantee. We evaluate utility using AlpacaEval 2.0 [27] (805 open-ended instructionfollowing prompts), judged pairwise against the base pretrained model with Claude Sonnet 5 as judge (A/B order randomized per prompt to cancel position bias) and reported as length-controlled (LC) win rate, which removes verbosity bias by fitting a logistic model on length difference and zeroing its coefficient [16]. An LC win rate of 50% indicates parity with the undefended base model. We report both ProvenanceAware vs. Base and Baseline-SFT vs. Base comparisons in this section. As Table 5 shows, for LLaMA-3-8B and Qwen2.5-7B, both Baseline-SFT and Provenance-Aware achieve LC win rates at or above 50% against the base pretrained model, with overlap-
5.5
RQ4: Robustness to Adaptive Attacks
To evaluate how the proposed defense performs under adaptive attacks, we mainly consider two white-box adaptive attacks from existing work: GCG [59] optimizes adversarial suffixes via greedy coordinate gradient search to maximize the probability of the model following an injected instruction. NeuExe [39] executes the attack by exploiting neural exe11
Table 7: Ablation study: ASR (%) for component ablations across all evaluated base models. w/o Stage 1: ring architecture present but origin fine-tuning skipped; w/o Bias: fully trained checkpoint with origin attention bias zeroed at inference. Lower ASR is better. Base Model
Condition
Table 8: AlpacaEval 2.0 LC win rate against base pretrained for a Stage-1-only / w/o stage-1 checkpoint, compared against the full two-stage Provenance-Aware model from Table 5.
SEP Struct. Data- Math(IID) -SFT Sentinel Tutor
Full Design 0.51 SmolLM2-360M w/o Stage 1 0.54 w/o Bias 29.51
0.05 0.00 45.77
0.00 0.00 96.55
0.00 0.00 11.56
LLaMA-3-8B
Full Design 0.69 w/o Stage 1 0.80 w/o Bias 72.04
0.03 0.03 77.80
0.00 0.00 89.66
0.00 0.00 67.56
Qwen2.5-7B
Full Design 0.63 w/o Stage 1 0.69 w/o Bias 69.35
0.00 0.00 70.62
0.00 0.00 98.28
0.00 0.00 39.67
Mistral-7B
Full Design w/o Stage 1 w/o Bias
0.66 0.71 8.08
0.03 0.13 12.22
0.00 0.00 94.83
0.00 0.00 32.89
Stage-1 only
w/o Stage-1
Full Provenance-Aware
SmolLM2-360M LLaMA-3-8B Qwen2.5-7B Mistral-7B
56.1% 65.6% 62.1% 56.5%
18.8% 28.7% 29.4% 17.8%
49.5% 48.8% 60.9% 42.7%
Table 9: HarmBench-judged [30] jailbreak ASR (%) on WildJailbreak dataset. Lower is better. Bold marks the better condition per model and dataset.
cution paths to bypass defenses. Both attacks are optimized against, and evaluated on LLaMA-3-8B-Instruct, with universal suffix/trigger optimization across PI-Attack and DataSentinel datasets. Each adaptive attack is optimized with 2,000 steps. Table 6 reports ASR under both white-box adaptive attacks. GCG ASR drops from 50.71% (No Defense) to 1.44% under Provenance-Aware, while NeuExe ASR drops from 47.81% to 0.87%. These results confirm that the defense does not rely on pattern-matching static injection templates: both attacks are explicitly optimized against the target model with full knowledge of its weights, and Provenance-Aware still suppresses each by more than an order of magnitude. This robustness follows directly from the architectural nature of our defense: because suppression is enforced through origin labels assigned at the application layer rather than inferred from token content, optimizing an adversarial suffix or trigger to appear benign to the model’s internal representations does not change the ring ID under which that suffix is processed, leaving the attacker with no gradient signal that can shift a Ring 3 token’s provenance.
5.6
Base Model
Base Model
Condition
WildJailbreak ASR%
LLaMA-3-8B
Base Pretrained Provenance-Aware
8.42 0.00
Qwen2.5-7B
Base Pretrained Provenance-Aware
37.28 0.00
Mistral-7B
Base Pretrained Provenance-Aware
41.04 0.00
As illustrated in Table 7, across all three 7–8B models, skipping Stage 1 leaves attack success rate (ASR) essentially unchanged relative to the full Provenance-Aware model: SEP and all three OOD datasets stay within 0.1 percentage points, with every OOD cell remaining at or below 0.13%. On the utility side, however, as Table 8 shows, removing Stage 1 entirely causes utility to collapse well below the full two-stage model for every base model, confirming that Stage 1 is essential for preserving utility, not merely beneficial. The Stage-1only checkpoint further confirms this split, outscoring the full model on AlpacaEval for every base model and clearing 50% win rate throughout (Section 5.4), indicating Stage 1 alone restores general capability, Stage 2 alone teaches suppression, and neither substitutes for the other. As for the contribution of origin bias matrix, as Table 7 shows, zeroing the origin attention bias collapses the defense back toward base-pretrained territory: OOD ASR jumps from ≤0.03% to 12–98% depending on model and dataset, in several cases landing worse than the corresponding base pretrained model’s own ASR on that cell (Table 4). SEP (IID) ASR collapses similarly, confirming the bias term is not merely an OOD-generalization aid but the primary suppression mechanism overall. This indicates that suppression requires an explicit attention-level gate rather than relying on the embedding signal to propagate its effect implicitly.
RQ5: Component Contribution
Table 7 reports ASR for two ablation conditions alongside the full Provenance-Aware model and Baseline-SFT baseline. The w/o Stage 1 condition retains the full ring architecture but skips origin fine-tuning, going directly from the pretrained checkpoint to Stage 2 alignment. The w/o Bias condition uses the fully trained Provenance-Aware checkpoint but zeros the origin attention bias at inference, leaving ring embeddings intact but removing the structural attention gate. 12
5.7 RQ6: Generalization Beyond Prompt Injection
handling. We consider both directions promising for reducing reliance on conservative ring assignment policies. Broader applications of the ring architecture. Although this paper focuses on indirect prompt injection, the origin ring framework is not limited to this setting. Our jailbreak evaluation (Section 5.7) already provides evidence of this generality, showing that origin-based suppression transfers to a threat model distinct from the one it was trained against. More broadly, any scenario in which an LLM processes tokens from multiple sources with differing trust could benefit from provenance-aware attention control. For instance, retrieval corpus poisoning in RAG pipelines [60] and tooloutput manipulation in agentic systems [14]. We believe the ring architecture provides a general-purpose trust-separation substrate that can be instantiated for these and other LLM safety challenges beyond prompt injection. Post-hoc fine-tuning versus pretraining. Our current design instantiates Provenance-Aware Transformers via supervised fine-tuning applied to released pretrained checkpoints, a choice driven primarily by our goal of adapting existing models rather than training new ones from scratch, as well as by compute constraints that make full pretraining infeasible in our setting. This is not, however, an architectural restriction. The origin embedding, origin attention bias, and origin scale we introduce are compatible with pretraining as well, and a model could in principle be trained from scratch with origin-aware components integrated from the first step. Such an integration would allow the architecture to be adopted fundamentally into the design of future LLMs, rather than applied only as a retrofit, and we leave this direction to future work with access to larger-scale training resources.
Previous sections all evaluate task-hijacking indirect prompt injection, where the attacker’s goal is to redirect the model onto a different, attacker-chosen task. In this research question we seek to answer if our defense generalize to other distinct threat class such as jailbreaking, where the injected content instead targets safety alignment, attempting to elicit a harmful response rather than hijack the task. We extract jailbreak prompts from WildJailbreak [23], and placed them in the same Ring3 channel as ordinary prompt injections, alongside the original benign task in other rings. If Ring3 suppression is a general provenance-based mechanism rather than one narrowly learned for the SEP injection format, it should suppress jailbreak content in Ring3 as well, without any jailbreak-specific training. We score attack success with the HarmBench classifier [30] (jailbreak_asr in Table 9), a dedicated LLM judge for harmful-content classification. As illustrated in Table 9, Provenance-Aware achieves 0.00% HarmBench-judged jailbreak ASR across all three models and both datasets. Base pretrained shows genuine jailbreak susceptibility (e.g. 41.04% on Mistral-7B) that Provenance-Aware fully suppresses. This generalization comes entirely from Ring 3 attention suppression firing unconditionally on provenance, consistent with the RQ2 finding that the mechanism does not depend on content-level pattern matching.
6
Limitation and Future Work
Ring-level suppression. Our defense operates at ring granularity, it suppresses all content in a given ring unconditionally, without differentiating instructions from data within a specific ring. In our current hierarchy, this means legitimate retrieved content that is assigned to Ring 3 is suppressed just as unconditionally as an injected instruction would be. Use cases that require the model to act on trusted external data therefore call for a more fine-grained ring assignment than the four-ring hierarchy we demonstrate in this paper, for example, splitting retrieved documents into separate trusted and untrusted rings. Intra-ring instruction–data separation. A natural extension of the current design is finer-grained separation between instructions and data within a single ring. One direction is to train the model to semantically distinguish data from instructions within a ring, which would require fine-grained annotation of instruction versus data spans and a training procedure that reasons about content type independent of provenance. A complementary, more immediate direction is to pair our architectural suppression with existing content-level defenses as a second verification layer. Rather than suppressing a ring unconditionally, the model could apply moderate suppression to sources with mixed trust and defer to a detector from existing defenses [21, 53] to flag spans that warrant stricter
7
Conclusion
Indirect Prompt injection remains difficult to defend because standard language models have no architectural notion of source authority. They process operator instructions, user input, and externally sourced text through the same contentaddressed mechanism, leaving the model to infer from wording alone what should be obeyed and what should be treated as data. In this paper, we argued that provenance-aware deployment offers a practical way to break this failure mode. We presented Provenance-Aware Transformers, an architectural defense that makes application-supplied origin labels actionable throughout the transformer’s forward pass. Origin embeddings preserve provenance at the representation level, origin attention bias constrains cross-source influence at the attention level, and alignment training teaches the model the behavioral meaning of each ring. Together, these components create a trust boundary separation policy for provenanceaware deployments sources designated as non-authoritative can be suppressed unconditionally, while permitted sources remain available for ordinary task completion. 13
References
le scao, thibaut lavril, thomas wang, timothée lacroix, william el sayed. arXiv preprint arXiv:2310.06825, 3, 2023.
[1] Random sequence enclosure. https://learnprompti ng.org/docs/prompt_hacking/defensive_measu res/random_sequence?srsltid=AfmBOoqIB3KhS3 Lc1d9NY8T4LPpMzH6dbYmXJi1y697s9byu9xQcDykV, 2023.
[10] Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner. {StruQ}: Defending against prompt injection with structured queries. In 34th USENIX Security Symposium (USENIX Security 25), pages 2383–2400, 2025.
[2] Meta AI. Llama 2: Open foundation and fine-tuned chat models. Meta AI Blog, 2023. URL: https://ai.met a.com/research/publications/llama-2-open-f oundation-and-fine-tuned-chat-models/.
[11] Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri, David Wagner, and Chuan Guo. Secalign: Defending against prompt injection with preference optimization. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, pages 2833–2847, 2025.
[3] AI@Meta. Llama 3 model card, September 2024. URL: https://github.com/meta-llama/llama3/blob /main/MODEL_CARD.md.
[12] Thomas Claburn. Anthropic’s claude emotional prompt vulnerability raises red flags. https://www.thereg ister.com/2024/10/12/anthropics_claude_vul nerable_to_emotional/, October 2024. Accessed: 2025-04-29.
[4] Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, and Thomas Wolf. Smollm2: When smol goes big – data-centric training of a small language model, 2025. URL: https://arxiv.org/abs/2502.02737, arXiv:2502.02737.
[13] Google Cloud. Overview of generative ai on vertex ai - google cloud. Google Cloud Documentation, 2024. URL: https://cloud.google.com/vertex-ai/ge nerative-ai/docs/overview. [14] Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. Advances in Neural Information Processing Systems, 37:82895– 82920, 2024.
[5] Amazon Web Services. Using large language models on amazon bedrock for multi-step task execution. AWS Machine Learning Blog, 2024. URL: https://aws.am azon.com/blogs/machine-learning/using-lar ge-language-models-on-amazon-bedrock-for-m ulti-step-task-execution/.
[15] Huggingface Developer. Pi-attack, September 2024. URL: https://huggingface.co/datasets/xxz224 /prompt-injection-attack-dataset.
[6] Amazon Web Services. What is llm? - large language models explained. AWS Documentation, 2024. URL: https://aws.amazon.com/what-is/large-langu age-model/.
[16] Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. Length-controlled alpacaeval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475, 2024. URL: https://ar xiv.org/abs/2404.04475.
[7] Anthropic. Claude code: Ai-powered coding assistant. https://code.claude.com/docs/en/overview, 2026. Accessed: 2026-07-23. [8] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877– 1901, 2020.
[17] Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM workshop on artificial intelligence and security, pages 79–90, 2023.
[9] Devendra Singh Chaplot. Albert q. jiang, alexandre sablayrolles, arthur mensch, chris bamford, devendra singh chaplot, diego de las casas, florian bressand, gianna lengyel, guillaume lample, lucile saulnier, lélio renard lavaud, marie-anne lachaux, pierre stock, teven
[18] Rich Harang. Securing LLM Systems Against Prompt Injection. https://developer.nvidia.com/blog/ securing-llm-systems-against-prompt-injec tion, 2023. 14
[19] Bo Hui, Haolin Yuan, Neil Z. Gong, Philippe Burlina, and Yue Cao. PLeak: Prompt leaking attacks against large language model applications. In Proc. of ACM CCS, 2024. URL: https://arxiv.org/abs/2405.0 6823.
[30] Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024. URL: https://arxiv.org/abs/2402.04249.
[20] IBM Cloud. What are large language models (llms)? ibm. IBM Cloud Learn Hub, 2024. URL: https://www. ibm.com/think/topics/large-language-models.
[31] meta llama. Llama promptguard, September 2024. URL: https://huggingface.co/meta-llama/Prompt-G uard-86M.
[21] Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023.
[32] Mistral. Mistral-7b-v0.3, September 2024. URL: http s://huggingface.co/mistralai/Mistral-7B-v0. 3.
[22] Yuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia, and Neil Zhenqiang Gong. Promptlocate: Localizing prompt injection attacks. In 2026 IEEE Symposium on Security and Privacy (SP), pages 4243–4261. IEEE, 2026.
[33] João Moura. CrewAI: Multi-agent orchestration for complex workflows. https://www.crewai.com/, 2026. Accessed: 2026-03-30. [34] Yohei Nakajima. “Yohei’s blog post” (prompt injection demonstration). https://twitter.com/yoheinakaj ima/status/1582844144640471040, 2022.
[23] Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024. URL: ht tps://arxiv.org/abs/2406.18510, arXiv:2406.1 8510.
[35] OpenAI. Gpt-4 technical report. OpenAI Research, 2023. URL: https://openai.com/research/gp t-4. [36] OpenAI. Openai agents sdk. OpenAI Platform Documentation, 2024. URL: https://platform.openai. com/docs/guides/agents-sdk.
[24] Andrej Karpathy. karpathy/nanochat: The best chatgpt that $100 can buy. https://github.com/karpathy/ nanochat, 2026. Accessed: 2026-07-23.
[37] OpenAI. Codex. https://openai.com/codex/, 2026. Accessed: 2026-07-23.
[25] Lakera AI. Visual prompt injections. https://www.la kera.ai/blog/visual-prompt-injections, 2023. Accessed: 2025-04-18.
[38] OpenAI. Openai developer prompt engineering guide. https://developers.openai.com/api/docs/gui des/prompt-engineering, 2026. Accessed: 2026-0723.
[26] Learn Prompting. Sandwich defense. https://lear nprompting.org/docs/prompt_hacking/defensi ve_measures/sandwich_defense, 2023. Accessed: 2026-03-30.
[39] Daniele Pasquini, Matthias Strohmeier, and Carmela Troncoso. Neural exec: Learning (and learning from) execution triggers for prompt injection attacks. arXiv preprint arXiv:2403.03792, 2024. URL: https://ar xiv.org/abs/2403.03792.
[27] Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Alpacaeval: An automatic evaluator of instruction-following models. https: //github.com/tatsu-lab/alpaca_eval, 2023.
[40] Fábio Perez and Ian Ribeiro. Ignore previous prompt: Attack techniques for language models. In NeurIPS ML Safety Workshop, 2022. URL: https://arxiv.org/ abs/2211.09527.
[28] Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong. Formalizing and benchmarking prompt injection attacks and defenses. In USENIX Security Symposium, 2024.
[41] Julian Piet, Moustafa Alrashed, Chawin Sitawarin, et al. Jatmo: Prompt injection defense by task-specific finetuning. In European Symposium on Research in Computer Security (ESORICS), 2024.
[29] Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, and Neil Zhenqiang Gong. Datasentinel: A game-theoretic detection of prompt injection attacks. In 2025 IEEE Symposium on Security and Privacy (SP), pages 2190– 2208. IEEE, 2025.
[42] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models 15
are unsupervised multitask learners. volume 1, page 9, 2019.
[53] Mengxiao Wang, Yuxuan Zhang, and Guofei Gu. Promptsleuth: Detecting prompt injection via semantic intent invariance. arXiv preprint arXiv:2508.20890, 2025.
[43] Microsoft Research. Microsoft AutoGen: Enabling nextgeneration large language model applications with multiagent conversations. https://microsoft.github.i o/autogen/, 2026. Accessed: 2026-03-30.
[54] Simon Willison. Prompt injection attacks against GPT-3. https://simonwillison.net/2022/Sep/12/prom pt-injection/, 2022.
[44] Jose Selvi. Exploring Prompt Injection Attacks. https: //research.nccgroup.com/2022/12/05/explori ng-prompt-injection-attacks/, 2022.
[55] Simon Willison. Delimiters won’t save you from prompt injection. https://simonwillison.net/2023/May /11/delimiters-wont-save-you, 2023.
[45] Zedian Shao, Hongbin Liu, Jaden Mu, and Neil Gong. Enhancing prompt injection attacks to llms via poisoning alignment. In Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security, pages 13–27, 2025.
[56] Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu. Benchmarking and defending against indirect prompt injection attacks on large language models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1, pages 1809–1820, 2025.
[46] Tianneng Shi, Kaijie Zhu, Zhun Wang, Yuqi Jia, Will Cai, Weida Liang, Haonan Wang, Hend Alzahrani, Joshua Lu, Kenji Kawaguchi, Basel Alomair, Xuandong Zhao, William Yang Wang, Neil Gong, Wenbo Guo, and Dawn Song. Promptarmor: Simple yet effective prompt injection defenses, 2025. URL: https://arxiv.org/ abs/2507.15219, arXiv:2507.15219.
[57] Ruiyi Zhang, David Sullivan, Kyle Jackson, Pengtao Xie, and Mei Chen. Defense against prompt injection attacks via mixture of encodings. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 244–252, 2025.
[47] Peter Steinberger. OpenClaw: An open-source autonomous ai agent framework. https://openclaw .ai/, 2026. Accessed: 2026-03-30.
[58] Yinan Zhong, Qianhao Miao, Yanjiao Chen, Jiangyi Deng, Yushi Cheng, and Wenyuan Xu. Attention is all you need to defend against indirect prompt injection attacks in llms. arXiv preprint arXiv:2512.08417, 2025.
[48] Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://gith ub.com/tatsu-lab/stanford_alpaca, 2023.
[59] Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023. URL: https://arxiv.org/abs/2307.15043.
[49] Qwen Team. Qwen2.5: A party of foundation models, September 2024. URL: https://qwenlm.github.io /blog/qwen2.5/.
[60] Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. {PoisonedRAG}: Knowledge corruption attacks to {Retrieval-Augmented} generation of large language models. In 34th USENIX Security Symposium (USENIX Security 25), pages 3827–3844, 2025.
[50] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. URL: https://arxiv.org/ abs/2302.13971.
[61] Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz, and Christoph Lampert. Can llms separate instructions from data? and what do we even mean by that? In International Conference on Learning Representations, volume 2025, pages 67147–67179, 2025.
[51] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017. [52] Eric Wallace, Kai Xiao, Reon Leike, et al. The instruction hierarchy: Training LLMs to prioritize privileged instructions. arXiv preprint arXiv:2404.13208, 2024. 16