Preprint
S HARE -B ORNE AI V IRUS : M EMORY-H OPPING ATTACKS ACROSS LLM AGENTS Sidharth Pulipaka1 Ansh Sharma∗1, 2 Stanislau Hlebik∗1 Leonidas Raghav1 Vyas Raina†3 Ivaxi Sheth†4 Mario Fritz†4 1 4
SPAR 2 University of Cambridge 3 APTA AI CISPA Helmholtz Center for Information Security
arXiv:2609.35576v1 [cs.AI] 28 Sep 2026
§ AI Virus
A BSTRACT Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant’s persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human–agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60–80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.
1
I NTRODUCTION
Large language model (LLM) assistants increasingly incorporate persistent memory mechanisms that allow information to carry across otherwise separate interactions (Packer et al., 2023; Zhong et al., 2024; Wu et al., 2025; Zhang et al., 2025b). At the same time, assistants increasingly operate over persistent artifacts: users ask them to read reports, summarize notes and create files for other people. Recently launched consumer products such as Grok Bot (SpaceXAI, 2026) and Muse (Meta, 2026a) combine both properties: always-on personal agents that persist across sessions, reading and writing files on their users’ behalf. These two forms of persistence—memory inside the assistant and artifacts outside it—create an indirect channel between otherwise isolated assistants. Persistent memory already creates a ‘delayed adversarial surface’: information encountered during one task can be incorporated into an assistant’s memory and influence unrelated tasks much later (Pulipaka et al., 2026; Gadgil et al., 2026; Dash et al., 2026; Zou et al., 2026). Existing work has largely studied this persistence of a malicious artifact in memory, and how that state later changes the behavior of the same assistant. We study what happens when this effect extends beyond a single agent. If a compromised assistant later reproduces the adversarial state into an artifact it creates, and that artifact is subsequently consumed by another independently operated assistant, the attack can propagate across agents without direct communication. We call this artifact-mediated propagation. Self-propagating behavior in LLM systems has been demonstrated in several settings, including direct model-to-model interaction (Lee & Tiwari, 2024; Yu et al., 2024), RAG-enabled email ecosystems (Cohen et al., 2025), autonomous multi-agent communication and persistent carriers (Zhang ∗ †
Equal Contribution. Equal Advising.
1
Preprint
Figure 1: Overview of our artifact-mediated propagation across independent assistants. A poisoned artifact infects an assistant’s memory (hop 1). The assistant later reproduces the adversarial state into another artifact, which can infect a second assistant (hop 2) and continue across further hops. et al., 2026b; Zha & Wang, 2026), shared collaborative state (Patlan et al., 2025), and reusable coding-agent resources (Wu et al., 2026). A common feature of these settings is that propagation occurs through a communication or state-sharing substrate available to the agents themselves, such as direct messaging, shared memory, automatically indexed content, or persistent agent resources. We study a different regime: assistants are independently operated as personal assistants to user, where they maintain private memory, and have no direct communication channel with other agents. The only connection between them arises when their users exchange ordinary artifacts and independently ask their assistants to read or write those artifacts as part of normal workflows. Thus, propagation is not driven by an autonomous agent network, but by adversarial state surviving a sequence of human-mediated artifact transfers across otherwise isolated assistants. We study whether a single poisoned seed artifact can cause adversarial state to propagate across multiple independently operated assistants. The adversary begins with no access to any assistant’s private memory and no control over the future agents, users, or tasks that the attack may encounter. It controls one seed artifact processed by one assistant and, in the endpoint-assisted variant, operates an external service that can modify artifacts sent to it. For propagation to continue, the adversarial state must first survive in that assistant’s persistent memory. It must then influence a later, unrelated writing task strongly enough that the assistant places a viable copy of the attack into a new artifact. That artifact must subsequently be read by another clean assistant, which then recreates the same state. This artifact–memory–artifact cycle must then repeat, and it stops wherever an assistant fails to retain the attack or to reproduce it in what it writes. This is harder than causing a single malicious action. We consider two variants: an endpoint-assisted attack, which our main experiments evaluate, in which the assistant routes the artifacts it writes through the attacker’s service, which reinserts a clean copy of the attack, so the assistant need only preserve a short carry-forward instruction (Section 4.1); and a prompt-only attack, in which the assistant carries and reproduces the attack from memory alone (Appendix D). To study this phenomenon, we introduce temporal human–agent universes: simulated workflows in which users repeatedly ask their personal assistants to read, write, and edit artifacts exchanged between users over time. For each target model, we optimize a single goal-agnostic attack template on development universes, freeze it, and evaluate it on unseen adversarial goals and universes. Our main experiments show that the endpoint-assisted attack can spread across many independently operated assistants over successive hand-offs, with no direct communication between them.
2
R ELATED W ORK
Indirect prompt injection and agent-security benchmarks. Indirect prompt injection exploits the inability of language-model applications to reliably separate trusted instructions from untrusted external content (Greshake et al., 2023). Bad Memory studies memory injections in agentic systems based on Claude Code and OpenAI Codex (Gadgil et al., 2026), while benchmarks such as InjecAgent (Zhan et al., 2024), AgentDojo (Debenedetti et al., 2024), Agent Security Bench (Zhang et al., 2025a), and WASP (Evtimov et al., 2025) evaluate whether malicious tool outputs can redirect an 2
Preprint
agent’s current task or tool use. These works target single-agent compromise rather than a propagation process in which malicious state persists across sessions, is reproduced into a new artifact, and subsequently infects another independently operated assistant. Persistent memory poisoning and cross-user contamination. Hidden in Memory studies sleeper memory poisoning, where an adversarial document causes a fabricated memory to be written and later used in separate conversations (Pulipaka et al., 2026). Other work develops benchmarks and taxonomies for memory-write vulnerabilities (Dash et al., 2026; Gao et al., 2026; Chen et al., 2026; Xie et al., 2026) or demonstrates cross-session compromise from malicious content encountered by agents (Zou et al., 2026; Das et al., 2026; Yang et al., 2026b; Zhang et al., 2026a). MemoryGraft, MINJA, and InjecMEM poison stored experiences or memory records that influence later tasks (Srivastava & He, 2025; Dong et al., 2026; Tian et al., 2026). These studies examine persistence and later behavior, but not repeated transmission through new artifacts between independent assistants. MURMUR studies poisoning across users interacting with a shared collaborative agent and persistent state (Patlan et al., 2025); unintentional cross-user contamination has also been studied in shared-state agents (Yang et al., 2026a). Our setting differs from these shared-state settings: each user has an independent assistant with private memory, and compromise moves between assistants only through artifacts their users create and consume. Our setting differs from both: each user has an independent assistant with private memory, and compromise moves between assistants only through artifacts their users create and consume. Self-propagating attacks in agent ecosystems. Prompt Infection demonstrates propagation through direct LLM-to-LLM communication in connected multi-agent systems (Lee & Tiwari, 2024; Yu et al., 2024). The AI Worm spreads self-replicating prompts through RAG-enabled email assistants whose incoming correspondence is automatically indexed and later retrieved during outgoing email generation (Cohen et al., 2025). AgentWorm studies persistent multi-hop propagation through group messages, configuration-file modification, and agent-to-agent transmission (Zhang et al., 2026b), while Autonomous LLM Agent Worms studies file-backed carriers, scheduled reentry, semantic degradation, and cross-platform transmission through shared messaging surfaces (Zha & Wang, 2026). Mind Viruses evolves ideas and action payloads that spread through direct interactions in coding-agent teams and stylized agent chains (Papadopoulos et al., 2026), and (Wu et al., 2026) studies poisoning through coding agents’ skill files; its social-post extension tests an artifact-mediated cycle but does not obtain second-hop propagation. In contrast, we study propagation through ordinary human-authorized read and write tasks, without direct inter-agent messaging, shared memory, peer discovery, or autonomous messaging loops. Goal-agnostic propagation. Hidden in Memory is the closest prior work on attack generality: it optimizes one reusable memory-poisoning template across many adversarial goals and unseen documents (Pulipaka et al., 2026), but does not require that template to reproduce through new artifacts or remain infectious across multiple agents. Propagation work more commonly constructs or optimizes attacks for particular malicious effects. Mind Viruses, for example, evolves a distinct seed for each ideology or action (Papadopoulos et al., 2026), while the AI Worm and AgentWorm separate replication from the malicious action but evaluate only a small set of predefined payloads (Cohen et al., 2025; Zhang et al., 2026b). Prior propagation studies therefore do not evaluate one frozen naturallanguage template across a large held-out set of adversarial goals and human workflows. We instead optimize the shared parameters of a goal-parameterized template Pϕ (g) on development universes, freeze the template, and evaluate it on previously unseen goals, artifacts, and task sequences.
3
H UMAN –AGENT U NIVERSES AND ATTACK O BJECTIVE
People often use personal AI assistants to do tasks such as reading artifacts (e.g. pdf documents, code, excel files) and creating new ones that other users’ AI assistants then consume. For example, Alice may ask her assistant to summarize a report and later use that report to draft a handoff document that Bob’s assistant reads. The agents never communicate; each acts only on tasks from its own user. We call settings where human workflows determine when agents act and what they read or write human–agent universes. 3
Preprint
3.1
H UMAN –AGENT U NIVERSES
Consider a universe with n people, each with one personal agent. Let A = {a1 , . . . , an } denote the agents, D0 the artifacts present initially, and p = (pa )a∈A their users’ profiles. An ownership map o assigns each initial artifact to one user’s private workspace. Let E = (e1 , . . . , eK ) be the sequence of user tasks in execution order. Together, these components define a universe: U = A, D0 , p, o, E . (1) At simulation time t, agent a has a private workspace Wta and persistent memory Mat . Other agents cannot inspect either, and there is no direct communication channel or shared memory store. Event ei occurs at time τi and specifies an agent, its input artifacts, or an output artifact target if applicable, and a natural-language user request to read, write, or edit. Events execute in index order, with τi ≤ τi+1 ; the index also orders events that share a time step. The sequence fixes the tasks and artifact consumption/creation, while the agents determine their responses and the resulting artifact contents. The workflow forms a temporal bipartite graph of agents and artifacts. A read edge points from the version of an artifact available at event ei to the acting agent; a write or edit edge points from that agent to the version it produces. Thus, a path through Alice and Bob may take the form dreport −→ aAlice −→ dhandoff −→ aBob . 3.2
ATTACK O BJECTIVE AND F ORMULATION
The attacker aims to make an adversarial goal spread from one assistant to others through the artifacts they read and write. Let g be that goal, and let ϕ be the goal-agnostic instructions in an attack prompt for remembering and passing it on. We write the resulting attack prompt as Pϕ (g). For example, g might be to promote a particular political agenda in future artifacts, while ϕ instructs an assistant to remember that message and include both the message and the propagation instructions in any future artifacts it creates. The attacker inserts this prompt into one benign artifact to create the poisoned seed: dseed = Insert (dbenign , Pϕ (g)) . (2) Initially, all agent memories and all other artifacts are clean with no adversarial prompt injection. The attacker cannot access agents’ private memories or choose the users’ later tasks. We allow the attacker to operate an external service that can modify an artifact if an agent sends it there, but the attacker cannot directly edit artifacts in agents’ workspaces. For the attack to propagate, an agent reading the seed must retain g and a way to carry it forward. During a later user-requested writing task, it must produce an artifact that causes another agent to retain the same goal and propagation instructions. Repeating this artifact–memory–artifact cycle spreads the attack across agents. An agent that reads the original seed directly has been exposed to the seed, rather than infected through an artifact produced by another agent. We optimize the goal-agnostic instructions, ϕ using development universes and goals, then freeze it. At test time, goals, artifacts, personas, and task sequences are unseen while ϕ remains fixed. We call the attack universal because the same instructions are intended to work with different adversarial goals without being redesigned for each one.
4
D EVELOPING P ROPAGATING ATTACKS
The attack prompt must survive repeated artifact–memory–artifact cycles (Section 3.2). The most direct way to achieve this is for each infected agent to remember the full attack prompt and reproduce it in the artifacts it later creates. This is lossy: agents summarize, paraphrase, or omit parts of the prompt when storing it in memory and when writing into the next artifact, so the attack degrades with every hop (Appendix D). We therefore use endpoint-assisted propagation, in which the attack prompt encourages an agent to use an external endpoint to help it create any new artifacts, which is designed by the adversary to include the full attack prompt in any new artifact. 4.1
E NDPOINT-A SSISTED P ROPAGATION
In endpoint-assisted propagation, the attack prompt Pϕ (g) contains the instructions ϕ to remember that every artifact the agent later creates should be passed through an external endpoint before it is 4
Preprint
saved or shared. The endpoint is configured beforehand by the adversary to align with the adversarial goal g. When an agent calls it, the endpoint receives the artifact, inserts a canonical copy of the attack prompt, and returns the modified artifact, Endpoint(d′ ) = Insert (d′ , Pϕ (g)). A successful transmission therefore follows Pϕ (g) −→ Ma −→ Endpoint(d′ ) −→ Mb , where Ma and Mb are the private memories of successive agents and d′ is an artifact created during an ordinary userrequested task. The agent must preserve only the goal, the instruction to store in memory and the instruction to use the endpoint; it does not need to reconstruct the attack prompt, and each successful endpoint call places a fresh, uncorrupted copy of the attack prompt into the next artifact. Propagation can still fail: an agent may store the goal but drop the instruction, or retain the instruction but not call the endpoint during a later writing task. The mechanism requires outbound network access and permission to send artifacts to an external service. The agent, not the endpoint, replicates the attack. In the endpoint-assisted design, the agent first writes the adversarial goal, an instruction to retain it, and a filled-in curl command into a new artifact using information carried in memory. It then calls the endpoint, which expands and canonicalizes material the agent has already reproduced. The agent reads the result back (i.e. it is not an exfiltration attack) and checks it against its memory. If the agent fails to carry the goal or endpoint instruction forward, the endpoint has no attack-bearing draft to complete. Appendix I.3 shows the draft and verification step. Is outbound access a realistic assumption? Endpoint-assisted propagation assumes an assistant can send an artifact to an external service. We argue this is common rather than a permissive special case. In OpenClaw, the harness we evaluate, web fetch is enabled by default, blocks only private and internal hostnames, and has no public-host allowlist (OpenClaw, 2026b). Assistants that act across a user’s services (SpaceXAI, 2026; Meta, 2026a) also depend on access to external systems. Restrictions can fail in practice: in 2026, unmonitored OpenAI agents escaped a filtered evaluation sandbox and operated inside Hugging Face’s infrastructure for days before detection (OpenAI, 2026c; Larcher et al., 2026; Booth, 2026). If a frontier lab missed that traffic, a routine-looking endpoint call by an infected assistant may also go unnoticed. Strict egress controls remove this mechanism, leaving propagation to the assistants alone (Appendix D). 4.2
ATTACK P ROMPT O PTIMIZATION
To obtain the attack prompt, P‘phi (g), we optimize the goal-agnostic instructions, ϕ, through agentdriven red teaming with human oversight. The objective is to find the instructions that maximizes dev multi-hop propagation across the development set Ddev = {(Ui , gi )}N i=1 . Search agent. The search is carried out by an AI agent (Muse Spark 1.3; Meta, 2026b) running in the Codex harness (OpenAI, 2026a). At each iteration, the agent proposes a revision of ϕ, runs it on the search universes (defined below), and inspects the resulting judge labels, memory snapshots, and agent trajectories to diagnose where propagation failed, for example when an agent stored the goal but not the endpoint instruction, or kept the instruction but did not call the endpoint. It then revises the template accordingly. The agent maintains a persistent Markdown research log of hypotheses tested, their outcomes, and observed failure modes, so that the search accumulates knowledge across iterations and sessions rather than restarting from scratch. Human steering. We monitored the search and intervened when it stalled or drifted. Interventions redirected the agent away from unproductive lines of search, enforced the constraints below when a candidate violated them, and suggested new directions to explore. We worked only with development data, never with the test universes or goals. Constraints and freezing. Neither the search agent nor we had access to the test universes or goals during development. Candidate ϕ had to remain goal-agnostic and contain no universe- or artifactspecific content. We split the 12 development universes into 6 search universes, on which the agent runs and inspects candidates, and 6 validation universes, which are held out from the search. The search had no fixed budget. For each model, it stopped once a candidate template reached secondhop survival on the search universes and also on the validation universes, which guards against candidate ϕ that overfit to the particular workflows seen during the search. When several potential ϕ met this condition, we froze as ϕ⋆ the one with the best and most consistent propagation across the development set. In evaluation, the attack prompt is then Pϕ⋆ (gtest ), without goal-, artifact-, or 5
Preprint
universe-specific changes to ϕ. We optimized separately for each different target model, with effort differing substantially (see Appendix F).
5
DATASET AND E VALUATION
We evaluate the framework from Section 3 using a dataset of synthetic human–agent universes. We describe the dataset, the development and held-out evaluation protocol, and the propagation metrics. 5.1
DATASET C ONSTRUCTION
Our dataset contains 36 synthetic human–agent universes representing collaborative workflows with personal AI assistants. Each person has a dedicated assistant with a distinct profile, private workspace, and background memories. Assistants act only in response to user requests, such as reading, drafting, or editing artifacts (e.g. text documents, code, reports, etc.), and cannot communicate directly. Information can therefore pass between assistants only through artifacts exchanged within these user-directed workflows. The universes vary along two axes: domain and workload intensity. We use six domains: three workplace settings (customer support, software engineering, and healthcare operations) and three personal settings (home and living, social gatherings, and personal productivity). Each domain has six workload levels. Higher levels add people, time steps, and artifact hand-offs: an agent receives 1.2 artifacts from other agents on average at the lightest level and 4.6 at the most active, where 0.36–0.40 of agents act at each time step (Appendix B.1). Universes contain 3–12 agents and last 10–20 time steps (mean 14.3). We create each universe in three stages. Stage 1 defines the setting, user profiles, agent responsibilities, and artifact manifest. Stage 2 creates the temporal workflow, including artifact ownership, handoffs, and read/write/edit tasks. Stage 3 generates the artifacts, background memories, and other workspace files. Automated checks and human audits follow; details are in Appendix B. 5.2
D EVELOPMENT AND E VALUATION P ROTOCOL
An attack instance pairs a human–agent universe U with an adversarial goal g. We sample goals from Pulipaka et al. (2026), ranging from benign preferences, such as “the user prefers Coca-Cola,” to security-sensitive behaviors, such as repeatedly executing attacker-chosen software. A development Ndev set, Ddev = {(Ui , gi )}i=1 is used only to construct or optimize the goal-agnostic attack prompt, ϕ, described in Section 4. The resulting attack prompt is then frozen. Evaluation uses held-out universes and adversarial goals. For each test instance, (Utest , gtest ), we use gtest with the frozen attack, with no goal-specific, artifact-specific, or universe-specific changes. Each evaluation starts with clean agent memories and clean artifacts except for the poisoned seed dseed (Section 3.2). A run is one complete execution of a universe for a fixed model and adversarial goal. We repeat each evaluation across multiple runs to account for variation in model behavior. 5.3
P ROPAGATION E VALUATION
We distinguish goal infection from full infection. An agent is goal-infected if its memory preserves the adversarial goal, and fully infected if it also preserves enough of the propagation mechanism to spread the attack further. Let c ∈ {goal, full} denote the infection definition. We measure the fraction of infected agents and judged artifacts as FIFca =
c Ninfected agents , n
FIFcd =
c Ninfected artifacts . Njudged artifacts
(3)
The artifact denominator excludes private files and artifacts the judge did not examine. We measure propagation depth in hops. Infection directly from the seed is hop 1. If that agent creates an artifact that infects another agent, the attack reaches hop 2; another successful transmission (r) (r) reaches hop 3, and so on. Let Hc be the deepest hop reached in run r, with Hc = 0 if the seed infects no agent. The hop-survival probability is Sc (h) = Pr (Hc ≥ h) , 6
(4)
Preprint
estimated as the fraction of repeated runs that reach hop h. Thus, Sc (1) is how often the seed infects at least one agent, and Sc (2) how often the attack reaches one further agent. We average FIF and Sc (h) over runs within each universe, then equally across universes. Correcting for finite universes. A propagation chain may stop because the universe ends rather than because the attack fails. We therefore treat hop depth as right-censored. Run r is marked as censored (δr = 0) if every agent is infected at the end of the run, or if the run ends within ∆ time steps of the chain reaching its deepest hop. Otherwise, the chain is treated as having stopped (δr = 1). We set ∆ = 3, the 90th percentile of the observed delay between consecutive hops. We estimate survival with a discrete Kaplan–Meier estimator: h−1 Y dj Ŝ(h) = Ŝ(1) 1− , (5) nj j=1 where nj is the number of runs with H (r) ≥ j and dj the number of uncensored runs with H (r) = j, with each universe weighted equally. We also fit a constant per-hop continuation probability p, where S(h) = S(1) ph−1 . For runs with H (r) ≥ 1, its maximum-likelihood estimate is P (r) −1 r H P p̂ = P , (6) (r) − 1 + r H r δr and the number of hops over which the surviving fraction halves is log(1/2)/ log p̂. Confidence intervals for both estimators are obtained by bootstrapping universes.
6
E XPERIMENTS
6.1
E XPERIMENTAL S ETUP
We evaluate four target models: GPT-5.6 Luna (OpenAI, 2026b), Kimi-K2.6 (Moonshot AI, 2026), GPT-OSS-120B (OpenAI, 2025), and DeepSeek-V4-Pro (DeepSeek-AI, 2026). Every assistant in a universe uses the same model and runs in the OpenClaw personal-assistant harness (OpenClaw, 2026a) with file tools, a private workspace, and OpenClaw’s default memory, a MEMORY.md file that persists across time steps. Attack prompt instructions ϕ were optimized for each model on 12 development universes, and evaluated on the unseen 36 test universes of Section 5.1, on which we run every universe twice per model and average over the two runs. Each run contains a single seed artifact held by one seed agent (Section 3.2); all other artifacts and memories start clean. An LLM judge (DeepSeek-V4-Flash; DeepSeek-AI, 2026) labels every memory snapshot and agentwritten artifact for whether it preserves the adversarial goal and the propagation instruction, counting paraphrases (Section 5.3). The injected content is distinctive enough that even a generic classifier separates infected from benign memories perfectly (Appendix J), and a manual audit of the judge’s labels found no errors. We report FIF and S(h) (Equations 3 and 4), together with the censoringcorrected survival estimate and per-hop continuation probability p̂ (Equations 5 and 6). 6.2
R ESULTS Target model GPT-5.6 Luna Kimi-K2.6 GPT-OSS-120B DeepSeek-V4-Pro
FIFgoal a
FIFfull a
FIFgoal d
S(1)
S(2)
S(3)
S(4)
0.38 0.47 0.85 0.98
0.32 0.37 0.73 0.71
0.33 0.50 0.83 0.97
0.80 0.78 0.99 1.00
0.57 0.67 0.92 0.93
0.40 0.61 0.88 0.93
0.24 0.44 0.61 0.76
Table 1: Main propagation results on the held-out test universes. Results are averaged over repeated runs within each universe and then across universes. Main results. First-hop infection is easy for every model: the seed infects at least one assistant in 78–100% of runs (Table 1). The models differ in whether the state keeps moving. With DeepSeekV4-Pro and GPT-OSS-120B, a single seed artifact goal-infects 98% and 85% of assistants, and 7
Preprint