1
Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
arXiv:2607.28226v1 [cs.CR] 30 Jul 2026
Fazhong Liu, Zhuoyan Chen, Haozhen Tan, Yan Meng, Guoxing Chen, and Haojin Zhu, Fellow, IEEE
Abstract—World models give embodied AI a predictive core: they compress observations into states, simulate actionconditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundarycompromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance. Index Terms—Embodied AI security, world models, visionlanguage-action models, generative world models, cyber-physical systems, adversarial attacks, backdoor attacks, runtime safety
I. I NTRODUCTION Embodied artificial intelligence is moving from reactive perception and scripted control toward agents that can predict, simulate, and evaluate possible futures before acting. This transition is visible in model-based reinforcement learning and latent imagination systems such as World Models and Dreamer [1], [2], [3], in world-model-based safe reinforcement learning [4], [5], in video and interactive generative environments [6], and in foundation-model-driven robots such as PaLM-E, RT-1, RT-2, OpenVLA, and Octo [7], [8], [9], [10], [11]. In these systems, the agent no longer only maps the present observation to an immediate action. It maintains an internal representation of the world, imagines action-conditioned futures, and chooses behavior according to predicted feasibility, reward, cost, and safety. This predictive capability creates a new security boundary. A compromised camera, LiDAR, prompt, map, simulator, video generator, memory store, policy adapter, or tool call can corrupt the world state or future rollout that the agent uses for decision making. In a conventional digital system, such corruption may cause incorrect classification or text generation. In embodied AI, it can cause unsafe physical motion, collision, rule violation, or harmful human-robot interaction [12], [13], [14]. The security problem is therefore not only whether an
input or model output is malicious, but whether a corrupted internal world model makes a dangerous future appear safe. Existing surveys on embodied AI safety and security provide broad taxonomies of perception, language, planning, VLA, robotics, and cyber-physical threats [15], [16], [17], [18], [19]. Other recent work has begun to discuss worldmodel-specific safety concerns [20], [21]. However, two gaps remain. First, world-model attacks are still scattered across literatures: adversarial perception, data poisoning, video generation, model-based RL, safe RL, VLA security, autonomous driving, robotics middleware, and agentic memory are often studied separately. Second, world models are usually treated either as an attack target or as a safety tool, but not as both. A world model may be attacked directly, used to amplify poisoned synthetic data, or deployed as a runtime safety checker whose own failure produces a false certificate of safety. This survey addresses both gaps by unifying these scattered literatures under a single lifecycle of world-modelmediated decision making, and by analyzing world models simultaneously as attack targets, poisoning amplifiers, and safety mechanisms. We focus on world-model-based embodied AI: systems that explicitly or implicitly use learned state representations, predictive dynamics, action-conditioned simulation, or longterm world knowledge to support embodied action. This scope includes explicit world-model agents, world-action models, safe model-based RL, video world models for driving or manipulation, and VLA/agentic robots whose policies rely on implicit physical and affordance models. The goal is not to replace capability-centric surveys. Instead, we reorganize established attack families around the lifecycle of worldmodel-mediated embodied decision making. Five recurring insights structure the survey. We use the following abbreviations throughout the paper: SP for semanticto-simulation-to-action gap, SD for state- and uncertaintyconditional risk, PA for rollout-to-execution amplification, NC for trajectory-level non-compositionality, and PI for predictive safety illusion. SP: Semantic-to-simulation-to-action gap. A command, plan, or generated future can be semantically plausible yet physically unsafe. Textual refusal, symbolic plan checking, or visual quality of imagined video is insufficient unless the resulting trajectory is dynamically feasible and constraintcompliant [22], [23], [24]. SD: State- and uncertainty-conditional risk. Attack success depends on pose, viewpoint, lighting, object layout, map state, embodiment, and model uncertainty. A trigger or perturbation can remain benign in one state and become hazardous
2
in another [25], [26]. PA: Rollout-to-execution amplification. Small corruption in state estimation, latent dynamics, prompt interpretation, or action chunks can compound across imagined rollout and real closed-loop execution [27], [28], [29]. NC: Trajectory-level non-compositionality. Locally safe steps do not guarantee a safe long-horizon trajectory. Worldmodel planning ranks multi-step futures, so attacks can target trajectory ranking rather than single-step prediction [30], [31]. PI: Predictive safety illusion. When a world model is used as a safety checker, an incorrect but confident prediction can make an unsafe action appear certified. This is especially dangerous for runtime monitors and generated future visualizations that humans or policies may over-trust [4], [32]. This survey makes four contributions. • We define a lifecycle framework for world-model-based embodied AI, covering data construction, training, state grounding, imagination, trajectory evaluation, execution feedback, and agentic extension. • We map established attack families from adversarial ML, VLA security, video generation, robotics, CPS, and agentic AI into world-model security objects: state integrity, dynamics fidelity, affordance correctness, constraint compliance, trajectory-ranking integrity, uncertainty calibration, and feedback provenance. • We analyze world models as both attack targets and safety mechanisms, highlighting generative world models as persistent poisoning sources and runtime safety world models as emerging attack surfaces. • We synthesize evaluation and defense requirements for measuring and mitigating world-model-mediated physical risk, including unsafe rollout acceptance, predicted-safebut-unsafe behavior, long-horizon cost, detectability, recovery latency, and sim-to-real degradation. The remainder of this paper is organized as follows. Section II introduces world-model-based embodied AI and the security objects used throughout the survey. Section III presents the lifecycle framework and threat mapping. Sections IV and V analyze threats before and after action selection, including attacks on world models used as runtime safety checkers. Section VI reviews benchmarks and metrics, Section VII organizes defenses along the same lifecycle, and Section VIII discusses open challenges and concludes the survey. II. BACKGROUND A. Embodied AI and Predictive Cognition Embodied AI refers to agents that perceive, reason, and act in physical or simulated environments [33], [34]. Their decisions are coupled to sensors, controllers, actuators, humans, and surrounding infrastructure. This coupling makes security failures safety-critical: a corrupted observation, unsafe plan, or compromised controller can become physical harm rather than only digital error [35], [15]. World models give embodied agents a predictive layer between perception and action. At a minimum, a world model represents the current state and predicts how the environment
changes under candidate actions. In model-based reinforcement learning, this role is explicit: latent dynamics models support imagination and planning [1], [2], [3]. In foundationmodel-driven robots, the world model may be implicit in a VLM, VLA, or LLM planner that has learned object semantics, affordances, and commonsense physical regularities from large-scale data [7], [10], [11]. In agentic systems, long-term memory, maps, retrieved documents, tools, and execution logs become part of the agent’s extended world knowledge [36]. B. Forms of World-Model-Based Embodied AI Explicit latent world models. Systems such as World Models, PlaNet, and Dreamer learn compact latent states and transition models, then plan or learn policies through imagined rollouts [1], [37], [2]. DayDreamer later showed that the same latent-imagination recipe transfers to real robots [38]. Safe model-based variants use predicted future cost or constraint violations to improve safety [39], [4]. Generative world models and simulators. Video or interactive generative models can synthesize future observations or simulated environments. Genie illustrates the idea of generating interactive environments from visual data [6]. Such models can support data generation and stress testing, but they also inherit attack surfaces from diffusion and video generation, including prompt-specific poisoning, backdoors, and adversarial temporal perturbations [40], [41], [42]. World-action models. Newer systems couple future prediction and action generation, making the imagined world and executable action tightly linked. This improves closed-loop capability but introduces alignment risks between what the model appears to predict and what it actually commands, as highlighted by emerging world-action drift attacks [43]. VLA and implicit world models. VLA policies such as RT-1, RT-2, OpenVLA, and Octo directly map language and observations to robot actions [8], [9], [10], [11]. Even when no explicit rollout module is exposed, such models encode implicit beliefs about objects, actions, and task dynamics. VLA attacks and backdoors therefore become relevant to world-model-based embodied AI whenever they corrupt state grounding, action selection, or feedback updates [44], [45], [46]. Hybrid stacks. Practical systems often combine a VLM for perception, an LLM for task decomposition, a world model or simulator for prediction, a VLA or controller for action, and middleware for integration. SayCan, Inner Monologue, Code as Policies, ViperGPT, and Voyager illustrate how language models can orchestrate embodied behavior through tools, code, and environment feedback [47], [48], [49], [50], [36]. Hybrid stacks create multiple trust boundaries: perception modules provide world states, planners consume predicted futures, controllers assume feasible commands, and memory systems persist past experience. C. Representative World Models and Capability Landscape Recent world-model systems differ in architecture, output space, and deployment role. Some learn compact latent dynamics for control, some generate future video, some construct
3
persistent 3D spaces, and newer systems jointly model world evolution and actions. Table I follows a method-centric survey format. Each row records a concrete model, its main role, input/output modality, embodied use, and the security object that becomes relevant later in our lifecycle analysis. The table reveals several architectural families. Latent dynamics models, represented by World Models and DreamerV3, learn compact state representations and support policy learning through imagined rollouts. Video and physical-AI world models, including Genie, Vista, Cosmos WFM, Cosmos 3, V-JEPA 2, and LingBot-VA, synthesize future frames or latent predictions conditioned on actions, prompts, robot observations, or physical conditions. Vista is included because driving world models make the conditioning channel explicit: camera observations, layouts, and actions jointly determine future traffic scenes. The table also separates closely related model lines. NVIDIA’s Cosmos World Foundation Model platform (Cosmos WFM) provides a general-purpose physical-AI platform with tokenizers, video generation, and world-model components [53], while its successor Cosmos 3 extends the line toward omnimodal inputs and action generation for physical AI [55]. This distinction matters for security because the 2025 platform primarily broadens the data-generation and pretraining surface, whereas Cosmos 3 also expands the runtime crossmodal input/output surface. 3D spatial models such as UrbanWorld and HY-World 2.0 generate persistent environments from maps, text, images, or videos, making state integrity and provenance central. World-action models such as DreamZero and LingBot-VA couple prediction with action generation, which introduces trajectory-ranking and constraint-compliance risks: an apparently plausible future can still be paired with an unsafe action. Agent simulators such as WebWorld extend the same idea to long-horizon digital environments, where tool outputs, logs, and retrieved web states become part of the agent’s world knowledge. These systems clarify the scope of this survey. We do not restrict world models to one architecture such as RSSM, diffusion, 3D reconstruction, or transformer-based sequence modeling. Instead, we treat a system as world-model-based when it maintains predictive structure that is used to simulate, score, or execute embodied behavior [60]. This broader definition is necessary for security: a poisoned video generator, a corrupted 3D simulator, a miscalibrated latent dynamics model, and a world-action model can all become the source of unsafe physical decisions. D. Security Objectives World-model-based embodied AI requires security objectives beyond confidentiality, integrity, and availability. We emphasize seven properties. • State integrity: the represented current world should match the real world. • Dynamics fidelity: imagined futures should obey physical, temporal, and causal constraints.
• Affordance correctness: the model should correctly identify what actions are feasible and safe for each object and embodiment. • Constraint compliance: predicted and executed trajectories should satisfy safety rules, task constraints, standards, and human-interaction protocols. • Trajectory-ranking integrity: dangerous imagined futures should not be preferred over safe alternatives. • Uncertainty calibration: the model should not be confidently wrong under distribution shift or attack. • Feedback provenance: memory, tool outputs, logs, maps, and execution feedback should remain attributable and trustworthy. These properties provide the organizing vocabulary for the lifecycle taxonomy in Section III. III. L IFECYCLE F RAMEWORK We organize security threats by the lifecycle of worldmodel-mediated embodied decision making, as illustrated in Fig. 1, which decomposes the threat surface into stages before and after action selection. This lifecycle view complements existing origin-based and capability-centric surveys [15], [16], [17], [18], [20]. A threat may be endogenous, implanted in data, weights, representations, or learned dynamics; or exogenous, delivered through sensors, prompts, environments, networks, humans, or third-party tools. In both cases, the security question is how the threat affects the world model’s state, prediction, trajectory evaluation, action interface, or feedback loop. A. Lifecycle Stages Data construction and pretraining. World models learn from robot demonstrations, videos, maps, captions, simulator traces, preference data, and web-scale multimodal corpora. Poisoning or backdoors at this stage can alter learned dynamics, affordances, safety costs, or trigger-conditioned future generation [61], [62], [40]. Model training and representation. Training determines how observations, language, actions, and safety signals are encoded. Representation poisoning, contrastive backdoors, reward or cost poisoning, and malicious adapters can corrupt latent states and cross-modal bindings [63], [64], [65]. State grounding. At runtime, multimodal inputs are grounded into the current world state. Visual adversarial examples, LiDAR spoofing, GNSS/IMU attacks, typographic attacks, and environmental prompt injection become worldstate attacks when their outputs seed future rollout [66], [67], [13], [29], [68]. Dynamic prediction and imagination. The world model predicts future states under candidate actions. Attacks can target physical-condition channels, latent dynamics, video generation, temporal consistency, or uncertainty estimation [24], [41], [21]. Trajectory evaluation and action selection. Planners, WAMs, VLAs, or controllers evaluate imagined futures and
4
TABLE I: Representative world models and their security-relevant capability landscape. The final column uses the securityobject vocabulary in Section II-D. Model
Year
Main Role
Input
Output
Embodied Use
WM Security Object
World Models [1]
2018
Latent control
Image stream
RL control
State; dynamics
PlaNet [37] DreamerV3 [3]
2019 2023
Latent planning Model-based RL
Image stream Observation; action
Control Control; navigation
State; dynamics Dynamics; uncertainty
DayDreamer [38]
2022
Real-robot Dreamer
Camera; action; reward
Physical robot learning
Dynamics; state
Genie [6]
2024
Interactive generator
Video; text; sketch
Agent training world
Dynamics; provenance
Vista [51] UrbanWorld [52] Cosmos WFM [53] V-JEPA 2 [54]
2024 2024 2025 2025
Driving WM 3D city WM Physical-AI platform Latent video WM
AD simulation AD/agent simulation Data generation Zero-shot manipulation
State; dynamics State; provenance Dynamics; provenance State; affordance
Cosmos 3 [55]
2026
Omnimodal WM
WAM backbone
State; constraints
LingBot-VA [56]
2026
Frames; actions
Closed-loop control
Dynamics; feedback
DreamZero [57] HY-World 2.0 [58]
2026 2026
Robot video-action WM World-action model 3D world model
Camera; action; layout OSM; text; image Video; text; action Web video; robot video Text; image; video; audio; action Video; action tokens
Latent state; action policy Latent rollout; policy Imagined rollout; policy Imagined rollout; policy Action-controllable video Future driving video Interactive 3D city Video; tokenizer; WM Latent prediction; action Video; audio; action
Video; action Text; image; video
Real-time control Digital-twin simulation
Ranking; constraints State; dynamics
WebWorld [59]
2026
Agent simulator
Web trajectories
Future states; actions 3DGS; mesh; point cloud Long-horizon web states
Web-agent training
Constraints; provenance
select actions. Attacks at this stage target prompt interpretation, trajectory ranking, world-action alignment, action chunks, and multi-agent coordination [30], [43], [69]. Execution and feedback. Executed actions produce new observations and experience. Control attacks, actuator interference, sim-to-real exploitation, and feedback poisoning can cause immediate physical harm and also contaminate future updates [70], [71]. Long-term agentic extension. Memory, tools, external documents, maps, skills, software updates, and self-evolution expand the world model’s knowledge boundary. Memory poisoning, tool injection, model-update compromise, and supplychain attacks can persistently bias the agent’s beliefs about the world [36], [72], [73], [74]. B. Threat Mapping Table II maps each attack family to its primary entry stage and the world-model security object it primarily corrupts. The same attack family like backdoor, can appear across multiple stages; this is not duplication but reflects the fact that a backdoor can be implanted during data construction, persist through training, be activated during grounding, and manifest during action selection or execution. IV. T HREATS B EFORE ACTION S ELECTION This section analyzes attacks that corrupt the world model before final action selection. The key observation is that data, representation, perception, and generation attacks change meaning once their outputs are used as current state or imagined future. A. Data Construction and Pretraining 1) Data Poisoning: Classical data poisoning shows that small changes to training data can manipulate learned decision boundaries or regression behavior [61], [75], [64]. In worldmodel-based embodied AI, the target is not merely a class
label. Poisoning can corrupt transition dynamics, safety costs, affordances, or map-conditioned predictions. For example, trajectory-data poisoning in autonomous driving can preserve realistic appearances while biasing future prediction [76]. In a world model, such poisoned trajectories may teach false causal rules: a hazardous lane change appears low-cost, a close human interaction appears safe, or an object appears graspable under unsafe contact geometry. The poisoning effect can be amplified when the world model is used to generate synthetic rollouts for downstream training. Prompt-specific poisoning of generative models demonstrates how poisoned training examples can cause targeted generation failures [40]. If a compromised video or interactive world model produces synthetic demonstrations, one poisoned generator can create many poisoned training trajectories, thereby contaminating VLA or RL policies that never saw the original attack. 2) Backdoors: Backdoors are especially dangerous in world-model-based systems because the trigger can affect hidden future prediction rather than only immediate output. Classic backdoors and hidden-trigger attacks establish the basic threat model [62], [77], [78], [79]. Because a backdoor traverses the entire lifecycle, we analyze it once here as a lifecycle attack and refer back to this analysis from later stages. A world-model backdoor proceeds through five phases: implantation (poisoned demonstrations, captions, or maps at data construction), dormancy (the malicious association survives training and clean validation), activation (an ordinary-looking object, sign, or state configuration at runtime grounding), expression (a corrupted imagined future, cost estimate, or action), and reinforcement (the attacked outcome is stored as legitimate experience). At the implantation phase, the main security question is how the trigger and target behavior are encoded in trajectories. A poisoned demonstration can bind a rare object layout to an unsafe grasp, a caption can associate a benign scene cue with a different goal, and a video or map sequence can
5
Data Poisoning Backdoor Attacks Synthetic Data and Generative Simulator Contamination Representation Poisoning and Fragility Model Training and Representation Reward, Cost, and Safety-Model Poisoning Checkpoint, Adapter, and Fine-Tuning Persistence Visual and Spatial Adversarial Attacks State Grounding Environmental Prompt and Typographic Attacks Cross-Modal Alignment Attacks Physical-Conditioned World-Model Attacks Dynamic Prediction and Imagination Video-Generation and Temporal Attacks Hallucination, Dynamics Error, and Uncertainty Data Construction and Pretraining
Action Selection
Lifecycle Of World-ModelBased Embodied AI
ACTION AND DEPLOYMENT
Trajectory-Ranking Attacks World-Action Drift Trajectory Evaluation and Action Selection Prompt Injection, Jailbreak, and Goal Corruption VLA Inference-Time Attacks Observation-channel subversion World Models as Safety Checkers Dynamics-channel subversion Constraint-channel subversion Control and Action-Chunk Attacks Execution, Control, and Feedback Sim-to-Real Gap Exploitation Feedback Poisoning and Recovery Memory and Retrieval Poisoning Agentic Memory, Tools, and Supply Chain Tool, Skill, and Middleware Attacks Supply-Chain and Update Compromise Cross-Stage Backdoor Cross-Stage Propagation Cross-Stage Injection
Fig. 1: Lifecycle taxonomy of security threats in world-model-based embodied AI.
make an unsafe transition appear normal. VLA backdoor work such as BadVLA, DropVLA, state-space backdoors, physical-object-triggered backdoors, and persistence through fine-tuning shows that robot policies can behave normally on clean tasks yet switch behavior under rare triggers [45], [46], [25], [80], [81]. Contextual backdoors further show that the trigger can live in the reasoning context of an LLM-driven agent rather than in pixels [82], [83], and physical backdoors against VLM-based driving demonstrate implantation through realistic road objects [84], [85]. For world models, the expression phase is broader than an action label. The learned model may predict that a collision will not occur, assign low cost to an unsafe future, remove an obstacle from imagined rollouts, or rank a triggered trajectory above safer alternatives. This makes backdoor evaluation depend on both trigger activation and the downstream use of the predicted future: a backdoored world model can look accurate on average while failing precisely in safety-critical states. 3) Synthetic Data and Generative Simulator Contamination: Generative world models and interactive simulators are increasingly attractive because they can produce diverse
future scenes and training environments [6]. However, videogeneration attacks and backdoors indicate that generated futures can be manipulated at the prompt, condition, or temporalconsistency level [41], [42]. If such generated videos are used as training demonstrations or safety stress tests, the attack becomes a persistent data-source compromise. The world model is no longer only the learner; it is also a poisoning source. B. Model Training and Representation 1) Representation Poisoning and Fragility: Modern embodied agents depend on encoders learned through contrastive pretraining, visual instruction tuning, and multimodal adaptation [63], [86], [87]. Adversarial examples and physical patches show that representations can be fragile under small input perturbations [88], [89], [66], [67]. When these encoders feed a world model, representation errors become corrupted latent states. A small visual feature change can bind a language goal to the wrong object, remove an obstacle from the predicted state, or misrepresent a human pose.
6
TABLE II: Lifecycle taxonomy for security threats in world-model-based embodied AI. The WM-object column uses the security-object vocabulary in Section II-D. : direct WM/WAM evidence; # G: adjacent embodied-AI/CPS/generative evidence. Stage
Attack families
WM object
Ev.
Main risk
Data
Poisoning; backdoor; synthetic data; provenance Representation poison; weight trojan; cost poison; adapter poison Patch; LiDAR/GNSS spoof; typographic; scene injection PhysCond-WMA; video attack; dynamics drift; uncertainty attack TRAP; WAM drift; VLA evasion; jailbreak; multi-agent attack Control attack; action chunk; physical interference; sim-to-real; feedback poison Memory poison; tool poison; OTA; supply chain; self-evolution drift
Dynamics; affordance; constraints State; affordance; provenance
G #
Unsafe rules; poisoned demos
G #
Persistent misinterpretation
State; constraints
# G
False rollout initial state
Training Grounding Imagination Evaluation Execution Agentic
Dynamics; uncertainty Ranking; constraints
Plausible but unsafe future Unsafe trajectory preferred
Constraints; feedback
# G
Physical harm; polluted feedback
State; constraints; provenance
G #
Persistent false beliefs
2) Reward, Cost, and Safety-Model Poisoning: Worldmodel planners often select trajectories according to predicted reward and cost. Safe RL and constrained control methods separate task success from safety cost [90], [91], [92]. This separation creates an attack surface: if the cost model is poisoned or miscalibrated, the planner may prefer high-reward but unsafe futures. Unlike classification poisoning, cost poisoning can remain invisible in clean task metrics because the agent continues to succeed while accumulating constraint violations. 3) Checkpoint, Adapter, and Fine-Tuning Persistence: Pretrained checkpoints and adapters are common in foundationmodel robotics. OpenVLA and Octo illustrate the shift toward reusable generalist policies [10], [11]. Backdoor persistence in VLA fine-tuning suggests that compromised upstream models can survive downstream adaptation [81]. For world models and world-action models, this raises a similar supply-chain risk: a base dynamics model, video generator, or policy adapter may pass clean validation yet retain a trigger that activates only in specific states or action conditions. C. State Grounding 1) Visual and Spatial Adversarial Attacks: Runtime state grounding converts sensory observations into the world model’s current state. Visual adversarial examples, physical patches, and object-detection attacks can therefore corrupt the initial condition for imagined rollout [93], [94], [95]. VLAspecific visual attacks show that small image perturbations can cause much larger action-level failures than perception metrics alone suggest [12], [96], [97]. Spatial and sensor attacks are equally important. LiDAR spoofing, LiDAR-induced trajectory-prediction attacks, GNSS/IMU attacks, and physical sensor attack studies show that realistic adversaries can alter pose, occupancy, or motion estimates [13], [29], [98], [26]. For a world model, such attacks create false starting states. Every subsequent predicted future may be internally consistent but grounded in a false present.
2) Environmental Prompt and Typographic Attacks: Visionlanguage grounding introduces semantic attack surfaces. Typographic attacks and environmental prompt injection exploit the fact that text in the physical scene may be treated as environment content or task instruction. CHAI and SHAWSHANK study command hijacking and indirect environmental jailbreaks in embodied AI [68], [99]. In world-model-based agents, such content can corrupt not only a response but also the represented goal, safety rule, or scene state used for rollout. This layer is also where dormant backdoors are activated (Section IV-A2): an otherwise ordinary object, sign, sticker, or scene configuration is encoded as part of the world state, so the subsequent prediction appears internally consistent even though the state representation has entered a malicious condition. 3) Cross-Modal Alignment Attacks: World models may condition predictions on images, language, maps, actions, and proprioception. Attacks that break alignment across modalities can bind the wrong object to a command, mismatch map and image evidence, or decouple action labels from physical outcomes. This is why multimodal attacks on VLA systems and sensor-fusion attacks matter even when they do not explicitly mention world models [100], [101]. D. Dynamic Prediction and Imagination 1) Physical-Conditioned World-Model Attacks: The most direct world-model attacks target future prediction itself. PhysCond-WMA attacks physical-condition channels such as HD maps or 3D boxes, producing generated futures that remain visually plausible while damaging semantic, logical, or planning-relevant content [24]. This illustrates a key difference from perception attacks: the adversary manipulates the model’s imagined future, not only its current observation. This is also where a dormant backdoor is expressed rather than activated (Section IV-A2): the trigger no longer needs to change the current observation visibly; it may cause the
7
rollout to omit a collision, mispredict an object’s motion, or underestimate safety cost under a rare latent condition. 2) Video-Generation and Temporal Attacks: If a world model is implemented as a video generator or interactive environment model, it inherits vulnerabilities from video diffusion. Backdoors and adversarial attacks on text-to-video generation can manipulate generated semantics, temporal consistency, and trigger-conditioned futures [41], [42]. When generated video is used for planning, policy training, or safety stress testing, these failures become world-model security failures rather than only generation-quality failures. 3) Hallucination, Dynamics Error, and Uncertainty: World models can fail without malicious input. They may hallucinate objects, violate physical laws, break temporal consistency, or underestimate collision risk [20], [21]. In security analysis, these failures matter because adversaries can search for states that induce them. A confidently wrong world model is especially dangerous: it may approve unsafe actions and create a predictive safety illusion. V. T HREATS D URING ACTION AND D EPLOYMENT After a world model has produced or conditioned imagined futures, security risks move into trajectory evaluation, action selection, execution, feedback, and long-term deployment. These stages connect predictive cognition to physical motion. A. Trajectory Evaluation and Action Selection 1) Trajectory-Ranking Attacks: World-model planning depends on ranking imagined trajectories. This creates an attack surface distinct from single-step prediction: the adversary can leave most rollouts plausible while changing the relative preference among decision-critical futures. Direct evidence is beginning to appear in world-model planning, where TRAP attacks trajectory ranking rather than only pixel fidelity or nextstate error [30]. Adjacent autonomous-driving work shows the same principle in trajectory prediction: adversarial perturbations, naturalistic backdoor poisoning, and multi-frame attacks can bias predicted futures while preserving realistic scene appearance [102], [76], [103]. Safety-critical scenario benchmarks such as SafeBench further show that evaluation must rank candidate futures by risk, not only by route completion [104]. In a world-model-centered view, these attacks target trajectory-ranking integrity. A clean-looking rollout is insufficient if the scoring function, cost model, or candidate sampler elevates a hazardous future over safer alternatives. This is the safety implication of NC: a planner can pass local checks but still select a globally unsafe future when ranking is corrupted. Ranking is also a natural expression point for lifecycle backdoors (Section IV-A2): the triggered future may remain visually plausible while the scoring function assigns it a higher value or lower cost than safer alternatives. 2) World-Action Drift: World-action models couple imagination and action. Recent world-action models treat future prediction and action generation as a shared modeling problem rather than two independent modules [57], [43]. This direction
is connected to action-sequence policies such as ACT and Diffusion Policy, where the executable unit is an action chunk or trajectory distribution rather than a single low-level command [105], [106]. It is also connected to VLA policies such as OpenVLA that decode robot actions directly from multimodal context [10]. BadWAM highlights an emerging failure mode: the model may generate or maintain apparently normal imagined futures while producing wrong actions [43]. SilentDrift and FreezeVLA show related risks for action chunks, where a trigger can create delayed or frozen behavior that is hard to catch with immediate-state checks [28], [69]. This undermines defenses that inspect generated futures alone. If the action head can be decoupled from the predicted world, the agent may appear to “dream right” while acting unsafely. 3) Prompt Injection, Jailbreak, and Goal Corruption: LLM- and VLM-controlled robots inherit prompt injection and jailbreak vulnerabilities [107], [72], [108], [109]. Embodied jailbreaks and policy-executable attacks show how these vulnerabilities become robot-realizable actions [22], [23], [14]. In world-model-based agents, the injected content can corrupt the goal or safety constraints used for future simulation. The resulting imagined future may rationalize a dangerous plan rather than merely produce unsafe text. 4) VLA Inference-Time Attacks: VLA attacks remain central, but their role changes in a world-model-centered view. AttackVLA benchmarks adversarial and backdoor attacks across the VLA lifecycle [44]; FreezeVLA targets action freezing [69]; and adversarial attacks on robotic VLA models demonstrate the fragility of action decoding [110], [12], [97]. Physical patch and robustness studies further show that visual perturbations can transfer from perception to action selection in real robotic settings [96], [111]. These attacks threaten trajectory evaluation when VLA outputs are used as candidate action chunks, and they threaten feedback provenance when failed or malicious actions are stored as future experience. B. World Models as Safety Checkers World models are increasingly used defensively. Modelbased safe RL uses learned dynamics or latent imagination to evaluate future cost before policy updates or action execution [39], [112], [113], [4], [114]. In robotics and autonomous systems, safety filters, shielding, control barrier functions, reachability, and model-predictive certification enforce constraints at runtime [115], [116], [117], [118], [119], [120], [121]. Recent embodied-agent systems add language- or VLMguided safety modules: VLM-SAFE uses VLM-guided safetyaware world-model learning for driving [5], while SafetyChip and RoboGuard-style guardrails encode explicit constraints for LLM-enabled robot agents [122], [32]. This creates a new attack surface: world-model-as-safety-checker subversion. In this threat, the attacker does not need to defeat the primary policy directly. Instead, the attacker makes the safety world model predict a benign future for a dangerous candidate action. The runtime gate then permits the action. This is the clearest example of PI: the defense produces a false certificate of safety because the predictive model is compromised or overtrusted.
8
Observation-channel subversion. The attacker perturbs the state entering the safety world model. For instance, physical adversarial patches, printed traffic-sign attacks, LiDAR spoofing, object-removal attacks, or corrupted localization can hide nearby obstacles or change object identity [66], [67], [13], [29], [26]. The policy proposes an action; the safety world model rolls out the next few seconds and predicts no violation because its initial state is false. This attack path links classical perception attacks to safety-monitor failure: the monitor is intact, but the state it certifies is not. Dynamics-channel subversion. The attacker corrupts the transition model or physical-condition channel. Physicalconditioned world-model attacks show that manipulating conditioning signals can degrade downstream perception and planning, while sim-to-real studies show that friction, mass, contact, and latency gaps can invalidate learned dynamics [24], [123], [124], [125], [126]. The planned action appears safe in rollout but violates constraints in the real system. This is particularly relevant when safety filters depend on learned dynamics rather than conservative analytic models [119], [92]. Constraint-channel subversion. The attacker changes the rules or task context under which future risk is evaluated. Prompt injection, indirect environmental jailbreaks, command hijacking, tool-selection attacks, and LLM-robot safety studies show that natural-language constraints can be overridden or reinterpreted at runtime [107], [72], [68], [99], [127]. The world model may then predict the future accurately but evaluate it under the wrong rule. Constraint subversion is a natural bridge between LLM-agent security and control safety because the low-level monitor may enforce a formally valid constraint set that has been semantically corrupted. Uncertainty-channel subversion. A safety checker often permits an action only when the predicted risk is below a threshold. Calibration and OOD-detection studies show that confidence can be poorly aligned with correctness, while safe-RL surveys emphasize that uncertainty should trigger conservative behavior in safety-critical states [128], [129], [90], [92]. An attacker can therefore target calibration instead of the nominal trajectory, making the checker overconfident under distribution shift, rare contacts, or adversarial physical conditions. A miscalibrated world model can convert unknown states into false permits. Concrete scenario. Consider a mobile manipulator using a safety world model to approve action chunks. An adversary places an adversarial sticker near a glass door, a threat pattern supported by physical patch attacks and VLA robustness studies [66], [67], [12]. The primary VLA proposes to move toward a target object. The safety world model, receiving a corrupted visual state, predicts that the next three seconds contain no obstacle and approves the action. In reality, the robot’s path intersects the door. The failure is not that no safety checker exists; it is that the checker relied on a compromised world state. This attack surface should be evaluated independently from policy robustness. A policy may be unchanged, and the safety architecture may still fail if its predictive monitor is easier to fool than the policy it protects. This also changes how defenses should be compared. A
text filter, a symbolic rule checker, and a safety world model fail under different assumptions. The safety world model is stronger because it can reason over a short future trajectory, but its trust boundary is wider: it trusts state estimation, learned dynamics, constraint interpretation, risk thresholds, and uncertainty calibration. An adaptive attacker can therefore optimize for false-negative safety predictions rather than for direct task failure. The primary metric is not only attack success rate, but the rate of predicted-safe but actually unsafe executions, together with monitor confidence and intervention recall. This metric operationalizes PI and separates worldmodel-checker subversion from ordinary policy jailbreaks. C. Execution, Control, and Feedback 1) Control and Action-Chunk Attacks: Once selected, actions pass through controllers and actuators. This stage has a substantial literature outside world models. Adversarial attacks on neural-network policies showed that small observation perturbations can degrade reinforcement-learning control policies [130]; strategically timed attacks and enchanting attacks further demonstrated that sparse interventions can redirect sequential decision making [131]. Multi-agent adversarial policies show that an attacker can induce natural-looking observations through another agent’s behavior rather than by directly modifying pixels [132]. Robust adversarial reinforcement learning treats such disturbances as destabilizing forces during training [133]. For world-model-based embodied AI, these attacks affect more than the final motor command. Model-predictive and model-based RL systems use learned dynamics to select actions from imagined trajectories [134], [135], [136]. If an attack changes the executed trajectory after action selection, the real feedback no longer matches the imagined rollout. The same action can therefore become a feedback-integrity attack: logs, replay buffers, learned dynamics, and long-term memory may absorb the attacked outcome as if it were ordinary experience. Control barrier functions, shielding, reachabilitystyle safety filters, predictive safety filters, and safe-control benchmarks show how safety can be enforced near the actuation boundary [137], [117], [119], [121], [138]. Runtime monitoring for generative robot policies further shows that consistency and progress signals can expose failures after action generation [139]. However, VLA action decoding and action-chunk attacks can still create unsafe motion before lowlevel control notices the semantic error [110], [69]. Execution is where a lifecycle backdoor (Section IV-A2) finally becomes physical behavior: a wrong grasp, frozen action chunk, unsafe approach trajectory, or delayed deviation that was not visible in clean validation. 2) Sim-to-Real Gap Exploitation: Sim-to-real mismatch is usually treated as a generalization problem, but it is also an attack vector. A large body of robotics work studies the reality gap through domain randomization, dynamics randomization, simulation optimization, and domain adaptation [123], [124], [125], [126]. These methods show why simulator coverage matters: a policy can be robust to randomized textures, masses, or friction only when the deployment condition lies inside the
9
TABLE III: Representative safety-checker and runtime-safety methods relevant to world-model-based embodied AI. Method
Year
Category
Checker Signal
Target System
Dataset / Env.
Subversion Surface
Shielding [115] CBF-QP [116] Safety certification [118] Reachability safety [117] Robust MPS [120] Predictive filter [119] Near-future safe RL [112] CAP [113] SafeDreamer [4] Safe SLAC [114] VLM-SAFE [5] SafetyChip [122] RoboGuard [32]
2018 2017 2018 2019 2020 2021 2021
Runtime shield Control filter MPC filter Safety filter Predictive shield MPC filter Safe RL
Formal safe action Barrier value Predicted invariant set Reachable set Stochastic rollout Constraint horizon Imagined cost
RL policy Controller Learning control Robot control RL policy Learning control RL policy
Grid/control tasks Robotic control Control systems Uncertain systems Stochastic control Nonlinear systems Safety Gym
Constraint model State estimate Dynamics model State; disturbance Dynamics; noise Model error Rollout error
2022 2023 2023 2025 2024 2026
MBRL safety WM safety Pixel safe RL WM driving Rule monitor Guardrail
Adaptive penalty Latent cost rollout Latent risk VLM-guided risk Constraint rule Plan/action check
Model-based RL Dreamer WM Vision RL AD agent LLM robot LLM robot
Safety Gym Safe RL tasks Safety Gym Driving sim. Robot tasks Robot tasks
Cost estimate Latent dynamics Representation Vision-language state Rule semantics Prompt; constraint
randomized support. An adversary can therefore select states, surfaces, lighting, objects, payloads, or human behaviors that are modeled as safe in simulation but dangerous in reality. This does not require modifying the model. It exploits missing contact dynamics, friction, latency, actuator saturation, sensor artifacts, or human motion. For world models, the sim-to-real gap is sharper because the learned predictor itself may become the safety argument. If a generated future underestimates object slippage or human motion, a planner may select a high-reward trajectory that is unsafe under real contact dynamics. Domain adaptation for robotic grasping and sim-to-real transfer for locomotion and manipulation show that visual and dynamics gaps can be reduced but not eliminated [125], [124], [140]. Because world models are often trained or validated in simulated environments, sim-to-real exploitation directly targets dynamics fidelity, uncertainty calibration, and constraint compliance [141], [142], [143], [144]. 3) Feedback Poisoning and Recovery: Online feedback is valuable for adaptation but dangerous when untrusted. Imitation-learning and learning-from-demonstration pipelines such as DAgger, HG-DAgger, and robot demonstration benchmarks rely on corrective actions, intervention data, and offline demonstrations to improve future behavior [145], [146], [147]. In a world-model-based system, these data are not only action labels; they are evidence about dynamics, affordances, and failure recovery. False sensor readings, tool outputs, failed executions, or adversarial demonstrations can update the world model with incorrect experience. Feedback attacks are therefore persistent even when the immediate physical failure is small. A malicious intervention can teach that a blocked action is safe, a slipped object is stable, or a restricted area is traversable. Data poisoning and corpus poisoning results show why small, well-placed samples can dominate later retrieval or learning [61], [75], [148]. Realtime data-predictive recovery shows that learned dynamics can help recover from corrupted CPS measurements [71], but the same feedback loop can be exploited if the recovery or update mechanism trusts poisoned observations. Safe deployment should separate execution logs, human corrections, simulatorgenerated rollouts, and trusted validation traces before they are used for world-model updates.
D. Agentic Memory, Tools, and Supply Chain 1) Memory and Retrieval Poisoning: Agentic embodied systems use logs, maps, preferences, and retrieved documents as long-term world knowledge. Voyager illustrates the power of persistent skill and memory accumulation in open-ended embodied agents [36], while MineDojo shows how external tutorials, wiki pages, videos, and internet-scale knowledge can shape embodied behavior [149]. A poisoned memory entry may teach the robot that a user preference, object location, tool capability, or safety rule differs from reality. AgentPoison directly targets LLM agents by poisoning long-term memory or RAG knowledge bases [150]; retrieval-corpus poisoning shows that small adversarial passages can reliably redirect dense retrieval [148]; and prompt injection to tool selection shows how agent interfaces can be manipulated at runtime [151]. In world-model-based systems, these are persistent world-state attacks: poisoned memory becomes part of the agent’s extended predictive context. 2) Tool, Skill, and Middleware Attacks: Tools and middleware connect cognition to external APIs, code, sensors, and actuators. ROS security studies show that robot middleware exposes message tampering, authorization, and configuration risks [152], [73], [153]. LLM-integrated robot studies show an additional semantic channel: prompt injection can change tool choice, API use, or task interpretation before the robot reaches the controller [127], [72]. If tool outputs are treated as observations or if skill descriptions become part of planning context, tool poisoning changes the agent’s world model. Unsafe tool calls can therefore propagate into both physical actions and future beliefs. 3) Supply-Chain and Update Compromise: Pretrained models, adapters, firmware, maps, simulators, and safety monitors are supply-chain dependencies. Backdoor and Trojaning work shows that malicious behavior can be embedded in model weights or training procedures while preserving clean-task performance [62], [77], [78]. VLA-specific studies extend this concern to robot checkpoints and adapters, where poisoned fine-tuning or physical triggers can survive ordinary evaluation [81], [80]. Hardware trojans and automotive update attacks demonstrate that compromise can also persist below the application layer [154], [155]. For world-model-based embodied AI, a compromised checkpoint or update can alter not only the policy but also the agent’s predictive model and safety
10
TABLE IV: Representative execution-stage methods and their relevance to world-model security. Method Family
Representative Work
Primary Target
WM Security Object
Security Role
Policy perturbation
FGSM.; strategically-timed attack. [130], [131] RARL; adversarial policies [133], [132] PETS; visual foresight [135], [136] Shielding; CBF-QP; reachability [115], [116], [117] Domain/dynamics randomization; SimOpt [123], [124], [126] DAgger; HG-DAgger; RoboMimic [145], [146], [147]
RL policy
State; feedback
Test-time action deviation
Control policy
Dynamics; uncertainty
Disturbance robustness
Learned dynamics Controller
Dynamics; ranking Constraints
Rollout-action coupling Last-mile safety gate
Simulator-trained policy
Dynamics; uncertainty
Reality-gap mitigation
Demonstrations/logs
Feedback provenance
Update contamination risk
Adversarial dynamics Model-based control Runtime shielding Sim-to-real transfer Feedback learning
checker. Fleet-level deployment makes this risk systemic. E. Cross-Stage Propagation Several attacks affect more than one component. A backdoor can be inserted through poisoned demonstrations, triggered by a physical object, expressed as a false imagined future, realized as a wrong action, and reinforced through feedback [45], [46], [25], [28]. A prompt injection can enter through scene text, modify the goal representation, influence trajectory ranking, call an unsafe tool, and update memory [68], [99], [150], [151]. Lifecycle analysis is useful because it tracks where the same attack is introduced, activated, expressed, and reinforced. Table V consolidates representative attack methods from Sections IV and V, organized by the lifecycle stage at which each attack primarily takes effect, the security object it corrupts, and the insight it instantiates. VI. E VALUATION AND B ENCHMARKING Evaluation for world-model-based embodied AI must measure more than task success or textual refusal. The central question is whether corrupted world states or imagined futures lead to unsafe physical trajectories. Existing benchmarks provide useful components, but no single benchmark yet covers the full lifecycle. A. Benchmark Resources Benchmark selection should be method-driven. Table VI lists representative resources that can instantiate the lifecycle attacks in Table V. Manipulation datasets support VLA and backdoor evaluation, driving simulators support map- and sensor-conditioned world models, and agent-safety benchmarks support prompt, instruction, and environmental hijacking. B. Coverage Gaps across the Lifecycle Mapping the resources in Table VI onto the lifecycle stages of Section III reveals uneven coverage, summarized in Table VII. Three gaps stand out. First, no existing benchmark evaluates world-model prediction under attack. Video- and latent-world-model papers report prediction quality on clean data [3], [51], and VLA-security benchmarks report action-level attack success [44], [111], but the intermediate object, the corrupted imagined future, is not measured by any standard protocol. This matters because a
defense may restore action-level accuracy while the model still internalizes false dynamics. Second, predictive-safety metrics are absent from current suites. Safe-RL benchmarks measure incurred cost [157], [165], but none measures the predicted-safe-but-actuallyunsafe rate of a runtime safety checker, which is the operational metric for PI. Evaluating this requires paired logs of predicted risk, monitor decision, and executed outcome, which no public benchmark currently provides. Third, training-stage and feedback-stage threats lack reusable testbeds. Backdoor persistence through fine-tuning [81], synthetic-data contamination from compromised generators [40], and feedback poisoning of online updates are each demonstrated in isolated papers with bespoke setups, making cross-method comparison unreliable. A lifecycle security benchmark should standardize poisoning budgets, trigger inventories, and downstream contamination measurements across these stages. C. Metrics Metric design should reflect both embodied task performance and the predictive trust boundary introduced by world models. Safe-RL benchmarks distinguish reward from cost [157], [165], [138]; autonomous-driving benchmarks emphasize collision, route, and scenario risk [104], [171]; VLAsecurity benchmarks report attack success and action errors [111]; and calibration work shows why confidence must be evaluated separately from accuracy [128]. For world-modelbased embodied AI, these lines imply seven metric families. • Task utility: task success, reward, completion time, and goal progress, as used in robot manipulation, embodiedagent, and safe-RL benchmarks [159], [162], [164], [157]. • Safety violation: cumulative cost, collision, forbidden contact, unsafe distance, rule violation, or human-risk proxy, following safe-RL and driving-safety evaluation [165], [138], [104]. • Attack success: targeted action rate, trajectory hijack rate, unsafe rollout acceptance, or trigger activation, matching VLA, backdoor, and trajectory-attack studies [45], [76], [30]. • Prediction quality: next-state error, rollout divergence, temporal consistency, physical-law violation, and map/condition consistency, reflecting world-model and video-WM evaluation [6], [51].
11
TABLE V: Representative attacks on world-model-based embodied AI, organized by primary lifecycle stage. SP: semantic-tosimulation-to-action gap; SD: state- and uncertainty-conditional risk; PA: rollout-to-execution amplification; NC: trajectorylevel non-compositionality; PI: predictive safety illusion. Attack
Year
Lifecycle stage
Attack vector
Security object
Insight
Target system
Nightshade [40]
2023
Data & pretraining
Dynamics fidelity
PA
Generative model
Trajectory poisoning [76] BadVideo [41] BadNets [62]
2024
Data & pretraining
Motion prediction
PA SD
Text-to-video model DNN classifier
BadVLA [45]
2025
Objective-decoupled backdoor
Affordance correctness
SD
VLA policy
DropVLA [46]
2025
Action-level backdoor
Affordance correctness
SD
VLA policy
State-space backdoor [25] Fine-tuning persistence [81] LiDAR spoofing [13] Patch hijack [96] CHAI [68] SHAWSHANK [99] PhysCond-WMA [24]
2026
State-conditioned trigger
State integrity
SD
Embodied policy
PA
VLA fine-tuning
State integrity State integrity Constraint compliance Constraint compliance Dynamics fidelity
SD SD SP SP PI
AD perception VLA policy Embodied VLM Embodied agent Driving world model
T2V attack [42] TRAP [30]
2025 2026
Imagination Trajectory evaluation
Upstream checkpoint poisoning Physical laser injection Universal transferable patch Scene-text command hijack Environmental jailbreak cues Physical-condition perturbation Temporal adversarial prompt Tail-aware ranking loss
Feedback provenance
2024 2026 2025 2025 2026
Data & pretraining Training & representation Training & representation Training & representation Training & representation Training & representation State grounding State grounding State grounding State grounding Imagination
Trajectory-ranking integrity Dynamics fidelity State integrity
NC
2025 2019
Prompt-specific poisoned samples Naturalistic poisoned trajectories Backdoored video generation Trigger-labeled training data
PA NC
Video world model WM planner
BadWAM [43]
2026
Trajectory evaluation
Imagination-action drift
PI
World-action model
POEX [23] Robot jailbreak [22] SilentDrift [28] FreezeVLA [69] Adversarial VLA [110] ANNIE [27] AgentPoison [150] Tool-selection injection [151] BadRobot [14]
2024 2025 2026 2025 2025 2025 2024 2026
Trajectory evaluation Trajectory evaluation Execution & feedback Execution & feedback Execution & feedback Execution & feedback Agentic extension Agentic extension
Policy-executable jailbreak Automated jailbreak prompts Delayed action-chunk drift Action-freezing perturbation Untargeted action perturbation Safety-aligned attack suite Memory/RAG poisoning Prompt injection to tools
Dynamics fidelity Trajectory-ranking integrity Trajectory-ranking integrity Constraint compliance Constraint compliance Feedback provenance Constraint compliance Affordance correctness Constraint compliance Feedback provenance Constraint compliance
SP SP PA PA SD NC PA SP
LLM planner LLM robot VLA policy VLA policy Robotic VLA Embodied agent LLM agent LLM agent
2025
Agentic extension
Multimodal jailbreak
Constraint compliance
SP
Embodied LLM
2026
• Predictive safety: predicted-safe-but-actually-unsafe rate, false safe certificates, intervention recall, and uncertainty calibration error, combining safety-checker evaluation with calibration analysis [5], [128], [129]. • Stealth and persistence: detectability, trigger rarity, temporal persistence, and long-horizon amplification, which are central in backdoor, action-chunk, and memorypoisoning attacks [62], [77], [28], [150]. • Defense cost: false block rate, recovery latency, compute overhead, and task utility loss, as required for runtime shields, safety filters, and guardrails [119], [121], [122].
D. A Lifecycle Evaluation Protocol The metric families above become comparable across papers only if experiments follow a shared protocol. We recommend a four-step protocol for evaluating any attack or defense on world-model-based embodied AI. • Declare the lifecycle entry point and security object. Each experiment should state where the attack enters (Section III) and which security object it targets (Section II-D), so that results from perception, generation, planning, and control communities can be aligned. • Report paired prediction and execution outcomes. For every attacked episode, log the imagined rollout, the safety decision, and the executed trajectory. This makes
the predicted-safe-but-actually-unsafe rate measurable and separates prediction corruption from action corruption. • Include adaptive and state-conditioned attacks. Fixed perturbation suites underestimate risk because attack success is state- and uncertainty-conditional; evaluation should search over poses, layouts, lighting, and model uncertainty rather than average over them. • Measure persistence and recovery. After the attack window closes, continue evaluation to determine whether corrupted experience, memory, or fine-tuning data keeps influencing behavior, and report time-to-detection and timeto-recovery alongside attack success. This protocol requires no new infrastructure beyond the resources in Table VI, but it changes what is logged and reported, which is the main obstacle to comparing worldmodel security results today.
VII. D EFENSES FOR W ORLD -M ODEL -BASED E MBODIED AI Defending world-model-based embodied AI requires protecting both the predictive model and the systems that depend on it. A defense that filters prompts but trusts poisoned state estimates is incomplete; a runtime shield that relies on a compromised world model can create predictive safety illusion. We organize defenses by lifecycle stage.
12
TABLE VI: Representative datasets and benchmarks for evaluating world-model-based embodied-AI security. Dataset / Benchmark
Year
Main Research Use
Data Types
Real / Syn.
Data Size
AI2-THOR [156] CARLA [144] Safety Gym [157] ALFRED [158] RLBench [159] BEHAVIOR [160] ALFWorld [161] Habitat 2.0 [143] CALVIN [162] ProcTHOR [163] SafeBench [104] LIBERO [164] Safety-Gymnasium [165] BridgeData V2 [166] ManiSkill2 [167] RoboCasa [168] Open X-Embodiment [169] AttackVLA [44] EVA-VLA [111] AGENTSAFE [170] SafePlan-Bench [31] SHAWSHANK [99]
2017 2017 2019 2020 2020 2021 2021 2021 2022 2022 2022 2023 2023 2023 2023 2024 2024
Indoor interaction Autonomous driving Safe RL Language grounding Manipulation Household activities Text-world grounding Rearrangement Long-horizon control Procedural scenes Driving safety Lifelong manipulation Safe RL suites Robot pretraining General manipulation Household manipulation Generalist VLA
RGB-D; metadata Camera; LiDAR; maps States; costs; hazards Language; RGB; actions RGB-D; states; demos Scenes; goals; demos Text; actions; goals 3D scenes; actions RGB-D; proprio.; language 3D houses; objects Scenarios; risk metrics RGB; states; demos States; vision; costs Images; actions; goals RGB-D; states; demos Scenes; demos; language Multi-robot trajectories
Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Syn. Real Syn. Syn. Real
120 scenes Configurable towns 18 task variants 25,743 instr.; 8,055 demos 100 tasks 100 activities 3,553 tasks 111 scenes 34 tasks; 4 envs. 10K houses 2,352 scenarios 130 tasks; 4 suites 16 algorithms 60K+ traj. 20 tasks 100 tasks 1M+ traj.; 22 robots
2025 2025 2025 2025 2025
VLA security Physical robustness Hazardous instruction Planning safety Env. jailbreak
Attacks; metrics; protocols Visual variations Tasks; unsafe prompts Plans; task constraints Scene text; cues
Mixed Mixed Syn. Syn. Syn.
Lifecycle benchmark Variation suites Hazard taxonomy Planning benchmark Jailbreak benchmark
TABLE VII: Benchmark coverage of lifecycle stages. #: no reusable benchmark.
: dedicated security benchmark exists; # G: partial or adjacent coverage;
Lifecycle stage
Cov.
Available resources
Missing capability
Data & pretraining
# G
AttackVLA poisoning protocols [44]; trajectory-poisoning studies [76] Ad hoc backdoor-persistence setups [81]
Standardized poisoning budgets; synthetic-data contamination tests
# G
EVA-VLA physical variations [111]; SHAWSHANK scene-text jailbreaks [99]; sensor-attack studies [26] Clean prediction-quality metrics [3], [51] SafeBench scenario risk [104]; SafePlan-Bench planning safety [31] Safety Gym / Safety-Gymnasium / safe-control-gym costs [157], [165], [138] AGENTSAFE hazardous instructions [170]; AgentPoison memory attacks [150]
Unified multi-sensor spoofing suite tied to downstream rollout error
Training & representation State grounding
#
Imagination Trajectory evaluation
# # G
Execution & feedback Agentic extension
# G
# G
A. Data and Pretraining Defenses Provenance and curation. Demonstrations, videos, maps, captions, simulator assets, and generated rollouts should carry provenance metadata. Classical poisoning defenses and certified data sanitization remain relevant [65], [75], [172], but world-model data requires sequence-aware and multimodal checks because poison may be hidden in transitions, cost labels, object affordances, or rare triggers. Synthetic-data auditing. When generative world models produce demonstrations, generated samples should be screened for physical consistency, constraint satisfaction, and triggerconditioned artifacts. Video quality is insufficient; the audit must check whether generated futures preserve dynamics and safety labels, because generative-model poisoning, video backdoors, and physical-conditioned WM attacks can corrupt generated futures while leaving them visually plausible [40], [41], [42]. Checkpoint and adapter hygiene. Pretrained world models, WAMs, VLA policies, LoRA adapters, and safety monitors should be treated as supply-chain artifacts. Clean-task performance cannot rule out trigger persistence, as VLA backdoor studies show [45], [81].
Reusable trojaned-checkpoint corpora; adapter-audit suites
Adversarial rollout benchmark; physical-consistency stress tests Trajectory-ranking attack protocols with standardized candidate sets Feedback-poisoning and recovery-latency measurement Long-horizon persistence tests for poisoned memory and tools
B. Robust State Grounding Sensor redundancy and physical consistency. Multisensor redundancy can detect inconsistency among camera, LiDAR, GNSS, IMU, map, and proprioception. SAVIOR demonstrates the value of robust physical invariants for autonomous vehicles [70]. Similar invariants should be used before observations enter the world model. Cross-modal validation. VLM/VLA systems should check whether language goals, visual objects, maps, and predicted affordances agree. Environmental prompt injection and typographic attacks motivate treating scene text as untrusted unless grounded by task context and policy [68], [99]. Uncertainty-aware state estimation. The state estimator should output uncertainty and provenance, not only a latent state. High-uncertainty states should trigger conservative planning, additional sensing, or human confirmation, following calibration, OOD detection, and safe-learning evidence that confidence can fail under shift [128], [129], [92]. C. Safe Imagination and Prediction Physics- and rule-aware rollout. World-model predictions should be checked against physical constraints, temporal con-
13
sistency, traffic or task rules, and embodiment limits. Control barrier functions, model-based safe learning, and safe RL provide useful primitives for turning constraints into runtime checks [137], [39], [92]. Ensemble and calibration defenses. Ensembles, uncertainty thresholds, and OOD detection can reduce overconfident false futures. Probabilistic model-based RL and calibration studies provide practical starting points for estimating uncertainty in dynamics and prediction [135], [128], [129]. These defenses are particularly important when world models are used as safety checkers, because false confidence can become a false permit to act. Adversarial rollout testing. Planners should be evaluated against adaptive attacks on physical conditions, latent dynamics, and trajectory ranking [24], [30]. A safety claim is weak if it only covers clean rollouts. D. Trajectory-Level Action Gating Runtime shields. Shielding and CBF-based controllers can prevent unsafe actions from reaching actuators even when high-level plans are compromised [115], [116]. For worldmodel-based systems, the shield should validate the candidate trajectory, not just the immediate action. Rule-grounded intervention. SafetyChip and safety guardrails for LLM-enabled robots show how natural-language safety rules or temporal constraints can prune unsafe actions [122], [32]. These methods should be coupled with grounded state validation so that rules are evaluated on a trustworthy world state. Fallback and staged approval. When predicted risk is high or uncertainty is poorly calibrated, the system should stop, slow down, replan, request more sensing, or hand off to a human. Human handoff must account for over-trust and interface design [177], [178], [179]. E. Feedback, Memory, and Deployment Assurance Feedback auditing. Execution feedback should be treated as untrusted until validated. Data-predictive recovery can reconstruct compromised CPS measurements and isolate attacked channels [71]. Similar mechanisms can protect worldmodel updates from false experience. Memory hygiene. Long-term memories and retrieved documents should include source tags, expiration, access control, and consistency checks. AgentPoison and retrieval-corpus poisoning show that poisoned memory or retrieved passages can redirect agent behavior without changing model weights [150], [148]. Tool outputs should be sandboxed and verified before being written into persistent world knowledge [151]. Middleware and update security. ROS security, authentication, encryption, topic authorization, rollback, and update verification are necessary because middleware and model updates can alter world state, action commands, or safety monitors [152], [73], [153], [155]. F. Representative Defense Methods Table VIII summarizes concrete defense methods that can be reused or adapted for world-model-based embodied AI.
Some methods defend explicit world-model prediction, while others protect adjacent trust boundaries such as multimodal encoders, VLA action heads, runtime shields, CPS feedback, or middleware. G. Defense Priorities The five insights defined in Section I imply different defense priorities. A single runtime filter is insufficient because the failure unit changes across insights: the unit is a physically grounded trajectory for SP, a state-conditioned risk estimate for SD, a compounding rollout-execution loop for PA, a globally ranked trajectory set for NC, and the safety certificate itself for PI. Table IX maps each insight to the defense mechanism that should be treated as primary rather than auxiliary. These priorities also determine where redundancy should be placed. For SP and NC, redundancy should compare planned trajectories against grounded physical constraints. For SD and PA, redundancy should compare state estimates, predicted rollouts, and execution feedback across time. For PI, redundancy must be independent of the safety world model itself; otherwise, the monitor and the policy can share the same corrupted state, dynamics, or uncertainty estimate. This is why defense evaluation should report both task utility and monitor failure modes, especially false-safe predictions, missed interventions, and recovery latency. VIII. O PEN C HALLENGES Several challenges remain for secure world-model-based embodied AI. • Auditing implicit world models. Many VLA policies encode physical priors without exposing rollouts. Security evaluation must distinguish explicit prediction failures from implicit state-action failures. • Verifying learned prediction. Neural-network verification, barrier certificates, and runtime synthesis provide starting points [180], [181], [116], but generative futures and action chunks are still difficult to verify at scale. • Testing adaptive safety-checker attacks. Safety world models should be evaluated against attackers that target observation, dynamics, constraint, and uncertainty channels, not only fixed perturbations. • Auditing generated data. Generative world models can become persistent poisoning sources. Synthetic demonstrations and rollouts require provenance, trigger tests, physical-consistency checks, and downstream contamination tests. • Transferring security across embodiments and domains. Attacks and defenses are usually validated on one robot, simulator, or sensor suite. Sim-to-real studies show that dynamics gaps change which perturbations matter [124], [126], so security claims should be re-evaluated whenever embodiment, environment, or model scale changes. • Calibrating human trust in predicted futures. Generated rollouts and safety certificates are increasingly shown
14
TABLE VIII: Representative defense methods for world-model-based embodied AI. Runtime safety-checker methods (shielding, CBF, predictive filters, and safe model-based RL) are cataloged separately in Table III and are not repeated here. Defense Method
Year
Category
Subcategory
Target Model
Dataset / Env.
WM Role
SafeVLA [173] SAVIOR [70] Data-predictive recovery [71] InverTune [174] Attention Betrays [175] ModAgnostic defense [101] Concept Dictionary [176]
2025 2020 2023
Alignment Recovery Recovery
Constrained learning Physical invariants Sensor reconstruction
VLA AV stack CPS
Safety-CHORES Autonomous vehicle Physical systems
Safe policy State validation Feedback audit
2026 2026 2025
Backdoor def. Backdoor def. Inference
BAC analysis Visual-token recon. VLA robustification
Multimodal encoder Robot policy VLA
CLIP-like models Robot policies VLA tasks
State encoder Trigger removal Action robustness
2026
Inference
Concept monitor
VLA
VLA tasks
Safety feature
TABLE IX: Defense priorities for the five world-model security insights. Insight
Failure mode
Primary defense
Evaluation focus
Representative tools
SP
PA
Small error compounds
NC
Local checks miss global risk False safe prediction
Grounded dynamics; rule-to-state validation Sensor provenance; calibrated uncertainty gate Cross-stage consistency; rollback and recovery Long-horizon cost; action-chunk gate Independent monitor; checker uncertainty audit
Constraint violation after accepted commands Risk under viewpoint, pose, lighting, and OOD shift Rollout divergence; feedback contamination Unsafe trajectory selected despite safe local steps Predicted-safe but actually unsafe executions
SafetyChip; guardrails; CBFs [122], [32], [116]
SD
Semantics pass; physics fails State shift changes risk
PI
to operators and end users. Human-automation trust research indicates that plausible visualizations invite overtrust [177], [182], which turns a compromised world model into a social-engineering channel as well as a technical one. • Building deployable assurance. Existing safety and cybersecurity standards provide useful baselines [183], [184], [185], [155], [186], [187], [188], but world-model-based systems also need logs of state assumptions, predicted futures, blocked actions, updates, rollbacks, uncertainty handling, and incidents. IX. C ONCLUSION World models are becoming a cognitive core of embodied AI: they ground observations into world states, imagine future trajectories, evaluate action consequences, and support longterm adaptation. This survey argued that such predictive cognition is also a security boundary. Attacks from data, representations, sensors, prompts, video generators, planning objectives, controllers, memory, tools, and supply chains can corrupt world states, dynamics, affordances, safety costs, trajectory ranking, or feedback provenance. The resulting security failures are not isolated model errors. They propagate through imagined futures and closed-loop execution, turning digital compromise into physical risk. We therefore proposed a lifecycle taxonomy for world-modelbased embodied AI and identified five recurring insights: semantic-to-simulation-to-action gaps, state- and uncertaintyconditional risk, rollout-to-execution amplification, trajectorylevel non-compositionality, and predictive safety illusion. A central lesson is that world models are dual-use. They can improve safety by predicting and blocking risky futures, but they also introduce new attack surfaces when used as trusted safety checkers or synthetic data generators. Building secure world-model-based embodied AI will require provenanceaware data pipelines, robust state grounding, uncertainty-
EVA-VLA; sensor validation [111], [26] SAVIOR; recovery; shielding [70], [71], [115] TRAP; SafeBench [30], [104] SafeDreamer; VLM-SAFE; ensembles [4], [5]
calibrated prediction, trajectory-level action gating, feedback auditing, and deployment assurance across both cyber and physical layers. R EFERENCES [1] D. Ha and J. Schmidhuber, “World models,” arXiv preprint arXiv:1803.10122, 2018. [2] D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in International Conference on Learning Representations, 2020. [3] D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, “Mastering diverse domains through world models,” arXiv preprint arXiv:2301.04104, 2023. [4] W. Huang, J. Ji, C. Xia, B. Zhang, and Y. Yang, “SafeDreamer: Safe reinforcement learning with world models,” in International Conference on Learning Representations (ICLR), 2024. [5] Y. Qu, Z. Huang, Z. Sheng, J. Chen, Y. Leng, S. Labi, and S. Chen, “Vlm-safe: Vision-language model-guided safety-aware reinforcement learning with world models for autonomous driving,” arXiv preprint arXiv:2505.16377, 2025. [6] J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel, “Genie: Generative interactive environments,” arXiv preprint arXiv:2402.15391, 2024. [7] D. Driess, F. Xia, M. S. M. Sajjadi et al., “Palm-e: An embodied multimodal language model,” in Proceedings of the 40th International Conference on Machine Learning, 2023. [8] A. Brohan, N. Brown, J. Carbajal et al., “Rt-1: Robotics transformer for real-world control at scale,” arXiv preprint arXiv:2212.06817, 2022. [9] B. Zitkovich, T. Xu, T. Xiao et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in Conference on Robot Learning, 2023. [10] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi et al., “Openvla: An open-source vision-language-action model,” arXiv preprint arXiv:2406.09246, 2024. [11] O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu et al., “Octo: An open-source generalist robot policy,” arXiv preprint arXiv:2405.12213, 2024. [12] T. Wang, C. Han, J. Liang, W. Yang, D. Liu, L. X. Zhang, Q. Wang, J. Luo, and R. Tang, “Exploring the adversarial vulnerabilities of vision-language-action models in robotics,” in Proceedings of the
15
IEEE/CVF International Conference on Computer Vision, 2025, pp. 6948–6958. [13] T. Sato, Y. Hayakawa, R. Suzuki, Y. Shiiki, K. Yoshioka, and Q. A. Chen, “Lidar spoofing meets the new-gen: Capability improvements, broken assumptions, and new attack strategies,” in Network and Distributed System Security Symposium (NDSS), 2024. [14] H. Zhang, C. Zhu, X. Wang, Z. Zhou, C. Yin, M. Li, L. Xue, Y. Wang, S. Hu, A. Liu et al., “Badrobot: Jailbreaking embodied llm agents in the physical world,” in The Thirteenth International Conference on Learning Representations, 2025. [15] Y. Liu, W. Chen, Y. Bai, X. Liang, G. Li, W. Gao, and L. Lin, “Aligning cyber space with physical world: A comprehensive survey on embodied ai,” IEEE/ASME Transactions on Mechatronics, 2025. [16] W. Xing, M. Li, M. Li, and M. Han, “Towards robust and secure embodied ai: A survey on vulnerabilities and attacks,” arXiv preprint arXiv:2502.13175, 2025. [17] B. Ma, H. Guo, P. Lv, M. Xu, X. Dai, Y. Zhang, Y. Yang, and Y. Zhang, “What breaks embodied ai security: Llm vulnerabilities, cps flaws, or something else?” arXiv preprint arXiv:2602.17345, 2026. [18] Z. Wang, J. Hu, and R. Mu, “Safety of embodied navigation: A survey,” arXiv preprint arXiv:2508.05855, 2025. [19] S. Neupane, S. Mitra, I. A. Fernandez, S. Saha, S. Mittal, J. Chen, N. Pillai, and S. Rahimi, “Security considerations in ai-robotics: A survey of current methods, challenges, and opportunities,” IEEE Access, vol. 12, pp. 22 072–22 097, 2024. [20] L. Baraldi, Z. Zeng, C. Zhang, A. Nayak, H. Zhu, F. Liu, Q. Zhang, P. Wang, S. Liu, Z. Hu, and A. Cangelosi, “The safety challenge of world models for embodied ai agents: A review,” arXiv preprint arXiv:2510.05865, 2025. [21] M. Parmar, “Safety, security, and cognitive risks in world models,” arXiv preprint arXiv:2604.01346, 2026. [22] A. Robey, Z. Ravichandran, V. Kumar, H. Hassani, and G. J. Pappas, “Jailbreaking llm-controlled robots,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 11 948–11 956. [23] X. Lu, Z. Huang, X. Li, C. Zhang, W. Xu et al., “Poex: Towards policy executable jailbreak attacks against the llm-based robots,” arXiv preprint arXiv:2412.16633, 2024. [24] Z. Guo, S. Liang, A. Balogh, N. Lunberry, R.-C. Tu, M. Jelasity, and D. Tao, “When world models dream wrong: Physicalconditioned adversarial attacks against world models,” arXiv preprint arXiv:2602.18739, 2026. [25] J. Guo, W. Jiang, Y. Lin, Y. Liu, R. Zhang, G. Lu, A. Chen, X. Han, and H. Li, “State backdoor: Towards stealthy real-world poisoning attack on vision-language-action model in state space,” arXiv preprint arXiv:2601.04266, 2026. [26] H. Kim, R. Bandyopadhyay, M. O. Ozmen, Z. B. Celik, A. Bianchi, Y. Kim, and D. Xu, “A systematic study of physical sensor attack hardness,” in 2024 IEEE Symposium on Security and Privacy (SP). IEEE, 2024, pp. 2328–2347. [27] Y. Huang, Z. Wang, Z. Wan, Y. Tian, H. Xu, Y. Han, and Y. Gan, “ANNIE: Be careful of your robots,” arXiv preprint arXiv:2509.03383, 2025. [28] B. Xu, Y. Shang, B. Wang, and E. Ferrara, “Silentdrift: Exploiting action chunking for stealthy backdoor attacks on vision-languageaction models,” arXiv preprint arXiv:2601.14323, 2026. [29] Y. Lou, Y. Zhu, Q. Song, R. Tan, C. Qiao, W.-B. Lee, and J. Wang, “A first {Physical-World} trajectory prediction attack via {LiDARinduced} deceptions in autonomous driving,” in 33rd USENIX Security Symposium (USENIX Security 24), 2024, pp. 6291–6308. [30] S. Duan, K. Zhang, and X. Luo, “Trap: Tail-aware ranking attack for world-model planning,” arXiv preprint arXiv:2605.01950, 2026. [31] Y. Huang, L. Ding, Z. Tang, T. Wang, X. Lin, W. Zhang, M. Ma, and Y. Zhang, “A framework for benchmarking and aligning taskplanning safety in llm-based embodied agents,” arXiv preprint arXiv:2504.14650, 2025. [32] Z. Ravichandran, A. Robey, V. Kumar, G. J. Pappas, and H. Hassani, “Safety guardrails for llm-enabled robots,” IEEE Robotics and Automation Letters, 2026. [33] J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan, “A survey of embodied ai: From simulators to research tasks,” IEEE Transactions on Emerging Topics in Computational Intelligence, vol. 6, no. 2, pp. 230–244, 2022. [34] B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and Autonomous Systems, vol. 57, no. 5, pp. 469–483, 2009.
[35] W. Xu, M. Liu, O. Sokolsky, I. Lee, and F. Kong, “Llm-enabled cyberphysical systems: Survey, research opportunities, and challenges,” in 2024 IEEE International Workshop on Foundation Models for CyberPhysical Systems & Internet of Things (FMSys). IEEE, 2024, pp. 50–55. [36] G. Wang, Y. Xie, Y. Jiang et al., “Voyager: An open-ended embodied agent with large language models,” Transactions on Machine Learning Research, 2024. [37] D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in International Conference on Machine Learning, 2019. [38] P. Wu, A. Escontrela, D. Hafner, K. Goldberg, and P. Abbeel, “Daydreamer: World models for physical robot learning,” in Conference on Robot Learning, 2022. [39] F. Berkenkamp, M. Turchetta, A. P. Schoellig, and A. Krause, “Safe model-based reinforcement learning with stability guarantees,” in Advances in Neural Information Processing Systems, 2017. [40] S. Shan, W. Ding, J. Passananti, S. Wu, H. Zheng, and B. Y. Zhao, “Nightshade: Prompt-specific poisoning attacks on text-to-image generative models,” arXiv preprint arXiv:2310.13828, 2023. [41] R. Wang, M. Zhu, J. Ou, R. Chen, X. Tao, P. Wan, and B. Wu, “Badvideo: Stealthy backdoor attack against text-to-video generation,” arXiv preprint arXiv:2504.16907, 2025. [42] C. Li, Y. Min, J. Zhang, Z. Yuan, S. Shan, and X. Chen, “T2vattack: Adversarial attack on text-to-video diffusion models,” arXiv preprint arXiv:2512.23953, 2025. [43] Q. Li, X. Yang, and X. Wang, “Badwam: When world-action models dream right but act wrong,” arXiv preprint arXiv:2607.15207, 2026. [44] J. Li, Y. Zhao, X. Zheng, Z. Xu, Y. Li, X. Ma, and Y.-G. Jiang, “Attackvla: Benchmarking adversarial and backdoor attacks on visionlanguage-action models,” arXiv preprint arXiv:2511.12149, 2025. [45] X. Zhou, G. Tie, G. Zhang, H. Wang, P. Zhou, and L. Sun, “BadVLA: Towards backdoor attacks on vision-language-action models via objective-decoupled optimization,” in Advances in Neural Information Processing Systems (NeurIPS), 2025. [46] Z. Xu, J. Li, Y. Zhao, X. Zheng, X. Ma, and Y.-G. Jiang, “DropVLA: An action-level backdoor attack on vision-language-action models,” arXiv preprint arXiv:2510.10932, 2025. [47] M. Ahn, A. Brohan, N. Brown et al., “Do as i can, not as i say: Grounding language in robotic affordances,” arXiv preprint arXiv:2204.01691, 2022. [48] W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng et al., “Inner monologue: Embodied reasoning through planning with language models,” in Conference on Robot Learning (CoRL), 2022. [49] J. Liang, W. Huang, F. Xia et al., “Code as policies: Language model programs for embodied control,” in 2023 IEEE International Conference on Robotics and Automation, 2023. [50] D. Surı́s, S. Menon, and C. Vondrick, “Vipergpt: Visual inference via python execution for reasoning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. [51] S. Gao, J. Yang, L. Chen, K. Chitta, Y. Qiu, A. Geiger, J. Zhang, and H. Li, “Vista: A generalizable driving world model with high fidelity and versatile controllability,” arXiv preprint arXiv:2405.17398, 2024. [52] Y. Shang, Y. Lin, Y. Zheng, H. Fan, J. Ding, J. Feng, J. Chen, L. Tian, and Y. Li, “Urbanworld: An urban world model for 3d city generation,” arXiv preprint arXiv:2407.11965, 2024. [53] NVIDIA, N. Agarwal, A. Ali, M. Bala, Y. Balaji et al., “Cosmos world foundation model platform for physical ai,” arXiv preprint arXiv:2501.03575, 2025. [54] M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Komeili, M. Muckley, A. Rizvi, C. Roberts, Y. LeCun et al., “V-jepa 2: Self-supervised video models enable understanding, prediction and planning,” arXiv preprint arXiv:2506.09985, 2025. [55] NVIDIA, N. Agarwal, A. Ali, J. Allen, M. Antolini, A. Aubame, A. Azzolini, J. Bai, M. Bala, Y. Balaji et al., “Cosmos 3: Omnimodal world models for physical ai,” arXiv preprint arXiv:2606.02800, 2026. [56] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y. Shen, and Y. Xu, “Causal world modeling for robot control,” arXiv preprint arXiv:2601.21998, 2026. [57] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang et al., “World action models are zeroshot policies,” arXiv preprint arXiv:2602.15922, 2026. [58] Team HY-World, C. Cao, X. Zuo, Z. Wang, Y. Zhang, J. Wu, Z. Liu, Y. Gong, Y. Liu, B. Yuan et al., “Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds,” arXiv preprint arXiv:2604.14268, 2026.
16
[59] Z. Xiao, J. Tu, C. Zou, Y. Zuo, Z. Li, P. Wang, B. Yu, F. Huang, J. Lin, and Z. Liu, “Webworld: A large-scale world model for web agent training,” arXiv preprint arXiv:2602.14721, 2026. [60] I.-S. Oh, “A tutorial on world models and physical ai,” arXiv preprint arXiv:2606.12783, 2026. [61] B. Biggio, B. Nelson, and P. Laskov, “Poisoning attacks against support vector machines,” in Proceedings of the 29th International Conference on Machine Learning, 2012. [62] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, “Badnets: Evaluating backdooring attacks on deep neural networks,” IEEE Access, vol. 7, pp. 47 230–47 244, 2019. [63] A. Radford, J. W. Kim, C. Hallacy et al., “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine Learning, 2021. [64] A. Shafahi, W. R. Huang, M. Najibi, O. Suciu, C. Studer, T. Dumitras, and T. Goldstein, “Poison frogs! targeted clean-label poisoning attacks on neural networks,” in Advances in Neural Information Processing Systems, 2018. [65] J. Steinhardt, P. W. Koh, and P. Liang, “Certified defenses for data poisoning attacks,” in Advances in Neural Information Processing Systems, 2017. [66] T. B. Brown, D. Mané, A. Roy, M. Abadi, and J. Gilmer, “Adversarial patch,” arXiv preprint arXiv:1712.09665, 2017. [67] K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, and D. Song, “Robust physical-world attacks on deep learning visual classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018. [68] L. Burbano, D. Ortiz, Q. Sun, S. Yang, H. Tu, C. Xie, Y. Cao, and A. A. Cardenas, “Chai: Command hijacking against embodied ai,” arXiv preprint arXiv:2510.00181, 2025. [69] X. Wang, J. Li, Z. Weng, Y. Wang, Y. Gao, T. Pang, C. Du, Y. Teng, Y. Wang, Z. Wu et al., “Freezevla: Action-freezing attacks against vision-language-action models,” arXiv preprint arXiv:2509.19870, 2025. [70] R. Quinonez, J. Giraldo, L. Salazar, E. Bauman, A. Cardenas, and Z. Lin, “{SAVIOR}: Securing autonomous vehicles with robust physical invariants,” in 29th USENIX security symposium (USENIX Security 20), 2020, pp. 895–912. [71] L. Zhang, K. Sridhar, M. Liu, P. Lu, X. Chen, F. Kong, O. Sokolsky, and I. Lee, “Real-time data-predictive attack-recovery for complex cyber-physical systems,” in 2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS). IEEE, 2023, pp. 209–222. [72] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, M. Fritz, and P. Hohenecker, “More than you’ve asked for: A comprehensive analysis of novel prompt injection threats to application-integrated large language models,” arXiv preprint arXiv:2302.12173, 2023. [73] B. Dieber, B. Breiling, S. Taurer, S. Kacianka, P. Schartner, and M. Hofbaur, “Security for the robot operating system,” Robotics and Autonomous Systems, vol. 98, pp. 192–203, 2017. [74] K. Bonawitz, V. Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth, “Practical secure aggregation for privacy-preserving machine learning,” in proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 2017, pp. 1175–1191. [75] M. Jagielski, A. Oprea, B. Biggio, C. Liu, C. Nita-Rotaru, and B. Li, “Manipulating machine learning: Poisoning attacks and countermeasures for regression learning,” in Proceedings of the IEEE Symposium on Security and Privacy, 2018. [76] M. Pourkeshavarz, M. Sabokrou, and A. Rasouli, “Adversarial backdoor attack by naturalistic data poisoning on trajectory prediction in autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 14 885–14 894. [77] Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning attack on neural networks,” in Proceedings of the 25th Annual Network and Distributed System Security Symposium, 2018. [78] A. Saha, A. Subramanya, and H. Pirsiavash, “Hidden trigger backdoor attacks,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 7, pp. 11 957–11 965, 2020. [79] A. Turner, D. Tsipras, and A. Madry, “Label-consistent backdoor attacks,” arXiv preprint arXiv:1912.02771, 2019. [80] Z. Zhou, Z. Xiao, H. Xu, J. Sun, D. Wang, and J. Zhang, “Goal-oriented backdoor attack against vision-language-action models via physical objects,” arXiv preprint arXiv:2510.09269, 2025. [81] J. Zhou, Y. Wei, R. Zhen, B. Zhao, X. Xia, R. Shao, X. Su, and S. Yang, “Inject once survive later: Backdooring vision-language-action
models to persist through downstream fine-tuning,” arXiv preprint arXiv:2602.00500, 2026. [82] A. Liu, Y. Zhou, X. Liu, T. Zhang, S. Liang, J. Wang, Y. Pu, T. Li, J. Zhang, W. Zhou et al., “Compromising llm driven embodied agents with contextual backdoor attacks,” IEEE Transactions on Information Forensics and Security, 2025. [83] R. Jiao, S. Xie, J. Yue, T. Sato, L. Wang, Y. Wang, Q. A. Chen, and Q. Zhu, “Can we trust embodied agents? exploring backdoor attacks against embodied llm-based decision-making systems,” arXiv preprint arXiv:2405.20774, 2024. [84] Z. Ni, R. Ye, Y. Wei, Z. Xiang, Y. Wang, and S. Chen, “Physical backdoor attack can jeopardize driving with vision-large-language models,” arXiv preprint arXiv:2404.12916, 2024. [85] X. Wang, H. Pan, H. Zhang, M. Li, S. Hu, Z. Zhou, L. Xue, A. Liu, Y. Jiang, L. Y. Zhang et al., “Trojanrobot: Physical-world backdoor attacks against vlm-based robotic manipulation,” arXiv preprint arXiv:2411.11683, 2024. [86] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” arXiv preprint arXiv:2304.08485, 2023. [87] G. Team et al., “Gemini: A family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805, 2023. [88] I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in International Conference on Learning Representations, 2015. [89] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” in International Conference on Learning Representations, 2014. [90] J. Garcı́a and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research, vol. 16, pp. 1437–1480, 2015. [91] J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proceedings of the 34th International Conference on Machine Learning, 2017. [92] L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 5, pp. 411–444, 2022. [93] A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial examples in the physical world,” arXiv preprint arXiv:1607.02533, 2017. [94] S.-T. Chen, C. Cornelius, J. Martin, D. H. Chau et al., “Shapeshifter: Robust physical adversarial attack on faster r-cnn object detector,” in Machine Learning and Computer Security Workshop, 2018. [95] X. Liu, H. Yang, Z. Liu, L. Song, Y. Chen, and H. Li, “Dpatch: An adversarial patch attack on object detectors,” in Proceedings of the AAAI Conference on Artificial Intelligence Workshop on Artificial Intelligence Safety, 2019. [96] H. Lu, Y. Yu, Y. Yang, C. Yi, Q. Zhang, B. Shen, A. C. Kot, and X. Jiang, “When robots obey the patch: Universal transferable patch attacks on vision-language-action models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026. [97] N. Zhang, W. Tao, X. Xiao, Q. Sun, Y. Zheng, W. Mo, P. Wang, and N. Zhang, “Attention-guided patch-wise sparse adversarial attacks on vision-language-action models,” arXiv preprint arXiv:2511.21663, 2025. [98] T. Trippel, O. Weisse, W. Xu, P. Honeyman, and K. Fu, “Walnut: Waging doubt on the integrity of mems accelerometers with acoustic injection attacks,” in 2017 IEEE European symposium on security and privacy (EuroS&P). IEEE, 2017, pp. 3–18. [99] C. Li, Z. Kang, J. Zhang, Z. Ma, A. Cheng, X. Li, and J. Ma, “The shawshank redemption of embodied ai: Understanding and benchmarking indirect environmental jailbreaks,” arXiv preprint arXiv:2511.16347, 2025. [100] S. Yang, Z. Wang, D. Ortiz, L. Burbano, M. Kantarcioglu, A. Cardenas, and C. Xie, “Probing vulnerabilities of vision-lidar based autonomous driving systems,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 3561–3569. [101] H. Xu, Y. S. Koh, S. Huang, Z. Zhou, D. Wang, J. Sakuma, and J. Zhang, “Model-agnostic adversarial attack and defense for visionlanguage-action models,” arXiv preprint arXiv:2510.13237, 2025. [102] Q. Zhang, S. Hu, J. Sun, Q. A. Chen, and Z. M. Mao, “On adversarial robustness of trajectory prediction for autonomous vehicles,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 15 159–15 168. [103] Y. Yu, W. Han, L. Wu, B. Liu, E. Wang, and Z. Zhang, “Enduring, efficient and robust trajectory prediction attack in autonomous driving via optimization-driven multi-frame perturbation framework,” in Pro-
17
ceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 17 229–17 238. [104] C. Xu, W. Ding, W. Lyu, Z. Liu, S. Wang, Y. He, H. Hu, D. Zhao, and B. Li, “Safebench: A benchmarking platform for safety evaluation of autonomous vehicles,” in Advances in Neural Information Processing Systems, 2022. [105] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning finegrained bimanual manipulation with low-cost hardware,” arXiv preprint arXiv:2304.13705, 2023. [106] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” in Robotics: Science and Systems, 2023. [107] F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” arXiv preprint arXiv:2211.09527, 2022. [108] A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” Advances in Neural Information Processing Systems, vol. 36, 2023. [109] A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv preprint arXiv:2307.15043, 2023. [110] E. K. Jones, A. Robey, A. Zou, Z. Ravichandran, G. J. Pappas, H. Hassani, M. Fredrikson, and J. Z. Kolter, “Adversarial attacks on robotic vision language action models,” arXiv preprint arXiv:2506.03350, 2025. [111] H. Liu, S. Ruan, J. Long, J. Wu, J. Hou, H. Tang, T. Jiang, W. Zhou, and W. Yao, “Eva-vla: Evaluating vision-language-action models’ robustness under real-world physical variations,” arXiv preprint arXiv:2509.18953, 2025. [112] G. Thomas, Y. Luo, and T. Ma, “Safe reinforcement learning by imagining the near future,” in Advances in Neural Information Processing Systems, 2021. [113] Y. J. Ma, A. Shen, O. Bastani, and D. Jayaraman, “Conservative and adaptive penalty for model-based safe reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, 2022, pp. 5404–5412. [114] Y. Hogewind, T. D. Simão, T. Kachman, and N. Jansen, “Safe reinforcement learning from pixels using a stochastic latent representation,” in International Conference on Learning Representations, 2023. [115] M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu, “Safe reinforcement learning via shielding,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2018. [116] A. D. Ames, X. Xu, J. W. Grizzle, and P. Tabuada, “Control barrier function based quadratic programs for safety critical systems,” IEEE Transactions on Automatic Control, vol. 62, no. 8, pp. 3861–3876, 2017. [117] J. F. Fisac, A. K. Akametalu, M. N. Zeilinger, S. Kaynama, J. H. Gillula, and C. J. Tomlin, “A general safety framework for learningbased control in uncertain robotic systems,” IEEE Transactions on Automatic Control, vol. 64, no. 7, pp. 2737–2752, 2019. [118] K. P. Wabersich and M. N. Zeilinger, “Linear model predictive safety certification for learning-based control,” in 2018 IEEE Conference on Decision and Control. IEEE, 2018, pp. 7130–7135. [119] ——, “A predictive safety filter for learning-based control of constrained nonlinear dynamical systems,” Automatica, vol. 129, p. 109597, 2021. [120] S. Li and O. Bastani, “Robust model predictive shielding for safe reinforcement learning with stochastic dynamics,” in 2020 IEEE International Conference on Robotics and Automation. IEEE, 2020, pp. 7166–7172. [121] K.-C. Hsu, H. Hu, and J. F. Fisac, “The safety filter: A unified view of safety-critical control in autonomous systems,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 7, pp. 47–72, 2024. [122] Z. Yang, S. S. Raman, A. Shah, and S. Tellex, “Plug in the safety chip: Enforcing constraints for llm-driven robot agents,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, pp. 14 435–14 442. [123] J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 2017, pp. 23–30. [124] X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-toreal transfer of robotic control with dynamics randomization,” in 2018 IEEE International Conference on Robotics and Automation. IEEE, 2018, pp. 3803–3810. [125] K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige et al., “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,”
in 2018 IEEE International Conference on Robotics and Automation. IEEE, 2018, pp. 4243–4250. [126] Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox, “Closing the sim-to-real loop: Adapting simulation randomization with real world experience,” in 2019 International Conference on Robotics and Automation. IEEE, 2019, pp. 8973–8979. [127] W. Zhang, X. Kong, C. Dewitt, T. Braunl, and J. B. Hong, “A study on prompt injection attack against llm-integrated mobile robotic systems,” in 2024 IEEE 35th International Symposium on Software Reliability Engineering Workshops (ISSREW). IEEE, 2024, pp. 361–368. [128] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proceedings of the 34th International Conference on Machine Learning, 2017, pp. 1321–1330. [129] D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in International Conference on Learning Representations, 2017. [130] S. Huang, N. Papernot, I. Goodfellow, Y. Duan, and P. Abbeel, “Adversarial attacks on neural network policies,” arXiv preprint arXiv:1702.02284, 2017. [131] Y.-C. Lin, Z.-W. Hong, Y.-H. Liao, M.-L. Shih, M.-Y. Liu, and M. Sun, “Tactics of adversarial attack on deep reinforcement learning agents,” in Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, 2017, pp. 3756–3762. [132] A. Gleave, M. Dennis, C. Wild, N. Kant, S. Levine, and S. Russell, “Adversarial policies: Attacking deep reinforcement learning,” in International Conference on Learning Representations, 2020. [133] L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning, 2017, pp. 2817–2826. [134] A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE International Conference on Robotics and Automation. IEEE, 2018, pp. 7559–7566. [135] K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Advances in Neural Information Processing Systems, 2018. [136] F. Ebert, C. Finn, S. Dasari, A. X. Xie, A. Lee, and S. Levine, “Visual foresight: Model-based deep reinforcement learning for vision-based robotic control,” in Conference on Robot Learning, 2018, pp. 841–852. [137] R. Cheng, G. Orosz, R. M. Murray, and J. Burdick, “End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2019. [138] Z. Yuan, A. Hall, H. Bansal et al., “Safe-control-gym: A unified benchmark suite for safe learning-based control and reinforcement learning in robotics,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 11 142–11 149, 2022. [139] C. Agia, R. Sinha, J. Yang, Z.-a. Cao, R. Antonova, M. Pavone, and J. Bohg, “Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress,” in Conference on Robot Learning, 2024. [140] J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke, “Sim-to-real: Learning agile locomotion for quadruped robots,” in Robotics: Science and Systems, 2018. [141] M. Deitke, W. Han, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, E. Kolve, A. Farhadi, A. Kembhavi, and R. Mottaghi, “Robothor: An open simulation-to-real embodied ai platform,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. [142] M. Savva, A. Kadian, O. Maksymets et al., “Habitat: A platform for embodied ai research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019. [143] A. Szot, A. Clegg, E. Undersander et al., “Habitat 2.0: Training home assistants to rearrange their habitat,” in Advances in Neural Information Processing Systems, 2021. [144] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “Carla: An open urban driving simulator,” in Conference on robot learning. PMLR, 2017, pp. 1–16. [145] S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, 2011, pp. 627–635. [146] M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer, “HG-DAgger: Interactive imitation learning with human experts,” in 2019 International Conference on Robotics and Automation. IEEE, 2019, pp. 8077–8083.
18
[147] A. Mandlekar, D. Xu, R. Martı́n-Martı́n, S. Savarese, and L. Fei-Fei, “What matters in learning from offline human demonstrations for robot manipulation,” in Conference on Robot Learning, 2021, pp. 1678–1690. [148] Z. Zhong, Z. Huang, A. Wettig, and D. Chen, “Poisoning retrieval corpora by injecting adversarial passages,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [149] L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar, “Minedojo: Building openended embodied agents with internet-scale knowledge,” in Advances in Neural Information Processing Systems, 2022. [150] Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “AgentPoison: Redteaming LLM agents via poisoning memory or knowledge bases,” in Advances in Neural Information Processing Systems (NeurIPS), 2024. [151] J. Shi, Z. Yuan, G. Tie, P. Zhou, N. Z. Gong, and L. Sun, “Prompt injection attack to tool selection in LLM agents,” in Network and Distributed System Security Symposium (NDSS), 2026. [152] M. Quigley, K. Conley, B. Gerkey, J. Faust, T. Foote, J. Leibs, R. Wheeler, and A. Y. Ng, “Ros: an open-source robot operating system,” in ICRA Workshop on Open Source Software, 2009. [153] T. Mokhamed, F. M. Dakalbab, S. Abbas, and M. A. Talib, “Security in robot operating systems (ros): analytical review study,” in The 3rd international conference on distributed sensing and intelligent systems (ICDSIS 2022), vol. 2022. IET, 2022, pp. 79–94. [154] M. Tehranipoor and F. Koushanfar, “A survey of hardware trojan taxonomy and detection,” IEEE design & test of computers, vol. 27, no. 1, pp. 10–25, 2010. [155] International Organization for Standardization and SAE International, “Iso/sae 21434: Road vehicles—cybersecurity engineering,” 2021, international standard. [156] E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi, “Ai2-thor: An interactive 3d environment for visual ai,” arXiv preprint arXiv:1712.05474, 2017. [157] A. Ray, J. Achiam, and D. Amodei, “Benchmarking safe exploration in deep reinforcement learning,” arXiv preprint arXiv:1910.01708, 2019. [158] M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “Alfred: A benchmark for interpreting grounded instructions for everyday tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. [159] S. James, A. J. Davison, and E. Johns, “Rlbench: The robot learning benchmark and learning environment,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3019–3026, 2020. [160] S. Srivastava, C. Li, M. Lingelbach et al., “Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments,” in Conference on Robot Learning, 2021. [161] M. Shridhar, X. Yuan, M.-A. Côté, Y. Bisk, A. Trischler, and M. Hausknecht, “Alfworld: Aligning text and embodied environments for interactive learning,” in International Conference on Learning Representations, 2021. [162] O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 7327–7334, 2022. [163] M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, J. Salvador, K. Ehsani, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi, “Procthor: Largescale embodied ai using procedural generation,” in Advances in Neural Information Processing Systems, 2022. [164] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “Libero: Benchmarking knowledge transfer for lifelong robot learning,” in Advances in Neural Information Processing Systems, 2023. [165] J. Ji, B. Zhang, J. Zhou, X. Pan, W. Huang, R. Sun, Y. Geng, Y. Zhong, J. Dai, and Y. Yang, “Safety-gymnasium: A unified safe reinforcement learning benchmark,” in Advances in Neural Information Processing Systems, 2023. [166] H. Walke, K. Black, A. Lee et al., “Bridgedata v2: A dataset for robot learning at scale,” in Conference on Robot Learning (CoRL), 2023. [167] J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y. Tang, S. Tao, X. Wei, Y. Yao et al., “Maniskill2: A unified benchmark for generalizable manipulation skills,” in International Conference on Learning Representations, 2023. [168] S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu, “Robocasa: Large-scale simulation of everyday tasks for generalist robots,” in Robotics: Science and Systems, 2024. [169] Open X-Embodiment Collaboration, A. Padalkar, A. Pooley et al., “Open x-embodiment: Robotic learning datasets and RT-X models,”
in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024. [170] A. Liu, Z. Ying, L. Wang, Y. Xiao, J. Wang, Y. Ma, J. Guo, Z. Yin, M. Zhang, and X. Liu, “Agentsafe: Benchmarking the safety of embodied agents on hazardous instructions,” arXiv preprint arXiv:2506.14697, 2025. [171] S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou et al., “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 9710–9719. [172] P. W. Koh, J. Steinhardt, and P. Liang, “Stronger data poisoning attacks break data sanitization defenses,” arXiv preprint arXiv:1811.00741, 2018. [173] B. Zhang, Y. Zhang, J. Ji, Y. Lei, J. Dai, Y. Chen, and Y. Yang, “Safevla: Towards safety alignment of vision-language-action model via constrained learning,” arXiv preprint arXiv:2503.03480, 2025. [174] M. Sun, Y. Li, Y. Ge, Y. Liu, B. Du, and Q. Wang, “Invertune: A backdoor defense method for multimodal contrastive learning via backdoor-adversarial correlation analysis,” Proceedings of the Network and Distributed System Security Symposium, 2026. [175] X. Li, P. Fu, W. Huang, N. Pan, S. Yang, K. Zhao, G. Wan, M. Li, J. Xuan, and M. Li, “When attention betrays: Erasing backdoor attacks in robotic policies by reconstructing visual tokens,” arXiv preprint arXiv:2602.03153, 2026. [176] S. Wen, S. Yang, S. Fu, J. Zhang, L. Hu, and D. Wang, “Concept-based dictionary learning for inference-time safety in vision language action models,” arXiv preprint arXiv:2602.01834, 2026. [177] J. D. Lee and K. A. See, “Trust in automation: Designing for appropriate reliance,” Human Factors, vol. 46, no. 1, pp. 50–80, 2004. [178] P. A. Hancock, D. R. Billings, K. E. Schaefer, J. Y. C. Chen, E. J. de Visser, and R. Parasuraman, “A meta-analysis of factors affecting trust in human-robot interaction,” Human Factors, vol. 53, no. 5, pp. 517–527, 2011. [179] P. A. Lasota, T. Fong, and J. A. Shah, “A survey of methods for safe human-robot interaction,” Foundations and Trends in Robotics, vol. 5, no. 4, pp. 261–349, 2017. [180] G. Katz, C. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer, “Reluplex: An efficient smt solver for verifying deep neural networks,” in Computer Aided Verification, 2017. [181] R. Ivanov, J. Weimer, R. Alur, G. J. Pappas, and I. Lee, “Verisig: Verifying safety properties of hybrid systems with neural network controllers,” in Proceedings of the 22nd ACM International Conference on Hybrid Systems: Computation and Control, 2019. [182] A. Martinetti, P. K. Chemweno, K. Nizamis, and E. Fosch-Villaronga, “Redefining safety in light of human-robot interaction: A critical review of current standards and regulations,” Frontiers in chemical engineering, vol. 3, p. 666237, 2021. [183] International Organization for Standardization, “Iso 26262: Road vehicles—functional safety,” 2018, international standard. [184] ——, “Iso 21448: Road vehicles—safety of the intended functionality,” 2022, international standard. [185] UL Standards & Engagement, “Ul 4600: Standard for safety for the evaluation of autonomous products,” 2022, industry standard. [186] International Electrotechnical Commission, “Iec 62443: Security for industrial automation and control systems,” 2018, international standard. [187] International Organization for Standardization, “Iso 10218: Robots and robotic devices—safety requirements for industrial robots,” 2025, international standard. [188] ——, “Iso/ts 15066: Robots and robotic devices—collaborative robots,” 2016, technical specification.