ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies Jianming Ma∗1,2 , Rongjun Jin∗1 , Xiaxi Si1 , Yang Zhang1 , Yiheng Li1 , Yue Gao1, 2† 1
arXiv:2609.11697v1 [cs.RO] 10 Sep 2026
2
Shanghai Jiao Tong University Shanghai Institute of Innovation
Abstract
Observa�on
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in generalpurpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones (π0.5 and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a 100% step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.
Language Instruc�on
Introduction Vision-Language-Action (VLA) models (Kim et al. 2024; Black et al. 2025; NVIDIA et al. 2025) and World-Action Models (WAMs) (Ye et al. 2026; Liao et al. 2025) have recently shown remarkable progress in general-purpose robotic manipulation. By combining large-scale visual-language representations with generative action heads, these embodied foundation models can adapt to diverse downstream tasks while retaining broad semantic and motor priors. Despite their strong task-solving capabilities, however, the actions generated by these models are not always directly executable. Predicted actions may violate hard physical constraints, such as position limits, velocity bounds, or workspace restrictions, leading to unstable execution, hardware damage, or unsafe robot behavior. For robotic manipulation, safety must be enforced at every execution step, because a single infeasible action can result in an unacceptable physical consequence. This is particularly problematic for flow-matching action heads, which generate ∗ †
These authors contributed equally. Corresponding author.
Founda�on Backbone (VLA/VAM)
Ac�on Head ( )
ODE Sampler
Unconstrained ac�on
Boundary-aware gradient
Differen�able (End-to-End Fine-tuning) Training-Inference Aligned
ActSafeGuard Reparameteriza�on:
Safe Discrete Update
ray-scaling operator
No Learnable Params (Minimal Invasive)
Convex Feasible Region
forward pass
Constrained ac�on
Implicit Oblique Projec�on
backward pass
Figure 1: Overview of ActSafeGuard. Given an observation and a language instruction, a foundation VLA/WAM backbone and its flow-matching action head predict an unconstrained velocity vt . ActSafeGuard transforms vt into an update direction dt and a safe scaling factor st , adaptively rescaling constraint-violating steps such that the trajectory remains within the convex feasible region. During backpropagation, its gradient transformation admits an implicit oblique-projection interpretation, providing boundary-aware learning signals.
action chunks in an unconstrained space without explicit feasibility guarantees. Thus, a practical safeguard must provide deterministic per-step feasibility while preserving the task competence and generative flexibility of pretrained foundation policies. Existing safety approaches generally address this challenge from either the training side or the inference side. Training-time constrained methods incorporate safety costs (Zhang et al. 2026) or regularization objectives (Wei, Wu, and Haghbayan 2026) into policy optimization, allowing the policy to become safety-aware but typically providing only probabilistic or expectation-based guarantees. In contrast, inference-time safeguards directly correct infeasible actions through projection, clipping, or optimization-based filtering (Hu et al. 2025). Although these methods can enforce hard constraints during deployment, the correction mechanism is absent from policy training. Consequently, the policy is optimized in an unconstrained action space but executed un-
der an externally modified action distribution, introducing a training–inference mismatch that may distort generated trajectories and degrade task performance. The central limitation is therefore not simply the absence of a safety mechanism, but the separation of hard constraint enforcement from policy learning. In this work, we introduce ActSafeGuard, a differentiable and training-aligned constraint operator for flow-matchingbased policies. Our key insight is that hard action feasibility can be integrated directly into the policy learning process, rather than being treated merely as an external inferencetime correction. ActSafeGuard is inserted at the output of the flow-matching action head and applies an analytical rayscaling operator to each discrete flow update. Starting from a feasible initial point, this construction ensures that every constrained flow step remains within the state-dependent convex feasible set. Crucially, the same safeguard operator is active during both training and inference. During training, its differentiable construction allows gradients to propagate through the constrained update, enabling the action head to receive boundary-aware learning signals. This training-aligned design encourages the policy to adapt its predicted directions to the geometry of the feasible action space, thereby preserving effective learning while enforcing hard feasibility. Moreover, ActSafeGuard introduces no additional learnable network and can be incorporated into pretrained flow-matching action heads with minimal modification. We evaluate ActSafeGuard on representative foundation backbones across multiple manipulation tasks with various constraints. The results show that ActSafeGuard consistently achieves a 100% step safety rate while preserving task success rates compared with other constrained methods. Our main contributions are summarized as follows: • We propose ActSafeGuard, a differentiable and trainingaligned constraint operator that integrates hard action feasibility directly into the learning and generation processes of flow-matching based policies. • We show that ActSafeGuard induces boundary-aware gradients through an implicit oblique projection, enabling constrained updates to slide along feasible boundaries instead of being merely truncated. • Extensive experiments with π0.5 and Fast-WAM demonstrate deterministic constraint satisfaction while preserving task success and action-generation quality.
Related Works Safety Constrained VLA and WAMs. Safety is a core desideratum for VLA and WAM policies deployed in openended environments, where unconstrained physical actions may lead to catastrophic failures (Li et al. 2026b,a; Kim et al. 2026a,b). Existing methods mainly address this issue from either the training side or the inference side: (1) Training-time Constrained Policies. One paradigm embeds safety directly into the policy optimization process (Zhang et al. 2026; Tang et al. 2026; Wei, Wu, and Haghbayan 2026). For instance, SafeVLA (Zhang et al. 2026) formulates the safety task as a constrained MDP and leverages safe reinforcement learning (RL) to achieve VLA safety
alignment. SafeDojo (Tang et al. 2026) utilizes a generative video world model to produce imaginary rollouts and introduces a Lagrangian-constrained GRPO objective for policy training. These methods can improve safety awareness, but they typically optimize expected costs or soft penalties and therefore do not guarantee deterministic per-step feasibility. Furthermore, incorporating constrained optimization techniques (e.g., Lagrangian multipliers) into the training pipeline introduces a non-trivial trade-off between task objective and constraint satisfaction. (2) Inference-time Safety Guards. Inference-time safeguards instead modify generated actions during deployment. One branch of work (Hu et al. 2025; English, Zheng, and Ewetz 2026) constructs control barrier functions (CBFs) and safety sets to project actions by solving a Quadratic Program (QP) optimization at each sampling step, ensuring that the generated actions strictly reside within the feasible region. Another branch adopts action masking (Beaudin et al. 2026) or heuristic-based hard truncation (Chandra, Damodaran, and Wang 2026) to filter out unsafe actions. While such methods can enforce hard constraints at execution time, they are absent from policy training, creating a training–inference mismatch that may distort the learned action distribution and degrade task performance. In summary, training-time constrained methods simultaneously optimize for task competence and safety regularizers but fall short of delivering deterministic constraint satisfaction. Conversely, inference-time safety guards mathematically enforce hard constraints during deployment but introduce a training-inference mismatch that often compromises the backbone’s execution capability. How to effectively balance task performance with strict constraint satisfaction through a training-inference aligned framework remains an open challenge.
Preliminary Policy Formulation. We formulate language-conditioned robotic manipulation as learning a mapping from multimodal observations to a sequence of future actions. At control step t, the observation tuple is defined as ot = [{Iit }ni=1 , qt ], comprising multi-view images and proprioceptive states. Guided by a language instruction l, a VLM/VGM backbone fθ first projects the inputs into a unified latent representation: et = fθ (ot , l). An action head πϕ then decodes this embedding into an action chunk at:t+H−1 , representing motor commands over a future horizon H. The policy is composed as Π(ot , l) = πϕ (et ), and the generated chunk is executed via a receding-horizon scheme. Flow-Matching Action Generation. Flow matching (Lipman et al. 2023) is a prevalent paradigm for training the generative action head πϕ . Let x1 denote the target expert action chunk with horizon H (i.e., at:t+H−1 ) and x0 ∼ p(x0 ) denote an initial noise sample. Flow matching constructs a probability path via linear interpolation: xτ = (1 − τ )x0 + τ x1 ,
τ ∈ [0, 1],
(1)
where τ denotes the continuous flow timestep. The action head, instantiated as a velocity field network vϕ (xτ , τ, et )
conditioned on the latent context et , is trained to approximate the target vector field u⋆ = x1 − x0 . The training objective is: i h 2 LCFM = Ex0 ,x1 ,τ ∥vϕ (xτ , τ, et ) − (x1 − x0 )∥2 . (2) During inference, starting from x0 , an Ordinary Differential Equation (ODE) solver integrates the learned vector field over τ ∈ [0, 1] to synthesize the final action chunk x1 . Constrained Action Space. In real-world robotic applications, physical and environmental limitations impose strict constraints on the action space. We characterize the feasible region as a state-dependent convex polytope: c
c
C(ot ) = {x | A(ot )x ≤ b(ot )},
(3)
where xc denotes the dimensions of the action chunk subject to hard constraints, and A(ot ), b(ot ) define observationconditioned boundaries. Our goal is to ensure that xc strictly resides within C(ot ) throughout the sampling process, thereby guaranteeing deterministic, zero-violation execution upon deployment.
Method To achieve zero-constraint-violation generation with minimal intervention to pretrained models, we propose ActSafeGuard. The core idea is to embed a differentiable, parameterfree ray-scaling operator directly into a discrete-time flowmatching process, ensuring per-step feasibility while facilitating boundary-aware gradient updates during training.
Discrete-Time Flow Matching with Feasible States Standard flow matching integrates a continuous vector field over time. However, verifying continuous-time safety analytically is often intractable. Instead, ActSafeGuard formulates the generation as an N -step discrete flow process. Let τk = k/N for k ∈ {0, . . . , N −1}. The exact probability path is discretized as: xk = (1 − τk )x0 + τk x1 , where the ideal one-step update is given by ∆⋆ = xk+1 −xk = (x1 −x0 )/N . Instead of learning the continuous velocity vϕ , the action head is parameterized to directly predict the discrete step ∆ϕ (xk , τk , et ). The model can be trained using a discretetime flow matching loss: h i 2 LDFM = Ek,x0 ,x1 ∥∆ϕ (xk , τk , et ) − ∆⋆ ∥2 . (4) The formulation of Eq. (4) is mathematically aligned with the CFM objective in Eq. (2) used by standard continuous flow matching. Although discretization introduces bounded approximation error (Ma et al. 2026), it preserves the vectorfield prior of continuous pretrained backbones by interpreting their predicted velocity as a nominal discrete update direction. The advantage of this formulation is that feasibility can be enforced step by step: if x0 ∈ C(ot ) and every constrained update respects the boundary, then the entire generated sequence x0 → · · · → xN remains feasible.
ActSafeGuard Parameterization Given the discrete formulation above, ActSafeGuard enforces feasibility by modifying each nominal flow update before it is applied. The design goal is to preserve the vector-field prior of the pretrained action head as much as possible: the predicted velocity still determines the update direction, while ActSafeGuard only rescales its magnitude when the update would leave the feasible set. For notational simplicity, we describe the operation on the constrained action dimensions. At step k, the pretrained action head predicts an unconstrained velocity vϕ (xk , τk , et ). We convert it into the following two components: Direction Vector: This vector dictates the intended moving direction along with its step magnitude: dϕ =
vϕ (xk , τk , et ) , N
(5)
where N denotes the total number of discretization steps. Safety Scaling Factor: To guarantee that the next state remains within the closed feasible domain, we adaptively constrain the magnitude of dϕ using a safety weight sϕ ∈ [0, 1]. ActSafeGuard modulates the nominal step to yield the final safe discrete step: ∆ϕ = sϕ dϕ .
(6)
Computation of sϕ via Ray Shooting. Given the direction vector dϕ and the current state xk , the ray-shooting operator computes the intersection z with the closest constraint boundary along the line of travel: α=
bi − a ⊤ i xk , ⊤ ⊤ ai dϕ i:ai dϕ >0 min
z = xk + αdϕ ,
(7)
where a⊤ i and bi denote the i-th row and element of the constraint matrix A and boundary vector b respectively, and α ≥ 0 is a scalar representing the maximum allowable scaling factor before hitting the nearest constraint facet i. Consequently, the safety scaling factor sϕ is defined as: 1, α ≥ 1, sϕ = min(1, α) = (8) α, α < 1. This piecewise formulation yields two regimes: • When α ≥ 1, the nominal update xk + dϕ is already feasible and ActSafeGuard leaves it unchanged. • When α < 1, the nominal update would cross a boundary, so ActSafeGuard shortens it exactly to the ray-boundary intersection: xk+1 = xk + ∆ϕ = xk + αdϕ = z ∈ ∂C.
(9)
Proposition 1 (Safety Guarantee) By construction, ActSafeGuard guarantees that the generated trajectory remains within the closed feasible region C(ot ) across all inference steps, provided that the initial noise satisfies x0 ∈ C(ot ).
∂C : a⊤ m x = bm
dϕ along
Gradient-Based Boundary Sliding Correction
∆⋆ (Target)
We now examine why ActSafeGuard provides boundaryaware learning signals rather than behaving as a nondifferentiable clipping module. The key lies in the local Jacobian of the safe update ∆ϕ with respect to the nominal direction dϕ . Consider the boundary-limited case α < 1, and assume that the active facet is locally unique:
Raw gradient
am (Normal)
Boundary-sliding gradient: a⊤ m Qm = 0, Qm dϕ = 0
dϕ (Unconstrained) xk+1 = z ∆ϕ = sϕ dϕ
xk
Feasible Region C(ot )
î = arg Figure 2: ActSafeGuard update and boundary-sliding gradient. When the nominal update dϕ crosses the active constraint facet a⊤ m x = bm , ActSafeGuard rescales the step to the boundary intersection z. Around this boundary-limited point, the local Jacobian acts as an oblique projection Qm , satisfying a⊤ m Qm = 0 and Qm dϕ = 0. Thus, raw gradients that would otherwise point outside the feasible region are converted into boundary-sliding gradients along the active facet, encouraging the policy to re-orient unsafe updates while preserving hard feasibility.
Proposition 2 (Differentiability) A key advantage of ActSafeGuard is its end-to-end differentiability, allowing it to be embedded directly into training loops without additional modification. The complete forward pass for the ActSafeGuard can be expressed as: ∆ϕ = min 1,
v bi − a ⊤ ϕ i xk , N · ⊤v N a i:a⊤ v >0 ϕ ϕ i i min
(10)
where vϕ is the shortcut notation for the velocity field prediction vϕ (xk , τk , et ). The gradient of the output ∆ϕ with respect to the input vϕ exists in closed-form everywhere except at the boundary corners, which can be handled via standard subgradient methods during backpropagation. The training and inference procedures of ActSafeGuard are summarized in Algorithm 1 and Appendix G. Algorithm 1: ActSafeGuard Training Require: Pretrained policy parameters ϕ, constraint sets C(ot ) 1: for each batch (ot , x1 ) guided by instruction l do 2: Encode context: et = fθ (ot , l) 3: Sample step k ∼ U (0, N − 1) and initial noise x0 ∼ p(x0 ) s.t. x0 ∈ C(ot ) k k 4: Compute interpolated state xk = (1 − N )x0 + N x1 5: Predict baseline velocity vϕ ← vϕ (xk , τk , et ) 6: Compute direction vector dϕ = vϕ /N 7: Compute scale sϕ using Eq (7) and (8) 8: Apply safe update ∆ϕ = sϕ dϕ 9: Compute discrete flow matching loss LDFM = ∥∆ϕ − (x1 − x0 )/N ∥22 10: Update policy parameters ϕ ← ϕ − η∇ϕ LDFM 11: end for
bi − a ⊤ i xk . ⊤d ⊤ a i:ai dϕ >0 i ϕ min
(11)
Let cî = bî − a⊤ x . In this regime, the safe update is î k ∆ϕ = αdϕ =
cî dϕ . d a⊤ î ϕ
(12)
Taking the local differential with respect to dϕ gives d∆ϕ = αQî ddϕ ,
Qî = I −
dϕ a⊤ î a⊤ d î ϕ
.
(13)
The matrix Qî is an oblique projection onto the tangent space of the active facet along the ray direction. It satisfies a⊤ Qî = 0⊤ , î
Qî dϕ = 0.
(14)
Therefore, once the nominal update reaches a boundary, first-order changes in the safe update cannot move outward through the active constraint. Instead, the Jacobian preserves only variations that slide the boundary intersection along the feasible facet. This property explains the boundary-sliding behavior of ActSafeGuard. Unlike hard clipping or stop-gradient correction, the ray-scaling operator remains differentiable in the boundary-limited regime. Gradients from the training objective can still flow through ∆ϕ to dϕ , but they are geometrically reshaped by Qî . As a result, the action head receives learning signals that encourage it to re-orient unsafe nominal directions along the feasible boundary, rather than merely increasing a step magnitude that would be truncated by the safeguard.
Experiments We conduct experiments to evaluate whether ActSafeGuard can enforce hard action feasibility without sacrificing the task competence of pretrained flow-matching policies. The experiments are organized around three research questions: • RQ1: Performance preservation. Can ActSafeGuard eliminate constraint violations while maintaining the task success rate of the original policy? • RQ2: Generalizability to complex constraints. Can ActSafeGuard handle high-dimensional, state-dependent, and multi-step dynamic constraints? • RQ3: Differentiability. How much does the differentiable ray-scaling design contribute to stable constrained policy learning?
Backbone
π0.5
Fast-WAM
Static Constraints (PosCons)
Method
Dynamic Constraints (PosCons + VelCons)
lp
ps
hm
pec
Mean SR
lp
ps
hm
pec
Mean SR
Baseline Projection Truncation GaugeFlow ActSafeGuard
100 100 99 100 100
90 90 91 82 90
21 21 27 25 34
90 88 92 93 97
75.25 74.75 77.25 75.00 80.25
100 100 99 N/A 100
90 84 90 N/A 98
21 26 27 N/A 31
90 95 91 N/A 97
75.25 76.25 76.75 N/A 81.50
Baseline Projection Truncation GaugeFlow ActSafeGuard
100 100 98 100 100
93 85 80 80 90
42 40 50 39 45
98 80 90 90 93
83.25 76.25 79.50 77.25 82.00
100 100 100 N/A 97
93 85 94 N/A 95
42 35 45 N/A 45
98 82 88 N/A 95
83.25 75.50 81.75 N/A 83.00
Table 1: Quantitative task success performance (SR, %) across all evaluation scenarios. We evaluate methods across four manipulation tasks: lift pot (lp), place shoe (ps), hanging mug (hm), and place empty cup (pec), under static position constraints (PosCons) and dynamic position + velocity constraints (PosCons + VelCons). Step Safety Rate (SSR) is omitted as all constrained methods deterministically achieve 100% SSR. Bold numbers indicate the best performance among constrained methods. Training and evaluation details are in Appendix D and E.
Experimental Setup Simulation environment and tasks. We evaluate ActSafeGuard in RoboTwin (Chen et al. 2025) with the bimanual ALOHA-Agilex platform. The robot has two 6-DoF arms and two grippers, yielding a 14-dimensional action space. We consider four manipulation tasks of different difficulty and contact structure: lift pot, place shoe, hanging mug, and place empty cup. Policy input and output. The policy uses absolute joint and gripper positions as actions. At each control step t, it receives a multimodal observation ot = [{Iit }3i=1 , qt , gt ], where {Iit }3i=1 are three camera views, qt ∈ R12 denotes the current arm-joint positions, and gt ∈ R2 denotes the gripper states. Given a language instruction l, the policy predicts an action chunk at:t+H−1 = [at , . . . , at+H−1 ] with horizon H. Safety constraints. We impose hard constraints on the 12DoF arm-joint subspace, denoted by the superscript q, while leaving the gripper dimensions unconstrained. Detailed parameterizations are provided in Appendix A. Position constraints (PosCons) define a static joint-limit box: qmin ≤ aqt+i ≤ qmax ,
∀i ∈ {0, . . . , H − 1}.
(15)
These bounds prevent the policy from producing actions outside the valid workspace. Velocity constraints (VelCons) bound one-step joint displacement: |qt − aqt | ≤ vmax ,
|aqt+i − aqt+i+1 | ≤ vmax ,
∀i ∈ {0, . . . , H − 2}.
(16) (17)
Unlike PosCons, VelCons depends on the current proprioceptive state qt , making the feasible set C(ot ) state-dependent and time-varying.
Backbones and baselines. We instantiate ActSafeGuard on two representative flow-matching policies: π0.5 (Black et al. 2025), a pretrained VLA model, and Fast-WAM (Yuan et al. 2026), a world-action model that uses world modeling during training while retaining direct action prediction at inference time. We compare against three constraint-handling baselines: Projection, an inference-time step-wise projection method; Truncation, a post-hoc clipping method applied after unconstrained generation; and GaugeFlow (Li, Liang, and Chen 2026), a training-aware method based on an invertible gauge map. Detail introduction of baselines are in Appendix B. GaugeFlow requires a static star-shaped feasible region, so it is only applicable to PosCons and is marked N/A under PosCons + VelCons. Evaluation metrics. We report Success Rate (SR), Step Safe Rate (SSR), Maximum Mean Discrepancy (MMD), and Log Dimensionless Jerk (LDLJ) (Balasubramanian, Melendez-Calderon, and Burdet 2011). SR measures task completion, SSR measures the percentage of executed control steps satisfying all constraints, MMD measures distributional deviation from expert action chunks, and LDLJ measures trajectory smoothness. MMD and LDLJ evaluate whether hard constraint enforcement damages the actiongeneration quality of the pretrained backbone. More details are in Appendix C.
Why Are Hard Constraints Necessary? We first examine whether safety can be obtained simply by training on feasible demonstrations. For each task, the PosCons and VelCons bounds are constructed from the minimum and maximum statistics of the training data, so the training demonstrations lie inside the feasible set by design. We then fine-tune the unconstrained baselines on these data and evaluate them in in-distribution RoboTwin environments. Table 2 shows that feasible training data alone is insufficient for deterministic safety. The mean SSR of the π0.5
PosCons
Task
PosCons + VelCons
π0.5
Fast-WAM
π0.5
Fast-WAM
lift pot place shoe hanging mug place empty cup
83.59 20.43 64.62 37.04
15.12 5.86 39.69 31.27
82.85 16.45 57.38 35.69
13.68 3.79 35.24 27.33
Mean SSR
51.42
22.99
48.09
20.01
Table 2: Step Safety Rate (SSR, %) of the unconstrained baseline with different backbones. Evaluated across four manipulation tasks under static position constraints (PosCons) and combined dynamic constraints (PosCons + VelCons). baseline is only 51.42% under PosCons and 48.09% under PosCons + VelCons, while Fast-WAM drops further to 22.99% and 20.01%, respectively. These violations occur despite using in-distribution evaluation settings, indicating that neural fitting error, flow sampling error, and distribution shift can push generated chunks outside the empirical feasible envelope. Therefore, safety cannot rely solely on data filtering or average-case regularization; it requires an explicit mechanism that guarantees per-step feasibility under the provided constraints.
Performance Preservation (RQ1) Table 1 evaluates task success rates under hard constraints across various scenarios. Under guaranteed safety, ActSafeGuard achieves the best success rate among constrained methods in most tasks, matching or even outperforming the unconstrained baseline in certain cases. The trajectoryquality results in Figure 3 further show that ActSafeGuard preserves the pretrained action distribution and smoothness better than post-hoc correction methods. Compared with inference-based methods, ActSafeGuard eliminates training–inference mismatch regarding constraint enforcement. Furthermore, its implicit oblique-projection gradients during training enable flexible boundary-sliding corrections, yielding lower fitting errors and smoother trajectories than simple truncation or projection. Compared with GaugeFlow, ActSafeGuard also provides a more conservative modification to the pretrained action head. GaugeFlow maps the original action space into a unit ball and trains the flow in the transformed coordinate system. This transformation can disrupt the vector-field priors already learned by the pretrained backbone. ActSafeGuard instead operates directly on the original flow update and only modifies the step when the nominal update would cross a boundary. This minimal intervention better preserves the pretrained capabilities.
Generalizability to Complex Constraints (RQ2) We next evaluate whether ActSafeGuard scales beyond static box constraints. Under PosCons + VelCons, the feasible region can be written as a high-dimensional polytope A(ot )x ≤ b(ot ) whose facets depend on the current robot
-log(MMD)
Norm(LDLJ)
lp (pos+vel)
pec (pos)
ps (pos+vel)
hm (pos)
pec (pos)
ps (pos+vel)
hm hm (pos+vel) (pos)
ps (pos)
pi05 Baseline
lp (pos+vel)
pec (pos+vel)
lp (pos) pi05 Projection
hm (pos+vel)
ps (pos)
pec (pos+vel)
lp (pos) pi05 Truncation
pi05 GaugeFlow
pi05 ASG (Ours)
(a) π0.5 backbone -log(MMD) pec (pos)
Norm(LDLJ)
lp (pos+vel)
ps (pos+vel)
hm (pos) ps (pos)
fastwam Baseline
lp (pos+vel)
pec (pos)
ps (pos+vel)
hm hm (pos+vel) (pos) pec (pos+vel)
lp (pos) fastwam Projection
hm (pos+vel)
ps (pos)
fastwam Truncation
pec (pos+vel)
lp (pos) fastwam GaugeFlow
fastwam ASG (Ours)
(b) Fast-WAM backbone
Figure 3: Trajectory quality and distribution alignment comparisons. Radar charts of − log(MMD) (higher is better) and normalized LDLJ (mean/std normalization, higher is better) across manipulation tasks: lift pot (lp), place shoe (ps), hanging mug (hm), and place empty cup (pec), under static position constraints (pos) and dynamic position + velocity constraints (pos + vel). Larger covered areas indicate superior performance in preserving original trajectory distributions and smoothness. Details in Appendix E. state and on adjacent actions within the generated chunk. This setting is substantially harder than static PosCons because the feasible set changes at every control step and couples multiple future actions through velocity bounds. As shown in Table 1 and Figure 3, ActSafeGuard remains effective in this dynamic regime. ActSafeGuard obtains the highest mean SR under PosCons + VelCons, outperforming other baselines while maintaining deterministic safety. This result highlights an important practical advantage of the ray-scaling formulation. Methods based on a fixed gauge map or other static coordinate transformations are difficult to apply when the feasible set changes with ot ; this is why GaugeFlow is not applicable to the PosCons + VelCons regime. CBF-style inference methods also become challenging for high-dimensional action chunks with many coupled linear inequalities. In contrast, ActSafeGuard only requires computing ray-boundary intersections with the current feasible polytope. As a result, it can naturally support both static and dynamic constraints without changing the pretrained backbone or solving an iterative optimization problem during
SR (%) ↑
LDLJ ↑
MMD ↓
Task 1: place shoe (PosCons + VelCons) π0.5 + ActSafeGuard 98.0 π0.5 + ActSafeGuard (w/ sg) 3.0
-17.171 -24.611
0.00242 0.03008 (a) Snapshots on pick green cube task.
-16.197 -16.991
0.00960 0.01171
Table 3: Ablation study on gradient differentiability. Evaluation is conducted across different tasks and constraint configurations. “w/ sg” denotes applying a stop-gradient operation on the safety scaling factor during training. Note that the step safety rate (SSR) is omitted from the table as both variants deterministically achieve 100% SSR across all evaluation episodes. each sampling step.
Ablation on Differentiability (RQ3) We isolate the contribution of differentiability by applying a stop-gradient operation to the safety scaling factor sϕ during fine-tuning. This variant, denoted ActSafeGuard (w/ sg), keeps the ray-scaling operator active during inference, so it still enforces 100% SSR. However, the action head no longer receives boundary-aware gradients through sϕ during training. The comparison is shown in Table 3. The results demonstrate that deterministic inference-time feasibility is not sufficient for learning a useful constrained policy. On place shoe under PosCons + VelCons, removing gradients through sϕ causes SR to collapse from 98.0% to 3.0%. This large gap indicates that dynamic constraints require the policy to learn how its action chunks interact with moving feasible boundaries.
Real-Robot Deployment To assess the practical deployment applicability of ActSafeGuard, we evaluate our approach on a physical AgileX ALOHA bimanual platform. We design two distinct realworld tasks (details are in Appendix F): pick green cube: The robotic arm is required to pick up a green cube and place it into a paper cup. The policy outputs absolute joint positions, and safety constraints are imposed as joint boundaries. guide the ball: Using a magnetic wand attached to the gripper, the arm guides a green ball along a physical track into a designated container. The policy outputs End-Effector (EEF) Cartesian coordinates, with safety constraints specified as an L-shaped corridor to prevent the EEF from slipping off the track. We adopt π0.5 as the backbone and fine-tune it using 50 demonstration trajectories collected for each task. Figure 4 illustrates real-world execution snapshots and the corresponding execution trajectories. As depicted in the trajectory plots (Figure 4c), ActSafeGuard strictly adheres to the prescribed boundaries across all steps. In contrast, the unconstrained baseline violates the corridor boundary in guide the ball. We evaluate each method across 5 rollouts per task. For pick green cube, both the baseline and ActSafeGuard achieve 5/5 successes. For the
(b) Snapshots on guide the ball task. pick green cube
−0.075
0.75 0.50 0.25
bound baseline ActSafeGuard
0.00 0
200
index
EEF y (m)
Task 2: place empty cup (PosCons) π0.5 + ActSafeGuard 97.0 π0.5 + ActSafeGuard (w/ sg) 95.0
joint pos(dim=7)
Method
guide the ball
−0.100 −0.125 constraint baseline ActSafeGuard
−0.150 −0.175
400
−0.10
−0.15
EEF x (m)
−0.20
(c) Action trajectories and constraints in two tasks.
Figure 4: Real Robot Deployment Results. ActSafeGuard strictly satisfies constraints while maintaining task performance.
more constrained guide the ball task, the baseline achieves 4/5 successes, whereas ActSafeGuard succeeds in all 5/5 trials. These empirical results show that ActSafeGuard can guarantee constraint satisfaction without compromising task capability. Notably, the feasible set in guide the ball forms a non-convex L-shaped corridor, further suggesting the generality of ActSafeGuard: as long as the ray-boundary intersection can be computed, the same ray-scaling method can be extended beyond convex polytopes.
Conclusion We present ActSafeGuard, a differentiable and trainingaligned safeguard layer for flow-matching based VLA and WAM policies. By applying a parameter-free ray-scaling operator to each discrete flow update, ActSafeGuard enforces hard action feasibility at every generation step. Its local Jacobian induces boundary-sliding gradients, enabling the action head to adapt to feasible-set geometry rather than relying on post-hoc correction alone. Across simulation and real robot tasks, multiple constraint regimes, and two foundation backbones, ActSafeGuard achieved deterministic step safety while maintaining strong task success and action-generation quality. ActSafeGuard also has some limitations. Currently it assumes that the feasible action space can be represented by explicitly specified constraints for which ray-boundary intersections are tractable. Moreover, the constraint set is still manually specified from task knowledge or demonstration statistics. Automatically discovering, validating, and updating such constraints from perception or interaction data is an important direction for future work.
References Balasubramanian, S.; Melendez-Calderon, A.; and Burdet, E. 2011. A robust and sensitive metric for quantifying movement smoothness. IEEE transactions on biomedical engineering, 59(8): 2126–2136. Beaudin, A.; Krasowski, H.; Nagpal, K.; Seshia, S. A.; Arcak, M.; and Mehr, N. 2026. Any-Body Guard: Universal Safeguarding for Manipulation Policies via Action Masking. arXiv:2606.22278. Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Galliker, M. Y.; Ghosh, D.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; LeBlanc, D.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Ren, A. Z.; Shi, L. X.; Smith, L.; Springenberg, J. T.; Stachowicz, K.; Tanner, J.; Vuong, Q.; Walke, H.; Walling, A.; Wang, H.; Yu, L.; and Zhilinsky, U. 2025. π0.5 : A Vision-Language-Action Model with Open-World Generalization. arXiv:2504.16054. Chandra, N.; Damodaran, S.; and Wang, L. 2026. PhysVLA: Towards Physically-Grounded VLA for Embodied Robotic Manipulation. arXiv preprint arXiv:2606.13886. Chen, T.; Chen, Z.; Chen, B.; Cai, Z.; Liu, Y.; Li, Z.; Liang, Q.; Lin, X.; Ge, Y.; Gu, Z.; Deng, W.; Guo, Y.; Nian, T.; Xie, X.; Chen, Q.; Su, K.; Xu, T.; Liu, G.; Hu, M.; ang Gao, H.; Wang, K.; Liang, Z.; Qin, Y.; Yang, X.; Luo, P.; and Mu, Y. 2025. RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation. arXiv:2506.18088. English, W.; Zheng, H.; and Ewetz, R. 2026. NeuroSymbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching. arXiv:2607.01378. Hu, S.; Liu, Z.; Liu, S.; Cen, J.; Meng, Z.; Wang, S.; Li, X.; and He, X. 2025. VLSA: Vision-LanguageAction Models with Plug-and-Play Safety Constraint Layer. arXiv:2512.11891. Kim, D.; Park, D.; Lee, S.; Kim, J.; Oh, Y.; Shin, J.; and Yoon, S. 2026a. Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation. arXiv:2606.05660. Kim, J.; Chen, W.; Soleymanzadeh, D.; Ding, Y.; Gao, X.; Tu, Z.; Zhang, R.; Fei, F.; Veer, S.; Lyu, Y.; Zheng, M.; and Gu, Y. 2026b. Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World. arXiv:2602.04056. Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2024. OpenVLA: An Open-Source Vision-Language-Action Model. arXiv:2406.09246. Li, Q.; Yin, B.; Huang, W.; Liu, R.; Zou, B.; Yu, R.; Ye, J.; Yu, W.; and Wang, X. 2026a. Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms. arXiv:2604.23775. Li, X.; Liang, E.; and Chen, M. 2026. Gauge Flow Matching: Efficient Constrained Generative Modeling over General
Convex Set and Beyond. In The Fourteenth International Conference on Learning Representations. Li, X.; Zheng, X.; Gao, Y.; Xia, X.; Wang, Y.; Wang, X.; Sun, Y.; Zhao, Y.; Wen, M.; Li, J.; Chen, Z.; Gong, X.; Liu, Y.; Li, Y.; Wu, Y.; Wang, C.; Sun, J.; Cao, Y.; Chen, Z.; Chen, J.; Gui, T.; Zhang, Q.; Wu, Z.; Qiu, X.; Huang, X.; Zhang, T.; Wei, Z.; Wang, K.; Li, X.; Huang, H.; Erfani, S.; Bailey, J.; Wang, J.; Xiao, C.; He, R.; Li, B.; Ma, X.; and Jiang, Y.-G. 2026b. Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses. arXiv:2605.02900. Liao, Y.; Zhou, P.; Huang, S.; Yang, D.; Chen, S.; Jiang, Y.; Hu, Y.; Cai, J.; Liu, S.; Luo, J.; Chen, L.; Yan, S.; Yao, M.; and Ren, G. 2025. Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation. arXiv:2508.05635. Lipman, Y.; Chen, R. T. Q.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2023. Flow Matching for Generative Modeling. In The Eleventh International Conference on Learning Representations. Ma, J.; Yang, Q.; Zhang, Y.; Yan, L.; Cao, Z.; Zhang, Y.; and Gao, Y. 2026. PolyFlow: Safe and Efficient PolytopeConstrained Flow Matching with Constraint Embedding and Projection-free Update. In Forty-third International Conference on Machine Learning. NVIDIA; :; Bjorck, J.; Castañeda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L. J.; Fang, Y.; Fox, D.; Hu, F.; Huang, S.; Jang, J.; Jiang, Z.; Kautz, J.; Kundalia, K.; Lao, L.; Li, Z.; Lin, Z.; Lin, K.; Liu, G.; Llontop, E.; Magne, L.; Mandlekar, A.; Narayan, A.; Nasiriany, S.; Reed, S.; Tan, Y. L.; Wang, G.; Wang, Z.; Wang, J.; Wang, Q.; Xiang, J.; Xie, Y.; Xu, Y.; Xu, Z.; Ye, S.; Yu, Z.; Zhang, A.; Zhang, H.; Zhao, Y.; Zheng, R.; and Zhu, Y. 2025. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv:2503.14734. Tang, K.; Jia, P.; Chu, Z.; Wu, J.; Ma, R.; Cao, J.; Zhao, F.; Chen, S.; Guo, Y.; Chi, X.; Fan, C.-K.; Zhang, K.; Xu, J.; Yang, F.; Mi, W.; Ju, X.; Tang, J.; and Zhang, S. 2026. SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model. arXiv:2606.20698. Wei, Y.; Wu, C.; and Haghbayan, H. 2026. Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study. arXiv:2604.17896. Ye, S.; Ge, Y.; Zheng, K.; Gao, S.; Yu, S.; Kurian, G.; Indupuru, S.; Tan, Y. L.; Zhu, C.; Xiang, J.; Malik, A.; Lee, K.; Liang, W.; Ranawaka, N.; Gu, J.; Xu, Y.; Wang, G.; Hu, F.; Narayan, A.; Bjorck, J.; Wang, J.; Kim, G.; Niu, D.; Zheng, R.; Xie, Y.; Wu, J.; Wang, Q.; Julian, R.; Xu, D.; Du, Y.; Chebotar, Y.; Reed, S.; Kautz, J.; Zhu, Y.; Fan, L.; and Jang, J. 2026. World Action Models are Zero-shot Policies. arXiv:2602.15922. Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026. Fast-WAM: Do World Action Models Need Test-time Future Imagination? arXiv:2603.16666. Zhang, B.; Zhang, Y.; Ji, J.; Lei, Y.; Dai, J.; Chen, Y.; and Yang, Y. 2026. Safevla: Towards safety alignment of visionlanguage-action model via constrained learning. Advances in Neural Information Processing Systems, 38: 153335– 153373.