Occluding the Solution Space: Planner-Agnostic Adversarial Attacks on Tolerance-Aware Manipulation
arXiv:2607.03758v1 [cs.RO] 4 Jul 2026
Keke Tang1,∗,† , Tianyu Hao1,∗ , Weilong Peng1,† , Hao Jiang2 , Feng Wu2 , Peican Zhu3,† , Jianmin Ji2 , and Zhihong Tian1
Abstract— Adversarial attacks on motion planning are crucial for evaluating and quantifying the intrinsic robustness of robotic manipulation. However, existing approaches are typically limited by restrictive exact-pose objectives and their reliance on planner-in-the-loop queries. To address these limitations, we propose a planner-agnostic attack framework for tolerance-aware manipulation. Our approach shifts the evaluation paradigm to task-level feasibility over goal regions, efficiently inserting adversarial obstacles without requiring oracle access to the victim system. Offline, we characterize the robot’s intrinsic workspace capabilities via a kinematic occupancy heatmap, which encodes the density of feasible trajectories and robustness priors without invoking a specific planner. Online, we formulate the attack as a budgeted maximum-coverage optimization, strategically deploying obstacles subject to explicit geometric constraints to occlude the solution space. Extensive experiments across simulation and real-world scenarios demonstrate that our method reliably induces planning failures, significantly outperforming planner-in-the-loop baselines in both computational efficiency and attack efficacy.
I. I NTRODUCTION Robotic manipulation relies fundamentally on motion planning to translate high-level task specifications into safe and feasible arm trajectories [1], [2], [3]. In safety-critical deployments, the reliability of this planning process is nonnegotiable, as failures can lead to operational hazards or costly interruptions. Adversarial attacks, which utilize targeted environmental modifications to provoke system failures, provide a rigorous framework for probing these safety boundaries. By systematically generating worst-case scene configurations, such methods serve as a necessary stress test to quantify the intrinsic robustness of planning algorithms. Despite their importance, adversarial evaluations targeting classical manipulation planners remain underexplored. Prior investigations have predominantly focused on perceptioninduced failures [4]. The limited body of work addressing planning robustness typically enforces highly restrictive success criteria, such as tracking specific pose targets within ∗ These authors contributed equally. † Corresponding authors. 1 Keke Tang, Tianyu Hao, Weilong Peng, and Zhihong Tian are with
Guangzhou University, Guangzhou 510006, China. 2 Hao Jiang, Feng Wu, and Jianmin Ji are with the University of Science and Technology of China, Hefei 230026, China. 3 Peican Zhu is with Northwestern Polytechnical University, Xi’an 710072, China. This work was supported in part by the National Natural Science Foundation of China (62472117, 62572400, U2436208, 62372129), the Guangdong Basic and Applied Basic Research Foundation (2025A1515010157, 2024A1515012064), the CCF-NetEase ThunderFire Innovation Research Funding (CCF-Netease 202514), the Science and Technology Projects in Guangzhou (2025A03J0137, 2024B0101010002).
Fig. 1. Concept of the proposed planner-agnostic attack on toleranceaware manipulation. Unlike conventional exact-pose grasping that targets fixed points (top-left), we consider a more realistic tolerance-aware task where the goal is a continuous region (bottom-left). Our framework (right) generates adversarial obstacles to occlude the robot’s potential motion paths without requiring oracle access to the victim’s internal planner.
tight tolerances [5]. However, these exact-pose objectives confound task-level feasibility: a plan functionally fails only if it misses the grasp region entirely, not merely a specific coordinate. Furthermore, existing approaches commonly rely on planner-in-the-loop optimization [5]. This dependency grants the adversary unrealistic oracle-like access to the victim planner, rendering the threat model impractical for black-box security assessments. To address these limitations, we argue that a practical adversarial framework must be fundamentally rethought across two critical dimensions: the task objective and the threat model. From a task perspective, manipulation is inherently tolerance-aware: operational success depends on reaching a feasible goal region, allowing the planner to utilize a continuous manifold of valid configurations, see Fig. 1. Consequently, an effective attack must occlude this entire solution space rather than merely obstructing a singular path. From a threat model perspective, a realistic adversary must be planner-agnostic, operating without oracle access to the victim system. Instead, we consider an attacker that leverages only public robot kinematics and static scene observability to deploy a limited budget of obstacles subject to explicit geometric constraints, see Fig. 1. In this paper, we propose a two-stage, planner-agnostic framework to execute adversarial attacks on manipulation tasks by systematically occluding the feasible solution space. In the offline stage, we characterize the robot’s intrinsic capabilities by constructing a kinematic occupancy heatmap. This volumetric prior encodes the density of valid manipulation trajectories, identifying critical spatial bottlenecks without invoking any specific planner. During the online
stage, we formulate the attack as a budgeted maximumcoverage optimization within a target-aligned search region. By strategically deploying obstacles, our method maximizes the occlusion of the task-level solution manifold. This approach efficiently yields robust adversarial constraints, bypassing the computational burden of iterative planner queries while ensuring rigorous geometric validity. Extensive experiments across diverse simulation environments and real-world robotic deployments validate the effectiveness and efficiency of the proposed method. Overall, our contributions are summarized as follows: • We formulate a task-level adversarial threat model for tolerance-aware manipulation, shifting the attack objective from targeting exact-pose configurations to occluding the entire feasible goal region. • We propose a two-stage, planner-agnostic framework that efficiently generates geometric constraints by combining an offline kinematic occupancy heatmap with an online budgeted maximum-coverage optimization. • We extensively validate our approach across diverse simulation environments and real-world robotic deployments, demonstrating superior attack efficacy and efficiency compared to state-of-the-art methods. II. R ELATED W ORK A. Adversarial Attacks on Robotic Systems Adversarial attacks have been extensively studied in computer vision and successfully extended to the robotics domain, particularly in 3D perception [6], [7], [8], [9], [10]. In the context of robotic manipulation, existing attacks predominantly target the grasp detection and evaluation stages. For instance, prior work [11] misleads image-based grasp evaluation networks via pixel-level perturbations or barely visible patches. Alternatively, other approaches [12], [13] tackle the problem by altering the target object’s physical geometry to diminish its overall graspability. Crucially, these paradigms aim to invalidate the grasp target itself. In contrast, our work addresses the oftenoverlooked vulnerability in the motion planning phase, where adversarial environmental modifications can obstruct the physical path necessary to approach and execute the task. B. Adversarial Attacks on Manipulation Planning Adversarial attacks on manipulation planning remain underexplored. Existing work induces failures by corrupting state estimation [4] or exploiting semantic vulnerabilities in high-level VLM-based agents [14], [15] and VLA models [16]. These methods perturb inputs rather than evaluate planner resilience to adversarially structured environments. Directly attacking planning algorithms via environmental constraints is a sparse field. To our knowledge, the only closely related work is by Wu et al. [5], which not only requires a planner-in-the-loop approach but also relies on rigid exact-pose objectives that overlook task-level tolerance. In contrast, our planner-agnostic framework addresses this by efficiently generating obstacles to occlude the solution space without requiring oracle access.
III. P ROBLEM F ORMULATION A. Tolerance-Aware Manipulation Planning We consider the problem of motion planning for manipulation, where the objective is to reach a goal region rather than a singular pose. Let Q ⊂ Rn denote the configuration space of an n-DoF manipulator, and let E denote the nominal static 3D scene geometry. The tolerance-aware goal is specified by a spatial subset G ⊂ SE(3), representing the volume of acceptable end-effector poses. This task constraint induces a feasible goal set in the joint space, denoted as QG = {q ∈ Q | FK(q) ∈ G}. The planning task is to find a collision-free path starting from a fixed qinit and terminating at any configuration qend that satisfies: qend ∈ QG . (1) Crucially, unlike classical point-to-point planning where the goal is a singleton, QG forms a continuous manifold. This redundancy provides the planner with multiple valid solutions, making the task inherently more resilient to simple constraints. B. Adversarial Threat Model The objective of the adversary is to force a task failure by strategically modifying the environment. Given the redundancy defined in Sec. III-A, a successful attack must effectively occlude the entire solution space that connects qinit to the manifold QG . We formulate this as a constrained geometric optimization problem. a) Attack Action: The adversary generates a set of M obstacles O(θ) = {o1 (θ1 ), . . . , oM (θM )} parameterized by θ, resulting in a perturbed environment: Eθ := E ∪ O(θ).
(2)
The attack is successful if, for a given environment Eθ , no path τ satisfies the conditions in Eq. (1). b) Insertion Constraints: To ensure the attack remains physically realistic and maintains scene consistency, the parameters θ are subject to the following constraints: C1) Cardinality Budget: The adversary is restricted to inserting at most M distinct obstacles. C2) Geometric Bounds: Each obstacle om is defined by standard geometric primitives (e.g., boxes or spheres) with a maximum size. C3) Placement Feasibility: Insertions must lie within a valid search region S ⊂ Efree and maintain a minimum clearance from qinit and G to avoid trivial infeasibility. c) Adversary Capabilities: We assume a planneragnostic adversary operating under the following conditions: A1) Known Kinematics: The adversary has access to the robot’s kinematic model and collision geometry. A2) Scene Observability: The adversary perceives the nominal scene E and the task specification G. A3) No Oracle Access: The adversary cannot query the victim planner or simulate the full planning process during
Fig. 2. Framework of the planner-agnostic attack. The offline phase constructs a kinematic occupancy heatmap to characterize workspace importance. The online phase then leverages this heatmap to perform budgeted maximum-coverage optimization for generating adversarial obstacles (red boxes) that disrupt tolerance-aware manipulation planning.
attack generation. This ensures the attack exploits intrinsic geometric vulnerabilities rather than over-fitting to a specific algorithm’s internal logic. IV. M ETHODOLOGY We propose a planner-agnostic framework to generate adversarial constraints that occlude the workspace volume most critical to the solution space of tolerance-aware manipulation. Our approach operates in two stages: (i) an offline phase that constructs a kinematic occupancy heatmap characterizing the robot’s intrinsic capabilities; (ii) an online phase that formulates the attack as a budgeted maximumcoverage optimization to strategically place obstacles. Please refer to Fig. 2 for an illustration. A. Offline: Kinematic Occupancy Heatmap Construction The offline stage maps the high-dimensional configuration space Q into a workspace scalar field H. This field identifies “spatial bottlenecks”—regions where valid and robust arm configurations are most densely concentrated. 1) Workspace Voxelization: The workspace is discretized into a voxel grid V with resolution δ×δ×δ. Let Vfree ⊂ Efree denote the set of voxels not occupied by static geometry E. We initialize a heatmap H : Vfree → R≥0 to zero, which will serve as the importance measure for obstacle placement. 2) Robust Configuration Sampling: We sample a large set of configurations q ∈ Q and retain only those that satisfy self-collision and scene-collision constraints: valid(q; E) = True. To ensure the heatmap reflects regions favored by robust motion planners, each valid sample is assigned a weight w(q) combining two quality metrics: w(q) = sman (q)γ1 · sclr (q)γ2 ,
(3)
where γ1 , γ2 ≥ 0 are hyperparameters. a) Manipulability Score p (sman ): We utilize the Yoshikawa index, m(q) = det(J(q)J(q)⊤ ), where J(q) is the geometric Jacobian. We apply a saturation function to bound the score: sman (q) = m(q)/(m(q) + κ). The purpose of this score is to identify regions where the robot possesses high control authority. By prioritizing these regions,
the adversary targets configurations furthest from kinematic singularities, which motion planners preferentially utilize to maintain trajectory responsiveness and maneuverability. b) Clearance-Robustness Score (sclr ): To prioritize configurations in open workspace areas, we compute the minimum distance between the robot’s occupied volume B(q) and static obstacles E, denoted dclr (q) = dist(B(q), E). This is normalized using a safety margin dsafe : sclr (q) = min(dclr (q)/dsafe , 1) . (4) By occluding these high-clearance regions, which planners naturally favor, the adversary effectively forces the robot into narrow, constrained corridors where planning is harder. 3) Swept-Volume Aggregation: For each weighted sample q, we let B(q) ⊂ R3 denote the volume occupied by the robot’s kinematic links. We project this volume onto the grid and accumulate the weight: H(v) ← H(v) + w(q),
∀v ∈ Voxelize(B(q)) ∩ Vfree . (5)
The resulting H encodes the density of robust trajectories within the workspace. B. Online: Adversarial Occupancy Maximization Given qinit and G, the adversary seeks to instantiate an obstacle set O(θ) that maximally covers the heatmap mass relevant to the current task. 1) Task-Relevant Candidate Selection: To satisfy placement feasibility (Constraint C3), we define a candidate set Xcand restricted to a task-aligned corridor connecting xstart and G. Voxels within a safety margin dsafe of qinit or G are pruned to avoid trivial infeasibility. 2) Budgeted Maximum-Coverage Optimization: We model obstacle placement as a submodular maximization problem [17]. Let U(c) be the voxels occupied by a primitive centered at c. The adversary solves: X H(v) max C⊆Xcand S v∈ c∈C U (c) (6) s.t. |C| ≤ M, ∀ci , cj ∈ C : dist(ci , cj ) ≥ dmin .
Algorithm 1 Adversarial Occupancy Maximization Require: Heatmap H; States qinit , G; Budget M ; Primitive size s Ensure: Adversarial Obstacles O 1: S ← DefineSearchRegion(qinit , G, dsafe ) ▷ Constraint C3 2: X ← {v ∈ Vfree | H(v) > 0} ∩ S 3: O ← ∅, Vcovered ← ∅ 4: for m = 1 to M do 5: if X = ∅ then break 6: end if P 7: c∗ ← arg maxc∈X v∈U (c)\Vcovered H(v) 8: O ← O ∪ {Primitive(c∗ , s)} 9: Vcovered ← Vcovered ∪ U(c∗ ) 10: X ← {c ∈ X | ∥c − c∗ ∥ ≥ dmin } ▷ Spatial diversity 11: end for 12: return O The constraint dmin ensures spatial diversity by preventing obstacles from overlapping excessively. This forces the generated constraints to obstruct multiple potential trajectories rather than a single path. 3) Greedy Attack Generation: The optimization in Eq. (6) is NP-hard; however, since the objective is submodular, a greedy strategy yields a (1 − 1/e) approximation. The iterative selection process is summarized in Algorithm 1. By selecting the center c∗ that maximizes the marginal heatmap coverage in each step, the adversary efficiently constructs a global constraint field that occludes the robot’s most robust motion options. V. E XPERIMENTAL R ESULTS We evaluate the proposed planner-agnostic adversarial framework in both high-fidelity simulation and real-robot deployments. Our experiments are designed to answer: (RQ1) How effective is the attack against classical planners? (RQ2) What obstacle budget is required to induce specific failure rates? (RQ3) Can kinematics-derived obstacles transfer to vision-language-action (VLA) policies? (RQ4) How does our efficiency compare to planner-in-the-loop baselines? (RQ5) Can the framework transfer to physical robots under real sensing and modeling noise? A. Experimental Setup a) Robotic Platform and Simulation: We use a 7-DoF Franka Emika Panda as the target platform. Simulations are conducted in NVIDIA Isaac Sim [18], which provides highfidelity physics and collision checking. The planning stack is built on ROS 2 [19] and MoveIt 2 [20] with OMPL [21] planners. Forward kinematics and Jacobians are parsed and evaluated via the Pinocchio library. b) Scene Design: We generate five tabletop manipulation scenes (Scene 1–5) using MotionBenchMaker [22], covering varying clutter patterns and goal regions. Due to space constraints, we report detailed quantitative benchmarks
on three representative scenes (Scene 1–3), while comprehensive qualitative evaluations across all five scenes are provided in Fig. 3. c) Planners and Protocol: To demonstrate planner agnosticism, we evaluate three widely used sampling-based planners: RRT [1], PRM∗ [2], and BKPIECE [23]. Each planner is given a strict time budget of 5.0 s with a single attempt per query. For each (scene, planner, method, budget) configuration, we report averages over 500 independent runs. d) Implementation Details: We construct the kinematic occupancy heatmap with a voxel resolution of δ = 0.02 m. Adversarial obstacles are modeled as axis-aligned box primitives with a fixed edge length of s = 0.05 m and a constant orientation. For the offline heatmap construction, we sample Noff = 15,000 collision-free joint configurations. The discrete workspace volume V is defined by the bounding box of the robot’s swept volume across these valid samples, augmented with a padding margin to fully encompass taskrelevant regions. For the heuristic weights in Eq. (3), we set the manipulability saturation parameter to κ = 0.01 and use equal weighting by setting γ1 = γ2 = 1. During the online candidate generation, we prune voxels within a safety margin of dsafe = 0.1 m from both qinit and the goal region G. Additionally, candidate centers are restricted to a taskaligned corridor of radius rcorr = 0.3 m connecting the start position xstart to G. To ensure spatial diversity among the generated constraints, we enforce a minimum separation of dmin = 0.05 m between obstacle centers. We evaluate our framework under budgets of 1, 3, and 5 obstacles. e) Baselines: We consider three baselines. • Random: obstacles uniformly sampled in the valid search region. • PFA: the planner-in-the-loop physical attack of Wu et al. [5], which iteratively queries a target planner to optimize obstacle placement. • PFA++: our tolerance-aware adaptation of PFA for fair comparison. Since PFA internally uses an RRTbased evaluation, PFA and PFA++ are mathematically equivalent when evaluated on RRT. f) Evaluation Metrics: We report five metrics capturing attack success and the quality of the surviving plans: • Planning Success Rate (PSR): Fraction of successful plans out of 500 trials; lower values indicate a more effective attack. −2 • Singularity Margin (SMR, ×10 p ): Minimum Yoshikawa manipulability m(q) = det(J(q)J(q)⊤ ) along a successful trajectory; lower values indicate the planner is driven closer to kinematic singularities. −2 • Clearance Margin (CMR, ×10 ): Estimated lower bound of joint-space distance to collision via randomized bi-directional search; lower values indicate narrower free-space corridors. • Joint Deviation Score (JDS): Average joint-space deviation from the initial configuration along the trajectory; higher values indicate larger detours forced by the
•
adversarial obstacles. Planning Time (PT): Total computation time required to return a planning result; failures are recorded as the maximum budget (5.0 second).
B. Attack Performance Against Classical Planners (RQ1) a) Overall Trends: Table I shows that our method consistently achieves the lowest PSR across almost all scene– planner combinations, and its advantage becomes dominant as the obstacle budget increases. This behavior aligns with our design: rather than optimizing against a specific planner’s sampling pattern, we directly target high-occupancy swept volumes that are shared across feasible configurations. As a result, obstacles placed by our method tend to obstruct the task-level feasible manifold more globally. b) Performance Across Obstacle Budgets: At a low budget (1 obstacle), absolute gaps between methods can be small in easy settings, but our framework already demonstrates consistent improvements in difficult cases. For instance, on PRM∗ in Scene 3, Random and PFA maintain a PSR close to 1, indicating that a single randomly placed or planner-coupled obstacle rarely hits a true bottleneck. In contrast, our method reduces the PSR to 0.712, proving that the heatmap can identify structurally important passages even with a single insertion. As the budget increases to 3 obstacles, this advantage is significantly amplified. In Scene 1 under BKPIECE, our PSR drops to 0.224 while the best baseline remains above 0.67. Finally, at a high budget (5 obstacles), our method achieves complete occlusion (PSR = 0) across all nine scene-planner configurations. In contrast, plannerin-the-loop baselines still leave notable residual success. This performance gap highlights that obstacles derived from kinematic occupancy can effectively cover global workspace bottlenecks, whereas iterative planner-coupled optimization often overfits to specific sampling patterns and misses alternative feasible corridors. c) Progressive Degradation of Plan Quality: Even before reaching PSR = 0, the auxiliary metrics demonstrate a consistent pattern of trajectory degradation: (i) CMR decreases, indicating that the remaining feasible motions are squeezed into precarious, narrow passages; (ii) SMR decreases, showing that the planner is driven toward less dexterous, near-singular postures; (iii) JDS increases, implying longer detours; and (iv) PT increases, reflecting substantially harder search scenarios. For example, under BKPIECE in Scene 3, increasing our budget from 1 to 3 reduces CMR from 1.808 to 0.942 while PT sharply increases, revealing that even when the victim planner succeeds, it does so under severely compromised safety and kinematic margins. C. Attack Efficiency on Classical Planners (RQ2) a) Interpretation of Budget Thresholds: Table II provides a direct “budget-to-failure” perspective, which is often more informative for safety evaluations than reporting PSR under a fixed budget. Across all scenes and planners, our method reaches each failure threshold using substantially
fewer obstacles. For example, to completely block RRT in Scene 1 (PSR = 0%), Random requires 50.8 obstacles and PFA requires 12.4, whereas our method requires only 4.1. In the most challenging configuration (Scene 3 under PRM∗ ), complete occlusion demands 93.2 obstacles for Random, 10.9 for PFA, and 10.0 for PFA++, compared to a mere 4.4 for our method. Crucially, our approach never requires more than 5 obstacles to achieve absolute planning failure across any tested configuration. This validates our core premise: a small number of optimally placed constraints over the kinematic occupancy field is sufficient to completely paralyze the manipulator. D. Transferability to VLA-based Policies (RQ3) a) Transfer Results and Implications: Table III reports the attack performance on two representative VLA policies, OpenVLA [24] and π0.5 [25]. Our obstacles are generated without access to the victim policy’s network architecture, training data, or gradients, relying solely on robot kinematic occupancy. Despite this strict black-box setting, the attack remains highly effective. Against OpenVLA, a single obstacle already reduces the PSR to 0.340 (compared to 0.364 for PFA). With three obstacles, our method achieves complete occlusion (PSR = 0), whereas PFA still leaves a residual success rate of 0.188. Against π0.5 , our method yields a PSR of 0.142 under a 3-obstacle budget, substantially outperforming PFA (0.426) and PFA++ (0.390). At a budget of 5 obstacles, the success rate collapses entirely. These results suggest that kinematic bottlenecks constitute a shared physical vulnerability across classical sampling-based planners and end-to-end learned policies: while their decision-making mechanisms differ, both ultimately operate within the exact same geometric and kinematic feasibility constraints. b) Qualitative Failure Modes of VLA Policies: We additionally observe qualitatively different failure behaviors compared to classical planners. When the task becomes infeasible, classical planners typically return failure explicitly within the planning budget. In contrast, attacked VLA policies often exhibit progressive, cascading failures: the robot first slows down and hesitates near the obstacle, then drifts into contact, and finally enters oscillatory or deadlocked states, see supplementary video. This highlights that current VLA policies may lack explicit geometric feasibility reasoning when the solution manifold is adversarially occluded. E. Computational Efficiency (RQ4) Table IV compares the computational runtimes across different attack methods. Our framework strategically shifts the heavy computational burden to a one-time offline heatmap construction (≈ 1251 s, visualized in Fig. 4), enabling extremely fast online attack generation (≈ 1.27 s). This paradigm contrasts sharply with planner-in-the-loop attacks: PFA absorbs its entire computational cost during the online phase, requiring ≈ 1224 s of iterative optimization for every new task query. PFA++ can be even more exorbitant, requiring ≈ 3978 s when coupled with PRM∗ . Crucially, because
TABLE I Q UANTITATIVE EVALUATION OF ATTACK PERFORMANCE ACROSS DIFFERENT SCENES , PLANNERS , AND OBSTACLE BUDGETS . SMR AND CMR VALUES ARE SCALED BY 10−2 . “–” DENOTES NON - APPLICABLE CONFIGURATIONS ; “N/A” INDICATES METRICS UNDEFINED DUE TO COMPLETE FAILURE (PSR = 0). B EST VALUES ARE IN BOLD . Scene
Scene 1
Scene 2
Scene 3
Planner
1 Obstacle
Method
3 Obstacles
5 Obstacles
PSR
SMR
CMR
JDS
PT
PSR
SMR
CMR
JDS
PT
PSR
SMR
CMR
JDS
PT
RTT
Random PFA PFA++ Ours
88.00 89.20 – 79.60
5.008 5.081 – 4.850
3.191 3.261 – 3.139
2.094 2.779 – 2.806
0.659 0.522 – 0.681
87.00 76.40 – 31.80
4.906 5.163 – 4.105
3.083 3.261 – 0.777
2.143 2.042 – 2.263
0.518 1.145 – 3.162
85.10 53.00 – 0
5.255 5.217 – N/A
2.638 1.680 – N/A
2.105 1.958 – N/A
0.520 2.355 – 4.992
PRM∗
Random PFA PFA++ Ours
81.00 80.20 78.60 80.60
6.281 6.765 6.544 6.194
3.571 3.286 3.289 3.274
1.624 1.582 1.619 1.645
4.980 4.990 4.980 4.990
79.40 82.20 79.80 38.60
6.249 6.348 6.244 4.346
3.226 3.275 3.143 0.871
1.609 1.582 1.661 2.291
4.970 4.990 4.950 5.000
82.20 75.20 71.20 0
6.384 6.054 5.876 N/A
3.520 1.865 1.820 N/A
1.620 1.589 1.592 N/A
4.970 4.940 5.000 5.000
BKPIECE
Random PFA PFA++ Ours
77.00 72.40 71.60 66.40
5.519 5.341 5.203 5.476
3.006 3.441 3.389 2.688
2.068 2.005 2.029 2.090
0.500 0.756 0.771 0.795
77.20 67.80 69.40 22.40
5.112 5.262 5.251 4.154
3.068 2.861 2.844 0.727
2.015 1.995 2.026 2.277
0.484 1.235 1.242 3.326
79.40 52.80 51.00 0
5.463 4.837 4.664 N/A
2.975 1.424 1.399 N/A
2.062 1.979 2.190 N/A
0.523 2.146 2.208 5.000
RTT
Random PFA PFA++ Ours
84.80 78.20 – 70.60
4.843 4.605 – 4.798
3.289 2.317 – 2.608
2.245 2.319 – 2.206
0.540 0.565 – 0.654
84.20 69.20 – 59.20
4.667 4.380 – 4.740
3.613 1.343 – 1.162
2.225 2.405 – 2.463
0.509 1.684 – 1.739
84.40 48.60 – 0
4.814 3.949 – N/A
3.280 1.090 – N/A
2.229 2.641 – N/A
0.554 1.424 – 5.000
PRM∗
Random PFA PFA++ Ours
78.00 73.80 71.90 74.60
5.905 6.134 5.893 6.304
3.007 2.680 2.471 2.285
1.805 1.912 1.962 1.802
4.980 4.990 5.000 5.000
79.80 67.20 64.00 60.80
6.016 5.834 5.793 5.718
3.031 2.042 2.034 1.943
1.791 2.041 2.138 2.199
4.686 4.972 5.000 5.000
74.60 49.00 47.60 0
6.002 5.258 5.170 N/A
2.816 1.654 1.598 N/A
1.844 2.365 2.471 N/A
5.000 5.000 5.000 5.000
BKPIECE
Random PFA PFA++ Ours
75.40 66.00 63.20 61.80
4.165 4.152 4.062 4.038
2.256 1.835 1.772 1.756
2.293 2.382 2.417 2.242
0.548 0.537 0.620 0.645
75.00 53.40 51.00 37.00
4.379 4.104 4.194 4.067
2.221 1.299 1.196 1.144
2.267 2.394 2.400 2.429
0.546 1.797 1.902 2.108
73.20 39.80 37.40 0
4.290 3.663 3.599 N/A
2.142 1.030 1.019 N/A
2.329 2.590 2.370 N/A
0.549 1.720 1.919 5.000
RTT
Random PFA PFA++ Ours
95.40 92.40 – 83.20
5.051 4.959 – 4.265
4.100 4.200 – 1.960
1.852 1.845 – 2.031
0.451 0.472 – 0.503
94.40 78.80 – 65.00
4.981 4.919 – 4.132
3.973 1.996 – 0.901
1.786 1.921 – 2.269
0.502 1.297 – 1.724
94.40 54.20 – 0
4.667 4.346 – N/A
3.544 2.459 – N/A
1.891 1.930 – N/A
0.527 2.504 – 4.999
PRM∗
Random PFA PFA++ Ours
98.40 98.80 95.20 71.20
6.899 6.768 6.506 6.113
4.513 4.732 4.546 1.832
1.065 1.062 1.177 1.319
4.874 4.799 4.836 4.972
98.80 86.20 84.40 47.40
6.780 6.849 6.709 5.246
4.593 1.916 1.860 0.878
1.066 1.222 1.404 1.848
4.841 4.871 4.941 4.836
99.00 59.60 57.40 0
6.742 5.635 5.450 N/A
4.486 2.580 2.600 N/A
1.070 1.142 1.203 N/A
4.894 4.990 5.000 5.000
BKPIECE
Random PFA PFA++ Ours
95.80 94.00 91.80 80.60
5.257 5.248 5.086 4.082
4.111 3.873 3.667 1.808
1.820 1.779 1.808 2.132
0.513 0.462 0.501 0.547
92.00 69.60 64.40 55.20
5.138 4.963 4.653 4.353
3.814 1.747 1.792 0.942
1.785 1.885 1.842 2.178
0.539 1.478 1.505 2.068
93.40 41.60 39.50 0
5.215 4.882 4.887 N/A
3.613 2.438 2.338 N/A
1.780 2.081 1.825 N/A
0.541 2.531 2.781 4.742
Fig. 3.
Qualitative visualization of adversarial obstacle placement in five tabletop scenes under budgets of 1, 3, and 5.
TABLE II C OMPARATIVE ANALYSIS OF ATTACK EFFICIENCY. M INIMUM OBSTACLE BUDGET REQUIRED TO REDUCE THE PSR TO SPECIFIED THRESHOLDS . Scene
Scene 1
Scene 2
Scene 3
Planner
Success Rate
Method 80%
60%
40%
20%
0%
RTT
Random PFA PFA++ Ours
6.2 2.5 – 1.0
15.7 4.5 – 1.9
25.6 6.5 – 2.6
35.3 9.0 – 3.6
50.8 12.4 – 4.1
PRM∗
Random PFA PFA++ Ours
2.5 1.5 1.2 1.1
12.4 8.6 6.3 1.8
20.2 10.2 9.6 2.9
38.4 14.1 13.7 3.8
55.6 18.4 16.6 4.7
BKPIECE
Random PFA PFA++ Ours
1.0 1.0 1.0 1.0
12.2 4.5 4.8 1.3
19.5 6.3 6.8 2.2
27.4 9.9 9.5 3.2
45.8 12.3 13.2 4.7
RTT
Random PFA PFA++ Ours
8.0 1.0 – 1.0
19.7 3.7 – 2.8
23.9 5.9 – 3.6
31.3 8.5 – 4.3
48.4 12.2 – 4.8
PRM∗
Random PFA PFA++ Ours
1.0 1.0 1.0 1.0
20.7 3.8 4.3 3.1
25.5 6.0 5.8 3.7
35.0 8.4 8.0 4.4
51.2 11.6 10.7 4.9
BKPIECE
Random PFA PFA++ Ours
1.0 1.0 1.0 1.0
13.0 2.0 1.8 1.1
22.8 4.9 4.5 2.8
30.4 7.8 7.1 3.9
39.4 10.3 9.5 4.5
RTT
Random PFA PFA++ Ours
14.9 2.8 – 1.3
32.1 4.4 – 3.4
42.7 6.3 – 4.1
69.7 8.7 – 4.6
91.1 11.9 – 4.9
PRM∗
Random PFA PFA++ Ours
21.0 3.6 3.4 1.0
39.7 4.9 4.8 1.9
54.8 6.7 6.4 3.3
71.7 8.5 8.0 4.2
93.2 10.9 10.0 4.4
BKPIECE
Random PFA PFA++ Ours
10.5 2.2 1.9 1.0
21.3 3.6 3.3 2.6
40.9 5.2 4.9 3.5
54.0 7.2 6.8 4.3
78.0 9.3 8.5 4.2
Fig. 4. Visualization of the kinematic occupancy heatmap (2D slice). Each voxel stores the swept-volume frequency under collision-free configurations, enabling budgeted maximum-coverage obstacle placement.
the offline heatmap is tied to the robot’s intrinsic kinematics rather than specific scene layouts, it is constructed once per robotic platform and reused indefinitely across varying tasks and goal regions. This massive amortized advantage makes our approach uniquely suited for near real-time safety evaluations and large-scale robustness benchmarking. F. Real-World Robotic Deployments (RQ5) To validate the sim-to-real transferability of our framework under realistic sensing constraints, we deployed the attack on a physical Rokae xMatePro7 manipulator (Fig. 5). The planning scene is constructed directly from real-world RGB-
Fig. 5. Real-robot deployment on Rokae xMatePro7. (a) Original scene without adversarial obstacles. (b–d) Adversarial scenes with 1/3/5 obstacles.
D sensing. Specifically, an Intel RealSense camera captures the workspace, from which we reconstruct a geometric collision mesh (visualized as the green overlay in MoveIt 2). Our adversarial obstacle placement is then executed over this sensed mesh, ensuring the attack is strictly grounded in the robot’s actual perception. To physically realize the attack, the generated adversarial obstacles are instantiated using rigidly mounted box primitives, ensuring spatial consistency with the computed geometric constraints while controlling for unwanted motion during repeated trials. The physical deployments closely mirror our simulation findings. As shown in Fig. 5(a), without adversarial obstacles, the robot smoothly plans and executes a collisionfree trajectory to the goal region. When a single adversarial obstacle is introduced (b), the planner may still succeed but is forced into a narrow, suboptimal path, visibly increasing joint deviations and execution complexity. However, as the obstacle budget increases to three and five (c–d), the planner consistently fails to find a feasible solution within the time budget and safely aborts. This behavior perfectly matches the “complete occlusion” phenomenon observed in simulation, confirming that our geometrically derived attack remains highly robust and effective in the real world despite raw sensor noise and execution uncertainties. VI. C ONCLUSION AND D ISCUSSION In this paper, we have presented a planner-agnostic adversarial framework targeting the robustness of tolerance-aware manipulation planning. By shifting the focus from rigid goal poses to task-level solution manifolds, our approach utilizes an offline kinematic occupancy heatmap to identify critical workspace regions and an online budgeted maximumcoverage optimization to strategically place obstacles. Unlike
TABLE III Q UANTITATIVE COMPARISON OF ATTACK TRANSFERABILITY TO VLA POLICIES (O PEN VLA AND π0.5 ). Model
1 Obstacle
Method
3 Obstacles
PSR
SMR
CMR
JDS
Random PFA(RRT) OpenVLA PFA++(PRM*) Ours
45.60 36.40 35.80 34.00
5.468 5.312 5.122 5.335
0.951 0.848 0.813 0.792
1.832 43.60 5.530 0.825 1.901 43.20 5.408 0.851 2.002 2.057 18.80 5.309 0.790 2.152 0 N/A N/A N/A 2.029 17.90 5.226 0.771 2.190 0 N/A N/A N/A 2.003 0 N/A N/A N/A 0 N/A N/A N/A
Random PFA(RRT) PFA++(PRM*) Ours
58.60 51.20 50.20 47.40
5.612 5.679 5.601 5.323
1.310 1.103 1.154 1.047
1.988 1.857 1.973 2.034
Pi 0.5
TABLE IV Q UANTITATIVE COMPARISON OF OFFLINE AND ONLINE RUNTIMES ( IN SECONDS ) ACROSS DIFFERENT METHODS . Method Random PFA PFA++ (PRM*) PFA++ (BKPIECE) Ours
Offline Time Online Time Total Time 0.10 – – – 1251.47
– 1224.75 3978.42 1326.67 1.27
0.10 1224.75 3978.42 1326.67 1252.74
prior methods, our framework operates without oracle access to the victim’s internal logic, ensuring high efficiency and broad applicability. Extensive experiments in both simulation and real-world scenarios demonstrate that our method effectively occludes the task-level solution space, significantly outperforming planner-in-the-loop baselines in both success rate and computational speed. Our results suggest that kinematic bottlenecks form a planner- and policy-agnostic vulnerability surface for tolerance-aware manipulation. Beyond attack generation, the proposed framework can serve as a diagnostic tool to expose where feasibility concentrates and how quickly it collapses under bounded perturbations, providing concrete signals for improving robustness. Future work will extend the framework to dynamic environments. R EFERENCES [1] J. J. Kuffner and S. M. LaValle, “Rrt-connect: An efficient approach to single-query path planning,” in ICRA, pp. 995–1001, 2000. [2] S. Karaman and E. Frazzoli, “Sampling-based algorithms for optimal motion planning,” IJRR, vol. 30, no. 7, pp. 846–894, 2011. [3] N. Lin, Y. Li, K. Tang, Y. Zhu, X. Zhang, R. Wang, J. Ji, X. Chen, and X. Zhang, “Manipulation planning from demonstration via goalconditioned prior action primitive decomposition and alignment,” RAL, vol. 7, no. 2, pp. 1387–1394, 2022. [4] J. Sadeghi, N. A. Lord, J. Redford, and R. Mueller, “Attacking motion planners using adversarial perception errors,” arXiv preprint arXiv:2311.12722, 2023. [5] W. Wu, F. Pierazzi, Y. Du, and M. Brandão, “Characterizing physical adversarial attacks on robot motion planners,” in ICRA, pp. 14319– 14325, 2024. [6] C. Xiang, C. R. Qi, and B. Li, “Generating 3d adversarial point clouds,” in CVPR, pp. 9136–9144, 2019. [7] K. Tang, J. Wu, W. Peng, Y. Shi, P. Song, Z. Gu, Z. Tian, and W. Wang, “Deep manifold attack on point clouds via parameter plane stretching,” in AAAI, vol. 37, pp. 2420–2428, 2023. [8] K. Tang, Z. Wang, W. Peng, L. Huang, L. Wang, P. Zhu, W. Wang, and Z. Tian, “Symattack: Symmetry-aware imperceptible adversarial attacks on 3d point clouds,” in MM, pp. 3131–3140, 2024. [9] K. Tang, Y. Shi, T. Lou, W. Peng, X. He, P. Zhu, Z. Gu, and Z. Tian, “Rethinking perturbation directions for imperceptible adversarial attacks on point clouds,” IoT-J, vol. 10, no. 6, pp. 5158–5169, 2022.
PSR
58.40 42.60 39.00 14.20
SMR
5.456 5.317 5.286 5.101
CMR
5 Obstacles
1.262 0.933 0.905 0.877
JDS
PSR
SMR
CMR
JDS
2.071 52.20 5.690 0.923 2.196 2.140 0 N/A N/A N/A 2.251 0 N/A N/A N/A 2.306 0 N/A N/A N/A
[10] K. Tang, T. Hao, X. Wang, W. Peng, D. Zhang, P. Zhu, and Z. Tian, “Less is more: Sparse and cooperative perturbation for point cloud attacks,” in AAAI, vol. 40, pp. 9430–9438, 2026. [11] N. W. Alharthi and M. Brandão, “Physical and digital adversarial attacks on grasp quality networks,” in ICRA, pp. 1907–1902, 2024. [12] D. Wang, D. Tseng, P. Li, Y. Jiang, M. Guo, M. Danielczuk, J. Mahler, J. Ichnowski, and K. Goldberg, “Adversarial grasp objects,” in CASE, pp. 241–248, 2019. [13] X. Wang, M. Han, T. Hao, C. Li, Y. Zhao, and K. Tang, “Advgrasp: adversarial attacks on robotic grasping from a physical perspective,” in IJCAI, pp. 547–555, 2025. [14] Y. Wang, H. Zhang, H. Pan, Z. Zhou, X. Wang, P. Guo, L. Xue, S. Hu, M. Li, and L. Y. Zhang, “Advedm: Fine-grained adversarial attack against vlm-based embodied agents,” in NeurIPS, 2025. [15] T. Wang, C. Han, J. Liang, W. Yang, D. Liu, L. X. Zhang, Q. Wang, J. Luo, and R. Tang, “Exploring the adversarial vulnerabilities of vision-language-action models in robotics,” in ICCV, pp. 6948–6958, 2025. [16] X. Wang, M. Han, T. Hao, Y. Yang, Y.-B. Zhao, and K. Tang, “Partially observable adversarial patch attacks on vision-language-action models in robotics,” RA-L, vol. 11, no. 8, pp. 9351–9358, 2026. [17] S. Khuller, A. Moss, and J. S. Naor, “The budgeted maximum coverage problem,” Information processing letters, vol. 70, no. 1, pp. 39–45, 1999. [18] V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al., “Isaac gym: High performance gpu based physics simulation for robot learning,” in NeurIPS, 2021. [19] S. Macenski, T. Foote, B. Gerkey, C. Lalancette, and W. Woodall, “Robot operating system 2: Design, architecture, and uses in the wild,” Science robotics, vol. 7, no. 66, p. eabm6074, 2022. [20] D. Coleman, I. A. Sucan, S. Chitta, and N. Correll, “Reducing the barrier to entry of complex robotic software: a moveit! case study,” Journal of Software Engineering for Robotics, vol. 5, no. 1, pp. 3–16, 2014. [21] I. A. Sucan, M. Moll, and L. E. Kavraki, “The open motion planning library,” IEEE Robotics & Automation Magazine, vol. 19, no. 4, pp. 72–82, 2012. [22] C. Chamzas, C. Quintero-Pena, Z. Kingston, A. Orthey, D. Rakita, M. Gleicher, M. Toussaint, and L. E. Kavraki, “Motionbenchmaker: A tool to generate and benchmark motion planning datasets,” IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 882–889, 2021. [23] I. A. Sucan and L. E. Kavraki, “A sampling-based tree planner for systems with complex dynamics,” IEEE Transactions on Robotics, vol. 28, no. 1, pp. 116–131, 2012. [24] M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language-action model,” arXiv preprint arXiv:2406.09246, 2024. [25] K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky, “π0.5 : a vision-language-action model with open-world generalization,” in CoRL, 2025.