Conceptio › Archive › arXiv CS
arXiv CSopen access

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? Zhenxuan Fan1 , Bo Zhang2 , Yutong Lin1 , Yuqian Yuan1 , Juekai Lin1 , Liang Liang1 Zhuoyi Huang3 , Wenqiao Zhang1 * , Juncheng Li1 * , Siliang Tang1 , Jun Xiao1 , Yueting Zhuang1 1

2

Zhejiang University University of Electronic Science and Technology of China [email protected]

Code

Abstract

Data

2024; Kawaharazuka et al., 2025). With large-scale pretraining and multimodal alignment, these models show encouraging performance on manipulation tasks, especially in structured, short-horizon settings (Shao et al., 2025; Zhong et al., 2025). However, deploying embodied agents in realworld scenarios requires capabilities beyond shorthorizon instruction following in simple, wellstructured scenes (Liu et al., 2025b; Zhang et al., 2025b; Wong et al., 2025; Yuan et al., 2025a; Dang et al., 2026). Everyday manipulation tasks pose three key challenges: (1) Fine-grained target disambiguation, requiring agents to identify the target among visually similar candidates through subtle spatial cues beyond category or color; (2) Temporally extended task execution, requiring multistep execution under temporal constraints and accumulated errors; and (3) Complexity-scalable embodied reasoning, where increasing candidates, horizons, and environmental diversity amplify grounding and planning failures. These challenges call for benchmarks that evaluate embodied reasoning under increasing spatial and procedural complexity. As shown in Table 1, existing robotic manipulation datasets and benchmarks leave several gaps for evaluating reasoning-oriented VLA models. First, fine-grained spatial reasoning is rarely evaluated explicitly, as prior works seldom test target disambiguation with subtle spatial cues. Second, long-horizon and step-level evaluation remain limited: LIBERO (Liu et al., 2023a) and RoboTwin 2.0 (Chen et al., 2026a) mainly focus on short-horizon tasks, while RoboCasa (Nasiriany et al., 2024) and MIKASA-Robo (Cherepanov et al., 2026) often lack step-level diagnosis. Third, most prior works lack controlled multi-level difficulty (Pumacay et al., 2024; Fei et al., 2026; Zhou et al., 2025), making it hard to analyze performance degradation as task complexity increases. Finally, dataset scale remains limited for broad

Vision-Language-Action (VLA) models have shown promising progress in languageconditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https: //github.com/fanzhenxuan/RoboSPA.

arXiv:2609.05324v1 [cs.RO] 4 Sep 2026

South China Normal University

[email protected]

Project Page

1

3

Introduction

Vision-Language-Action (VLA) models (Zitkovich et al., 2023; Kim et al., 2025; Black et al., 2025b; NVIDIA et al., 2025; Black et al., 2025a; Gao et al., 2026) have emerged as a promising paradigm for general-purpose embodied agents, integrating visual perception, language understanding, and action generation in a unified framework (Ma et al., * Corresponding author.

1

Name

Categories Tasks Trajectories Embodiments

Spatial Long-Horizon Step-level MultiScene Reasoning Tasks Evaluation Difficulty Diversity

BC-Z (Jang et al., 2022) RT-1 (Brohan et al., 2023) BridgeData V2 (Walke et al., 2023) Open X-Embodiment (O’Neill et al., 2024) DROID (Khazatsky et al., 2024)

3 8 – 527 –

100+ 700+ 13 160,266 86

26K 130K 60.1K 1.4M 76K

1 1 1 22 1

✗ ✗ ✗ ✗ ✗

✗ ✓ ✗ ✗ ✗

✗ ✗ ✗ ✗ ✗

✗ ✗ ✗ ✗ ✗

✗ ✓ ✓ ✓ ✓

RLBench (James et al., 2020) CALVIN (Mees et al., 2022) LIBERO (Liu et al., 2023a) ManiSkill2 (Gu et al., 2023) RoboCasa (Nasiriany et al., 2024) SimplerEnv (Li et al., 2025) MIKASA-Robo (Cherepanov et al., 2026) RoboTwin 2.0 (Chen et al., 2026a) LIBERO-Pro (Zhou et al., 2025) RMBench (Chen et al., 2026b)

– – 4 4 – – 12 – 4 2

100 34 130 20 100 8 32 50 40 9

– N/A† 6.5K N/A‡ ∼100K – 32K ≤137.5K – 450

1 1 1 1 1 2 1 5 1 1

✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗

✓ ✓ ✗ ✗ ✓ ✗ ✓ ✗ ✗ ✓

✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗

✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✗

✗ ✓ ✓ ✓ ✓ ✓ ✗ ✓ ✓ ✗

RoboSPA (Ours)

10

280

527K

5

✓

✓

✓

✓

✓

Table 1: Comparison of representative robotic manipulation datasets and benchmarks with RoboSPA. The upper and lower parts summarize datasets and benchmarks, respectively. † CALVIN reports approximately 24 hours of teleoperated demonstrations. ‡ ManiSkill2 reports over 4M demonstration frames.

reasoning evaluation. Most widely used datasets contain fewer than 200K trajectories, limiting task, embodiment, and scene coverage. To address these limitations, we introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark. To systematically study embodied reasoning under increasing task complexity, RoboSPA is guided by three core design principles:

and action horizons. Beyond task-level success, RoboSPA reports step-level progress for detailed diagnosis of model failures. Together, these designs provide a controlled foundation for large-scale VLA data collection and evaluation. RoboSPA provides demonstrations in clean and domain-randomized scenes, covering 56 base tasks across 10 capability categories and five difficulty levels, yielding 280 variants. Each task supports step-level evaluation beyond final success rates. Across five embodiments, RoboSPA contains 527K trajectories and 997 hours of videos, forming a comprehensive benchmark for embodied reasoning under increasing task complexity. We evaluate representative VLA models (Liu et al., 2025a; AgiBot-World-Contributors et al., 2025; Black et al., 2025a; Zheng et al., 2026) on RoboSPA and find that they struggle as task complexity increases. On the hardest tasks, all models achieve an average success rate below 25%, with some tasks dropping to 0%. Diagnostic analyses reveal failures in target grounding, low-level manipulation, long-horizon tracking, and memory-based reasoning, highlighting RoboSPA as a challenging diagnostic testbed for embodied reasoning.

• Fine-Grained Spatial Reasoning. As shown in Fig. 1, this component evaluates whether VLA models can ground instructions in complex spatial structures, such as geometric attributes, distances, cross-view cues, relations, and canonical indexing. This capability is essential for cluttered manipulation. To our knowledge, RoboSPA is the first VLA dataset to make fine-grained spatial reasoning a core evaluation dimension. • Long-Horizon Procedural Planning. This component evaluates whether VLA models can execute manipulation tasks with multi-step decisions. As shown in Fig. 1, it covers repetitive procedures, order-constrained and order-free execution, composite coordination, and memoryintensive planning. These capabilities are crucial for real-world manipulation requiring temporal consistency and stepwise progress.

2

• Multi-Level Hierarchical Evaluation. This component measures how VLA model performance changes with increasing task complexity. Each base task is instantiated across five difficulty levels, enabling analysis of performance degradation under growing spatial complexity

Related Work

Vision-Language-Action Models. Large language models (LLMs) (Grattafiori et al., 2024; Yang et al., 2025; Cao et al., 2026; Fan et al., 2026) and multimodal large language models (MLLMs) (Liu et al., 2023b; Bai et al., 2025; Yuan et al., 2025b; Lin et al., 2025; Wang et al., 2026) 2

Fine-grained Spatial Reasoning ① Geometric Attribute Cognition ⑤

①

② ④

② Spatial Distance Estimation ③

③

③ Canonical Position Indexing

⑥

"Lift the 2nd largest block with one arm."

Row 1

“Lift the 4th object in the 1st row counted from near to far, with positions counted from right to left, with one arm.”

“Pick up the 3rd farthest pill bottle from the brown pill bottle.”

⑤ Cross-View Reasoning Right

④ Referential Relational Reasoning

Col. 4

“Pick the nearest one directly at the back of the middle-left soap bar.”

⑥ Repetitive Procedure Following

Left

“Viewed from the opposite side of the table, take the rightmost object.”

Step-Level Eval

5 Difficulty Levels

56 Base Tasks

280 Task Variants

527k Trajectories

997h Videos

Long-Horizon Procedural Planning ⑦ Order-Free Execution

⑧ Order-Constrained Execution

⑨ Composite Action Coordination ④

①

“Stack each blue bowl on the smooth ceramic plate.”

⑤

②

④

③

③

①

“Pick up the small handheld fan and put it down 5 times.”

⑩ Memory-Intensive Planning

⑤ ②

Left “(1) Place the bell over the electronic scale, (2)then place the toy car over the display stand, (3) and tap the card box, (4) then put the bell back on the table, (5) and put the toy car back on the table."

“Click the frontmost top bell, then circle clockwise through the other top bell.”

Right

"Memorize the front T-shaped blocks, then once they are hidden, make the rear T-shaped blocks match the front ones in orientation from left to right."

Figure 1: Overview of RoboSPA. RoboSPA centers on Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task has 5 difficulty levels, yielding 280 variants. The dataset includes 527K trajectories across multiple embodiments and scene settings, supporting step-level evaluation of VLA limitations in spatial reasoning, long-horizon execution, and scalable embodied reasoning.

Robotic Manipulation Datasets and Benchmarks. The rapid development of VLA models has also accelerated the development of robotic manipulation datasets and evaluation benchmarks. Large-scale datasets, including BridgeData V2 (Walke et al., 2023), Open XEmbodiment (O’Neill et al., 2024), DROID (Khazatsky et al., 2024), and AgiBot World (AgiBotWorld-Contributors et al., 2025), expand realworld, cross-embodiment manipulation coverage. Benchmark suites such as RLBench (James et al., 2020), CALVIN (Mees et al., 2022), ManiSkill2 (Gu et al., 2023), and LIBERO (Liu et al., 2023a) provide structured settings for evaluating diverse skills, language-conditioned control, and sequential manipulation. Recent benchmarks further examine scene and capability generalization, sim-to-real transfer, bimanual manipulation, and memory-dependent tasks (Pumacay et al., 2024; Nasiriany et al., 2024; Li et al., 2025; Sedlacek et al., 2026; Garcia et al., 2025; Zhang et al., 2025c; Zhou et al., 2025; Zhang et al., 2025a; Chen et al., 2026a; Cherepanov et al., 2026; Chen et al., 2026b). However, they rarely systematically evaluate fine-grained spatial reasoning and long-horizon procedural planning with step-level diagnostics. RoboSPA addresses this gap with tasks specifically designed around these two capabilities.

have demonstrated strong capabilities for understanding and reasoning over language and multimodal inputs, respectively. Building on these advances, VLA models map visual observations and language instructions to robot actions through a unified perception-to-control framework. Early methods, such as RT-1 (Brohan et al., 2023), RT2 (Zitkovich et al., 2023), and OpenVLA (Kim et al., 2025), cast action prediction as autoregressive discrete token generation. Subsequent VLA models largely adopt diffusion or flow-matching policies for continuous action generation, as in RDT (Liu et al., 2025a) and π0 (Black et al., 2025b). To improve cross-task generalization, VLA models such as π0.5 (Black et al., 2025a) and GR00T N1 (NVIDIA et al., 2025) adopt hierarchical policies to bridge high-level language understanding with low-level motor execution, while other works enhance action and embodiment generalization through implicit action modeling and cross-embodiment data (AgiBot-WorldContributors et al., 2025; Zheng et al., 2026). From a different perspective, another line of work improves VLA models through spatial guidance, memory mechanisms, and future-frame prediction, further enhancing their capability and usability (Chen et al., 2025; Shi et al., 2026; Torne et al., 2026; Zhang et al., 2025d; Bi et al., 2026). 3

More Objects

Repetitive Procedure Following

Longer Action Sequence ①

①→②

①→②→③

①→②→③→④

①→②→③→④→⑤

RoboSPA

(b) Multi-difficulty task design

ing nn la

Canonical Position Indexing

d Spatial R ea ine ra

Lon gH

n Procedura lP izo or

g nin so

(c) Domain-randomized scenes

Fin e-G

(a) Capability taxonomy

Figure 2: Design of RoboSPA. (a) Capability taxonomy with two dimensions and 10 categories. (b) Multi-difficulty task design through increasing object counts or action-sequence length. (c) Domain-randomized scenes with diverse layouts, distractors, lighting, textures, and tabletops.

3

RoboSPA

3.1

Overview

objects or to a reference entity, reasoning about proximity, remoteness, and relative distance. • Canonical Position Indexing (CPI): Evaluates whether the model can identify objects by rowcolumn indices under different counting directions and spatial scanning orders.

As shown in Fig. 1, we introduce RoboSPA, a largescale robotic manipulation dataset and benchmark for evaluating reasoning-oriented VLA models. RoboSPA is built using the SAPIEN (Xiang et al., 2020) simulator and RoboTwin 2.0 (Chen et al., 2026a) framework. Each task is instantiated across multiple difficulty levels and scene settings. 3.2

Benchmark Construction

3.2.1

Capability Taxonomy

• Referential Relational Reasoning (RRR): Assesses whether the model can locate targets through directional relations to reference objects, covering basic and compositional positions. • Cross-View Reasoning (CVR): Examines whether the model can interpret spatial instructions from non-egocentric viewpoints by transforming spatial references across perspectives.

To systematically evaluate embodied reasoning in VLA models, we build a hierarchical capability taxonomy instead of treating manipulation tasks as isolated instances. As shown in Fig. 2(a), it includes two core dimensions: Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning.

Long-Horizon Procedural Planning. This dimension evaluates a model’s ability to complete extended manipulation tasks with diverse actions, multiple subgoals, states, and procedural constraints. It also consists of five categories:

Fine-Grained Spatial Reasoning. This dimension assesses instruction grounding in complex spatial configurations. Tasks require selecting the correct target through fine-grained spatial reasoning in complex scenes. It includes five categories:

• Repetitive Procedure Following (RPF): Assesses whether the model can repeat a specified operation the required number of times while tracking progress and stopping correctly.

• Geometric Attribute Cognition (GAC): Evaluates whether the model can identify and manipulate objects by geometric or shape-related attributes beyond category-level recognition.

• Order-Free Execution (OFE): Measures whether the model can complete multiple subgoals in flexible order, covering all targets without omission or unnecessary repetition.

• Spatial Distance Estimation (SDE): Measures whether the model can compare distances among 4

(a) Embodiment Distribution

6008 600

130k

6 400 400

400

Randomized ARX-X5 Randomized 463K 133k 463K

200

0 L1

Piper 65k UR5 73k L4 L3

L2 Difficulty Level

4 200 Randomized 200 2 463K

0 00 L5

(b) Scene Distribution

Object Countvs. vs. Difficulty Object Count Timesteps vs.Difficulty Difficulty

L2 L3 L4 L5 L1L1 L1 L2L2 L3L3 L4L4 L5L5 DifficultyLevel Level Difficulty Difficulty Level

(c) Average Object Count

Object Count v

1010 800

10

Average Number

Clean 64K

88 600 66 400 44

2 2200

00 0

8 6 4 2

L1L1L1 L2L2L2 L3L3L3 L4L4L4 L5L5L5 Difficulty Level Difficulty Level Difficulty Level

0

(d) Average Timesteps

Figure 3: Dataset statistics of RoboSPA. (a) Trajectory distribution across five embodiments. (b) Trajectories across clean and domain-randomized scenes. (c) Average object count by difficulty for spatial reasoning tasks. (d) Average Object Count vs. Difficulty 10 trajectory length by difficulty for long-horizon planning tasks. Average Number

UR5 73k

ARX-X5 ARX-X5 Randomized 133k 133k 463K UR5 Piper 73k 65k

113k

800 800 10

Average Number Number AverageNumber Average

Piper 65k

5

Franka Piper 130k 65k

600

Timesteps vs. Difficulty Timesteps vs. Difficulty Object Count vs. Difficulty

TimestepsClean vs. Difficulty

Clean 64K Aloha-AgileX64K Franka

800

64K

Average Number Average Number Average Number

Franka 130k

Clean Aloha-AgileX UR5 Franka Aloha-AgileX 64K 113k 130k 114k 85k

Average Number

AgileX 3k

64K 64K

64K

8 6

• Order-Constrained Execution (OCE): Evalu- descriptions and manually checked 4 to ensure diverates whether the model can execute subgoals in sity under controlled semantics. 2 0 a specified order, as correct actions in the wrong L1 L2 L3 L4 3.2.3 Data Collection Robotic order can still fail. After task design, we execute expert code in simula• Composite Action Coordination (CAC): Ex- tion to collect trajectories in both clean and domainamines whether the model can coordinate hetero- randomized scenes. Clean scenes contain only taskgeneous manipulation skills across multi-stage relevant objects, while domain-randomized scenes procedures with different action types. vary clutter, textures, lighting, and tabletop configurations, as shown in Fig. 2(c), enabling evalua• Memory-Intensive Planning (MIP): Tests tion of both task competence and scene robustness. whether the model can retain information and We collect data across five embodiments: Alohause it to guide later actions when cues must be AgileX, ARX-X5, Piper, Franka, and UR5. recalled or are no longer available. 3.2.4 Evaluation Metrics 3.2.2 Task Suite Construction To enable more diagnostic evaluation beyond Task Construction. Based on the proposed taxSuccess Rate (SR), we further introduce two onomy, we design each task to instantiate a spedimension-specific complementary metrics. cific reasoning requirement. Tasks are defined by For fine-grained spatial reasoning tasks, we expert code specifying the scene, execution prointroduce Object-Normalized Target Accuracy cedure, and success condition. We reuse part of (ON T A), inspired by Cohen’s kappa (Cohen, RoboTwin 2.0 (Chen et al., 2026a) action prim1960), to fairly compare spatial reasoning ability itives. Ten trained graduate students design the under varying scene complexity. ON T A removes tasks, with each student responsible for one cathe object-count-induced chance baseline: pability category. Each task is reviewed by two   SR(n) − 1/n additional students for quality assurance. ON T A(n) = × 100. (1) 1 − 1/n Hierarchical Difficulty Design. To measure model performance under increasing reasoning bur- where n is the number of objects. ON T A is 100 den, we instantiate each task across five difficulty for perfect target selection, 0 for random guessing, levels. As shown in Fig. 2(b), difficulty is task- and negative for below-chance performance. specific: spatial tasks increase candidate objects, For long-horizon procedural planning tasks, we while procedural tasks extend action sequences report Progress Score (PS) to capture partial comwith more intermediate steps. pletion. It is computed on a 0–100 scale: Instruction Design. For each task, we use GPT5.2 (OpenAI, 2025) to generate 60 non-overlapping instruction templates with diverse linguistic forms, using 50 for training and 10 for testing. These templates are instantiated with task-specific object

PS =

N 1 X ci N i=1 T

!

× 100.

(2)

where T is the total number of subtasks, and ci is the number completed in episode i. 5

L1

L2 L3 Difficulty

Table 2: Performance comparison of four baseline VLA models on RoboSPA. We use Success Rate (SR) as the evaluation metric. Drop denotes the L1-to-L5 decrease. Blue shading marks the best model for each metric. RDT

Task L1

π0.5

GO-1

L5

Drop

L1

L5

Drop

X-VLA

L1

L5

Drop

L1

L5

Drop

42.8 35.0 44.0 33.2 41.2 39.2

12.5 23.4 23.6 9.8 37.6 21.4

30.3 11.6 20.4 23.4 3.6 17.8

57.8 41.4 31.8 40.4 38.2 41.9

25.3 34.6 8.8 14.2 36.6 23.9

32.5 6.8 23.0 28.2 1.6 18.0

8.0 35.5 38.8 21.6 7.6 22.3

80.3 85.8 77.4 58.6 54.0 71.2

65.0 26.1 12.0 12.5 0.0 23.1

15.3 59.7 65.4 46.1 54.0 48.1

34.8 86.0 70.0 58.9 44.6 58.8

22.5 43.6 7.7 5.9 0.0 15.9

12.3 42.4 62.3 53.0 44.6 42.9

16.3

55.2

22.3

32.9

50.4

19.9

30.5

Fine-Grained Spatial Reasoning Geometric Attribute Cognition Spatial Distance Estimation Canonical Position Indexing Referential Relational Reasoning Cross-View Reasoning Average

10.0 15.2 15.0 4.4 3.0 9.5

4.8 9.4 8.6 4.4 3.4 6.1

5.2 5.8 6.4 0.0 -0.4 3.4

27.8 20.2 22.0 17.2 17.4 20.9

11.5 11.6 7.2 6.6 16.2 10.6

16.3 8.6 14.8 10.6 1.2 10.3

Long-Horizon Procedural Planning Repetitive Procedure Following Order-Free Execution Order-Constrained Execution Composite Action Coordination Memory-Intensive Planning Average

40.8 20.3 13.6 27.5 18.2 24.1

35.0 0.8 0.1 2.4 0.0 7.7

5.8 19.5 13.5 25.1 18.2 16.4

31.8 41.1 41.9 24.2 7.6 29.3

23.8 5.6 3.1 2.6 0.0 7.0

Overall Benchmark 16.8

Overall Average Geometric Attribute Cognition

100

6.9

Spatial Distance Estimation

100

9.9

25.1

8.8

Canonical Position Indexing

100

100

Referential Relational Reasoning

80

80

80

80

80

60

60

60

60

60

40

40

40

40

40

20

20

20

20

20

0

0

0

0

L1

L2

L3

L4

L5

Repetitive Procedure Following

100

L1

L2

L3

L4

L5

Order-Free Execution

100

L1

L2

L3

L4

L5

Order-Constrained Execution

100

0 L1

L2

L3

L4

L5

Composite Action Coordination

100

L1

80

80

80

80

60

60

60

60

60

40

40

40

40

40

20

20

20

20

20

0

0

0

0

L2

L3

L4

L5

L1

L2

L3

L4

L5

L1

RDT

L2

L3

GO-1

L4

L5

π0.5

L2

L3

L4

L5

Memory Intensive Planning

100

80

L1

Cross-View Reasoning

100

0 L1

L2

L3

L4

L5

L1

L2

L3

L4

L5

X-VLA

100 baseline VLA models across 10 task categories over five difficulty levels. Figure 4: Success rate trends of four

Order-Constrained Execution

80 60

3.3

Benchmark Statistics

40

4

20 0 L1

L2

Experiments

L5 4.1L4 Experimental Setup

L3

RoboSPA contains 56 base tasks across two capability dimensions. Each base task is instantiated at 5 difficulty levels, yielding 280 task variants. In total, we collect approximately 527K trajectories, corresponding to over 997 hours of interaction videos and 108M timesteps. More dataset statistics are provided in the Appendix C. Fig. 3(a) presents the trajectory distribution across the 5 robotic embodiments, while Fig. 3(b) shows the trajectory distribution across clean and domain-randomized scenes. Fig. 3(c) and Fig. 3(d) further illustrate the hierarchical difficulty design along the two dimensions, respectively.

We evaluate four representative VLA models, RDT (Liu et al., 2025a), GO-1 (AgiBot-WorldContributors et al., 2025), π0.5 (Black et al., 2025a), and X-VLA (Zheng et al., 2026), using the cleanscene data collected with the Aloha-AgileX embodiment for both training and evaluation. All models are trained and evaluated under a single-task setting. For each base task, we train a separate model using data from all five difficulty levels and evaluate it independently at each level. During evaluation, each model is tested with 100 rollout trials per task variant. 6

Table 3: Comparison of Object-Normalized Target Accuracy (ON T A) on fine-grained spatial reasoning tasks. RDT

Task

π0.5

GO-1

X-VLA

L1

L3

L5

L1

L3

L5

L1

L3

L5

L1

L3

L5

Geometric Attribute Cognition Spatial Distance Estimation Canonical Position Indexing Referential Relational Reasoning Cross-View Reasoning

-80.0 -27.2 -27.5 -91.2 -45.5

-25.0 -9.1 -4.9 -25.1 -28.8

-14.3 -4.6 0.3 -7.6 -28.8

-44.5 -19.7 -17.0 -65.6 -23.9

-7.0 -6.7 -3.3 -18.9 -14.7

-6.2 -2.1 -1.2 -5.1 -11.7

-14.5 2.5 16.0 -33.6 11.8

1.3 13.6 10.2 0.5 16.0

-5.0 11.6 16.7 -1.5 16.8

15.5 12.1 -2.3 -19.2 7.3

11.0 18.2 2.2 0.3 17.3

10.3 24.4 0.5 3.5 15.5

Average

-54.3

-18.6

-11.0

-34.1

-10.1

-5.3

-3.6

8.3

7.7

2.7

9.8

10.8

Table 4: Performance comparison in terms of Progress Score (PS) on long-horizon procedural planning tasks. RDT

Task

π0.5

GO-1

X-VLA

L1

L3

L5

L1

L3

L5

L1

L3

L5

L1

L3

L5

Repetitive Procedure Following Order-Free Execution Order-Constrained Execution Composite Action Coordination Memory-Intensive Planning

40.8 20.3 13.6 27.5 18.2

46.3 33.0 25.3 32.5 5.0

44.5 24.9 14.7 26.5 3.6

31.8 41.1 41.9 24.2 7.6

35.0 39.7 24.1 26.2 4.9

32.0 25.7 14.8 22.7 6.0

80.3 85.8 77.4 58.6 54.0

78.7 75.0 44.3 60.5 18.7

77.3 62.1 28.0 48.0 11.4

34.8 86.0 70.0 58.9 44.6

34.8 84.4 51.1 49.0 12.2

31.9 73.9 25.1 40.1 7.3

Average

24.1

28.4

22.8

29.3

26.0

20.2

71.2

55.4

45.4

58.8

46.3

35.7

4.2

performing model, π0.5 , decreasing from 71.2% at L1 to 23.1% at L5. Models perform best on Repetitive Procedure Following (RPF), while the other four categories show substantially lower performance. Notably, all models fail on the hardest Memory-Intensive Planning (MIP), revealing severe limitations in maintaining task-relevant memory over extended horizons. Fig. 4 further shows that performance drops from L1 to L5, exposing failures in spatial discrimination and long-horizon tracking.

Main Results

Table 2 reports the Success Rates (SR) of four representative VLA models. Overall, model performance declines from L1 to L5, indicating limited ability to reason under complex task settings. Model Comparison. Among the evaluated models, π0.5 performs best overall, but reaches only 22.3% success at L5, followed by X-VLA with 19.9%. Both drop by over 30 percentage points as difficulty increases. GO-1 and RDT perform worse, achieving only 8.8% and 6.9% success at L5, respectively. These results show that stronger models handle easier cases better, but all models still lack robust reasoning under complex conditions.

4.3

Diagnostic Evaluation

Object-Normalized Target Accuracy. Table 3 compares models using the ON T A metric. Overall, all models achieve low scores. RDT and GO-1 obtain negative ON T A on most categories and difficulty levels, showing little spatial reasoning beyond random selection. π0.5 and X-VLA show fewer negative scores and achieve better average ON T A, but their overall scores remain close to the random-selection baseline. These results suggest that current VLA models still struggle to identify the correct target when success depends on fine-grained spatial cues.

Fine-Grained Spatial Reasoning. X-VLA shows the best performance in this dimension, but its success rate still drops from 41.9% at L1 to only 23.9% at L5. This decline shows that increasing spatial complexity substantially weakens target grounding. Models struggle particularly with Canonical Position Indexing (CPI) and Referential Relational Reasoning (RRR), indicating that current VLA models remain unreliable when target selection depends on fine-grained spatial relations rather than simple object recognition.

Progress Score. Table 4 reports Progress Score for long-horizon procedural planning tasks, providing a progress-based view beyond binary Success Rate. The results show that Progress Score is generally higher than final Success Rate, indicating that

Long-Horizon Procedural Planning. For LongHorizon Procedural Planning, the performance drop is even more pronounced, with the best7

Figure 5: Category-wise performance of four baseline VLA models on RoboSPA. Left: Success Rate across 10 capability categories at L1. Right: Success Rate across 10 capability categories at L5. ① Target Execution Error

② Target Grounding Error

③ Manipulation Error

Viewed from the left side of the table, pick up the backmost object.

Pick the 6th highest block.

Click the bell, open the microwave, click the bell, place the playing cards left of the blue soap, click the bell.

④ Memory Error

⑤ Temporal Ordering Error

Based on the colors of the blocks on the display platform, when they are hidden, move matching blocks closer to you.

⑥ Redundant Repetition Error ②

①

⑤ ④

③

①

②

Click the alarm clock, the can, the stapler, the playing cards, the bread sequentially

①

②

③

First stamp the left seal twice on the left pad, then stamp the right seal twice on the right pad, then press the stapler with the red top once.

Figure 6: Representative failure modes observed on RoboSPA. Failure modes vary across different RoboSPA task categories, showing that our benchmark can distinguish model weaknesses along different capability dimensions.

models can often execute part of a procedure but fail due to accumulated errors, incorrect ordering, or incomplete later-stage execution. π0.5 achieves the strongest average completion rate, followed by X-VLA. Nevertheless, Memory-Intensive Planning (MIP) remains weak at higher difficulty levels, indicating that current models still have limited memory ability during long-horizon execution. 4.4

details provided in the Appendix A. Observed Failure Modes. For fine-grained spatial reasoning, target execution error denotes failed manipulation after correct grounding, while target grounding error denotes selecting a distractor. For long-horizon procedural planning, manipulation error refers to failed intermediate execution, memory error to forgetting earlier task-relevant information, temporal ordering error to executing subtasks in the wrong order, and redundant repetition error to repeating completed actions instead of progressing.

Qualitative Results

Fig. 5 visualizes category-wise model performance across task difficulty. The radar plots show that model performance is highly uneven across categories and shrinks markedly under harder settings. This reveals persistent limitations in handling increasingly complex embodied reasoning tasks. 4.5

Implications for Future Models. These findings suggest several directions for future VLA models. Improving spatial performance requires stronger object-centric grounding and relation awareness, together with reliable low-level execution. Longhorizon tasks require stronger progress tracking, memory, and execution monitoring to maintain task state. These abilities help models follow temporal

Failure Analysis

We further analyze model failures. Fig. 6 presents six representative failure modes, with quantitative 8

constraints and reduce ordering errors, redundant repetitions, and cascading failures.

5

collected demonstrations to ensure that the dataset does not contain personally identifying information or offensive, hateful, or explicit content. Therefore, no additional anonymization is required.

Conclusion

We present RoboSPA, a reasoning-focused, largescale robotic manipulation dataset and benchmark. Centered on Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, RoboSPA covers ten task categories, five difficulty levels, multiple embodiments, and diverse scenes, with finegrained metrics for spatial grounding and subtask completion. Experiments show that existing VLA models still struggle with complex spatial relations, precise execution, and memory-intensive planning, especially at higher difficulty levels. We hope RoboSPA can advance the development of more capable and reliable embodied agents.

Artifact Licensing. RoboSPA builds upon thirdparty open-source resources. We use the SAPIEN simulator, which is released under the Apache License 2.0, and reuse code and 3D assets from RoboTwin 2.0 and RMBench, which are released under the MIT License. Baseline models and pretrained checkpoints used in our experiments are obtained from their official releases and used in accordance with their original license terms. These third-party artifacts are used consistently with their intended research purposes for simulationbased robotic manipulation, benchmark construction, data generation, and evaluation. We retain the original copyright and license notices of all third-party resources. The artifacts introduced by RoboSPA, including task definitions, benchmark code, evaluation scripts, and collected demonstration data, will be released under the MIT License.

Limitations Although our dataset covers diverse fine-grained spatial reasoning and long-horizon procedural planning tasks, several limitations remain. First, all tasks are constructed in simulation, and the sim-toreal gap may limit direct transfer to physical robots. Second, our benchmark focuses on tabletop manipulation, which may not fully capture the diversity and open-endedness of real-world environments. Third, while we include multiple robotic embodiments and domain-randomized scenes, broader settings such as deformable object manipulation, human-robot interaction, and open-ended task instructions remain underexplored.

Responsible Use. We acknowledge that simulator design, task selection, object categories, and embodiment choices may introduce biases and limit transferability, which we mitigate through multiple embodiments, diverse task categories, controlled difficulty levels, and both clean and domainrandomized scenes. We also encourage responsible use of RoboSPA, including efficient training, transparent reporting of computational costs, and consideration of environmental impact.

Ethics Statement

Acknowledgments

Scope and Safety. RoboSPA is designed to diagnose and evaluate VLA models in simulated robotic manipulation environments, rather than to serve as a directly deployable real-world robotic system. Data collection is conducted entirely in simulation and does not involve real-world robot deployment. Since models evaluated or developed with RoboSPA may eventually be transferred to physical environments, such deployment should include appropriate safety constraints, human oversight, and task-specific validation.

This work was supported by the National Key R&D Program of China (2025ZD0123100), the NSFC (62272411), the Zhejiang NSF (LRG25F020001), and the Key R&D Program of Zhejiang Province (2025C01030).

References AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, Shu Jiang, Yuxin Jiang, Cheng Jing, Hongyang Li, Jialu Li, Chiming Liu, Yi Liu, Yuxiang Lu, and 33 others. 2025. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. Preprint, arXiv:2503.06669.

Data Content and Privacy. RoboSPA does not involve human-subject experiments, personal data, or real-world private data, and thus poses no privacy risks to human participants. We checked the task instructions, object categories, scene assets, and

Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei

9

Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025. Qwen3-vl technical report. Preprint, arXiv:2511.21631.

Kaixuan Wang, Yue Chen, Hongcheng Wang, Junjie Wang, Tianhang Yang, Renjing Xu, Ruihai Wu, Yao Mu, Yaodong Yang, Hao Dong, and Ping Luo. 2026b. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design. Preprint, arXiv:2603.01229.

Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, and Jun Zhu. 2026. Motus: A unified latent action world model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 35101–35113.

Xinyi Chen, Yilun Chen, Yanwei Fu, Ning Gao, Jiaya Jia, Weiyang Jin, Hao Li, Yao Mu, Jiangmiao Pang, Yu Qiao, Yang Tian, Bin Wang, Bolun Wang, Fangjing Wang, Hanqing Wang, Tai Wang, Ziqin Wang, Xueyuan Wei, Chao Wu, and 10 others. 2025. Internvla-m1: A spatially guided visionlanguage-action framework for generalist robot policy. Preprint, arXiv:2510.13778.

Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, brian ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, and 16 others. 2025a. π0.5 : a vision-language-action model with open-world generalization. In Proceedings of The 9th Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, pages 17–40. PMLR.

Egor Cherepanov, Nikita Kachaev, Alexey Kovalev, and Aleksandr Panov. 2026. Memory, benchmark & robots: A benchmark for solving complex tasks with reinforcement learning. In The Fourteenth International Conference on Learning Representations. Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1):37–46. Ronghao Dang, Jiayan Guo, Bohan Hou, Sicong Leng, Kehan Li, Xin Li, Jiangpin Liu, Yunxuan Mao, Zhikai Wang, Yuqian Yuan, Minghao Zhu, Xiao Lin, Yang Bai, Qian Jiang, Yaxi Zhao, Minghua Zeng, Junlong Gao, Yuming Jiang, Jun Cen, and 7 others. 2026. Rynnbrain: Open embodied foundation models. Preprint, arXiv:2602.14979.

Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, and 6 others. 2025b. π0 : A Vision-Language-Action Flow Model for General Robot Control. In Proceedings of Robotics: Science and Systems, LosAngeles, CA, USA.

Zhenxuan Fan, Jie Cao, Yang Dai, Zheqi Lv, Wenqiao Zhang, Zhongle Xie, Peng LU, and Beng Chin Ooi. 2026. Ctrlcot: Dual-granularity chain-of-thought compression for controllable reasoning. Preprint, arXiv:2601.20467.

Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, and 32 others. 2023. RT-1: Robotics Transformer for RealWorld Control at Scale. In Proceedings of Robotics: Science and Systems, Daegu, Republic of Korea.

Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. 2026. Libero-plus: A progressive robustness benchmark for visual-language-action models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 38574–38583.

Jie Cao, Tianwei Lin, Zhenxuan Fan, Bo Yuan, Ziyuan Zhao, Rolan Yan, Wenqiao Zhang, and Siliang Tang. 2026. Draft-thinking: Learning efficient reasoning in long chain-of-thought llms. Preprint, arXiv:2603.00578.

Mingjian Gao, Wenqiao Zhang, Yuqian Yuan, Yang Dai, Binhe Yu, Zheqi Lv, Haoyu Zheng, Jiaqi Zhu, Zhiqi Ge, Zixuan Wan, Siliang Tang, and Yueting Zhuang. 2026. Visualthink-vla: Visual intermediate reasoning for effective and low-latency vision-language-action policies. Preprint, arXiv:2605.30011.

Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, Weiliang Deng, Yubin Guo, Tian Nian, Xuanbing Xie, Qiangyu Chen, Kailun Su, Tianling Xu, Guodong Liu, Mengkang Hu, and 7 others. 2026a. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. In Fortythird International Conference on Machine Learning.

Ricardo Garcia, Shizhe Chen, and Cordelia Schmid. 2025. Towards generalizable vision-language robotic manipulation: A benchmark and llm-guided 3d policy. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 8996–9002. IEEE. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad AlDahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh

Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang,

10

Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783.

2025. Healthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation. In Fortysecond International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 of Proceedings of Machine Learning Research. PMLR / OpenReview.net.

Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen, and Hao Su. 2023. Maniskill2: A unified benchmark for generalizable manipulation skills. In The Eleventh International Conference on Learning Representations.

Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023a. Libero: Benchmarking knowledge transfer for lifelong robot learning. In Advances in Neural Information Processing Systems, Datasets and Benchmarks Track.

Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J. Davison. 2020. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 5(2):3019–3026.

Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023b. Visual instruction tuning. Advances in neural information processing systems, 36:34892– 34916.

Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. 2022. Bc-z: Zero-shot task generalization with robotic imitation learning. In conference on Robot Learning, pages 991–1002. PMLR.

Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. 2025a. RDT-1b: a diffusion foundation model for bimanual manipulation. In The Thirteenth International Conference on Learning Representations.

Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner, and Yuke Zhu. 2025. Vision-language-action models for robotics: A review towards real-world applications. IEEE Access, 13:162467–162504.

Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. 2025b. Aligning cyber space with physical world: A comprehensive survey on embodied ai. IEEE/ASME Transactions on Mechatronics, 30(6):7253–7274.

Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Jason Ma, Patrick Tree Miller, Jimmy Wu, Suneel Belkhale, Shivin Dass, and 79 others. 2024. DROID: A large-scale in-the-wild robot manipulation dataset. In Proceedings of Robotics: Science and Systems, Delft, Netherlands.

Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. 2024. A survey on vision-languageaction models for embodied ai. arXiv preprint arXiv:2405.14093. Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. 2022. Calvin: A benchmark for language-conditioned policy learning for longhorizon robot manipulation tasks. IEEE Robotics and Automation Letters, 7(3):7327–7334.

Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2025. Openvla: An open-source vision-language-action model. In Proceedings of The 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pages 2679–2713. PMLR.

Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. 2024. RoboCasa: Large-Scale Simulation of Household Tasks for Generalist Robots. In Proceedings of Robotics: Science and Systems, Delft, Netherlands. NVIDIA, :, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, and 24 others. 2025. Gr00t n1: An open foundation model for generalist humanoid robots. Preprint, arXiv:2503.14734.

Xuanlin Li, Kyle Hsu, Jiayuan Gu, Oier Mees, Karl Pertsch, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. 2025. Evaluating real-world robot manipulation policies in simulation. In Proceedings of The 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pages 3705–3728. PMLR.

OpenAI. 2025. Introducing gpt-5.2. https://openai. com/index/introducing-gpt-5-2/.

Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li, Wanggui He, Hao Jiang, Mengze Li, Xiaohui Song, Siliang Tang, Jun Xiao, Hui Lin, Yueting Zhuang, and Beng Chin Ooi.

Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar,

11

Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew Wang, Anikait Singh, and 260 others. 2024. Open x-embodiment: Robotic learning datasets and rt-x models : Open x-embodiment collaboration0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903.

vances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, pages 26787–26795. AAAI Press. Lik Hang Kenny Wong, Xueyang Kang, Kaixin Bai, and Jianwei Zhang. 2025. A survey of robotic navigation and manipulation with physics simulators in the era of embodied ai. arXiv preprint arXiv:2505.01458. Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, Li Yi, Angel X. Chang, Leonidas J. Guibas, and Hao Su. 2020. Sapien: A simulated part-based interactive environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Wilbert Pumacay, Ishika Singh, Jiafei Duan, Ranjay Krishna, Jesse Thomason, and Dieter Fox. 2024. The Colosseum: A Benchmark for Evaluating Generalization for Robotic Manipulation. In Proceedings of Robotics: Science and Systems, Delft, Netherlands. Martin Sedlacek, Pavlo Yefanov, Georgy Ponimatkin, Jai Bardhan, Simon Pilc, Mederic Fourmy, Evangelos Kazakos, Cees G. M. Snoek, Josef Sivic, and Vladimir Petrik. 2026. Realm: A real-to-sim validated benchmark for generalization in robotic manipulation. IEEE Robotics and Automation Letters, 11(7):8315–8322.

An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388.

Rui Shao, Wei Li, Lingsen Zhang, Renshan Zhang, Zhiyang Liu, Ran Chen, and Liqiang Nie. 2025. Large vlm-based vision-language-action models for robotic manipulation: A survey. arXiv preprint arXiv:2508.13073.

Yuqian Yuan, Ronghao Dang, long li, Wentong Li, Dian Jiao, Xin Li, Deli Zhao, Fan Wang, Wenqiao Zhang, Jun Xiao, and Yueting Zhuang. 2025a. Eoc-bench: Can mllms identify, recall, and forecast objects in an egocentric world? In Advances in Neural Information Processing Systems, volume 38, Main Conference. Curran Associates, Inc.

Hao Shi, Bin Xie, Yingfei Liu, Lin Sun, Fengrong Liu, Tiancai Wang, Erjin Zhou, Haoqiang Fan, Xiangyu Zhang, and Gao Huang. 2026. MemoryVLA: Perceptual-cognitive memory in vision-languageaction models for robotic manipulation. In The Fourteenth International Conference on Learning Representations.

Yuqian Yuan, Hang Zhang, Wentong Li, Zesen Cheng, Boqiang Zhang, Long Li, Xin Li, Deli Zhao, Wenqiao Zhang, Yueting Zhuang, Jianke Zhu, and Lidong Bing. 2025b. Videorefer suite: Advancing spatialtemporal object understanding with video llm. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18970–18980.

Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. 2026. Mem: Multi-scale embodied memory for vision language action models. Preprint, arXiv:2603.03596.

Borong Zhang, Jiahao Li, Jiachen Shen, Yishuai Cai, Yuhao Zhang, Yuanpei Chen, Juntao Dai, Jiaming Ji, and Yaodong Yang. 2025a. VLA-arena: An opensource framework for benchmarking vision-languageaction models. arXiv preprint arXiv:2512.22539.

Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine. 2023. Bridgedata v2: A dataset for robot learning at scale. In Proceedings of The 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 1723– 1736. PMLR.

Kun Zhang, Peng Yun, Jun Cen, Junhao Cai, Didi Zhu, Hangjie Yuan, Chao Zhao, Tao Feng, Michael Yu Wang, Qifeng Chen, Jia Pan, Wei Zhang, Bo Yang, and Hua Chen. 2025b. Generative artificial intelligence in robotic manipulation: A survey. Preprint, arXiv:2503.03464. Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, and Xipeng Qiu. 2025c. Vlabench: A large-scale benchmark for languageconditioned robotics manipulation with long-horizon reasoning tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11142–11152.

Zhuonan Wang, Zhenxuan Fan, Siwen Tan, Yu Zhong, Yuqian Yuan, Haoyuan Li, Hao Jiang, Wenqiao Zhang, Feifei Shao, Hongwei Wang, and Jun Xiao. 2026. MAU-GPT: enhancing multi-type industrial anomaly understanding via anomaly-aware and generalist experts adaptation. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Ad-

Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, XinQiang Yu, Jiazhao Zhang, Runpei Dong, Jiawei He, He Wang, Zhizheng Zhang, Li Yi, Wenjun Zeng,

12

model, we analyze approximately 1,120 evaluation videos, sampling up to 20 error examples per L5 task; if fewer than 20 errors occur, we include all available examples. For fine-grained spatial reasoning tasks, we define three error types to distinguish target selection failures from execution failures:

and Xin Jin. 2025d. Dreamvla: A vision-languageaction model dreamed with comprehensive world knowledge. In Advances in Neural Information Processing Systems. Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, Tai Wang, Ya-Qin Zhang, Jingjing Liu, and Xianyuan Zhan. 2026. XVLA: Soft-prompted transformer as scalable crossembodiment vision-language-action model. In The Fourteenth International Conference on Learning Representations.

• Target Execution Error: The model selects the correct target but fails to execute the grasp or pick action properly. • No-Contact Grasp Error: The model attempts to grasp but fails to contact any object.

Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, and Yaodong Yang. 2025. A survey on vision-language-action models: An action tokenization perspective. Preprint, arXiv:2507.01925.

• Target Grounding Error: The model selects the wrong object instead of the intended target. For long-horizon procedural planning tasks, we define five error types. One of them is the Target Grounding Error (as above), and the remaining four are:

Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Guowen Zhang, Duanfeng Chu, Pan Zhou, and Lichao Sun. 2025. Libero-pro: Towards robust and fair evaluation of vision-language-action models beyond memorization. Preprint, arXiv:2510.03827.

• Manipulation Error: Low-level execution of a subtask fails (e.g., pick, pull, or move).

Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, and 35 others. 2023. Rt-2: Visionlanguage-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, pages 2165–2183. PMLR.

• Memory Error: The model fails to retain relevant information from previous steps, affecting subtask completion. • Temporal Ordering Error: Subtasks are executed in an incorrect order. • Redundant Repetition Error: The model repeats an already completed action instead of progressing to the next subgoal.

Appendix A.2 In this document, we provide additional details about RoboSPA. The appendix is organized as follows:

Figure 7 and Figure 8 show the distribution of error types for the π0.5 and X-VLA models, respectively. Overall, the two models exhibit broadly similar error patterns across the two capability dimensions, although the exact proportions vary.

• §A: Detailed Failure Analysis. • §B: Training and Evaluation Details.

Fine-Grained Spatial Reasoning. In these tasks, both models are primarily challenged by Target Grounding Errors, which account for over 70% of all errors. The remaining errors are mostly Target Execution Errors, while No-Contact Grasp Errors occur very rarely. Among individual tasks, Canonical Position Indexing (CPI) exhibits the highest proportion of Target Grounding Errors, while CrossView Reasoning (CVR) has a relatively larger share of Target Execution Errors compared with the other tasks. This indicates that the main challenge lies in insufficient spatial reasoning, while low-level execution errors are comparatively rare.

• §C: Dataset Details. • §D: Basic Grounding Tasks. • §E: Task List and Descriptions. • §F: Detailed Experimental Results. • §G: Task Visualizations.

A

Detailed Failure Analysis

A.1

Analysis Setup

Failure Analysis Results

We perform a detailed error analysis on the two best-performing models, π0.5 and X-VLA. For each 13

Distribution of Error Types Fine-Grained Spatial Reasoning GAC SDE

27%

CPI 1% 72%

RRR CVR

Long-Horizon Procedural Planning

RPF

13% 10%

OFE

OCE

12%

CAC

15%

50%

MIP 0%

20%

40%

60%

80%

Target Execution Error

No-Contact Grasp Error

Target Grounding Error

Manipulation Error

Memory Error

Redundant Repetition Error

100%

Temporal Ordering Error

Figure 7: Distribution of error types for the π0.5 model.

Distribution of Error Types Fine-Grained Spatial Reasoning GAC 20%

SDE

2%

CPI RRR

78%

CVR Long-Horizon Procedural Planning

RPF

14%

OFE

OCE

9%

14%

CAC

12%

51%

MIP 0%

20%

40%

60%

80%

Target Execution Error

No-Contact Grasp Error

Target Grounding Error

Manipulation Error

Memory Error

Redundant Repetition Error

100%

Temporal Ordering Error

Figure 8: Distribution of error types for the X-VLA model.

Long-Horizon Procedural Planning. Within this dimension, errors are diverse. Overall, Manipulation Errors account for the largest share, approximately 50%, while the remaining error types occur at roughly similar rates.

cution (OFE) and Composite Action Coordination (CAC) are dominated by Manipulation Errors, highlighting difficulties in low-level execution across multiple subgoals. Order-Constrained Execution (OCE) shows that, aside from Manipulation Errors, most remaining failures are Temporal Ordering Errors, indicating that models struggle to correctly sequence subtasks. Memory-Intensive Planning (MIP) is primarily affected by Memory Errors, indicating limitations in retaining task-relevant infor-

Repetitive Procedure Following (RPF) primarily exhibits Redundant Repetition Errors, suggesting that models lack effective tracking of task progress and may decide the next action based solely on the current visual observation. Order-Free Exe14

pi05

Table 5: Asset categories and the number of variants used in RoboSPA. Asset

#Var. Asset

#Var. Asset

#Var. Asset

bottle bowl cover plate fluted-block french-fries hamburg T_block shoe-box tray kettle pen dustbin plant-pot dumbbell-rack bookcase laptop oven calculator microphone coaster hammer cup cup-with-liquid tissue-box scanner chips-tub pet-collar table-tennis roll-paper

27 7 1 1 2 4 6 2 1 4 9 7 1 5 4 4 11 4 6 4 1 1 13 1 7 5 4 4 2 4

5 7 5 1 1 1 2 1 1 4 13 1 10 1 2 2 3 6 3 7 7 2 5 7 5 2 3 8 6 6

7 7 6 11 11 6 5 3 3 4 7 4 6 5 3 5 7 5 5 7 7 5 3 4 4 5 6 5 8 5

olive-oil drill jam-jar screwdriver fork knife apple cabinet box milk-box mug rack shoe wooden_box book microwave sand-clock alarm-clock mouse stapler shampoo bell candlestick dumbbell teanet baguette small-speaker switch toycar markpen

pencup kitchenpot battery plasticbox tabletrashbin msg soy-sauce vinegar steamer boxdrink vagetable paymentsign can electronicscale rubikscube displaystand bread breadbasket phone phonestand remotecontrol pillbottle playingcards smallshovel brush woodenmallet gong woodenblock waterer wineglass

Table 6: Model size and training compute budget per base task. All experiments were conducted on NVIDIA RTX PRO 6000 GPUs. Model

Params.

#GPUs

Time

GPU Hours.

RDT GO-1 π0.5 X-VLA

1.2B 2.8B 3.3B 0.9B

1 2 2 1

2.5 h 6h 6h 1.5 h

2.5 12 12 1.5

B

globe trophy notebook brush-pen rest glue cleaner screen speaker fan seal milk-tea roller fruit board sauce-can skillet soap block hydrating-oil basket callbell tea-box coffee-box perfume keyboard whiteboard-eraser tooth-paste mini-chalkboard plant

#Var. 2 5 3 6 4 6 4 4 6 7 6 6 3 7 5 5 4 4 7 4 4 6 6 7 4 4 1 1 1 1

Training and Evaluation Details

Training. For each base task, we use 50 trajectories per difficulty level, resulting in 250 trajectories across five levels. For each task variant, we generate 60 templates, with 50 used for training and 10 held out for evaluation. These templates are instantiated with task-specific object descriptions to form 100 distinct training instructions and 100 non-overlapping evaluation instructions. During training, one instruction is randomly sampled from the corresponding task-variant instruction pool for each trajectory. To ensure a fair and controlled comparison, all baseline models are initialized from their officially released pretrained weights and fully fine-tuned on our dataset, with all parameters updated during training. We follow their official implementations, recommended environments, preprocessing pipelines, and evaluation settings, with only the training data and evaluation tasks replaced by RoboSPA. We adopt a unified training configuration across all models, using a batch size of 16 and training for 20,000 steps. All experiments are conducted

mation across extended action sequences.

Summary. These task-specific error patterns demonstrate that different tasks in RoboSPA correspond to distinct failure modes, providing a benchmark with good task differentiation for diagnosing VLA model capabilities. They further suggest that future VLA models should strengthen spatial grounding, ensure precise low-level execution, implement effective progress tracking, and incorporate memory mechanisms to improve long-horizon tasks. 15

Figure 9: The five robotic embodiments supported by RoboSPA. From left to right: Aloha-AgileX, ARX-X5, Piper, Franka, and UR5.

on NVIDIA RTX PRO 6000 GPUs under consistent hardware settings. Table 6 summarizes the model sizes and training compute budgets per base task, including the number of parameters, GPU configuration, training time, and GPU hours for each evaluated baseline. Across all 56 base tasks and four baseline models, the total training budget is approximately 1,568 GPU-hours.

camera viewpoint, and execution behavior. This design enables evaluation of VLA models under diverse robotic configurations and tests whether models can generalize beyond a single embodiment. C.3

To evaluate robustness beyond clean environments, we collect additional trajectories under domainrandomized scenes, as shown in Figure 10. The task semantics and success conditions remain unchanged, while the visual and environmental conditions are varied. Compared with clean scenes, the randomized setting introduces four types of variations: background appearance, tabletop clutter, table height, and lighting conditions. Background randomization exposes models to diverse visual contexts, while tabletop clutter introduces distractor objects. Table-height perturbation adds mild geometric variation, and lighting randomization changes illumination conditions.

Evaluation. Each task variant is evaluated over 100 test trials, and the Success Rate is computed accordingly. To assess generalization ability, all evaluation instructions are strictly disjoint from those used during training, ensuring no overlap between training and test instructions. In addition to binary success, each task script records the completion status of individual actions or subtasks during execution. This enables step-level fine-grained evaluation for long-horizon tasks, where a rollout may fail the final goal while still completing part of the required procedure.

C

Dataset Details

C.1

Object Asset Details

D

The object assets used in RoboSPA are collected from the RoboTwin-OD asset library (Chen et al., 2026a) and RMBench (Chen et al., 2026b). In total, RoboSPA includes 120 object categories and 581 object variants, covering diverse daily-use objects with different appearances, geometries, and semantic types. These assets are used to instantiate task scenes across both fine-grained spatial reasoning and long-horizon procedural planning tasks, as summarized in Table 5. C.2

Domain-Randomized Scenes

Basic Grounding Tasks

We additionally evaluate three basic grounding variants as lower-level references for spatial reasoning. These variants match the object categories and candidate counts of the corresponding RoboSPA tasks, namely SDE, CVR, and CPI, respectively. ColorBased Grounding relies only on color cues, Egocentric Position Grounding uses robot-view positions without cross-view transformation, and Simple Ordinal Indexing fixes the counting directions to leftto-right and far-to-near, unlike CPI where indexing directions vary. As shown in Table 7, these tasks achieve relatively high Success Rates across difficulty levels, indicating that basic grounding tasks may be insufficient to expose the reasoning limitations of current VLA models. In contrast, RoboSPA introduces fine-grained spatial reasoning settings, such as distance estimation,

Multi-Embodiment Setting

RoboSPA supports multiple robotic embodiments, including Aloha-AgileX, ARX-X5, Piper, Franka, and UR5, as shown in Figure 9. These embodiments share the same task semantics, language instructions, object configurations, and success conditions, while differing in morphology, workspace, 16

Figure 10: Examples of domain-randomized scenes in RoboSPA. Table 7: Performance on basic spatial task variants.

F

Detailed Experimental Results

Task

F.1

Statistical Uncertainty

L1

L2

L3

L4

L5

Color-Based Grounding 80.6 77.8 79.8 71.2 74.0 Egocentric Position Grounding 67.8 69.4 74.8 72.2 72.4 Simple Ordinal Indexing 87.0 86.8 83.8 87.0 84.8

cross-view reasoning, relational grounding, and canonical position indexing with more complex spatial layouts. These tasks provide a more challenging and diagnostic evaluation of spatial reasoning ability.

We use a distinct seed for each evaluation episode, which determines the scene and object layout. Given the same checkpoint and seed, the simulated rollout is largely deterministic. To quantify the remaining finite-sample uncertainty, we report bootstrap 95% confidence intervals for the L5 results in Table 10. The results show that the main performance trends remain unchanged after accounting for statistical uncertainty.

E

F.2

Task List and Descriptions

Tables 8 and 9 provide the full list of base tasks in RoboSPA and their objectives. Each base task is instantiated into 5 difficulty variants, where finegrained spatial reasoning tasks vary the number of candidate objects and long-horizon procedural planning tasks vary the number of procedural steps.

Multi-Embodiment Evaluation

To evaluate consistency across robotic platforms, we evaluate X-VLA on five embodiments. Due to computational resource constraints, we select one representative task from each of RoboSPA’s ten capability categories. For each task and embodiment, the model is independently trained and 17

F.5

evaluated using data from that embodiment. Tasks are denoted by their category abbreviation and index (e.g., GAC-1), with the index corresponding to their order within the category in Tables 8 and 9.

Tables 21–24 report the Progress Scores of all evaluated models on long-horizon procedural planning tasks. Unlike binary Success Rate, the Progress Score measures partial subtask completion during execution. These results provide a more finegrained view of long-horizon failures by showing how much of the required procedure each model completes before failing. The results highlight the limited long-horizon planning ability of current models. RDT and GO-1 obtain consistently low Progress Scores, with most tasks scoring below 50, suggesting that they often fail to complete even half of the required procedure. π0.5 and X-VLA achieve higher Progress Scores, indicating better ability to complete intermediate subtasks. However, their scores still decline clearly as difficulty increases, showing that current VLA models remain far from robust long-horizon procedural planning.

As shown in Tables 11 and 12, absolute performance varies across embodiments, but the L5 success rate is lower than the L1 success rate for every evaluated task and embodiment. This demonstrates that RoboSPA exhibits a consistent difficulty trend across robotic embodiments. F.3

Success Rates

To provide a complete view of model performance, we report the detailed per-task Success Rates for all evaluated models across the five difficulty levels. Tables 13–16 list the results of RDT, GO-1, π0.5 , and X-VLA on every task in RoboSPA. These tables complement the aggregated results in the main paper by showing how each model performs on individual task variants, making it easier to identify task-specific strengths and failure patterns.

G

Task Visualizations

We provide visualizations for the most difficult variant (L5) of each task in RoboSPA, as shown in Figures 11–20. Each row shows representative rollout frames together with the task category, task name, and instruction, illustrating the scene layout, target objects, and required execution process.

Overall, the detailed results show that models generally perform better in settings with fewer distractor objects and shorter action sequences. However, performance drops become much sharper as difficulty increases, especially for tasks involving fine-grained spatial disambiguation, strict ordering, compositional execution, or memory-intensive planning. F.4

Progress Score

Object-Normalized Target Accuracy

Tables 17–20 report the Object-Normalized Target Accuracy (ON T A) of all evaluated models on finegrained spatial reasoning tasks. These results help identify whether failures are caused by weak target disambiguation rather than low-level manipulation errors. The ON T A results reveal clear differences in spatial grounding ability across models. RDT and GO-1 achieve poor scores, with many values below zero, indicating below-chance target selection on numerous spatial tasks. π0.5 and X-VLA perform better, but their scores still remain low on many variants and often stay close to the randomguessing baseline. These results suggest that current VLA models still have substantial room for improvement in fine-grained spatial reasoning, especially under increasing spatial ambiguity and scene complexity. 18

Table 8: Task list and descriptions for fine-grained spatial reasoning. Category abbreviations: GAC = Geometric Attribute Cognition, SDE = Spatial Distance Estimation, CPI = Canonical Position Indexing, RRR = Referential Relational Reasoning, and CVR = Cross-View Reasoning. Cat.

Task Name

Task Description

GAC GAC GAC GAC SDE

Pick Blocks Size Pick Blocks Height Pick Blocks Length Pick Blocks Area Pick Blocks Distance

SDE

Pick Mugs Distance

SDE

Pick Pill Bottles Distance

SDE

Pick Mixed Objects Distance A

SDE

Pick Mixed Objects Distance B

CPI

Pick Blocks Canonical

CPI CPI

Pick Cups Canonical Pick Rubik Cubes Canonical

CPI CPI RRR

Pick Mixed Objects Canonical A Pick Mixed Objects Canonical B Pick Tea Box Relational

RRR

Pick Cans Relational

RRR

Pick Soaps Relational

RRR

Pick Mixed Objects Relational A

RRR

Pick Mixed Objects Relational B

CVR

Pick Sauce Can Multi View

CVR

Pick Seals Multi View

CVR

Pick Breads Multi View

CVR

Pick Mixed Objects Multi View A

CVR

Pick Mixed Objects Multi View B

Pick the target block from multiple blocks based on the size. Pick the target block from multiple blocks based on the height. Pick the target block from multiple blocks based on the length. Pick the target block from multiple blocks based on the base area. Pick the target block from multiple blocks using both color and relative distance cues. Pick the target mug from multiple mugs using both color and relative distance cues. Pick the target pill bottle from multiple pill bottles using both color and relative distance cues. Pick the target object from mixed-object set A using both color and relative distance cues. Pick the target object from mixed-object set B using both color and relative distance cues. Pick the target block according to its ordinal position among multiple blocks. Pick the target cup according to its ordinal position among multiple cups. Pick the target Rubik cube according to its ordinal position among multiple Rubik cubes. Pick the target object from mixed-object set A according to ordinal position. Pick the target object from mixed-object set B according to ordinal position. Pick the target tea box according to its relative position to another reference tea box. Pick the target can according to its relative position to another reference can. Pick the target soap according to its relative position to another reference soap. Pick the target object in a mixed-object scene according to another reference object. Pick the target object in a mixed-object scene according to another reference object. Pick up the target sauce can specified by its absolute position under a given viewpoint. Pick up the target seal specified by its absolute position under a given viewpoint. Pick up the target bread specified by its absolute position under a given viewpoint. Pick up the target object from mixed-object group A specified by its absolute position under a given viewpoint. Pick up the target object from mixed-object group B specified by its absolute position under a given viewpoint.

19

Table 9: Task list and descriptions for long-horizon procedural planning. Category abbreviations: RPF = Repetitive Procedure Following, OFE = Order-Free Execution, OCE = Order-Constrained Execution, CAC = Composite Action Coordination, and MIP = Memory-Intensive Planning. Cat.

Task Name

RPF RPF RPF RPF OFE

Press Stapler Repeat Lift Pot Repeat Lift Fan Repeat Click Bell Repeat Put Bottles Dustbin

OFE OFE OFE OFE OFE OFE OFE OCE OCE OCE OCE OCE OCE OCE CAC CAC CAC CAC CAC CAC CAC CAC MIP MIP MIP MIP

MIP

Task Description

Press the stapler repeatedly. Lift and put down the pot repeatedly. Lift the fan repeatedly. Click the bell repeatedly. Put multiple bottles into the dustbin without requiring a fixed execution order. Move Blocks Apart Move multiple blocks apart without requiring a fixed execution order. Move Playing Cards Away Move multiple playing cards away without requiring a fixed execution order. Place Bowls Plates Place multiple bowls on plates without requiring a fixed execution order. Separate Fries Bread Separate fries and bread into different areas without requiring a fixed execution order. Rank Blocks Height Arrange multiple blocks by height while allowing flexible intermediate execution order. Rank Blocks Color Arrange multiple blocks by color while allowing flexible intermediate execution order. Rank Blocks Size Arrange multiple blocks by size while allowing flexible intermediate execution order. Hit Blocks Hammer Order Hit multiple blocks with a hammer in the specified color order. Click Objects Order Click multiple target objects in the specified order. Stamp Seals Order Stamp multiple seals in the specified order. Click the bell and then execute the clockwise movement sequence in order. Click Bell Clockwise Order Stack Blocks Color Order Stack multiple blocks in the specified color order. Stack Blocks Size Order Stack multiple blocks in the specified size order. Stack Blocks Length Order Stack multiple blocks in the specified length order. Place Phone Press Stapler Execute a manipulation task involving phone placement, repeated pressing, and dumping. Place Bottle Cup Execute a manipulation task involving throwing bottles into a dustbin and placing cups on coasters. Hang Mug Stack Blocks Execute a manipulation task involving hanging a mug on a rack and stacking blocks from bottom to top. Place Object Scale Click Execute a manipulation task involving object placement, pressing, and returning objects to the table. Place Burger Fries Click Bell Execute a manipulation task involving placing a hamburger and fries on a tray, clicking a bell, and returning them. Click Can Place Items Execute a manipulation task involving interleaved can clicking with placing bread in a basket and a pill bottle on a pad. Stamp Seals Press Stapler Execute a manipulation task involving stamping left and right seals twice on their pads and pressing a stapler. Click Bell Open Microwave Place Ob- Execute a manipulation task involving interleaved bell clicking, opening a ject microwave, and placing an object left of another. Observe Blocks Move Memory First remember the colors of the display blocks, then move matching blocks on the table closer after the display blocks are hidden. Observe Objects Click Memory First remember the objects on the display platform from left to right, then click matching objects on the table in the same order after they are hidden. Remember Color Cover First remember the block colors, then uncover the blocks in the given color order after they are covered. Remember Orientation Restore First remember the orientations of the front T-shaped blocks, then adjust the rear T-shaped blocks to match them from left to right after they are hidden. Press Stapler Memory Remember how many times each stapler should be pressed and then press them.

20

Table 10: Bootstrap 95% confidence intervals for the average L5 Success Rates (%). Model

Fine-Grained Spatial Reasoning Avg.

Long-Horizon Procedural Planning Avg.

Overall Avg.

RDT GO-1 π0.5 X-VLA

6.1 [5.2, 7.1] 10.6 [9.4, 11.8] 21.4 [19.8, 22.9] 23.9 [22.3, 25.5]

7.7 [6.8, 8.5] 7.0 [6.2, 7.7] 23.1 [21.9, 24.3] 15.9 [14.8, 17.0]

6.9 [6.3, 7.5] 8.8 [8.1, 9.5] 22.3 [21.3, 23.2] 19.9 [18.9, 20.8]

Table 11: Multi-embodiment evaluation of X-VLA on Fine-Grained Spatial Reasoning tasks. Embodiment

GAC-1

SDE-5

CPI-4

RRR-3

CVR-4

L1 L3 L5 L1 L3 L5 L1 L3 L5 L1 L3 L5 L1 L3 L5 Aloha-AgileX ARX-X5 Piper Franka UR5

48 51 48 22 74

22 18 34 12 43

16 48 31 27 29 11 36 29 26 32 16 18 20 17 40 8 14 10 10 38 25 52 41 34 67

13 7 43 31 14 41 30 22 14 7 50 24 13 28 11 13 18 10 40 18 20 50 44 39 18 8 48 18 14 12 8 7 35 26 58 45 38 59 55 49

Table 12: Multi-embodiment evaluation of X-VLA on Long-Horizon Procedural Planning tasks. Embodiment

RPF-1

OFE-1

OCE-5

L1 L3 L5 L1 L3 L5 Aloha-AgileX ARX-X5 Piper Franka UR5

49 54 74 66 83

30 46 66 40 58

27 40 40 24 49

79 55 64 59 96 9 26 20 99 24

L1

L3 L5

CAC-6 L1

28 99 67 18 54 27 100 54 22 99 0 100 16 9 52 11 92 18 0 98 5 93 33 17 100

L3 L5

MIP-3 L1

L3 L5

37 24 100 8 53 29 99 12 31 19 99 0 40 0 100 0 51 48 100 16

0 0 0 0 0

Table 13: RDT per-task Success Rate across difficulty levels.

Task

L1 L2 L3 L4 L5

Task

L1 L2 L3 L4 L5

Pick Blocks Size Pick Blocks Height Pick Blocks Length Pick Blocks Area Pick Blocks Distance Pick Mugs Distance Pick Pill Bottles Distance Pick Mixed Objects Distance A Pick Mixed Objects Distance B Pick Blocks Canonical Pick Cups Canonical Pick Rubik Cubes Canonical Pick Mixed Objects Canonical A Pick Mixed Objects Canonical B Pick Tea Box Referential Pick Cans Referential Pick Soaps Referential Pick Mixed Objects Referential A Pick Mixed Objects Referential B Pick Sauce Can Multi View Pick Seals Multi View Pick Breads Multi View Pick Mixed Objects Multi View A Pick Mixed Objects Multi View B Press Stapler Repeat Lift Pot Repeat Lift Fan Repeat Click Bell Repeat

3 28 6 3 1 15 25 15 20 3 35 7 11 19 7 2 8 2 3 3 3 3 1 5 33 76 39 15

Put Bottles Dustbin Move Blocks Apart Move Playing Cards Away Place Bowls Plates Separate Fries Bread Rank Blocks Height Rank Blocks Color Rank Blocks Size Hit Blocks Hammer Order Click Objects Order Stamp Seals Order Click Bell Clockwise Order Stack Blocks Color Order Stack Blocks Size Order Stack Blocks Length Order Place Phone Press Stapler Place Bottle Cup Hang Mug Stack Blocks Place Object Scale Click Place Burger Fries Click Bell Click Can Place Items Stamp Seals Press Stapler Click Open Place Observe Blocks Move Memory Observe Objects Click Memory Remember Color Cover Remember Orientation Restore Press Stapler Memory

51 6 8 48 24 6 14 5 40 20 3 5 11 9 7 18 56 10 39 33 62 2 0 0 14 47 0 30

1 21 1 1 6 17 15 16 13 5 18 6 4 6 2 1 12 2 2 5 5 5 5 1 37 68 33 17

3 20 1 1 1 12 14 14 14 4 21 6 5 5 5 7 9 7 3 4 5 2 2 4 37 63 26 22

6 15 4 5 3 13 13 6 13 1 17 7 1 6 6 5 5 7 4 3 4 3 3 3 34 56 9 22

1 12 3 3 4 10 13 10 10 5 22 7 5 4 3 6 7 4 2 7 4 0 4 2 32 68 17 23

21

37 2 5 15 9 1 6 0 4 0 0 0 0 0 0 6 29 0 33 20 24 3 1 0 0 14 0 3

22 1 0 12 6 0 1 0 0 0 1 0 0 0 0 1 26 0 26 4 34 0 2 0 0 6 0 0

8 2 0 0 0 0 6 2 1 2 1 0 0 0 0 0 0 0 0 0 1 1 0 0 0 0 0 0 0 0 0 0 6 0 0 0 9 17 0 0 0 2 0 0 0 0 0 0 0 0 0 0 0 0 0 0

Table 14: GO-1 per-task Success Rate across difficulty levels.

Task

L1 L2 L3 L4 L5

Task

L1 L2 L3 L4 L5

Pick Blocks Size Pick Blocks Height Pick Blocks Length Pick Blocks Area Pick Blocks Distance Pick Mugs Distance Pick Pill Bottles Distance Pick Mixed Objects Distance A Pick Mixed Objects Distance B Pick Blocks Canonical Pick Cups Canonical Pick Rubik Cubes Canonical Pick Mixed Objects Canonical A Pick Mixed Objects Canonical B Pick Tea Box Referential Pick Cans Referential Pick Soaps Referential Pick Mixed Objects Referential A Pick Mixed Objects Referential B Pick Sauce Can Multi View Pick Seals Multi View Pick Breads Multi View Pick Mixed Objects Multi View A Pick Mixed Objects Multi View B Press Stapler Repeat Lift Pot Repeat Lift Fan Repeat Click Bell Repeat

33 71 3 4 3 15 36 26 21 16 24 26 21 23 21 16 12 20 17 28 18 21 16 4 27 61 23 16

Put Bottles Dustbin Move Blocks Apart Move Playing Cards Away Place Bowls Plates Separate Fries Bread Rank Blocks Height Rank Blocks Color Rank Blocks Size Hit Blocks Hammer Order Click Objects Order Stamp Seals Order Click Bell Clockwise Order Stack Blocks Color Order Stack Blocks Size Order Stack Blocks Length Order Place Phone Press Stapler Place Bottle Cup Hang Mug Stack Blocks Place Object Scale Click Place Burger Fries Click Bell Click Can Place Items Stamp Seals Press Stapler Click Open Place Observe Blocks Move Memory Observe Objects Click Memory Remember Color Cover Remember Orientation Restore Press Stapler Memory

59 21 52 37 44 34 41 41 36 62 12 61 54 38 30 48 16 8 12 37 77 24 12 3 19 5 2 9

13 62 11 3 2 14 8 19 17 11 17 17 12 15 17 12 8 18 15 33 15 12 8 4 21 53 17 12

12 50 15 2 6 14 15 16 14 8 13 14 6 7 18 4 4 13 15 31 14 9 4 12 12 83 17 3

10 47 7 0 4 8 15 19 8 6 8 13 6 4 9 14 3 15 14 29 12 27 18 12 8 84 21 1

9 32 5 0 3 9 26 12 8 7 7 15 4 3 7 8 3 9 6 27 11 15 18 10 7 77 11 0

77 10 39 23 32 3 14 22 11 11 15 48 12 11 7 14 12 2 14 29 11 17 13 0 1 0 0 2

42 3 33 21 19 0 8 13 3 6 10 34 3 1 5 14 34 0 10 11 14 2 7 0 0 0 0 0

33 11 1 0 13 17 13 5 21 7 0 0 5 2 9 3 3 1 4 3 9 6 21 12 1 0 0 0 0 0 5 1 23 4 0 0 2 3 3 5 5 3 0 0 0 2 0 0 0 0 0 0 0 0 0 0

Table 15: π0.5 per-task Success Rate across difficulty levels.

Task

L1 L2 L3 L4 L5

Task

Pick Blocks Size Pick Blocks Height Pick Blocks Length Pick Blocks Area Pick Blocks Distance Pick Mugs Distance Pick Pill Bottles Distance Pick Mixed Objects Distance A Pick Mixed Objects Distance B Pick Blocks Canonical Pick Cups Canonical Pick Rubik Cubes Canonical Pick Mixed Objects Canonical A Pick Mixed Objects Canonical B Pick Tea Box Referential Pick Cans Referential Pick Soaps Referential Pick Mixed Objects Referential A Pick Mixed Objects Referential B Pick Sauce Can Multi View Pick Seals Multi View Pick Breads Multi View Pick Mixed Objects Multi View A Pick Mixed Objects Multi View B Press Stapler Repeat Lift Pot Repeat Lift Fan Repeat Click Bell Repeat

58 67 22 24 16 39 46 33 41 41 61 51 35 32 33 27 53 33 20 47 32 56 33 38 79 60 89 93

Put Bottles Dustbin 95 Move Blocks Apart 73 Move Playing Cards Away 81 Place Bowls Plates 94 Separate Fries Bread 92 Rank Blocks Height 72 Rank Blocks Color 88 Rank Blocks Size 91 Hit Blocks Hammer Order 81 Click Objects Order 60 Stamp Seals Order 40 Click Bell Clockwise Order 98 Stack Blocks Color Order 91 Stack Blocks Size Order 89 Stack Blocks Length Order 83 Place Phone Press Stapler 46 Place Bottle Cup 97 Hang Mug Stack Blocks 15 Place Object Scale Click 19 Place Burger Fries Click Bell 87 Click Can Place Items 78 Stamp Seals Press Stapler 77 Click Open Place 50 Observe Blocks Move Memory 21 Observe Objects Click Memory 97 Remember Color Cover 100 Remember Orientation Restore 20 Press Stapler Memory 32

27 54 5 13 16 51 36 37 29 22 37 35 15 15 31 21 44 17 21 48 31 51 27 29 41 43 80 84

20 48 11 25 16 52 31 27 21 14 40 33 10 10 31 28 38 12 18 43 31 54 27 30 42 63 73 87

17 30 4 12 10 49 33 29 22 10 41 36 11 10 12 17 19 12 18 42 31 52 33 22 36 54 71 95

13 24 3 10 11 38 33 23 12 11 45 44 11 7 15 8 10 8 8 42 30 61 32 23 45 63 62 90

22

L1 L2 L3 L4 L5 85 51 64 85 84 45 50 66 30 29 40 95 58 43 27 14 79 8 36 89 46 82 22 0 35 48 0 3

81 24 54 61 58 19 46 37 14 16 30 45 25 4 18 17 69 3 56 57 47 8 37 0 3 11 1 0

74 35 52 52 59 24 36 15 7 13 36 31 5 3 3 12 67 0 6 14 14 7 5 0 0 2 0 0

61 18 47 22 36 6 11 8 3 15 35 28 1 0 2 1 30 0 5 34 18 1 11 0 0 0 0 0

Table 16: X-VLA per-task Success Rate across difficulty levels.

Task

L1 L2 L3 L4 L5

Task

Pick Blocks Size Pick Blocks Height Pick Blocks Length Pick Blocks Area Pick Blocks Distance Pick Mugs Distance Pick Pill Bottles Distance Pick Mixed Objects Distance A Pick Mixed Objects Distance B Pick Blocks Canonical Pick Cups Canonical Pick Rubik Cubes Canonical Pick Mixed Objects Canonical A Pick Mixed Objects Canonical B Pick Tea Box Referential Pick Cans Referential Pick Soaps Referential Pick Mixed Objects Referential A Pick Mixed Objects Referential B Pick Sauce Can Multi View Pick Seals Multi View Pick Breads Multi View Pick Mixed Objects Multi View A Pick Mixed Objects Multi View B Press Stapler Repeat Lift Pot Repeat Lift Fan Repeat Click Bell Repeat

48 88 43 52 39 44 34 42 48 36 38 26 29 30 43 36 43 36 44 54 30 37 41 29 49 0 61 29

Put Bottles Dustbin 79 Move Blocks Apart 86 Move Playing Cards Away 97 Place Bowls Plates 67 Separate Fries Bread 78 Rank Blocks Height 91 Rank Blocks Color 93 Rank Blocks Size 97 Hit Blocks Hammer Order 84 Click Objects Order 39 Stamp Seals Order 34 Click Bell Clockwise Order 57 Stack Blocks Color Order 99 Stack Blocks Size Order 78 Stack Blocks Length Order 99 Place Phone Press Stapler 51 Place Bottle Cup 98 Hang Mug Stack Blocks 23 Place Object Scale Click 57 Place Burger Fries Click Bell 98 Click Can Place Items 54 Stamp Seals Press Stapler 75 Click Open Place 3 Observe Blocks Move Memory 46 Observe Objects Click Memory 18 Remember Color Cover 100 Remember Orientation Restore 18 Press Stapler Memory 41

27 81 30 27 44 42 42 40 40 23 23 18 10 15 38 22 30 21 34 53 38 37 28 27 35 0 70 24

22 65 21 25 33 25 42 36 31 10 20 14 13 15 27 19 31 26 23 62 32 33 30 33 30 0 59 20

23 56 28 28 29 29 34 34 37 10 19 5 7 13 23 24 15 14 16 57 30 41 28 36 29 0 51 19

16 55 14 16 33 23 64 26 27 8 11 12 7 6 14 13 14 17 13 59 35 34 22 33 27 0 43 20

L1 L2 L3 L4 L5 64 74 96 44 87 95 92 88 36 6 71 15 93 87 74 28 44 5 22 59 41 86 2 11 0 48 6 1

55 81 92 29 71 84 91 65 16 0 65 1 67 29 59 25 19 3 19 18 37 4 0 5 0 8 0 0

31 17 87 15 79 72 82 32 7 0 65 0 33 2 14 22 16 2 5 12 34 2 0 3 0 0 0 0

28 31 83 8 66 46 63 24 4 0 31 0 18 1 0 7 6 0 2 4 24 1 0 0 0 0 0 0

Table 17: RDT per-task Object-Normalized Target Accuracy (ON T A) across difficulty levels. Fine-grained spatial reasoning tasks only. Task

L1

L2

L3

L4

L5

Pick Blocks Size -94.0 -48.5 -29.3 -17.5 -18.8 Pick Blocks Height -44.0 -18.5 -6.7 -6.3 -5.6 Pick Blocks Length -88.0 -48.5 -32.0 -20.0 -16.4 Pick Blocks Area -94.0 -48.5 -32.0 -18.8 -16.4 Pick Blocks Distance -48.5 -12.8 -13.1 -9.1 -6.7 Pick Mugs Distance -27.5 -10.7 -10.0 -4.4 -5.0 Pick Pill Bottles Distance -12.5 -13.3 -7.5 -4.4 -1.5 Pick Mixed Objects Distance A -27.5 -12.0 -7.5 -12.8 -5.0 Pick Mixed Objects Distance B -20.0 -16.0 -7.5 -4.4 -5.0 Pick Blocks Canonical -45.5 -14.0 -9.7 -11.4 -3.6 Pick Cups Canonical 2.5 1.6 9.7 6.6 14.9 Pick Rubik Cubes Canonical -39.5 -12.8 -7.4 -4.6 -1.5

Task

L1

L2

L3

L4

L5

Pick Mixed Objects Canonical A -33.5 -15.2 -8.6 -11.4 -3.6 Pick Mixed Objects Canonical B -21.5 -12.8 -8.6 -5.8 -4.7 Pick Tea Box Referential -86.0 -47.0 -26.7 -12.8 -9.1 Pick Cans Referential -96.0 -48.5 -24.0 -14.0 -5.8 Pick Soaps Referential -84.0 -32.0 -21.3 -14.0 -4.6 Pick Mixed Objects Referential A -96.0 -47.0 -24.0 -11.6 -8.0 Pick Mixed Objects Referential B -94.0 -47.0 -29.3 -15.2 -10.3 Pick Sauce Can Multi View -45.5 -26.7 -28.0 -29.3 -24.0 Pick Seals Multi View -45.5 -26.7 -26.7 -28.0 -28.0 Pick Breads Multi View -45.5 -26.7 -30.7 -29.3 -33.3 Pick Mixed Objects Multi View A -48.5 -26.7 -30.7 -29.3 -28.0 Pick Mixed Objects Multi View B -42.5 -32.0 -28.0 -29.3 -30.7

23

Table 18: GO-1 per-task Object-Normalized Target Accuracy (ON T A) across difficulty levels. Fine-grained spatial reasoning tasks only. Task

L1

L2

L3

L4

L5

Pick Blocks Size -34.0 -30.5 -17.3 -12.5 -9.2 Pick Blocks Height 42.0 43.0 33.3 33.8 18.4 Pick Blocks Length -94.0 -33.5 -13.3 -16.3 -14.0 Pick Blocks Area -92.0 -45.5 -30.7 -25.0 -20.0 Pick Blocks Distance -45.5 -17.6 -7.4 -8.0 -7.8 Pick Mugs Distance -27.5 -14.7 -7.5 -10.4 -6.2 Pick Pill Bottles Distance 4.0 -22.7 -6.3 -2.0 13.7 Pick Mixed Objects Distance A -11.0 -8.0 -5.0 2.8 -2.7 Pick Mixed Objects Distance B -18.5 -10.7 -7.5 -10.4 -7.3 Pick Blocks Canonical -26.0 -6.8 -5.1 -5.8 -1.5 Pick Cups Canonical -14.0 0.4 0.6 -3.5 -1.5 Pick Rubik Cubes Canonical -11.0 0.4 1.7 2.1 7.3

Task

L1

L2

L3

L4

L5

Pick Mixed Objects Canonical A -18.5 -5.6 -7.4 -5.8 -4.7 Pick Mixed Objects Canonical B -15.5 -2.0 -6.3 -8.0 -5.8 Pick Tea Box Referential -58.0 -24.5 -9.3 -9.2 -4.6 Pick Cans Referential -68.0 -32.0 -28.0 -3.2 -3.5 Pick Soaps Referential -76.0 -38.0 -28.0 -16.4 -9.1 Pick Mixed Objects Referential A -60.0 -23.0 -16.0 -2.0 -2.4 Pick Mixed Objects Referential B -66.0 -27.5 -13.3 -3.2 -5.8 Pick Sauce Can Multi View -8.0 10.7 8.0 5.3 2.7 Pick Seals Multi View -23.0 -13.3 -14.7 -17.3 -18.7 Pick Breads Multi View -18.5 -17.3 -21.3 2.7 -13.3 Pick Mixed Objects Multi View A -26.0 -22.7 -28.0 -9.3 -9.3 Pick Mixed Objects Multi View B -44.0 -28.0 -17.3 -17.3 -20.0

Table 19: π0.5 per-task Object-Normalized Target Accuracy (ON T A) across difficulty levels. Fine-grained spatial reasoning tasks only. Task

L1

L2

L3

L4

L5

Pick Blocks Size 16.0 -9.5 -6.7 -3.8 -4.4 Pick Blocks Height 34.0 31.0 30.7 12.5 8.8 Pick Blocks Length -56.0 -42.5 -18.7 -20.0 -16.4 Pick Blocks Area -52.0 -30.5 0.0 -10.0 -8.0 Pick Blocks Distance -26.0 -0.8 4.0 -1.2 1.1 Pick Mugs Distance 8.5 34.7 40.0 38.8 27.7 Pick Pill Bottles Distance 19.0 14.7 13.8 19.6 21.8 Pick Mixed Objects Distance A -0.5 16.0 8.8 14.8 10.2 Pick Mixed Objects Distance B 11.5 5.3 1.3 22.0 -2.7 Pick Blocks Canonical 11.5 6.4 1.7 -1.2 2.9 Pick Cups Canonical 41.5 24.4 31.4 33.6 40.0 Pick Rubik Cubes Canonical 26.5 22.0 23.4 28.0 38.9

Task

L1

L2

L3

L4

L5

Pick Mixed Objects Canonical A 2.5 -2.0 -2.9 -0.1 2.9 Pick Mixed Objects Canonical B -2.0 -2.0 -2.9 -1.2 -1.5 Pick Tea Box Referential -34.0 -3.5 8.0 -5.6 4.4 Pick Cans Referential -46.0 -18.5 4.0 0.4 -3.5 Pick Soaps Referential 6.0 16.0 17.3 2.8 -1.2 Pick Mixed Objects Referential A -34.0 -24.5 -17.3 -5.6 -3.5 Pick Mixed Objects Referential B -60.0 -18.5 -9.3 1.6 -3.5 Pick Sauce Can Multi View 20.5 30.7 24.0 22.7 22.7 Pick Seals Multi View -2.0 8.0 8.0 8.0 6.7 Pick Breads Multi View 34.0 34.7 38.7 36.0 48.0 Pick Mixed Objects Multi View A -0.5 2.7 2.7 10.7 9.3 Pick Mixed Objects Multi View B 7.0 5.3 6.7 -4.0 -2.7

Table 20: X-VLA per-task Object-Normalized Target Accuracy (ON T A) across difficulty levels. Fine-grained spatial reasoning tasks only. Task

L1

L2

L3

L4

L5

Pick Blocks Size -4.0 -9.5 -4.0 3.8 -0.8 Pick Blocks Height 76.0 71.5 53.3 45.0 46.0 Pick Blocks Length -14.0 -5.0 -5.3 10.0 -3.2 Pick Blocks Area 4.0 -9.5 0.0 10.0 -0.8 Pick Blocks Distance 8.5 32.8 23.4 20.1 25.6 Pick Mugs Distance 16.0 22.7 6.3 14.8 10.2 Pick Pill Bottles Distance 1.0 22.7 27.5 20.8 58.0 Pick Mixed Objects Distance A 13.0 20.0 20.0 20.8 13.7 Pick Mixed Objects Distance B 22.0 20.0 13.8 24.4 14.8 Pick Blocks Canonical 4.0 7.6 -2.9 -1.2 -0.4 Pick Cups Canonical 7.0 7.6 8.6 8.9 2.9 Pick Rubik Cubes Canonical -11.0 1.6 1.7 -6.9 4.0

Task

L1

L2

L3

L4

L5

Pick Mixed Objects Canonical A -6.5 -8.0 0.6 -4.6 -1.5 Pick Mixed Objects Canonical B -5.0 -2.0 2.9 2.1 -2.5 Pick Tea Box Referential -14.0 7.0 2.7 7.6 3.3 Pick Cans Referential -28.0 -17.0 -8.0 8.8 2.1 Pick Soaps Referential -14.0 -5.0 8.0 -2.0 3.3 Pick Mixed Objects Referential A -28.0 -18.5 1.3 -3.2 6.6 Pick Mixed Objects Referential B -12.0 1.0 -2.7 -0.8 2.1 Pick Sauce Can Multi View 31.0 37.3 49.3 42.7 45.3 Pick Seals Multi View -5.0 17.3 9.3 6.7 13.3 Pick Breads Multi View 5.5 16.0 10.7 21.3 12.0 Pick Mixed Objects Multi View A 11.5 4.0 6.7 4.0 -4.0 Pick Mixed Objects Multi View B -6.5 2.7 10.7 14.7 10.7

24

Table 21: RDT per-task step Progress Score across difficulty levels. Step tables include long-horizon tasks only. Task

L1

L2

L3

L4

L5

Press Stapler Repeat 33.0 37.0 39.3 36.5 35.2 Lift Pot Repeat 76.0 75.5 75.7 72.0 80.0 Lift Fan Repeat 39.0 41.5 41.3 29.8 34.0 Click Bell Repeat 15.0 21.5 28.7 26.5 28.6 Put Bottles Dustbin 51.0 56.0 48.3 33.8 33.6 Move Blocks Apart 6.0 7.0 12.7 11.0 8.0 Move Playing Cards Away 8.0 13.0 8.0 9.0 10.6 Place Bowls Plates 48.0 37.5 39.0 32.3 20.6 Separate Fries Bread 24.0 26.0 33.0 28.0 35.4 Rank Blocks Height 6.0 50.5 40.0 37.8 31.4 Rank Blocks Color 14.0 53.0 44.3 37.5 31.4 Rank Blocks Size 5.0 50.0 38.3 33.8 28.0 Hit Blocks Hammer Order 40.0 14.5 9.3 6.5 7.2 Click Objects Order 20.0 23.5 22.7 17.5 0.0 Stamp Seals Order 3.0 3.0 5.7 6.3 6.6 Click Bell Clockwise Order 5.0 50.0 38.3 33.8 28.0

Task

L1

L2

L3

L4

L5

Stack Blocks Color Order 11.0 50.0 33.7 25.0 20.2 Stack Blocks Size Order 9.0 0.0 33.7 25.3 20.4 Stack Blocks Length Order 7.0 50.0 34.0 25.3 20.4 Place Phone Press Stapler 18.0 40.0 44.0 49.3 38.2 Place Bottle Cup 56.0 52.5 52.7 44.3 33.6 Hang Mug Stack Blocks 10.0 9.0 17.7 13.0 9.6 Place Object Scale Click 39.0 41.5 41.3 29.8 34.0 Place Burger Fries Click Bell 33.0 23.5 17.7 15.3 11.4 Click Can Place Items 62.0 55.5 75.7 55.0 65.0 Stamp Seals Press Stapler 2.0 5.0 1.3 0.8 3.0 Click Open Place 0.0 16.5 10.0 22.8 17.2 Observe Blocks Move Memory 0.0 0.0 0.0 0.0 0.0 Observe Objects Click Memory 14.0 7.0 5.3 6.3 5.8 Remember Color Cover 47.0 23.5 15.0 5.5 6.0 Remember Orientation Restore 0.0 0.5 0.7 0.0 0.2 Press Stapler Memory 30.0 7.0 4.0 3.3 5.8

Table 22: GO-1 per-task step Progress Score across difficulty levels. Step tables include long-horizon tasks only. Task

L1

L2

L3

L4

L5

Press Stapler Repeat 27.0 25.5 14.7 8.3 8.8 Lift Pot Repeat 61.0 54.0 92.0 92.8 91.6 Lift Fan Repeat 23.0 22.5 24.7 26.5 21.2 Click Bell Repeat 16.0 19.0 8.7 9.3 6.3 Put Bottles Dustbin 59.0 84.5 68.0 60.3 48.4 Move Blocks Apart 21.0 21.5 12.0 8.3 6.6 Move Playing Cards Away 52.0 43.5 45.0 20.5 21.0 Place Bowls Plates 37.0 36.0 35.0 26.8 20.8 Separate Fries Bread 44.0 37.5 26.7 27.8 14.6 Rank Blocks Height 34.0 51.5 39.3 36.0 30.0 Rank Blocks Color 41.0 57.0 44.7 36.5 30.6 Rank Blocks Size 41.0 61.0 47.3 40.5 34.0 Hit Blocks Hammer Order 36.0 29.0 24.3 27.5 27.0 Click Objects Order 62.0 42.0 30.3 24.3 3.0 Stamp Seals Order 12.0 29.5 24.0 23.3 21.4 Click Bell Clockwise Order 61.0 52.0 38.3 30.8 25.2

Task

L1

L2

L3

L4

L5

Stack Blocks Color Order 54.0 56.0 36.7 25.8 20.0 Stack Blocks Size Order 38.0 16.5 5.0 4.3 4.2 Stack Blocks Length Order 30.0 12.0 10.3 4.3 3.0 Place Phone Press Stapler 48.0 33.0 34.3 30.0 28.8 Place Bottle Cup 16.0 16.5 50.3 68.5 57.0 Hang Mug Stack Blocks 8.0 23.3 18.5 18.3 13.1 Place Object Scale Click 12.0 19.0 14.3 6.8 8.2 Place Burger Fries Click Bell 37.0 32.5 13.7 5.0 6.6 Click Can Place Items 77.0 45.5 53.3 51.0 55.4 Stamp Seals Press Stapler 24.0 23.0 8.3 5.3 5.2 Click Open Place 12.0 13.0 7.0 0.0 5.8 Observe Blocks Move Memory 3.0 3.0 1.3 11.0 11.6 Observe Objects Click Memory 19.0 6.5 8.3 5.8 5.2 Remember Color Cover 5.0 0.0 0.0 0.3 0.0 Remember Orientation Restore 2.0 6.0 5.7 6.5 5.4 Press Stapler Memory 9.0 6.5 9.3 9.5 7.6

Table 23: π0.5 per-task step Progress Score across difficulty levels. Step tables include long-horizon tasks only. Task

L1

L2

L3

L4

L5

Press Stapler Repeat 79.0 61.5 63.0 59.0 64.8 Lift Pot Repeat 60.0 64.5 77.0 71.8 74.8 Lift Fan Repeat 89.0 85.0 84.0 80.8 75.4 Click Bell Repeat 93.0 89.0 90.7 98.0 94.2 Put Bottles Dustbin 95.0 90.5 93.0 87.8 82.4 Move Blocks Apart 73.0 67.5 62.0 61.5 55.6 Move Playing Cards Away 81.0 78.5 77.3 81.5 77.8 Place Bowls Plates 94.0 91.5 78.0 69.5 50.4 Separate Fries Bread 92.0 88.5 83.0 86.8 74.4 Rank Blocks Height 72.0 72.5 63.7 68.0 53.0 Rank Blocks Color 88.0 77.5 73.0 68.5 55.4 Rank Blocks Size 91.0 84.0 69.7 57.5 47.8 Hit Blocks Hammer Order 81.0 33.0 27.3 23.3 19.8 Click Objects Order 60.0 35.0 27.0 28.0 15.0 Stamp Seals Order 40.0 50.5 43.0 43.6 46.6 Click Bell Clockwise Order 98.0 96.5 52.7 40.8 38.6

Task

L1

L2

L3

L4

L5

Stack Blocks Color Order 91.0 80.0 58.7 37.0 23.8 Stack Blocks Size Order 89.0 75.0 49.3 36.3 25.0 Stack Blocks Length Order 83.0 65.5 52.0 33.3 27.0 Place Phone Press Stapler 46.0 31.0 34.7 32.5 26.6 Place Bottle Cup 97.0 90.0 87.3 86.3 75.8 Hang Mug Stack Blocks 15.0 48.0 44.3 26.5 23.6 Place Object Scale Click 19.0 50.5 56.0 22.8 30.2 Place Burger Fries Click Bell 87.0 90.0 78.0 61.8 72.0 Click Can Place Items 78.0 69.0 78.3 69.0 73.0 Stamp Seals Press Stapler 77.0 86.0 50.0 48.8 45.0 Click Open Place 50.0 48.5 55.0 30.8 37.4 Observe Blocks Move Memory 21.0 15.0 16.0 13.5 14.4 Observe Objects Click Memory 97.0 58.5 33.0 25.5 14.4 Remember Color Cover 100.0 50.5 20.3 12.0 7.0 Remember Orientation Restore 20.0 12.0 14.7 13.3 12.0 Press Stapler Memory 32.0 10.0 9.7 8.8 9.4

25

Table 24: X-VLA per-task step Progress Score across difficulty levels. Step tables include long-horizon tasks only. Task

L1

L2

L3

L4

L5

Press Stapler Repeat 49.0 42.5 40.3 38.8 36.6 Lift Pot Repeat 0.0 0.0 0.0 0.0 0.0 Lift Fan Repeat 61.0 79.5 72.0 71.8 66.2 Click Bell Repeat 29.0 28.5 27.0 23.3 24.8 Put Bottles Dustbin 79.0 80.0 78.3 71.0 74.0 Move Blocks Apart 86.0 80.5 89.7 65.0 71.2 Move Playing Cards Away 97.0 98.0 95.7 93.6 93.6 Place Bowls Plates 67.0 62.5 52.7 39.3 30.4 Separate Fries Bread 78.0 92.5 87.3 89.5 82.2 Rank Blocks Height 91.0 97.5 93.3 91.0 83.0 Rank Blocks Color 93.0 96.5 95.0 92.5 90.0 Rank Blocks Size 97.0 94.5 83.3 71.0 67.0 Hit Blocks Hammer Order 84.0 57.5 46.7 40.5 30.8 Click Objects Order 39.0 20.0 10.3 6.5 0.0 Stamp Seals Order 34.0 77.0 78.3 78.3 53.2 Click Bell Clockwise Order 57.0 26.0 14.7 9.8 7.2

Task

L1

L2

L3

L4

L5

Stack Blocks Color Order 99.0 96.5 79.0 54.5 37.8 Stack Blocks Size Order 78.0 93.5 55.0 29.3 22.2 Stack Blocks Length Order 99.0 87.5 73.7 41.8 24.2 Place Phone Press Stapler 51.0 45.5 46.7 43.5 43.8 Place Bottle Cup 98.0 71.0 61.3 49.0 40.0 Hang Mug Stack Blocks 23.0 51.0 51.7 41.3 31.4 Place Object Scale Click 57.0 35.0 30.3 27.3 20.0 Place Burger Fries Click Bell 98.0 78.5 55.7 57.3 49.8 Click Can Place Items 54.0 64.5 67.7 69.3 64.2 Stamp Seals Press Stapler 75.0 89.0 61.0 49.3 42.8 Click Open Place 15.0 19.5 18.0 24.8 28.8 Observe Blocks Move Memory 46.0 29.5 27.7 27.3 19.4 Observe Objects Click Memory 18.0 4.5 5.0 2.8 2.6 Remember Color Cover 100.0 50.0 18.3 6.5 4.0 Remember Orientation Restore 18.0 19.5 7.0 6.3 5.6 Press Stapler Memory 41.0 6.5 3.0 4.3 4.8

26

Geometric Attribute Cognition

Geometric Attribute Cognition

Task name: Pick Blocks Size Instruction: Grab the block ranked 5th by size with one arm.

Task name: Pick Blocks Height Instruction: Lift the block that is 3rd in height order.

Geometric Attribute Cognition

Geometric Attribute Cognition

Task name: Pick Blocks Length Instruction: Use one arm to lift the 2nd longest block.

Task name: Pick Blocks Area Instruction: Grab the block ranked 1st by base area with one arm.

Spatial Distance Estimation

Spatial Distance Estimation

Task name: Pick Blocks Distance Instruction: Lift the block that is the 2nd farthest from the orange block with one arm.

Task name: Pick Mugs Distance Instruction: Lift the 1st farthest mug from the green mug with one arm.

Spatial Distance Estimation

Spatial Distance Estimation

Task name: Pick Pill Bottles Distance Instruction: Take hold of the 3rd farthest pill bottle from the orange pill bottle using one arm.

Task name: Pick Mixed Objects Distance A Instruction: Find the 6th farthest object from the fan and grab it.

Spatial Distance Estimation

Canonical Position Indexing

Task name: Pick Mixed Objects Distance B Instruction: Lift the object that is the 3rd farthest from the cup.

Task name: Pick Blocks Canonical Instruction: Pick up the block in the 1st row counted from far to near at the 1st position counted from left to right.

Figure 11: Fine-grained spatial reasoning task examples in RoboSPA (Part 1).

27

Canonical Position Indexing

Canonical Position Indexing

Task name: Pick Cups Canonical Instruction: Find the 4th cup in the 1st row counted from near to far, with positions counted from left to right, and pick it up.

Task name: Pick Rubik Cubes Canonical Instruction: Pick up the 2nd rubik’s cube counted from right to left in the 1st row counted from near to far.

Canonical Position Indexing

Canonical Position Indexing

Task name: Pick Mixed Objects Canonical A Instruction: Pick up the object in the 2nd row counted from far to near at the 2nd position counted from left to right.

Task name: Pick Mixed Objects Canonical B Instruction: Use one arm to take the target object in the 2nd row counted from far to near at the 1st position counted from right to left.

Referential Relational Reasoning

Referential Relational Reasoning

Task name: Pick Tea Box Relational Instruction: Front is the side farther from the robot. Grab the nearest one to the front-right of the center tea box using one arm.

Task name: Pick Cans Relational Instruction: The side farther from the robot is front. Use an arm to lift the nearest one directly at the back of the middle-left can.

Referential Relational Reasoning

Referential Relational Reasoning

Task name: Pick Soaps Relational Instruction: The side farther from the robot is front. Pick the nearest one directly in front of the back-right soap bar.

Task name: Pick Mixed Objects Relational A Instruction: Back means the side nearer to the robot. Take the closest object to the front-left of the back-middle object.

Referential Relational Reasoning

Cross-View Reasoning

Task name: Pick Mixed Objects Relational B Instruction: Back means the side nearer to the robot. Lift the nearest object directly in front of the middle-left object.

Task name: Pick Sauce Can Multi View Instruction: From the robot's side of the table, pick up the leftmost object.

Figure 12: Fine-grained spatial reasoning task examples in RoboSPA (Part 2).

28

Cross-View Reasoning

Cross-View Reasoning

Task name: Pick Seals Multi View Instruction: Seen from the opposite side of the table, take the rightmost object with an arm.

Task name: Pick Breads Multi View Instruction: Viewed from the right side, use one arm to pick up the leftmost object.

Cross-View Reasoning

Cross-View Reasoning

Task name: Pick Mixed Objects Multi View A Instruction: From the left side of the table, pick up the frontmost object.

Task name: Pick Mixed Objects Multi View B Instruction: With the camera on the opposite side of the table, grab the backmost object with an arm.

Figure 13: Fine-grained spatial reasoning task examples in RoboSPA (Part 3).

29

Repetitive Procedure Following

Task name: Press Stapler Repeat Instruction: Press the blue stapler with shiny silver parts directly 5 times.

Repetitive Procedure Following

Task name: Lift Pot Repeat Instruction: Lift and put down the cooking kitchen pot with side grips with arms 5 times.

Repetitive Procedure Following

Task name: Lift Fan Repeat Instruction: Lift the white fan and set it down 5 times.

Repetitive Procedure Following

Task name: Click Bell Repeat Instruction: Press the top center of the compact tabletop bell 5 times.

Order-Free Execution

Task name: Put Bottles Dustbin Instruction: Lift the five beverage bottles and drop them into the plastic dustbin.

Figure 14: Long-horizon procedural planning task examples in RoboSPA (Part 1).

30

Order-Free Execution

Task name: Move Blocks Apart Instruction: Set every magenta block on the right and every blue block on the left.

Order-Free Execution

Task name: Move Playing Cards Away Instruction: Move all boxes of playing cards to the left and right.

Order-Free Execution

Task name: Place Bowls Plates Instruction: Stack each medium-sized bowl for holding food one by one onto the dish plate.

Order-Free Execution

Task name: Separate Fries Bread Instruction: Place every bread at the front and every fries at the back.

Order-Free Execution

Task name: Rank Blocks Height Instruction: Order the blocks in height order, tallest left to shortest right.

Figure 15: Long-horizon procedural planning task examples in RoboSPA (Part 2).

31

Order-Free Execution

Task name: Rank Blocks Color Instruction: Order red, green, blue, yellow, and purple blocks from left to right.

Order-Free Execution

Task name: Rank Blocks Size Instruction: Sort the blocks by size from left to right, largest to smallest.

Order-Constrained Execution

Task name: Hit Blocks Hammer Order Instruction: Lift the grippy handle hammer using the right arm to hit the red, green, blue, yellow, then purple block.

Order-Constrained Execution

Task name: Click Objects Order Instruction: Press the rectangular alarm-clock, followed by the blue stapler, the red beverage can with curved top, the playingcards storage box, and the soft golden bread.

Order-Constrained Execution

Task name: Stamp Seals Order Instruction: Grab the brown seal, stamp on Olive, Coral, Teal, Tan, then Beige.

Figure 16: Long-horizon procedural planning task examples in RoboSPA (Part 3).

32

Order-Constrained Execution

Task name: Click Bell Clockwise Order Instruction: Click all five bells starting from the frontmost one and moving clockwise.

Order-Constrained Execution

Task name: Stack Blocks Color Order Instruction: Stack the blocks at the center from bottom to top as red, green, blue, yellow, and purple.

Order-Constrained Execution

Task name: Stack Blocks Size Order Instruction: Stack the blocks at the center from largest to smallest, bottom to top.

Order-Constrained Execution

Task name: Stack Blocks Length Order Instruction: Use the arms to build a center stack from longest to shortest, bottom to top.

Composite Action Coordination

Task name: Place Phone Press Stapler Instruction: Place the sleek phone on the flat base phone stand, press the curved black stapler thrice, then empty tabletop trash bin into the trash bin.

Figure 17: Long-horizon procedural planning task examples in RoboSPA (Part 4).

33

Composite Action Coordination

Task name: Place Bottle Cup Instruction: Move three bottles in the trash bin one after another, and place the two plastic cups for drinks on the two flat round wooden coasters one by one.

Composite Action Coordination

Task name: Hang Mug Stack Blocks Instruction: First hang the solid black ceramic mug on the metallic rack with smooth finish, then stack red block, green block, blue block, and yellow block in order from bottom to top.

Composite Action Coordination

Task name: Place Object Scale Click Instruction: Start by placing the bell with black flat base on the black and white electronic scale, then put the toy car made of plastic material on the smooth textured black display stand, and press the box with cards inside, then put the bell with black flat base back on the table, and put the toy car made of plastic material back on the table.

Composite Action Coordination

Task name: Place Burger Fries Click Bell Instruction: Deposit the orange hamburger with white underside on the orange tray, then deposit palm-sized red fries container on the orange tray, tap the plastic and metal bell once, then place both back.

Composite Action Coordination

Task name: Click Can Place Items Instruction: Press the brown can, move the bread block with light patches to the beige plastic breadbasket, click the brown can, set the bottle with large orange label on the pad, and finally press the brown can.

Figure 18: Long-horizon procedural planning task examples in RoboSPA (Part 5).

34

Composite Action Coordination

Task name: Stamp Seals Press Stapler Instruction: Tap the left seal twice on left, tap the right seal twice on right, then press the stapler with curved blue top once.

Composite Action Coordination

Task name: Click Bell Open Microwave Place Object Instruction: Ring the plastic and metal bell, open the countertop microwave, ring the plastic and metal bell again, position the coffee-box with smooth card board left of the dark gray mouse, and ring the plastic and metal bell.

Memory-Intensive Planning

Task name: Observe Blocks Move Memory Instruction: Observe the colors of the blocks on the display platform, then after they are hidden, move matching blocks on the table closer to you.

Memory-Intensive Planning

Task name: Observe Objects Click Memory Instruction: Observe the objects on the display platform from left to right, then after they are hidden, click the matching objects on the table in the same order.

Memory-Intensive Planning

Task name: Remember Color Cover Instruction: Memorize the colors of the blocks, then once they are covered, uncover the blocks in the order of orange, black, green, magenta, red.

Figure 19: Long-horizon procedural planning task examples in RoboSPA (Part 6).

35

Memory-Intensive Planning

Task name: Remember Orientation Restore Instruction: Watch the front T-shaped blocks, then after they are hidden, make the rear T-shaped blocks match the front ones in orientation from left to right.

Memory-Intensive Planning

Task name: Press Stapler Memory Instruction: Press the five staplers in order moving left to right, 3, 5, 4, 1, and 2 times, respectively.

Figure 20: Long-horizon procedural planning task examples in RoboSPA (Part 7).

36

Record · ID 660852 · SHA-256 bef67138e60a16b3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.