Conceptio › Archive › arXiv CS
arXiv CSopen access

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies Xingyu Lin1 , Zhuang Li1 , Zhongrun Wu2 , Shouquan Zhou3 , and Dehui Du1

arXiv:2609.21659v1 [cs.RO] 18 Sep 2026

Abstract—Vision-language-action (VLA) policies solve the same manipulation task through different action interfaces, but task success alone does not establish whether their physical executions agree. We study cross-policy end-effector geometry in 15,000 closed-loop LIBERO rollouts from four policies. The primary clean-condition analysis forms 3,600 configuration-matched, and therefore dependent, policy pairs. Both-success pairs have a median normalized dynamic time warping distance of 0.0120 m versus 0.0380 m when exactly one policy succeeds. This ordering holds in every task, every policy pair, and nine sampling and bandlimited representations; however, the ratio varies severalfold across representations, so we report the direction rather than a fixed multiple. Both-failure pairs are more separated again but rest on thin, uneven support, so we report them as exploratory. Within successful executions, partner replacements separate more across tasks than across initial states. A matched baseline still reveals measurable, heterogeneous residual policy differences, so a low cross-policy distance does not imply interchangeability. Successful executions sit about as far from same-task demonstrations as those demonstrations sit from each other, compatible with task-associated geometry without separating training-data overlap from task constraints. A common 72-action window preserves the ordering but reduces its magnitude; endpoint and duration adjustment likewise leaves a positive mixed-outcome coefficient relative to both-success pairs, though its magnitude is specificationdependent. Under composite visual stress, policy rankings and pair composition change together.

I. I NTRODUCTION

Vision-language-action (VLA) policies map images and task instructions to robot actions, and their action interfaces now diverge sharply. OpenVLA produces discretized action tokens [1], OpenVLA-OFT (OFT) predicts continuous action chunks [2], UniVLA decodes learned latent actions [3], and π0.5 uses a flow-based action model [4]. These policies are compared almost exclusively by a task success rate. A leaderboard number, however, is silent about how the arm actually moves: two policies that each succeed on 90% of trials can reach the goal through different routes, and two that fail can diverge or fail after following similar routes. This gap matters in deployment: replacing one policy with another at the same success rate can silently change motion and workspace use. It also matters scientifically, because success counts do not show whether heterogeneous VLA pipelines trace similar end-effector paths toward a goal or follow distinct paths. Recent work enriches evaluation with per-policy behavioral metrics and robustness profiles [5], [6], [7], [8], [9], comparing policies mainly through per-policy scores or the sets of scenes each solves. We ask a complementary question about pairs: under a matched task, initial state, instruction, and visual condition, how close are the end-effector paths of two different policies, and how does that closeness depend on their joint outcome? Task success alone does not specify how closely two end-effector paths agree. We measure their agreement and its sensitivity to endpoint proximity and episode duration, in an

offline comparison rather than online failure prediction or safety certification. Conditioning on the joint outcome is central, not cosmetic. Each pair enters a both-success (SS), mixed (SF), or bothfailure (FF) stratum. Pooling by “agreeing outcome” is misleading because successes dominate that pool, so a low pooled distance would describe successful execution while appearing to describe failure. Comparing successes against failures also changes endpoints, termination times, and task completion— properties of the measurement problem, not nuisances that alignment removes. On 3,600 configuration-matched pairs from the clean LIBERO-Spatial grid [10], SS paths have a median distance of 0.0120 m versus 0.0380 m for SF. This ordering holds in every task, policy pair, and representation examined, but a common observation window attenuates the gap. Within successful executions, matched controls locate the agreement as largely task-associated, while a within-policy baseline reveals small, heterogeneous residual policy differences rather than equivalence. FF has a larger median (0.0439 m), but its thin support and representation-dependent rank relative to SF make it exploratory. Our contribution is this finding and the comparison design that supports it: outcome-conditioned strata; matched task/state controls; a within-policy baseline; a demonstration reference at matched task identity; task–state block resampling; fixed observation windows; explicit alignment sensitivity; and an analysis of how visual stress changes conditional geometry and which pairs enter the comparison. II. R ELATED W ORK A. Behavioral Evaluation and Robustness VLATest mutates manipulation scenes and instructions [11], LIBERO-Plus expands evaluation across environmental and input dimensions [9], NEBULA separates capability tests from operational probes such as timing and action stability [6], RoboEval adds per-policy behavioral metrics that discriminate policies of similar success rate [5], and Distracted Robot reports distinct vulnerabilities and low agreement on which cluttered scenes policies solve [8]. Each characterizes one policy, or its set of solved scenes, at a time. We instead compare the paired endeffector paths of two different policies at a shared configuration and condition every comparison on their joint outcome; our composite image ladder serves only to expose how performance and pair composition change within one recorded grid. Benchmark audits of narrow evaluation distributions [12] and deployment reports of differing behavior across implementations [13] motivate retaining task-level outcomes and reporting the execution interface rather than treating a policy name as a complete experimental specification.

a

b Enumerate the six cross-policy pairs

One configuration, four physical paths

Spatial T7 · state 9 · template A · clean L0

Outcome class and normalized DTW (m) OpenVLA

0.0489

UniVLA

0.0641

z y x

0.1 m / axis

π0.5

OpenVLA F UniVLA S

c

OFT F π0.5 S

FF

OFT

SF

SF

0.0652

d = 0.0117 m

Partner: T7 / state 9

UniVLA

SF

0.0180 SF

0.0165

SS

0.0117

One example; these are not population medians.

Anchor: π0.5 Partner: UniVLA

Keep a successful anchor; replace only its partner Same task + same state

OFT S: success F: failure

Other state, same task

d = 0.0156 m

Partner: T7 / state 0

Other task

d = 0.0789 m

Partner: T5 / state 0

Policy pair, template, L0 and success are fixed. The three control views share a world frame and projection scale.

Fig. 1. Primary comparison design. (a) Recorded end-effector paths from a selected clean configuration with two successes and two failures. (b) The six unordered cross-policy pairs; blank cells omit self- and duplicate pairings, and distances describe this example, not population medians. (c) A successful π0.5 anchor is retained while its UniVLA partner changes state or task, holding policy pair, template, level, and successful outcomes fixed. Views are orthographic projections; DTW uses the original 3D coordinates.

Language-focused benchmarks provide stronger controls than a few fixed templates. LIBERO-Para decomposes meaningpreserving linguistic variation over object and action axes [14], multilingual evaluation examines the changing influence of language during execution [15], LIBERO-CF changes feasible instructions under familiar layouts to probe visual dominance [16], and wording-sensitivity studies show action outputs shift with instruction phrasing [17]. Our template contrast supplies limited supporting evidence about one checkpoint; the primary controls ask how recorded physical paths differ when task or state matching changes.

of both executions, not one execution’s outcome. Controls: success-preserving task and state replacements, plus a withinpolicy baseline. Inference: descriptive conditional geometry, with no distance threshold classifying a failure mechanism. Several methods derive failure signals from policy representations, uncertainty, or information-theoretic quantities [20], [21], [22]; these internal signals and our physical-path distances describe different spaces, since latent agreement can coexist with path separation. Our fixed-window analysis tests whether a geometric contrast is observable within a fixed horizon; it evaluates no detector.

B. Physical Trajectories and Failure Signals

III. S TUDY D ESIGN AND M EASUREMENTS A. Tasks, Policies, and the Archive The common benchmark is LIBERO-Spatial: ten tasks, the first twenty states in each task’s initial-state bank, three instruction templates, and five visual levels. Each policy contributes 3,000 rollouts, for 12,000 Spatial executions; a further 3,000 Object rollouts cover π0.5 only, so cross-policy geometry is restricted to Spatial. The data have three units: an episode is one policy execution, a configuration fixes task, initial state, template, and visual level, and a comparison is one unordered pair of policies at that configuration. At L0, 600 configurations yield 2,400 episodes and 3,600 pairs; over L0–L4 the counts are 3,000, 12,000, and 18,000. Every policy/configuration cell has exactly one archived execution (Fig. 1). Spatial instructions use a task-specific relation r in three templates (e.g., A, “pick up the black bowl r and place it on the plate”); Object templates analogously name an object and the basket.

Valle et al. evaluate action uncertainty and physical execution quality of individual runs [7]; Krüger et al. compare VLA inspection trajectories against reference paths in position and orientation [18]; and dynamic time warping (DTW) itself is an established sequence-alignment method [19]. Our comparison object is two executed end-effector paths from different policies, grouped by joint outcome and compared against successful task/state replacements. The closest positional comparison appears in LIBERO-Para’s trajectory appendix, which builds a pseudo-reference from successful runs, resamples paths, applies DTW to end-effector positions, compares by outcome, and thresholds the distance to label failure types [14]. We therefore claim no novelty in positional DTW comparison, resampling, or outcome-based grouping. Four differences define the increment. Object: matched executions of two different policies, not one execution against a successful pseudo-reference. Conditioning: the joint outcome

TABLE I I NTERFACES USED BY THE COLLECTOR . E/W DENOTE EXTERNAL / WRIST VIEWS AND P DENOTES PROPRIOCEPTION . T HE PIXEL GRID IS WHERE VISUAL CORRUPTION IS APPLIED . C ADENCE IS EXECUTED ACTIONS PER POLICY QUERY. Policy

Inputs

Pixel grid

Cadence

OpenVLA OFT UniVLA π0.5

E E,W,P E E,W,P

2242 2562 2242 2562

1 8 1 5

TABLE II F INAL OPERATORS PER VISUAL LEVEL . B RIGHTNESS IS UNIFORM ON THE INTERVAL . N OISE IS IN 8- BIT PIXEL UNITS ; SHIFT IS IN PIXELS . M ASK NOTATION GIVES COUNT AND SIDE LENGTH . Level

Brightness

σ

Shift

Masks

L0 L1 L2 L3 L4

1 [0.8,1.2] [0.6,1.4] [0.45,1.6] [0.3,1.9]

0 10 20 35 55

0 ±8 horiz. ±16 both ±24 both ±32 both

none none 1 × 402 2 × 602 3 × 802

The collector renders 256 × 256 observations, takes ten initialization steps, and permits up to 240 control actions, recording end-effector positions, emitted actions, gripper state, and tracked-object positions after each step. Success is the simulator’s task-completion signal and terminates the episode; all archived failures reach the control cap. Array lengths define time, and initialization records hold zero action vectors. The archive contains 15,000 distinct keys and matching trajectory files; we checked array dimensions, finite values, record lengths, and source-file identity. The first three policies use Spatial-finetuned checkpoints; the π0.5 launcher selects its LIBERO configuration. OpenVLA and UniVLA query every control step. OpenVLA decodes tokens deterministically; UniVLA retains latent-action history and samples at temperature 0.75, top-p 0.9. OFT uses a deterministic L1-trained continuous head with an eight-action queue; π0.5 samples a flow-based chunk and executes five actions before replanning (Table I). OpenVLA and OFT therefore decode without sampling, UniVLA and π0.5 are stochastic. Training data, checkpoints, horizons, and preprocessing also differ from one another and from official recipes, so these results describe the recorded implementations. B. Composite Visual Operators L0 leaves observations unchanged. Higher levels combine brightness scaling, Gaussian pixel noise, and wraparound shifts that alter content without moving the camera. From L2 onward, they also add possibly overlapping square masks—grayscale 128 at L2 and random at L3–L4—before clipping pixels to [0,255] (Table II). These are fixed nominal settings, not an adaptive search. OpenVLA and UniVLA receive corruption after resizing while OFT and π0.5 receive it on the original render, so equal mask dimensions cover different area fractions across policies; both queried views are corrupted for two-view policies. Each rollout’s corruption seed hashes suite, policy, task, template, level, and state. Because we saved neither realized

seeds nor queried images, matching a nonzero level matches operator settings, not the realized corruption across policies. The stress comparison therefore concerns these complete pipelines, folding together decoder, input, preprocessing, and cadence. C. Outcome-Conditioned Geometry Let Xm,c be the recorded end-effector path of policy m at configuration c, and let qm,c ∈ {0, 1} denote its observed outcome. For every configuration, we enumerate all six unordered policy pairs. A pair is SS when qm,c + qn,c = 2, SF when the sum is 1, and FF when it is 0, so the strata are defined by executions rather than by permanent groups of policies (Fig. 1(b)). For each path, retain every fifth recorded position in the common simulator frame. Given sampled sequences X = (x1 , . . . , xLX ) and Y = (y1 , . . . , yLY ), the local cost is cij = ∥xi − yj ∥2 . We compute Dij = cij + min{Di−1,j , Di,j−1 , Di−1,j−1 }, DLX ,LY , d(X, Y ) = LX + LY

(1) (2)

with D00 = 0 and the remaining boundary entries infinite. The denominator is the sum of sequence lengths, not the alignmentpath length. The statistic is in meters and measures position discrepancy under unconstrained temporal warping; it omits orientation, contact, and force, so it is a geometric distance, not a tracking error. Among the 68 clean configurations with two successful and two failed policies, Fig. 1 shows the one whose three stratum medians best match the pooled L0 medians; population results use all eligible pairs, not this example. Our primary estimate is the empirical distribution of d within each outcome stratum and level. An SS median therefore describes only pairs that both succeeded, not every pair on the grid; strata also differ in task and policy-pair composition. We retain those counts and inspect task-, template-, and policy-pair-specific results alongside pooled summaries. D. Successful Controls and State-Variation Baseline For each clean SS pair, keep the first policy in alphabetical archive-key order as an anchor and replace its partner with a successful trajectory of the same partner policy and template. The within-task control selects another initial state and the across-task control another task, both retaining L0 and successful outcomes. One eligible partner per control, sampled with seed 20260908, yields 2,588 matched triples, and twenty fixed seeds assess partner-choice sensitivity. These controls separate exact state matching from the joint effect of task geometry and goal specification. A second control asks whether cross-policy discrepancy exceeds within-policy state variation. For anchor policy A at state s, select another state s′ at which both A and partner policy B succeed on the same task and template. Compare dwithin = d(XA,s , XA,s′ ), δ = dcross − dwithin .

dcross = d(XA,s , XB,s′ ), (3)

The same replacement state is used on both sides. Four SS pairs have no eligible common state, leaving 2,584 comparisons;

the primary seed is 20260909 and twenty seeds test partnersampling sensitivity. The fixed alphabetical anchor makes this control directional, measuring state variation among archived successes. Swapping which policy is retained leaves the pooled increment positive. A third reference set is external: the official LIBERO demonstrations for these ten tasks, fifty per task and 500 in total, from a pinned dataset revision checked file by file against published hashes. The stored demonstrations use different upstream trimming from our archive, so this comparison resamples both sides to fifty equally spaced normalized-time positions before applying the same recurrence. Its values are consequently not term-by-term comparable with the original-sampling distances above and are never subtracted from them. For each execution we take the median distance over the reference set, then report the median over executions, using same-task and other-task references and the leave-one-out distance among demonstrations as the reference scale. No state matching to demonstrations is available, and we make no claim about which files trained any checkpoint. E. Dependence, Alignment, and the Observation Window An episode appears in multiple policy pairs, so we resample task–state blocks, not individual pair distances. Each of 2,000 bootstrap replicates draws twenty states with replacement within each of the ten fixed tasks, yielding percentile 95% intervals conditional on those tasks. Successful controls reuse replacement trajectories across blocks, so we report descriptive medians and sampling ranges, not independent-pair confidence intervals. To test sampling and spatial alignment, we first remove initialization and recompute distances. We then linearly resample each path at fifty equally spaced cumulative arc-length positions (same recurrence, denominator 100), reducing the influence of differing record counts and dwell times. Subtracting each resampled path’s centroid additionally removes translation while preserving scale and orientation. A stricter test resamples to fifty equally spaced normalized-time positions and then restricts the warping path to a Sakoe–Chiba band; because a band on a fifty-point grid has an integer half-width, we report the realized fraction rather than the requested one. None of these operations removes task phases or endpoint constraints. A separate diagnostic fixes the observation window at 72 control records, or 30% of the control cap. We retain only pairs whose two episodes remain active beyond that point. After dropping ten startup records, every-fifth sampling gives fifteen positions per path; recomputing full-episode distances on the same cohort separates cohort selection from truncation. The cutoff was fixed before this reanalysis and outcomes still refer to eventual success, making this an explanatory comparison, not an online classifier. IV. R ESULTS : S UCCESS -A SSOCIATED T RAJECTORY P ROXIMITY A. Successful Paths Agree More Closely The clean SS median is 0.0120 m (n = 2,588), compared with 0.0380 m for SF (n = 923) and 0.0439 m for FF (n = 89), as shown in Fig. 2(a). Task-stratified bootstrap intervals for the median contrasts are [0.0235,0.0284] m for SF minus SS and [0.0245,0.0506] m for FF minus SS.

The SS–SF ordering holds in all ten tasks, six policy pairs and three templates (Fig. 3), and in nine representations (Section IV-D). Policy-pair SS medians range from 0.0109 to 0.0143 m and SF medians from 0.0341 to 0.0431 m, and the per-task counts show the pooled contrast is not a single task-invariant effect size. The SS–SF ordering also holds in each of the three policy pairs that exclude OFT (Fig. 3). The result is a stable ordering over the representations examined, not a universal distance scale: a common observation window attenuates the gap, and successful-control comparisons reveal residual policy differences. The distributions overlap despite their different medians: the empirical 5th–95th percentile ranges are 0.0076–0.0207 m for SS and 0.0167–0.0820 m for SF, so a single distance in their overlap does not identify a joint outcome: the finding concerns distributions, not a per-path classifier. FF evidence is far less balanced: 64 of its 89 pairs are OpenVLA–OFT, other pairs contribute 1–11 each, and three tasks have no FF pair at all (Fig. 3). Excluding that dominant pair leaves 25 FF pairs with a median of 0.0381 m. Within tasks, the FF median already falls below the SF median at the two tasks holding one and three FF pairs, and Section IV-D shows the same reversal at the pooled level under other representations. We therefore treat FF as exploratory: it is more separated than SS, but it supports neither a stable SS<SF<FF ordering nor a taxonomy of failure mechanisms. Conditioning matters even before comparing distances. Pooling pairs by agreeing outcome mixes the SS and FF distributions, placing weight w = nSS /(nSS + nFF ) on the successful component. At L0, w = 0.967, so a low pooled distance characterizes successful execution and cannot be read as evidence of a common failed path. B. Task Identity and Exact State Matching The successful-control medians are 0.0120 m for the original task/state match, 0.0157 m for another state within the same task, and 0.0817 m for another task (Fig. 2(b)). Across twenty partner samplings the within-task median lies between 0.0155 and 0.0158 m and the across-task median between 0.0798 and 0.0835 m, so the larger between-task contrast is stable under these partner choices. The interpretation is joint: replacing a task changes the arrangement, required motion, and instruction goal together, so shared demonstrations, learned priors, and manipulation constraints may all contribute. Within this archive, successful execution is more geometrically consistent within a task than across tasks. The official demonstrations give an external reference at matched task identity. Under normalized-time resampling, successful executions (n = 2,033) have a median same-task reference distance of 0.0167 m, against 0.0164 m for leave-oneout distances among the 500 demonstrations; other-task values are 0.0837 and 0.0841 m respectively, and failed executions (n = 367) sit at 0.0374 m from same-task demonstrations. These similar magnitudes are compatible with task-associated positional geometry, not statistical equivalence. The comparison does not separate training-data overlap, task constraints, and termination rules, and it establishes no checkpoint’s training-set

a

b

Outcome-conditioned distances

Successful partner controls

Median

Median

Both succeed SS

0.0120

Same task same state

0.0120

Mixed SF

0.0380 n = 923

Same task other state

0.0157

Both fail FF

0.0439

Other task

0.0817

n = 2,588

n = 89

.005

.01

.02

.05

.1

.2

n = 2,588

n = 2,588

n = 2,588

.005

Normalized DTW (m; log scale) Thin line: 5th–95th pct.

.01

.02

.05

.1

.2

Normalized DTW (m; log scale)

Thick line: 25th–75th pct.

Point: median

Rug: all pairs

Fig. 2. Clean L0 distances. (a) Configuration-matched pairs by outcome. (b) Successful partner controls. Marks are defined in the figure key; both panels use a logarithmic scale and describe observed distributions, not confidence bounds.

SS medians < SF medians in all 19 subgroups SS

SF

FF

FF: no fixed rank vs SF n SS·SF·FF

FF n<10

Task T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Policy pair OFT–OpenVLA OFT–π0.5 OFT–UniVLA OpenVLA–π0.5 OpenVLA–UniVLA π0.5–UniVLA Template A B C

295·64·1 145·187·28 339·21·— 348·12·— 239·113·8 179·167·14 336·24·— 218·128·14 285·72·3 204·135·21 325·211·64 402·195·3 392·197·11 456·140·4 440·154·6 573·26·1 896·281·23 829·328·43 863·314·23

.01

.02

.05

.1

Median normalized DTW (m; log)

Fig. 3. Subgroup ordering at L0. Each row gives the SS and SF median on one logarithmic axis for ten tasks, six policy pairs and three templates from the same 3,600 pairs; SS lies left of SF in every row. FF is shown for reference (hollow: n < 10; dash: absent) and holds no fixed rank against SF. Counts are SS·SF·FF; rows are subgroups, not independent samples.

membership. The demonstration comparison provides a reference scale within the resampled metric, rather than a tolerance for policy replacement. C. Residual Policy Differences at a Common Replacement State The common-state control in Eq. 3 gives same-policy and cross-policy medians of 0.0133 and 0.0156 m over 2,584 comparisons. The median paired difference is 0.0021 m (distinct from the 0.0023 m gap between the two medians); twenty partner samplings span 0.00195–0.00219 m. The pooled increment hides policy-pair structure (Table III): the three OFT-anchored comparisons have larger increments than the other three. The OFT–OpenVLA increment (0.0051 m) arises between two deterministic decoders, so it cannot be attributed to inference-sampling variance, while the smallest increments involve a stochastic policy. A pre-registered repeat

TABLE III S UCCESSFUL STATE - VARIATION CONTROLS . W ITHIN AND CROSS DENOTE e IS THE MEDIAN OF PAIRED DIFFERENCES , THE TWO DISTANCES IN E Q . 3; δ NOT A DIFFERENCE OF MEDIANS . D ISTANCES ARE IN METERS . T HE FIRST LISTED POLICY IS THE ANCHOR . Anchor / partner

n

Within

Cross

δe

OFT / OpenVLA OFT / π0.5 OFT / UniVLA OpenVLA / π0.5 OpenVLA / UniVLA π0.5 / UniVLA Pooled

323 401 391 456 440 573 2584

0.0117 0.0124 0.0121 0.0142 0.0144 0.0141 0.0133

0.0169 0.0176 0.0162 0.0147 0.0146 0.0146 0.0156

0.0051 0.0059 0.0043 0.0007 0.0003 0.0006 0.0021

experiment tests this directly: 1,200 executions rerun every policy five times at each of ten tasks, three initial states and two templates under five fixed seed labels. Within a configuration the non-sampling decoders reproduce byte-identical paths (withinpolicy median 0 over 480 and 380 successful rerun pairs), while the stochastic policies do not (0.0053 m over 596 pairs for UniVLA, 0.0067 m over 588 for π0.5 ). Across the 265 of 360 cells where both policies succeeded at least twice, the cross-policy median exceeds the mean of the two within-policy medians by 0.0082 m (state-block 95% interval [0.0077,0.0085] m). Crosspolicy separation thus exceeds rerun variability, although the result remains conditional on these tasks, states, seed labels, and success-dependent eligibility. Because the anchor convention is directional these are not a symmetric ranking, but they establish that shared successful geometry coexists with measurable, heterogeneous policy differences: low cross-policy distance does not imply equivalence. D. Sensitivity to Sampling and Spatial Alignment Successful clean executions use a median of 102 control actions while failures all use 240, so full-episode comparisons span different durations. Removing initialization, resampling at equal arc length, and additionally centering each path all leave SS below SF (Table IV). Centering reduces the across-task successful-control median from 0.0794 to 0.0556 m, still above the same-task/state median

TABLE IV C LEAN METRIC SENSITIVITY: MEDIAN NORMALIZED DTW ( M ). SS/SF/FF COUNTS ARE 2,588/923/89. S TATE / TASK COLUMNS REPLACE THE SUCCESSFUL PARTNER WITHIN / ACROSS TASKS , USING THE SAME 2,588 TRIPLES . Metric Original No startup Arc-length + Centering

SS

SF

FF

State

Task

0.0120 0.0132 0.0117 0.0094

0.0380 0.0401 0.0301 0.0305

0.0439 0.0458 0.0478 0.0375

0.0157 0.0163 0.0146 0.0109

0.0817 0.0885 0.0794 0.0556

TABLE V O BSERVATION - WINDOW COMPARISON ON THE SAME ACTIVE COHORT. E NTRIES ARE MEDIAN DTW ( M ). F ULL RETAINS INITIALIZATION ; TRIMMED REMOVES IT; PREFIX USES THE FIRST 72 CONTROL RECORDS . B OTH EPISODES ARE ACTIVE BEYOND THE CUTOFF .

TABLE VI E MPIRICAL SUCCESS RATES , 600 ROLLOUTS PER POLICY AND LEVEL . O BJECT IS EVALUATED ONLY FOR π0.5 . Policy / suite

L0

L1

L2

L3

L4

OpenVLA / Spatial OFT / Spatial UniVLA / Spatial π0.5 / Spatial π0.5 / Object

0.762 0.673 0.962 0.992 0.977

0.837 0.650 0.957 0.983 0.987

0.683 0.522 0.697 0.978 0.968

0.068 0.320 0.110 0.567 0.637

0.000 0.013 0.000 0.007 0.003

the entire separation, though full-episode geometry magnifies it. This is an explanatory comparison on a survived-to-72 cohort, not a failure-onset estimate.

E. A Common Control Window Attenuates the Gap

F. Residual Outcome Association After Endpoint and Duration Adjustment Success shortens episodes and pulls endpoints together, so the raw ordering could partly reflect duration and endpoint proximity. These are consequences of outcome, not pre-treatment confounders, and endpoint mismatch shares terms with d, so this is a mechanical adjustment, not causal control. Regressing d on SF/FF indicators over the 3,600 clean pairs and adding regressors in stages, the SF coefficient falls from 0.0294 (raw) to 0.0277 (task and pair fixed effects), 0.0156 (adding endpoint mismatch), and 0.0048 (adding length gap and mean length); FF falls from 0.0508 to 0.0141. Both stay positive with task–state cluster-bootstrap 95% intervals excluding zero at every stage (full model SF [0.0017, 0.0077], FF [0.0034, 0.0260]). The full specification is strongly collinear, which bounds what it can settle: measured over the complete design rather than the continuous covariates alone, variance inflation factors reach 17.1 for mean length, 11.9 for the length gap and 10.1 for the SF indicator, because outcome, duration, and endpoint proximity carry largely overlapping information. The surviving coefficient is accordingly specification-dependent rather than an identified effect: two pre-specified simpler duration controls (length gap only, mean length only) leave SF at 0.0082 and 0.0058, both above the full-model value. A permutation test reassigning outcomes within each task–state–template block (preserving per-block success counts and pair structure) places the observed SF−SS median contrast (0.0260) beyond all 5,000 permutations (null median 0.0130; p = 0.0002; per-task Cliff’s δ 0.76–0.997). The association therefore does not vanish under the measured endpoint and duration variables, and Section IV-D shows it is not an artifact of unconstrained warping: these data establish a residual positional association, not a decomposition of it.

The 72-control diagnostic retains 3,503 pairs (2,496 SS, 918 SF, all 89 FF), excluding 97 whose episodes have already ended. Full-episode medians on this cohort are almost unchanged, while the SF–SS median contrast contracts markedly (Table V). The fixed-prefix SF-minus-SS contrast is 0.01007 m (task-stratified 95% interval [0.00842,0.01172] m); FF-minusSS is 0.02005 m ([0.01494,0.02875] m). The SS–SF ordering remains in all ten tasks and six policy pairs, but the pooled contrast is only about 39% of its samecohort full-episode value. Matching observed prefix length while conditioning on survival to 72 records preserves the ordering, so unequal duration and later behavior do not create

V. P ERFORMANCE AND C OMPOSITION UNDER V ISUAL S TRESS A. The Clean Ranking Does Not Persist The clean Spatial ordering is π0.5 > UniVLA > OpenVLA > OFT. At L3 it becomes π0.5 > OFT > UniVLA > OpenVLA (Table VI). UniVLA falls from 577/600 to 66/600 successes and OFT from 404/600 to 192/600, and the surviving L3 successes concentrate on a few tasks (e.g., 31 of OpenVLA’s 41 occur on T8). This rank change applies only to the tested pipelines under the nominal operators. All four policies lose success between

Outcome

n

Full

Trimmed

Prefix

SS SF FF

2496 918 89

0.0120 0.0380 0.0439

0.0132 0.0402 0.0458

0.0140 0.0240 0.0340

of 0.0094 m, so translation contributes to but does not account for the contrast. Restricting temporal warping tests whether the contrast is an artifact of free alignment. On the fifty-point normalized-time grid, integer half-widths of 9, 4 and 2 realize band fractions of 18.37%, 8.16% and 4.08%; we quote realized fractions because a band set as a percentage of sequence length can be widened by a length difference until it no longer binds. We compare nine representations: unbanded length-sum normalization (dividing by the summed sample counts), unbanded alignmentpath averaging (dividing by the number of aligned cells), synchronous same-index comparison (mean pointwise distance with no warping at all), and the first two under each of the three bands. Across all nine the SS median stays within 0.0115– 0.0378 m and SF within 0.0368–0.1696 m, and SS lies below SF in every task and policy pair. The SF/SS ratio nevertheless ranges from 2.53 to 6.13, since narrowing the band raises the SF median sharply once long temporal detours are forbidden. The direction is therefore robust to these choices while the magnitude is a property of the representation, and no single multiple should be read as an effect size. FF exceeds SF in the two unbanded resampled normalizations but falls below it under synchronous comparison and all six banded settings, which is why we report no three-stratum ordering.

TABLE VII N ORMALIZED DTW MEDIANS ( M ) AND COUNTS . E ACH LEVEL CONTAINS 3,600 CONFIGURATION - MATCHED PAIRS . SS

SF

FF

Level

d

n

d

n

d

n

L0 L1 L2 L3 L4

0.0120 0.0122 0.0133 0.0206 –

2588 2652 1928 223 0

0.0380 0.0372 0.0354 0.0382 0.0580

923 864 1328 1471 36

0.0439 0.0443 0.0301 0.0446 0.0678

89 84 344 1906 3564

L0 and L3, and π0.5 leads at every level up to L3. The result does not establish statistical independence between capability and robustness, which theory suggests are coupled [23], and the differing views, preprocessing, corruption draws, and query schedules prevent attribution to an action representation alone. L1 is not uniformly harmful—OpenVLA rises from 0.762 to 0.837—so the response is not monotonic, while at L4 all Spatial policies fall below 0.014. On Object, π0.5 likewise holds through L2 and degrades at L3–L4, extending one policy’s qualitative profile without replicating the four-policy analysis across suites. OFT’s clean-grid success rate is below its published LIBERO-Spatial result [2]; the evaluations use different instruction/state grids and control horizons, and this archive does not identify the source of the discrepancy, so the reported comparisons concern the recorded pipelines, not a reproduction of published leaderboard performance. B. Changing Outcomes Change the Comparison Population SS distance changes modestly through L2, then rises to 0.0206 m at L3 (Table VII). The population also turns over: an agreeing-outcome distribution changes through both its component distributions and their weights, and the 223 surviving SS pairs at L3 are not a longitudinal panel of all clean successes. Matching pair identities across levels separates composition from geometry: 195 of those 223 pairs are also SS at L0 and 28 are new, and on that shared support the median paired increase is 0.0076 m at L3, against 0.0016 m at L2 and 0.0003 m at L1. SS is the least separated group wherever it exists, but SF and FF do not follow one universal hierarchy: at L2 FF has a lower median than SF, and at L4 there are no SS pairs and 3,564 of the 3,600 comparisons are FF, leaving 36 SF pairs. Distances and sample composition must therefore be read together. C. Instruction Templates: a Limited Clean Contrast At L0, template A exceeds B by 0.160 for OpenVLA across 200 matched task–state cells (exact McNemar p = 4.22 × 10−5 , Holm-adjusted 1.69 × 10−4 ); OFT, UniVLA, and π0.5 differences are −0.035, +0.005, +0.015 (adjusted p 0.75, 1.00, 0.75). One execution per cell cannot separate wording from inference variability, so this is exploratory support for systematic language evaluation [14], [15]. VI. D ISCUSSION AND R ESEARCH I MPLICATIONS A. What the Geometry Establishes The measurements establish a specific regularity whose direction, but not size, is stable across the representations examined. Within successful executions the structure is primarily taskassociated, while changing task also changes goal geometry

and instruction; it is not policy equivalence, since a withinpolicy baseline exposes a small, heterogeneous cross-policy increment. Plausible sources include a shared bowl-placement form, common reach/grasp/place constraints, and overlapping demonstrations. The reference in Section IV-B places successful executions at the scale of the demonstrations’ own mutual spread, consistent with all three and separating none. Path agreement is also not execution quality: similar positions can accompany different orientations, contacts, forces, or unnecessary motion, absent from our position-only distance [7], nor correct language grounding [16]. B. Implications for Cross-Policy Evaluation For offline cross-policy comparison, success rates and outcome-conditioned path distances answer different questions: a low SS distance does not establish interchangeability, and cross-level changes must be read together with changes in which pairs enter each outcome stratum. Cross-policy trajectory reports should state the outcome-conditioned population, show support by task and policy pair, preserve outcomes in partner controls, separate complete episodes from common windows, and report the alignment constraint with the resulting magnitude. These retrospective analyses are not deployable monitors: prospective use requires a decision clock, task-heldout evaluation, and threshold calibration [20], [21]. C. Scope and Reproducibility The evidence remains conditional on four checkpoints, ten simulated tasks, and twenty initial states per task, with directional anchoring and trajectory reuse in the controls. The main archive provides one execution per cell, so repeatedinference behaviour rests on the separate repeat experiment, whose eligibility is success-dependent. The demonstration reference compares magnitudes in a resampled representation without state matching. Realized corruption seeds and queried images are absent. Both-failure comparisons are the thinnest evidence here and should not be extended beyond the recorded pairs. VII. C ONCLUSION Across four VLA policy pipelines on LIBERO-Spatial, executions that both succeed are more similar in end-effector positional geometry than executions with mixed outcomes. That direction persists under band-limited alignment, a common observation window, and endpoint and duration adjustment, while its magnitude depends on the representation. Within successful executions, matched task and state replacements locate the agreement as largely task-associated, and a withinpolicy baseline still reveals residual policy differences; bothfailure comparisons stay too thinly supported to order reliably. Cross-policy trajectory evaluation is thus most informative when conditioning, control matching, observation duration, alignment constraints, and sample support are reported together. To support reproduction and community use, our code and data will be made publicly available upon publication. ACKNOWLEDGMENT OpenAI Codex and GPT-5.6-sol assisted with drafting, code for analysis and figures, and language refinement; GLM assisted

earlier drafting; MiniMax speech-2.8-hd generated the video narration. Rollouts were collected by the evaluated policies in simulation; all scientific content, analysis, and conclusions are the authors’ own. R EFERENCES [1] M. J. Kim, et al., “OpenVLA: An open-source vision-language-action model,” in Proceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 270. PMLR, 2025, pp. 2679–2713. [2] M. J. Kim, C. Finn, and P. Liang, “Fine-tuning vision-language-action models: Optimizing speed and success,” in Proceedings of Robotics: Science and Systems, 2025. [3] Q. Bu, et al., “Learning to act anywhere with task-centric latent actions,” in Proceedings of Robotics: Science and Systems, 2025. [4] Physical Intelligence, et al., “π0.5 : a vision-language-action model with open-world generalization,” arXiv preprint arXiv:2504.16054, 2025. [5] Y. R. Wang, et al., “RoboEval: Where robotic manipulation meets structured and scalable evaluation,” arXiv preprint arXiv:2507.00435, 2026. [6] J. Peng, Y. Zhang, Y. Duan, T. Liang, V. Chaudhary, and Y. Yin, “NEBULA: Do we evaluate vision-language-action agents correctly?” arXiv preprint arXiv:2510.16263, 2025. [7] P. Valle, C. Lu, S. Ali, and A. Arrieta, “Evaluating uncertainty and quality of vision-language-action-enabled robots,” arXiv preprint arXiv:2507.17049, 2025. [8] A. Rasouli, et al., “Distracted Robot: How visual clutter undermine robotic manipulation,” arXiv preprint arXiv:2511.22780, 2025. [9] S. Fei, et al., “LIBERO-Plus: A progressive robustness benchmark for visual-language-action models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026, pp. 38 574–38 583. [10] B. Liu, et al., “LIBERO: Benchmarking knowledge transfer for lifelong robot learning,” in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 44 776–44 791. [11] Z. Wang, Z. Zhou, J. Song, Y. Huang, Z. Shu, and L. Ma, “VLATest: Testing and evaluating vision-language-action models for robotic manipulation,” Proceedings of the ACM on Software Engineering, vol. 2, pp. 1615–1638, 2025. [12] T. Jiang, X. Tan, S. Wheeler, L. Sun, T. W. Ayalew, and M. Walter, “What are we actually benchmarking in robot manipulation?” arXiv preprint arXiv:2606.04233, 2026. [13] Y. Zhang, Y. Qi, and X. Zheng, “Experiences from benchmarking vision-language-action models for robotic manipulation,” arXiv preprint arXiv:2511.11298, 2025. [14] C. Kim, M. Kim, M. Kang, H. Kim, and D. Jung, “LIBERO-Para: A diagnostic benchmark and metrics for paraphrase robustness in VLA models,” arXiv preprint arXiv:2603.28301, 2026. [15] X. Dong, Z. Han, T. Niu, Q. Zhu, and W. Che, “When does language matter? multilingual instructions reveal step-wise language sensitivity in vision-language-action models,” in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2026, pp. 44 615– 44 629. [16] Y. Fang, et al., “When vision overrides language: Evaluating and mitigating counterfactual failures in VLAs,” arXiv preprint arXiv:2602.17659, 2026. [17] J. Woo, “Task-dependent sensitivity of VLA models to instruction wording,” Authorea preprint, 2026. [18] M. Krüger, M. Salem, and M. Reischl, “Assessment of a fine-tuned visionlanguage-action model for robotic feature-following inspection,” Neural Processing Letters, vol. 58, p. 41, 2026. [19] H. Sakoe and S. Chiba, “Dynamic programming algorithm optimization for spoken word recognition,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 26, no. 1, pp. 43–49, 1978. [20] Q. Gu, et al., “SAFE: Multitask failure detection for vision-languageaction models,” in Advances in Neural Information Processing Systems, vol. 38, 2025, pp. 40 041–40 076. [21] S. Park, et al., “Hide-and-seek in trajectories: Discovering failure signals for VLA runtime monitoring,” arXiv preprint arXiv:2605.30834, 2026. [22] J. Yang, et al., “Tri-Info: Generalizable, interpretable failure prediction for VLA models via information theory,” arXiv preprint arXiv:2606.19998, 2026. [23] J. Tai, “Capability and robustness cannot both be free: An informationtheoretic bound for vision-language-action models,” arXiv preprint arXiv:2605.25889, 2026.

Record · ID 1006909 · SHA-256 727d60c0b53c0ba3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.