Conceptio › Archive › arXiv CS
arXiv CSopen access

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies Vivek Chavan1,2,∗

Pengtao Xie2,∗,†

Yahuan Shi2,†

Oliver Heimann1

Kevin Haninger1,‡

Jörg Krüger1,2,‡

1 Fraunhofer Institute for Production Systems and Design Technology IPK 2 Technische Universität Berlin

arXiv:2609.05376v1 [cs.RO] 4 Sep 2026

∗ Equal contribution.

† Work conducted as part of student projects at Fraunhofer IPK.

Abstract—Visuomotor imitation policies can achieve high performance in curated environments yet fail when visually similar objects compete with task-relevant entities. We study this behavior as a problem of conditional visual grounding: which visual entity matters depends on the current manipulation phase and, in more complex tasks, on the inferred task state. Using Action Chunking with Transformers (ACT), we systematically vary object and receptacle competitors and localize failures to picking and placement. The resulting errors are cue- and phase-specific and are accompanied by corresponding changes in learned visual representations. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for strengthening task-relevant grounding while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same diagnosisintervention principle in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the correct destination depends on the observed instrument state. Together, the results support conditional visual grounding as a useful framework for diagnosing and improving robustness across distinct visuomotor policy-learning regimes. 8 Page paper! Index Terms—Robot Imitation Learning, Visual Grounding, Visual Distractors, ACT, Vision-Language-Action Models.

I. I NTRODUCTION Vision-based imitation policies are commonly trained and evaluated under curated visual conditions, while deployment scenes can contain visually similar objects, receptacles, and irrelevant clutter [1], [2]. Adding such competitors changes neither the manipulation objective nor the underlying robot dynamics, yet can substantially reduce policy success [2], [3]. Aggregate success alone, however, cannot distinguish between a policy that has lost the ability to execute the manipulation and one that retains the motor behavior but applies it to the wrong entity. We study this distinction as conditional visual grounding. Given candidate referents Rt , an instruction ℓ, and execution context ct , correct action requires selecting a contextdependent referent rt⋆ ∈ Rt . Crucially, the relevant referent is not necessarily static. For Action Chunking with Transformers (ACT) [4], we examine phase-conditioned relevance: the manipulated object is critical during target acquisition

‡ Supervision

and grasping, whereas the receptacle becomes critical during placement. In our vision-language-action (VLA) case study, relevance is additionally state-conditioned: the observed state of a medical instrument determines which receptacle is the correct destination. We ask whether visual failures can be localized to the cue, referent, and execution stage at which grounding becomes incorrect, and whether this diagnosis suggests effective interventions. Controlled ACT experiments first expose distinct object- and receptacle-selection bottlenecks. We then evaluate lightweight object-centric interventions and validate their effect in simulation and on a physical UR3e. Finally, we examine whether the same diagnosis-intervention principle remains useful in a pretrained VLA policy on a state-conditioned routing task. II. R ELATED W ORK Visuomotor imitation policies are vulnerable to spurious observation-action correlations [5], and recent work has systematically quantified generalization failures under changes in object appearance, scene configuration, and visual distractors [1]. In particular, Causal-ACT shows that ACT can exploit irrelevant distractor correlations [3], while ImitDiff and DRAIL improve robustness through semantic guidance and task-relevant region-aware representations, respectively [2], [6]. Complementary approaches provide explicit spatial or visual guidance through grounding masks and visual prompts [7], [8]. Visual robustness is likewise unresolved in pretrained visionlanguage-action policies: BYOVLA demonstrates sensitivity to task-irrelevant visual content and mitigates it through runtime observation interventions [9], while recent work explicitly supervises task-relevant factors in VLA action representations [10]. Our focus is complementary; we jointly diagnose which visual cues induce incorrect grounding, which task referent is affected, and when the resulting error becomes behaviorally consequential, and use this diagnosis to motivate interventions across distinct visuomotor policy-learning regimes.

III. M ETHOD a) Controlled diagnosis with ACT.: We study two simulated pick-and-place tasks designed to separate target and destination grounding. In Task 1, a randomly placed object is moved to a fixed receptacle; in Task 2, both object and receptacle positions are randomized, requiring visual localization during both picking and placement. ACT is trained from 100 clean scripted demonstrations per task. At evaluation, held-out object and receptacle competitors match the target in color, shape, or neither, with one to three competitors per class. The full mixed condition contains one competitor from each object and receptacle class. In addition to end-to-end success, we measure P (pick), P (lift | pick), and P (place | pick, lift), allowing failures to be localized within the manipulation sequence. Simulation results aggregate three training seeds, four evaluation seeds, and 50 closed-loop rollouts per seed. b) Diagnosis-driven intervention.: The diagnosis motivates three complementary mechanisms targeting incorrect grounding. Image-space copy-paste augmentation inserts synthetic competitors into otherwise clean demonstrations, reducing the reliability of incidental appearance correlations. Phase-dependent attention regularization supervises one decoder cross-attention head toward the currently relevant object or receptacle and away from synthetic distractors. Appearancebased visual prompting provides positionless crops of the target object and receptacle, with a jointly trained phase predictor selecting the appropriate prompt during execution. Figure 1 summarizes the resulting ACT-Modified architecture. Let A denote the supervised head’s normalized attention, Lϕ the phase-prediction loss, Mϕ the phase-relevant mask, and Md the distractor mask. Training uses L = LACT + λϕ Lϕ + λ+ (1 − ⟨A, Mϕ ⟩) + λ− ⟨A, Md ⟩. (1) The visual crop specifies target appearance but contains no explicit full-frame coordinates. Simulation crops and masks use privileged segmentation, while hardware prompts are obtained from an initial manual annotation followed by SAM 2 segmentation [11]. Because prompt-bearing policies receive additional target information, augmentation-only and augmentation-plusattention variants are evaluated separately to isolate robustness gains that do not depend on explicit target specification. c) Pretrained-VLA case study.: To examine whether the same grounding perspective remains useful beyond taskspecific ACT policies, we fine-tune π0.5 [12] on 231 teleoperated episodes of a seven-step instrument-handling procedure observed through fixed-base and wrist cameras. Five subgoals are evaluated, including two spatially ambiguous routing subgoals in which the correct destination depends on the observed state of the medical instrument. We compare a prompt-trained checkpoint with and without its expected RGB cue and a regularized checkpoint evaluated on clean RGB input. Five trials per subgoal and condition yield 75 trials. For the routing subgoals, the prompt variants receive the generic destination instruction highlighted, whereas the regularized variant receives explicit left/right destination text; the conditions are therefore not directly comparable.

TABLE I E ND - TO - END SUCCESS UNDER FULL MIXED DISTRACTORS . S IMULATION CONTAINS 600 ROLLOUTS PER CELL ; HARDWARE CONTAINS 20 TRIALS PER CELL . T1 USES A FIXED RECEPTACLE AND T2 A RANDOMIZED RECEPTACLE .

Method

Simulation (%) T1 T2

UR3e (%) T1 T2

Standard ACT Augmentation only ACT-Modified

39.5 100.0 94.5

0.0 65.0

14.0 64.0 88.5

0.0 60.0

IV. E XPERIMENTS AND R ESULTS a) ACT failures are cue-specific and stage-localized.: Standard ACT reaches 98.5% and 99.7% success in clean Tasks 1 and 2, confirming that both manipulation routines are learned. Under three color-matched object competitors in Task 2, P (pick) falls to 39.2%, while P (lift | pick) and P (place | pick, lift) remain 93.8% and 95.5%. In contrast, two shape-matched receptacle competitors leave P (pick) at 97.5% but reduce conditional placement to 33.9%. Within the evaluated assets, object selection is therefore most sensitive to color, whereas receptacle selection is most sensitive to shape. The effect increases with competitor count and compounds when both bottlenecks are active: under full mixed distractors, end-to-end success falls to 39.5% on Task 1 and 14.0% on Task 2. The preservation of downstream conditional success after correct selection indicates that much of the degradation arises from incorrect grounding rather than loss of the learned manipulation routine. b) Object-centric interventions recover robustness.: Table I summarizes the central behavioral results under full mixed distractors. Augmentation alone reaches 100.0% on Task 1 and 64.0% on Task 2 without explicit target prompts, showing that substantial robustness can be recovered by disrupting spurious appearance correlations during training. Its remaining Task 2 deficit is concentrated in randomizedreceptacle localization. ACT-Modified reaches 94.5% and 88.5% on Tasks 1 and 2, respectively. On the physical UR3e, standard and modified ACT remain comparable in clean scenes: 17/20 versus 17/20 successes on Task 1 and 16/20 versus 15/20 on Task 2. Under mixed distractors, standard ACT fails all tested trials, whereas ACTModified succeeds in 13/20 and 12/20 trials. Given the sample size, these experiments establish that the behavioral ordering observed in simulation persists on hardware rather than reproducing the complete mechanism analysis in the physical setting. c) Robust representations require invariance without loss of task geometry.: To examine what changes inside the policy, we compare standard ACT, augmentation-only ACT, and ACTModified at the vision encoder, transformer memory, and decoder state. Matched clean and full-distractor activations are evaluated using cosine shift, pick/place centroid separation, and Task 2 left/right container structure.

...

Transformer Encoder

Transformer Encoder

Transformer Decoder

...

GT Action Chunk

... Predicted Action Chunk Attention Regularization

CLS Phase-Gated Selection

Phase Predictor

Robot Proprioception

Vision Encoder Prompt 1

Copy-Paste Augmentation

...

Learned Query

Learned Query

Vision Encoder

Vision Encoder Prompt 2

Cross-Attention

Full View

Cross-Attention

...

Addition with Positional Embedding

Training & Inference

Only for Training

Fig. 1. ACT-Modified architecture. The full camera observation and phase-selected visual prompt are encoded into the visual memory consumed by the ACT transformer decoder. A phase predictor selects the object or receptacle prompt according to the current manipulation phase, while robot proprioception (qpos) provides state conditioning. The decoder combines visual memory, learned queries, cross-attention, and positional embeddings to predict an action sequence. The action-sequence encoder and CLS latent path form the CVAE posterior during training only; during inference, future actions are unavailable and the latent is fixed to the prior mean z = 0.

TABLE II TASK 2 DECODER - STATE DIAGNOSTICS . S HIFT IS THE PLACE - PHASE CLEAN - TO - FULL COSINE DISTANCE . P HASE RETENTION MEASURES FULL - TO - CLEAN PICK / PLACE CENTROID SEPARATION . G EOMETRY IS THE FULL - DISTRACTOR CONTAINER - POSITION SILHOUETTE . Metric

Standard DataAug Modified

Place shift ↓ Phase retained ↑ Container silhouette ↑ Conditional placement ↑

0.1993 13.1% 0.0095 33.3%

0.0011 94.8% 0.0889 66.3%

0.0125 108.5% 0.2037 100.0%

Standard ACT’s decoder state shifts substantially under distractors during placement and retains only 13.1% of its clean pick/place separation. Both object-centric variants strongly suppress this shift and preserve phase structure. The augmentation-only model provides a critical counterexample to distractor invariance as a sufficient explanation of robustness: it is more invariant than ACT-Modified, yet retains weaker container-position geometry and reaches only 66.3% conditional placement compared with 100.0% for ACT-Modified. Across the three policies, destination-geometry preservation follows placement performance more closely than clean-to-distractor invariance alone. Robust behavior is therefore associated with suppressing nuisance-induced variation while preserving the spatial structure required by the currently relevant referent. The representation analyses remain correlational, and attention maps are likewise treated as diagnostic rather than

causal explanations. We therefore complement them with a direct prompt intervention. In a separate cup-selection probe, replacing the target prompt with a distractor prompt increases distractor selection from 0-1% to 35-91%. Because the visual crop contains target appearance but no explicit full-frame coordinates, this intervention shows that prompt identity can actively redirect closed-loop target selection rather than merely correlate with successful behavior. d) The VLA case study provides preliminary cross-regime evidence.: The prompt-trained π0.5 checkpoint succeeds in 20/25 trials when evaluated without its expected RGB cue and in 25/25 when the cue is restored. The entire deficit is concentrated in the ambiguous routing subgoals: routing success increases from 5/10 to 10/10, while the remaining instrument-handling subgoals are unaffected. The regularized clean-input checkpoint also reaches 25/25 overall and 10/10 on routing. Thus, the observed deficit again localizes to selection of a context-dependent destination rather than to execution of the remaining manipulation sequence. Because the experiment is small and the guided conditions are not equally informed, it provides convergent behavioral evidence for the conditionalgrounding perspective rather than evidence that ACT and π0.5 share the same internal shortcut. V. D ISCUSSION AND O UTLOOK The ACT and VLA studies instantiate conditional visual grounding differently. In ACT, manipulation phase determines whether the object or receptacle is relevant; in the VLA task, the observed instrument state determines the relevant

(a) Standard: clean/full shift

(b) DataAug: clean/full shift

(c) Modified: clean/full shift

(d) Standard: container geometry

(e) DataAug: container geometry

(f) Modified: container geometry

Fig. 2. Robust placement requires both distractor invariance and retained target geometry. The top row compares clean and full-distractor Task 2 place-phase decoder states. The bottom row colors the corresponding representations by the target container’s left/right position, with clean and full conditions shown in a shared within-model projection. Augmentation-only ACT exhibits the strongest clean/full invariance, whereas ACT-Modified preserves substantially clearer container-position structure. The t-SNE projections are illustrative; quantitative comparisons in Table II are computed in the original feature space.

destination. Across both settings, competent motor behavior can coexist with incorrect selection of a context-dependent referent. The contribution is therefore not that both architectures fail identically, but that analyzing what must be grounded when provides a common framework for localizing visual failures and designing targeted interventions. The ACT representation analysis further shows that robustness requires not only suppressing distractor-induced variation but also preserving the task-relevant geometry required for control. Several limitations bound these conclusions. The observed ACT color-shape hierarchy is specific to the evaluated assets and requires validation with fully counterbalanced visual factors. Prompt-bearing ACT variants receive additional target information, although augmentation-only results demonstrate substantial robustness gains without this advantage. Attention and representation analyses remain correlational, the hardware evaluation contains only 20 trials per cell, and the exploratory VLA study contains only five trials per subgoal with nonequivalent guidance conditions. The VLA task additionally uses a visible contamination surrogate and fixed routing rule rather than general contamination understanding. R EFERENCES [1] A. Xie, L. Lee, T. Xiao, and C. Finn, “Decomposing the generalization gap in imitation learning for visual robotic manipulation,” in 2024 IEEE International Conference on Robotics and Automation (ICRA), 2024, pp. 3153–3160. [2] Y. Dong, H. Ge, Y. Zeng, J. Zhang, B. Tian, H. Zhu, Y. Jia, R. Wang, Z. Xue, G. Zhou, L. Ma, and G. Tian, “ImitDiff: Transferring foundationmodel priors for distraction robust visuomotor policy,” arXiv preprint arXiv:2502.09649, 2025. [3] Y. Chen, Y. Zhang, G. D’urso, N. Lawrance, and B. Tidd, “Improving generalization ability of robotic imitation learning by resolving causal confusion in observations,” arXiv preprint arXiv:2507.22380, 2025.

[4] T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” 2023, arXiv:2304.13705. [5] P. de Haan, D. Jayaraman, and S. Levine, “Causal confusion in imitation learning,” in Advances in Neural Information Processing Systems, vol. 32, 2019. [Online]. Available: https://arxiv.org/abs/1905.11979 [6] S. Hattori, H. Sasaki, T. Hachimine, Y. Mizutani, and T. Matsubara, “Task-relevant and irrelevant region-aware augmentation for generalizable vision-based imitation learning in agricultural manipulation,” arXiv preprint arXiv:2603.04845, 2026. [7] H. Huang, X. Chen, Y. Chen, H. Li, X. Han, Z. Wang, T. Wang, J. Pang, and Z. Zhao, “RoboGround: Robotic manipulation with grounded vision-language priors,” arXiv preprint arXiv:2504.21530, 2025. [Online]. Available: https://arxiv.org/abs/2504.21530 [8] M. A. Muttaqien, T. Motoda, R. Hanai, and Y. Domae, “Visual prompting for robotic manipulation with annotation-guided pick-andplace using ACT,” arXiv preprint arXiv:2508.08748, 2025. [Online]. Available: https://arxiv.org/abs/2508.08748 [9] A. J. Hancock, A. Z. Ren, and A. Majumdar, “Run-time observation interventions make vision-language-action models more visually robust,” arXiv preprint arXiv:2410.01971, 2024. [Online]. Available: https://arxiv.org/abs/2410.01971 [10] X. Jia, B. Yang, Z. Ge, X. Nie, Y. Zhou, C. Fan, Y. Li, Y. Chai, C. Jing, Z. Liang, Q. Bu, H. Cao, C. Wu, Q. Li, Z. Yang, C. Zhang, H. Li, Z. Wu, J. Yan, and Y.-G. Jiang, “Guidedvla: Specifying task-relevant factors via plug-and-play action attention specialization,” 2026. [Online]. Available: https://arxiv.org/abs/2605.12369 [11] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer, “Sam 2: Segment anything in images and videos,” 2024. [Online]. Available: https://arxiv.org/abs/2408.00714 [12] K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky, “π0.5 : A vision-language-action model with open-world generalization,” in Proceedings of the 9th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 305. PMLR, 2025, pp. 17–40. [Online]. Available: https://proceedings.mlr.press/v305/black25a.html

Record · ID 660845 · SHA-256 d788b397d591d7eb
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.