ConceptioArchivearXiv CS
arXiv CSopen access

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

1

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Abstract—Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the Dynamics-Aware Meta-Imitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.

Generalization Task Specific

Training on seen data

Multi-Task

arXiv:2607.15880v1 [cs.RO] 17 Jul 2026

Zhenduo Shang, Xiyao Liu, Bohan Li, Xudong Wang, Teng Ren, Lianqing Liu, and Zhi Han

Unseen Task

In-Context

Init

Align

Close Gripper

Press

Action Misinterpretation

Fig. 1. Generalization gap in conventional methods. Standard approaches overfit seen data and fail on unseen tasks. The bottom sequence shows Action Misinterpretation: the robot fails to push the cylinder with an open gripper.

I. I NTRODUCTION Note to Practitioners—Deploying robot manipulators in factories and service environments is often limited by the cost of collecting demonstrations for every new task. This work is motivated by the practical need to reuse experience from previously learned tasks and adapt a robot to a new manipulation task with only a small number of demonstrations. The proposed DAMI framework combines the robot’s current observation, a language instruction, and a complete reference demonstration to identify the intended behavior and adapt the control policy. Tests in simulated benchmarks and on a physical robot show that this approach can improve task success and reduce action misinterpretation compared with representative baseline methods, while requiring less than two minutes for the reported real-world adaptation procedure. The method is most applicable when a complete and relevant reference demonstration, reliable three-dimensional observations, and a small task-specific demonstration set are available. Its present limitations include the need for task-specific fine-tuning, additional computation compared with the baseline, and evaluation on a limited set of manipulation tasks. Future work should reduce adaptation time and validate reliability under broader object, sensing, and operating conditions. Index Terms—Few-shot learning, Imitation learning, Metalearning, Robotic manipulation.

Zhenduo Shang, Xiyao Liu, Xudong Wang, Lianqing Liu, and Zhi Han are with the State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China. Zhenduo Shang and Xudong Wang are also with the University of Chinese Academy of Sciences, Beijing 100049, China. Bohan Li and Teng Ren are with Shenyang University of Technology. Corresponding author: Xiyao Liu (e-mail: [email protected]).

Robotic manipulation aims to enable agents to interact with and change their environments [1], [2]. Imitation Learning (IL) is effective for acquiring complex robotic manipulation skills from demonstrations [3], [4]. However, IL policies learned from limited expert demonstrations can suffer from distribution shift and compounding errors, while collecting additional demonstrations is often costly [5]–[7]. As a result, many methods train specialized policies for individual tasks [8], [9] or use multi-task models that may memorize training patterns [10], [11]. These methods often generalize poorly to out-of-distribution tasks and still fall short of the needs of general-purpose robots [12]. Large-scale human videos [13] and synthesized data [14] can reduce the need for robot demonstrations, but they still require costly training and suffer from domain shifts [15]. Fewshot generalization provides a more data-efficient direction by using prior knowledge to solve new tasks [16], [17]. However, traditional meta-learning methods often struggle to infer the dynamics of novel tasks under large distribution shifts. Recently, In-Context Imitation Learning (ICIL) has emerged as a gradient-free paradigm that uses demonstrations directly as context during inference [18], [19]. Although efficient, these non-parametric methods often rely on visual similarity rather than task dynamics or action semantics, which limits their performance when tasks have different structures or require semantic understanding. Recent advances in 3D policy learning show that explicit 3D representations, including point clouds and fused geometric features, are essential for capturing spatial structure in unstructured

cond

特征空间指标 concat相同来源的特征,计算cos sim 任务语义匹配 时间窗口KL散度 帧之间的权重和时间有关,指标之间的权重先用平均

text

2

Text

CLIP Text Encoder

Positional Embedding

Door open Assembly ……

U2T

VMT

Text Instruction

Action

Time n

Vision Encoder

Add & Norm Multi-Head Self-Attention

...

...

...

Time 1

Feed-Forward Network

...

Complete Demonstration

Add & Norm

Demo Vision Feature

TCFM

Up Sample

Current Vision Feature

Trajectory

Down Sample U-Net

Gaussian Noise

Support Observation

Fig. 2. Overview of the DAMI Architecture. Given the current observation, a language instruction, and a complete reference demonstration, DAMI predicts an action trajectory using a meta-learned 3D diffusion policy.

environments [20]–[22]. Within this area, diffusion policies have emerged as a powerful paradigm for modeling multimodal action distributions and generating complex visuomotor behaviors [23]–[25]. However, their performance often depends on dense demonstrations. When only a few demonstrations are available, directly fine-tuning a high-capacity 3D diffusion policy creates a mismatch between data availability and model capacity, making the policy prone to overfitting support trajectories rather than learning transferable task dynamics [26], [27]. Therefore, combining expressive 3D diffusion policies with data-efficient adaptation remains a key challenge. Meeting this challenge requires extracting task-specific dynamics from sparse demonstrations while preserving transferable knowledge. To this end, we present the Dynamics-Aware Meta-Imitation (DAMI) framework, which couples an expressive 3D diffusion policy with a data-efficient meta-learning objective. Instead of learning each novel task from scratch, DAMI meta-learns a shared initialization [28] that can be rapidly updated from only a few demonstrations. The policy is conditioned on a text instruction, a complete reference demonstration, and the current observation, enabling task-specific behavior to be inferred from sparse data. Specifically, the Visual-Motor Trajectory (VMT) module jointly encodes visual features and motor actions from the demonstration with temporal position information, capturing task dynamics rather than static appearance. The Unpaired Unified Task (U2T) block then uses Transformer-based fusion to align the demonstration and text tokens with the observation–timestep condition at the first U-Net downsampling block. Finally, Task-Conditioned Feature Modulation (TCFM) injects this fused condition into the low-level 3D features of the diffusion U-Net [29], [30] in a task-conditioned manner. Together, these components reduce reliance on dense demonstrations and support rapid adaptation to unseen tasks. Extensive experiments show that DAMI outperforms state-of-the-art baselines, especially on unseen tasks. In summary, our contributions are as follows: We propose the Dynamics-Aware Meta-Imitation (DAMI) framework, which uses meta-learning to build a shared skill space for rapid few-shot adaptation and cross-task generalization. • We introduce the Visual-Motor Trajectory (VMT) module

to capture spatio-temporal dynamics in a compact latent Training Task2 θ  space, helping the policy model task logic. • We design the Unpaired Unified Task (U2T) block with a TCFM mechanism to fuse unstructured multimodal observations and modulate low-level 3D features. • Experiments on Meta-World, RLBench, and real-world tasks demonstrate stronger direct inference on seen tasks and rapid adaptation to unseen tasks. II. R ELATED W ORK A. Few-Shot Imitation Learning Conventional imitation learning often assumes independent and identically distributed data, making policies brittle under distribution shifts and novel tasks. BC-Z [31] and RAM [32] use large-scale data or retrieval for zero-shot transfer, but require high computational cost. OmniManip [33] and Funcanon [34] introduce structured constraints. Alignment-based methods address mismatched execution in one-shot imitation [35]. Recent zero-shot imitation approaches learn viewpoint-invariant objectcentric representations through contrastive alignment, enabling cross-environment deployment with fewer demonstrations [36]. Closed-loop learning-from-demonstration frameworks further evaluate and automatically optimize skill-transfer generalization to unseen tasks [37]. Despite these advances, balancing crossdomain generality with low-level control precision remains challenging. B. 3D Visual-Motor Policies 3D representations provide geometry and viewpoint invariance, improving policy robustness in robotic manipulation [38]. DP3 [24] established a strong 3D diffusion baseline, but highdimensional 3D inputs remain costly to process. Later methods strengthen 3D perception and manipulation through visual pretraining [39], feature fusion [40], active scene perception with joint viewpoint planning and depth completion in cluttered environments [41], or multi-view graph-attention modeling of visual manipulation relationships for robotic grasping [42]. Language- and affordance-aware grasping further uses partaffordance grounding and multimodal large language model reasoning for task-driven selection [43], [44]. Despite better 3D

3

Temporal Downsampling

Obs Feature

Feed-Forward

… Action Sequence

Feed-Forward

Temporal Embeddings

Multi-Head Self-Attention

Demo Feature

Multi-Head Self-Attention

… … …

Fig. 3. The illustration of VMT.

perception and execution, these methods are still mainly taskspecific or in-domain, with limited support for data-efficient adaptation to unseen tasks. C. Multi-task Robotic Manipulation

windows to Bisup or Biqry . At each control step, the policy πθ takes the current observation Ot (a 3D point cloud and robot state), a text instruction ℓi , and τi as inputs, and predicts an H-step action trajectory A0 = [at , . . . , at+H−1 ]. Our goal is to learn a policy initialization that enables strong base-task performance and rapid few-shot adaptation to unseen tasks. B. Diffusion-based Meta-Imitation Formulation DAMI adopts the Diffusion Policy paradigm [23] to model action trajectories. Given the clean action trajectory A0 , the forward diffusion process gradually perturbs it into a noisy trajectory Ak . The reverse process is implemented by a conditional U-Net that predicts the clean trajectory from the noisy sample, diffusion step, current observation feature fobs , demonstration tokens Hdemo , and text feature ftext . Following DP3 [24], we use the action-prediction parameterization:   2 Ldif f = Ek,A0 ,ϵ A0 − Â0,θ (Ak , k, fobs , Hdemo , ftext ) .

Multi-task robotic manipulation requires effective geometric and semantic features from high-dimensional observations. RVT-2 [45] improves precision with a multi-view transformer, (1) and GNFactor [46] uses neural feature fields for semantic To support rapid adaptation, we integrate this diffusion understanding. ManiGaussian [47] models scene dynamics objective with the full Model-Agnostic Meta-Learning (MAML) with dynamic Gaussian Splatting [48], while diffusion-based procedure [28]. For a task Ti , the inner loop adapts the current frameworks [49], [50] and FlowRAM [51] improve trajectory initialization θ on the support set: modeling and efficiency. However, these methods still depend θi′ = θ − α∇θ Ldif f (θ, Bisup ), (2) strongly on training distributions and generalize poorly to unseen tasks, limiting their use when rapid adaptation is needed. where α is the inner learning rate. At each outer step, the task batch B is constructed by enumerating all tasks assigned to III. M ETHOD the current GPU worker. The shared initialization is updated Overview. This section details the core design of DAMI. from the average query-set meta-gradient using AdamW: 1 X As shown in Figure 2, DAMI conditions a 3D diffusion policy ḡt = ∇θ Ldif f (θi′ , Biqry ), on the current observation, a language instruction, and a |B| (3) Ti ∈B complete reference demonstration. The current point cloud θ ← AdamW(θ, ḡt ; ηt ), and state are encoded by a visual encoder, the instruction is encoded by a frozen CLIP text encoder, and the reference where ηt follows a linear-warmup and cosine-decay schedule. demonstration is encoded by the proposed VMT module. The We retain the computation graph through the inner-loop update, diffusion-timestep embedding and current observation feature so the meta-gradient includes the second-order derivative terms are first combined into a U-Net condition. U2T then fuses of full MAML. During evaluation, base tasks are handled by this condition with projected demonstration and text tokens direct inference using the EMA policy obtained during metaat the first downsampling block, while the remaining blocks training, whereas novel tasks are initialized from the same retain the observation–timestep condition. The resulting policy EMA policy, fine-tuned on few-shot demonstrations, and then predicts an action trajectory, while the meta-learning objective evaluated with the adapted policy. encourages the learned initialization to adapt rapidly to novel tasks with only a few demonstrations. C. The VMT Module To distill the intricate spatio-temporal dynamics embedded within expert demonstrations, we propose the Visual-Motor We consider manipulation tasks Ti ∼ p(T ), each associated Trajectory, visually detailed in Figure 3. Purely visual demonwith a task-specific expert dataset DTi containing ten complete stration encoding can over-emphasize static object appearance, expert episodes. At each meta-training step, a batch of 16-frame which is insufficient for distinguishing manipulation intents observation–action windows is sampled from DTi and split such as pushing, pulling, and pressing. In contrast, the action by index into a support batch Bisup for inner-loop adaptation sequence exposes how the scene should evolve under the and a query batch Biqry for the outer-loop update. A complete expert’s control. VMT therefore jointly models visual states 200-frame reference demonstration τi is obtained by randomly and motor commands, allowing the representation to capture selecting one complete episode from the same task dataset. “what changes under what action” rather than relying only Thus, the demonstration, support batch, and query batch all on static visual cues. Formally, we consider a reference originate from DTi ; no episode-level isolation is enforced, demonstration trajectory τ = {(fv,t , at )}Tt=1 , comprising a and the selected demonstration episode may also contribute sequence of high-dimensional visual representations fv,t ∈ RDv A. Problem Formulation

4

Algorithm 1 DAMI Meta-Training Algorithm Require: Policy πθ , inner learning rate α, outer optimizer Ometa , EMA schedule {γt }, per-task datasets {Di } 1: Load the frozen CLIP encoder; initialize θema ← θ 2: repeat 3: Construct B from all tasks assigned to the current GPU worker; set gtotal ← 0 4: for each task Ti ∈ B do 5: Sample a window batch Bi ∼ Di ; split it by index into Bisup and Biqry 6: Sample ei ∼ Uniform{0, . . . , 9}; load the 200-frame trajectory τi ← FullEpisode(Di , ei ) 7: Hsup demo ← VMTθ (VisEncθ (τi .obs), τi .action); ftext ← CLIP(ℓi ) sup 8: fobs ← VisEncθ (Bisup .obs) 9: Set A0 ← Bisup .action; sample ϵ, k; construct Ak ; ek ← TimeEmbθ (k) sup 10: hsup i,k ← MLPθ ([ek ⊕ fobs ]) sup sup sup sup 11: Within U-Net, set cDown1 ← U2Tθ ([Projθ (Hsup demo ) || Projθ (ftext ) || hi,k ]) and cother ← hi,k b sup ← UNetθ (Ak , k; global cond = f sup , fuse feature = [Hsup , ftext ]) 12: A obs demo 0,θ b sup ∥2 13: Lin ← ∥A0 − A 0,θ 2 14: θi′ ← θ − α∇θ Lin ′ ′ 15: Hqry demo ← VMTθi (VisEncθi (τi .obs), τi .action) qry 16: fobs ← VisEncθi′ (Biqry .obs) qry 17: Lout ← Ldif f (θi′ , Biqry , Hqry demo , ftext , fobs ) 18: gtotal ← gtotal + ∇θ Lout 19: end for 20: θ ← Ometa (θ, gtotal /|B|) ▷ AdamW with warmup and cosine decay 21: θema ← γt θema + (1 − γt )θ 22: until the maximum number of epochs is reached 23: return πθema

and proprioceptive actions at ∈ RDa , where T denotes the temporal horizon. Unified Modality Embedding. Since vision and action reside in heterogeneous feature spaces, direct interaction is suboptimal. We first bridge this semantic gap by projecting the concatenated features onto a shared latent manifold via a learnable linear projection ϕproj : et = ϕproj ([fv,t ⊕ at ]),

(4)

may correspond to different states, object poses, or execution phases. In the implemented data flow, the demonstration and text form the explicit fusion tokens, whereas the current observation follows the U-Net global-conditioning branch and is combined with the diffusion timestep before fusion. The frozen CLIP encoder [52] produces the text feature ftext , and VMT produces the demonstration tokens Hdemo . We use modality-specific affine projections to align only these two fusion inputs:

where ⊕ denotes concatenation, and et ∈ RDmodel represents em = Wm · LN(fm ) + bm , m ∈ {demo, text}, (6) the unified token at step t. This results in a raw token sequence E = [e1 , . . . , eT ]. where fdemo = Hdemo and LN(·) denotes LayerNorm. The Spatio-Temporal Representation Learning. To capture current observation feature fobs is not passed through this trajectory-level evolution, we add learnable positional em- modality-projection branch. beddings P, implemented with an embedding layer, to E to Observation–Timestep-Conditioned Fusion. At diffusion retain sequential order. The sequence is then processed by a step k, the timestep embedding ek = TimeEmb(k) is concateTransformer encoder to model global temporal dependencies: nated with the current observation feature and transformed by H = Transformer(E + P). (5) the U-Net condition MLP: demo

The resulting token sequence Hdemo serves as a motioncentric context, preserving the fine-grained dynamics required for downstream policy conditioning. D. The U2T Module

hk = MLP([ek ⊕ fobs ]).

(7)

Distinct learnable positional embeddings preserve the structure and source of the demonstration and text tokens. The U2T input and the resulting first-downsampling-block condition are

Zin = [edemo + Pdemo || etext + Ptext || hk ], (8) Cross-Modal Semantic Alignment. To harmonize an unCDown1 = U2T(Zin ) = TransformerEncoder(Zin ), structured text instruction, a structured reference demonstration, and the current observation, we propose the Unpaired Unified where || denotes sequence concatenation. Thus, U2T performs Task encoder. Here, “unpaired” indicates that the reference Transformer-based fusion of the demonstration and text tokens demonstration and current observation need not be temporally with the observation–timestep condition specifically for the aligned or frame-wise matched; they share task semantics but first downsampling block.

5

Algorithm 2 Novel-Task Few-Shot Adaptation and Evaluation Algorithm Require: Pre-trained EMA policy πθema , target-task expert dataset D, text instruction ℓ, steps N , learning rate α, weight decay λ, clipping threshold c, EMA schedule {γi } 1: Initialization 2: πϕ ← Copy(πθema ); ϕema ← ϕ 3: Fit observation and action normalizers on D 4: ftext ← CLIP(ℓ) 5: Adaptation 6: for i = 1 to N do 7: Sample a complete reference trajectory d ∼ D and a batch B of observation-action windows from D 8: Hdemo ← VMTϕ (VisEncϕ (d.obs), d.action) 9: fobs ← VisEncϕ (B.obs) 10: Set A0 ← B.action; sample ϵ, k; construct Ak ; ek ← TimeEmbϕ (k) 11: hk ← MLPϕ ([ek ⊕ fobs ]) 12: Within U-Net, set cDown1 ← U2Tϕ ([Projϕ (Hdemo ) || Projϕ (ftext ) || hk ]) and cother ← hk b 0,ϕ ← UNetϕ (Ak , k; global cond = fobs , fuse feature = [Hdemo , ftext ]) 13: A b 0,ϕ ∥2 14: L ← ∥A0 − A 2 15: ϕ ← AdamW(ϕ, Clip(∇ϕ L, c); α, λ) 16: ϕema ← γi ϕema + (1 − γi )ϕ 17: end for 18: Evaluation 19: For each Meta-World episode, sample deval ∼ D and hold it fixed throughout the rollout 20: Evaluate RLBench by sampling one deval ∼ D and reusing it for all episodes 21: return evaluation results

A DataLoader samples 16-frame observation–action windows We implement the denoising network as a modified 1D from this dataset, and each batch is split by index into support conditional U-Net with residual convolutional blocks. The and query subsets. The reference demonstration is sampled current observation feature fobs is supplied as the U-Net global from the same task dataset by randomly selecting one complete condition, while [Hdemo , ftext ] is supplied as the fusion feature. episode and loading all 200 frames. The code does not enforce Since early downsampling layers perform geometric grounding episode-level separation, so the selected demonstration episode and local contact reasoning, we apply U2T fusion only at the may also contribute support or query windows. VMT encodes first downsampling block. The remaining blocks use the lighter the demonstration, the frozen CLIP encoder processes the text instruction, and the visual encoder extracts the current observation–timestep condition hk from Equation (7). Let hj denote the feature at the j-th U-Net block. We observation feature. The observation and timestep embeddings are transformed into hk ; U2T fuses hk with the projected modulate it using FiLM: demonstration and text tokens only at the first downsampling e j = γ (cj ) ⊙ hj + β (cj ), h (9) block, while the other blocks use hk . The inner loop performs j j one action-prediction gradient update on all trainable policy where γ j and β j are learned projections that produce the parameters. The complete query forward pass, including VMT, modulation coefficients. The layer-specific condition is the visual encoder, U2T, and the U-Net, is recomputed using θi′ ; ( only the frozen CLIP feature is reused. Full-MAML gradients CDown1 , j = Down1 , cj = (10) are averaged across the task batch and applied with AdamW hk , otherwise, under a linear-warmup and cosine-decay learning-rate schedule. An EMA copy of the policy is maintained with a step-dependent Here, ek is the diffusion timestep embedding, not the sampled decay schedule. diffusion noise ϵ. Consequently, every U-Net block remains conditioned on both the current observation and diffusion Novel-task Adaptation and Evaluation. The few-shot timestep, while only the first downsampling block performs adaptation procedure is summarized in Algorithm 2. For every Transformer fusion with the demonstration and text features. evaluated task, the runner receives that task’s own expert dataset D. For a novel task, the same D supplies the fine-tuning F. Training and Adaptation Procedures windows, full reference demonstrations, and normalization Meta-training. The complete procedure is summarized in statistics for the point cloud, agent state, and action. Each Algorithm 1. At each outer step, every GPU worker constructs adaptation iteration samples an observation–action batch and a its task batch by enumerating all tasks assigned to that worker complete reference trajectory from D, rebuilds the observation– rather than randomly drawing tasks from p(T ). Each task has timestep and U2T conditions, and updates the policy using one expert dataset Di containing ten complete expert episodes. AdamW with weight decay and gradient clipping. A scheduled E. Task-Conditioned Feature Modulation

6

TABLE I K EY H YPERPARAMETERS U SED IN O UR E XPERIMENTS . Training and Data

Optimization

Hyperparameter

Value

Hyperparameter

Training Epochs Episodes per Epoch Prediction Horizon Action Pred. Step Observation Step Diff. Training Steps Diff. Inference Steps Support / Episodes per Task Query Set Size Total Batch Size

3,000 18 16 8 2 100 10 20 / 10 88 108

Outer Optimizer Outer Learning Rate Optimizer Betas Weight Decay Inner Optimizer Inner Learning Rate Inner Loop Steps LR Scheduler EMA Power / Max Decay FT Diffusion LR

Architecture and Adaptation Value

Hyperparameter

AdamW 1.0 × 10−4 [0.95, 0.999] 1.0 × 10−6 SGD 1.0 × 10−4 1 Warmup + Cosine 0.75 / 0.9999 3.0 × 10−5

Value

Point Feature Dim Action Dim VMT Attention Heads VMT Layers U2T Attention Heads U2T Pos. Emb. (Sim) U2T Pos. Emb. (Real) Diff. Step Embed Dim FT Vision LR FT Other LR

64 128 4 2 8 200 500 128 1.0 × 10−5 1.0 × 10−4

TABLE II S UCCESS R ATES (%) ON BASE M ETA -W ORLD TASKS . B OLD AND U NDERLINED I NDICATE THE B EST AND S ECOND -B EST R ESULTS . ML10 Tasks

ML45 Tasks

Average

Alg. / Tasks

E(5)

M(3)

H(2)

E(26)

M(9)

H(5)

VH(5)

ML10

ML45

Meta-World DP3 Mamba FreqPolicy FlowPolicy

74.40 87.07 85.07 84.33 83.73

1.33 80.11 76.22 82.11 74.44

10.00 72.83 33.83 69.83 69.17

13.23 59.47 64.18 57.88 61.21

1.41 34.93 36.48 32.63 28.81

0.00 23.60 32.80 26.00 15.80

0.00 29.40 25.87 27.07 26.73

39.60 82.13 72.17 80.77 78.03

7.93 47.24 50.90 45.87 45.85

Ours(DAMI)

90.27

91.78

72.83

87.19

71.04

61.60

68.60

87.23

79.05

TABLE III S UCCESS R ATES (%) ON N OVEL M ETA -W ORLD TASKS . B OLD AND U NDERLINED I NDICATE THE B EST AND S ECOND -B EST R ESULTS . ML10 Tasks

ML45 Tasks

Average

Alg. / Tasks

Door Close

Drawer Open

Lever Pull

Sweep Into

Shelf Place

Door Lock

Door Unlock

Bin Picking

Box Close

Hand Insert

ML10

ML45

Meta-World DP3 Mamba FreqPolicy FlowPolicy

94.00 100.00 100.00 93.00 100.00

22.67 70.33 32.33 98.00 13.33

5.33 44.67 27.67 18.67 17.67

24.67 16.67 9.67 15.67 6.67

0.00 12.67 10.00 7.00 0.00

42.67 51.00 22.67 11.33 15.00

7.67 94.00 34.67 6.00 22.00

0.33 9.33 9.67 11.00 9.00

1.67 26.67 12.67 25.33 13.33

10.33 6.67 5.33 13.67 6.00

29.33 48.87 35.93 46.47 27.53

12.53 37.53 17.00 13.47 13.07

Ours(DAMI)

100.00

100.00

45.33

28.00

10.67

39.33

98.67

13.67

22.33

13.67

56.80

37.53

TABLE IV S UCCESS R ATES (%) ON N OVEL RLB ENCH FS25 N OVEL TASKS . B OLD AND U NDERLINED I NDICATE THE B EST AND S ECOND -B EST R ESULTS . Task

DP3

Mamba

FreqPolicy

FlowPolicy

Ours(DAMI)

Avg.

9.93

0.00

1.20

0.20

10.80

EMA copy of the adapted policy is maintained to smooth the updates. For Meta-World, one complete demonstration is randomly selected from the current task’s ten episodes at the start of each rollout, held fixed throughout that rollout, and resampled for the next rollout. Base-task evaluation uses the same task dataset that was used during training. RLBench instead samples one demonstration and reuses it across its evaluation episodes. IV. E XPERIMENT A. Experiment Setup Simulations. In our simulation experiments, we adopt the standard ML10 and ML45 protocols from Meta-World [53]

and the FS25 setting from RLBench [54] to systematically evaluate the few-shot adaptation and generalization capabilities of our method. These benchmarks comprise a broad spectrum of robotic skills ranging from simple object interactions to complex articulations and tool use. Task Splits. The benchmarks are structured as follows: • Metaworld ML10: This suite is designed to test rapid adaptation to novel tasks with limited data. It consists of 10 training tasks and 5 held-out testing tasks. • Metaworld ML45: This larger-scale setting contains 45 training tasks and 5 held-out testing tasks, which evaluate the ability of the model to generalize from a diverse set of behaviors. • RLBench FS25: This setting trains policies on 25 base tasks and evaluates adaptation on 5 held-out tasks after few-shot fine-tuning. Expert Demonstration. We collect expert trajectories using RL agents [55], [56], achieving 100% success on most tasks. Crucially, we maintain strict data consistency across all evaluation methods. This standardization eliminates discrepancies in

7

TABLE V D ETAILED S UCCESS R ATES (%) ON S ELECTED R EPRESENTATIVE BASE S IMULATION TASKS FROM THE ML10 B ENCHMARK . W E R EPORT R ESULTS ON A S UBSET OF TASKS : 5 E ASY, 3 M EDIUM , AND 2 H ARD TASKS , AVERAGED OVER 3 E VALUATION S EEDS .

Meta-World (Easy)

Meta-World (Medium)

Meta-World (Hard)

Alg \ Task

Door Button Press Drawer Window Peg Insert Pick Reach Basketball Sweep Open Topdown Close Open Side Place

Push

Meta-World DP3 MambaPolicy FreqPolicy FlowPolicy DAMI

74.00 82.33 86.33 96.67 91.33 92.67

16.00 77.00 67.67 64.67 77.33 75.33

Door Open

Peg Insert Side

85.00 100.00 100.00 89.00 100.00 100.00

98.33 100.00 100.00 100.00 100.00 100.00

Sweep Into

16.00 54.67 39.00 42.33 27.33 58.67

98.67 98.33 100.00 93.67 100.00 100.00

Drawer Open

Peg Fall

Object Fall

Task Confusion

Success

Success

Success

Success

DAMI

DP3

Error Pos

Fig. 4. Qualitative comparison on Meta-World. Execution frames of DP3 (top) and DAMI (bottom) on base tasks (left) and novel-task adaptation (right).

0.33 100.00 100.00 100.00 96.67 100.00

0.00 70.33 62.33 65.33 69.33 88.00

3.67 70.00 66.33 81.00 57.33 87.33

4.00 68.67 0.00 75.00 61.00 70.33

generation, and LayerNorm to enhance training stability. We utilize the pre-trained CLIP model with a ViT-B/32 backbone [52] as the text encoder. The model weights are kept frozen during training to preserve the learned semantic alignment. Point clouds are downsampled using Farthest Point Sampling (FPS). Specifically, we retain 512 points for simulation tasks and 1024 points for real-world experiments to accommodate the varying complexity of environmental details. For the metalearning setup, we train the model for 3000 epochs, configuring a support set size of 20 and a query set size of 88 per task. The inner loop performs a single gradient update step for each support batch to enable rapid adaptation. A comprehensive list of hyperparameters is provided in Table I. Key Hardware and Software. All methods are implemented within the PyTorch framework. Experiments are conducted on a computational node equipped with the Intel Xeon Platinum 8358 CPU and the NVIDIA L40S GPU (48GB).

data quality or quantity, ensuring that performance differences are solely attributable to the policy architectures. Baselines. We benchmark DAMI against a diverse suite of state-of-the-art methods to assess both adaptation and multi-task proficiency. Our comparison includes DP3, a representative 3D diffusion policy, alongside cutting-edge architectures such as Mamba Policy [57], FlowPolicy [55], and FreqPolicy [56]. We extend these originally single-task baselines to the multi-task B. Quantitative Comparison setting via joint training on the aggregated dataset. This rigorous We benchmark DAMI against prior state-of-the-art methods protocol allows us to examine whether advanced backbones can on the Meta-World and RLBench simulation suites. Following inherently absorb complex multi-task distributions via naive the protocol established in DP3 [24], we stratify the Metadata aggregation, or if explicit meta-adaptation mechanisms World tasks within ML10 and ML45 into Easy (E), Medium remain indispensable. Finally, the standard Meta-World RL (M), Hard (H), and Very Hard (VH) levels. In addition, we baseline [53] is included for completeness. evaluate cross-task adaptation on RLBench under the FS25 Evaluation Protocols. To ensure statistical reliability, all setting. Accordingly, we report the quantitative results grouped simulation experiments are conducted across three seeds, by these benchmark-specific evaluation protocols. numbered 0, 1, and 2. Adhering to protocols [24], [40], [57], 1) Novel Tasks Evaluation: As presented in Table III and we evaluate imitation learning algorithms every 200 epochs, Table IV, DAMI ranks first or joint first across all three novelperforming 20 evaluation episodes per task at each checkpoint. task evaluation settings. It achieves an average success rate of For the Meta-World RL baseline, evaluations are performed 56.80% on ML10, outperforming all baselines, and reaches every 20 epochs. We employ distinct strategies based on task 37.53% on ML45, tying DP3 for the highest average success set. Direct inference is used for base tasks, while few-shot fine- rate. Notably, this joint-best ML45 adaptation performance tuning for 200 steps is applied to novel tasks before inference, is achieved while DAMI attains 79.05% on ML45 base following Algorithm 2. For each seed, we calculate the average tasks, substantially exceeding DP3’s 47.24% (Table II). This success rate of the top-5 checkpoints. We then report the mean 31.81-percentage-point improvement demonstrates that DAMI of these per-seed averages across all three seeds. preserves substantially stronger base-task proficiency without Implementation Details. We standardize the visual encoder sacrificing novel-task adaptation. On RLBench FS25, DAMI architecture across all methods. This encoder comprises a further obtains the highest average success rate of 10.80%. three-layer MLP followed by a max-pooling layer for fea- Overall, these results highlight DAMI’s stronger balance ture aggregation, a projection head for compact 3D feature between base-task proficiency and cross-task adaptation.

8

Close Drawer Object Place

D435i

Press Button

Pull Ball

Sweep Cylinder Push Cube

Stack Cube

Init

Init

Init

Init

Init

Init

Init

Move Gripper

Move Gripper

Move Gripper

Align

Align

Close Gripper

Align

Start

Grasp

Align

Touch

Touch

Align

Grasp

In Process

Move

Close Gripper

In Process

In Process

Pull

Move

End

End

End

End

End

End

End

UR10e

AG95

Fig. 5. Real-world setup and tasks. The left panel shows the UR10e manipulator with an AG95 gripper, while the right panels show representative execution frames for training and unseen tasks. TABLE VI D ETAILED S UCCESS R ATES (%) ON A LL 45 BASE S IMULATION TASKS FROM THE ML45 B ENCHMARK . T HE TASKS A RE C ATEGORIZED BY D IFFICULTY L EVELS : 26 E ASY, 9 M EDIUM , 5 H ARD , AND 5 V ERY H ARD TASKS . W E R EPORT THE AVERAGE S UCCESS R ATE OVER 3 E VALUATION S EEDS . T HESE R ESULTS C ORRESPOND TO THE S UMMARIZED P ERFORMANCE IN TABLE II.

Alg \ Task

Button Press Topdown

Button Press Topdown Wall

Button Press

Button Press Wall

Meta-World DP3 Mamba FreqPolicy FlowPolicy DAMI

0.00 91.00 100.00 73.33 99.67 100.00

0.00 30.00 73.33 18.33 35.67 100.00

0.00 51.67 50.33 57.67 47.00 100.00

6.00 100.00 100.00 100.00 100.00 100.00

Alg \ Task

Handle Press

Handle Pull Side

Handle Pull

Lever Pull

Peg Unplug Side

Meta-World DP3 Mamba FreqPolicy FlowPolicy DAMI

95.33 86.00 92.33 74.00 92.33 86.67

15.33 26.33 22.33 35.00 5.00 37.33

2.00 2.33 5.33 5.33 8.67 16.00

0.00 8.00 16.67 36.67 41.67 75.00

0.67 24.00 30.00 13.67 18.67 89.00

Meta-World (Easy) - Part I Coffee Dial Door Door Button Turn Close Open

Drawer Close

Drawer Open

Faucet Close

Faucet Open

Handle Press Side

25.33 96.67 92.33 88.00 100.00 100.00

94.67 100.00 100.00 100.00 100.00 100.00

0.00 95.33 98.67 88.00 93.00 100.00

0.00 2.33 4.00 7.33 2.67 100.00

0.67 74.33 100.00 76.00 99.67 100.00

73.33 35.00 55.67 21.33 26.00 66.67

27.33 18.67 22.00 23.33 21.33 73.67

1.33 100.00 100.00 100.00 100.00 100.00

Meta-World (Easy) - Part II Plate Slide Plate Slide Plate Slide Back Side Back Side 0.00 100.00 100.00 100.00 100.00 100.00

Alg \ Task

Basketball

Coffee Pull

Coffee Push

Hammer

Meta-World DP3 Mamba FreqPolicy FlowPolicy DAMI

0.00 96.00 100.00 95.33 69.00 97.33

0.00 0.00 0.00 0.00 0.00 82.33

8.00 65.67 58.67 68.67 59.33 92.67

0.00 30.00 27.67 16.00 15.33 92.00

2.00 100.00 100.00 100.00 97.33 100.00

Meta-World (Medium) Peg Insert Side 0.00 2.33 3.00 5.33 2.67 59.67

Meta-World (Hard) Alg \ Task

Assembly

Pick Out Of Hole

Meta-World DP3 Mamba FreqPolicy FlowPolicy DAMI

0.00 34.67 39.00 42.67 24.00 88.00

0.00 28.67 36.33 40.00 8.33 31.67

0.00 100.00 99.33 90.33 100.00 95.00

0.00 0.00 0.00 0.00 0.00 99.33

Plate Slide

Reach

Reach Wall

Window Close

Window Open

0.00 99.67 98.67 98.67 92.00 98.33

0.00 41.33 33.33 35.00 39.67 63.00

0.00 38.67 41.33 53.67 40.67 68.67

0.00 98.00 100.00 100.00 100.00 100.00

0.00 27.00 33.00 9.33 30.33 98.33

Push Wall

Soccer

Sweep Into

Sweep

0.00 47.67 56.67 41.00 43.00 80.00

3.33 18.33 13.00 13.33 7.33 28.33

1.33 15.00 23.33 18.33 17.33 15.00

0.00 39.33 46.00 35.67 45.33 92.00

Meta-World (Very Hard)

Pick Place

Push Back

Push

0.00 0.00 36.33 2.67 0.00 60.33

0.00 46.67 46.33 38.33 43.00 44.67

0.00 8.00 6.00 6.33 3.67 83.33

2) Base Tasks Evaluation: Complementing our analysis on novel tasks, we evaluate the performance on base tasks. Table II summarizes the results by difficulty category, while Table V reports 10 representative ML10 base tasks (5 Easy, 3 Medium,

Avg.

Disassemble

Pick Place Wall

Shelf Place

Stick Pull

Stick Push

0.00 28.00 4.67 14.33 14.00 65.00

0.00 22.00 21.67 27.00 20.00 86.33

0.00 5.67 4.67 3.67 10.67 23.67

0.00 16.67 28.33 13.67 16.33 69.33

0.00 74.67 70.00 76.67 72.67 98.67

7.93 47.24 50.90 45.87 45.85 79.05

and 2 Hard) and Table VI reports all 45 ML45 base tasks. All per-task success rates are averaged over 3 random seeds. DAMI achieves average success rates of 87.23% on ML10 and 79.05% on ML45. These results significantly surpass all

9

Config 2

Config 3

Config 4

Config 5

Config 6

Config 7

Config 8

Config 9

Config 10

Sweep Cylinder Push Cube Stack Cube

Pull Ball

Press Button Object Place Close Drawer

Config 1

Fig. 6. Initial environmental configurations for the seven real-world manipulation tasks used in the evaluation. The figure illustrates the starting states for the five base tasks and two novel tasks. All object and target positions correspond to the initialization protocols during the testing phase.

baselines, showing that DAMI retains strong proficiency on base tasks. C. Qualitative Comparison Figure 4 illustrates qualitative results across representative base and novel tasks. On base tasks, DP3 frequently fails to achieve task completion. While it appears to replicate expert trajectories, it suffers from execution instability. For instance, in the Door Open task, the gripper prematurely loses contact with the handle, leading to failure. In contrast, DAMI demonstrates superior physical robustness. Regarding novel tasks, DP3 exhibits limited adaptation capabilities after fine-tuning. This is evident in the Sweep Into task, where performance remains suboptimal. Furthermore, DP3 suffers from Action Misinterpretation, erroneously executing actions consistent with Door Open when presented with the Drawer Open task. DAMI adapts effectively and yields robust performance. D. Real-world Experiments We conduct experiments on a UR10e manipulator equipped with an AG95 gripper and a RealSense D435i camera, adopting the DP3 visual format [24]. We design seven tasks to evaluate, divided into a base training set and a novel adaptation set. Each task collects 10 demonstrations via teleoperation. The

real-world experimental setup and representative executions of the designed tasks are shown in Figure 5. The five base tasks are defined as follows: • Close Drawer: The robot closes the top and middle drawers of a three-tier drawer unit, whose position is kept uniform across ten experimental trials. • Object Place: The robot grasps a plastic fruit and deposits it into a basket, with the initial fruit position distributed uniformly across experiments. •

Press Button: The robot presses a tabletop button whose successful activation produces an auditory signal; the button is placed at uniform positions throughout the experiments.

Pull Ball: The robot moves a ball into a designated goal whose position remains uniform, while the initial ball position is randomized relative to the goal.

Sweep Cylinder: The robot pushes a cylinder past a yellow marker line. Training uses only a standard green cylinder, whereas testing introduces three types of irregular cylinders to assess generalization; the cylinder is placed uniformly in all experiments. The two novel tasks are defined as follows: • Push Cube: A small cube rests on a larger cube and must be pushed off. The adaptation data contain only a purple •

10

TABLE VII R EAL -W ORLD TASK S UCCESS R ATES (%). Type

Task

DP3

DAMI (Ours)

Base

Close Drawer Object Place Press Button Pull Ball Sweep Cylinder

90 10 70 100 10

90 50 80 100 100

Novel

Push Cube Stack Cube

60 50

80 80

55.71

82.86

Average

Standard cylinder

Irregular cylinder 1 Irregular cylinder 2 Irregular cylinder 3

Test configs: 3

Test configs: 2

Test configs: 2

Test configs: 3

Init

Move Gripper

Close Gripper

Press Down

Init

Move Gripper

Close Gripper

Press Down

cube, whereas testing additionally introduces a smaller 2 × 2 Rubik’s cube. Stack Cube: The robot picks up a small cube from the Fig. 7. Generalization on Sweep Cylinder. The top row shows the standard table and stacks it onto a larger cube. The supporting large training cylinder and irregular unseen variants. The bottom rows illustrate cube is positioned uniformly, while the initial position of DP3’s action misinterpretation failure. “Test config” denotes trials per cylinder. the small cube is randomized across experiments. The initial test configurations for all seven tasks are shown an 80% success rate, outperforming DP3 at 60%. This result in Figure 6. To ensure a fair comparison, we use identical demonstrates that DAMI adapts more reliably to novel tasks and sequences of initial environmental configurations for DAMI generalizes more effectively to variations in object appearance. and all baseline methods across trials. Comprehensive video recordings provided with the supplementary material document the complete executions of both DAMI and the baseline E. Ablation Study methods on all evaluation tasks. To isolate the effect of meta-learning, we additionally train For all real-world evaluations, we utilize the model check- MAML-enhanced versions of DP3, Mamba, FreqPolicy, and point obtained from the final training epoch. We fine-tune on FlowPolicy. Table VIII reports category-level performance on novel tasks for 500 steps to assess adaptability. the ML10 base tasks and per-task performance on its five novel Quantitative Results. We evaluate model performance based tasks. Compared with their conventionally trained counterparts, on task completion criteria. We conduct 10 trials for each MAML improves the average novel-task success rate of all four task, with environmental configurations uniformly distributed. baselines, confirming the benefit of meta-learning for rapid Table VII reports results on both base and novel tasks. On adaptation. DAMI nevertheless achieves the highest overall base tasks, DAMI matches DP3 on Close Drawer and Pull averages on both base tasks (87.23%) and novel tasks (56.80%). Ball, while improving the success rate from 10% to 50% on On the base split, DAMI performs best on the Easy and Object Place and from 10% to 100% on Sweep Cylinder. On Medium categories and remains competitive with the strongest novel tasks, DAMI achieves an 80% success rate on both Push method on the Hard category. On novel tasks, DAMI attains Cube and Stack Cube, compared with 60% and 50% for DP3, the best or tied-best results on Door Close, Drawer Open, respectively. Overall, DAMI improves the average success rate Sweep Into, and Shelf Place, while Mamba performs best on from 55.71% to 82.86%. Lever Pull. These results demonstrate that DAMI’s gains cannot Analysis. To evaluate out-of-distribution generalization, be attributed to meta-learning alone, but also arise from its we test Sweep Cylinder with irregular objects, including task-aware representation and conditioning mechanisms. beverage and cosmetic bottles, that differ substantially from We further conduct a component-level ablation study on the standard cylinder used during training. As shown in the novel tasks of ML10 to evaluate the effectiveness of our Figure 7, DAMI maintains a 100% success rate across the proposed modules, as summarized in Table IX. For the baseline tested object configurations, whereas DP3 fails in the majority configuration, we remove all proposed modules and substitute of trials. The qualitative results further reveal an Action the VMT with a standard MLP to aggregate demonstration Misinterpretation failure in DP3: it incorrectly closes the information. The empirical results indicate that all modules gripper and applies downward force, exhibiting behavior contribute positively to the overall performance. Notably, the associated with Press Button rather than the intended sweeping integration of TCFM yields the optimal performance when motion. In contrast, DAMI preserves the correct task semantics applied to the down-sampling layer. and robustly executes the sweeping behavior across different object geometries. We additionally analyze object-level generalization on Push F. Computational Efficiency Analysis Cube. As shown in Figure 8, the fine-tuning data contain only Beyond task success rates, we rigorously evaluate the a purple cube, whereas evaluation additionally includes an temporal cost of the adaptation process to assess real-world unseen 2 × 2 Rubik’s cube. The figure also illustrates DP3’s deployability. We measure the fine-tuning time once for action misinterpretation during evaluation. DAMI achieves each of the 45 tasks and report the average across tasks. •

11

TABLE VIII MAML A BLATION ON M ETA -W ORLD ML10. S UCCESS R ATES (%) C OMPARE MAML-E NHANCED BASELINES WITH DAMI ON BASE -TASK C ATEGORIES AND I NDIVIDUAL N OVEL TASKS . B OLD AND U NDERLINED I NDICATE THE B EST AND S ECOND -B EST R ESULTS , R ESPECTIVELY. Base Tasks

Novel Tasks

Alg. / Tasks

E(5)

M(3)

H(2)

Avg.

Easy Door Close

DP3 Mamba FreqPolicy FlowPolicy

86.13 87.33 85.20 84.53

80.22 79.11 78.22 77.56

73.00 33.33 70.33 62.17

81.73 74.07 80.13 77.97

100.00 100.00 100.00 97.00

100.00 100.00 99.33 55.00

46.67 55.33 35.67 32.67

20.33 12.67 18.67 12.00

8.33 3.67 3.67 0.00

55.07 54.33 51.47 39.33

Ours (DAMI)

90.27

91.78

72.83

87.23

100.00

100.00

45.33

28.00

10.67

56.80

TABLE IX C OMPONENT A BLATION ON ML10 N OVEL TASKS . S UCCESS R ATES (%) A RE R EPORTED FOR VMT AND TCFM VARIANTS . L1–L3 D ENOTE TCFM IN THE D OWN -S AMPLING , M IDDLE , AND U P -S AMPLING U-N ET L AYERS . Method

Door Close

Drawer Open

Lever Pull

Sweep Into

Shelf Place

Avg.

Baseline + VMT + TCFM (L1) + TCFM (L2) + TCFM (L3)

100.00 100.00 100.00 100.00 100.00

100.00 100.00 100.00 100.00 100.00

40.00 53.00 50.00 41.00 41.00

31.00 26.00 36.00 28.00 31.00

12.00 10.00 10.00 10.00 7.00

56.60 57.80 59.20 55.80 55.80

Init

Purple Cube

Rubik’s Cube

Test configs: 3

Test configs: 7

Close Gripper

Move

Open Gripper

Fig. 8. Generalization on Push Cube. The top row compares the purple cube used for fine-tuning with the unseen Rubik’s cube used for evaluation. The bottom row shows DP3’s action misinterpretation during evaluation. “Test config” denotes trials per object.

In the 200 update steps simulation, DAMI requires 39.27 seconds compared to 32.18 seconds for the DP3 baseline. Similarly, in the 500 update steps real-world setting, our method incurs a slight overhead, requiring 74.73 seconds, while 60.56 seconds for DP3. This marginal increase in latency is primarily attributed to the additional multi-modal fusion operations within the U2T and VMT modules. However, considering the substantial gains in generalization capability and semantic consistency, this computational cost is negligible. V. C ONCLUSION In this work, we present the Dynamics-Aware MetaImitation (DAMI) framework to address the critical challenges of data scarcity and limited generalization in robotic imitation learning. It establishes a robust initialization that facilitates rapid few-shot adaptation to unseen tasks. Central to our approach is the integration of the Visual-Motor Trajectory (VMT) module, which effectively captures complex spatiotemporal dynamics, enabling the agent to learn intrinsic task logic. Furthermore, the proposed Unpaired Unified Task (U2T) block, coupled with TCFM, achieves dynamic modulation

Easy Drawer Open

Easy Lever Pull

Medium Sweep Into

Very Hard Shelf Place

Avg.

of low-level 3D features by fusing unstructured multimodal observations. Extensive experiments on the Meta-World and RLBench benchmarks and real-world manipulation tasks demonstrate that DAMI significantly outperforms state-ofthe-art baselines in both direct inference on seen tasks and adaptation to unseen tasks. ACKNOWLEDGMENTS This work is supported in part by the National Natural Science Foundation of China under Grant 62303447, Grant U23A20343, T2596045, in part by the Fundamental research project of SIA, No. 2023JC1K01 and No. 2023JC3K05 and in part by the Natural Science Foundation of Liaoning Province, No. 2024-MSBA-81. R EFERENCES [1] Y. Ji, H. Tan, J. Shi, X. Hao, Y. Zhang, H. Zhang, P. Wang, M. Zhao, Y. Mu, P. An, et al., “Robobrain: A unified brain model for robotic manipulation from abstract to concrete,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 1724–1734. [2] Z. Li, A. Chapin, E. Xiang, R. Yang, B. Machado, N. Lei, E. Dellandrea, D. Huang, and L. Chen, “Robotic manipulation via imitation learning: Taxonomy, evolution, benchmark, and challenges,” arXiv preprint arXiv:2508.17449, 2025. [Online]. Available: https: //arxiv.org/abs/2508.17449 [3] Z. Lin, Y. Chen, Z. Li, X. Zhang, B. Liang, and Z. Liu, “Adaptive video-conditioned imitation learning via bidirectional cross-domain skill transfer,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 23 214–23 227, 2025. [4] H. Xu, Y. Chen, D. Yu, Y. Ren, and J. Pan, “Bikc+: Bimanual hierarchical imitation with keypose-conditioned coordination-aware consistency policies,” IEEE Transactions on Automation Science and Engineering, vol. 23, pp. 1064–1079, 2025. [5] Z. Lin, Z. Chen, G. Zhu, L. Wang, and J. Li, “Toward reliable imitation learning with limited expert demonstrations via search-based inverse dynamic learning,” IEEE Transactions on Automation Science and Engineering, 2026. [6] H. Wang, Z. Dong, T. Zhu, H. Lei, W. Shi, Z. Zhang, W. Luo, W. Wan, X. Chen, and J. Huang, “Robot deformable object manipulation via nmpc-generated demonstrations in deep reinforcement learning,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 23 566– 23 578, 2025. [7] Z. Jiang, Y. Xie, K. Lin, Z. Xu, W. Wan, A. Mandlekar, L. J. Fan, and Y. Zhu, “Dexmimicgen: Automated data generation for bimanual dexterous manipulation via imitation learning,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 16 923–16 930. [8] S. Xia, H. Fang, C. Lu, and H.-S. Fang, “Cage: Causal attention enables data-efficient generalizable robotic manipulation,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 13 242–13 249. [9] Y. Su, N. Liu, D. Chen, Z. Zhao, K. Wu, M. Li, Z. Xu, Z. Che, and J. Tang, “Freqpolicy: Efficient flow-based visuomotor policy via frequency consistency,” Advances in Neural Information Processing Systems, vol. 38, pp. 27 769–27 797, 2026.

12

[10] A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox, “Rvt: Robotic view transformer for 3d object manipulation,” in Conference on Robot Learning. PMLR, 2023, pp. 694–710. [11] T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki, “Act3d: 3d feature field transformers for multi-task robotic manipulation,” in Proceedings of The 7th Conference on Robot Learning. PMLR, 2023, pp. 3949–3965. [12] Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. Tan, L. Chen, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” in Proceedings of Robotics: Science and Systems, 2024. [13] Z. Chen, S. Chen, E. Arlaud, I. Laptev, and C. Schmid, “Vividex: Learning vision-based dexterous manipulation from human videos,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 3336–3343. [14] J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y. Fang, F. Hu, S. Huang, K. Kundalia, Y.-C. Lin, et al., “Dreamgen: Unlocking generalization in robot learning through video world models,” in Proceedings of The 9th Conference on Robot Learning. PMLR, 2025, pp. 5170–5194. [15] H. Fang, C. Wang, Y. Wang, J. Chen, S. Xia, J. Lv, Z. He, X. Yi, Y. Guo, X. Zhan, et al., “Airexo-2: Scaling up generalizable robotic imitation learning with low-cost exoskeletons,” in Proceedings of The 9th Conference on Robot Learning. PMLR, 2025, pp. 198–220. [16] Y. Liu, J. Mao, J. B. Tenenbaum, T. Lozano-Pérez, and L. P. Kaelbling, “One-shot manipulation strategy learning by making contact analogies,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 15 387–15 393. [17] J. Wang, K. Liu, D. Guo, Z. Xian, and C. G. Atkeson, “One-shot video imitation via parameterized symbolic abstraction graphs,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 10 552–10 560. [18] H. Oh, A. M. Salcedo-Vázquez, I. G. Ramirez-Alpizar, and Y. Domae, “Robust instant policy: Leveraging student’s t-regression model for robust in-context imitation learning of robot manipulation,” in 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025, pp. 7973–7980. [19] V. Vosylius and E. Johns, “Instant policy: In-context imitation learning via graph diffusion,” in The Thirteenth International Conference on Learning Representations (ICLR), 2025. [20] Q. Zhao, M. Zheng, Z. Li, S. Huang, and W. Shi, “Robot dexterous grasping in cluttered scenes based on single-view point cloud,” IEEE Transactions on Automation Science and Engineering, 2025. [21] X. Zhong, T. Gong, J. Yu, J. Luo, C. Zhou, X. Zhong, and Q. Liu, “Region-aware grasping for stacked workpieces: A 6d-wise label selfgeneration method and robust evaluation strategy,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 16 899–16 912, 2025. [22] S. Yu, J. Yin, D.-H. Zhai, and Y. Xia, “Fast and accurate category-level object pose estimation without shape priors for robotic grasp detection,” IEEE Transactions on Automation Science and Engineering, 2025. [23] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research, vol. 44, no. 10-11, pp. 1684–1704, 2025. [24] Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu, “3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,” in Proceedings of Robotics: Science and Systems (RSS), 2024. [25] Y. Ze, Z. Chen, W. Wang, T. Chen, X. He, Y. Yuan, X. B. Peng, and J. Wu, “Generalizable humanoid manipulation with 3d diffusion policies,” in 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 2873–2880. [26] D. Chen, Z. Chen, X. Zheng, W. Xu, C. Ma, and C. Mao, “Adp: Adaptive diffusion policy energizes robots thinking in both learning and practice,” IEEE Transactions on Automation Science and Engineering, 2025. [27] J. Zhao, J. Liu, X. Su, B. He, and S. Jiang, “Robot few-shot manipulation skills learning based on meta imitation learning and mixture of experts model,” IEEE Transactions on Automation Science and Engineering, 2026. [28] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning. PMLR, 2017, pp. 1126–1135. [29] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.

[30] E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence, 2018. [31] E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning. PMLR, 2022, pp. 991– 1002. [32] Y. Kuang, J. Ye, H. Geng, J. Mao, C. Deng, L. Guibas, H. Wang, and Y. Wang, “Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation,” arXiv preprint arXiv:2407.04689, 2024. [33] M. Pan, J. Zhang, T. Wu, Y. Zhao, W. Gao, and H. Dong, “Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 17 359–17 369. [34] H. Xu, L. Zhang, X. Hu, B. Zhong, K. Bai, Z.-C. Márton, Z. Bing, Z. Chen, A. C. Knoll, and J. Zhang, “FUNCanon: Learning pose-aware action primitives via functional object canonicalization for generalizable robotic manipulation,” arXiv preprint arXiv:2509.19102, 2025. [Online]. Available: https://arxiv.org/abs/2509.19102 [35] K. Kedia, P. Dan, A. Chao, M. A. Pace, and S. Choudhury, “One-shot imitation under mismatched execution,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 15 649–15 656. [36] J. Luo, S. Yang, C. Zheng, J. Fan, Y. Gao, R. Song, and W. Zhang, “Covail: Zero-shot imitation learning via contrastive viewpoint alignment on object-centric representation,” IEEE Transactions on Automation Science and Engineering, 2026. [37] S. Wu, Y. Wang, and Y. Huang, “Autolfd: Closing the loop for learning from demonstrations,” IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 11 124–11 138, 2025. [38] C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in neural information processing systems, vol. 30, 2017. [39] C. Hou, Y. Ze, Y. Fu, Z. Gao, S. Hu, Y. Yu, S. Zhang, and H. Xu, “4d visual pre-training for robot learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 8451–8461. [40] Z. Liu, Y. Wang, K. Wang, L. Liang, X. Xue, and Y. Fu, “Spatialtemporal aware visuomotor diffusion policy learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 7122–7131. [41] Y. Liu, T. Jia, H. Zhang, G. Yang, H. Wang, and D. Chen, “Rasp: Robot active scene perception with joint viewpoint planning and depth completion in cluttered environments,” IEEE Transactions on Automation Science and Engineering, 2025. [42] M. Ding, Y. Liu, Y. Shi, X. Lan, and N. Zheng, “Dual graph attention networks for multi-view visual manipulation relationship detection and robotic grasping,” IEEE Transactions on Automation Science and Engineering, 2025. [43] Y. Song, P. Sun, P. Jin, Y. Ren, Y. Zheng, Z. Li, X. Chu, Y. Zhang, T. Li, and J. Gu, “Learning 6-dof fine-grained grasp detection based on part affordance grounding,” IEEE Transactions on Automation Science and Engineering, 2025. [44] Z. Zhao, J. Gao, and D. Zheng, “Affordance-guided robotic grasping via multimodal large language model reasoning,” IEEE Transactions on Automation Science and Engineering, 2026. [45] A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox, “Rvt-2: Learning precise manipulation from few demonstrations,” in Proceedings of Robotics: Science and Systems, 2024. [46] Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang, “Gnfactor: Multi-task real robot learning with generalizable neural feature fields,” in Conference on robot learning. PMLR, 2023, pp. 284–301. [47] G. Lu, S. Zhang, Z. Wang, C. Liu, J. Lu, and Y. Tang, “Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation,” in European Conference on Computer Vision. Springer, 2024, pp. 349– 366. [48] B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.” ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023. [49] T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” arXiv preprint arXiv:2402.10885, 2024. [50] J. Tian, L. Wang, S. Zhou, S. Wang, J. Li, H. Sun, and W. Tang, “Pdfactor: Learning tri-perspective view policy diffusion field for multitask robotic manipulation,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 15 757–15 767.

13

[51] S. Wang, L. Wang, S. Zhou, J. Tian, J. Li, H. Sun, and W. Tang, “Flowram: Grounding flow matching policy with region-aware mamba framework for robotic manipulation,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 12 176–12 186. [52] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning. PmLR, 2021, pp. 8748–8763. [53] T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on robot learning. PMLR, 2020, pp. 1094–1100. [54] S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “RLBench: The robot learning benchmark & learning environment,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3019–3026, 2020. [55] Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu, “Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2025, pp. 14 754–14 762. [56] Y. Zhong, Y. Liu, C. Xiao, Z. Yang, Y. Wang, Y. Zhu, Y. Shi, Y. Sun, X. Zhu, and Y. Ma, “Freqpolicy: Frequency autoregressive visuomotor policy with continuous tokens,” Advances in Neural Information Processing Systems, vol. 38, pp. 56 493–56 526, 2026. [57] J. Cao, Q. Zhang, J. Sun, J. Wang, H. Cheng, Y. Li, J. Ma, K. Wu, Z. Xu, Y. Shao, et al., “Mamba policy: Towards efficient 3d diffusion policy with hybrid selective state models,” in 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2025, pp. 11 359–11 366.

Record · ID 381758 · SHA-256 3a7f77f85df321be
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.