UniMate: One Unified Model to Animate Diverse Skeletons
arXiv:2609.05415v1 [cs.CV] 4 Sep 2026
LINZHAN MOU, Princeton University, USA JIAHUI LEI, University of California, Berkeley, USA ZHIYANG DOU, Massachusetts Institute of Technology, USA CHENYUE CAI, Princeton University, USA CHAOYUE SONG, Nanyang Technological University, Singapore ADAM FINKELSTEIN, Princeton University, USA SZYMON RUSINKIEWICZ, Princeton University, USA
Fig. 1. Given a rigged 3D asset and a text prompt, UniMate generates animations for characters with diverse skeletal topologies within a single unified model. Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on categoryspecific templates or require per-skeleton fine-tuning and reference motions at inference. We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining. UniMate introduces a topology-aware diffusion transformer, which integrates skeletal topology into attention via three mechanisms: (1) a graph-aware attention bias from pairwise joint relations and geodesic distances; (2) a spectral rotary position embedding generalizing RoPE to arbitrary kinematic trees via the graph Laplacian; and (3) a global topological conditioner attention-pooled from the rest-pose skeleton. We also curate UniML3D, 13,006 motion sequences spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects with unified canonicalization and text pairing. Trained on this dataset, UniMate outperforms state-of-theart baselines in quality, generalization, and efficiency, and supports zero-shot cross-topology transfer, in-betweening, expansion, and text-guided editing. Our project page is available at https://linzhanmou.com/unimate/. CCS Concepts: • Computing methodologies → Animation; Artificial intelligence; Machine learning. Authors’ Contact Information: Linzhan Mou, Princeton University, USA, linzhan@ princeton.edu; Jiahui Lei, University of California, Berkeley, USA, leijh@berkeley. edu; Zhiyang Dou, Massachusetts Institute of Technology, USA, [email protected]; Chenyue Cai, Princeton University, USA, [email protected]; Chaoyue Song, Nanyang Technological University, Singapore, [email protected]; Adam Finkelstein, Princeton University, USA, [email protected]; Szymon Rusinkiewicz, Princeton University, USA, [email protected].
Additional Key Words and Phrases: Motion Synthesis, Skeletal Animation, 3D Character Animation, Diffusion Models, Topology-Aware Learning
1
Introduction
Character animation is fundamental to 3D content creation, central to film, gaming, virtual reality, and robotics simulation. Traditional pipelines define a unique skeleton per character, over which artists craft motion through manual keyframing, motion capture, or slow per-asset optimization. Recent advances in 3D content creation [Li et al. 2024c; Liu et al. 2023c,b; Nichol et al. 2022; Wang et al. 2023; Xiang et al. 2025; Zhao et al. 2025a] and automatic rigging [Liu et al. 2025; Song et al. 2025a,b; Xu et al. 2020; Zhang et al. 2025b] now deliver skeleton-ready 3D assets across a broad range of categories, from humans and animals to articulated rigid objects. Yet while rigged assets can be generated at scale, the motion that drives them cannot—animation remains the laborious final bottleneck in an otherwise automated 3D content creation pipeline. The bottleneck lies in the design of existing learned animators. State-of-the-art motion generators [Karunratanakul et al. 2023; Petrovich et al. 2021; Raab et al. 2023; Rempe et al. 2026; Tevet et al. 2025, 2023; Wen et al. 2025; Zhao et al. 2025b] are typically topologyconstrained, relying on fixed category-specific skeleton templates such as SMPL [Loper et al. 2015] or SMAL [Zuffi et al. 2017]. Topologyagnostic models [Gat et al. 2025; Li et al. 2022a; Raab et al. 2024b] relax these constraints but still require per-skeleton fine-tuning or
2
•
Mou et al.
reference motions at inference time. Mesh-based methods avoid explicit skeleton modeling altogether, but either depend on costly per-asset distillation [Chen et al. 2025a; Jiang et al. 2024; Uzolas et al. 2025] or regress kinematically unconstrained vertex-wise deformations [Shi et al. 2025; Wu et al. 2025; Zhang et al. 2025c,a]. These limitations motivate a unified foundation model that can synthesize motion for arbitrary skeletal topologies directly from high-level descriptions. Such a model would animate any rig in a single feed-forward pass and share motion priors across topologies, generalizing to unseen rigs and supporting a range of downstream applications (Figs. 12 to 15). However, developing such a unified animator poses two fundamental challenges. The first is modeling: real-world skeletons are highly heterogeneous—bipedal humans, multi-legged insects, winged animals, and articulated rigid objects all exhibit distinct kinematic trees, joint counts, and motion patterns. This challenge is further compounded by the diverse motion behaviors associated with different morphologies. A general-purpose model must therefore treat skeletal topology as an explicit input, rather than baking it into an architectural prior, and reason jointly over structure and motion. The second is data: text-paired motion corpora spanning diverse skeletal topologies remain scarce, with existing benchmarks dominated by humans [Guo et al. 2022; Mahmood et al. 2019; Plappert et al. 2016] and a limited number of quadrupeds [Yang et al. 2024]. Meanwhile, raw rigged 4D assets [Deitke et al. 2023a,b; Truebones 2022] are often noisy and inconsistent, and lack unified preprocessing and canonicalization across topologies, leaving datadriven approaches without coherent supervision for cross-topology generalization. In this work, we propose UniMate, a unified foundation model that animates diverse skeletons. Given a rigged 3D asset and a natural-language prompt, UniMate synthesizes plausible articulated motion with no test-time fitting or per-skeleton specialization (see Fig. 1). Joint training across a wide range of skeletons lets the model learn motion patterns that are shared and transferable across topologies, enabling stronger generalization to unseen rigs and motion transfer between heterogeneous structures. At the core of UniMate is the Topology-Aware Diffusion Transformer (TADiT), a flow-matching architecture in which attention layers jointly reason over rest-pose kinematics and motion manifolds through a shared token stream. To encode heterogeneous topologies, we equip TADiT with three key design choices. First, vanilla self-attention is blind to the underlying kinematic graph. We therefore inject a graph-aware attention bias [Ying et al. 2021] derived from pairwise joint relations and geodesic distances, so anatomically nearby joints attend more strongly while the model retains its capacity for long-range, full-body coordination. Second, we introduce Spec-RoPE, a spectral rotary position embedding that generalizes RoPE [Su et al. 2024] to arbitrary kinematic trees by deriving rotary angles from the graph Laplacian spectrum. With provable translation invariance in spectral coordinates and equivariance under joint permutation, Spec-RoPE adapts to skeletons of varying size and connectivity—a property that index- or coordinatebased encodings cannot provide. Third, a global topological conditioner, attention-pooled from the skeleton tokens, modulates every transformer block through AdaLN-Zero [Peebles and Xie 2023], so
layer-wise feature statistics adapt to the input skeleton and provide global structural context complementing the local signals above. To support training at scale, we curate UniML3D, a heterogeneous motion dataset of 13,006 animation sequences (roughly 20 hours) drawn from Truebones [Truebones 2022], Mixamo [Adobe 2022], and Objaverse-XL [Deitke et al. 2023a,b], pairing thousands of distinct rigs across bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects with 3,584 unique text prompts that span a broad action spectrum, including locomotion, combat, idle, mechanical articulation, and object manipulation. Rigorous filtering followed by unified canonicalization yields a shared representation across skeleton types, while online skeletal augmentation further broadens topological coverage during training. This scale and coverage substantially exceed those of prior work [Gat et al. 2025] and are essential for cross-skeleton generalization. Extensive experiments demonstrate that UniMate achieves stateof-the-art performance on topology-agnostic motion generation and mesh animation, surpassing prior methods in quality, generalization, and efficiency. UniMate also supports zero-shot downstream tasks such as cross-topology motion transfer, in-betweening, expansion, and text-guided editing, serving as a controllable engine for scalable 3D character animation. In summary, our contributions are: • We present UniMate, a unified foundation model that synthesizes articulated motion for skeletons of arbitrary topology from a rigged 3D asset and a text prompt. • We propose TADiT, which couples motion and skeletal structure within shared attention layers through a graph-aware attention bias, the Spec-RoPE spectral rotary position embedding with provable structural properties, and a global topological conditioner. • To facilitate training and benchmarking, we curate UniML3D, comprising 13,006 motion sequences over thousands of skeletons spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects, unified by a canonicalization pipeline and online skeletal augmentation. • We conduct various experiments showing that UniMate improves over prior methods in quality, generalization, and runtime, and enables zero-shot cross-topology motion transfer, in-betweening, expansion, and real-time text-guided editing.
2 Related Work 2.1 3D Animation A growing body of work animates 3D assets by distilling image- and video-generative priors or by reconstructing motion from generated videos [Chen et al. 2025a; Huang et al. 2025; Jiang et al. 2024; Liang et al. 2024; Lyu et al. 2026; Mou et al. 2025; Uzolas et al. 2025; Yun et al. 2025]. These methods typically require costly per-asset test-time optimization and tend to produce jittery, unstable trajectories, making them ill-suited to large-scale or interactive use. A second class of methods [Chen et al. 2026; Sabathier et al. 2026; Shi et al. 2025; Wang et al. 2026b; Wu et al. 2025; Yenphraphai et al. 2025; Zhang et al. 2025c,a] trains feed-forward networks that regress per-vertex or per-point deformations directly. Operating in raw geometry space sacrifices the compactness of skeletal representations and breaks
UniMate: One Unified Model to Animate Diverse Skeletons
•
3
Fig. 2. Animations generated by UniMate. Our method generalizes across heterogeneous skeletons and diverse motion prompts.
native compatibility with the rig-driven ecosystem: linear blend skinning [Li et al. 2021; Magnenat-Thalmann et al. 1988], physics-based controllers [Tevet et al. 2025], physics simulators [Makoviychuk et al. 2021; Todorov et al. 2012], and motion-capture pipelines [Gong et al. 2025]. Recent advances in automatic rigging [Liu et al. 2025; Song et al. 2025a,b; Xu et al. 2020; Zhang et al. 2025b] have made high-quality skeletons widely accessible, yet driving the resulting rigs still demands manual keyframing or expensive per-asset optimization [Li et al. 2025; Song et al. 2025a; Xie et al. 2025]. UniMate closes this gap with a data-driven generative model that operates directly on arbitrary skeletons.
2.2
Cross-Topology Motion Generation and Retargeting
The dominant family of learned motion generators, from human motion models [Chen et al. 2024; Dou et al. 2023; Fan et al. 2025; Karunratanakul et al. 2023; Li et al. 2024a; Lu et al. 2025; Meng et al. 2025; Raab et al. 2024a, 2023; Rempe et al. 2026; Sawdayee et al. 2026; Shafir et al. 2024; Tevet et al. 2025, 2023; Wan et al. 2024; Wen et al. 2025; Zhao et al. 2025b; Zhou et al. 2024] to species-specific animal models [Sun et al. 2024; Wang et al. 2026a, 2025], assumes a single fixed skeleton template, typically inherited from parametric body models such as SMPL [Loper et al. 2015; Pavlakos et al. 2019] and SMAL [Zuffi et al. 2017], and therefore cannot handle characters whose topology departs from the template. A separate family of example-based methods sidesteps neuralnetwork training by stitching patches from exemplar motion clips: generative motion matching [Li et al. 2023] carries the patch nearestneighbor synthesis of Drop-the-GAN [Granot et al. 2022] over to motion, and Motion2Motion [Chen et al. 2025b] extends it to cross-topology transfer through sparse correspondences. Although topology-flexible, these methods still require an exemplar motion for every target and cannot synthesize motion from a static rigged asset alone. Closer to our setting, GANimator [Li et al. 2022a] and SinMDM [Raab et al. 2024b] learn neural generators on arbitrary topologies, but train a separate model per skeleton and so do not generalize across structures. Cross-topology retargeting [Aberman et al. 2020; Lee et al. 2023; Li et al. 2024b; Liu et al. 2026; Zhao et al. 2024] bypasses the template constraint from a different angle,
but again only by transferring an existing source motion onto a target skeleton. The closest prior work, AnyTop [Gat et al. 2025], jointly trains a single diffusion model over heterogeneous animal skeletons, but is limited to a small animal corpus [Truebones 2022], requires motion data of the target skeleton at inference to estimate its normalization statistics, and offers no text conditioning. In contrast, a single UniMate model covers a far broader range of skeletons, from humans and animals to general articulated objects, accepts text conditioning, and animates a rigged mesh end-to-end, without reference motion or per-skeleton training.
3
Method
Given a rigged 3D asset and a text prompt, our goal is to synthesize a plausible motion sequence that animates the input mesh (Fig. 3). We first introduce a unified representation for heterogeneous skeletons and motion sequences (Section 3.1), then present our TopologyAware Diffusion Transformer (Section 3.2), and finally describe the training objective and inference procedure (Section 3.3).
3.1
Skeleton and Motion Representation
An articulated rigged 3D asset is animated by a skeleton—a kinematic tree with a joint hierarchy and bone lengths, whose joints drive mesh deformation through forward kinematics. A rest pose of a skeleton is the neutral undeformed reference configuration. Our model takes as input the skeleton at rest pose and builds a diffusion model over a motion sequence defined on it. Skeleton Definition. We model an articulated object as a rooted kinematic tree with 𝐽 joints. To represent skeletons with heterogeneous topologies in a shared transformer space, we first assign each skeleton a canonical joint ordering. Concretely, we linearize the tree using breadth-first search (BFS) given the known root node JS = ( 𝑗1, pa1 ), ( 𝑗2, pa2 ), . . . , ( 𝑗 𝐽 , pa 𝐽 ) , (1) where 𝑗𝑘 ∈ R3 denotes the rest-pose position of the 𝑘-th joint and pa𝑘 ∈ {1, . . . , 𝐽 } denotes its parent index. The BFS ordering places the root at 𝑘 = 1; since it has no parent, we adopt the self-parent convention pa1 = 1.
4
•
Mou et al.
Fig. 3. UniMate pipeline. The proposed TADiT operates on a joint token space that combines per-frame motion features with rest-pose skeletal descriptors, while injecting skeletal graph-structured bias into both positional encoding and attention computation. Conditioned on input text prompts, our unified model produces realistic, coherent animations while demonstrating strong cross-topology generalization and efficient runtime performance.
Topology-Diameter Normalization. Skeletons in our dataset vary significantly in absolute size and topological extent. To place them into a comparable representation space, we normalize each rest-pose skeleton by its topology diameter, defined as geo
𝑑 topo = max disttree (𝑖, 𝑗),
(2)
𝑖,𝑗
geo
where disttree (𝑖, 𝑗) denotes the geodesic distance between joints 𝑖 and 𝑗 along the kinematic tree. This normalization removes scale differences while preserving the relative kinematic structure. Local Topology Descriptors. Given the canonicalized skeleton, we describe its rest-pose geometry and local topology using several descriptors. Rest-pose joint positions are stacked in PS ∈ R 𝐽 ×3 . Local kinematic context is encoded by a pairwise relation matrix R S ∈ N0𝐽 ×𝐽 , whose entries specify relation types such as parent, child, sibling, or ancestor; a pairwise graph-distance matrix GS ∈ N0𝐽 ×𝐽 on the kinematic tree; a per-joint depth vector D S ∈ N0𝐽 ; and a per-joint name index NS ∈ N0𝐽 into a joint-name vocabulary that captures semantic identity. Spectral Coordinates. To complement the discrete topological descriptors above with a continuous encoding of skeletal structure, we additionally compute spectral features from the kinematic graph. Let 𝐴 ∈ {0, 1} 𝐽 ×𝐽 denote the adjacency matrix of the skeleton and 𝐿 S = diag(𝐴1) − 𝐴 its graph Laplacian. Let 𝐿 S = 𝑈 Λ𝑈 ⊤ be its eigendecomposition. Discarding the trivial constant eigenvector 𝑢 0 , we define the spectral feature of joint 𝑗 using the first 𝑚 non-trivial eigenvectors: FS ( 𝑗) = [𝑢 1 ( 𝑗), 𝑢 2 ( 𝑗), . . . , 𝑢𝑚 ( 𝑗)] ∈ R𝑚 .
(3)
S = {JS , PS , R S , GS , D S , NS , FS },
(4)
These spectral coordinates provide a continuous encoding of joint location on the kinematic graph (see Fig. 7), complementing the discrete relation and distance descriptors. Overall, we represent a skeleton as
where JS specifies the ordered kinematic tree, PS the rest-pose geometry, R S , GS , D S the discrete topological structure, NS the semantic joint identity, and FS the global spectral descriptor. Motion Representation. Given a skeleton S, a motion sequence is 𝑋 = {m𝑡𝑗 } ∈ R𝑇 ×𝐽 ×𝐷 over 𝑇 frames and 𝐽 joints. Following [Gat et al. 2025; Guo et al. 2022], each joint feature has dimension 𝐷 = 12: m𝑡𝑗 = [p𝑡𝑗 , r𝑡𝑗 , v𝑡𝑗 ] ∈ R𝐷 .
(5)
p𝑡𝑗 ∈ R3 : Position. For non-root joints, horizontal (𝑥, 𝑧) components are taken relative to the root and rotated into the facingcanonical frame, with vertical 𝑦 kept in the world frame. The root joint p𝑡1 retains its global position. r𝑡𝑗 ∈ R6 : Rotation. Represented in the continuous 6D format [Zhou et al. 2019]. v𝑡𝑗 ∈ R3 : Velocity. Computed in the world frame as the temporal derivative of global joint positions, preserving absolute motion cues that the canonical-frame projection of p𝑡𝑗 would otherwise discard.
3.2
Topology-Aware Diffusion Transformer (TADiT)
Our goal is to generate a plausible motion sequence 𝑥ˆ1 conditioned on a rest-pose skeleton S and a text prompt 𝑐. This is challenging because the model must generalize across heterogeneous skeletons with different kinematic trees, joint counts, and motion patterns. To address this, we introduce the Topology-Aware Diffusion Transformer (TADiT), which injects topology through a graph-aware attention bias, a spectral rotary position embedding (Spec-RoPE), and a global topological conditioner. We train TADiT with conditional flow matching [Liu et al. 2023a], learning a velocity field 𝑣𝜃 (𝑥𝜏 , 𝜏, 𝑐, S) that transforms Gaussian noise 𝑥 0 ∼ N (0, 𝐼 ) into a valid motion sequence 𝑥 1 . Skeleton and Motion Tokenization. For each joint 𝑗, we form the skeleton token T 𝑗 ∈ R𝑑 by concatenating MLP-projected rest-pose
UniMate: One Unified Model to Animate Diverse Skeletons
•
5
Fig. 4. Cross-topology animation. UniMate generates prompt-aligned motions for diverse characters and articulated objects within a single unified model.
positions of the joint and its parent and applying a fusion MLP: (6)
and a joint-name text embedding as hierarchical and semantic priors: M𝑡𝑗 = MLP(m𝑡𝑗 ) + 𝐸 depth D S ( 𝑗) + 𝐸 name NS ( 𝑗) . (7)
where PS ( 𝑗) is the rest-pose position of joint 𝑗. Each motion feature m𝑡𝑗 ∈ R𝐷 (defined in Section 3.1) is projected to dimension 𝑑 and augmented with a learnable depth embedding
We prepend the skeleton tokens to the motion tokens along the temporal axis, forming 𝑍 = Concat(T, M) ∈ R (𝑇 +1) ×𝐽 ×𝑑 , such that every transformer block operates on a unified token space containing both static skeletal structure and dynamic per-frame motion.
T 𝑗 = MLPfuse Concat MLPjt (PS ( 𝑗)), MLPpa (PS (pa 𝑗 )) ,
6
•
Mou et al.
Fig. 5. One skeleton, diverse prompts. Given a single skeleton, UniMate synthesizes distinct, prompt-faithful motions for different input text prompts.
Fig. 6. One skeleton, one prompt, diverse motions. Given the same skeleton and text prompt, UniMate generates diverse plausible motion samples.
Skeletal-Temporal Transformer Blocks. Each block operates on 𝑍 ∈ R (𝑇 +1) ×𝐽 ×𝑑 and is conditioned on the diffusion timestep, text prompt, and global topology embeddings. For tractable cost on heterogeneous skeletons, each block uses a factorized attention with a joint branch (across joints at each frame) and a temporal branch (across frames at each joint), followed by a feed-forward sublayer: 𝑍 (ℓ+1) = FFN TempAttn JointAttn(𝑍 (ℓ ) , S), 𝑐 , 𝑐 . (8) The temporal branch is a standard multi-head self-attention with 1D rotary position embedding (RoPE) [Su et al. 2024] on the frame index. Kinematic structure is exposed exclusively to the joint branch, through a graph-aware attention bias and the Spectral Rotary Position Embedding (Spec-RoPE) described next. Graph-Aware Attention Bias. We inject the kinematic graph into attention through a learned bias added to the joint-attention logits, exposing pairwise structural relations that are hard for vanilla selfattention to recover from token features alone. Following Ying et al.
[2021], the pairwise graph-distance and relation descriptors GS , R S are embedded by lookup tables 𝑒𝑑 (𝑖, 𝑗) = 𝐸𝑑 (GS (𝑖, 𝑗)) ,
𝑒𝑟 (𝑖, 𝑗) = 𝐸𝑟 (R S (𝑖, 𝑗)) ,
and projected to a per-head scalar bias
(ℎ) ⊤ (ℎ) ⊤ 𝐵𝑖(ℎ) 𝑗 = (𝑤𝑑 ) 𝑒𝑑 (𝑖, 𝑗) + (𝑤 𝑟 ) 𝑒𝑟 (𝑖, 𝑗),
(10)
with head-specific projection vectors 𝑤𝑑(ℎ) , 𝑤𝑟(ℎ) ∈ R𝑑𝑒 . Letting 𝑞˜ (ℎ) , 𝑘˜ (ℎ) ∈ R𝑑ℎ denote the Spec-RoPE-rotated queries and keys 𝑖
𝑗
for joints 𝑖 and 𝑗, the joint-attention logits read ⊤ Attn (ℎ) (𝑖, 𝑗) = √1 𝑞˜𝑖(ℎ) 𝑘˜ 𝑗(ℎ) + 𝐵𝑖(ℎ) 𝑗 .
(11)
𝑑ℎ
The bias is shared across frames, so its memory cost is independent of sequence length. Because it is parameterized by graph-distance and relation-type embeddings rather than absolute joint indices, it transfers to unseen topologies (Section 5.6). Spectral Rotary Position Embedding (Spec-RoPE). Kinematic trees have no canonical ordering, so the index used by 1D RoPE is illdefined for joints. We instead derive rotary angles from the spectrum of the graph Laplacian [Dwivedi and Bresson 2021; Rampášek et al. 2022], applied to the joint branch only: 𝜽 𝑗 = 𝝎 𝑗 −→ 𝜽 𝑗 = 𝑓 (s 𝑗 ), | {z } | {z } 1D RoPE
Fig. 7. Spectral visualization. From left to right, Laplacian eigenvectors increase in frequency. Low-frequency modes capture global kinematic structure, while higher-frequency modes encode finer local relationships.
(9)
(12)
Spec-RoPE
where 𝑗 ∈ N is the token index, 𝝎 ∈ R𝑑ℎ /2 the standard frequency vector, s 𝑗 = FS ( 𝑗) ∈ R𝑚 the joint’s spectral coordinate from Section 3.1, and 𝑓 a learned angle map specified below. Intuition. RoPE relies on positional coordinates to define relative phase offsets in attention. For temporal tokens, the frame index
UniMate: One Unified Model to Animate Diverse Skeletons
•
7
Fig. 8. Long-horizon generation. A variant trained at 180 frames generates long sequences that stay temporally coherent and drift-free.
is a natural causal position; for kinematic-tree joints, the BFS index is arbitrary: two joints adjacent in the index can lie on opposite limbs. The spectral coordinate s 𝑗 replaces it with an intrinsic position on the graph: the low-frequency Laplacian eigenvectors capture the coarse global organization of the kinematic tree, while higher-frequency eigenvectors progressively encode finer structural variation (Fig. 7). The rotary phase therefore depends on where a joint sits on the skeleton, not on how it is serialized. The spectral coordinate s 𝑗 = FS ( 𝑗) ∈ R𝑚 comprises the leading 𝑚 non-trivial eigenvectors of 𝐿 S . Since these are determined only up to sign, we realize 𝑓 as a SignNet [Lim et al. 2023]: 𝑚 𝜽 𝑗 = MLPang Concat MLPsym (𝑢𝑛 ( 𝑗)) + MLPsym (−𝑢𝑛 ( 𝑗)) 𝑛=1 . (13) The angles 𝜽 𝑗 ∈ R𝑑ℎ /2 then drive the standard RoPE block-diagonal rotation R(𝜽 𝑗 ) [Su et al. 2024], applied to per-joint queries and keys as 𝑞˜ 𝑗 = R(𝜽 𝑗 ) 𝑞 𝑗 and 𝑘˜ 𝑗 = R(𝜽 𝑗 ) 𝑘 𝑗 . Spec-RoPE satisfies two structural properties, formalized in Appendix C: translation invariance in spectral coordinates and equivariance under joint permutation. Together, these allow Spec-RoPE to adapt to skeletons of varying size and connectivity. We empirically validate the effect of Spec-RoPE in Section 5.6. Global Topological Conditioner. Beyond the pairwise graphaware attention bias and per-joint Spec-RoPE topology signals, each transformer block also requires a single, joint-count-invariant skeleton summary. We obtain this global topological condition 𝑐 topo via attention pooling [Lee et al. 2019] over the skeleton token sequence T, producing a fixed-size summary independent of joint count and, after the final mean aggregation, of joint ordering. A small set of 𝑛𝑞 learnable query tokens Qpool ∈ R𝑛𝑞 ×𝑑 serves as a content-adaptive readout: each query attends to all skeleton tokens through crossattention, with T providing the keys and values, H = softmax √1 (Qpool𝑊𝑞 )(T𝑊𝑘 ) ⊤ T𝑊𝑣 , (14) 𝑑
where 𝑊𝑞 ,𝑊𝑘 ,𝑊𝑣 ∈ R𝑑 ×𝑑 are learnable projections. The query outputs H ∈ R𝑛𝑞 ×𝑑 are aggregated by mean pooling into 𝑐 topo , which
is fused with the timestep and text embeddings and injected into every block via AdaLN-Zero [Peebles and Xie 2023] (detailed in Appendix B.2), yielding a holistic skeletal context.
3.3
Training and Inference
Training Objective. We supervise 𝑣𝜃 with three complementary losses, with full expressions deferred to Appendix B.4. Flow-matching MSE. The base loss Lmse is a masked mean-squared error against the target velocity 𝑣 ∗ = 𝑥 1 − 𝑥 0 , with padded joint slots zeroed out in heterogeneous-skeleton batches. Geodesic rotation loss. Because joint rotations live on the nonEuclidean manifold 𝑆𝑂 (3), an isotropic MSE on their 6D channels is geometrically misaligned. We therefore add a geodesic loss Lgeo on the one-step denoised rotations 𝑅ˆ𝑡𝑗 ∈ 𝑆𝑂 (3) from 𝑥ˆ1 = 𝑥𝜏 + (1−𝜏) 𝑣𝜃 , which penalizes their angular deviation from the ground truth. Velocity smoothness. To suppress high-frequency jitter, a smoothness regularizer Lsmooth penalizes the temporal acceleration of the denoised velocity channels of 𝑥ˆ1 . The final objective is a weighted combination of these three losses: L = Lmse + 𝜆geo Lgeo + 𝜆smooth Lsmooth .
(15)
Inference. At inference time, we draw 𝑥 0 ∼ N (0, 𝐼 ) and integrate the learned velocity field 𝑥¤𝜏 = 𝑣𝜃 (𝑥𝜏 , 𝜏, 𝑐, S) from 𝜏 = 0 to 1 using a fixed-step Euler solver, yielding the generated motion features 𝑥ˆ1 . Each step applies classifier-free guidance [Ho and Salimans 2022].
4
UniML3D Dataset
Training a truly generalizable animation model for diverse object categories requires a large-scale dataset with varied skeletal structures and plausible motion sequences. However, raw 4D motion sources are noisy and inconsistent, often containing disconnected or scene-level skeletons, broken roots, non-functional joints, physically implausible motions, and mismatched coordinate frames or facing directions, making them unsuitable for direct cross-topology motion learning without rigorous preprocessing.
8
•
Mou et al.
Fig. 9. Samples from the UniML3D dataset. Our dataset spans diverse skeletons across bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects, with detailed skeleton annotations and coherent text prompts paired with motion sequences.
To address this, we curate UniML3D from three complementary sources: Truebones [Truebones 2022], Mixamo [Adobe 2022], and Objaverse-XL [Deitke et al. 2023a,b]. After filtering and canonicalization, UniML3D comprises 13,006 motion sequences and 2,140,232 frames over thousands of skeletons, spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects (Fig. 9). Fig. 10 summarizes this pipeline; per-source statistics and per-stage details are in Appendix A. Skeleton-Based Filtering. We apply a multi-stage filter so that every retained skeleton is connected and kinematically valid, and every retained clip carries plausible, non-trivial motion: (1) Single-tree pruning. Multiple or disconnected kinematic trees are pruned, keeping only the primary skeleton: the tree with the largest cumulative skinning weight. (2) Root-joint realignment. Spurious or misaligned root joints are corrected by propagating their global transformation onto the semantic root via forward kinematics. (3) Phantom-joint removal. Non-functional joints, e.g., IK controllers and helper bones with zero skinning weight, are recursively pruned to streamline the topology. (4) Static-clip removal. Clips with negligible activity, quantified by bone-length-normalized global displacement, are discarded. (5) Implausibility filtering. Clips with out-of-distribution root velocities or per-joint angular jitter above an anatomical threshold are discarded. Truebones animal motions
Mixamo human motions
Objaverse-XL rigged-asset motions
1. Skeleton-based filtering: single-tree pruning · root-joint realignment · phantom-joint removal · static-clip removal · implausibility filtering 2. Semantic annotation: LLM joint-name standardization & symmetric-pair selection (facing) · four-view rendering at 30 FPS · multimodal-LLM captioning 3. Motion canonicalization: BFS serialization · topology-diameter scaling · canonical root frame · rest-pose-relative rotations · feature normalization UniML3D: canonicalized skeleton–motion–caption training samples Training time (on the fly): square-root-balanced sampling · augmentations (joint removal · joint addition · skeleton pooling · bone-length perturbation)
Fig. 10. UniML3D data-processing pipeline. Source assets pass through skeleton-based filtering, semantic annotation, and motion canonicalization; balanced sampling and augmentation happen on the fly during training.
Motion Canonicalization. Building on the BFS serialization and topology-diameter normalization of Section 3.1, we map every motion into a unified canonical space: (1) Each animation is placed in a canonical coordinate frame with 𝑦-axis up and the initial root position at the origin. (2) The initial facing direction is aligned with the positive 𝑧-axis, estimated as f = proj𝑥𝑧 (e𝑦 × vhip ), where vhip is the left-toright hip (or any symmetric-joint pair) direction. (3) Joint rotations are expressed relative to the rest pose, yielding a unified kinematic basis across heterogeneous skeletons. (4) We compute global statistics (𝜇 global, 𝜎global ) ∈ R𝐷 over rootjoint features and local statistics (𝜇 local, 𝜎local ) ∈ R𝐷 over non-root features, and apply them to normalize all motions.
5
Experiments
We evaluate UniMate from five perspectives: presenting qualitative results across diverse rigs, comparing with skeleton-based motion generation methods, benchmarking against skeleton-free mesh animation baselines, demonstrating broader applications, and ablating our topology-aware design choices.
5.1
Implementation Details
Our motion diffusion transformer consists of 8 blocks with hidden dimension 512. We use a frozen FLAN-T5 [Chung et al. 2024] text encoder and train the model with AdamW [Loshchilov and Hutter 2019] at a learning rate of 1 × 10−4 . To address the long-tailed topology distribution of UniML3D and improve generalization, we adopt square-root-balanced sampling together with four kinematicspreserving on-the-fly augmentations: joint removal, joint addition, skeleton pooling [Aberman et al. 2020], and bone-length perturbation. Training is performed with a global batch size of 256 on 8 NVIDIA H100 GPUs for one day, and inference runs at 50 FPS.
5.2
Qualitative Results
Figs. 2 and 4 showcase animations generated by UniMate across heterogeneous rigs, spanning humanoids, quadrupeds, avians, insects, and articulated rigid objects: a single unified model produces prompt-faithful, temporally coherent motion while adapting to each skeleton’s structure. Figs. 5 and 6 show that a single rig faithfully follows distinct prompts and yields diverse yet prompt-consistent samples under the same prompt, reflecting both the controllability and the generative diversity of the model, and Fig. 8 demonstrates temporally coherent, drift-free long-horizon generation.
UniMate: One Unified Model to Animate Diverse Skeletons
•
9
Fig. 11. Comparison with AnyTop on unseen Truebones skeletons. Ours (left) vs. AnyTop (right): our deer falls and lies on its side and our raptor walks forward as prompted, whereas AnyTop’s deer never falls and its raptor remains nearly stationary.
Fig. 12. Text-mediated motion transfer. The source behavior is abstracted into a text prompt (shown atop each group); the same prompt then animates four target rigs of differing topology, conditioned directly on each target skeleton.
5.3
5,861 frames. Specifically, (1) Motion quality is measured by the Fréchet Inception Distance (FID) between extracted kinematic features [Chen et al. 2025b; Li et al. 2022b] of generated and groundtruth motions, and (2) Diversity is computed as the average pairwise joint distance among 5 generated samples. The held-out skeletons, prompts, and seed appear in Appendix D.2. Results. UniMate lowers FID from 2.711 to 0.757 and raises diversity from 8.139 to 9.200 (Tab. 1), so fidelity is not bought with mode collapse, and the margin reflects the generalization of our topology-aware design to unseen skeletons. Fig. 11 qualitatively demonstrates that UniMate produces high-quality, prompt-faithful motion, whereas AnyTop either remains nearly static or ignores the prompt; Appendix E.1 adds further morphologies and motions.
Comparison on Topology-Aware Motion Generation
Baselines. Our task is to generate a motion sequence conditioned on an input skeleton and a text prompt. To our knowledge, no existing baseline directly addresses this setting. The closest prior method is AnyTop [Gat et al. 2025], which supports unconditional motion generation for diverse skeletons on Truebones [Truebones 2022]. To adapt it to our setting and enable a fair comparison, we augment AnyTop with a cross-attention module for text conditioning, following MDM [Tevet et al. 2023]. We train and evaluate both the extended AnyTop and our model on Truebones. Metrics. We randomly hold out 7 skeleton types as unseen test skeletons, covering bipedal, quadrupedal, avian, marine, insectoid, and serpentine categories, for a total of 58 motion sequences and
Table 1. Comparison on motion gen- Table 2. Comparison on mesh animation. Our Table 3. User study on mesh animation. Our model eration. Our model clearly outperforms method outperforms prior state-of-the-art approaches. achieves the best text–motion alignment and motion quality. AnyTop in both quality and diversity. Method
Method AnyTop [2025] UniMate (Ours)
FID↓
2.711 0.757
Div.↑ 8.139 9.200
V2M4 [2025a] AnimateAnyMesh [2025] UniMate (Ours)
OC ↑
0.167 0.151 0.186
MS ↑
0.991 0.995 0.993
DD ↑
0.667 0.352 0.833
AQ ↑
0.506 0.498 0.544
Time ↓
1.641h 15.542s 1.214s
Method
V2M4 [2025a] AnimateAnyMesh [2025] UniMate (Ours)
TA ↑
2.378 2.193 4.617
MP ↑
2.341 2.930 4.568
ME ↑
2.596 2.362 4.646
SP ↑
2.336 3.747 4.630
Avg. ↑ 2.413 2.808 4.615
10
•
Mou et al.
Fig. 13. Motion in-betweening. Given the start and end poses and a text prompt, UniMate synthesizes smooth and plausible intermediate motions; the opaque poses mark the given boundary constraints, while the translucent poses are the synthesized in-betweens.
Fig. 14. Motion editing. Given a source motion (left), a subset of joints stays fixed while the rest (dashed boxes) are resampled under a new prompt (right).
Fig. 15. Motion expansion. Given sequential text prompts—here standing up, walking forward, then turning in place—UniMate extends a motion with smooth transitions; dark gray marks the boundary frames shared by consecutive segments.
5.4
Comparison on Text-Conditioned Mesh Animation
Baselines. To further evaluate UniMate in a broader animation setting, we compare against two state-of-the-art skeleton-free and vertex-wise mesh animation baselines: (1) AnimateAnyMesh [Wu et al. 2025], a feed-forward, text-conditioned framework; and (2) V2M4 [Chen et al. 2025a], an optimization-based, monocular videoconditioned method. For a fair comparison, all methods are evaluated using the same text prompts. For V2M4, we use Wan2.2 [Wan Team, Alibaba Group 2025] to generate the driving videos. Metrics. Our benchmark consists of 18 randomly selected meshes spanning bipeds, quadrupeds, avians, marine life, insects, and general articulated objects. Following [Huang et al. 2025; Wu et al. 2025], we render 512 × 512 multi-view videos from fixed viewpoints and assess perceptual quality with VBench [Huang et al. 2024] along four axes: overall consistency (OC), motion smoothness (MS), dynamic degree (DD), and aesthetic quality (AQ). We also report the average generation time per mesh animation. We further conduct a user study with 32 participants, each reviewing 12 test cases and rating each animation on a 5-point Likert scale (1 = very poor, 5 = excellent) along text-to-motion agreement (TA), motion plausibility (MP), motion expressiveness (ME), and shape preservation (SP). Results. In Tabs. 2 and 3, UniMate outperforms prior methods on most VBench metrics and achieves the highest scores across all user-study criteria, indicating stronger text–motion alignment and
more expressive, plausible, and coherent motion; AnimateAnyMesh attains higher motion smoothness mainly due to near-static outputs with much lower dynamic degree and expressiveness, while V2M4 relies on separately generated driving videos and is substantially slower and less practical.
5.5
More Applications
All four applications below are zero-shot: they reuse the same pretrained model without fine-tuning or auxiliary networks, differing only in which motion tokens are held fixed during sampling. Text-Mediated Motion Transfer. UniMate naturally supports cross-topology motion transfer, as shown in Fig. 12: a source motion is first abstracted into a language prompt, which then animates a target rig of different topology. Since generation is directly conditioned on the target skeleton, no joint correspondence, exemplar alignment, or per-skeleton optimization is required. Motion In-Betweening. UniMate further supports zero-shot motion in-betweening, as shown in Fig. 13. Given a target skeleton, a text prompt, and prescribed start and end poses, we perform sampling-time pose guidance by keeping the boundary-frame pose tokens fixed during the flow integration, while classifier-free text guidance enforces the motion semantics. This produces coherent transitions that satisfy the endpoint constraints while language specifies the transition’s style and intent. The same guidance extends beyond endpoints: keyframes of arbitrary number and position can
UniMate: One Unified Model to Animate Diverse Skeletons
•
11
Fig. 16. Foot sliding. In a generated walk, the hind-paw contacts drift along the ground (red traces) instead of staying planted.
Fig. 17. Rare-topology failure. Prompted to fall onto the desk, the lamp only dips and recovers its head; the base never tips over.
be held fixed as a sampling-time mask, and UniMate infills coherent, prompt-consistent motion between them. Motion Expansion. UniMate can extend an animation by chaining text prompts, as shown in Fig. 15: each new segment is generated with the preceding segment’s final pose as a boundary condition, so the sequence grows smoothly while its semantics evolve. Chaining complements the long-horizon variant of Fig. 8: the variant widens the temporal window of one sampling pass, while chaining composes arbitrarily many prompted segments. Text-Guided Motion Editing. Fig. 14 showcases text-guided motion editing. Starting from an existing animation, we keep the unchanged parts fixed and resample selected joints under a new text prompt. This enables localized edits such as changing action intensity, modifying limb behavior, or altering motion direction without regenerating the entire sequence. Because the model jointly reasons over motion and topology, the edits remain temporally smooth and structurally consistent with the input rig. Since only the selected joints are resampled, edits complete at interactive rates, supporting iterative prompt-driven refinement.
constraint-guided sampling; where contacts are well defined, foot locking or IK post-processing can be applied to the predicted rig. Appendix E.5 quantifies foot sliding on held-out legged skeletons and reports the effect of IK-based foot locking. Rare Topologies and Motions. UniMate is less reliable on rare skeletal topologies and out-of-distribution motions, where results can become static, jittery, or semantically inaccurate (Fig. 17). This stems from the scarce, long-tailed 4D animation data, which favor humanoids and common locomotion. Future work could distill Internet-scale video priors to broaden topology and motion coverage while retaining UniMate’s topology-aware backbone. Agentic asset-generation systems [Zhou et al. 2026] offer a complementary data-side remedy: pairing automatically generated assets with scripted, simulated, or distilled motion could help densify the rare topology and motion regions where captured data is scarce.
5.6
Ablation Study
In Tab. 4, we ablate the three topology-aware components of TADiT. Removing the graph-aware attention bias increases FID, showing the value of explicit structural relations in joint attention. Removing Spec-RoPE further degrades both fidelity and diversity, indicating weaker generalization to unseen skeletons. Removing the global topological conditioner causes the largest drop in motion quality; although diversity increases, the generated motions become less stable and often exhibit jitter. Appendix E complements these numbers with qualitative renderings of each ablated variant, along with further ablations of the data-processing pipeline. Table 4. Ablation. Every topology-aware component improves quality. Method w/o graph-aware attention bias w/o Spec-RoPE w/o global topological conditioner Full Model
6
FID↓
Diversity↑
0.757
9.200
0.773 0.798 0.825
8.799 8.424 9.475
Limitations and Future Work
Contact and Foot Sliding. UniMate can produce foot sliding, drift, hovering, or ground penetration in contact-rich motions (Fig. 16), because it imposes no unified contact model: foot–ground contact is meaningful for bipeds and quadrupeds, but ill-defined for snakes, swimming fish, birds in flight, and many articulated objects. Future work could introduce morphology-aware contact objectives during
7
Conclusion
We presented UniMate, a unified foundation model that synthesizes articulated motion for skeletons of arbitrary topology from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton specialization. At its core, the Topology-Aware Diffusion Transformer couples motion and skeletal structure through a graph-aware attention bias, the Spec-RoPE spectral rotary position embedding, and a global topological conditioner, while UniML3D provides large-scale motion supervision across thousands of heterogeneous skeletons. UniMate achieves state-of-the-art quality, generalization, and efficiency, and the same pretrained model supports motion transfer, in-betweening, expansion, and editing zero-shot. We hope UniMate and UniML3D provide a foundation for scalable, controllable animation of arbitrary rigged assets, and that coupling them with video priors and agentic data generation will further broaden their coverage.
References
Kfir Aberman, Peizhuo Li, Dani Lischinski, Olga Sorkine-Hornung, Daniel Cohen-Or, and Baoquan Chen. 2020. Skeleton-aware networks for deep motion retargeting. ACM Transactions on Graphics 39, 4 (2020), 62:1–62:14. doi:10.1145/3386569.3392462 Adobe. 2022. Mixamo. https://www.mixamo.com/. Hongyuan Chen, Xingyu Chen, Youjia Zhang, Zexiang Xu, and Anpei Chen. 2026. Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis. arXiv preprint arXiv:2601.14253. Jianqi Chen, Biao Zhang, Xiangjun Tang, and Peter Wonka. 2025a. V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 11643–11653. Ling-Hao Chen, Yuhong Zhang, Zixin Yin, Zhiyang Dou, Xin Chen, Jingbo Wang, Taku Komura, and Lei Zhang. 2025b. Motion2Motion: Cross-topology Motion Transfer with Sparse Correspondence. In SIGGRAPH Asia 2025 Conference Papers. Association for Computing Machinery, 1–11. doi:10.1145/3757377.3763811 Rui Chen, Mingyi Shi, Shaoli Huang, Ping Tan, Taku Komura, and Xuelin Chen. 2024. Taming Diffusion Probabilistic Models for Character Control. In ACM SIGGRAPH 2024 Conference Papers. Association for Computing Machinery, 1–10. doi:10.1145/ 3641519.3657440
12
•
Mou et al.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1–53. Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. 2023a. Objaverse-XL: A Universe of 10M+ 3D Objects. In Advances in Neural Information Processing Systems. 35799–35813. Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. 2023b. Objaverse: A universe of annotated 3D objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13142–13153. Zhiyang Dou, Xuelin Chen, Qingnan Fan, Taku Komura, and Wenping Wang. 2023. C·ASE: Learning Conditional Adversarial Skill Embeddings for Physics-based Characters. In SIGGRAPH Asia 2023 Conference Papers. Association for Computing Machinery, 1–11. doi:10.1145/3610548.3618205 Vijay Prakash Dwivedi and Xavier Bresson. 2021. A generalization of transformer networks to graphs. In AAAI Workshop on Deep Learning on Graphs: Methods and Applications. Ke Fan, Shunlin Lu, Minyue Dai, Runyi Yu, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, and Jingbo Wang. 2025. Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13336–13348. Inbar Gat, Sigal Raab, Guy Tevet, Yuval Reshef, Amit Haim Bermano, and Daniel CohenOr. 2025. AnyTop: Character Animation Diffusion with Any Topology. In ACM SIGGRAPH 2025 Conference Papers. Association for Computing Machinery, 1–10. doi:10.1145/3721238.3730621 Kehong Gong, Zhengyu Wen, Weixia He, Mingxi Xu, Qi Wang, Ning Zhang, Zhengyu Li, Dongze Lian, Wei Zhao, Xiaoyu He, et al. 2025. MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos. arXiv preprint arXiv:2512.10881. Niv Granot, Ben Feinstein, Assaf Shocher, Shai Bagon, and Michal Irani. 2022. Drop the GAN: In Defense of Patches Nearest Neighbors as Single Image Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13460–13469. Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. 2022. Generating Diverse and Natural 3D Human Motions from Text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5152–5161. Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598. Zehuan Huang, Haoran Feng, Yang-Tian Sun, Yuan-Chen Guo, Yan-Pei Cao, and Lu Sheng. 2025. AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models. In SIGGRAPH Asia 2025 Conference Papers. Association for Computing Machinery, 1–13. doi:10.1145/3757377.3763885 Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. 2024. VBench: Comprehensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21807–21818. Yanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang, Weiming Hu, and Jin Gao. 2024. Animate3D: Animating Any 3D Model with Multi-view Video Diffusion. In Advances in Neural Information Processing Systems, Vol. 37. 125879–125906. doi:10.52202/ 079017-3999 Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, and Siyu Tang. 2023. Guided motion diffusion for controllable human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 2151–2162. Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for attention-based permutation-invariant neural networks. In International Conference on Machine Learning. PMLR, 3744– 3753. Sunmin Lee, Taeho Kang, Jungnam Park, Jehee Lee, and Jungdam Won. 2023. SAME: Skeleton-Agnostic Motion Embedding for Character Animation. In SIGGRAPH Asia 2023 Conference Papers. Association for Computing Machinery, 1–11. doi:10.1145/ 3610548.3618206 Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. 2024c. Instant3D: Fast Text-to3D with Sparse-View Generation and Large Reconstruction Model. In International Conference on Learning Representations. Peizhuo Li, Kfir Aberman, Rana Hanocka, Libin Liu, Olga Sorkine-Hornung, and Baoquan Chen. 2021. Learning skeletal articulations with neural blend shapes. ACM Transactions on Graphics 40, 4 (2021), 1–15. doi:10.1145/3450626.3459852 Peizhuo Li, Kfir Aberman, Zihan Zhang, Rana Hanocka, and Olga Sorkine-Hornung. 2022a. GANimator: Neural Motion Synthesis from a Single Sequence. ACM Transactions on Graphics 41, 4 (2022), 1–12. doi:10.1145/3528223.3530157 Peizhuo Li, Sebastian Starke, Yuting Ye, and Olga Sorkine-Hornung. 2024b. WalkTheDog: Cross-Morphology Motion Alignment via Phase Manifolds. In ACM
SIGGRAPH 2024 Conference Papers. Association for Computing Machinery, 1–10. doi:10.1145/3641519.3657508 Siyao Li, Weijiang Yu, Tianpei Gu, Chunze Lin, Quan Wang, Chen Qian, Chen Change Loy, and Ziwei Liu. 2022b. Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11050–11059. Tianyu Li, Calvin Qiao, Guanqiao Ren, KangKang Yin, and Sehoon Ha. 2024a. AAMDM: accelerated auto-regressive motion diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1813–1823. Weiyu Li, Xuelin Chen, Peizhuo Li, Olga Sorkine-Hornung, and Baoquan Chen. 2023. Example-based motion synthesis via generative motion matching. ACM Transactions on Graphics 42, 4 (2023), 1–12. doi:10.1145/3592395 Xuan Li, Qianli Ma, Tsung-Yi Lin, Yongxin Chen, Chenfanfu Jiang, Ming-Yu Liu, and Donglai Xiang. 2025. Articulated kinematics distillation from video diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference. 17571– 17581. Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang, Zhangyang Wang, Konstantinos N Plataniotis, Yao Zhao, and Yunchao Wei. 2024. Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models. In Advances in Neural Information Processing Systems, Vol. 37. 110854–110875. doi:10.52202/079017-3519 Derek Lim, Joshua David Robinson, Lingxiao Zhao, Tess Smidt, Suvrit Sra, Haggai Maron, and Stefanie Jegelka. 2023. Sign and basis invariant networks for spectral graph representation learning. In International Conference on Learning Representations. Isabella Liu, Zhan Xu, Wang Yifan, Hao Tan, Zexiang Xu, Xiaolong Wang, Hao Su, and Zifan Shi. 2025. RigAnything: Template-Free Autoregressive Rigging for Diverse 3D Assets. ACM Transactions on Graphics 44, 4 (2025), 1–12. doi:10.1145/3731149 Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Mukund Varma T, Zexiang Xu, and Hao Su. 2023c. One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without PerShape Optimization. In Advances in Neural Information Processing Systems, Vol. 36. 22226–22246. Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023b. Zero-1-to-3: Zero-shot One Image to 3D Object. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 9298–9309. Siqi Liu, Maoyu Wang, Bo Dai, and Cewu Lu. 2026. PALUM: Part-based Attention Learning for Unified Motion Retargeting. arXiv preprint arXiv:2601.07272. Xingchao Liu, Chengyue Gong, and Qiang Liu. 2023a. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. In International Conference on Learning Representations. Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2015. SMPL: A Skinned Multi-Person Linear Model. ACM Transactions on Graphics 34, 6 (2015), 248:1–248:16. doi:10.1145/2816795.2818013 Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. Shunlin Lu, Jingbo Wang, Zeyu Lu, Ling-Hao Chen, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai, and Ruimao Zhang. 2025. ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model. In Proceedings of the Computer Vision and Pattern Recognition Conference. 27872–27882. Yanzhe Lyu, Chen Geng, Karthik Dharmarajan, Yunzhi Zhang, Hadi Alzayer, Shangzhe Wu, and Jiajun Wu. 2026. Choreographing a World of Dynamic Objects. arXiv preprint arXiv:2601.04194. Nadia Magnenat-Thalmann, Richard Laperrière, and Daniel Thalmann. 1988. Jointdependent local deformations for hand animation and object grasping. In Proceedings of Graphics Interface. 26–33. Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black. 2019. AMASS: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 5442–5451. Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al. 2021. Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning. arXiv preprint arXiv:2108.10470. Zichong Meng, Zeyu Han, Xiaogang Peng, Yiming Xie, and Huaizu Jiang. 2025. Absolute coordinates make motion generation easy. arXiv preprint arXiv:2505.19377. Linzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu, and Kostas Daniilidis. 2025. DIMO: Diverse 3D Motion Generation for Arbitrary Objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 14357–14368. Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. 2022. Point-E: A System for Generating 3D Point Clouds from Complex Prompts. arXiv preprint arXiv:2212.08751. Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. 2019. Expressive Body Capture: 3D Hands, Face, and Body from a Single Image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10975–10985. William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 4195–4205.
UniMate: One Unified Model to Animate Diverse Skeletons Mathis Petrovich, Michael J Black, and Gül Varol. 2021. Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 10985–10995. Matthias Plappert, Christian Mandery, and Tamim Asfour. 2016. The KIT MotionLanguage Dataset. Big Data 4, 4 (2016), 236–252. Sigal Raab, Inbar Gat, Nathan Sala, Guy Tevet, Rotem Shalev-Arkushin, Ohad Fried, Amit Haim Bermano, and Daniel Cohen-Or. 2024a. Monkey see, monkey do: Harnessing self-attention in motion diffusion for zero-shot motion transfer. In SIGGRAPH Asia 2024 Conference Papers. Association for Computing Machinery, 1–13. doi:10.1145/3680528.3687579 Sigal Raab, Inbal Leibovitch, Peizhuo Li, Kfir Aberman, Olga Sorkine-Hornung, and Daniel Cohen-Or. 2023. MoDi: Unconditional Motion Synthesis from Diverse Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13873–13883. Sigal Raab, Inbal Leibovitch, Guy Tevet, Moab Arar, Amit H Bermano, and Daniel Cohen-Or. 2024b. Single motion diffusion. In International Conference on Learning Representations. Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini. 2022. Recipe for a general, powerful, scalable graph transformer. Advances in Neural Information Processing Systems 35 (2022), 14501– 14515. Davis Rempe, Mathis Petrovich, Ye Yuan, Haotian Zhang, Xue Bin Peng, Yifeng Jiang, Tingwu Wang, Umar Iqbal, David Minor, Michael de Ruyter, et al. 2026. Kimodo: Scaling Controllable Human Motion Generation. arXiv preprint arXiv:2603.15546. Remy Sabathier, David Novotny, Niloy J Mitra, and Tom Monnier. 2026. ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion. arXiv preprint arXiv:2601.16148. Haim Sawdayee, Chuan Guo, Guy Tevet, Bing Zhou, Jian Wang, and Amit H Bermano. 2026. Dance like a chicken: Low-rank stylization for human motion diffusion. Computer Graphics Forum (2026), e70365. doi:10.1111/cgf.70365 Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. 2024. Human motion diffusion as a generative prior. In International Conference on Learning Representations. Yahao Shi, Yang Liu, Yanmin Wu, Xing Liu, Chen Zhao, Jie Luo, and Bin Zhou. 2025. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video. arXiv preprint arXiv:2506.07489. Chaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin, and Jianfeng Zhang. 2025a. Puppeteer: Rig and Animate Your 3D Models. arXiv preprint arXiv:2508.10898. Chaoyue Song, Jianfeng Zhang, Xiu Li, Fan Yang, Yiwen Chen, Zhongcong Xu, Jun Hao Liew, Xiaoyang Guo, Fayao Liu, Jiashi Feng, et al. 2025b. MagicArticulate: Make Your 3D Models Articulation-Ready. In Proceedings of the Computer Vision and Pattern Recognition Conference. 15998–16007. Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing 568 (2024), 127063. Keqiang Sun, Dor Litvak, Yunzhi Zhang, Hongsheng Li, Jiajun Wu, and Shangzhe Wu. 2024. Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos. In European Conference on Computer Vision. Springer, 100–119. Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit H Bermano, and Michiel van de Panne. 2025. CLoSD: Closing the Loop between Simulation and Diffusion for Multi-Task Character Control. In International Conference on Learning Representations. Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. 2023. Human motion diffusion model. In International Conference on Learning Representations. Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012. MuJoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, 5026–5033. Truebones. 2022. Truebones Zoo Dataset. https://truebones.gumroad.com/. Lukas Uzolas, Elmar Eisemann, and Petr Kellnhofer. 2025. MotionDreamer: Exploring Semantic Video Diffusion features for Zero-Shot 3D Mesh Animation. In International Conference on 3D Vision. Weilin Wan, Zhiyang Dou, Taku Komura, Wenping Wang, Dinesh Jayaraman, and Lingjie Liu. 2024. TLControl: Trajectory and Language Control for Human Motion Synthesis. In European Conference on Computer Vision. Springer, 37–54. Wan Team, Alibaba Group. 2025. Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314. Miaowei Wang, Qingxuan Yan, Zhi Cao, Yayuan Li, Oisin Mac Aodha, Jason J Corso, and Amir Vaxman. 2026b. BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation. arXiv preprint arXiv:2602.18873. Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, and Baining Guo. 2023. RODIN: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4563–4573. Xuan Wang, Kai Ruan, Liyang Qian, Zhizhi Guo, Chang Su, and Gaoang Wang. 2026a. X-MoGen: Unified Motion Generation Across Humans and Animals. Proceedings
•
13
of the AAAI Conference on Artificial Intelligence 40, 12 (2026), 10234–10242. doi:10. 1609/aaai.v40i12.37992 Xuan Wang, Kai Ruan, Xing Zhang, and Gaoang Wang. 2025. AniMo: Species-Aware Model for Text-Driven Animal Motion Generation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 1929–1939. Yuxin Wen, Qing Shuai, Di Kang, Jing Li, Cheng Wen, Yue Qian, Ningxin Jiao, Changhai Chen, Weijie Chen, Yiran Wang, et al. 2025. HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation. arXiv preprint arXiv:2512.23464. Zijie Wu, Chaohui Yu, Fan Wang, and Xiang Bai. 2025. AnimateAnyMesh: A FeedForward 4D Foundation Model for Text-Driven Universal Mesh Animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13557– 13568. Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. 2025. Structured 3D Latents for Scalable and Versatile 3D Generation. In Proceedings of the Computer Vision and Pattern Recognition Conference. 21469–21480. Tianyi Xie, Yunuo Chen, Yaowei Guo, Yin Yang, Bolei Zhou, Demetri Terzopoulos, Ying Jiang, and Chenfanfu Jiang. 2025. AnimaMimic: Imitating 3D Animation from Video Priors. arXiv preprint arXiv:2512.14133. Zhan Xu, Yang Zhou, Evangelos Kalogerakis, Chris Landreth, and Karan Singh. 2020. RigNet: Neural Rigging for Articulated Characters. ACM Transactions on Graphics 39, 4 (2020), 1–14. doi:10.1145/3386569.3392379 Zhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen, Ziwei Xuan, Mitch Hill, Junjie Bai, Guo-Jun Qi, and Yalin Wang. 2024. OmniMotionGPT: Animal Motion Generation with Limited Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1249–1259. Jiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou, Sergey Tulyakov, Raymond A Yeh, Peter Wonka, and Chaoyang Wang. 2025. ShapeGen4D: Towards High Quality 4D Shape Generation from Videos. arXiv preprint arXiv:2510.06208. Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems 34 (2021), 28877– 28888. Kwan Yun, Seokhyeon Hong, Chaelin Kim, and Junyong Noh. 2025. AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models. In Proceedings of the Computer Vision and Pattern Recognition Conference. 27838–27848. Bowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao, Dong Chen, and Baining Guo. 2025c. Gaussian Variation Field Diffusion for High-Fidelity Video-to4D Synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 12502–12513. Jia-Peng Zhang, Cheng-Feng Pu, Meng-Hao Guo, Yan-Pei Cao, and Shi-Min Hu. 2025b. One Model to Rig Them All: Diverse Skeleton Rigging with UniRig. ACM Transactions on Graphics 44, 4 (2025), 1–18. doi:10.1145/3730930 Xinyi Zhang, Naiqi Li, and Angela Dai. 2025a. DNF: Unconditional 4D Generation with Dictionary-based Neural Fields. In Proceedings of the Computer Vision and Pattern Recognition Conference. 26047–26056. Kaifeng Zhao, Gen Li, and Siyu Tang. 2025b. DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control. In International Conference on Learning Representations. Qingqing Zhao, Peizhuo Li, Yifan Wang, Olga Sorkine-Hornung, and Gordon Wetzstein. 2024. Pose-to-Motion: Cross-Domain Motion Retargeting with Pose Prior. Computer Graphics Forum 43, 8 (2024), e15170. doi:10.1111/cgf.15170 Zibo Zhao, Zeqiang Lai, Qingxiang Lin, Yunfei Zhao, Haolin Liu, Shuhui Yang, Yifei Feng, Mingxin Yang, Sheng Zhang, Xianghui Yang, et al. 2025a. Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation. arXiv preprint arXiv:2501.12202. Matt Zhou, Ruining Li, Xiaoyang Lyu, Zhaomou Song, Zhening Huang, Chuanxia Zheng, Christian Rupprecht, Andrea Vedaldi, and Shangzhe Wu. 2026. Articraft: An Agentic System for Scalable Articulated 3D Asset Generation. arXiv preprint arXiv:2605.15187. Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu. 2024. EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation. In European Conference on Computer Vision. Springer, 18–38. Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019. On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5745–5753. Silvia Zuffi, Angjoo Kanazawa, David W Jacobs, and Michael J Black. 2017. 3D Menagerie: Modeling the 3D shape and pose of animals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6365–6373.
14
•
Mou et al.
APPENDIX
Code and data availability. The public project repository is available at https://github.com/Friedrich-M/UniMate. The repository currently serves as the permanent project entry point; upon publication, we will release the complete training and data-preprocessing code, pretrained model checkpoints, and the UniML3D dataset at this URL. Interactive demo. An interactive demo with more results is available at https://linzhanmou.com/unimate/interactive.
A Dataset Construction A.1 Data Sources and Statistics UniML3D aggregates three complementary motion sources. Truebones [Truebones 2022] contributes 1,094 animal motion sequences spanning 74 distinct skeletons. Mixamo [Adobe 2022] adds 2,425 high-quality human sequences. From Objaverse-XL [Deitke et al. 2023a,b] we filter and deduplicate 6,965 articulated assets with both rigging [Song et al. 2025] and animation [Liang et al. 2024] annotations, yielding 9,487 valid action sequences. In total, UniML3D comprises 13,006 sequences and 2,140,232 frames over thousands of unique skeletons. Algorithm 1 lists the per-stage operations of the curation pipeline summarized in the main paper. Figures 18 and 19 visualize the distribution of UniML3D across morphology categories and joint counts. The category distribution is heavily long-tailed: bipedal characters dominate at 84.1%, while serpentine (0.3%) and marine (0.8%) categories are sparsely populated, which motivates the square-root-balanced sampler defined in Section A.6. The joint-count distribution further shows that the dataset spans skeletons of widely varying complexity, with concentrations around the typical humanoid rigging conventions.
A.2
Joint Name Standardization
Raw joint labels in our source datasets are highly inconsistent: they mix DCC-tool namespaces (mixamorig:, QuickRigCharacter_, Bip01_), Maya/Blender suffixes (_jnt, .001), chain indices (Spine1, Thumb1), 3ds Max Biped numeric finger codes (Finger01, Finger21), Japanese romaji roots used by some animal rigs (momo,
Fig. 18. Morphology distribution of UniML3D. Share of motion sequences per skeletal morphology; bipedal characters dominate, while serpentine and marine rigs are rare. Distribution of Joint Counts across Skeletons median 38
mean 40.4
Kernel density Histogram
12%
Proportion of skeletons
This appendix supplies material that the main paper defers for space. Section A covers the dataset statistics, the LLM system prompts used for joint-name standardization, facing-direction joint-pair selection, and motion captioning, the four-view rendering camera setup, and the online skeletal augmentation pipeline. Section B reports the architectural and training/inference hyperparameters for UniMate and the full expressions for the three training losses. Section C gives a self-contained theoretical analysis of Spec-RoPE, covering its relative-coordinate structure, permutation equivariance, connection to standard RoPE, and effective-resistance interpretation. Section D details our evaluation protocols, baselines, and the userstudy setup, and Section E collects additional results: further AnyTop comparisons, direct comparisons with the skeleton-free baselines, qualitative ablations of the architecture and of the dataset-curation stages, and the foot-sliding analysis. Finally, Section F expands the discussion of limitations and future work.
10%
8%
5%
2%
0%
15
30
45
60
75
90
Number of joints
Fig. 19. Joint-count distribution across skeletons. Histogram and kernel density of the number of joints per retained skeleton in UniML3D (median 38, mean 40.4); the bimodal spread shows that the dataset covers rigs of widely varying complexity.
munabire), and idiosyncratic per-asset placeholders. ObjaverseXL exports additionally append a per-joint index to every name, sometimes on top of the original chain index (Hip_01_41, Index.R.001_013_6). To make the per-joint name embeddings used by UniMate comparable across rigs, we standardize every joint label to a fixed anatomical vocabulary using DeepSeek-V4Flash [DeepSeek-AI 2026] with the deterministic system prompt below. Each rig is processed as a whole, so the model can resolve a joint’s role from its neighbors in the hierarchy, and every input joint must map to exactly one label. Responses that violate this oneto-one correspondence are rejected and re-queried, and rigs that repeatedly fail fall back to a deterministic rule-based cleaner over the same vocabulary. An optional refinement pass with GPT-5 [OpenAI 2025] then re-examines each label alongside its raw name and corrects residual errors. Standardize 3D rig joint names to canonical anatomical labels. Inputs come from Mixamo, Maya, Blender, Unreal, Truebones and custom rigs. Focus on semantic meaning, not surface syntax.
UniMate: One Unified Model to Animate Diverse Skeletons
•
15
Algorithm 1 UniML3D data processing. The three stages of Section 4 act at three granularities: skeleton-level steps run once per rig, clip-level steps once per clip, and feature normalization once over the whole set. A clip either passes every check or is discarded. Require: rigged assets from Truebones, Mixamo, and Objaverse-XL, each a skeleton S with a set of animation clips X Ensure: the training set of canonical skeleton–motion–caption samples (S, 𝑥, 𝑐) 1: for all rigged assets (S, X) in the source pool do Skeleton-based filtering, skeleton level (Section 4) 2: S ← the tree with the largest cumulative skinning weight; realign a spurious root onto the semantic root via forward kinematics; prune zero-weight helper joints Semantic annotation, skeleton level (Sections A.2 and A.3) 3: NS ← joint names standardized to the anatomical vocabulary by an LLM 4: ( 𝑗L, 𝑗R ) ← the symmetric joint pair defining the lateral axis, selected by an LLM from NS 5: for all clips 𝑥 ∈ X do Skeleton-based filtering, clip level (Section 4) 6: if 𝑥 is static or physically implausible under the criteria of Section 4 then 7: discard 𝑥 and continue with the next clip 8: end if Semantic annotation, clip level (Sections A.4 and A.5) 9: resample 𝑥 to 30 FPS; render four synchronized views of 𝑥 and of the rest pose; 𝑐 ← caption of the rendering by a multimodal LLM Motion canonicalization, clip level (Sections 3.1 and 4) 10: serialize S in breadth-first order from the root; scale S and 𝑥 by the topology diameter 𝑑 topo 11: place 𝑥 in a 𝑦-up frame with the initial root at the origin; align the initial facing direction f, derived from the 𝑗L → 𝑗R direction as in Section 4, with +𝑧 12: express joint rotations relative to the rest pose; assemble the per-joint features m𝑡𝑗 = [p𝑡𝑗 , r𝑡𝑗 , v𝑡𝑗 ]; add (S, 𝑥, 𝑐) to the training set 13: end for 14: end for Motion canonicalization, dataset level (Section 4) 15: normalize root features with (𝜇 global , 𝜎global ) and non-root features with (𝜇 local , 𝜎local ), both computed over the whole training set OUTPUT: a single JSON array of strings, same length and order as the input. No prose, no fences, no extra text. One input entry -> one output entry; never dedupe, merge, skip or reorder, even when neighbours produce identical labels. CLEAN each name by removing rig noise and extracting meaning: 1. Digits and Blender '.NNN' counters are meaningless – rig-internal bookkeeping (chain index, mirror id, duplicate counter). Ignore them when matching, and never include a digit, underscore, or dot in the final label. 'Spine', 'Spine1', 'Spine_02', 'Spine.003' all map to 'Spine'. Finger-chain segments ('Thumb1/Thumb2/Thumb3') all map to 'Thumb Finger'. The Objaverse export pipeline also stamps a trailing global index on EVERY name ('_NN' or '_0NN': '_01', '_010', '_063'); indices can STACK ('Hip_01_41', 'Head_1_016', 'Index.R.001_013_6', 'Bone.001_01', 'Spine_1_013') – strip ALL of them, in any position. Editor decorations are noise too: strip '(mirrored)' anywhere and a trailing '.x' center marker ('spine_01.x' -> 'Spine'). 2. Drop any leading '<Word>:' namespace (case-insensitive, trailing digits in the namespace OK): mixamorig:, Mixamorig1:, Mutant:, Sif:. The same words are noise without the colon too – leading 'mixamorig_'/'Mixamorig'/'Character'/'Rig'/'QuickRigCharacter_' segments are dropped, never echoed into the label. The same goes for embedded ASSET/CHARACTER names and their decorations – 'rp_karl_animated_006_warmingUp_spine_01' -> 'Spine', 'CMan0205-M4-CS_Hips L Finger0' -> 'Left Thumb Finger': keep only the anatomical tokens, drop every name-like or counter token. 3. Drop rig prefixes (match before ignoring digits so 'Bip01_' still strips cleanly): any Bip<digits> container with any separator (Bip01_, Bip002 , Bip01-, and separator-free Bip001LFinger0), BN_Bip01_, BN_, Bn_, NPC_, jt_, Elk, Sabrecat_, QuickRigCharacter_, Bind_, Skeleton_, Root_, DEF-, def_. 4. Drop Maya suffixes: _jnt, _jt, _Jt, _JNT, _joint, _bone, _bn, _C. 5. Extract side as explicit 'Left '/'Right ' prefix. Recognise: prefix L_ / R_, Lt_ / Rt_, Left_ / Right_, Left<UpperWord> / Right<UpperWord> (LeftHand, RightArm), L<UpperLetter> / R<UpperLetter> (LArm, RHand) – but NOT when followed by a lowercase letter (Lower, Ribcage). 3ds Max Biped space-separated single-letter side: 'Bip001 L UpperArm' -> Left Upper Arm, 'Bip001 R Thigh' -> Right Thigh. Token boundaries are spaces, not underscores. suffix _L / _R / _l / _r, .L / .R, Japanese trailing L/R.
PRECEDENCE: a trailing .L/.R/_L/_R token OVERRIDES a leading side word – Blender's symmetrize renames only the suffix, leaving the prefix text stale: 'mixamorig:LeftShoulder.R' -> 'Right Shoulder', 'r_toe.L' -> 'Left Toe'. quadruped F_/B_ = Front/Back (e.g. 'F_R_Shoulder' -> 'Right Front Shoulder', 'B_L_Foot' -> 'Left Back Foot'). 6. Common body-part roots – translate case-insensitively, accepting Unreal snake_case AND CamelCase short-forms: pelvis/spine/neck/head/jaw/eye -> Pelvis/Spine/Neck/Head/Jaw/Eye; clavicle/collar -> Shoulder; upperarm / UpperArm / UpArm -> Upper Arm; lowerarm / LowerArm / LowArm / ForeArm -> Forearm; hand -> Hand; thigh / upleg / UpLeg -> Thigh; calf / lowleg / LowLeg / lowerleg -> Shin; upperleg -> Thigh; toes -> Toe; in a mixamo chain (UpLeg -> Leg -> Foot) the mid-bone 'Leg' is the shin -> Shin; foot -> Foot; toebase / Toe0 / toe -> Toe; ball -> Toe; eyelid -> Eyelid. Finger roots index/middle/ring/pinky/thumb -> '<Root> Finger'. *_twist -> '<Root> Twist'. 6b. 3ds Max Biped fingers use numeric codes: Finger0*=Thumb, Finger1*=Index, Finger2*=Middle, Finger3*=Ring, Finger4*=Pinky. Any trailing digits after that code are chain position – ignore. 'Bip001 L Finger0' and 'Bip001 L Finger01' both -> 'Left Thumb Finger'; 'Bip001 R Finger21' -> 'Right Middle Finger'; 'Bip001 R Toe0' -> 'Right Toe'. 6c. Mocap-segmented fingers are 1-BASED and carry a segment word: Finger1..Finger5 + Metacarpal/Proximal/Medial/Distal/Tip, with Finger1=Thumb ... Finger5=Pinky. 'LeftFinger1Metacarpal' -> 'Left Thumb Finger'; the Tip segment -> '<Root> Finger End'. Rule 6b's 0-based codes apply only to BARE Finger<digit> names with no segment word. Segment words after a NAMED finger ('IndexDistal', 'thumb_proximal_l', 'RingIntermediate') are likewise chain position – drop them: 'IndexDistal' -> 'Index Finger'. 7. Other direction words: Top->Upper, Low->Lower ('Topjaw'->'Upper Jaw'). 'HeadTop_End' / '*_End' / '*Nub' -> '<Root> End'. Animal 'Hair*/Mane*' -> 'Mane'; humanoid accessories 'Ponytail*/Cape*/Cloth*/Skirt*' -> 'Appendage'. 8. Placeholders -> 'Bone': Bone, joint, Xtra*, MagicEffectsNode, and any token with no clear anatomy (meshok, Capuche, ...). Do NOT fabricate body parts. EXCEPTION: a name that is ONLY digits, optionally with a leading underscore ('_00', '12'), is copied through UNCHANGED – it marks a rig with unnamed bones. 9. Japanese roots (Alligator/Pirrana/Tukan): body momo=Thigh, hiza=Knee, ashi=Foot, hiji=Elbow, te=Hand, kata=Shoulder, mune=Chest, hara=Abdomen, koshi/kosi=Hips, kubi=Neck, atama/kao=Head, ago=Jaw
16
•
Mou et al.
tail sippo/shippo=Tail, o=Tail fish munabire=Pectoral Fin, harabire=Pelvic Fin, sebire=Dorsal Fin, obire=Caudal Fin, shiribire=Anal Fin, era=Gill Trailing L/R on any of these -> Left/Right prefix. 10. Output Title Case, single spaces. Prefer the CANONICAL vocabulary; if nothing fits, use 'Bone'. ALIASES & TYPOS (map to the canonical term): Spline=Spine, Scull=Skull, Nek=Neck, Tai=Tail, Tone/Thouge/Tunge=Tongue, Eyeleds=Eyelid, HorseLink=Fetlock, LargeCannon=Cannon, PhalanxPrima=Pastern, PhalangesManus=Phalanges, Foreleg=Front Leg, Hindleg=Hind Leg, Digit=Finger, Hair=Mane, Little=Pinky (finger), locator/Trajectory/Cog=Root, Clavicle/Collarbone=Shoulder. Species terms: insect 'Clip'/'Shall'=Mandible, 'Pliers'/'Piers'=Pincer, cricket 'Feeler'=Antenna but fish 'Feelers'=Barbel, bird 'ponitail'=Crest. CANONICAL VOCABULARY (use these exact terms verbatim): Core: Pelvis, Hips, Spine, Ribcage, Neck, Head, Skull, Skull Base, Head End, Body, Upper Body, Lower Body, Chest, Abdomen, Waist, Collar, Hip, Belly, Root, Center Arm : Shoulder, Scapula, Arm, Upper Arm, Forearm, Elbow, Wrist, Hand, Palm Leg : Thigh, Shin, Leg, Knee, Ankle, Foot, Heel, Toe, Paw, Hoof, Fetlock, Cannon, Metacarpus, Phalanges, Pastern Finger: Finger, Thumb Finger, Index Finger, Middle Finger, Ring Finger, Pinky Finger Head: Jaw, Upper Jaw, Lower Jaw, Tongue, Ear, Eye, Eyeball, Eyebrow, Eyelid, Mouth, Lip, Upper Lip, Lower Lip, Nose, Muzzle, Chin, Cheek Appendage: Tail, Wing, Feather, Antenna, Barbel, Tentacle, Claw, Hand Claw, Fang, Mandible, Large Mandible, Lower Mandible, Pincer, Stinger, Appendage Fins: Fin, Pectoral Fin, Pelvic Fin, Dorsal Fin, Caudal Fin, Anal Fin, Gill Coat/Equipment: Mane, Fur, Whisker, Shell, Dorsal Plate, Crest, Horn, Reins, Halter Quadruped: Front Leg, Middle Leg, Hind Leg, Front Paw, Back Paw, Front Hoof, Rear Hoof, Front Shoulder, Back Hip Physics: Twist, Upper Arm Twist, Forearm Twist, Thigh Twist, Shin Twist, Thigh Muscle, Neck Muscle, Tail Twist, Jiggle, Handle, IK Chain, Bone Side goes first ('Left Upper Arm', 'Right Pinky Finger'); quadruped markers sit between side and part ('Right Front Shoulder'). COMPOSED labels are also valid: '<part> End' for chain tips/Nub bones, 'Inner/Middle/Outer <part>' (raptor toes, claws, fingers), 'Upper/Lower Eyelid', 'Upper/Lower Left/Right/Front Lip', 'Wrist Back', 'Elbow Back', 'Front/Middle/Hind Leg End'. EXAMPLES (one per rig family): mixamorig:LeftUpLeg -> Left Thigh Mutant:RightHandThumb1 -> Right Thumb Finger Sif:calf_twist_01_r -> Right Shin Twist index_01_l -> Left Index Finger Bip01_L_Thigh -> Left Thigh BN_Bip01_R_Forearm_03 -> Right Forearm NPC_L_Finger02 -> Left Finger Elk_RearHoof_L -> Left Rear Hoof jt_FrontLeg1_R_C -> Right Front Leg Lt_Thumb1_jt -> Left Thumb Finger R_toeBase_jnt -> Right Toe Eye.R.001 -> Right Eye mixamorig:LeftShoulder.R -> Right Shoulder LeftHandIndex1 -> Left Index Finger mixamorig:RightLeg -> Right Shin LeftFinger2Distal -> Left Index Finger BN_Spline_03 -> Spine Bone.001 -> Bone F_R_Shoulder -> Right Front Shoulder B_L_Foot -> Left Back Foot Topjaw -> Upper Jaw momoR -> Right Thigh munabireL -> Left Pectoral Fin joint12 / Xtra01 / meshok -> Bone OBJAVERSE (trailing global '_NN', often stacked with a chain index): mixamorig:LeftUpLeg_056 -> Left Thigh mixamorig:LeftHandThumb1_012 -> Left Thumb Finger QuickRigCharacter_LeftForeArm_014 -> Left Forearm Bip001 L UpperArm_07 -> Left Upper Arm Bip001 L Finger01_011 -> Left Thumb Finger Bip001 R Finger21_036 -> Right Middle Finger Bip001 R Toe0_054 -> Right Toe UpArm.R_010_15 -> Right Upper Arm LowLeg.L_038_39 -> Left Shin Hip_01_41 -> Pelvis Index.R.001_013_6 -> Right Index Finger Rt_Eyelid_jt_08 -> Right Eyelid Skeleton_Root_02 -> Root joint1_2 / Bone.001_01 -> Bone BATCH EXAMPLE (notice identical outputs are preserved, not merged):
input : ["Spine", "Spine1", "Spine2", "Spine3", "Neck", "Tail_01", "Tail_02", "Tail_03"] output : ["Spine", "Spine", "Spine", "Spine", "Neck", "Tail", "Tail", "Tail"] WRONG : ["Spine", "Neck", "Tail"] (8 inputs must yield 8 outputs – deduping is a failure)
A.3
Facing-Direction Joint-Pair Selection
The canonicalization stage of the main paper aligns each clip’s initial facing direction with the positive 𝑧-axis using the left-to-right direction of a bilaterally symmetric joint pair. Which pair plays that role differs from rig to rig (the thighs of a humanoid, the front shoulders of a quadruped, the pectoral fins of a fish), so we select it per rig with DeepSeek-V4-Flash, reasoning over the standardized joint names of Section A.2. A deterministic rule-based resolver first proposes a pair from a fixed priority order of body parts; the model then reviews the candidate joints (right-side, left-side, head, and tail) together with that proposal and follows a three-step decision. It pairs the right and left joints that mirror the highest-priority body part present on both sides; if no mirrored pair exists but head and tail joints do, as for serpentine rigs, it returns the head and tail chain endpoints as a longitudinal body axis instead; otherwise it returns no pair, which marks rigs without a meaningful lateral axis, such as vehicles and props. Invalid responses fall back to the rule-based proposal, and an optional GPT-5 refinement pass revisits the rigs left without a pair. The system prompt follows. # TASK Pick the joint pair that defines one 3D rig's lateral facing axis. Rigs come from Objaverse (Mixamo, 3ds Max Biped, Maya QuickRig, Unreal mannequin, CAT, custom) and Truebones animals. All clean names are already normalised to a canonical vocabulary – reason on CLEAN names; RAW names are only for verbatim copy-back. # INPUT (user message) Four pre-filtered buckets (any may be empty): RIGHT : rows whose clean name starts with 'Right ' LEFT : rows whose clean name starts with 'Left ' HEAD : rows whose clean name (minus any trailing ' N') is one of {Head, Skull, Skull Base, Head End, Jaw, Upper Jaw, Lower Jaw, Tongue, Muzzle, Nose, Chin} TAIL : rows whose clean name (minus any trailing ' N') is 'Tail' or 'Tail Twist' Plus a 'Rule-based hint' JSON object – the deterministic resolver's pick. Use it as a sanity check, not ground truth (see HINT below). # ALGORITHM Execute the steps in order. Do not skip ahead. STEP 1. Compute OVERLAP. - For each row in RIGHT, suffix = clean_name minus 'Right ' minus any trailing ' <digits>'. Collect into set R_SUFFIXES. - For each row in LEFT, suffix = clean_name minus 'Left ' minus any trailing ' <digits>'. Collect into set L_SUFFIXES. - OVERLAP = R_SUFFIXES ∩ L_SUFFIXES. STEP 2. If OVERLAP is non-empty – APPLY RULE A AND RETURN. a. Pick suffix `s` = first element of OVERLAP that appears in PRIORITY (see below). If none appear, pick the alphabetically first element of OVERLAP. b. R_ROWS = RIGHT rows whose suffix == s; L_ROWS = LEFT rows whose suffix == s. c. If len(R_ROWS)==1 and len(L_ROWS)==1, pair them. Otherwise (chain duplicates, e.g. scorpion legs) match by DIGIT SIGNATURE: the tuple of ALL digit groups in the RAW name ('Bip01_R_Thigh_4' -> (01,4)). Tier 1 – pair rows whose full signatures are equal. Tier 2 – no tier-1 match: drop the LAST group (the Objaverse per-joint global index: 'Bip01_R_Thigh_1_053' and 'Bip01_L_Thigh_1_054' both reduce to (01,1)) and pair on the rest. If neither tier matches, take R_ROWS[0] and L_ROWS[0]. d. Output: source = s.lower(); body_axis = false. e. DO NOT consider HEAD or TAIL. Even if HEAD+TAIL look perfect, rule A wins whenever OVERLAP is non-empty. STEP 3. OVERLAP is empty. If HEAD is non-empty AND TAIL is non-empty – APPLY RULE B AND RETURN. - r_hip = LAST HEAD row; l_hip = LAST TAIL row (chain endpoints).
UniMate: One Unified Model to Animate Diverse Skeletons
- source = 'body_axis'; body_axis = true. STEP 4. Otherwise – APPLY RULE C (empty). - Both hips are {raw: '', clean: ''}. source = 'empty'; body_axis = false. - Vehicles, props, and abstract rigs land here. NEVER pair an unrelated Right/Left entry (e.g. lone 'Right Eye' without left mate) just to avoid emitting empty. # PRIORITY (highest -> lowest; first match in OVERLAP wins) Hip-level : thigh, shoulder, front shoulder, back hip, hip, scapula Whole-limb : upper arm, arm, front leg, hind leg, middle leg, back leg, wing, leg Aquatic : pectoral fin, pelvic fin, fin, gill Arthropod : pincer, mandible, large mandible, lower mandible, stinger, claw, hand claw, antenna Mid-limb : forearm, shin, knee, elbow, ankle, wrist Extremity : hand, palm, foot, heel, front paw, back paw, paw, front hoof, rear hoof, hoof, fetlock, cannon, metacarpus, pastern, toe Fingers : thumb finger, index finger, middle finger, ring finger, pinky finger, finger, neck Weak : eye, eyeball, eyelid, eyebrow, ear, horn, cheek, whisker, fang, barbel, tentacle, feather Last : tail Suffix comparison is case-insensitive but uses the clean-name form ('Upper Arm', 'Front Shoulder', 'Pectoral Fin'). Read the list above LITERALLY, left to right, top to bottom – it is exactly the order the rule-based resolver uses, so do not re-rank quadruped markers ('Front Shoulder', 'Back Hip', 'Front Leg') against their plain counterparts ('Shoulder', 'Hip', 'Leg'). Note in particular that 'shoulder' comes BEFORE 'front shoulder', while 'front leg' comes BEFORE 'leg'. # HINT The hint is the rule-based resolver's output. It is usually right but can err on edge cases: - Hint says 'body_axis' or 'empty' yet OVERLAP is non-empty -> hint is wrong, apply STEP 2 (rule A wins). - Hint swapped sides (Right/Left in wrong fields) -> fix. - Hint chose a lower-priority suffix than STEP 2a finds -> override with the higher-priority one. - Hint paired mismatched chain indices on a multi-leg rig -> fix via STEP 2c raw-suffix match. - Your algorithm yields the same answer as the hint -> return the hint verbatim. # OUTPUT Exactly one JSON object, no prose, no markdown fences, no comments: {"r_hip": {"raw": "<exact>", "clean": "<exact>"}, "l_hip": {"raw": "<exact>", "clean": "<exact>"}, "source": "<lowercase suffix or 'body_axis' or 'empty'>", "body_axis": <true|false>} # INVARIANTS (verify before emitting) I1. r_hip.raw != l_hip.raw, UNLESS both are '' (empty case). I2. r_hip.raw and l_hip.raw each appear verbatim in the rig's input rows (copied character-for-character). I3. r_hip always holds the Right (or head) side; l_hip always holds the Left (or tail) side. Never swap. I4. body_axis == true IFF source == 'body_axis'. I5. source is lowercase. If rule A fired, source is the suffix lowercased with single spaces (e.g. 'thigh', 'front shoulder', 'pectoral fin'). I6. Unless body_axis, r_hip and l_hip mirror the SAME part: their clean names minus the side prefix and any trailing digits are identical. Never pair different parts ('Right Thigh' with 'Left Shoulder' is invalid even though the sides are correct). # WORKED EXAMPLES ## Ex1 – Mixamo humanoid (rule A, hip-level pick) RIGHT has 'Right Thigh' + 'Right Shoulder'; LEFT has the mirrors. OVERLAP = {Thigh, Shoulder}. Thigh outranks Shoulder -> pick Thigh. {"r_hip": {"raw": "mixamorig:RightUpLeg_056", "clean": "Right Thigh"}, "l_hip": {"raw": "mixamorig:LeftUpLeg_056", "clean": "Left Thigh"}, "source": "thigh", "body_axis": false} ## Ex2 – Objaverse quadruped (Front Shoulder beats Back Hip) No Thigh pair; OVERLAP = {Front Shoulder, Back Hip}. Front Shoulder outranks Back Hip. {"r_hip": {"raw": "F_R_Shoulder_012", "clean": "Right Front Shoulder"}, "l_hip": {"raw": "F_L_Shoulder_012", "clean": "Left Front Shoulder"}, "source": "front shoulder", "body_axis": false} ## Ex3 – Scorpion multi-leg (STEP 2c raw-suffix match) RIGHT has three 'Right Thigh' rows (Bip01_R_Thigh_1/_2/_4); LEFT the mirrors. OVERLAP = {Thigh}. Digit signatures: (01,1) on both sides -> tier-1 match. {"r_hip": {"raw": "Bip01_R_Thigh_1", "clean": "Right Thigh"}, "l_hip": {"raw": "Bip01_L_Thigh_1", "clean": "Left Thigh"}, "source": "thigh", "body_axis": false} ## Ex4 – OVERRIDE a mistaken body_axis hint
•
17
RIGHT has 'Right Thigh'; LEFT has 'Left Thigh'; HEAD has 'Tongue'; TAIL has 'Tail'. Hint = body_axis Tongue+Tail. OVERLAP = {Thigh} is non-empty -> STEP 2 wins, hint is wrong. {"r_hip": {"raw": "RightUpLeg_033", "clean": "Right Thigh"}, "l_hip": {"raw": "LeftUpLeg_028", "clean": "Left Thigh"}, "source": "thigh", "body_axis": false} ## Ex5 – Snake / serpent (rule B fires) RIGHT and LEFT are empty; HEAD has 'Head'; TAIL has 'Tail_30'. STEP 3 applies. {"r_hip": {"raw": "Head", "clean": "Head"}, "l_hip": {"raw": "Tail_30", "clean": "Tail"}, "source": "body_axis", "body_axis": true} ## Ex6 – Vehicle / prop (rule C fires) All sections empty, or only a lone 'Right Eye' with no left mate. {"r_hip": {"raw": "", "clean": ""}, "l_hip": {"raw": "", "clean": ""}, "source": "empty", "body_axis": false}
A.4
Motion Rendering
For caption supervision and qualitative inspection, every retained clip is rendered to a four-view synchronized video with an automated Blender pipeline that follows the camera convention of AnimaX [Huang et al. 2025] and DIMO [Mou et al. 2025]. For each clip we additionally render the rest-pose asset under the same camera setup, providing the captioner with a static reference frame against which articulated motion can be read out. Four cameras are placed at fixed elevation 0◦ and orthogonal azimuths 𝑎 ∈ {0◦, 90◦, 180◦, 270◦ } on a sphere of radius 2 m around the rig, with a fixed field of view of 33.9◦ and a pinhole projection. All motions are resampled to 30 FPS prior to rendering so that clip-level temporal statistics, frame-ratedependent motion descriptors, and the captioning model’s perceived motion speed are comparable across sources.
A.5
Motion Captioning
We caption each animation clip by querying a multimodal LLM (Qwen3.5-9B [Qwen Team 2026]) on the four-view rendering described in Section A.4, paired with a task-specific system prompt. Three prompts target the human (Mixamo), generic (Objaverse-XL), and animal (Truebones) subsets of UniML3D. All three share the same structure (camera description, observation protocol, caption rules, and style examples) and differ in the subject term, heading cues, body-part vocabulary, and dataset-specific rules; for Mixamo and Truebones, the source catalogue’s motion label is supplied as a hint that the model may use to disambiguate but must not copy. We reproduce the Objaverse-XL prompt verbatim below as the most general of the three, since it must cover humanoids, quadrupeds, avians, marine creatures, insectoids, serpentines, and articulated rigid objects within a single instruction. To ensure caption fidelity, we additionally perform a manual quality-control pass in which annotators inspect each motion sequence side-by-side with its generated caption, and revise or discard any pair in which the caption misidentifies the dominant action, attributes motion to the wrong body part, or disagrees with the underlying clip; this human-in-theloop check guarantees that both the motion data and the textual supervision used to train UniMate meet a consistent quality bar. Clips with and without root translation are both retained for training; the captions make the distinction explicit, labeling stationary motions with an “in place” qualifier so that the prompt disambiguates in-place articulation from root-translating locomotion. # Role
18
•
Mou et al.
You caption clips from the Objaverse 3D animation dataset – humanoids (majority), quadrupeds, avians, marine creatures, insectoids, serpentines, and articulated rigid objects. # Input Four synchronized cameras 90 degrees apart at fixed elevation, labeled only by their relative azimuth about the vertical axis (0°, 90°, 180°, 270°; 0° and 180° are opposite cameras, as are 90° and 270°). The labels carry NO information about the asset's canonical heading – the subject's world orientation is arbitrary, so any camera may be seeing the front, back, side, or an oblique angle. Determine facing from the body itself (protocol step 1). The cameras are STATIC: when the subject shifts across the frame or grows/shrinks, that is the subject translating, never camera motion. A view looking straight along an elongated body may show only a compact silhouette – rely on the other views. The frames of each view are uniformly sampled across the clip in chronological order, and the views are synchronized: frame k of every view shows the same moment. Pose can change noticeably between consecutive sampled frames. # Task ONE short sentence describing the dominant motion in OBJECT-RELATIVE terms. Subject is 'An object', body parts use natural anatomy words (arm, wing, tail, ...). When topology is ambiguous, default to humanoid vocabulary – most assets in this dataset are humanoid. # Observation protocol (silent – write only the caption) 1. ESTABLISH HEADING. Cues, in priority order: (a) head/face/eye/snout direction, (b) spine direction shoulders->hips for quadrupeds, (c) beak/head for avians, (d) head vs. tail end for serpentines, (e) principal translation axis for faceless rigid assets. 2. CHIRALITY. From the front, the object's left side is on the viewer's right (mirror). Apply consistently – never label a limb by which side of the camera frame it sits on. 3. SCAN every frame across all four views; do not infer from first/last alone. Note which body parts change pose and how. 4. CLASSIFY: translation (moves through space), rotation in place, articulation only (limbs/wings/tail move, body fixed), or held pose. Direction is read relative to the heading from step 1. Say 'in place' only for rotation without translation, or for locomotion-style movement without translation ('walks in place'); never as a default. Compare the subject's facing at the START and END of the clip: if it differs, a turn happened and belongs in the caption ('turns around and walks away'). 5. PICK the most specific verb that fits the kinematics. # Body-part vocabulary (descriptions of shape, not category labels) - Humanoid/bipedal: arm, hand, leg, foot, torso, head, hip, shoulder - Quadruped: front leg, hind leg, head, torso, tail - Winged/flying: wing, head, torso, tail, leg (if visible) - Serpentine: head, body, tail - Aquatic: fin, tail, head, body - Insectoid: leg, body, head - Articulated rigid: part, segment, base, top, arm (mechanical) If shape fits no row: upper/lower part, left/right side, front, back. # Rules - Format: 'An object [action].' (one sentence, <12 words) - ONE dominant action – or a short two-phase sequence ('X and then Y') when the clip clearly has two stages. - A cyclic motion (walk cycle, idle sway) is described ONCE as a continuous action, never per repetition; reserve the two-phase form for genuinely distinct stages. - A concurrent pose or secondary movement may be attached with 'while ...' / 'with ...' ('walks forward with arms swinging', 'rotates its base while extending an arm'). - Mention direction (forward/backward/left/right/up/down/in place) only if clearly observed; omit rather than guess. - All directions and chirality are object-relative (step 1), never camera-relative. If heading is uncertain, omit chirality. - Mention a body part only when essential to disambiguate. - No category names ('person', 'dog', 'dragon', 'car', 'robot', ...); anatomy words (arm, wing, tail) are allowed. - No adverbs, no appearance/color/material/texture, no scene/lighting. - Only say 'stands still' if the pose is truly unchanged across ALL frames; subtle sway, breathing, or limb shifts still count as motion. # Style examples (format only – do NOT copy unless the motion matches) - An object walks forward with arms swinging. - An object kicks with the right leg while pivoting. - An object flaps its wings and rises. - An object slithers forward. - An object rotates its upper segment in place. - An object crouches down and springs upward. Respond with ONLY the caption sentence – plain text, no quotation marks.
A.6
Balanced Sampling and Augmentation
UniML3D is heavily long-tailed: humanoids dominate while rare species are underrepresented. We mitigate this with a re-weighted sampler that assigns each instance of skeletal type 𝑖 (with 𝑛𝑖 samples) the weight 𝑤𝑖 = 𝑛𝑖−𝛼 . Setting 𝛼 = 0.5 yields square-root sampling, which trades off uniform sampling (𝛼 = 0, which underexposes rare structures) against full balancing (𝛼 = 1, which overfits tiny categories). On top of sampling, we apply four kinematics-preserving augmentations on the fly during training, each with a fixed probability. Joint removal prunes a random subset of leaf joints, preferring short bones. Joint addition inserts a kinematically neutral joint between an existing joint and its parent, splitting the bone so that forward kinematics is preserved. Skeleton pooling [Aberman et al. 2020] collapses degree-2 joints by composing each single-child joint’s local transform into its child, compressing linear chains without changing articulated motion. Bone-length perturbation scales each non-root bone offset by a factor near unity, varying limb proportions without modifying rotations. After augmentation, motion features are recomputed through forward kinematics and checked for self-consistency; clips that fail the check are skipped. Because augmentation changes the kinematic tree, the topology descriptors of Section 3.1 (Laplacian eigenvectors, graph-distance and relation matrices, and joint depths) are recomputed as well, so that the spectral coordinates and graph biases always reflect the augmented skeleton.
B Implementation Details B.1 Architecture The motion diffusion transformer comprises 𝑁 = 8 Skeletal-Temporal Transformer blocks with hidden dimension 𝑑 = 512, 𝑛 head = 8 attention heads (per-head dimension 𝑑ℎ = 64), and a SwiGLU [Shazeer 2020] feed-forward network of intermediate dimension 4𝑑 = 2048. We use RMS normalization [Zhang and Sennrich 2019] and querykey normalization [Dehghani et al. 2023; Henry et al. 2020] throughout. The graph distance and relation-type embeddings have dimension 𝑑𝑒 = 128, and Spec-RoPE uses the 𝑚 = 8 leading non-trivial Laplacian eigenvectors with the SignNet [Lim et al. 2023] backend by default. The pooled topological condition 𝑐 topo is obtained by crossattention pooling with 𝑛𝑞 = 4 learnable queries. Per-joint name embeddings are enabled by default; they are precomputed once on the unique joint-name vocabulary and cached. The conditioning signal, classifier-free guidance, and the AdaLN-Zero injection scheme are described in Section B.2.
B.2
Conditioning
TADiT is conditioned on the diffusion timestep, the input text prompt, and the skeleton topology. The flow time 𝜏 is mapped to a timestep embedding through a sinusoidal positional encoding followed by an MLP, 𝑐𝜏 = MLP(PE(𝜏)).
(16)
The text prompt is encoded into a latent text condition 𝑐 text by a frozen FLAN-T5-Base [Chung et al. 2024] encoder, whose output token sequence is linearly projected to dimension 𝑑 = 512. To enable
UniMate: One Unified Model to Animate Diverse Skeletons
classifier-free guidance [Ho and Salimans 2022], 𝑐 text is independently dropped to a learnable null embedding with probability 𝑝 cf during training. The topology-specific term 𝑐 topo is defined in the main paper. The three condition vectors are fused into a single global signal 𝑐 = 𝑐𝜏 + 𝑐 text + 𝑐 topo , which is injected into every transformer sublayer (joint attention, temporal attention, MLP) and the output head through adaptive layer normalization with zero-initialized gating (AdaLN-Zero) [Peebles and Xie 2023]. For a sublayer 𝑓 acting on a hidden state ℎ, AdaLN-Zero produces ℎ ′ = ℎ + 𝛼 (𝑐) ⊙ 𝑓 (1 + 𝛾 (𝑐)) ⊙ LN(ℎ) + 𝛽 (𝑐) , (17)
where (𝛼, 𝛽, 𝛾) = 𝑊𝑐 𝑐 are zero-initialized linear projections of the conditioning vector, so 𝛼 =0 at initialization and each block starts from the identity transform.
B.3
Output Head
After the final transformer block, the latent tensor takes the shape 𝑍 (𝑁 ) ∈ R (𝑇 +1) ×𝐽 ×𝑑 . The output head maps latent tokens back to the motion feature space. Because root and non-root joints follow distinct feature conventions, we decode them with separate twolayer MLPs, both AdaLN-Zero modulated by 𝑐. Prior to decoding, the root token at each frame attends to the non-root joint tokens of the same frame through a zero-initialized cross-attention block, aggregating whole-body information into the root channel as needed. Letting Π(·) denote the final decoder, the predicted velocity field is 𝑉ˆ = Π(𝑍 (𝑁 ) , 𝑐) ∈ R𝑇 ×𝐽 ×𝐷 . The prepended topology slot is discarded after decoding, and the remaining output is reshaped to the original motion layout.
B.4
(18)
with 𝜏 ∼ U (0, 1) and Ω ∈ {0, 1}𝑇 ×𝐽 ×𝐷 a binary validity mask that zeros out padded joint slots in heterogeneous-skeleton batches. The geodesic loss on 𝑆𝑂 (3) is applied to the one-step denoised rotations rather than to the velocity itself, so that supervision is performed in the clean motion space. From the predicted velocity, the clean motion estimate is 𝑥ˆ1 = 𝑥𝜏 + (1 − 𝜏) 𝑣𝜃 (𝑥𝜏 , 𝜏, 𝑐, S). Let 𝑅ˆ𝑡𝑗 ∈ 𝑆𝑂 (3) denote the rotation matrix recovered from the 6D rotation channels of 𝑥ˆ1 at frame 𝑡 and joint 𝑗, and 𝑅𝑡𝑗 ∈ 𝑆𝑂 (3) the corresponding ground-truth rotation. Then "Í # ˆ𝑡 𝑡 𝑗,𝑡 𝜔 𝑗,𝑡 𝑑 geo 𝑅 𝑗 , 𝑅 𝑗 Í Lgeo = E𝑥 1 , 𝑥 0 , 𝜏 , 𝑗,𝑡 𝜔 𝑗,𝑡 (19) tr 𝑅⊤𝑎 𝑅𝑏 − 1 𝑑 geo (𝑅𝑎 , 𝑅𝑏 ) = arccos , 2 where 𝜔 𝑗,𝑡 ∈ {0, 1} is the per-joint validity mask induced by Ω and the arccos argument is clamped to [−1 + 𝜖, 1 − 𝜖] for numerical stability. The velocity smoothness regularizer applies to the temporal acceleration of the denoised velocity channels. Letting 𝑣ˆ𝑡𝑗 ∈ R3 denote
19
(20)
where 𝜔 Δ𝑗,𝑡 = 𝜔 𝑗,𝑡 𝜔 𝑗,𝑡 +1 restricts the finite difference to pairs of consecutive valid frames within a clip.
B.5
Training
We train UniMate with the flow-matching objective in Section B.4, weighting the geodesic regularizer at 𝜆geo = 0.5 and the smoothness regularizer at 𝜆smooth = 0.1. Optimization uses AdamW [Loshchilov and Hutter 2019] with a learning rate of 1×10−4 , (𝛽 1, 𝛽 2 ) = (0.9, 0.99), weight decay 1 × 10−5 , and gradient clipping at norm 1.0. The learning rate follows a cosine schedule with linear warmup over the first 3% of training and a floor of 5% of the peak rate. We maintain an exponential moving average of model parameters with a step-dependent decay 1+𝑛 𝜂𝑛 = min 0.9999, , (21) 10 + 𝑛 where 𝑛 is the optimization step, so that the EMA tracks fast updates early and saturates to 0.9999 later in training. Training is distributed across 8 NVIDIA H100 GPUs using bfloat16 mixed precision, with a per-GPU batch size of 32 for a global batch size of 256. The classifierfree guidance dropout probability is 𝑝 cf = 0.1 on the text condition. We train for 100k optimization steps in total, which corresponds to approximately one day of wall-clock time on the configuration above.
B.6
Training Objective
The masked flow-matching MSE loss is " 2# Ω ⊙ 𝑣𝜃 (𝑥𝜏 , 𝜏, 𝑐, S) − (𝑥 1 − 𝑥 0 ) 2 Lmse = E𝑥 1 , 𝑥 0 , 𝜏 , ∥Ω∥ 1
the velocity block of 𝑥ˆ1 at joint 𝑗 and frame 𝑡, # " ∑︁ 1 Δ 𝑡 +1 𝑡 2 𝜔 𝑗,𝑡 𝑣ˆ 𝑗 − 𝑣ˆ 𝑗 2 , Lsmooth = E𝑥 1 , 𝑥 0 , 𝜏 Í 3 𝑗,𝑡 𝜔 Δ𝑗,𝑡 𝑗,𝑡
•
Inference
At inference time we evaluate the EMA copy of the model. Given a rigged 3D asset and a text prompt, we first canonicalize the skeleton, compute its topology descriptors (rest-pose joint positions, graph distance and relation matrices, joint depths, Laplacian eigenvectors, and joint-name embeddings), and encode the text prompt with the frozen FLAN-T5 encoder; these conditioning tensors are computed once per asset/prompt pair and reused across samples. We then draw an initial noise tensor 𝑥 0 ∼ N (0, 𝐼 ) in the padded motion shape and integrate the learned velocity field 𝑣𝜃 (𝑥𝜏 , 𝜏, 𝑐, S) from 𝜏 = 0 to 𝜏 = 1 with a fixed-step Euler ODE solver using 50 steps, which we found to be sufficient for visually clean trajectories. Classifier-free guidance is applied at every solver step: we run one forward pass with the text condition active and one with the null embedding, and combine them as 𝑣ˆ𝜃 = 𝑣𝜃uncond + 𝑠 (𝑣𝜃cond − 𝑣𝜃uncond ), (22)
with guidance scale 𝑠 = 3.0 by default; the unconditional and conditional branches are batched into a single forward call to avoid wallclock overhead. After the final solver step we discard the prepended skeleton tokens, slice off the padded joint slots using the jointvalidity mask, and de-normalize the predicted features with the dataset-level statistics computed during preprocessing. Joint rotations are converted from the continuous 6D representation to 𝑆𝑂 (3) matrices, the root trajectory is reconstructed by integrating the yaw-canonical root velocities and re-applying the initial facing direction, and the remaining global joint poses are obtained through
20
•
Mou et al.
forward kinematics. The resulting skeleton-space motion is finally driven onto the input mesh via linear blend skinning to produce the rendered animation. Sampling a 60-frame clip on a single NVIDIA H100 GPU takes roughly 1.2 seconds end-to-end, including text encoding and forward-kinematic recovery.
B.7
Animation Module
From 𝑥ˆ1 , we recover the global root trajectory and per-joint rotations and apply forward kinematics with the rest-pose offsets to obtain per-joint rigid transformations in 𝑆𝐸 (3), which drive the mesh via linear blend skinning [Magnenat-Thalmann et al. 1988] or dualquaternion blending [Kavan et al. 2007].
C
Theoretical Analysis of Spec-RoPE
This section discusses several useful properties of Spec-RoPE, the spectral rotary positional encoding employed in our topology-aware motion transformer. Unlike standard RoPE, which is defined on canonical 1D or 2D coordinates, our setting operates on the kinematic tree induced by a skeleton. We therefore construct rotary coordinates from the graph spectrum of that skeleton.
C.1
Setup
Let TS = (V, E) denote the kinematic tree of a skeleton with |V | = 𝐽 joints. Let 𝐴 ∈ {0, 1} 𝐽 ×𝐽 be its adjacency matrix. The combinatorial graph Laplacian is 𝐿 S = diag(𝐴1) − 𝐴, (23) which is symmetric and positive semidefinite, and admits the eigendecomposition 𝐿 S = 𝑈 Λ𝑈 ⊤,
𝑈 = [𝑢 0, 𝑢 1, . . . , 𝑢 𝐽 −1 ],
(24)
with Λ = diag(𝜆0, . . . , 𝜆 𝐽 −1 ) and 0 = 𝜆0 ≤ 𝜆1 ≤ · · · ≤ 𝜆 𝐽 −1 . Since the kinematic tree is connected, 𝜆0 is simple and 𝑢 0 is constant; we therefore discard 𝑢 0 and form the spectral coordinate of joint 𝑖 from the leading 𝑚 non-trivial eigenvectors, s𝑖 = [𝑢 1 (𝑖), 𝑢 2 (𝑖), . . . , 𝑢𝑚 (𝑖)] ∈ R𝑚 .
Given a per-head query or key vector 𝑧𝑖 the block-diagonal form RoPE(s𝑖 ) 𝑧𝑖 =
𝑑ℎ /2 Ê 𝑛=1
𝜌 (𝜃 𝑛 (s𝑖 )) [𝑧𝑖 ] 2𝑛−2:2𝑛−1,
cos 𝜃 𝜌 (𝜃 ) = − sin 𝜃
sin 𝜃 , cos 𝜃
𝜃 𝑛 (s) = 𝜔𝑛⊤ s,
𝜔𝑛 ∈ R𝑚 ,
(25)
∈ R𝑑ℎ , Spec-RoPE takes
(26)
where 𝜃 𝑛 : R𝑚 → R is a learned per-channel angle map. The analysis below studies the analytically tractable linear parameterization (27)
which admits the cleanest algebraic structure. The SignNet backend, used as our default, replaces (27) with the sign-symmetric construction of the main paper: each spectral coordinate is passed through a shared MLP at both signs and the symmetrized features are concatenated and mapped to the angles, 𝜃 𝑛 (s) = 𝜓𝑛 Concat {𝜙 (𝑠𝑘 ) + 𝜙 (−𝑠𝑘 )}𝑚 (28) 𝑘=1 , which is invariant to independent sign flips of the individual eigenvectors; this trades the strict relative-coordinate property below
for invariance to the spectral sign ambiguity, while preserving the qualitative behavior.
C.2
Proposition 1 (Translation Invariance in Spectral Coordinates)
Under the linear parameterization (27), the inner product between rotary-encoded queries and keys depends on the spectral coordinates only through their difference: RoPE(s𝑖 ) ⊤ RoPE(s 𝑗 ) = RoPE(s 𝑗 − s𝑖 ),
⟨RoPE(s𝑖 ) 𝑞, RoPE(s 𝑗 ) 𝑘⟩ = 𝑞 ⊤ RoPE(s 𝑗 − s𝑖 ) 𝑘.
(29)
Proof. Each 2 × 2 block satisfies 𝜌 (𝛼) ⊤ 𝜌 (𝛽) = 𝜌 (𝛽 − 𝛼) since 𝜌 is a one-parameter subgroup with 𝜌 (𝛼) ⊤ = 𝜌 (−𝛼). With linear angles, 𝜔𝑛⊤ s 𝑗 − 𝜔𝑛⊤ s𝑖 = 𝜔𝑛⊤ (s 𝑗 − s𝑖 ), so each block reduces to 𝜌 (𝜔𝑛⊤ (s 𝑗 − s𝑖 )). The block-diagonal structure of RoPE then yields the matrix identity, and the inner-product form follows immediately. □ Consequently the rotary encoding injects relative structural information from the kinematic tree into attention, mirroring the central design principle of standard RoPE on regular sequences.
C.3
Proposition 2 (Permutation Equivariance)
A spectral graph encoding should depend on the kinematic tree itself, not on any particular enumeration of its joints. Our construction satisfies this property up to the standard ambiguities intrinsic to eigendecompositions. Let 𝑃 ∈ R 𝐽 ×𝐽 be the permutation matrix induced by a reordering of the joints. Then the reordered Laplacian satisfies 𝐿 ′S = 𝑃 𝐿 S 𝑃 ⊤ . Its eigenvalues coincide with those of 𝐿 S , and its eigenvectors transform as 𝑢𝑘′ = ±𝑃𝑢𝑘 when 𝜆𝑘 is simple, and as 𝑈 Λ′ = 𝑃 𝑈 Λ 𝑄 for some 𝑄 ∈ 𝑂 (dim EΛ ) within any degenerate eigenspace EΛ . Consequently, Spec-RoPE is equivariant to joint permutation modulo this standard sign-and-basis ambiguity. Proof. Permuting the joint set conjugates both the adjacency and degree matrices by 𝑃, hence 𝐿 ′S = 𝑃𝐿 S 𝑃 ⊤ . Conjugation preserves the spectrum, so the eigenvalues are invariant. If 𝐿 S𝑢𝑘 = 𝜆𝑘 𝑢𝑘 then 𝐿 ′S (𝑃𝑢𝑘 ) = 𝑃𝐿 S 𝑃 ⊤ 𝑃𝑢𝑘 = 𝜆𝑘 (𝑃𝑢𝑘 ), identifying 𝑃𝑢𝑘 as an eigenvector of 𝐿 ′S . For a simple eigenvalue, eigenvectors are determined up to sign, giving 𝑢𝑘′ = ±𝑃𝑢𝑘 . For a degenerate eigenvalue, any orthonormal basis of the eigenspace is admissible, leaving the residual orthogonal freedom 𝑄 ∈ 𝑂 (dim EΛ ). The spectral coordinates s𝑖 inherit the same equivariance up to these intrinsic ambiguities, and the SignNet backend further removes the sign ambiguity by construction: writing 𝜋 for the relabeling, simple 𝜆1, . . . , 𝜆𝑚 give s𝜋′ (𝑖 ) = [±𝑢 1 (𝑖), . . . , ±𝑢𝑚 (𝑖)], and since 𝜙 (𝑠) + 𝜙 (−𝑠) is even in each coordinate, the SignNet angles satisfy 𝜃 𝑛 (s𝜋′ (𝑖 ) ) = 𝜃 𝑛 (s𝑖 ), so the rotary matrices permute with the joints and the attention logits are consistent under relabeling. □ Repeated and near-repeated eigenvalues. SignNet resolves the independent sign ambiguity of each eigenvector, but it is not invariant to an arbitrary orthogonal change of basis within a repeated eigenspace. Near-repeated eigenvalues can likewise yield numerically unstable eigenvectors: small perturbations may rotate their basis substantially even when the underlying subspace changes little. Spec-RoPE alone therefore does not provide complete basis invariance in these cases. The truncation at 𝑚 inherits the same caveat:
UniMate: One Unified Model to Animate Diverse Skeletons
when 𝜆𝑚 = 𝜆𝑚+1 , the cut falls inside a repeated eigenspace, so even the subspace spanned by the leading 𝑚 non-trivial eigenvectors is ambiguous. In the full architecture, however, spectral coordinates are only one of several complementary structural cues. Graph distance and relation-type biases are independent of the eigenspace basis, while rest-pose geometry, joint depth, and joint-name embeddings provide additional geometric, hierarchical, and semantic information. The model can therefore retain structural identifiability when individual spectral axes are ambiguous rather than relying exclusively on their orientation. A fully basis-invariant encoding of degenerate eigenspaces remains an interesting direction for future work.
C.4
Proposition 3 (Connection to Standard RoPE)
Spec-RoPE generalizes ordinary RoPE from regular grids to arbitrary skeletal graphs. When the underlying graph is a path graph on 𝐽 vertices, the first non-trivial Laplacian eigenvector 𝑢 1 is a strictly monotone function of the token index, and, restricted to this leading spectral coordinate, SpecRoPE recovers standard 1D RoPE up to a monotone reparameterization of the position coordinate. More generally, on Cartesian products of path graphs the axis-aligned first-harmonic eigenvectors vary along one axis each and recover the multi-axis RoPE used in image and video transformers, up to per-axis reparameterization. Proof. For a path graph 𝑃 𝐽 the Laplacian eigenpairs admit the closed form 𝑢𝑘 ( 𝑗) ∝ cos ( 𝑗 + 12 )𝜋𝑘/𝐽 with 𝜆𝑘 = 2 − 2 cos(𝜋𝑘/𝐽 ) for 𝑘 = 0, . . . , 𝐽 − 1 [Dwivedi and Bresson 2021; Rampášek et al. 2022]. The first non-trivial eigenvector 𝑢 1 is therefore a strictly monotone function of the index 𝑗, so an angle map supported on the leading coordinate, 𝜃 𝑛 = 𝜔𝑛,1 𝑢 1 ( 𝑗), is a monotone reparameterization of the standard rotary angle 𝜔𝑛,1 𝑗; the higher coordinates 𝑢 2, . . . , 𝑢𝑚 oscillate in 𝑗, so the reduction concerns this leading component. For a Cartesian product 𝑃 𝐽1 × · · · × 𝑃 𝐽𝑑 , the Laplacian eigenvectors Î factorize as 𝑢𝑘1 ,...,𝑘𝑑 ( 𝑗1, . . . , 𝑗𝑑 ) = 𝑎 𝑢𝑘𝑎 ( 𝑗𝑎 ), so the axis-aligned first harmonics vary along one axis each and play the role of axiswise positional coordinates; which eigenvectors are leading depends on the side lengths—for strongly skewed products, several higher harmonics of the longest axis precede the first harmonic of a shorter axis. In both cases Spec-RoPE thus reduces to existing RoPE constructions under a smooth, monotone change of coordinates rather than introducing a fundamentally different mechanism. □
C.5
Interpretation via Effective Resistance
A useful way to understand the structural bias induced by SpecRoPE is through the effective resistance of the kinematic tree. Let 𝐿 †S denote the Moore–Penrose pseudoinverse of the Laplacian. The effective resistance between joints 𝑖 and 𝑗 is 𝑅eff (𝑖, 𝑗) = (𝐿 †S )𝑖𝑖 + (𝐿 †S ) 𝑗 𝑗 − 2(𝐿 †S )𝑖 𝑗 , and the spectral expansion 𝐿 †S = 𝑅eff (𝑖, 𝑗) =
Í 𝐽 −1 𝑘=1
𝜆𝑘−1𝑢𝑘 𝑢𝑘⊤ yields
𝐽 −1 ∑︁ (𝑢 (𝑖) − 𝑢 ( 𝑗)) 2 𝑘
𝑘
𝜆𝑘
𝑘=1
(30)
.
(31)
Defining the scaled spectral coordinate " # 𝑢 𝐽 −1 (𝑖) 𝑢 1 (𝑖) 𝑢 2 (𝑖) s̃𝑖 = √ , √ , . . . , √︁ , 𝜆1 𝜆2 𝜆 𝐽 −1
•
21
(32)
this becomes the squared Euclidean distance in the scaled spectral embedding, 𝑅eff (𝑖, 𝑗) = ∥ s̃𝑖 − s̃ 𝑗 ∥ 22 .
(33)
In other words, differences in scaled spectral coordinates encode a graph-geometric distance. For a tree with unit edge weights—which our kinematic skeletons are—effective resistance coincides with shortest-path distance, so joints farther apart along the articulated hierarchy are also farther apart in scaled spectral space.
C.6
Small-Variance Interpretation
The effective-resistance identity offers a transparent picture of how rotary attention behaves in scaled spectral space. Suppose we use scaled coordinates s̃𝑖 in place of s𝑖 and average the rotary kernel over isotropic Gaussian frequencies 𝜔 ∼ N (0, 𝜎 2 𝐼 𝐽 −1 ). Standard properties of the characteristic function of a Gaussian give the closed form 2 E𝜔 cos 𝜔 ⊤ (s̃𝑖 − s̃ 𝑗 ) = exp − 𝜎2 ∥ s̃𝑖 − s̃ 𝑗 ∥ 22 (34) 2 = exp − 𝜎2 𝑅eff (𝑖, 𝑗) , so the expected rotary kernel decays exponentially in the effective resistance, and—to leading order in 𝜎 2 —its first-order correction is proportional to 𝑅eff (𝑖, 𝑗). This motivates viewing Spec-RoPE as a soft structural prior that suppresses attention between joints that are far apart on the kinematic tree. In our model this picture should be read as an interpretive guide rather than a literal theorem about the deployed architecture, for three reasons: (i) we use a truncated set of 𝑚 leading spectral coordinates rather than the full spectrum; (ii) the angles are produced by a learned deterministic map—linear in the analytically tractable case, sign-symmetric and nonlinear in our default SignNet construction— rather than by averaging over random Gaussian frequencies; and (iii) attention additionally incorporates an explicit graph bias through relation- and distance-type embeddings. The closed-form expectation therefore does not transfer verbatim. Empirically, however, the qualitative effect persists: rotations driven by spectral coordinates encourage attention to depend on relative graph geometry rather than on an arbitrary joint ordering.
C.7
Summary
Taken together, Propositions 1–3 justify our use of Spec-RoPE in a topology-aware motion model. The encoding is well-defined on arbitrary kinematic trees, is insensitive to joint permutation up to intrinsic spectral ambiguities (with SignNet removing eigenvector sign ambiguity), reduces to standard RoPE on regular structures, and admits a natural distance-based interpretation on trees through effective resistance. These properties make it particularly suitable for our setting, where a single model must generalize across heterogeneous skeleton topologies while preserving meaningful structural relationships between joints.
22
•
D
Mou et al.
Experimental Protocols
This section documents the evaluation protocols behind the experiments in the main paper. Section D.1 distinguishes the metrics used for skeletal motion and rendered mesh animation, and Section D.2 lists the exact held-out skeletons, per-sequence prompts, and random seed. Section D.3 describes our out-of-distribution evaluation, Section D.4 discusses the compared methods and the adaptations required for a fair comparison, and Section D.5 summarizes the complementary strengths of skeleton-based and skeleton-free animation. Finally, Section D.6 details the protocol of our anonymous perceptual study.
D.1
Evaluation Metrics
We evaluate the two comparison settings in their native output domains. For skeleton-based motion generation, motion quality is measured directly in skeleton space using FID over kinematic features, together with diversity over generated joint trajectories. VBench is used only for the mesh-animation comparison: following the protocol of the skeleton-free baselines [Huang et al. 2025; Wu et al. 2025], we render all methods as fixed-view videos and evaluate their perceptual quality in video space. This separation avoids using a video metric as a proxy for skeletal motion quality while retaining a common output representation for methods that do not produce skeletons.
D.2
Held-Out Evaluation Set
The motion-generation comparison in the main paper holds out seven Truebones skeleton types that span six morphology categories: Raptor (bipedal), Leapord and Raindeer (quadrupedal), Parrot2 (avian), Jaws (marine), Crab (insectoid), and KingCobra (serpentine). No motion of these skeletons is seen during training. Together they contribute 58 evaluation sequences; Tab. 5 lists, for each skeleton type, the exact text prompts used to condition generation. Skeleton names keep the original Truebones spelling (e.g., Leapord, Raindeer), while prompts refer to each character by its species name. All quantitative results are generated with random seed 10.
D.3
Out-of-Distribution Evaluation
In addition to the held-out Truebones [Truebones 2022] skeletons and the mesh-animation benchmark of the main paper, we evaluate UniMate on a small curated set of AI-generated and automatically rigged meshes. These assets exhibit non-canonical proportions, irregular skeleton topologies, and rigging conventions that are not represented in any of our training sources, and therefore probe the model’s ability to generalize beyond artist-authored rigs.
D.4
Baselines
AnyTop. For a fair comparison, we extend AnyTop [Gat et al. 2025] with the cross-attention text-conditioning module described in the main paper, since the published model is purely unconditional. Beyond this adaptation, three structural limitations of AnyTop are worth highlighting. (i) It relies on skeleton-specific statistical estimation and normalization, which requires the user to supply at least one reference motion sequence for every target skeleton before generation can begin. (ii) It is trained on a comparatively small corpus
of artist-crafted animal skeletons, which constrains the topological diversity it can express. (iii) It is fundamentally an unconditional generator and therefore lacks any native mechanism for textual control, which limits its practical utility for user-driven animation. Our cross-attention extension partly addresses (iii) for the purpose of comparison, but does not remedy (i) or (ii). How to Move Your Dragon. Closest to our text-conditioned setting, How to Move Your Dragon [Lee et al. 2025] annotates the Truebones corpus with text descriptions and trains a topology-adaptive motion diffusion model on it, using rig augmentation to broaden the skeletal configurations seen during training. It therefore natively addresses limitation (iii) above, but shares limitation (ii): its training domain remains the same small, artist-crafted animal corpus, whereas UniML3D spans humans, animals, and general articulated objects over thousands of distinct skeletons. To our knowledge, neither code nor pretrained models had been released at the time of writing, so we do not include it in our quantitative comparison. AnimaX. We do not include AnimaX [Huang et al. 2025] in our mesh-animation comparison: the authors have not publicly released their code or pretrained models, and the technical report does not provide enough detail for a faithful reimplementation.
D.5
Skeleton-Based vs. Skeleton-Free Animation
The two paradigms make complementary representation choices. A skeletal representation is compact and semantically meaningful: a motion is expressed as a small set of joint transformations rather than dense vertex trajectories. It decouples motion from a particular mesh geometry, allowing one generated sequence to drive different skinned assets with compatible articulation, and it is efficient to edit using established rig-based tools. It also integrates naturally with inverse kinematics, physics simulators, motion-capture pipelines, and interactive controllers. These properties make skeleton-based methods especially suitable for articulated characters and production workflows that require reusable and controllable motion. Their main cost is the need for a valid rig, and the kinematic tree constrains deformation to motions that its joints and skinning model can express. Skeleton-free methods instead predict deformation directly in mesh, point, or implicit geometry space. By avoiding a fixed joint hierarchy, they can better accommodate objects that are not naturally articulated, highly non-rigid deformation, and fine-grained surface dynamics that would require an impractically dense rig. This flexibility comes with a substantially higher-dimensional output, weaker built-in kinematic structure, and less direct compatibility with standard rig-based editing and control. Thus, skeleton-free methods are not superseded by our formulation; they are preferable when deformation cannot be captured well by a kinematic skeleton, whereas UniMate targets efficient, controllable animation of rigged assets. Hybrid models that retain skeletal control while learning residual non-rigid deformation are a promising direction for combining both strengths; Section E.2 contrasts the two paradigms visually.
UniMate: One Unified Model to Animate Diverse Skeletons
•
23
Table 5. The held-out evaluation set. The 58 evaluation prompts of the seven held-out Truebones skeletons (Crab 10, Jaws 8, KingCobra 8, Parrot2 2, Leapord 12, Raindeer 8, Raptor 10), grouped by skeleton type; each prompt is the exact text used to condition generation. Skeleton
Prompt
Skeleton
Prompt
Crab
A crab attacks forward. A crab lunges forward to attack. A crab strikes forward with its claws. A crab snaps its claws and attacks. A crab lowers its head and lies down. A crab collapses backward and lies down. A crab eats. A crab lies on its back. A crab strafes. A crab walks forward using its legs. A parrot preens its feathers. A parrot flaps its wings and glides forward.
Leapord
A leopard attacks forward. A leopard crouches low and bites forward. A leopard collapses and lies on its side. A leopard falls backward and lies on its side. A leopard crouches low. A leopard growls with its head lowered. A leopard gets up from lying down. A leopard pounces forward. A leopard pounces forward and swipes with its front legs. A leopard runs forward. A leopard stands still. A leopard walks forward.
Raindeer
A reindeer attacks forward and strikes with its head. A reindeer charges forward with its antlers lowered. A reindeer thrusts its head forward to attack. A reindeer falls backward and lies on its side. A reindeer runs forward. A reindeer stands still. A reindeer walks forward. A reindeer lowers its head and snorts.
KingCobra
A king cobra strikes forward with tail curled. A king cobra rears up and strikes forward. A king cobra strikes forward. A king cobra circles and bites its own tail. A king cobra raises its head. A king cobra stands still. A king cobra raises its head and holds a still pose. A king cobra throws its head back and falls forward.
Jaws
A shark bites to the left. A shark bites to the right. A shark thrashes in place with its jaws open. A shark swims forward with its jaws open. A shark swims forward and turns around. A shark swims forward. A shark swims to the left. A shark swims to the right.
Raptor
A raptor attacks forward and strikes with its head. A raptor falls backward and lies on its side. A raptor walks forward at a brisk pace. A raptor lowers its head and lunges forward. A raptor stands still. A raptor holds a still standing pose. A raptor scratches with its right foreleg. A raptor roars while standing still. A raptor walks forward slowly. A raptor yawns with its head raised upward.
Parrot2
D.6
User Study
To complement the automatic metrics with a human judgment of perceptual quality, we conducted an anonymous user study on textdriven mesh animations. As illustrated in Fig. 20, each form presents a single test object together with three side-by-side videos, each generated by a different method from the same input mesh and text prompt; participants then answer the four evaluation questions shown in Fig. 20b. To eliminate ordering and naming bias, the mapping between methods and video labels was anonymized and independently randomized per participant and per example. Figure 21 plots the resulting per-criterion ratings.
E
Additional Results
This section collects additional experimental results that complement the main-paper evaluation: Section E.1 extends the qualitative AnyTop comparison to further held-out morphologies and motion types, Section E.2 adds direct visual comparisons against the skeleton-free mesh-animation baselines, Section E.3 visualizes the architecture ablations, Section E.4 ablates the dataset-curation components, and Section E.5 quantifies foot sliding and the effect of foot-locking post-processing.
E.1
Additional Comparisons with AnyTop
Fig. 22 extends the AnyTop comparison beyond quadruped and biped locomotion to two further held-out skeletons and motion types from Section D.2: an avian rig on Parrot2-Land and a crustacean rig on Crab-Die. The same pattern holds across these morphologies: UniMate executes the prompted action on the unseen rig, while AnyTop’s outputs remain close to an idle, in-place motion regardless of the prompt.
E.2
Comparisons with Skeleton-Free Baselines
Fig. 23 contrasts the three methods on two cases from the user-study test set (Section D.6). The frame strips visualize the failure modes behind the quantitative gap in the main paper: AnimateAnyMesh’s high motion smoothness comes from near-static outputs with little articulated motion, and V2M4, driven by a separately generated video, both deviates from the prompt and accumulates mesh distortion over time. UniMate produces prompt-faithful, kinematically coherent motion on both rigs; the project page shows the full clips.
E.3
Qualitative Ablations
Fig. 24 complements the quantitative ablation table of the main paper by rendering the same held-out rig and prompt with each ablated variant. The visual differences track the numbers: the graphaware attention bias mainly sharpens joint coordination, Spec-RoPE is what carries the prompted action onto an unseen skeleton, and the global topological conditioner stabilizes the motion—removing it raises diversity but visibly degrades stability, which matches its higher diversity score alongside the largest FID increase.
E.4
Dataset-Curation Ablation
We ablate the two dataset-curation components that admit a controlled comparison on a fixed dataset. Removing rest-pose rotation rebasing degrades FID from 0.757 to 1.179: expressing joint rotations relative to the rest pose yields a unified kinematic basis across heterogeneous skeletons, and dropping it costs the most motion quality of any curation stage. Disabling the on-the-fly augmentations reduces diversity from 9.200 to 8.265, confirming that they mainly broaden topological coverage rather than per-clip fidelity. The filtering and canonicalization stages change which clips and skeletons enter the dataset, so a direct FID comparison against the
24
•
Mou et al.
(a) Participant instruction page. Fig. 21. User study results on text-conditioned mesh animation. Mean Likert ratings (1 = very poor, 5 = excellent, with error bars) per criterion and averaged overall; UniMate receives the highest rating on all four criteria— text-to-motion agreement, motion plausibility, motion expressiveness, and shape preservation—as well as overall.
rig-level foot-locking post-process removes. Ground contact is only well defined for morphologies with a support pattern, so we evaluate on 20 held-out legged skeletons (10 bipedal and 10 quadrupedal) drawn from the Truebones and Mixamo test splits; serpentine, marine, in-flight, and articulated-object rigs are excluded. For each rig, the set of contact joints C is selected automatically from the standardized joint vocabulary of Section A.2: the distal-most joint of every limb chain whose canonical name contains a ground-contact term (Foot, Toe, Ball, Paw, Hoof, or Pastern). No per-rig manual annotation is involved. (b) Representative test trial. Fig. 20. User-study interface and protocol. (a) Instruction page presented at the start of each session, with institution-identifying information masked. (b) A representative trial: three side-by-side videos generated by different methods from the same input mesh and text prompt, followed by the four evaluation questions used for subjective scoring; method–label assignments are anonymized and randomized per participant.
unfiltered corpus is not distribution-matched; their per-stage criteria and effects are documented in Section 4 and Algorithm 1; as a diagnostic statistic, implausibility filtering discards 8.7% of the candidate motion clips. Fig. 25 complements these statistics with before/after qualitative examples: it renders the same held-out rig and prompt from models trained without implausibility filtering (the last filtering stage) and without dataset-level statistics normalization (the last canonicalization stage), alongside the full pipeline. Dropping the filtering stage lets physically implausible training clips leak into the model, which reproduces their erratic, airborne dynamics; dropping the normalization destabilizes ground contact, producing drift and hovering; the full pipeline executes the prompted action with plausible, grounded motion.
E.5
Foot-Sliding Evaluation and Foot-Locking Post-Processing
We quantify the foot-sliding artifacts discussed in the limitations section of the main paper, and measure how much of this sliding a
Normalization. Because the evaluated skeletons differ widely in scale and proportion, all lengths are reported in units of the character’s rest-pose root height 𝐿, the vertical distance from the skeleton root to the ground plane; the ground plane itself is the horizontal plane through the lowest rest-pose joint. For a contact joint 𝑗 at frame 𝑡, let ℎ 𝑗,𝑡 denote its signed height above the ground plane (negative below it) and 𝑑 𝑗,𝑡 the horizontal (ground-parallel) displacement of the joint between frames 𝑡−1 and 𝑡, both in units of 𝐿, with all sequences sampled at 30 FPS. A joint is near the ground when ℎ 𝑗,𝑡 < 𝛿ℎ with 𝛿ℎ = 0.05, and a near-ground frame counts as skating when 𝑑 𝑗,𝑡 > 𝛿𝑠 with 𝛿𝑠 = 0.025; these thresholds correspond to the 5 cm and 2.5 cm used at human scale by Karunratanakul et al. [2023], where the hip height is roughly one meter. Metrics. We report two sliding measures, each averaged over all generated sequences of the evaluation set (lower is better). The skating ratio (Skate) follows Karunratanakul et al. [2023]: the fraction of near-groundframes of contact joints that are skating, i.e. Pr 𝑑 𝑗,𝑡 > 𝛿𝑠 | ℎ 𝑗,𝑡 < 𝛿ℎ , which captures how often planted limbs drift. The sliding distance (Slide) is the height-weighted horizontal drift of Zhang et al. [2018], 𝑠 𝑗,𝑡 = 𝑑 𝑗,𝑡 2 − 2 ℎ 𝑗,𝑡 /𝛿ℎ , averaged over near-ground frames and reported in units of 10−2 𝐿; unlike the binary ratio, it measures how far the limbs slide. Foot-locking post-processing. Where contact is well defined, foot locking applies zero-shot to the predicted rig without retraining. Per contact joint, frames with ℎ 𝑗,𝑡 < 𝛿ℎ and 𝑑 𝑗,𝑡 < 𝛿𝑠 are marked as contact, the binary signal is median-filtered with a five-frame
UniMate: One Unified Model to Animate Diverse Skeletons
Ours
•
25
AnyTop
Fig. 22. Additional comparisons with AnyTop on unseen Truebones skeletons. Rest-pose asset and prompt (left), our result (middle), and AnyTop (right), extending the main-paper comparison to an avian glide and a crustacean collapse. Our parrot flaps and glides forward and our crab collapses backward onto its shell as prompted, whereas AnyTop’s parrot hovers flapping in place and its crab keeps stepping with raised claws without collapsing.
Fig. 23. Qualitative comparison with skeleton-free baselines. Four evenly spaced frames per clip on two user-study cases, cropped around the subject for visibility. UniMate executes the prompted action with dynamic, coherent motion—the lion’s leap-and-bite and the anaconda’s upward coil; AnimateAnyMesh remains near-static on both cases, while V2M4 distorts the lion’s mane geometry and raises only the anaconda’s head, never producing the prompted coil.
window, and segments shorter than three frames are discarded. Each remaining segment is assigned an anchor: the mean horizontal position of the joint over the segment, projected onto the ground plane. The joint is then pinned to its anchor by damped least-squares inverse kinematics over its limb chain (the contact joint and up to three parent joints), leaving the root trajectory and all other chains untouched, and the correction is blended in and out linearly over five frames at each segment boundary to avoid pops. Table 6. Foot-sliding evaluation on held-out legged skeletons. Skate is the skating ratio and Slide the height-weighted sliding distance (Section E.5), the latter in units of 10 −2 𝐿 for a rest-pose root height 𝐿. FL denotes the foot-locking post-process. Method
Skate↓
Slide↓
UniMate (Ours) UniMate + FL
0.107 0.023
0.542 0.191
Tab. 6 reports the results. Foot locking reduces the skating ratio from 0.107 to 0.023 and the sliding distance from 0.542 to 0.191.
Within detected contact segments the correction eliminates sliding by construction; the residual stems from near-ground frames outside detected segments, the blended segment boundaries, and brief touchdowns discarded by the minimum-length filter.
F
Limitations and Future Work
The main paper summarizes the principal contact, scope, and datacoverage limitations of UniMate; here we provide a more detailed discussion and outline additional directions for controllability. Supervision is bounded by 4D animation data. UniMate is trained end-to-end on UniML3D, a curated corpus of paired skeletal motions and text prompts assembled from artist-authored 4D animation sources. Although UniML3D is, to our knowledge, the largest crossspecies rigged-motion dataset assembled to date, it remains orders of magnitude smaller than the visual corpora available to image and video foundation models, and its category distribution is heavily long-tailed: bipedal humanoids dominate, while serpentine, marine, and insectoid morphologies are sparsely populated (Section A.1).
26
•
Mou et al.
Fig. 24. Qualitative ablation of the topology-aware components. The same rig and prompt rendered from the already-trained ablation models of the main paper, with each variant’s FID and diversity inset. The full model executes a clear forward attack; without the graph-aware attention bias the wing–body coordination degrades; without Spec-RoPE the motion loses the prompted action and collapses toward in-place wing flailing; without the global topological conditioner the motion stays dynamic but becomes unstable, with jittery, poorly grounded poses.
Even after the square-root-balanced sampler and the kinematicspreserving online augmentations described in Section A.6, rare motions (extreme gymnastic skills, fine manipulation, complex social interactions) and underrepresented species (large invertebrates, exotic aquatic locomotion) are visibly harder for the model to cover than well-represented humanoid locomotion. A natural next step is to distill Internet-scale video or video-generative priors [Chen et al. 2026; Mou et al. 2025] into our unified framework—for example, by using video-derived motion fields as auxiliary supervision, or by aligning UniMate’s latent dynamics with a pretrained video diffusion model. This would let cross-species generalization scale with passive video data rather than with the cost of new 4D capture, while keeping the topology-aware backbone of UniMate intact. Agentic content-generation systems offer a complementary dataside remedy: Articraft [Zhou et al. 2026], for example, uses an LLM to synthesize diverse 3D assets at scale, and pairing such generated assets with scripted, simulated, or distilled motion could help densify the rare topology and motion regions that current 4D corpora leave uncovered. Conditioning is restricted to skeletons and text. The current conditioning interface accepts only a rigged skeleton and a naturallanguage prompt. This is sufficient for the text-to-animation setting we evaluate, but it leaves out several input modalities that artists
Fig. 25. Before/after qualitative examples for the curation stages. The same held-out rig and prompt rendered from models trained with the full curation pipeline (top) and with one stage removed. Without implausibility filtering, the implausible clips that survive in the corpus surface at generation time: the reindeer tumbles through airborne, twisted poses instead of attacking. Without dataset-level statistics normalization, the attack is roughly executed but the motion drifts and hovers with degraded ground contact. The full pipeline performs the prompted antler attack with stable, grounded motion.
and downstream systems regularly want to drive animation with: monocular video (motion capture from a single camera), exemplar motion clips (“animate this rig in the style of that clip”), shapeonly inputs (a mesh without an authored rig), and partial skeletal demonstrations (e.g., a keyframed root trajectory, a hand pose, or a footstep schedule). Each of these is straightforward to express as an additional conditioning stream consumed by the same AdaLNZero injection pathway used for text and topology (Section B.2); the substantive work lies in collecting paired data and designing modality-specific encoders that share UniMate’s topology-aware token layout. Adding these channels would move UniMate closer to a general motion-capture and animation system for arbitrary 3D assets, in which the same backbone covers text-driven synthesis, video-based mocap, and example-based retargeting under one model. No native channel for fine-grained spatial motion control. Text is a powerful but coarse controller: a prompt can specify “walk forward” or “perform a spin kick,” but cannot precisely place a foot on a target, route a hand along a desired curve, or impose contact constraints with the environment. UniMate currently exposes no first-class mechanism for such spatial directives, which is a real limitation for production use cases where animators need framelevel control over a few key joints. Two complementary directions
UniMate: One Unified Model to Animate Diverse Skeletons
can address this without retraining the backbone. First, samplingtime guidance methods—for example, the projection-based flow guidance of Watanabe et al. [2026]—let users specify per-joint spatial targets and have the ODE solver project each integration step onto the constraint manifold, trading a small amount of fidelity for hard control. Second, in-context learning over partial trajectory tokens, in the spirit of Cong et al. [2026], would let artists “paint” a desired motion onto a subset of joints and frames and have UniMate infill the remainder while respecting the topology-aware priors learned during training. Combining these mechanisms with the existing text and topology conditions would strengthen UniMate’s value as an animation foundation model that allows artists to both describe and directly control desired motions.
References
Kfir Aberman, Peizhuo Li, Dani Lischinski, Olga Sorkine-Hornung, Daniel Cohen-Or, and Baoquan Chen. 2020. Skeleton-aware networks for deep motion retargeting. ACM Transactions on Graphics 39, 4 (2020), 62:1–62:14. doi:10.1145/3386569.3392462 Adobe. 2022. Mixamo. https://www.mixamo.com/. Honglin Chen, Karran Pandey, Rundi Wu, Matheus Gadelha, Yannick Hold-Geoffroy, Ayush Tewari, Niloy J Mitra, Changxi Zheng, and Paul Guerrero. 2026. ViPS: Videoinformed Pose Spaces for Auto-Rigged Meshes. arXiv preprint arXiv:2604.17623. Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1–53. Xiaoyan Cong, Zekun Li, Zhiyang Dou, Hongyu Li, Omid Taheri, Chuan Guo, Abhay Mittal, Sizhe An, Taku Komura, Wojciech Matusik, et al. 2026. UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors. arXiv preprint arXiv:2603.15975. DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. arXiv preprint arXiv:2606.19348. Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. 2023. Scaling Vision Transformers to 22 Billion Parameters. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202). PMLR, 7480–7512. Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. 2023a. Objaverse-XL: A Universe of 10M+ 3D Objects. In Advances in Neural Information Processing Systems. 35799–35813. Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. 2023b. Objaverse: A universe of annotated 3D objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13142–13153. Vijay Prakash Dwivedi and Xavier Bresson. 2021. A generalization of transformer networks to graphs. In AAAI Workshop on Deep Learning on Graphs: Methods and Applications. Inbar Gat, Sigal Raab, Guy Tevet, Yuval Reshef, Amit Haim Bermano, and Daniel CohenOr. 2025. AnyTop: Character Animation Diffusion with Any Topology. In ACM SIGGRAPH 2025 Conference Papers. Association for Computing Machinery, 1–10. doi:10.1145/3721238.3730621 Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. 2020. Query-Key Normalization for Transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020. 4246–4253. Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598. Zehuan Huang, Haoran Feng, Yang-Tian Sun, Yuan-Chen Guo, Yan-Pei Cao, and Lu Sheng. 2025. AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models. In SIGGRAPH Asia 2025 Conference Papers. Association for Computing Machinery, 1–13. doi:10.1145/3757377.3763885 Korrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, and Siyu Tang. 2023. Guided motion diffusion for controllable human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 2151–2162. Ladislav Kavan, Steven Collins, Jiří Žára, and Carol O’Sullivan. 2007. Skinning with dual quaternions. In Proceedings of the 2007 Symposium on Interactive 3D Graphics and Games. 39–46. Wonkwang Lee, Jongwon Jeong, Taehong Moon, Hyeon-Jong Kim, Jaehyeon Kim, Gunhee Kim, and Byeong-Uk Lee. 2025. How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects. In Proceedings of the 42nd International
•
27
Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 267). PMLR, 33110–33128. Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang, Zhangyang Wang, Konstantinos N Plataniotis, Yao Zhao, and Yunchao Wei. 2024. Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models. In Advances in Neural Information Processing Systems, Vol. 37. 110854–110875. doi:10.52202/079017-3519 Derek Lim, Joshua David Robinson, Lingxiao Zhao, Tess Smidt, Suvrit Sra, Haggai Maron, and Stefanie Jegelka. 2023. Sign and basis invariant networks for spectral graph representation learning. In International Conference on Learning Representations. Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. Nadia Magnenat-Thalmann, Richard Laperrière, and Daniel Thalmann. 1988. Jointdependent local deformations for hand animation and object grasping. In Proceedings of Graphics Interface. 26–33. Linzhan Mou, Jiahui Lei, Chen Wang, Lingjie Liu, and Kostas Daniilidis. 2025. DIMO: Diverse 3D Motion Generation for Arbitrary Objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 14357–14368. OpenAI. 2025. OpenAI GPT-5 System Card. arXiv preprint arXiv:2601.03267. doi:10. 48550/arXiv.2601.03267 William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 4195–4205. Qwen Team. 2026. Qwen3.5-Omni Technical Report. arXiv preprint arXiv:2604.15804. Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini. 2022. Recipe for a general, powerful, scalable graph transformer. Advances in Neural Information Processing Systems 35 (2022), 14501– 14515. Noam Shazeer. 2020. GLU Variants Improve Transformer. arXiv preprint arXiv:2002.05202. Chaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin, and Jianfeng Zhang. 2025. Puppeteer: Rig and Animate Your 3D Models. arXiv preprint arXiv:2508.10898. Truebones. 2022. Truebones Zoo Dataset. https://truebones.gumroad.com/. Akihisa Watanabe, Qing Yu, Edgar Simo-Serra, and Kent Fujiwara. 2026. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control. arXiv preprint arXiv:2602.22742. Zijie Wu, Chaohui Yu, Fan Wang, and Xiang Bai. 2025. AnimateAnyMesh: A FeedForward 4D Foundation Model for Text-Driven Universal Mesh Animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13557– 13568. Biao Zhang and Rico Sennrich. 2019. Root mean square layer normalization. Advances in Neural Information Processing Systems 32 (2019). He Zhang, Sebastian Starke, Taku Komura, and Jun Saito. 2018. Mode-adaptive neural networks for quadruped motion control. ACM Transactions on Graphics 37, 4 (2018), 1–11. doi:10.1145/3197517.3201366 Matt Zhou, Ruining Li, Xiaoyang Lyu, Zhaomou Song, Zhening Huang, Chuanxia Zheng, Christian Rupprecht, Andrea Vedaldi, and Shangzhe Wu. 2026. Articraft: An Agentic System for Scalable Articulated 3D Asset Generation. arXiv preprint arXiv:2605.15187.