ConceptioArchivearXiv CS
arXiv CSopen access

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neuralnetworks
machine learning, deep learning, neural networks

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

arXiv:2605.13838v1 [cs.CV] 13 May 2026

ZIJIE WU, Huazhong University of Science and Technology, China and Tencent Hunyuan, China LIXIN XU∗ , Tencent Hunyuan, China PUHUA JIANG, Tencent Hunyuan, China SICONG LIU, Tencent Hunyuan, China CHUNCHAO GUO, Tencent Hunyuan, China XIANG BAI† , Huazhong University of Science and Technology, China

Fig. 1. Video-Guided 3D Animation via Rectified Dynamic Mesh (R-DMesh). Given a monocular reference video (left), our method synthesizes high-fidelity, motion-aligned 4D meshes. Unlike traditional methods relying on skeletal rigging, R-DMesh directly predicts vertex trajectories, enabling the animation of topological varying characters (top) and common objects (bottom left) without explicit shape priors. Trained on same-identity video-4D pairs, our framework generalizes robustly to drive 3D models using reference videos from different identities, including in-the-wild footage. Furthermore, our approach supports versatile applications such as pose retargeting (bottom center) and holistic video-to-4D generation (bottom right). ∗ Project lead. † Corresponding author.

Authors’ Contact Information: Zijie Wu, Huazhong University of Science and Technology, Wuhan, China and Tencent Hunyuan, Shanghai, China, [email protected]; Lixin Xu, Tencent Hunyuan, Shanghai, China, [email protected]; Puhua Jiang, Tencent Hunyuan, Shenzhen, China, [email protected]; Sicong Liu, Tencent Hunyuan, Shenzhen, China, [email protected]; Chunchao Guo, Tencent Hunyuan, Shenzhen, China, [email protected]; Xiang Bai, Huazhong University of Science and Technology, Wuhan, China, [email protected].

This work is licensed under a Creative Commons Attribution 4.0 International License. SIGGRAPH Conference Papers ’26, Los Angeles, CA, USA © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2554-8/2026/07 https://doi.org/10.1145/3799902.3811135

Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical deployment faces a critical yet frequently overlooked hurdle: the pose misalignment dilemma. In real-world scenarios, the initial pose of a userprovided static mesh rarely aligns with the starting frame of a reference video. Naively forcing a mesh to follow a mismatched trajectory inevitably leads to severe geometric distortion or animation failure. To address this, we present Rectified Dynamic Mesh (R-DMesh), a unified framework designed to generate high-fidelity 4D meshes that are “rectified” to align with video context. Unlike standard motion transfer approaches, our method introduces a novel VAE that explicitly disentangles the input into a conditional base mesh, relative motion trajectories, and a crucial rectification jump offset. This offset is learned to automatically transform the arbitrary pose of the input mesh to match the video’s initial state before animation begins. We process these components via a Triflow Attention mechanism, which leverages vertex-wise geometric features to modulate the three orthogonal flows, ensuring physical consistency and local rigidity during the rectification

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

2

Wu, Z. et al

and animation process. For generation, we employ a Rectified Flow-based Diffusion Transformer conditioned on pre-trained video latents, effectively transferring rich spatio-temporal priors to the 3D domain. To support this task, we construct Video-RDMesh, a large-scale dataset of over 500k dynamic mesh sequences specifically curated to simulate pose misalignment. Extensive experiments demonstrate that R-DMesh not only solves the alignment problem but also enables robust downstream applications, including pose retargeting and holistic 4D generation. Code and pre-trained weights will be available at: https://github.com/Tencent-Hunyuan/R-DMesh.

Reference video

w/o Initial pose rectification

Pose misalignment

CCS Concepts: • Computing methodologies → Motion capture. Additional Key Words and Phrases: 3D animation, rectified flow, feed-forward, video-driven ACM Reference Format: Zijie Wu, Lixin Xu, Puhua Jiang, Sicong Liu, Chunchao Guo, and Xiang Bai. 2026. R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers (SIGGRAPH Conference Papers ’26), July 19–23, 2026, Los Angeles, CA, USA. ACM, New York, NY, USA, 11 pages. https://doi.org/10.1145/3799902.3811135

1

Introduction

While generative models have revolutionized static 3D content creation [Lai et al. 2025; Tochilkin et al. 2024; Zhang et al. 2024b], extending this success to the 4D domain remains formidable. The scarcity of high-quality 4D data, coupled with the immense complexity of modeling joint spatio-temporal distributions, hinders the development of generalizable 4D generative models. Existing solutions typically fall into two categories. Holistic 4D generation approaches, which rely on SDS [Poole et al. 2022] or multi-view video generation, struggle with spatio-temporal consistency. Without explicit 4D supervision, they often suffer from temporal flickering and fail to preserve local structural details, especially in unseen views. In parallel, Mesh-based Animation methods leverage high-quality static 3D assets. While promising, they are often confined to specific templates (e.g., SMPL [Loper et al. 2023]), require costly per-scene optimization, or lack precise control mechanisms, limiting their applicability to diverse, open-world objects. To overcome these limitations, we advocate for Dynamic Mesh (DMesh) as the ideal representation, disentangling motion from geometry to leverage the high fidelity of 3D assets while focusing the learning capacity solely on motion dynamics. Furthermore, we identify monocular video as the most intuitive and informationrich signal for precise motion control. However, a critical yet overlooked hurdle in video-guided 3D animation is the pose misalignment dilemma, as illustrated in Fig. 2. In practical scenarios, the initial pose of a user-provided mesh rarely matches the starting pose of a reference video. Naively forcing a mesh to follow a mismatched video leads to severe distortion or animation failure. Therefore, a robust system must not only generate motion but also “rectify” the mesh to align with the video’s context before animation begins. To address these challenges, we present Rectified Dynamic Mesh (R-DMesh), a unified framework for generating high-fidelity, motion-aligned 4D meshes. Our approach hinges on two core designs. First, we propose a novel R-DMesh VAE that disentangles the input into a conditional mesh, vertex jump offsets, and relative motion trajectories. The jump offsets are the key to our rectification SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Initial frame

Static 3D model

w/ Initial pose rectification (Ours)

Fig. 2. The challenge of pose misalignment in video-guided 3D animation. Significant discrepancies often exist between the initial video frame and the input 3D model. Directly transferring motion without addressing this leads to mismatched deformations or static outputs (top right). In contrast, our method performs pose rectification prior to motion transfer, laying a solid foundation for high-fidelity video-guided 3D animation (bottom right).

capability, learning to transform the arbitrary pose of a conditional mesh to align precisely with the video’s start frame. To coordinate these components, we introduce a Triflow Attention mechanism, which utilizes vertex-wise geometric features to modulate the three decoupled flows, enforcing physical priors of local rigidity and motion synergy. Second, for generation, we employ a Diffusion Transformer (DiT) [Peebles and Xie 2023] architecture based on Rectified Flow [Liu et al. 2022]. We condition the 4D DiT on video latent extracted from a pre-trained video generation model [Wan et al. 2025]. By harnessing the rich spatio-temporal priors inherent in large-scale video models, we achieve coherent 4D generation with significantly reduced computational overhead. To facilitate the above training process, we introduce VideoRDMesh, a large-scale dataset comprising over 500k high-fidelity dynamic sequences derived from Objaverse [Deitke et al. 2024, 2023]. Designed to simulate real-world inference scenarios (e.g., handling pose misalignment), this dataset provides a robust foundation for our framework. Consequently, our approach extends beyond standard video-guided 3D animation to support diverse downstream applications, including pose/motion retargeting and holistic 4D generation (as shown in Fig. 1), positioning it as a versatile solution for high-quality dynamic 3D content creation. • We propose R-DMesh, a unified framework for video-guided 3D animation that effectively resolves the critical pose misalignment dilemma. By explicitly formulating a learnable jump offset, our method enables the seamless animation of static meshes with arbitrary initial poses using unaligned monocular videos. • We design a novel VAE architecture incorporating the Triflow Attention mechanism. By leveraging vertex-wise geometric features to modulate the disentangled flows of base geometry, jump offsets, and motion trajectories, this mechanism enforces physical priors of local rigidity and motion synergy, significantly enhancing generation fidelity. • We develop a scalable generative pipeline employing a Rectified Flow-based DiT conditioned on pre-trained VDM’s latents. This design efficiently transfers rich spatio-temporal priors from video

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

generation models to the 4D mesh domain, ensuring temporal coherence while reducing computational overhead. • We construct Video-RDMesh, a large-scale dataset comprising over 500k high-quality dynamic sequences with paired misalignment simulation. This dataset not only facilitates robust training but also empowers our framework to support diverse applications, including pose/motion retargeting and holistic 4D generation.

2 Related Works 2.1 Holistic 4D Generation Due to the scarcity of 4D data, pioneering approaches [Bahmani et al. 2023; Jiang et al. 2023; Singer et al. 2023; Wu et al. 2024b] focus on distilling spatio-temporal priors from pre-trained image [Liu et al. 2023; Shi et al. 2023], video [Blattmann et al. 2023; Cerspense 2023], or 3D generative models [Hong et al. 2023; Tang et al. 2024] to facilitate 4D generation. Typically, these methods employ Score Distillation Sampling (SDS) [Poole et al. 2022] to compute gradients on renderings from specific viewpoints, which are then back-propagated to update the parameters of the underlying 4D representations (e.g., Dynamic NeRFs [Pumarola et al. 2021] or 4DGS [Wu et al. 2024a]). Despite demonstrating the feasibility of 4D generation, the stochastic nature of SDS often leads to compromised spatio-temporal consistency. V2M4 [Chen et al. 2025] adopts a multi-stage scheme to optimize a dynamic mesh However, the requirement for per-scene optimization significantly hinders the practical scalability of these methods. To improve consistency, recent works [Jiang et al. 2024; Liang et al. 2024; Zhang et al. 2025] propose synthesizing multi-view videos of dynamic objects. However, these methods require an additional lifting step to reconstruct 4D representations. Moreover, without genuine 4D supervision, rendering from viewpoints outside the generated video’s scope often leads to instability and artifacts.

2.2

3D Animation

Instead of generating 4D content from scratch, another paradigm animates existing static 3D meshes. These approaches decouple motion from geometry, leveraging high-quality static 3D assets and requiring significantly less 4D training data. We categorize these methods into optimization-based and learning-based approaches. Methods like [Chen et al. 2024; Millán et al. 2025; Uzolas et al. 2024] utilize SDS and reference-view losses to distill motion from video foundation models. While flexible, they suffer from the same bottlenecks: slow per-scene optimization and artifacts caused by the stochastic nature of SDS. To achieve efficient inference, recent approaches focus on training feed-forward models. While some methods [Guo et al. 2024; Zhang et al. 2023b, 2024a] [Te et al. 2022] have achieved impressive results on parametric models (e.g., SMPL [Loper et al. 2023]) or specific skeletons, their generalizability is severely limited. They are unable to animate models beyond the scope of their training templates or skeletons, let alone drive general-category objects. Recent approaches targeting general skeletons [Gat et al. 2025; Gong et al. 2025] predict joint movements but struggle with nonrigid objects (e.g., fluids, clothing) and require pre-defined skeletons, preventing end-to-end animation of arbitrary meshes. Closely related to our work is AnimateAnyMesh [Wu et al. 2025, 2026], which devise a text-to-trajectory rectified flow model. While

3

innovative, text conditioning suffers from ambiguity and is limited by the expressiveness of textual descriptions. In contrast, we utilize video as a deterministic driving signal, ensuring precise motion control and enabling motion retargeting. Alongside our work, DriveAnyMesh [Shi et al. 2025] and the concurrent Mesh4D [Jiang et al. 2026] also adopt video guidance for dynamic mesh generation. However, utilizing video as a driving signal introduces a unique challenge absent in text-driven methods: the pose misalignment between the static mesh and the reference video’s initial frame. While these existing video-driven approaches largely overlook this critical issue, We address it by designing a specialized VAE and Rectified Flow framework, effectively resolving the discrepancy and streamlining the animation pipeline.

3

Methodology

We propose R-DMesh, a framework synthesizing high-fidelity motion dynamics while rectifying the mesh’s initial pose. It comprises two components: the R-DMesh VAE (Sec.3.1), which compresses vertex trajectories into a disentangled latent space handling pose misalignment, and the R-DMesh RF Model (Sec.3.2), a Rectified Flow generator conditioned on video foundation model latents.

3.1

R-DMesh VAE

To address the spatial discontinuity between a static conditional mesh 𝑀𝑐𝑜𝑛𝑑 and a target dynamic sequence 𝐷, we introduce a hierarchical VAE that explicitly disentangles initial pose rectification from continuous dynamics (Fig. 3). DMesh Decomposition and Representation. We consider a static conditional mesh 𝑀𝑐𝑜𝑛𝑑 = (𝑉𝑐𝑜𝑛𝑑 ∈ R𝑁 ×3, 𝐹 ∈ Z𝑀 ×3 ) and a target dynamic sequence 𝐷 = (𝑉1:𝑇 ∈ R𝑇 ×𝑁 ×3, 𝐹 ∈ Z𝑀 ×3 ), which share an identical topology with 𝑁 vertices and 𝑀 faces. Direct reconstruction of the absolute vertex sequence 𝑉1:𝑇 presents significant challenges. First, modeling absolute coordinates inherently entangles the subject’s intrinsic geometry with its motion dynamics. This forces the VAE to implicitly encode static geometric features within the motion distribution, unnecessarily increasing its complexity and hindering the convergence of subsequent generative models. Second, the spatial misalignment between the condition 𝑉𝑐𝑜𝑛𝑑 and the sequence start state 𝑉1 introduces abrupt discontinuities. Without explicit decomposition, these “jumps” contaminate the motion representation with static offsets, degrading temporal smoothness and complicating the learning of continuous dynamics. To address these issues, we decompose the conditional mesh and target sequence into four disentangled components: the faces (𝐹 ∈ Z𝑀 ×3 ), the initial vertices (𝑉𝑐𝑜𝑛𝑑 ∈ R𝑁 ×3 ), the jump offset (∆𝐽 ∈ R𝑁 ×3 ), and the relative trajectories (𝑇𝑟𝑒𝑙 ∈ R𝑁 × (𝑇 ·3) ), as illustrated in Fig. 3. Specifically, the jump offset (∆𝐽 ) and the relative trajectories (𝑇𝑟𝑒𝑙 ) are modeled as: ∆𝐽 = 𝑉1 − 𝑉𝑐𝑜𝑛𝑑 , 𝑇𝑟𝑒𝑙 = 𝑉1:𝑇 − 𝑉1 .

(1)

Prior to this decomposition, we employ a canonicalization strategy. 𝑉𝑐𝑜𝑛𝑑 is centered around its own centroid, whereas the target sequence 𝑉1:𝑇 is centered relative to the centroid of the first frame 𝑉1 . Subsequently, both components are normalized by the maximum absolute coordinate of the centered 𝑉𝑐𝑜𝑛𝑑 . This strategy is necessitated by the lack of a corresponding control frame for 𝑉𝑐𝑜𝑛𝑑 in the SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Wu, Z. et al

R-DMesh VAE Encoder PEpts

Decomposition Consecutive Sequence (𝐷) 2. Initial Vertices (𝑉!"#$ )

Triflow Attention × N

𝑉%!"#$

𝑉!"#$

×N

R-DMesh VAE Decoder

FPS Q

Consecutive Sequence (𝐷)

Q

K,V

mask

𝑥!"#$

K

Cross Attention

1. Faces (𝐹)

Self Attention

Conditional Mesh (𝑀!"#$ )

Cross Attention

Self Attention

4

𝑀!"#$ (Given) Recomposition

index

𝑧(

KL Regularization

3. Jump Offset (Δ𝐽)

Adjacency Matrix (𝐴𝑑𝑗)

𝑧)%*+

KL Regularization

𝐹

PEpts

⋮ Δ𝐽)

Δ𝐽

4. Trajectories (𝑇%&' )

index

PEtraj

𝑇%&'

𝑇%%&'

PEpts: Point Positional Encoding PEtraj: Trajectory Positional Encoding FPS: Farthest Point Sampling

: Conditional frame token

Δ𝐽

𝑇%&' : Jump offset token

: Trajectory token

Reconstruction Results

: Attention map

: Matrix multiplication

Fig. 3. Illustration of our proposed R-DMesh VAE. It compresses and reconstructs dynamic mesh sequences conditioned on a static mesh of the same object in an arbitrary pose. (Left) Decomposition: The input sequence is decoupled into vertices 𝑉𝑐𝑜𝑛𝑑 , face 𝐹 , global offsets ∆𝐽 , and relative motion 𝑇𝑟𝑒𝑙 . (Middle) Encoder: The Triflow Attention mechanism jointly processes these components to capture spatio-temporal correlations, producing compact latent codes. (Right) Decoder: The 4D sequence is reconstructed via a Triflow cross-attention, queried by the condition’s vertex features.

video-conditioning stage, which renders global translation unlearnable. By applying separate centering, we eliminate this non-essential displacement to stabilize inputs. Consequently, this effectively disentangles local alignment (∆𝐽 ) from motion generation (𝑇𝑟𝑒𝑙 ). Hierarchical Encoder with Triflow Attention. As illustrated in Fig. 3, we first project 𝑉𝑐𝑜𝑛𝑑 , ∆𝐽 , and 𝑇𝑟𝑒𝑙 into higher dimensions using positional encodings [Zhang et al. 2023a] tailored for vertices ˆ , and 𝑇ˆ𝑟𝑒𝑙 . To encode and trajectories, respectively, yielding 𝑉ˆ𝑐𝑜𝑛𝑑 , ∆𝐽 local geometry, we perform self-attention on 𝑉ˆ𝑐𝑜𝑛𝑑 masked by the mesh adjacency matrix (𝐴𝑑 𝑗), inspired by [Wu et al. 2025]. We employ Farthest Point Sampling (FPS) [Qi et al. 2017] on 𝑉ˆ𝑐𝑜𝑛𝑑 to determine 𝑛 representative indices, which are then used to gather corresponding feature subsets across all three modalities 𝑛 , ∆ 𝐽ˆ𝑛 , 𝑇ˆ 𝑛 ). Subsequently, we introduce Triflow Attention (𝑉ˆ𝑐𝑜𝑛𝑑 𝑟𝑒𝑙 to aggregate global contexts. Instead of computing separate attention maps, we generate a shared attention map using the sampled 𝑛 geometry 𝑉ˆ𝑐𝑜𝑛𝑑 as the query and the full set 𝑉ˆ𝑐𝑜𝑛𝑑 as the key. This geometry-guided map is then applied to simultaneously aggregate features from all three modalities. By explicitly aligning motion and trajectory aggregation with geometric topology, this design encodes the intrinsic correlations between local rigidity and dynamics without compromising feature disentanglement, effectively reducing the learning difficulty: ! 𝑛 𝑇 𝑉ˆ𝑐𝑜𝑛𝑑 · 𝑉ˆ𝑐𝑜𝑛𝑑 𝐴 = Softmax , √ 𝑑𝑘

𝑟𝑒𝑐 L = L1𝑟𝑒𝑐 + L𝑟𝑒𝑙 + 𝜂 1 · L𝑘𝑙∆ + 𝜂𝑟𝑒𝑙 · L𝑘𝑙𝑡𝑟𝑎 𝑗 ,

(2)

We stack multiple Triflow Attention layers to strengthen the semantic information of each compressed token, ultimately obtaining the compressed vertex features 𝑥𝑐𝑜𝑛𝑑 , the jump features 𝑥 ∆ , and the relative trajectory features 𝑥𝑡𝑟𝑎 𝑗 . Since our task targets mesh

(3)

where 𝜂 1, 𝜂𝑟𝑒𝑙 are set to 1e-6 as default. Note that the reconstruction terms are averaged over all 𝑁 vertices.

3.2

𝑛 𝑛 𝑛 𝑛 (𝑉ˆ𝑐𝑜𝑛𝑑 , ∆ 𝐽ˆ𝑛 , 𝑇ˆ𝑟𝑒𝑙 ) = 𝐴 · (𝑉ˆ𝑐𝑜𝑛𝑑 , ∆ 𝐽,ˆ 𝑇ˆ𝑟𝑒𝑙 ) + (𝑉ˆ𝑐𝑜𝑛𝑑 , ∆ 𝐽ˆ𝑛 , 𝑇ˆ𝑟𝑒𝑙 ).

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

animation where 𝑥𝑐𝑜𝑛𝑑 is always given, we only need to model the distributions of the probabilistic components. To this end, we model the probabilistic components 𝑥𝑘 ∈ {𝑥 ∆, 𝑥𝑡𝑟𝑎 𝑗 } using Gaussian distributions 𝑧𝑘 = 𝜇𝑘 + 𝜎𝑘 · 𝜖𝑘 . The network is optimized using the standard KL divergence loss L𝑘𝑙𝑘 to regularize the latent space. Decoder and Reconstruction. In the decoding phase, we project the three latent codes to a unified dimension and process them via stacked Triflow attention layers to reinforce spatial correlations. To propagate features back to the dense mesh topology, we employ a cross-attention mechanism using the fine-grained encoder features 𝑉ˆ𝑐𝑜𝑛𝑑 as queries and the processed latents 𝑥𝑐𝑜𝑛𝑑 as keys. Finally, specific projection heads are used to reconstruct the jump offsets 𝑟𝑒𝑐 , respectively. ∆𝐽 𝑟𝑒𝑐 and relative trajectories 𝑇𝑟𝑒𝑙 To ensure stable convergence for both the jump offset and relative motion, we apply separate supervision to the reconstructed outputs. The reconstruction objectives are formulated as the Mean Squared Error (MSE) between the predicted and ground-truth values. The overall optimization objective for the R-DMesh VAE is a weighted sum of reconstruction and KL divergence losses:

R-DMesh RF Model

Rectified Flow Formulation for Dynamic Meshes. We employ Rectified Flow [Liu et al. 2022] to model the distribution of dynamic components 𝑍𝑑 𝑦𝑛 = [𝑧 ∆, 𝑧𝑡𝑟𝑎 𝑗 ] conditioned on the static structural latent 𝑥𝑐𝑜𝑛𝑑 and video features 𝐹 𝑣𝑖𝑑 . Since the initial mesh is given during inference, we treat the its representation 𝑥𝑐𝑜𝑛𝑑 as a fixed clean condition. During training, we construct the time-dependent state 𝑍𝑡 via linear interpolation: 𝑍𝑡 = 𝑡 · 𝑍𝑑 𝑦𝑛 + (1 − 𝑡) · 𝜖, where

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow ×𝑁 Transformer blocks

🔥

𝛾#

🔥 🧊

𝛼# , 𝛽# 𝛾"

trainable frozen

VDM DiT Self Attn

🧊

𝛼" , 𝛽"

𝛾!

𝑘𝑡ℎ block

Cross Attn 𝛼! , 𝛽!

⋮ Channel-wise Concatenation

t

𝑥$%&'

𝑧(

𝑧)*+,

R-DMesh VAE Encoder

VDM 🧊 VAE Encoder

Dynamic Mesh Sequence

Rendered Video

5

exhibiting diminished temporal consistency for motion dynamics. Consequently, we target intermediate blocks to capture rich spatiotemporal structures. These features are subsequently injected into our 4D branch via cross-attention. Model Architecture. Our velocity estimator 𝑣𝜃 employs a DiTbased architecture specifically tailored for sequence modeling. As illustrated in Fig. 4, the network comprises 𝑁 stacked Transformer blocks, each containing a Cross-Attention layer, a Self-Attention layer, and a Feed-Forward Network (FFN). To effectively inject the time step 𝑡 and stabilize training, we utilize AdaLN-Zero [Yang et al. 2024] modulation. The time embedding is projected to regress layer-specific scaling and shifting parameters (𝛼, 𝛽), as well as gating parameters (𝛾). Prior to each attention or FFN operation, input features are modulated via 𝛼 and 𝛽, functioning as a time-dependent Adaptive Layer Normalization (AdaLN) [Perez et al. 2018]. Furthermore, the output of each sub-module is scaled by the learnable gate 𝛾 before being added to the residual stream. In terms of feature interaction, the input latents first aggregate motion cues from the video latents via cross-attention. Subsequently, self-attention layers model local motion dependencies, ensuring the generated motion is smooth and globally coherent. During training, we randomly drop the video conditioning latents with a probability of 𝑝 = 0.1 to enable Classifier-Free Guidance (CFG) [Ho and Salimans 2022] for flexible control during inference. This approach ultimately yields a highfidelity, video-aligned network for video-guided mesh animation.

scale & shift gate

FFN

Fig. 4. Illustration of our proposed R-DMesh RF model. We leverage a pretrained, frozen Video Diffusion Model (VDM) as a strong visual prior. The VDM processes the reference video to extract rich semantic and dynamic features. These features are injected into the trainable Transformer blocks via Cross-Attention, guiding the generation of mesh dynamics (𝑧 Δ , 𝑧𝑡𝑟𝑎 𝑗 ).

𝜖 ∼ N (0, 𝐼 ) and 𝑡 ∈ [0, 1]. The network input is the channel-wise concatenation of the noisy dynamic state 𝑍𝑡 and the clean condition 𝑥𝑐𝑜𝑛𝑑 , trained to minimize the flow matching objective: 2

𝐿𝑅𝐹 = E𝑡,𝑍𝑑 𝑦𝑛 ,𝜖 [ 𝑣𝜃 ([𝑍𝑡 ; 𝑥𝑐𝑜𝑛𝑑 ], 𝑡, 𝐹 𝑣𝑖𝑑 ) − (𝑍𝑑 𝑦𝑛 − 𝜖) ],

(4)

where 𝐹 𝑣𝑖𝑑 represents the video guidance features. Leveraging Pre-trained VDM Priors. To ensure the generated mesh motion aligns with the input video, we leverage a pre-trained Video Diffusion Model (VDM) [Wan et al. 2025] as the feature extractor. We adopt the VDM for two main reasons: First, generative models are trained on massive-scale video datasets to synthesize realistic pixels, forcing them to learn superior spatiotemporal correlations and physics priors. Second, modern VDMs employ a latent diffusion architecture that compresses video into compact spatio-temporal tokens. This significantly reduces computational costs while retaining richer semantic and dynamic information than pixel-space encoders. Furthermore, we strategically extract features 𝐹 𝑣𝑖𝑑 from a specific intermediate layer, rather than the final output. We observe that shallower blocks primarily encode low-level spatial details, lacking global temporal integration. Conversely, the deepest layers, become overly specialized for the low-level denoising objective, thereby

4 Experiments 4.1 Experimental Setup Implementation Details. Our R-DMesh VAE features a symmetric architecture with 8 Triflow Attention layers in both the encoder and decoder. The encoder downsamples vertices by 8× via Farthest Point Sampling (FPS) to generate queries. Input features (vertices, offsets, 64-frame trajectories) are projected to 256 dimensions and compressed into latents of sizes 64, 16, and 64. The R-DMesh RF model comprises 12 Transformer blocks with a dimension of 512. Training sequences are sliced into 64-frame clips and filtered to exclude static samples (displacement < 0.01). Data samples are filtered to < 8, 192 vertices and a face-to-vertex ratio < 2.5, then padded to fixed sizes (8, 192 vertices, 20, 480 faces) during training. To improve robustness against pose discrepancies, we employ a misalignment simulation strategy by conditioning on a random frame from the sequence rather than the first frame. We condition the RF model using features from the 10-th layer of the Wan2.2-TI2V5B [Wan et al. 2025] DiT, extracted from a 256 × 256 silhouette video rendering (tensor shape: 1088 × 3072). Training was performed on 32 NVIDIA H20 GPUs: the VAE for 200k iterations (cosine schedule, 2e-4 → 2e-5, ∼ 54h) and the RF for 300k iterations (constant 1e-4, ∼ 120h). Please refer to the 𝑆𝑢𝑝𝑝𝑙 .1 for further details. Datasets and Evaluation Metrics. We train our model on VideoRDMesh, a large-scale dataset curated from 252, 823 unique dynamic assets (largely rig-based animations) sourced from Objaverse [Deitke et al. 2024, 2023]. Through a pipeline of extraction, slicing, and rigorous filtering, we processed these assets to yield 513, 690 high-quality, 64-frame vertex trajectory-based clips paired 1 𝑆𝑢𝑝𝑝𝑙 .: the supplemental file

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

6

Wu, Z. et al

Table 1. Quantitative comparison with state-of-the-art methods. We evaluate rendering quality (PSNR), temporal consistency (Subject Consistency "SC." and Motion Smoothness "SM." from VBench [Huang et al. 2024]), geometric accuracy (Euclidean Distance "EucD"), and inference time.

Exp

Rendering Metrics

4D Metrics

Time (↓)

PSNR (↑)

SC. (↑)

SM. (↑)

EucD (↓)

SC4D L4GM AAM PUPT

23.4 22.3 13.5 19.8

0.933 0.915 0.948 0.932

0.993 0.995 0.994 0.991

/ / / 0.035

∼40m ∼30s ∼8s ∼25m

Ours

25.8

0.949

0.995

0.012

∼10s

with reference videos. To ensure visual consistency, the corresponding ground-truth videos are rendered via Blender. Please refer to Sec. 1 of the supplemental file for detals of data curation. To benchmark performance in video-guided mesh animation, we introduce the Video-RDMesh test set. This dataset comprises two distinct subsets, each containing 64 examples. The first subset consists of ground-truth (GT) dynamic mesh sequences paired with their corresponding frontal-view rendered videos. The second subset pairs conditional meshes with reference videos generated by WAN2.2-I2V14B [Wan et al. 2025]. For the latter, we synthesize reference videos by conditioning the video generator on front-view renderings of the input meshes and corresponding text prompts. The Video-RDMesh test set encompasses a diverse range of subjects, including humans, reptiles, flying creatures, and dynamic inanimate objects. For evaluation, we employ specific protocols for each subset. For examples with GT dynamic meshes, we quantify performance using the average Euclidean Distance (EucD) of vertices relative to the GT at each timestamp. Conversely, for the subset utilizing generated reference videos, we render our results using identical camera settings and compute the frame-by-frame Peak Signal-toNoise Ratio (PSNR) to assess visual alignment. Furthermore, we adopt the Subject Consistency (SC.) and Motion Smoothness (SM.) metrics from VBench [Huang et al. 2024] to evaluate the temporal consistency and motion fluidity of the generated 4D objects.

4.2

Comparisons

Due to the scarcity of methods specifically targeting video-guided universal mesh animation, we compare our approach against four state-of-the-art methods from closely related tasks. These baselines include two video-to-4D generation methods, SC4D [Wu et al. 2024b] and L4GM [Ren et al. 2025], the text-driven mesh animation method AnimateAnyMesh [Wu et al. 2025] (AAM), and the video-guided mesh animation method Puppeteer [Song et al. 2025] (PUPT). While other methods directly utilize the reference video as a condition, for AAM, we employ Qwen-VL-2.5 [Bai et al. 2025] to caption the reference videos. These generated captions serve as the text prompts to ensure semantic alignment during the evaluation. We evaluate all methods using their official implementations on the datasets described in Sec. 4.1, covering scenarios both with and without ground-truth 4D meshes. The qualitative and quantitative comparisons are presented in Fig. 5 and Tab. 1, respectively. Qualitative Comparison. As shown in Fig. 5, while SC4D and L4GM generate renderings that closely align with the video in the SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

reference view, they consistently exhibit color and shape distortions in novel views. This issue is particularly pronounced for reference videos featuring slender foreground objects (e.g., the pterosaur in Fig. 5), where the generated results often suffer from severe deformation or drift significantly from the scene center, making optimization intractable. Although AAM generates high-quality animation results, it struggles to achieve fine-grained control via text prompts. Furthermore, for rare object categories, the learned distribution often fails to cover the corresponding motion manifold, leading to deformations or semantic motion artifacts (e.g., the pterosaur in Fig. 5). PUPT, relying on 2D priors like optical flow, degrades in complex scenarios. It specifically fails in radial motions due to the lack of explicit depth (cols 5-6) and cannot handle pose misalignment between the mesh and the video’s initial frame (cols 1-4). In contrast, our proposed R-DMesh, trained on a large-scale dataset of high-quality video-DMesh pairs, generates dynamic mesh sequences that maintain high correspondence with the reference video while preserving global and local geometric fidelity. Benefiting from our sophisticated model design and curated training data, our method effectively resolves pose misalignment and ensures robust motion transfer. As demonstrated in Fig. 5, our method significantly outperforms comparative approaches in terms of alignment with the reference video, shape preservation, and motion plausibility. Quantitative Comparison. Tab. 1 summarizes the numerical results. We first evaluate the alignment fidelity between the rendered results and the reference videos. Our method achieves the highest PSNR (25.8), indicating that our driven meshes align most accurately with the target motion. regarding temporal consistency, our approach outperforms all baselines in both SC. (0.949) and SM. (0.995), ensuring that the generated 4D sequences maintain robust subject identity and fluid motion. For geometric accuracy, we report the Euclidean Distance (EucD). Our method yields a significantly lower error (0.012) compared to PUPT (0.035), demonstrating more precise geometry deformation. In terms of efficiency, our approach takes only about 10 seconds, which is comparable to the fastest baseline (AAM) and orders of magnitude faster than optimization-based methods. Overall, these quantitative metrics align well with the qualitative visualizations, further validating the superiority of our proposed framework.

4.3

Ablation Studies

R-DMesh VAE Components. We investigate four key designs in our VAE: (1) Dual-Norm, which normalizes the conditional mesh and the subsequent sequence independently; (2) Jump-Decomp, which explicitly models the displacement between the condition and the sequence start; (3) Tri-Attn, which disentangles the learning of geometry, pose, and motion; and (4) Decoup-Loss, which supervises offset and trajectory reconstruction separately. Tab. 2 quantifies the contribution of four core design elements in our R-DMesh VAE. The Jump-Decomp module emerges as the most pivotal factor. As evidenced in Tab. 2 and Fig. 6, removing this module results in severe reconstruction degradation; notably, the generated initial frame (𝑡 0 ) fails to capture the motion transition, remaining virtually identical to the conditional mesh pose. This failure occurs because, without explicit jump modeling, vertex trajectories

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

t1

t2

t1

t2

t1

t2

N/A

N/A

t1

t2

t1

t2

N/A

N/A

7

Ref Vid

Novel View GT

Ref View

Severe Drift

Severe Drift

Novel View

Severe Drift

Severe Drift

SC4D

Ref View

L4GM Novel View

Ref View

AAM Novel View

Ref View

PUPT Novel View

Ref View

Ours Novel View

Fig. 5. Qualitative comparison with state-of-the-art methods. We evaluate against video-to-4D methods (SC4D [Wu et al. 2024b], L4GM [Ren et al. 2025]), text-driven AAM (AnimateAnyMesh [Wu et al. 2025]), and video-guided PUPT (Puppeteer [Song et al. 2025]). Columns show renderings at initial (𝑡 1 ) and random (𝑡 2 ) timestamps across reference and novel views. Blue boxes highlight the input mesh’s initial pose. "Severe Drift" marks an SC4D failure case due to excessive spatial deviation. "N/A" indicates missing ground truth for synthetic data. Zoom in for a better view.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

8

Wu, Z. et al Conditional Mesh

t0

t1

Conditional Mesh

t0

t1

GT

Table 3. Ablation study on video feature extraction layers. We utilize the Wan2.2-TI2V-5B model for feature extraction. DiT-L0, L10, L20, and L30 represent features from the corresponding intermediate blocks. DiT-L0

w/o

DiT-L10

DiT-L20

DiT-L30

EncD (↓)

0.059 0.012 0.024 0.037 0.018 0.014 0.014

Decomp

✓ ✓

w/o

Tri-Attn

✓ ✓ ✓

✓ ✓

features from the 10th transformer block of Wan2.2-TI2V-5B as the condition for our R-DMesh RF model.

Full

Fig. 6. Visual ablation on Jump Decomposition and Triflow Attention. The start and end frame of the generated sequences are shown under 𝑡 0 , 𝑡 1 . Table 2. Ablation studies of technical components. Dual-Norm: Dual-Center Norm, Jump-Decomp: Jump Offset Decomposition, Tri-Attn: Triflow Attention, Decoup-Loss: Decoupled Reconstruction Loss. Dual-Norm

Jump-Decomp

Tri-Attn

Decoup-Loss

EncD (↓)

✗ ✓ ✓ ✓ ✓

✓ ✗ ✓ ✓ ✓

✓ ✓ ✗ ✓ ✓

✓ ✓ ✓ ✗ ✓

0.007 0.025 0.020 0.007 0.005

are not de-centered. Consequently, the large coordinate offsets of the jump frame dominate the compression latent space, hindering effective trajectory clustering and separation. Furthermore, treating the jump frame indistinguishably during optimization neglects its foundational role; a poorly reconstructed initial state propagates errors, rendering the alignment of subsequent frames intractable. The efficacy of Tri-Attn is also highlighted in our results. By leveraging priors of local rigidity and motion correlation, this module effectively disentangles the processing of jump vectors from relative motion trajectories. This isolation prevents mutual interference between these distinct information flows during optimization, significantly boosting reconstruction fidelity. Finally, Tab. 2 confirms the positive impact of Dual-Norm and Decoup-Loss. Together, these components enable our full model to achieve high-fidelity compression and reconstruction for dynamic mesh sequences involving significant initial state transitions. Video Feature Extraction. As outlined in Sec. 3.2, we leverage the pre-trained Wan2.2-TI2V-5B [Wan et al. 2025] video model as our reference video feature extractor. To determine the optimal conditioning source, we conduct an ablation study on the feature layers. Given that Wan2.2-TI2V-5B comprises 30 transformer blocks, we experiment with latent outputs from the 0th (input), 10th, 20th, and 30th layers. These latents are injected into our R-DMesh RF model and trained for an identical number of iterations. As reported in Tab. 3, the latent output from the 10th layer proved most effective for our task. We also investigate concatenating latents from multiple layers; however, this approach yields inferior results compared to using the 10th layer alone. Consequently, we adopt the video SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

4.4

Application

Beyond the primary task of video-guided mesh animation, our framework exhibits significant versatility and potential for downstream tasks. We highlight three key applications: pose retargeting, motion retargeting, and holistic video-to-4D generation. Pose Retargeting. As demonstrated in Fig. 7, our method robustly transfers poses from diverse sources to the conditional mesh. We showcase inputs ranging from synthetic 3D asset renderings (first two examples) to unconstrained in-the-wild images (last three examples). Even when presented with complex real-world poses, our approach successfully deforms the mesh to match the target posture. This robustness—particularly in handling the domain gap between synthetic and real data—is largely attributed to our explicit decoupling and modeling of vertex jump offsets, which effectively isolates large-scale pose changes from local details. Reference Poses Conditional Mesh

Ref View

Novel View

Fig. 7. Pose retargeting application examples of our method.

Motion Retargeting. Fig. 8 illustrates our capability to transfer motion sequences from a driving video to a target 3D character. Our method generates highly consistent animation sequences that faithfully follow the driving motion while preserving the local geometry and identity of the target mesh. Notably, this is achieved without specific fine-tuning for different identities. The model generalizes well even when there are significant body shape discrepancies or complex accessories between the source and target. We attribute this zero-shot generalization capability to the high-quality distribution of our training data and the rich semantic priors inherited from the video generation model.

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

Reference Video

9

Retargeting Results

Fig. 8. Motion retargeting application examples of our method. Left: Reference videos generated by video generation models. Right: Generated 3D animations. The conditional mesh (T-pose) is shown at the top-left of each sequence as the input condition, followed by the generated motion frames and the final pose. Time

Time

Ref Video

Ref View

Novel View

Fig. 9. Holistic video-to-4D generation application examples of our method. The top row shows the reference videos. The middle row displays the reconstructed dynamic 3D content rendered from the reference camera view. The bottom row demonstrates the rendered novel views.

Holistic Video-to-4D Generation. By integrating with state-ofthe-art 3D generation models, our framework facilitates a complete video-to-4D pipeline. We utilize real-world videos as input, employing Hunyuan3D [Lai et al. 2025] to generate a static 3D mesh from the first frame as the conditional input. Using the segmented video [Carion et al. 2025] as guidance, our method then animates this static asset. As shown in Fig. 9, this pipeline produces highfidelity dynamic mesh sequences that are closely aligned with the reference video, successfully handling complex scenarios ranging from human dance moves to animal locomotion.

5

Limitation and Conclusion

Limitation. As shown in Fig. 10, our method currently faces two challenges: (1) Mesh Interpenetration: Occasional self-collisions appear in the output, attributed to noisy ground-truth data containing self-intersecting geometry that is difficult to fully purge. (2) Generalization to Rare Objects: The synthesis quality degrades under the guidance of rare categories, where data sparsity leads to unnatural deformations. Future work will focus on data filtering to eliminate intersections, post-hoc optimization, and dataset augmentation to improve robustness across diverse object categories. Conclusion. In this work, we present R-DMesh, a robust framework for generating high-quality dynamic meshes controlled by

(a) Mesh Interpenetration

(b) Failure case under rare object guidance

Fig. 10. Limitations of our method. (a) Mesh Interpenetration. Our generated results may occasionally exhibit self-intersection artifacts. (b) Generalization to Rare Objects. When guided by rare objects (the armored cat), the plausibility of the generated pose and motion may be compromised.

monocular video. By identifying and addressing the pose misalignment dilemma—a pervasive issue in practical animation workflows, we introduce a disentangled representation that explicitly models the rectification process via a learnable jump offset. Our technical core, the Triflow Attention mechanism, successfully encodes physical priors into the generative process, ensuring that the synthesized motion remains geometrically consistent and locally rigid. Furthermore, by conditioning the Rectified Flow DiT on pre-trained video latents, we effectively bridge the gap between 2D video priors and 3D motion generation, achieving both high-fidelity and computational efficiency. The release of the Video-RDMesh dataset provides a substantial foundation for future research in this domain. We SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

10

Wu, Z. et al

believe that R-DMesh not only advances the state of the art in videoguided 3D animation but also offers a versatile, general-purpose solution for creating diverse and coherent 4D assets.

References Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gordon Wetzstein, Leonidas Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, and David B Lindell. 2023. 4d-fy: Text-to-4d generation using hybrid score distillation sampling. arXiv preprint arXiv:2311.17984 (2023). Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. 2025. Qwen2. 5-VL Technical Report. arXiv preprint arXiv:2502.13923 (2025). Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023). Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719 (2025). Cerspense. 2023. Zeroscope text-to-video model. https://huggingface.co/cerspense/ zeroscope_v2_576w. Accessed: 2023-10-31. Ce Chen, Shaoli Huang, Xuelin Chen, Guangyi Chen, Xiaoguang Han, Kun Zhang, and Mingming Gong. 2024. Ct4d: Consistent text-to-4d generation with animatable meshes. arXiv preprint arXiv:2408.08342 (2024). Jianqi Chen, Biao Zhang, Xiangjun Tang, and Peter Wonka. 2025. V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video. ArXiv abs/2503.09631 (2025). Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. 2024. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Information Processing Systems 36 (2024). Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. 2023. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13142–13153. Inbar Gat, Sigal Raab, Guy Tevet, Yuval Reshef, Amit Haim Bermano, and Daniel CohenOr. 2025. Anytop: Character animation diffusion with any topology. In Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. 1–10. Kehong Gong, Zhengyu Wen, Weixia He, Mingxi Xu, Qi Wang, Ning Zhang, Zhengyu Li, Dongze Lian, Wei Zhao, Xiaoyu He, et al. 2025. MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos. arXiv preprint arXiv:2512.10881 (2025). Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang, and Li Cheng. 2024. Momask: Generative masked modeling of 3d human motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1900–1910. Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022). Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. 2023. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400 (2023). Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. 2024. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21807–21818. Yanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang, Weiming Hu, and Jin Gao. 2024. Animate3d: Animating any 3d model with multi-view video diffusion. arXiv preprint arXiv:2407.11398 (2024). Yanqin Jiang, Li Zhang, Jin Gao, Weimin Hu, and Yao Yao. 2023. Consistent4d: Consistent 360 {\deg } dynamic object generation from monocular video. arXiv preprint arXiv:2311.02848 (2023). Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, and Andrea Vedaldi. 2026. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video. ArXiv abs/2601.05251 (2026). Zeqiang Lai, Yunfei Zhao, Haolin Liu, Zibo Zhao, Qingxiang Lin, Huiwen Shi, Xianghui Yang, Mingxin Yang, Shuhui Yang, Yifei Feng, et al. 2025. Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details. arXiv preprint arXiv:2506.16504 (2025). Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang, Zhangyang Wang, Konstantinos N Plataniotis, Yao Zhao, and Yunchao Wei. 2024. Diffusion4d: Fast spatial-temporal consistent 4d generation via video diffusion models. arXiv preprint arXiv:2405.16645 (2024). Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

of the IEEE/CVF International Conference on Computer Vision. 9298–9309. Xingchao Liu, Chengyue Gong, and Qiang Liu. 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 (2022). Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2023. SMPL: A skinned multi-person linear model. Seminal Graphics Papers: Pushing the Boundaries, Volume 2 (2023), 851–866. Marc Benedí San Millán, Angela Dai, and Matthias Nießner. 2025. Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models. arXiv preprint arXiv:2503.15996 (2025). William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision. 4195–4205. Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022). Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. 2021. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10318–10327. Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems 30 (2017). Jiawei Ren, Cheng Xie, Ashkan Mirzaei, Karsten Kreis, Ziwei Liu, Antonio Torralba, Sanja Fidler, Seung Wook Kim, Huan Ling, et al. 2025. L4gm: Large 4d gaussian reconstruction model. Advances in Neural Information Processing Systems 37 (2025), 56828–56858. Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. 2023. Zero123++: a single image to consistent multi-view diffusion base model. arXiv preprint arXiv:2310.15110 (2023). Yahao Shi, Yang Liu, Yanmin Wu, Xing Liu, Chen Zhao, Jie Luo, and Bin Zhou. 2025. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video. ArXiv abs/2506.07489 (2025). Uriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual, Iurii Makarov, Filippos Kokkinos, Naman Goyal, Andrea Vedaldi, Devi Parikh, Justin Johnson, et al. 2023. Textto-4d dynamic scene generation. arXiv preprint arXiv:2301.11280 (2023). Chaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin, and Jianfeng Zhang. 2025. Puppeteer: Rig and animate your 3d models. arXiv preprint arXiv:2508.10898 (2025). Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. 2024. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. In European Conference on Computer Vision. Springer, 1–18. Gusi Te, Xiu Li, Xiao Li, Jinglu Wang, Wei Hu, and Yan Lu. 2022. Neural Capture of Animatable 3D Human from Monocular Video. ArXiv abs/2208.08728 (2022). Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-Pei Cao. 2024. Triposr: Fast 3d object reconstruction from a single image. arXiv preprint arXiv:2403.02151 (2024). Lukas Uzolas, Elmar Eisemann, and Petr Kellnhofer. 2024. Motiondreamer: Zero-shot 3d mesh animation from video diffusion models. arXiv preprint arXiv:2405.20155 (2024). Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. 2025. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025). Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 2024a. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 20310–20320. Zijie Wu, Chaohui Yu, Yanqin Jiang, Chenjie Cao, Fan Wang, and Xiang Bai. 2024b. Sc4d: Sparse-controlled video-to-4d generation and motion transfer. In European Conference on Computer Vision. Springer, 361–379. Zijie Wu, Chaohui Yu, Fan Wang, and Xiang Bai. 2025. AnimateAnyMesh: A FeedForward 4D Foundation Model for Text-Driven Universal Mesh Animation. arXiv preprint arXiv:2506.09982 (2025). Zijie Wu, Chaohui Yu, Fan Wang, and Xiang Bai. 2026. AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation. arXiv:2604.26917 [cs.CV] https://arxiv.org/abs/2604.26917 Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. 2024. Cogvideox: Textto-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072 (2024). Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 2023a. 3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models. ACM Transactions On Graphics (TOG) 42, 4 (2023), 1–16. Haiyu Zhang, Xinyuan Chen, Yaohui Wang, Xihui Liu, Yunhong Wang, and Yu Qiao. 2025. 4diffusion: Multi-view video diffusion model for 4d generation. Advances in

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

11

Neural Information Processing Systems 37 (2025), 15272–15295. Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan. 2023b. Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 14730–14740. Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. 2024b. Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG) 43, 4 (2024), 1–20. Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2024a. Motiondiffuse: Text-driven human motion generation with diffusion model. IEEE transactions on pattern analysis and machine intelligence 46, 6 (2024), 4115–4128.

SIGGRAPH Conference Papers ’26, July 19–23, 2026, Los Angeles, CA, USA.

Record · ID 180638 · SHA-256 aa21dabb5e2663df
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.