arXiv:2605.06667v1 [cs.CV] 7 May 2026
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation OMAR EL KHALIFI, Kinetix, France THOMAS ROSSI, Kinetix, France OSCAR FOSSEY, Kinetix, France THIBAULT FOUQUE, Kinetix, France ULYSSE MIZRAHI, Kinetix, France PHILIP TORR, University of Oxford, United Kingdom IVAN LAPTEV, MBZUAI, United Arab Emirates FABIO PIZZATI, MBZUAI, United Arab Emirates BAPTISTE BELLOT-GURLET, Kinetix, France
Fig. 1. Overview. ActCam enables zero-shot joint control of acting motion and camera motion for single-image video generation from a reference image, assuming only widespread conditioning capability of the backbone model on depth and keypoints. Given a reference image, an acting video representing the desired motion, and a target per-frame camera trajectory, ActCam generates a video that preserves identity while following both motion and cinematography. For artistic applications, video generation requires fine-grained control over both performance and cinematography—i.e., the actor’s motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly (i) transfers character motion from a driving video into a new scene and (ii) enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase Authors’ Contact Information: Omar El Khalifi, [email protected], Kinetix, France; Thomas Rossi, [email protected], Kinetix, France; Oscar Fossey, oscar@kinetix. tech, Kinetix, France; Thibault Fouque, [email protected], Kinetix, France; Ulysse Mizrahi, [email protected], Kinetix, France; Philip Torr, [email protected], University of Oxford, United Kingdom; Ivan Laptev, [email protected], MBZUAI, United Arab Emirates; Fabio Pizzati, [email protected], MBZUAI, United Arab Emirates; Baptiste Bellot-Gurlet, [email protected], Kinetix, France. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM 1557-7368/2026/5-ART https://doi.org/10.1145/nnnnnnn.nnnnnnn
conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose+camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations—especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/. CCS Concepts: • Networks → Network algorithms; • Computing methodologies → Computer vision. ACM Reference Format: Omar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque, Ulysse Mizrahi, Philip Torr, Ivan Laptev, Fabio Pizzati, and Baptiste Bellot-Gurlet. 2026. ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation. ACM Trans. Graph. 1, 1 (May 2026), 12 pages. https://doi.org/10. 1145/nnnnnnn.nnnnnnn
1
Introduction
Driven by the intrinsic ambiguities of natural language, video generation has rapidly progressed from text-guided synthesis to strong conditioned generation pipelines [Jiang et al. 2025; Zhang et al. ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
2
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
2023], capable of accepting additional signals in the generation process to improve fidelity to the user’s instructions. This has already opened up possibilities for use in content production. For example, in experimental filmmaking [Juhi Marzia 2025; Tangermann 2025], visually appealing, consistent shots of synthetic humans can be seamlessly combined to form a short movie. However, results are still far from perfect, mostly due to the limited control over the stylistic characteristics of the output videos. Among other factors, in cinematography, a compelling shot of an actor is defined not only by what the subject does but also by how the camera moves (trajectory, parallax, viewpoint changes). Yet, the capability of controlling both the subject acting performance and the camera movements remains largely unexplored in the current literature. To mitigate this problem, a class of motion-control approaches conditions video generation on 2D signals, typically depending on keypoints of human bodies [Huang et al. 2025; Wang et al. 2025a]. While effective under mild viewpoint changes, such controls can become ambiguous under moving cameras: the same 2D signal can correspond to multiple 3D motions, and the control signal may no longer remain consistent once the camera rotates or translates. The closest approach enabling acting videos of humans with a moving camera is Uni3C [Cao et al. 2025], in which 3D representations of humans in motion are combined with camera information thanks to a custom architecture and ad-hoc finetuning. However, we argue that the necessity of using a specific model for generating videos with acting and camera control is a fundamental limitation: First, because the finetuning procedure can be expensive and limitedly transferable to newer models using different architectures. Second, because using a specialized model for this task, potentially combined with other models for shots in which acting is not necessary, may introduce stylistic inconsistencies in the resulting movie, ultimately harming the quality of the end result. Unlike Uni3C, requiring taskspecific training, ActCam operates entirely at inference time and can be applied to any compatible video backbone without modification. Hence, we introduce ActCam, a zero-shot framework exploiting only existing dense 2D conditioning mechanisms, such as depth and human keypoints [Zhang et al. 2023], to generate videos of acting humans with 3D camera control. We show some compelling results of ActCam in Figure 1. Assuming as input an acting video and the first frame of the target scenario, ActCam is able to generate multiple scenarios of the same acting, in the target scene, with arbitrary camera movements, using a pre-trained model with no finetuning. Our main intuition is that the main limiting factor for this task is the current lack of a camera-aligned conditioning representation for joint motion and camera control. ActCam addresses this by constructing a multi-modal, camera-aligned conditioning signal, including per-frame pose maps that encode acting motion under the desired camera, and per-frame depth maps that provide a coarse scene geometry proxy under the same camera. For controlling human motion, we benefit from 3D reconstruction of the acting video, avoiding inconsistencies related to the 2D keypoints. For camera control, instead, we benefit from depth reprojection, similarly to related approaches [Cao et al. 2025; Ren et al. 2025]. However, one core characteristic of ActCam is how these conditioning signals ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
interact. To avoid interference between static (depth) and dynamic (human motion) elements in the scene, we remove the reference character from the scene geometry and insert the animated character with a novel geometry-aware placement and depth alignment strategy. Finally, we introduce a two-phase conditioning schedule that exploits depth information only in early high-noise steps, while we use only pose information in later steps. This still provides the necessary guidance, removing artifacts due to the over-constraining of the diffusion model with a coarse depth. With both static cameras and dynamic cameras, ActCam outperforms the state-of-the-art under multiple metrics on visual quality and control fidelity. Our contributions are: • Zero-shot joint control. We introduce ActCam, a trainingfree method for joint acting-motion and camera-trajectory control in image-conditioned video generation. • Condition construction for joint control. We design a novel geometry-grounded conditioning pipeline that aligns motion (pose) and scene geometry (depth) to the target camera while preventing static/dynamic interference via referencecharacter removal, geometry-aware placement, and depth alignment. • Two-phase conditioning pipeline. We propose a two-phase inference pipeline that adapts the conditioning information depending on the denoising step, leading to a flexible conditioning that preserves dynamic aspects of the scene.
2
Related Work
Control signals for video generation. A growing body of work improves controllability of diffusion-based image and video generation by injecting external conditions through adapters or control branches, and by conditioning on dense spatial signals such as edges, depth, or semantic layouts [Guo et al. 2024, 2023; Lin et al. 2024; Zhang et al. 2023]. These approaches show that dense conditioning can strongly steer generative models. Some focus on motion control, either with reference videos [Ling et al. 2024; Pondaven et al. 2025; Xiao et al. 2024; Yatim et al. 2024] or with additional control signals such as trajectories [Mou et al. 2024; Wu et al. 2024; Yin et al. 2023; Zhang et al. 2025b; Zhou et al. 2025a] or sparse optical flow [Geng et al. 2025]. Some impose motion control in a zero-shot manner, although they do not support precise camera control [Burgert et al. 2025]. Pardo et al. [2025] manipulate noise at inference time to generate videos with similar motion. However, none of those methods solve the joint requirement of controlling an articulated performance while enforcing a non-trivial camera trajectory. Human-oriented video generation. Recent reference-based animation methods condition video diffusion on human-centric signals (e.g., keypoints, pose, dense pose) to reproduce an intended performance while preserving identity [Cheng et al. 2025; Hu et al. 2024; Jiang et al. 2025; Tan et al. 2024; Wang et al. 2025c, 2024; Xu et al. 2025b; Zhang et al. 2025a, 2024]. Wang et al. [2025a] uses keypoints as intermediate condition to render more realistic videos starting from text. Xu et al. [2024] exploits dense human part masks for guiding generation. All these approaches may suffer from occlusions and perspective-induced errors. Recent efforts [Gan et al. 2025] have also focused on increasing the length of synthesized human videos.
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
Reference Image
Acting Video
Inpaint character
•
3
Flow Matching process
2D Pose
Depth
4D Motion
Depth
Pose timesteps
VACE
Depth timesteps
Input noise
Text Prompt
Output
Reference Image
Camera Rotate Right 45°
Unproject
Unproject
Scene Transfer
Character 3D Fitting
Rendering
Control Signals
Fig. 2. ActCam pipeline. Given a reference image, an acting video, and a target camera trajectory, we (1) estimate background depth from an inpainted reference, (2) recover motion and align it to the background scene via fitting, and (3) rasterize pose and depth+pose control signals under the target viewpoint. A two-phase denoising schedule conditions early steps on depth+pose for stronger camera control, then refines with pose-only to encourage motion adherence.
Differently, some use reconstructed 3D humans to condition video generation [Liang et al. 2025; Zhu et al. 2024], although they offer no camera control. Similarly, Kulal et al. [2023] insert humans into scenes by inferring affordance-aware poses, but without motion or camera control. Our proposal, instead, offers unified camera control and human motion conditioning by exploiting 3D humans. Camera control and joint camera–motion control. There is an interest in controlling camera trajectories in generated videos. A first line of work exploits Plucker embeddings and additional trained branches to enforce camera control [Bahmani et al. 2025a,b; He et al. 2025; Kuang et al. 2024; Xu et al. 2025a]. Alternatively, camera trajectories can be enforced by geometry-aware conditions that convey viewpoint changes more directly [Hou and Chen 2024; Meng et al. 2025; Ren et al. 2025]. These systems do not allow for joint human and camera conditioning simultaneously. Concurrently, Pulp Motion [Courant et al. 2026] generates camera trajectories and human motions, but focuses on trajectory generation rather than video synthesis. The closest work to ours is Uni3C [Cao et al. 2025], that allows joint camera and human motion control. However, it relies on additional training/finetuning to unify camera and motion control. In contrast, ActCam outperforms Uni3C exploiting joint control by constructing camera-aligned pose and depth conditions from a reference image, an acting video, and a camera preset.
3 Method 3.1 Problem setup Inputs and goals. We study image-conditioned video generation with joint control over the character’s motion and the camera trajectory, using a pretrained conditioned video diffusion backbone.
We assume a sequence length 𝑇 and three inputs: a reference image 𝐼 ref ∈ R𝐻 ×𝑊 ×3 defining the target character identity and the environment appearance, an acting video 𝑉act = {𝐼𝜏act }𝑇𝜏=1 providing the target performance, and a target per-frame camera trajectory C = {(𝐾𝜏 , 𝑅𝜏 , t𝜏 )}𝑇𝜏=1 , with intrinsics 𝐾𝜏 ∈ R3×3 and extrinsics (𝑅𝜏 , t𝜏 ) ∈ 𝑆𝑂 (3) × R3 . Let 𝑉 = {𝐼𝜏 }𝑇𝜏=1 be the generated video. Our goal is to generate 𝑉 such that the identity and appearance match 𝐼 ref , the motion in the output video matches the performance in 𝑉act , and finally, such that each frame’s viewpoint is in accordance with the desired camera parameters C. Formalisation of the setup. We formalize the pretrained VACE [Jiang et al. 2025] model as a mapping 𝑓 : ℭ → V that transforms a set of dense video-aligned signals 𝐶 into a video sequence 𝑉 . VACE is trained on a video dataset combining various video-aligned conditions such as: depth, optical flow, character poses, and activation masks. We restrict the conditioning input to the tuple 𝐶 = (𝐼 ref , 𝑐 pose+depth, 𝑐 pose ). Our objective is to synthesize novel, geometrically consistent dense conditions 𝑐 pose and 𝑐 depth derived jointly from the acting video 𝑉act and the target camera C, such that the generated video 𝑉 = 𝑓 (𝐶) satisfies the desired motion and trajectory constraints. ActCam is a pure inference-time method: we keep the backbone fixed and design conditioning signals that jointly encode motion and camera. Continuous flow formulation. Let 𝑧𝑡 denote the latent video state at continuous time 𝑡 ∈ [0, 1] and let 𝐶 denote the conditioning input. We model the generative dynamics as a flow governed by an ordinary differential equation (ODE). The backbone network 𝑣𝜃 approximates the instantaneous velocity field, defining the temporal ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
4
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
evolution of the latent state as d𝑧𝑡 = 𝑣𝜃 (𝑧𝑡 , 𝑡, 𝐶). (1) d𝑡 The target latent 𝑧 1 is obtained by integrating this flow over time from 𝑧 0 ∼ N (0, 1) and is subsequently decoded into the video 𝑉 . Conditioning Requirements for Camera-Aligned Motion Control. To jointly ensure camera control and pose reproduction, the conditioning must provide spatially grounded, per-frame cues in the target camera view: it must both enforce the intended motion and stabilize the scene layout under viewpoint changes. To enforce camera control, we preferred a dense per frame aligned depth signal making sure the model understands the 3D geometry and camera motion rather than numerical camera control using Plücker rays since dense superiority has been proven in [Ren et al. 2025]. In practice, simply combining off-the-shelf controls is brittle: view-locked motion cues (e.g., 2D keypoints) become ambiguous under camera motion, and conditioning on depth estimated from 𝐼 ref entangles the static reference character with the dynamic character we want to animate, which can cause duplicated characters, freezing, or violations of motion/camera constraints. ActCam addresses this by constructing camera-aligned pose/depth conditions and by using a two-phase conditioning schedule (Sections 3.2–3.3).
3.2
Camera-aligned condition construction
Intuition. Our goal is to provide the diffusion model with a targetview-aligned condition that jointly encodes character motion and camera motion. To this end, we leverage depth as a geometric prior: once rasterized under the target camera trajectory, the depth signal implicitly conveys viewpoint changes as apparent background motion. However, naively using the reference image depth introduces a static character that conflicts with the dynamic pose signal (see Figure 2). We address this by inpainting the character out of the reference image and estimating a background-only depth map Dbg . ActCam then constructs a unified 3D scene from Dbg and the recovered poses, from which we rasterize two control signals under the target camera: a depth+pose condition and a pose-only condition obtained by omitting the depth. Camera motion and 3D anchor. In order to provide a 3D anchor to VACE, we construct a 3D background environment which we argue, once rasterized, will provide enough information to the model to effectively encode the camera motion as an equivalent background motion, and disentangle the character motion from camera motion, reducing conflicts between the two signals. We estimate a depth map 𝐷 ref from 𝐼 ref using a monocular depth estimator (MoGe [Wang et al. 2025b]). This leads to the creation of a 3D mesh that can be rendered from different viewpoints. Mesh-based rendering provides stronger geometric consistency than point clouds, which tend to exhibit sparsity and visual artifacts under camera motion. Nonetheless, solely relying on the reference image depth proves insufficient. Indeed, the presence of a static character in the 3D mesh provides a conflicting signal when rendered with a dynamic pose control signal, as in Figure 2. To tackle this, we propose to inpaint the character out of the reference image 𝐼 ref to extract a background-only depth map Dbg , yielding a background 3D mesh, ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
suppressing depth-static character related issues, such as character duplication in the output video. Also, since VACE relies on dense conditioning, enforcing strict pixel correspondence between the control signals and the generated video is crucial. A key challenge is the alignment of the dynamic actor pose with the background 3D scene geometry. In practice, the reference depth map Dref and the background depth map Dbg are estimated in two independent passes, leading to inconsistencies in 3D space. To resolve this discrepancy, we estimate a geometric alignment allowing the character to be correctly registered within the background 3D mesh. We refer to this stage as the scene transfer, described in the next paragraph, and visible in Figure 2. Scene transfer. Let M ⊂ ⟦1, 𝐻 ⟧ × ⟦1,𝑊 ⟧ denote the binary character segmentation mask in Iref , obtained using [Liu et al. 2024]. Our goal is to find the new position of the reference character when transferred into the 3D background mesh, while preserving pixel reprojection with Iref ; therefore, only the depth of the character points is adjusted, since its scale and image plane coordinates will be imposed. We exploit the set of non-inpainted pixels (𝑢, 𝑣) ∉ M, for which reliable correspondences exist between the two depth maps. ref ∈ R3 and xbg ∈ R3 denote the 3D points reconstructed Let x𝑢,𝑣 𝑢,𝑣 from Dref and Dbg , respectively. To emphasize constraints near the character boundary, we assign each an importance weight: not all environment points contribute equally to the scene transfer, as points closer to the character are more informative for rendering correct scene interactions. We assign higher weight to closer points: ref 𝑤 (𝑢, 𝑣) = exp − dist(x𝑢,𝑣 , M) , (2) where dist(·, M) denotes the Euclidean distance in the image plane to the closest pixel in the mask M. We then compute the weighted centroids of the character relative to each set of points in Dref and Dbg respectively: Í Í bg ref (𝑢,𝑣)∉M 𝑤 (𝑢, 𝑣) x𝑢,𝑣 (𝑢,𝑣)∉M 𝑤 (𝑢, 𝑣) x𝑢,𝑣 pref = Í , pbg = Í . (𝑢,𝑣)∉M 𝑤 (𝑢, 𝑣) (𝑢,𝑣)∉M 𝑤 (𝑢, 𝑣) (3) Assuming that the relative position of the character with respect to these centroids is preserved up to a global depth scaling, we align the character by applying an affine transformation along the depth char , the aligned axis. For any character point with reference depth 𝑧 ref depth in the background coordinate system is given by 𝑧
char char 𝑧 bg = 𝑧 ref − 𝑝 𝑧ref
𝑝 bg + 𝑝 𝑧bg, 𝑝 𝑧ref
(4)
where 𝑝 𝑧ref and 𝑝 𝑧bg denote the 𝑧-coordinates of pref and pbg , respectively. This transformation jointly accounts for scale and translation mismatches, enabling consistent addition of the character points to the background 3D scene, while effectively taking characterenvironment proximities (e.g. contacts) into account via the importance weighting. 4D motion recovery (acting). We recover a motion sequence of 3D humans from 𝑉act using a monocular 3D human motion estimator (GVHMR [Chu et al. 2024]). We denote the recovered articulated state at frame 𝜏 as S𝜏 (e.g., SMPL parameters and root pose). Unlike
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
•
5
2D keypoints, {S𝜏 } reduces depth ambiguity and provides a stable motion signal under viewpoint changes.
VACE [Jiang et al. 2025] as backbone. We set the same resolution (𝐻 ×𝑊 ), number of steps (𝑁 ), and scheduler across methods.
Character 3D fitting. As in [Cao et al. 2025], we align the dynamic poses S with the actual replaced character in the assembled background 3D scene (see Figure 2). To that extent, we adopt the same approach as in [Cao et al. 2025] using a least squares estimation based on rigid transformation at time 𝜏 = 0 between the S0 and the extracted keypoints from 𝐼 ref . This method [Umeyama 1991] provides a rotation matrix 𝑅 ∈ 𝑆𝑂 (3), a translation 𝑡ˆ ∈ R3 and a scale 𝑠 that we will apply to the pose sequence as follows:
Datasets. For moving camera, we design an evaluation inspired by Uni3C [Cao et al. 2025]. We select 4 camera presets with common cinematic motions and evaluate on 100 reference clips from RealisDance-Val [Zhou et al. 2025b] per preset (total 4×100 tests). For each clip, we sample a reference 𝐼 ref , extract an acting-motion signal from the original clip, and generate a new video under one of the 4 camera presets. For baselines, we input the same setup, meaning same 𝐼 ref , extracted motion, camera preset. For static cameras, we use RealisDance-Val [Zhou et al. 2025b] under fixed viewpoint.
Ŝ𝜏 = 𝑠.𝑅S𝜏 + 𝑡ˆ This new sequence Ŝ𝜏 is the one being rendered. Rendering control signals. VACE control signal is a dense, imagelike representation that can be directly processed by standard video encoders. We thus directly rasterize both the pose and the depth+pose control signals as videos. We denote R C the rasterization operator under the camera C. For the pose, we rasterize the animated skeleton Ŝ𝜏 following the standard openpose representation in other works (see [Jiang et al. 2025; Zhang et al. 2025a]) encoding 2D joint locations and limb connectivity. This yields a pose control video 𝑐 pose = R C ( Ŝ𝜏 ) ∈ R𝑇 ×𝐻 ×𝑊 ×3 , rendered on a black background, following the target camera viewpoint (see Figure 2). For the depth+pose control signal, we simply render the background 3D mesh built from Dbg using min-max depth normalized grayscale color values, following the format that VACE uses at train time. To obtain actual depth+pose joint signal, we finally superimpose the previously rendered pose signal with the background depth rasterization to obtain 𝑐 pose+depth = R C ( Ŝ𝜏 , Dbg ) ∈ R𝑇 ×𝐻 ×𝑊 ×3 .
3.3
Two-phase conditioning
Monocular depth estimates are often coarse and locally inaccurate. Conditioning solely on actor poses fails to capture camera motion, as the model interprets such motion as part of the body rather than the camera. On the other hand, conditioning on both depth and poses throughout all denoising steps can over-constrain the generation, leading to static backgrounds and the propagation of depth artifacts into high-frequency details (see Ablation, 5). To address this, we adopt a two-phase conditioning schedule that gradually relaxes depth conditioning. Let 𝑡 denote the diffusion timestep. We define the conditioning signal 𝑐 (𝑡)𝜏 using a cutoff threshold 𝑡 stop : ( R C𝜏 ( Ŝ𝜏 , Dbg ), if 𝑡 ≤ 𝑡 stop, 𝑐 (𝑡)𝜏 = (5) R C𝜏 ( Ŝ𝜏 ), if 𝑡 > 𝑡 stop . The pose signal is derived from 3D motion recovery and targetcamera alignment, and provides a stable motion constraint. We therefore keep pose conditioning throughout the full denoising process.
4 Experiments 4.1 Setup We evaluate ActCam having an image 𝐼 ref , an acting video 𝑉act , and optionally a target per-frame camera trajectory as inputs. We adopt
Baselines. For moving camera, we compare with Uni3C [Cao et al. 2025], the most similar method in literature. Both ActCam and Uni3C are based on the same Wan 2.1 14B backbone, ensuring a fair comparison. We also test ActCam in a static camera setup by comparing against strong motion control methods without explicit camera control: Moore-AnimateAnyone [Hu et al. 2024], HumanVid [Wang et al. 2024], MimicMotion [Zhang et al. 2024], Animate-X [Tan et al. 2024], Hyper-Motion [Xu et al. 2025b], UniAnimate-DiT [Wang et al. 2025c], VACE [Jiang et al. 2025], Wan-Animate [Cheng et al. 2025], and SteadyDancer [Zhang et al. 2025a]. We intentionally do not compare to camera-control-only methods, as our goal is joint control with motion. Metrics. We follow Uni3C [Cao et al. 2025] for the evaluation protocol. First, we use VBench [Huang et al. 2024a] for evaluating the visual quality of the generated videos. We evaluate Subject Consistency (SC), Background Consistency (BC), Appearance Fidelity (AF), Imaging Quality (IQ), Temporal Consistency (TC) and Motion Smoothness (MS), all higher is better. This quantifies qualityoriented generation capabilities. We also calculate the Mean Per Joint Error (MPJPE) between the estimated 3D humans in the acting video and in the generated one, to assess the quality of the joint camera and motion control. We evaluate the Sampson Error (SE) [Sampson 1982] for geometric consistency in presence of moving camera. We also report the 3D Consistency (3D-C) and Object Control (OC) scores from WorldScore [Chen et al. 2025]: 3D-C measures multi-view coherence of the generated scene, while OC quantifies the fidelity of object appearance across frames. In the comparison against methods with no camera control, we evaluate only motion quality with VBench metrics.
4.2
Quantitative evaluation
4.2.1 Joint camera and motion control. We now present our main results. From Table 1, ActCam achieves higher quality/consistency scores than Uni3C and reduces motion/geometry errors under controlled camera trajectories. ActCam outperforms Uni3C in MPJPE (0.2087 vs 0.2121) and SE (0.4546 vs 0.5665), showcasing the superior potential of our conditioning mechanism for the moving camera and character. Interestingly, we also outperform Uni3C in the majority of VBench metrics, reporting significant improvements, especially in Subject Consistency (0.9212) and Imaging Quality (0.7212). In our zero-shot setup, we avoid loss of performance due to finetuning on restricted ad-hoc data for motion control. We propose also a qualitative comparison with Uni3C in Figure 9. ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
6
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
Table 1. Joint camera and character control. We evaluate against Uni3C both on VBench, focusing on generation quality, and assessing control quality (MPJPE) and geometric consistency (SE). We also use WorldScore [Chen et al. 2025] to evaluate the 3D consistency (3D-C) and object control (OC) of the generations. We outperform in all cases Uni3C, the closest baseline in our setup. Model Uni3C [Cao et al. 2025] ActCam (Ours)
VBench Average↑ 0.8370 0.8497
SC↑ 0.9084 0.9212
BC↑ 0.9380 0.9350
AF↑ 0.5688 0.5767
IQ↑ 0.6640 0.7212
TC↑ 0.9607 0.9571
MS↑ 0.9821 0.9872
MPJPE↓ 0.2121 0.2087
SE↓ 0.5665 0.4546
3D-C↑ 0.539 0.6304
OC↑ 0.9878 0.9953
Table 2. Static camera comparison. We evaluate on RealisDance-Val [Zhou et al. 2025b] with a static camera using VBench [Huang et al. 2024a,b]. The improved performance of ActCam compared to alternatives using 2D keypoints as conditions advocates for the superiority of our 3D-based pipeline.
Model Moore-AnimateAnyone [Hu et al. 2024] HumanVid [Wang et al. 2024] MimicMotion [Zhang et al. 2024] Animate-X [Tan et al. 2024] Hyper-Motion [Xu et al. 2025b] UniAnimate-DiT [Wang et al. 2025c] VACE [Jiang et al. 2025] Wan-Animate [Cheng et al. 2025] SteadyDancer [Zhang et al. 2025a] ActCam (Ours)
Average↑ 83.78 84.68 82.27 82.93 84.04 84.29 85.33 84.38 85.15 86.47
User evaluation. As a further comparison on joint motion and camera control, we use a two-alternative forced choice (2AFC) study on 17 users with anonymized video pairs comparing ActCam against Uni3C. We use videos generating the same 4 camera presets as in Table 1, and a small subset of RealisDance-Val clips. Each trial shows two generated videos conditioned on the same inputs, alongside the reference acting video and a textual description of the camera motion. Participants then reply to questions evaluating: (1) Camera adherence: “Which video better follows the specified camera motion (viewpoint changes and stable background/parallax)?”, (2) Motion faithfulness: “Which video better matches the motion in the reference acting video (pose accuracy and smoothness)?”, and (3) Visual quality: “Which video looks more visually realistic and pleasing overall (fewer artifacts and less flicker)?”. We report the results in Figure 3. As visible, we considerably outperform Uni3C in all questions, strongly suggesting that the user preference for
Preference
Ours
Uni3C
Tie
Motion
Video
Fig. 3. User study. We compare with Uni3C on camera adherence (Camera) and motion faithfulness (Motion) with respect to the conditioning input, alongside overall visual quality (Visual). We considerably outperform Uni3C, the closest method to ours. ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
BC↑ 94.90 94.94 93.60 95.11 94.97 95.44 95.03 94.52 95.18 95.83
AF↑ 51.56 55.58 52.09 51.72 52.97 52.18 57.81 54.47 56.80 58.66
IQ↑ 66.34 67.45 59.67 60.91 65.52 65.52 70.61 66.87 68.45 70.83
TC↑ 97.16 97.87 97.46 97.79 98.19 98.78 96.74 98.42 97.99 98.88
MS↑ 98.07 98.52 98.61 98.68 99.01 99.24 98.25 98.96 99.02 99.34
generated videos are aligned with the performance boost reported in Table 1. 4.2.2 Motion control with static camera. We now compare against previous methods for motion control, isolating the quality of our 3D human-based motion conditioning. In Table 2, we show VBench metrics on RealisDance-Val. For fairness with others, we assume a static camera, and render only videos of characters in motion following the reference acting video. ActCam is consistently better than strong baselines in this setting, improving subject/background consistency and temporal/motion metrics while maintaining high imaging quality. For example, ActCam improves TC from 98.78 (UniAnimate-DiT, second best) to 98.88 and AF from 57.81 (VACE, second best) to 58.66, while also improving IQ from 70.61 (VACE, second best) to 70.83. We attribute this result to the structural characteristics of our 3D-based conditioning: while methods based on 2D keypoints are subject to bone deformation and occlusion issues, exploiting a 3D signal regularizes subject proportions across the video.
4.3
1 0.8 0.6 0.4 0.2 0 Camera
SC↑ 94.65 93.69 92.21 93.39 93.58 94.56 93.56 93.06 93.48 95.28
Qualitative evaluation
We now report qualitative results. Besides the reported frames, we strongly suggest to visualize the supplementary video. We include multiple results showcasing the different degrees of control of ActCam. In Figure 10, we illustrate how we can render different cameras for the same scene. Those scenarios prove the flexibility of our method across scenes, characters, and reference motion. In Figure 11, we show how the same motion with the same camera can be rendered on different scenes. In Figure 12, we include a camera variation, showing that even complex motion is preserved across scenarios. Moreover, our method is adaptable to multi-character scenarios if the backbone supports it, as shown in Figure 13.
7
Pose conditioning
0.86 0.85 0.84 0.83 0.82 0.81 0.8 0 0.1 0.2 Normalized Timesteps ND
1.0
Fig. 4. Effect of 𝑁𝐷 on VBench score. The figure shows the average VBench scores as a function of 𝑁𝐷 , where the conditioning switches from pose+depth to pose-only. Early switching under-constrains the generation, while late switching (low 𝑡 ) can propagate depth artifacts into highfrequency details, harming results. We set an optimal 𝑁𝐷 = 0.2.
Without Condition Schedule
Fig. 6. Importance of depth. Providing only pose information (𝑁𝐷 = 0, top) for conditioning creates ambiguities between camera and character motion. Conversely, using depth yields the correct character and camera motions (𝑁𝐷 = 0.2, bottom).
With Condition Schedule
Balance of depth conditioning. We vary the number of initial diffusion steps conditioned on both pose and depth (𝑁𝐷 ) and observe a trade-off. This is reflected in VBench metrics reported in Figure 4. We select a canonical 𝑁𝐷 that balances these effects; unless stated otherwise, we use 𝑁𝐷 = 0.2𝑁 (e.g., 𝑁𝐷 =2 when 𝑁 =10). From our evaluation, introducing 𝑁𝐷 denoising iterations with depth conditioning improves environment/camera stability, improving VBench-based evaluations. However, setting 𝑁𝐷 too high can overconstrain late-stage refinement, as we show in Figure 5. In there we set 𝑁𝐷 = 1, resulting in guidance on depth for all diffusion steps. As visible, in presence of dynamic elements or interacting objects, the rigid depth conditioning prevents motion, limiting the realism of the output scene. On the contrary, not benefiting from depth conditioning (𝑁𝐷 = 0) results in ambiguities between character and camera motion, as we demonstrate in Figure 6. There, providing only pose information results in a moving character on a fixed background, failing to capture adequately the camera motion. Reference character removal. As described in Sec. 3.2, we remove the static reference character from the reconstructed scene depth before inserting the animated character. Without this step, the static geometry of the reference character interferes with the dynamic motion conditioning in the depth map, often leading to duplicated characters. Figure 7 shows a representative example: the static character imprint in the depth signal interferes with the animated actor, yielding duplication and inconsistent compositing. This proves the importance of our design decision based on character inpainting.
Fig. 7. Character removal. Without removal, the reference character is captured in the depth map, yielding duplicate subjects. Input image depth Side-view
No alignment
Ablation studies
Uniform weighting
4.4
Importance weighting (ours)
Fig. 5. Importance of conditioning schedule. Excessive depth guidance (setting 𝑁𝐷 = 1) can overly constrain the scene, producing static backgrounds under camera motion (center, red circle). Instead, 𝑁𝐷 < 1 allows to flexibly move the barbell to follow the human motion (right).
With removal
No removal
Depth Map
•
Depth + Pose conditioning
Average (VBench)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
Fig. 8. Importance of scene transfer. Without scene transfer (No alignment), the condition does not respect 3D coherence. Uniform weighting improves placement but importance weighting (ours) is required to achieve best results. The red arrows (right column) show depth/positions offsets.
Scene transfer. In Section 3.2, we describe how we align the composed character depth with the rendered environment depth to stabilize occlusions and prevent depth-inconsistent decomposition under difficult motions or camera changes. Without alignment, reusing ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
8
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
the first-frame depth map can produce implausible occlusions and unstable layout when the scene contains strong depth variation or large viewpoint change. Figure 8 illustrates that depth alignment improves occlusion ordering and reduces layout/tearing artifacts around the actor during strong viewpoint changes. Our importance weighting produces depth-faithful placements, as shown in Figure 8.
5
Conclusion
We presented ActCam, a zero-shot method for joint camera and motion control in image-conditioned video generation. ActCam constructs camera-aligned conditioning cues from a reference image, an acting video, and a target camera preset by combining target-view pose control with target-view depth-based scene guidance, while preventing static/dynamic interference through reference character removal, geometry-aware placement, and depth alignment. To mitigate depth-induced artifacts, we introduced a two-phase conditioning schedule that uses depth only in early denoising steps to lock global structure and viewpoint changes, then relies on pose-only control for high-frequency refinement. Experiments on both static cameras and moving camera benchmarks assess the capabilities of ActCam to render high-quality videos of humans in motion. We also provided ablations and failure cases to clarify the contribution of each design choice and the remaining limitations.
ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
First frame
9
Outputs
ActCam
Uni3C
ActCam
Uni3C
ActCam
Uni3C
Camera
•
Fig. 9. Comparison with Uni3C. Uni3C yields suboptimal camera control (top, middle) and unrealistic character motion (bottom). In the insets, a visualization of the control signal for both Uni3C and ActCam.
Fig. 10. Different cameras. We first show the conditioning signal and ActCam results (top two rows). In the next three rows, we variate camera movements. As visible, the character appearance and motion remain consistent.
ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
10
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
Fig. 11. Different scenes. We display two outputs of ActCam showing the same motion rendered on two characters in different scenes, using the same camera controls.
Output
Conditioning
Output
Conditioning
Fig. 12. Different scenes and different cameras. To show the flexibility of our approach, we apply the same motion to two characters in different scenes, by also varying the camera control. ActCam still renders the correct motion.
Fig. 13. Multi-character results. ActCam handles multiple characters by applying the scene transfer and motion fitting independently per character.
ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
References Sherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin, Willi Menapace, Andrea Tagliasacchi, David B Lindell, and Sergey Tulyakov. 2025a. Ac3d: Analyzing and improving 3d camera control in video diffusion transformers. In CVPR. Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace, Guocheng Qian, Michael Vasilkovsky, Hsin-Ying Lee, Chaoyang Wang, Jiaxu Zou, Andrea Tagliasacchi, et al. 2025b. Vd3d: Taming large video diffusion transformers for 3d camera control. In ICLR. Ryan Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Li Ma, Yitong Deng, Lingxiao Li, Mohsen Mousavi, et al. 2025. Go-with-the-flow: Motion-controllable video diffusion models using real-time warped noise. In CVPR. Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang, Yanwei Fu, and Xiangyang Xue. 2025. Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation. arXiv (2025). Handi Chen, Hongming Zhang, Tianyu Pang, Chao Du, and Min Lin. 2025. WorldScore: A Unified Evaluation Benchmark for World Generation. In IEEE/CVF International Conference on Computer Vision (ICCV). Gang Cheng, Xin Gao, Li Hu, Siqi Hu, Mingyang Huang, Chaonan Ji, Ju Li, Dechao Meng, Jinwei Qi, Penchong Qiao, Zhen Shen, Yafei Song, Ke Sun, Linrui Tian, Feng Wang, Guangyuan Wang, Qi Wang, Zhongjian Wang, Jiayu Xiao, Sheng Xu, Bang Zhang, Peng Zhang, et al. 2025. Wan-Animate: Unified Character Animation and Replacement with Holistic Replication. arXiv (2025). Honglin Chu et al. 2024. GVHMR: Human Motion Recovery via Gravity-View Coordinates. In SIGGRAPH Asia. Robin Courant, David Loiseaux, Xi Wang, Marc Christie, and Vicky Kalogeiton. 2026. Pulp Motion: Framing-aware multimodal camera and human motion generation. In International Conference on Learning Representations (ICLR). Qijun Gan, Yi Ren, Chen Zhang, Zhenhui Ye, Pan Xie, Xiang Yin, Zehuan Yuan, Bingyue Peng, and Jianke Zhu. 2025. Humandit: Pose-guided diffusion transformer for longform human motion video generation. arXiv (2025). Daniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole, Serena Zhang, Tobias Pfaff, Tatiana Lopez-Guevara, Yusuf Aytar, Michael Rubinstein, Chen Sun, et al. 2025. Motion prompting: Controlling video generation with motion trajectories. In CVPR. Xun Guo, Mingwu Zheng, Liang Hou, Yuan Gao, Yufan Deng, Pengfei Wan, Di Zhang, Yufan Liu, Weiming Hu, Zhengjun Zha, et al. 2024. I2v-adapter: A general image-tovideo adapter for diffusion models. In SIGGRAPH. Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2023. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning. In ICLR. Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. 2025. Cameractrl: Enabling camera control for text-to-video generation. In ICLR. Chen Hou and Zhibo Chen. 2024. Training-free camera control for video generation. arXiv (2024). Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, Liefeng Bo, et al. 2024. Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation. In CVPR. Hsin-Ping Huang, Yang Zhou, Jui-Hsien Wang, Difan Liu, Feng Liu, Ming-Hsuan Yang, and Zhan Xu. 2025. Move-in-2d: 2d-conditioned human motion generation. In CVPR. Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. 2024a. VBench: Comprehensive Benchmark Suite for Video Generative Models. In CVPR. Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. 2024b. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models. arXiv (2024). Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, Yu Liu, et al. 2025. VACE: All-in-One Video Creation and Editing. arXiv (2025). Juhi Marzia. 2025. “Only Gets Better Here”: Elon Musk Reacts to AI-Generated Short Film Harry Potter Set in Vietnam War Going Viral. Sportskeeda. https://www.sportskeeda.com/pop-culture/news-only-gets-better-here-elonmusk-reacts-ai-generated-short-film-harry-potter-set-vietnam-war-goes-viral Accessed: 2026-01-22. Zhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu, Hongsheng Li, Leonidas J Guibas, and Gordon Wetzstein. 2024. Collaborative video diffusion: Consistent multi-video generation with camera control. NeurIPS (2024). Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A. Efros, and Krishna Kumar Singh. 2023. Putting People in Their Place: AffordanceAware Human Insertion into Scenes. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Jingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao, Lei Sun, Yichen Qian, Weihua Chen, and Fan Wang. 2025. Realismotion: Decomposed human motion control and
•
11
video generation in the world space. arXiv (2025). Han Lin, Jaemin Cho, Abhay Zala, and Mohit Bansal. 2024. Ctrl-adapter: An efficient and versatile framework for adapting diverse controls to any diffusion model. arXiv (2024). Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, and Yi Jin. 2024. Motionclone: Training-free motion cloning for controllable video generation. arXiv (2024). Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. 2024. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In ECCV. YOU Meng, Zhiyu Zhu, LIU Hui, and Junhui Hou. 2025. NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer. In ICLR. Chong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang, Ying Shan, and Jian Zhang. 2024. Revideo: Remake a video with motion and content control. NeurIPS (2024). Alejandro Pardo, Fabio Pizzati, Tong Zhang, Alexander Pondaven, Philip Torr, Juan Camilo Perez, and Bernard Ghanem. 2025. MatchDiffusion: Training-free Generation of Match-cuts. In ICCV. Alexander Pondaven, Aliaksandr Siarohin, Sergey Tulyakov, Philip Torr, and Fabio Pizzati. 2025. Video motion transfer with diffusion transformers. In CVPR. Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin NimierDavid, Thomas Müller, Alexander Keller, Sanja Fidler, and Jun Gao. 2025. GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control. In CVPR. Paul D Sampson. 1982. Fitting conic sections to “very scattered” data: An iterative refinement of the Bookstein algorithm. Computer graphics and image processing (1982). Shuai Tan, Biao Gong, Xiang Wang, Shiwei Zhang, Dandan Zheng, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, Ming Yang, et al. 2024. Animate-X: Universal Character Image Animation with Enhanced Motion Representation. In ICLR. Victor Tangermann. 2025. OpenAI Says It’s Making a Full Hollywood Movie Using AI. https://futurism.com/openai-full-hollywood-movie-using-ai Accessed: 2026-01-22. S. Umeyama. 1991. Least-squares estimation of transformation parameters between two point patterns. IEEE T-PAMI (1991). Boyuan Wang, Xiaofeng Wang, Chaojun Ni, Guosheng Zhao, Zhiqin Yang, Zheng Zhu, Muyang Zhang, Yukun Zhou, Xinze Chen, Guan Huang, et al. 2025a. HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation. In CVPR. Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, Jiaolong Yang, et al. 2025b. MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision. In CVPR. Xiang Wang, Shiwei Zhang, Longxiang Tang, Yingya Zhang, Changxin Gao, Yuehuan Wang, Nong Sang, et al. 2025c. UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer. arXiv (2025). Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Youqing Fang, Yuwei Guo, Wenran Liu, Jing Tan, Kai Chen, Tianfan Xue, Bo Dai, Dahua Lin, et al. 2024. HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation. In NeurIPS. Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. 2024. Draganything: Motion control for anything using entity representation. In ECCV. Zeqi Xiao, Yifan Zhou, Shuai Yang, and Xingang Pan. 2024. Video diffusion models are training-free motion interpreter and controller. NeurIPS (2024). Dejia Xu, Weili Nie, Chao Liu, Sifei Liu, Jan Kautz, Zhangyang Wang, and Arash Vahdat. 2025a. Camco: Camera-controllable 3d-consistent image-to-video generation. In ICLR. Shuolin Xu, Siming Zheng, Ziyi Wang, H. C. Yu, Jinwei Chen, Huaqi Zhang, Bo Li, Peng-Tao Jiang, et al. 2025b. HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions. arXiv (2025). Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2024. Magicanimate: Temporally consistent human image animation using diffusion model. In CVPR. Danah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten, and Tali Dekel. 2024. Spacetime diffusion features for zero-shot text-driven motion transfer. In CVPR. Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. 2023. Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory. arXiv (2023). Jiaming Zhang, Shengming Cao, Rui Li, Xiaotong Zhao, Yutao Cui, Xinglin Hou, Gangshan Wu, Haolan Chen, Yu Xu, Limin Wang, Kai Ma, et al. 2025a. SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation. arXiv (2025). Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. In ICCV. Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, Fangyuan Zou, et al. 2024. MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance. In ICML. Zhenghao Zhang, Junchao Liao, Menghao Li, Zuozhuo Dai, Bingxue Qiu, Siyu Zhu, Long Qin, and Weizhi Wang. 2025b. Tora: Trajectory-oriented diffusion transformer
ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
12
•
O. El Khalifi, T. Rossi, O. Fossey, T. Fouque, U. Mizrahi, P. Torr, I. Laptev, F. Pizzati, and B. Bellot-Gurlet
for video generation. In CVPR. Haitao Zhou, Chuang Wang, Rui Nie, Jinlin Liu, Dongdong Yu, Qian Yu, and Changhu Wang. 2025a. Trackgo: A flexible and efficient method for controllable video generation. In AAAI. Jingkai Zhou, Yifan Wu, Shikai Li, Min Wei, Chao Fan, Weihua Chen, Wei Jiang, Fan Wang, et al. 2025b. RealisDance-DiT: Simple yet Strong Baseline towards
ACM Trans. Graph., Vol. 1, No. 1, Article . Publication date: May 2026.
Controllable Character Animation in the Wild. arXiv (2025). Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. 2024. Champ: Controllable and consistent human image animation with 3d parametric guidance. In ECCV.