ConceptioArchivearXiv CS
arXiv CSopen access

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories

Wonbong Jang 1 Shikun Liu 1 Soubhik Sanyal 1 Juan Camilo Perez 1 Kam Woh Ng 1 Sanskar Agrawal 1 Juan-Manuel Perez-Rua 1 Yiannis Douratsos 1 Tao Xiang 1 Input Image

Generated Frame 0

Generated Frame 64

arXiv:2604.09429v1 [cs.CV] 10 Apr 2026

Arc Left

Input Image

Generated Frame 55

Generated Frame 35

Move Forward

Input Video …

Generated Camera Trajectories

Figure 1. Rays as Pixels: Unifying Video Generation and Camera Pose Estimation. The first two rows show trajectory-controlled video generation from user-defined camera paths, and the last row shows pose estimation recovering a camera trajectory from a raw input video. By representing camera parameters as dense raxels (rays as pixels), our model learns a joint distribution of videos and camera trajectories, enabling both camera pose prediction and novel view synthesis within a single model.

Abstract

coupled Self-Cross Attention mechanism. A single trained model handles three tasks: predicting camera trajectories from video, jointly generating video and camera trajectory from input images, and generating video from input images along a target camera trajectory. Because the model can both predict trajectories from a video and generate views conditioned on its own predictions, we evaluate it through a closed-loop self-consistency test, demonstrating that its forward and inverse predictions agree. Notably, trajectory prediction requires far fewer denoising steps than video generation, even a few denoising step suffices for self-consistency. We report results on pose estimation and camera-controlled video generation.

Recovering camera parameters from images and rendering scenes from novel viewpoints have long been treated as separate tasks in computer vision and graphics. This separation breaks down when image coverage is sparse or poses are ambiguous, since each task needs what the other produces. We propose Rays as Pixels, a Video Diffusion Model (VDM) that learns a joint distribution over videos and camera trajectories. We represent each camera as dense ray pixels (raxels) and denoise them jointly with video frames through De1 Meta AI, London, United Kingdom. Webpage: Link . Correspondence to: Wonbong Jang <[email protected]>.

Preprint. April 13, 2026.

1

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

1. Introduction

both cases, cameras serve as fixed inputs, not generative targets. We instead re-encode the cameras themselves: we represent rays as pixels (raxels), dense maps where every pixel encodes the origin and direction of the corresponding camera ray. Because raxels share the spatial structure of video frames, the same spatio-temporal VAE compresses them, and the same diffusion backbone processes them without modification. We further introduce decoupled self-cross attention to better couple raxel and video latents.

Videos implicitly encode 3D geometry through motion and parallax. In 3D computer vision and neural rendering, this manifests as two complementary problems: recovering camera trajectories from images (the inverse process), and rendering images from those trajectories and the underlying geometry (the forward process). Despite this duality, traditional pipelines treat these problems in isolation. Structure-from-Motion (SfM) acts as a prerequisite, demanding dense overlapping views to estimate the camera parameters that downstream methods such as Neural Radiance Fields (NeRFs) (Mildenhall et al., 2020) or 3D Gaussian Splatting (3DGS) (Kerbl et al., 2023a) consume as ground truth. This separation creates a fragility: when image coverage is sparse or camera motion is ambiguous, the pipeline breaks. This motivates a unified approach: learning the forward and inverse processes jointly within a single generative model.

Our model, Rays as Pixels, performs three tasks with a single set of weights: i) Camera Pose Estimation: Recovering camera trajectories from video, without ground-truth poses or external pose estimators. ii) Joint Video and Pose Generation: Generating video and the corresponding camera trajectory jointly from one or more sparse input images.

While optimization-based methods like NeRF and 3DGS excel at reconstructing observed regions, they struggle when input views are scarce. Because they reconstruct rather than generate, they cannot plausibly synthesize occluded regions, resulting in blurry artifacts or floaters. Recent 3D-aware generative models (Gao et al., 2025; Liu et al., 2026) address this by leveraging data-driven priors to sample plausible completions for unseen regions. In parallel, camera-controlled VDMs (Liang et al., 2024; Bai et al., 2025) generate video along a specified camera trajectory by conditioning on camera parameters alongside dense input frames. However, these approaches share a limitation: they decouple pose estimation from the generation process. They treat camera poses as a solved prerequisite, relying on offthe-shelf estimators, whether classical, such as COLMAP (Schönberger & Frahm, 2016a), or learned, such as DUSt3R (Wang et al., 2024b), to provide conditioning. Consequently, the generative model inherits the fragility of these upstream estimators, which often fail to converge given the very same sparse inputs the generative model is designed to handle.

iii) Pose-Conditioned Video Generation: Generating video that follows a user-specified camera trajectory, from sparse input images. We evaluate Rays as Pixels on pose estimation and cameracontrolled video generation, and validate that the model’s forward and inverse predictions agree through a closed-loop self-consistency test.

2. Related Work Reconstructing 3D from Images. Optimization-based frameworks such as COLMAP (Schönberger & Frahm, 2016b), ORB-SLAM (Mur-Artal & Tardos, 2017), and DROID-SLAM (Teed & Deng, 2021) remain the gold standard for recovering 3D structure and serve as the backbone for creating most real-world 3D datasets. However, these methods rely on dense feature matching, making them brittle under sparse views or limited overlap. Recent feed-forward approaches such as DUSt3R (Wang et al., 2024a), MASt3R (Leroy et al., 2024), VGGT (Wang et al., 2025), Pow3R (Jang et al., 2025) and MapAnything (Keetha et al., 2025) predict 3D pointmaps and camera parameters from sparse views in a single forward pass. While effective at handling sparse inputs, these methods reconstruct rather than generate: they cannot synthesize plausible content for unobserved regions.

To overcome this dependency, we introduce a unified diffusion model that handles both directions of the duality with a single joint distribution. Unlike existing 3D-aware generative models, which treat camera parameters as fixed conditioning inputs, our model learns a joint distribution over video frames and camera trajectories. This enables both forward and inverse inference within a single model.

Novel View Synthesis. NeRF (Mildenhall et al., 2020) and its successors such as 3DGS (Kerbl et al., 2023b) achieve high-quality rendering through per-scene optimization. However, they rely on dense multi-view supervision and lack the generative priors needed to plausibly synthesize occluded regions in sparse-view settings. Feedforward approaches such as LRM (Hong et al., 2023), pixelSplat (Charatan et al., 2024), and NViST (Jang & Agapito,

The primary challenge is that pretrained video diffusion models operate on dense spatial tensors, while camera parameters are typically low-dimensional global vectors. Prior camera-controlled VDMs handle this mismatch by adding camera-conditioning components, either by injecting camera matrices through learned encoders (Wang et al., 2024c) or by adopting Plücker embeddings (He et al., 2024). In 2

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories Training

DiT (w/ Decoupled Self-Cross Attn)

Add Noise

… …

Input Image -> Video + Camera Motion

<latexit sha1_base64="+JF6fWaoHM/QrWL3z8Hj0/VKTp8=">AAADEXicbVLNbtNAEN6Yv9b8tXDkYhFV4oCiBKHSS6UKOMAFBUHaSokVjddjd5Vd72p3TJVYfgROSPAs3BDXPgFvwpF14oqm7Uj2jr7vm52fncRI4ajf/9MJbty8dfvOxmZ49979Bw+3th8dOl1ajiOupbbHCTiUosARCZJ4bCyCSiQeJbM3DX/0Ba0TuvhMc4OxgrwQmeBAHvq0mNJ0q9vv9ZcWXXUGrdNlrQ2n252/k1TzUmFBXIJz40HfUFyBJcEl1uGkdGiAzyDHsXcLUOjiallrHe14JI0ybf1XULREL0ZUoJybq8QrFdCJu8w14HXcuKRsL65EYUrCgq8SZaWMSEdN41EqLHKSc+8At8LXGvETsMDJj2ctS3O3gYVe66Si3PhDSliHTZ4ZqcnV4c5FuGnOGeRenGLmH2fZaKXmifSKurJ5Uld+ynvPo///5g6HpEAUTfx4+Xz7bUxcvX3/IRpaX1YYTs7TermfrDJA+wWeoqO2upaWIrFg51VudWnaOq/jUyBsgBVLYrY4Z8Bafep6Csk37ndlcHkzrjqHL3qD3d7ux5fdg9ft1mywJ+wpe8YG7BU7YO/YkI0YZzn7yr6zH8G34GfwK/i9kgadNuYxW7Pg7B8LBwEX</latexit>

Video Target Latents zv

… <latexit sha1_base64="cdrJ7Ms6X/WsVu70L3mrJ7qMO1c=">AAADEXicbVLNbtNAEN6Yv2L+WjhysYgqcUBRjKrSS6Wq5QCXKgjSVkqiaLweu6vsele7Y6rUyiNwQoJn4Ya48gS8CUfWjiuatiPZO/q+b3Z+dhIjhaN+/08nuHX7zt17a/fDBw8fPX6yvvH0yOnSchxyLbU9ScChFAUOSZDEE2MRVCLxOJkd1PzxZ7RO6OITzQ1OFOSFyAQH8tDHw6mbrnf7vX5j0XUnbp0ua20w3ej8HaealwoL4hKcG8V9Q5MKLAkucRGOS4cG+AxyHHm3AIVuUjW1LqJNj6RRpq3/Cooa9HJEBcq5uUq8UgGduqtcDd7EjUrKdiaVKExJWPBloqyUEemobjxKhUVOcu4d4Fb4WiN+ChY4+fGsZKnvNnCuVzqpKDf+kBJWYZNnRmpyi3DzMlw35wxyL04x84/TNFqpeSK9YlHZPFlUfso7r6L///oOh6RAFHX8qHm+3TZmUr19fxgNrC8rDMcXab3cT1YZoN0Cz9BRW11LS5FYsPMqt7o0bZ038SkQ1sCSJTE7v2DAWn3megrJN+53Jb66Gdedo9e9eLu3/WGru7ffbs0ae85esJcsZm/YHnvHBmzIOMvZF/aNfQ++Bj+Cn8GvpTTotDHP2IoFv/8BkpEA6g==</latexit>

Ns

Input Images + Camera Motion -> Video

<latexit sha1_base64="WpG0ZC7/i+/TOKM3ZX2Q+fqgHuU=">AAADEXicbVLNbtNAEN6YvxL+WjhysYgqcUCRjarSS6UKOMAFBUHaSokVjddjd5Vd72p33Cqx8gickOBZuCGuPAFvwpF14oqm7Uj2jr7vm52fndRI4SiK/nSCGzdv3b6zcbd77/6Dh482tx4fOl1ZjkOupbbHKTiUosQhCZJ4bCyCSiUepdM3DX90itYJXX6mmcFEQVGKXHAgD32aT04nm72oHy0tvOrErdNjrQ0mW52/40zzSmFJXIJzozgylNRgSXCJi+64cmiAT6HAkXdLUOiSelnrItz2SBbm2vqvpHCJXoyoQTk3U6lXKqATd5lrwOu4UUX5XlKL0lSEJV8lyisZkg6bxsNMWOQkZ94BboWvNeQnYIGTH89aluZuA3O91klNhfGHlLAOmyI3UpNbdLcvwk1zziD34gxz/zjLRms1S6VXLGpbpIvaT3nvRfj/39zhkBSIsokfLZ9vv41J6rfvP4QD68vqdsfnab3cT1YZoP0Sz9BRW11LS5FasLO6sLoybZ3X8RkQNsCKJTGdnzNgrT5zfYXkG/e7El/ejKvO4ct+vNvf/bjTO3jdbs0Ge8qesecsZq/YAXvHBmzIOCvYF/aNfQ++Bj+Cn8GvlTTotDFP2JoFv/8BEF0BGQ==</latexit>

nv

Nt

Sparse Target Latents zt

<latexit sha1_base64="49FU7gVWmi3Ps5NFmbg+qrFrzHM=">AAADEXicbVLNbtNAEN6Yv9b8tXDkYhFV4oCiBKHSS6UKOMAFBUHaSokVjddjd5Vd72p3TJVYfgROSPAs3BDXPgFvwpF14oqm7Uj2jr7vm52fncRI4ajf/9MJbty8dfvOxmZ49979Bw+3th8dOl1ajiOupbbHCTiUosARCZJ4bCyCSiQeJbM3DX/0Ba0TuvhMc4OxgrwQmeBAHvq0mLrpVrff6y8tuuoMWqfLWhtOtzt/J6nmpcKCuATnxoO+obgCS4JLrMNJ6dAAn0GOY+8WoNDF1bLWOtrxSBpl2vqvoGiJXoyoQDk3V4lXKqATd5lrwOu4cUnZXlyJwpSEBV8lykoZkY6axqNUWOQk594BboWvNeInYIGTH89aluZuAwu91klFufGHlLAOmzwzUpOrw52LcNOcM8i9OMXMP86y0UrNE+kVdWXzpK78lPeeR///zR0OSYEomvjx8vn225i4evv+QzS0vqwwnJyn9XI/WWWA9gs8RUdtdS0tRWLBzqvc6tK0dV7Hp0DYACuWxGxxzoC1+tT1FJJv3O/K4PJmXHUOX/QGu73djy+7B6/brdlgT9hT9owN2Ct2wN6xIRsxznL2lX1nP4Jvwc/gV/B7JQ06bcxjtmbB2T8IXAEW</latexit>

Source Latents zs

Images +Videos + Ray Latents

Input Images + Videos + Raxel Images

TAE VAE Encoder

Inference

<latexit sha1_base64="jLNOLnHSRJeQCn68itK5cVB3zuY=">AAADEXicbVJLb9NAEN6YVzGvthy5WESVOKAoQVXppVJFOcAFBUHaSkkUjTdjd5V9aXfcKrX8EzghwW/hhrjyC/gnHFknrmjajmTv6Pu+2XnspFYKT93un1Z06/adu/fW7scPHj56/GR9Y/PQm8JxHHAjjTtOwaMUGgckSOKxdQgqlXiUzg5q/ugUnRdGf6a5xbGCXItMcKAAfdKT08l6u9vpLiy57vQap80a6082Wn9HU8MLhZq4BO+Hva6lcQmOBJdYxaPCowU+gxyHwdWg0I/LRa1VshWQaZIZFz5NyQK9HFGC8n6u0qBUQCf+KleDN3HDgrLdcSm0LQg1XybKCpmQSerGk6lwyEnOgwPciVBrwk/AAacwnpUs9d0Wzs1KJyXlNhxSwips88xKQ76Kty7DdXPeIg/iKWbhcRaNlmqeyqCoSpenVRmmvPsy+f+v7/BICoSu44eL59trYsbl2/cfkr4LZcXx6CJtkIfJKgu0p/EMPTXVNbQUqQM3L3NnCtvUeRM/BcIaWLIkZucXDDhnznxHIYXGw670rm7GdefwVae309n5uN3ef9NszRp7xp6zF6zHXrN99o712YBxlrMv7Bv7Hn2NfkQ/o19LadRqYp6yFYt+/wPwMgEN</latexit>

<latexit sha1_base64="rDp/AxUCy3chLbPwXnLt69ytkZs=">AAADEXicbVLNbtNAEN6Yv2L+WjhysYgqcUBRjKrSS6Wq5QCXKgjSVkqiaLweu6vsele7Y6rUyiNwQoJn4Ya48gS8CUfWjiuatiPZO/q+b3Z+dhIjhaN+/08nuHX7zt17a/fDBw8fPX6yvvH0yOnSchxyLbU9ScChFAUOSZDEE2MRVCLxOJkd1PzxZ7RO6OITzQ1OFOSFyAQH8tDHwylN17v9Xr+x6LoTt06XtTaYbnT+jlPNS4UFcQnOjeK+oUkFlgSXuAjHpUMDfAY5jrxbgEI3qZpaF9GmR9Io09Z/BUUNejmiAuXcXCVeqYBO3VWuBm/iRiVlO5NKFKYkLPgyUVbKiHRUNx6lwiInOfcOcCt8rRE/BQuc/HhWstR3GzjXK51UlBt/SAmrsMkzIzW5Rbh5Ga6bcwa5F6eY+cdpGq3UPJFesahsniwqP+WdV9H/f32HQ1Igijp+1Dzfbhszqd6+P4wG1pcVhuOLtF7uJ6sM0G6BZ+iora6lpUgs2HmVW12ats6b+BQIa2DJkpidXzBgrT5zPYXkG/e7El/djOvO0etevN3b/rDV3dtvt2aNPWcv2EsWszdsj71jAzZknOXsC/vGvgdfgx/Bz+DXUhp02phnbMWC3/8AlTwA6w==</latexit>

Input Latent Format

Input Video -> Camera Motion

Figure 2. Overview of the Rays as Pixels Framework. Training (Left): We jointly encode video frames and corresponding raxel images into a shared latent space using a frozen VAE Encoder. Video inputs undergo 4× temporal compression, while image inputs remain temporally uncompressed. The DiT is trained to co-denoise these latents with Decoupled Self-Cross Attention, conditioned on clean source inputs and noisy sparse target signals. Inference (Right): This joint distribution formulation supports multiple inference modes: (1) generating a synchronized video and motion trajectory from input image(s); (2) performing camera-controlled video generation given a motion trajectory; and (3) recovering camera parameters directly from a input video.

2024) amortize 3D reconstruction across scenes by training transformers on large multi-view datasets, regressing explicit 3D representations from sparse inputs in a single forward pass. Because these methods regress rather than sample, they produce a single deterministic output and cannot capture the multiple plausible completions that occluded regions admit.

ation. Foundational works such as Make-A-Video (Singer et al., 2022) and Video LDM (Blattmann et al., 2023) established the paradigm of inflating 2D priors to achieve temporal consistency. More recently, large-scale commercial systems like Sora (Brooks et al., 2024) and Veo3 (Google DeepMind, 2025), alongside open-source models like LTX (HaCohen et al., 2024) and Wan (Wan et al., 2025), have demonstrated that training on massive video data yields high-fidelity, physically plausible results. Beyond pure synthesis, pretrained VDMs have emerged as effective priors with surprisingly broad emergent capabilities (Wiedemer et al., 2025).

3D-aware Diffusion Models. Diffusion models have been adapted for 3D tasks through several distinct strategies. Early approaches lifted 2D priors to 3D: DreamFusion (Poole et al., 2022) uses Score Distillation Sampling to optimize NeRFs from text, while Zero-1-to-3 (Liu et al., 2023a) conditions image diffusion models on relative camera poses. More recent multi-view diffusion models (Shi et al., 2024; Liu et al., 2023b; Long et al., 2024; Wu et al., 2024; Gao et al., 2025; Zhou et al., 2025; Liu et al., 2026) generate consistent novel views from sparse inputs by jointly denoising multiple target views. All of these approaches, however, treat camera poses as a fixed conditioning input rather than a generative variable. Unified frameworks for joint pose estimation and view synthesis remain rare; the closest prior work is Matrix3D (Lu et al., 2025), which jointly models RGB, pose, and depth across sparse views using a multi-modal diffusion transformer, then renders final outputs through a 3DGS optimization step. Rays as Pixels differs in two respects. First, we build on a pretrained video diffusion model and adapt it minimally through our raxel representation, inheriting the temporal priors learned from large-scale video data. Second, we generate video frames directly from the diffusion model, without an intermediate 3D representation or post-processing rendering step.

Repurposing Generative Models for Different Tasks. A growing body of work fine-tunes pretrained generative models for tasks beyond their original objective. Marigold (Ke et al., 2024) adapts Stable Diffusion (Rombach et al., 2021) for monocular depth estimation, and DepthCrafter (Hu et al., 2025) extends this idea to consitent video depth estimation. More broadly, large-scale generative training has been repurposed for interactive world simulation (Bruce et al., 2024). Recovering camera trajectories is harder than these per-pixel tasks: depth has a one-to-one correspondence with image pixels, while camera trajectories are defined relative to a reference frame and require joint reasoning across the full sequence. Rays as Pixels extends the repurposing pattern to camera geometry by adapting a pretrained video diffusion model to recover camera trajectories from video while preserving its generative capability, treating the video prior as a basis for both generation and inference. Camera-Controlled Video Diffusion. Existing cameracontrolled VDMs differ in how they inject camera information. Several methods feed camera matrices through learned adapter modules (He et al., 2024; 2025; Bai et al., 2025),

Video Diffusion Models. Scaling diffusion models to the temporal domain has driven rapid progress in video gener3

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

while others adopt dense Plücker embeddings (Bahmani et al., 2025) or epipolar attention (Xu et al., 2024; Kuang et al., 2024). A separate line of work conditions VDMs on explicit 3D representations such as point clouds or 3D Gaussians produced by external estimators (Yu et al., 2025b;a; Liang et al., 2024). Each of these strategies treats cameras as fixed conditioning inputs, often requiring additional adapter layers or external 3D pre-processing. Rays as Pixels instead encodes cameras as raxels, letting the same architecture process them alongside video frames as a generative target rather than a conditioning input.

and pointmaps proposed by DUSt3R (Wang et al., 2024a) require per-pixel depth that is not available at inference time. Table 1. Comparison of Camera Representations. We define the ray direction d = X/∥X∥, where X denotes the unprojected pixel coordinates obtained from the intrinsic matrix K. R and T denote rotation and translation, D denotes depth, and × denotes the cross product. Raxels are the only representation that matches both the spatial structure and the 3-channel input format of a pretrained video diffusion VAE.

3. Methodology 3.1. Preliminaries and Notations

Representation

Formulation

Dimension

Limitation

Camera Matrices Plücker Embedding Raymap Pointmap

K, R, T [Rd, Rd × T ] [T, Rd] RX · D + T

(4, 4) (H, W, 6) (H, W, 6) (H, W, 3)

spatially unaligned input channel mismatch input channel mismatch depth required

Raxels (Ours)

Rd + T

(H, W, 3)

At training time, the input consists of three sets of frames: a target video V of N frames, Ns source images, and Nt sparse target images, with the total N + Ns + Nt held fixed across training examples. The source images provide clean conditioning, the sparse target images keep the input length fixed regardless of how many source images are provided, and the target video provides dense temporal supervision.

Coordinate Canonicalization. To ensure the model learns generalizable relative poses rather than memorizing absolute global coordinates, we canonicalize the scene to a randomly selected reference frame, indexed s. We then transform the extrinsic matrices of all video, source, and target frames into its coordinate system:

Each frame I ∈ RH×W ×3 comes with an intrinsic matrix K ∈ R3×3 and an extrinsic (camera-to-world) matrix P ∈ R4×4 . We use subscripts s and t to denote source and sparse target frames respectively, so source frames have parameters (Ks , Ps ) and sparse target frames have (Kt , Pt ).

Prel = Ps−1 Pj .

(j)

(1)

This places the reference camera at the origin with identity (s) rotation (Prel = I), decoupling the camera parameters from the arbitrary global coordinate frame of the original dataset.

We encode the visual inputs into a lower-dimensional latent space using a pretrained VAE encoder E. The target video V is encoded into video latents zv ∈ Rnv ×h×w×c , with 4× temporal compression so that N = 4(nv − 1) + 1. The source and sparse target images are encoded by the same VAE without temporal compression, yielding source latents zs of length Ns and sparse target latents zt of length Nt . The source latents zs remain clean throughout training and serve as conditioning, while the sparse target latents zt are noised and supervised together with the video latents zv .

Raxel Image Construction. For every pixel coordinate u = [u, v]⊺ in a frame Ij , we un-project it to the camera coordinate system using the intrinsic matrix Kj , normalize the resulting direction, and then transform it to the canonical (j) world space using the relative extrinsic Prel , composed of (j) (j) rotation Rrel and translation Trel . This yields the worldspace ray direction d ∈ R3 and origin o ∈ R3 : (j)

d = Rrel

3.2. Representing Rays as Pixels

Kj−1 ũ Kj−1 ũ 2

,

(j)

o = Trel ,

(2)

where ũ is the homogeneous pixel coordinate. We define the raxel at coordinate u as the vector sum d + o, a compact 3-channel encoding that combines the ray’s direction and origin into a single RGB-compatible representation. These per-pixel vectors form the dense raxel image Rj ∈ RHr ×Wr ×3 , which we compute on a coarser pixel grid (Hr = H/2, Wr = W/2) to reduce token overhead. We then encode each raxel image using the shared VAE encoder E, producing latents rv , rs , rt that spatially align with the video latents zv , zs , zt . Because raxel latents and video latents share the same spatial grid, we apply the same Rotary Positional Embedding (RoPE) (Su et al., 2021) to both, with an additional learnable modality embedding to distinguish video latents from raxel latents.

To make camera parameters compatible with a pretrained video diffusion model, we represent each camera as a raxel image: a dense per-pixel map where every pixel encodes the ray origin and direction associated with that pixel. Because raxel images share the same 3-channel spatial structure as RGB video frames, we can encode them using the same pretrained VAE encoder E, placing raxel images and their corresponding video frames in a shared latent space. Table 1 compares raxels with alternative camera encodings: lowdimensional representations such as camera matrices lack spatial structure, Plücker embeddings and raymaps (used in Depth Anything 3 (Lin et al., 2025)) are dense but their 6 channels are incompatible with 3-channel pretrained VAEs, 4

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Input Image

Kaleido

Ours

Ground Truth

Figure 3. Qualitative Results on DL3DV. We visualize the NVS performance of our model against the state-of-the-art baseline Kaleido (Liu et al., 2026) on DL3DV. Given a single reference image (first column), our model synthesizes target views (third column) that exhibit superior structural fidelity and lighting consistency, closely matching the ground truth (fourth column). Note that ours predicts the camera parameters by itself.

ing r at t = 1 for camera-conditioned video generation while denoising z. We find that r converges much faster than z during sampling, and a few denoising steps on r achieves comparable results to multi-step inference on selfconsistency test (Table 3).

3.3. Joint Denoising via Flow Matching Given the visual latents z ∈ {zv , zs , zt } and their corresponding ray latents r ∈ {rv , rs , rt }, we model the joint distribution p(z, r). Previous methods typically learn the conditional distributions p(z|r) for view synthesis or p(r|z) for pose estimation; we instead learn the joint density, which lets a single model recover camera trajectories from video, generate video from camera trajectories, and jointly generate both from input images.

Training Objective. We parameterize the velocity field with a neural network vθ (xt , t) that predicts the target velocity ut = x1 − x0 . The velocity vector has both a magnitude and a direction; while the MSE loss penalizes both jointly, we add a cosine similarity term that isolates the direction component:

Flow Matching Formulation. We adopt Flow Matching (Lipman et al., 2023) as our generative framework: it integrates cleanly with our pretrained backbone and provides straight-line sampling paths that require fewer steps than DDPM-style diffusion at inference. Let x = [z, r] denote the concatenation of visual and ray latents. We define a time-dependent probability density path pt (x) that transforms a Gaussian prior p0 (x) = N (0, I) at t = 0 into the data distribution p1 (x) ≈ pdata (z, r) at t = 1. The conditional probability path is a linear interpolation:

  L(θ) = Et,x0 ,x1 ∥vθ (xt , t) − ut ∥2 + λ 1 −

vθ (xt , t)⊤ ut ∥vθ (xt , t)∥ ∥ut ∥ (4)



where λ weights the cosine term (set to 0.5 in our experiments). We find that the cosine term improves selfconsistency in our ablations (Table 3). 3.4. Decoupled Self-Cross Attention

xt = (1 − t)x0 + tx1 ,

t ∈ [0, 1],

x0 ∼ p0 .

(3)

While ray latents r and video latents z inhabit the same VAE latent space, they have different structural characteristics. Visual latents z encode dense, high-frequency texture, while ray latents r encode smooth camera ray information anchored to a canonical reference frame. The two modalities also differ in their temporal profiles: visual content can change rapidly between frames due to camera motion, occlusion, and lighting, while the underlying camera trajectory evolves smoothly relative to the reference frame s. We

This corresponds to a constant conditional velocity field ut (x | x1 ) = x1 − x0 , guiding the flow from noise to data along a straight line. Unified Tokenization. We concatenate visual latents z and ray latents r along the sequence dimension. Because the two modalities occupy distinct token positions, we can apply asymmetric inference schedules: for example, fix5

,

Self Attention (Video to Video)

Cross Attention (Video to Ray)

By decoupling these operations, we encourage the model to learn the marginal priors (via self-attention) and the conditional dependencies (via cross-attention) separately. This is consistent with a related observation from Seeing without Pixels (Xue et al., 2025), which shows that camera information can be incorporated into video models as a lightweight modality on top of the existing visual backbone.

Feed Forward (Video)

Q

Q

Ray Latents

KV

Ray Latents

Video Latents

Ray Latents

Video Latents

Video Latents

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Self Attention (Ray to Ray)

KV

Cross Attention (Ray to Video)

Video Branch

Feed Forward (Ray)

3.5. From Raxels to Camera Parameters

Q

Q Ray Latents

KV

Video Latents

Pose Recovery via Procrustes Alignment. We decode the ray latents rv , rs , rt using the pretrained VAE decoder, obtaining reconstructed raxel images R̂v , R̂s , R̂t . Recall that raxel images are expressed in a canonicalized coordinate system anchored at the source frame Is (Section 3.2). The reference raxel image R̂s represents the bundle of cam(s) era rays at the identity pose (Prel = I), and any other frame k contains the same bundle transformed by the relative pose (k) Prel . We recover this pose by aligning the predicted bundle R̂k to the reference bundle R̂s as a rigid registration problem:

Ray Branch KV

Figure 4. Decoupled Self-Cross Attention. We replace standard self-attention in the video diffusion backbone with a Decoupled Self-Cross Attention block that processes video and ray latents in parallel branches. Within each branch, intra-modal self-attention operates on tokens of the same modality, followed by inter-modal cross-attention where queries from one modality attend to keys and values from the other. This separation encourages stable training and allows the video and ray latents to attend to each other.

find that applying a single global self-attention across the concatenated sequence under-utilizes this asymmetry: the model tends to fit each modality somewhat independently rather than learning rich cross-modal dependencies.

(k)

P̂rel = argmin

To address this, we replace the standard self-attention with a Decoupled Self-Cross Attention. Each transformer block applies attention in two stages: i) Self-Attention (IntraModal): we apply self-attention separately to z and r. This enforces consistency within each modality, ensuring temporal smoothness in video frames and trajectory coherence in ray paths, without interference between modalities. ii) Symmetric Cross-Attention (Inter-Modal): cross-attention layers then exchange information between modalities, with visual tokens attending to ray tokens (z ← r) and ray tokens attending to visual tokens (r ← z). This guides the video generation to follow the predicted camera rays while allowing the ray latents to refine their trajectory based on visual context.

P ∈SE(3) i,j

(i,j)

R̂k

− P R̂s(i,j)

2

,

(6)

where (i, j) ranges over the spatial grid of the raxel image. This admits a closed-form solution via Orthogonal Procrustes (Luo & Hancock, 1999; Brégier, 2021). Focal Length Recovery via Median-of-Ratios. Once (k) P̂rel is recovered, we transform R̂k back into its own camera coordinate system by applying the inverse pose, yielding (i,j) (k) (i,j) per-pixel local ray vectors dlocal = (P̂rel )−1 R̂k . Under a pinhole camera model, each local ray (x, y, z) relates to its centered pixel coordinate (u, v) by u/fx = x/z and v/fy = y/z. To recover the focal lengths robustly against decoding noise, we use the Median-of-Ratios estimator:

One subtlety: without positional encoding, cross-attention treats keys and values as an unordered set. Because raxel latents and video latents are spatially aligned at every position, we need to preserve this alignment in the cross-attention. We apply RoPE to both queries and keys, ensuring that each visual token attends to its spatially corresponding raxel token (and vice versa).

fˆx = median i,j



ui,j · ẑi,j x̂i,j



, fˆy = median i,j



vi,j · ẑi,j ŷi,j

 (7)

where (ui,j , vi,j ) are pixel coordinates centered at the principal point (assumed to be the image center).

4. Experiments

Probabilistic Interpretation. This decomposition has a clean probabilistic motivation. By the chain rule of probability, the joint log-likelihood log p(z, r) factorizes in two equivalent ways:

Architecture Design. We build on the Wan 2.1 14B Textto-Video model (Wan et al., 2025). We choose the T2V checkpoint rather than the Image-to-Video (I2V) variant because I2V is post-trained from T2V with a fixed frame ordering that conflicts with our flexible source-frame setup. We replace the standard self-attention in each transformer block with our decoupled self-cross attention (Section 3.4), and add a dedicated ray branch consisting of independent

log p(z, r) = log p(r) + log p(z|r) ≡ log p(z) + log p(r|z) . | {z } | {z } | {z } | {z } Self-Attn(r) Cross-Attn(z←r)

X

Self-Attn(z) Cross-Attn(r←z)

(5)

6

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Input Image

Arc Left

Input Image

Arc Right

Figure 5. Qualitative Results on DL3DV-140 following the predefined camera trajectory. We visualize pose-conditioned video generation given a single input image from DL3DV-140 (Ling et al., 2024) and a specific camera path. These scenes present distinct challenges: the top example features a cluttered layout with an off-center subject, while the bottom example contains highly reflective metallic surfaces. Top: The model renders the scene from “Arc Left” camera movement while maintaining the geometric consistency of sofa and many objects within a frame. Bottom: The model is capable of synthesizing the scene from “Arc Right” trajectory, notably our model can synthesize view-dependent reflections on the metallic surfaces while following the camera path.

layer normalizations, feed-forward networks, and linear embedding layers for the ray latents. The ray branch is initialized from the corresponding pretrained video layers and adds 6B parameters, for a total of 20B. We fine-tune all parameters during training.

visible degradation. Unlike I2V models, our approach is flexible in the temporal position of the input image, which can serve as the first, middle, or last frame of the generated sequence.

Training Datasets and Strategies. We train on two realworld datasets: RealEstate10K (Re10K) (Zhou et al., 2018) and DL3DV (Ling et al., 2024). The camera parameters for these datasets, obtained from ORB-SLAM and COLMAP respectively, are reconstructed at arbitrary per-scene scales. We align all scenes to a common metric scale using (Hu et al., 2024; Keetha et al., 2025).

4.1. Self-Consistency and Ablation Study Our model learns a joint distribution over video latents z and camera trajectories r, which enables a test that no conditional-only model can pass: cycle self-consistency. A model that learns only p(z | r) or only p(r | z) cannot round-trip through both directions, because its forward and inverse predictions are not constrained to agree. We use cycle consistency both as a direct evaluation of our joint distribution learning and as the metric for our ablation study.

We resize and center-crop both datasets to 480 × 832 while preserving the original aspect ratio. To encourage scene–trajectory disentanglement, we apply time-reversal augmentation: each scene’s trajectory is augmented with its reverse, giving the model two distinct trajectories per scene during training.

Cycle Setup. We sample a ground-truth pair (z, r) and select three source images from z as conditioning. We first predict the camera trajectory by sampling r′ ∼ p(r | z), then re-generate the video conditioned on the predicted trajectory and the three source images, z ′ ∼ p(z | r′ , Is ). A self-consistent model satisfies two properties: (i) r′ is close to the ground-truth trajectory r, and (ii) z ′ is close in distribution to z. We measure (i) with rotation error Rerr and translation error Terr computed against ground truth, and (ii) with FID and FVD on the re-generated set. We evaluate on the DL3DV-140 test set, sampled at 12 FPS directly from the raw videos rather than the standard DL3DV-140 COLMAP split, which uses non-uniform subsampling.

Qualitative Results. Figures 1, 3, 5, and 6 together demonstrate the visual quality and versatility of our model. Figure 1 shows the model generating consistent frames from artwork and our test set following pre-defined camera trajectories, as well as predicting camera trajectories within the same model. Figure 3 compares Rays as Pixels against Kaleido (Liu et al., 2026) on DL3DV-140 test set, where our model produces better results. Figure 5 shows our model following predefined camera trajectories on complex scenes from DL3DV-140 test set: the top row demonstrates consistency in cluttered spatial arrangements, and the bottom row shows view-dependent effects such as specular reflections on metallic surfaces. Figure 6 visualizes the cycle selfconsistency test under different ablation conditions, where replacing raxels with Plücker embeddings causes the most

Ablation Study. We ablate three design choices using cycle self-consistency as the metric (Table 2): the decoupled self-cross attention (DSCA), the cosine similarity loss, and the choice of ray representation. For the ray-representation ablation, we replace raxels with Plücker embeddings: since 7

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Plücker

No DSCA

No Cosine

Ours

Ground Truth

Figure 6. Self-Consistency and Ablation Study. Re-generated frames from the cycle self-consistency test on three scenes from DL3DV. From left to right: with Plücker embeddings replacing raxels, without DSCA, without cosine similarity loss, our full model, and ground truth. Replacing raxels with Plücker embeddings causes the most severe degradation, consistent with the quantitative results in Table 2.

4.2. Camera Pose Estimation

Plücker embeddings have 6 channels and cannot be encoded by the pretrained VAE, we instead embed them directly via an MLP and concatenate the result with the VAE-encoded video latents.

Experimental Setup. We evaluate pose estimation on three benchmarks: 975 quality-filtered clips from Re10K (Zhou et al., 2018), all 299 test clips from DL3DV140 (Ling et al., 2024), and all 211 clips from Tanks and Temples (T&T) (Knapitsch et al., 2017). These benchmarks cover two distinct regimes: Re10K and T&T contain temporally continuous video clips where consecutive frames are close viewpoints, while DL3DV-140 follows the standard COLMAP subsampling convention, producing wide-baseline frames with large pose differences. Our model is trained on relatively continuous video sequences, so DL3DV-140 probes generalization to a sampling regime different from training. We compare against VGGT (Wang et al., 2025), a feed-forward pose estimator trained on widebaseline multi-view data, evaluated on the same frames for fair comparison.

Results. Table 2 shows that the Plücker-embedding variant dramatically underperforms raxels across every metric: FID (Heusel et al., 2017), FVD (Unterthiner et al., 2018; Skorokhodov et al., 2022), rotation error, and translation error. These gaps hold despite the architecture being otherwise identical except for the ray encoding. The raxel representation, by encoding both ray origin and direction in a form compatible with the pretrained VAE, gives the model a substantially better starting point than Plücker’s 6-channel input-level embedding. Removing the cosine similarity loss or DSCA also hurts selfconsistency, though less dramatically across every metric. The metric-scale training data provides a well-conditioned signal for cross-modal learning, making DSCA a useful refinement rather than a strict requirement. Still, the raxel representation and metric-scale training are the primary load-bearing design choices.

We report mean Relative Rotation Accuracy (mRRA@30), computed over all frames in each clip. Note that we do not report Relative Translation Accuracy: both the groundtruth camera parameters and VGGT are scale-ambiguous, so angular translation distance is very sensitive to small changes near the reference frame Is .

Table 2. Self-Consistency and Ablation Study. We evaluate cycle consistency on DL3DV (sampled at 12 FPS from raw videos) with three source images. Each row shows the effect of removing or replacing one design choice. Replacing raxels with Plücker embeddings degrades all four metrics by large margins, while removing DSCA or the cosine similarity loss produces smaller but consistent drops. Lower is better for all metrics.

Table 3. Camera Pose Estimation. Mean Relative Rotation Accuracy (mRRA@30, ×100) across diffusion step counts on three benchmarks, with VGGT (Wang et al., 2025) as a reference. Our model reaches its best accuracy at 2 diffusion steps across all benchmarks. Higher is better.

Method

FID ↓

FVD ↓

Rerr ↓

Terr ↓

Dataset

1 step

2 steps

5 steps

20 steps

VGGT

Ours w/o DSCA w/o Cosine Sim. Loss Plücker Embedding

7.33 8.69 9.48 21.97

68.17 77.08 97.84 333.56

0.020 0.048 0.058 0.241

0.018 0.052 0.094 0.430

Re10K DL3DV-140 T&T

95.39 82.78 92.34

95.91 88.37 93.51

95.08 87.30 93.43

94.78 85.90 93.01

98.07 91.86 97.70

8

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories Table 4. Quantitative Comparison of Camera-Controlled Video Generation Evaluations. We evaluate visual quality (FID ↓, FVD ↓) and trajectory adherence (Rerr ↓, Terr ↓) across three benchmarks, and achieve the best performance on visual qualities (FID, FVD) with temporal coherence on generated videos.

Results. Table 3 shows two main findings. First, across all three benchmarks, our model reaches its best rotation accuracy at 2 diffusion steps, with further steps producing slightly worse results. The pattern is consistent: 1 step underperforms, 2 steps is the peak, and 5 or 20 steps degrade gradually. This is consistent with the observation in Section 3.3 that ray latents converge much faster than video latents under joint flow matching. Second, VGGT (Wang et al., 2025) achieves higher rotation accuracy across all benchmarks. This is expected: VGGT predicts 3D pointmaps, which inherently encode depth and camera poses, so pose parameters can be derived directly from the pointmap output. Our model predicts videos, which do not contain explicit geometric information; camera parameters are learned as a separate modality (raxels) jointly with video generation but without explicit 3D representations.

Method

4.3. Camera-Controlled Video Generation Experimental Setup. We evaluate visual quality FID (Heusel et al., 2017) and temporal coherence FVD (Unterthiner et al., 2018; Skorokhodov et al., 2022) on Re10K, DL3DV-140, and Tanks and Temples, following the protocols of Wonderland (Liang et al., 2024) and Kaleido (Liu et al., 2026), using VGGT (Wang et al., 2025) to validate the trajectory adherence. Since the ground-truth trajectories are scale-ambiguous, we use the variant of our model trained with per-scene normalized camera trajectories for this evaluation. We compare against MotionCtrl (Wang et al., 2023) and VD3D (Bahmani et al., 2025), which condition video diffusion models on camera matrices and Plücker embeddings; ViewCrafter (Yu et al., 2025b) and Wonderland (Liang et al., 2024), which integrate explicit 3D representations; and Kaleido (Liu et al., 2026), an image-based generative novel view synthesis model with camera positional embedding.

Metrics

Dataset

FID ↓

FVD ↓

Rerr ↓

Terr ↓

RealEstate10K MotionCtrl VD3D ViewCrafter Wonderland Kaleido

22.58 21.40 20.89 16.16 18.04

229.34 187.55 203.71 153.48 103.03

0.231 0.053 0.054 0.046 0.049

0.794 0.126 0.152 0.093 0.181

Ours

15.76

98.72

0.056

0.115

DL3DV-140 MotionCtrl VD3D ViewCrafter Wonderland Kaleido

25.58 22.70 20.55 17.74 41.18

248.77 232.97 210.62 169.34 458.60

0.467 0.094 0.092 0.061 0.011

1.114 0.237 0.243 0.130 0.026

Ours

9.73

102.52

0.098

0.192

Tanks and Temples MotionCtrl VD3D ViewCrafter Wonderland Kaleido

30.17 24.33 22.41 19.46 14.84

289.62 244.18 230.56 189.32 245.09

0.834 0.117 0.125 0.094 0.016

1.501 0.292 0.306 0.172 0.086

Ours

13.02

187.03

0.105

0.192

5. Conclusion Camera pose estimation and camera-controlled video generation have traditionally been treated as separate tasks, with pose estimators and video generators trained and evaluated independently. We presented Rays as Pixels, a unified framework that learns the joint distribution of video frames and camera trajectories in a single model. By encoding camera parameters as dense raxel images compatible with a pretrained video VAE, and by introducing decoupled selfcross attention to couple the two modalities, we enable joint denoising of video and ray latents. Our experiments show visual quality on camera-controlled video generation benchmarks, competitive pose estimation accuracy. We also demonstrate strong cycle self-consistency: the model can predict a camera trajectory from a video and re-generate the video from its own predicted trajectory with minimal degradation, a property that single conditional models cannot achieve.

Results. Table 4 shows that our model achieves the best FID and FVD on all three benchmarks without using explicit 3D representations or camera-specific positional embeddings. 4.4. Limitations Several limitations remain. First, our training data consists of static scenes with smooth camera trajectories (Re10K and DL3DV), so the model may not generalize well to rapid camera motion or scenes containing dynamic objects. Second, the 4× temporal compression of the VAE compresses consecutive frames into shared latent positions, which limits the temporal resolution at which we can recover camera poses. Third, our approach currently uses image conditioning only; integrating text-based control alongside camera control is a natural extension that we leave to future work.

More broadly, we view the raxel representation as an instance of a general pattern: a non-visual modality (such as camera, segmentation) re-encoded as an image-compatible tensor that shares a pretrained visual backbone’s latent space. More speculatively, the joint distribution of visual observations and camera motion that our model captures is one of the primitives relevant to embodied perception, where agents must reason about both what they see and how they are moving through the world. 9

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Impact Statement

Create anything in 3d with multi-view diffusion models. Advances in Neural Information Processing Systems (NeurIPS), 2025.

This work models the joint distribution of video sequences and camera trajectories, with applications to cameracontrolled video generation, pose estimation, and view synthesis. Like other advances in generative video models, our method could be misused to create misleading or fabricated content, although it does not introduce risks beyond those already present in existing video generation systems. We support continued work on detection tools, provenance tracking, and safety protocols to address these concerns as generative models become more capable.

Google DeepMind. Veo 3 technical report. Technical report, Google DeepMind, 2025. URL https://storage. googleapis.com/deepmind-media/veo/ Veo-3-Tech-Report.pdf. HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., Panet, P., Weissbuch, S., Kulikov, V., Bitterman, Y., Melumian, Z., and Bibi, O. Ltxvideo: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103, 2024.

References Bahmani, S., Skorokhodov, I., Siarohin, A., Menapace, W., Qian, G., Vasilkovsky, M., Lee, H.-Y., Wang, C., Zou, J., Tagliasacchi, A., et al. Vd3d: Taming large video diffusion transformers for 3d camera control. In Proceedings of the International Conference on Learning Representations (ICLR), 2025.

He, H., Xu, Y., Guo, Y., Wetzstein, G., Dai, B., Li, H., and Yang, C. Cameractrl: Enabling camera control for textto-video generation. arXiv preprint arXiv:2404.02101, 2024. He, H., Yang, C., Lin, S., Xu, Y., Wei, M., Gui, L., Zhao, Q., Wetzstein, G., Jiang, L., and Li, H. Cameractrl ii: Dynamic scene exploration via camera-controlled video diffusion models. arXiv preprint arXiv:2503.10592, 2025.

Bai, J., Xia, M., Fu, X., Wang, X., Mu, L., Cao, J., Liu, Z., Hu, H., Bai, X., Wan, P., et al. Recammaster: Cameracontrolled generative rendering from a single video. In Proceedings of the International Conference on Computer Vision (ICCV), 2025.

Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, volume 30, 2017.

Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: Highresolution video synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.

Hong, Y., Zhang, K., Gu, J., Bi, S., Zhou, Y., Liu, D., Liu, F., Sunkavalli, K., Bui, T., and Tan, H. Lrm: Large reconstruction model for single image to 3d. In Proceedings of the International Conference on Learning Representations (ICLR), 2023.

Brégier, R. Deep regression on manifolds: A 3d rotation case study. In Proceedings of the International Conference on 3D Vision (3DV), 2021. Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024.

Hu, M., Yin, W., Zhang, C., Cai, Z., Long, X., Chen, H., Wang, K., Yu, G., Shen, C., and Shen, S. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):10579–10596, 2024. doi: 10.1109/TPAMI.2024.3444912.

Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al. Genie: Generative interactive environments. In Proceedings of the International Conference on Machine Learning (ICML), 2024.

Hu, W., Gao, X., Li, X., Zhao, S., Cun, X., Zhang, Y., Quan, L., and Shan, Y. Depthcrafter: Generating consistent long depth sequences for open-world videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2005–2015, 2025.

Charatan, D., Li, S., Tagliasacchi, A., and Sitzmann, V. pixelsplat: 3d gaussian splats from epipolar feature pairs for scalable generalizable 3d reconstruction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

Jang, W. and Agapito, L. Nvist: In the wild new view synthesis from a single image with transformers. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

Gao, R., Holynski, A., Henzler, P., Brussee, A., Martin Brualla, R., Srinivasan, P., Barron, J., and Poole, B. Cat3d: 10

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Jang, W., Weinzaepfel, P., Leroy, V., Agapito, L., and Revaud, J. Pow3r: Empowering unconstrained 3d reconstruction with camera and scene priors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025.

Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. Liu, R., Wu, R., Van Hoorick, B., Tokmakov, P., Zakharov, S., and Vondrick, C. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the International Conference on Computer Vision (ICCV), 2023a.

Ke, B., Obukhov, A., Huang, S., Metzger, N., Daudt, R. C., and Schindler, K. Repurposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9492–9502, 2024.

Liu, S., Ng, K. W., Jang, W., Guo, J., Han, J., Liu, H., Douratsos, Y., Pérez, J. C., Zhou, Z., Phung, C., et al. Scaling sequence-to-sequence generative neural rendering. Proceedings of the International Conference on Learning Representations (ICLR), 2026.

Keetha, N., Müller, N., Schönberger, J. L., Porzi, L., Zhang, Y., Fischer, T., Knapitsch, A., Zauss, D., Weber, E., Antunes, N., Luiten, J., Lopez-Antequera, M., Bulò, S. R., Richardt, C., Ramanan, D., Scherer, S., and Kontschieder, P. Mapanything: Universal feed-forward metric 3d reconstruction. arXiv preprint arXiv:2509.13414, 2025. URL https://arxiv.org/abs/2509.13414.

Liu, Y., Lin, C., Zeng, Z., Long, X., Liu, L., Komura, T., and Wang, W. Syncdreamer: Generating multiviewconsistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023b.

Kerbl, B., Kopanas, G., Leimkühler, T., and Drettakis, G. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (TOG), 2023a.

Long, X., Guo, Y.-C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.-H., Habermann, M., Theobalt, C., et al. Wonder3d: Single image to 3d using cross-domain diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9970–9980, 2024.

Kerbl, B., Kopanas, G., Leimkühler, T., and Drettakis, G. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 2023b.

Lu, Y., Zhang, J., Fang, T., Nahmias, J.-D., Tsin, Y., Quan, L., Cao, X., Yao, Y., and Li, S. Matrix3d: Large photogrammetry model all-in-one. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025.

Knapitsch, A., Park, J., Zhou, Q.-Y., and Koltun, V. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (TOG), 2017. Kuang, Z., Cai, S., He, H., Xu, Y., Li, H., Guibas, L., and Wetzstein, G. Collaborative video diffusion: Consistent multi-video generation with camera control. arXiv preprint arXiv:2405.17414, 2024.

Luo, B. and Hancock, E. R. Procrustes alignment with the em algorithm. In CAIP, 1999. Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.

Leroy, V., Cabon, Y., and Revaud, J. Grounding image matching in 3d with mast3r. In European conference on computer vision, pp. 71–91. Springer, 2024. Liang, H., Cao, J., Goel, V., Qian, G., Korolev, S., Terzopoulos, D., Plataniotis, K., Tulyakov, S., and Ren, J. Wonderland: Navigating 3d scenes from a single image. arXiv preprint arXiv:2412.12091, 2024.

Mur-Artal, R. and Tardos, J. D. ORB-SLAM2: An opensource SLAM system for monocular, stereo, and RGB-D cameras. IEEE Transactions on Robotics, 33(5):1255– 1262, 2017.

Lin, H., Chen, S., Liew, J. H., Chen, D. Y., Li, Z., Shi, G., Feng, J., and Kang, B. Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025. URL https:// arxiv.org/abs/2511.10647.

Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models, 2021.

Ling, L., Sheng, Y., Tu, Z., Zhao, W., Xin, C., Wan, K., Yu, L., Guo, Q., Yu, Z., Lu, Y., et al. Dl3dv-10k: A largescale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

Schönberger, J. L. and Frahm, J.-M. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016a. 11

Rays as Pixels: Learning A Joint Distribution of Video and Camera Trajectories

Schönberger, J. L. and Frahm, J.-M. Structure-from-motion revisited. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016b.

Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., and Revaud, J. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024b.

Shi, Y., Wang, P., Ye, J., Long, M., Li, K., and Yang, X. Mvdream: Multi-view diffusion for 3d generation. In Proceedings of the International Conference on Learning Representations (ICLR), 2024.

Wang, Z., Yuan, Z., Wang, X., Chen, T., Xia, M., Luo, P., and Shan, Y. Motionctrl: A unified and flexible motion controller for video generation. arXiv preprint arXiv:2312.03641, 2023.

Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y. Make-a-video: Text-to-video generation without text-video data, 2022. URL https: //arxiv.org/abs/2209.14792.

Wang, Z., Yuan, Z., Wang, X., Li, Y., Chen, T., Xia, M., Luo, P., and Shan, Y. Motionctrl: A unified and flexible motion controller for video generation. In Proceedings of SIGGRAPH, 2024c. Wiedemer, T., Li, Y., Vicol, P., Gu, S. S., Matarese, N., Swersky, K., Kim, B., Jaini, P., and Geirhos, R. Video models are zero-shot learners and reasoners. arXiv preprint arXiv:2509.20328, 2025.

Skorokhodov, I., Tulyakov, S., and Elhoseiny, M. Styleganv: A continuous video generator with the price, image quality and perks of stylegan2. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3626–3636, 2022.

Wu, R., Mildenhall, B., Henzler, P., Park, K., Gao, R., Watson, D., Srinivasan, P. P., Verbin, D., Barron, J. T., Poole, B., et al. Reconfusion: 3d reconstruction with diffusion priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 21551–21561, 2024.

Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 2021. Teed, Z. and Deng, J. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in Neural Information Processing Systems (NeurIPS), 2021.

Xu, D., Nie, W., Liu, C., Liu, S., Kautz, J., Wang, Z., and Vahdat, A. Camco: Camera-controllable 3dconsistent image-to-video generation. arXiv preprint arXiv:2406.02509, 2024.

Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018.

Xue, Z., Grauman, K., Damen, D., Zisserman, A., and Han, T. Seeing without pixels: Perception from camera trajecories. arXiv preprint arXiv:2511.21681, 2025.

Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.-W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Feng, M., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W., Wang, W., Shen, W., Yu, W., Shi, X., Huang, X., Xu, X., Kou, Y., Lv, Y., Li, Y., Liu, Y., Wang, Y., Zhang, Y., Huang, Y., Li, Y., Wu, Y., Liu, Y., Pan, Y., Zheng, Y., Hong, Y., Shi, Y., Feng, Y., Jiang, Z., Han, Z., Wu, Z.-F., and Liu, Z. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025.

Yu, M., Hu, W., Xing, J., and Shan, Y. Trajectorycrafter: Redirecting camera trajectory for monocular videos via diffusion models. arXiv preprint arXiv:2503.05638, 2025a. Yu, W., Xing, J., Yuan, L., Hu, W., Li, X., Huang, Z., Gao, X., Wong, T.-T., Shan, Y., and Tian, Y. Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2025b. Zhou, J., Gao, H., Voleti, V., Vasishta, A., Yao, C.-H., Boss, M., Torr, P., Rupprecht, C., and Jampani, V. Stable virtual camera: Generative view synthesis with diffusion models. Proceedings of the International Conference on Computer Vision (ICCV), 2025.

Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., and Novotny, D. Vggt: Visual geometry grounded transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2025.

Zhou, T., Tucker, R., Flynn, J., Fyffe, G., and Snavely, N. Stereo magnification: learning view synthesis using multiplane images. ACM Transactions on Graphics (TOG), 2018.

Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., and Revaud, J. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2024a. 12

Record · ID 5957 · SHA-256 c312530328b3411f
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.