2026-5-15
Quantitative Video World Model Evaluation for Geometric-Consistency 1
1
2
3
1
Jiaxin Wu , Yihao Pi , Yinling Zhang , Yuheng Li and Xueyan Zou
arXiv:2605.15185v1 [cs.CV] 14 May 2026
1
2
3
Tsinghua University - IEI Lab, UW-Madison, Adobe Research
Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned graders, which can be subjective and weakly diagnostic for geometric failures. We introduce PDI-Bench (Perspective Distortion Index), a quantitative framework for auditing geometric coherence in generated videos. Given a generated clip, we obtain object-centric observations via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3), lift them to 3D world-space coordinates via monocular reconstruction, and compute a set of projective-geometry residuals capturing three failure dimensions: scale–depth alignment, 3D motion consistency, and 3D structural rigidity. To support systematic evaluation, we build PDI-Dataset, covering diverse scenarios designed to stress these geometric constraints. Across state-of-the-art video generators, PDI reveals consistent geometry-specific failure modes that are not captured by common perceptual metrics, and provides a diagnostic signal for progress toward physically grounded video generation and physical world model. Our code and dataset can be found at https://pdi-bench.github.io/. Videos GT
Seedance
CogVideoX-3
Veo 3.1
Wan2.2
Sora
Hunyuan n
0.2480
0.4521
0.5595
0.8255
0.8825
PDI (Geometric Error) ↓ 0.1206
0.2422
Figure 1 ∣ Overview of the PDI-Bench Evaluation. (Top) Qualitative samples from our dataset, featuring real-world Ground Truth (GT) videos and generated sequences from state-of-the-art models. (Bottom) The corresponding PDI-Scores for GT and each model. Lower scores indicate better adherence to 3D physical laws (scale alignment, motion consistency, and structural rigidity).
1. Introduction High-fidelity generative video models like Seedance [6], Veo 3.1 [9], and Sora [17, 19] have reshaped content creation with their unprecedented visual realism. This impressive quality has led many to view them as early "World Models". This terminology implies a critical shift from simple 2D pixel interpolation to a deeper, latent understanding of 3D structures and physical laws.
Correspondence: [email protected]
SCALE-DEPTH ALIGNMENT Scale-Depth Invariant Preserved
3D MOTION CONSISTENCY
3D STRUCTURAL RIGIDITY
Smooth Trajectory
Consistent 3D Pairwise Distances (ldeal Rigid Motion)
h
...
t=2
h3 h2 h1
Z(depth) Hallucinated Scale
trajectory
Spatial Jitter / Unnatural Reversal
t=0
Non-Euclidean Deformation (Jello-Effect)
h
... t=2
h3 h1
t=i
t=1
h2 t=i
Z(depth)
wrong trajectory
t=1
t=0
Figure 2 ∣ The Three key perspectives that PDI-Bench is evaluating for geometric consistency. (Left) Scale-Depth Alignment: Based on the pinhole camera model, we verify if the product of projected height ℎ and depth 𝑍 remains constant. (Middle) 3D Motion Consistency: We audit the smoothness of 3D world trajectories. (Right) 3D Structural Rigidity: By monitoring internal 3D pairwise distances 𝑖 𝑗 between anchors (skeleton q , q ), we detect non-Euclidean deformations. Despite their visual realism, a significant gap remains between visual plausibility and geometric rigor. Current state-of-the-art models frequently struggle with spatial scale and perspective consistency. While sometimes subtle, these artifacts clearly violate fundamental Euclidean properties. Common failure modes include “volume breathing,” where rigid objects unrealistically expand or contract, and “skating,” where an object’s motion is decoupled from the ground plane’s perspective. These geometric flaws stem from a lack of explicit structural constraints during generation, exacerbated by the absence of metrics capable of evaluating 3D properties from 2D videos. Existing metrics, such as Fréchet Video Distance (FVD) [24] or CLIP-based scores [20], rely on pixel distributions or semantic features, rendering them inherently “geometry-blind.” They fail to penalize physical errors, such as a train shrinking disproportionately to its velocity. To progress toward true physical simulation, the field needs a new evaluation standard that assesses 2D video motion using the strict rules of 3D projective geometry. While evaluating physical common sense has emerged as a key research frontier, existing suites remain insufficient for rigorous geometric verification. Current pipelines typically depend on human judgment or rely on Large Multimodal Models (LMMs) as proxy evaluators, which are prone to subjective bias. Automated benchmarks like PhysBench [7] and WorldBench [25] address scalability but primarily focus on high-level categorical phenomena (e.g., gravity, buoyancy), leaving them largely “geometry-agnostic.” Conversely, emergent geometric metrics such as MEt3R [2] and TRAJAN [1] assess implicit consistency via learned representations, but lack the transparency to diagnose specific physical violations. Our work, PDI-Bench, fills this critical gap by operationalizing explicit physical laws—such as scale–depth alignment, 3D motion consistency, and 3D structural rigidity—as hard quantitative constraints. By lifting 2D video
2
Figure 3 ∣ Overview of the PDI-Dataset and Experimental Setup. Our benchmark comprises 183 highquality videos generated from 28 diverse text prompts, evaluated against real-world Ground Truth (GT). We benchmark six state-of-the-art video generators, categorized into Open-source and Closed-source models . The dataset is meticulously curated to cover five critical physical scenarios: (1) Longitudinal Convergence; (2) Dynamic Tracking; (3) Biological Motion; (4) Curved Motion; and (5) Partial Occlusion. dynamics into verifiable world-space residuals, PDI-Bench provides a precise diagnostic signal to expose the underlying geometric hallucinations of generative world models. As shown in Fig. 2, PDI-Bench evaluates physical realism via three orthogonal geometric metrics. Scaledepth alignment enforces the inverse correlation between an object’s projected 2D height and its 3D depth to penalize unnatural scale hallucinations. 3D motion consistency evaluates the object’s 3D centroid motion in world coordinates, penalizing non-physical accelerations and abrupt angular shifts (e.g., spatial jitter or teleports) while fully decoupling camera movements. 3D structural rigidity tracks internal 3D point-pair distances over time to penalize localized, non-Euclidean deformations (e.g., the “jello effect”) and preserve rigid-body integrity. To move beyond subjective visual assessment, PDI-Bench extracts 3D geometric evidence from 2D pixels via a “perception-to-reasoning” Target-Uplift-Anchor workflow. First, Semantic Targeting via SAM 2 [21] isolates the subject to establish precise 2D spatial boundaries and scale priors (ℎ). Next, 3D Geometric Uplifting via MegaSaM [16] reconstructs world-coordinate pointmaps and camera poses directly from the monocular sequence, lifting 2D observations into a unified 3D physical environment while fully decoupling camera ego-motion. Finally, 3D Structural Anchoring via CoTracker3 [13] deploys dense point anchors within the subject mask. By utilizing the 2D trajectories as precise spatial indices into 𝑛 the aforementioned 3D pointmaps, we lift visual cues into structurally meaningful 3D trajectories (q𝑡 ), enabling camera-motion-invariant evaluation of both Motion consistency and structural rigidity. In summary, we claim the following contributions: • We propose PDI-Bench, a quantitative framework that translates 2D pixel dynamics into 3D geometric reasoning to detect physical hallucinations. • We introduce the Perspective Distortion Index, a metric system that utilizes 3D lifting to decouple kinematic errors from geometric alignment. • We present PDI-Dataset, dedicated to geometric consistency, featuring real-world data and 6 SoTA open/closed-source video models. • We demonstrate that PDI-Bench effectively identifies subtle spatial inconsistencies across five stresstest scenarios, offering critical insights for developing the next generation of space-aware generative systems. 3
2. Related Work Video Generation and World Model. Recent breakthroughs in diffusion-based generative models have significantly advanced the state of video synthesis. Early works focused on temporal extensions of 2D diffusion processes [4, 10]. More recently, large-scale models such as Sora [17], and Wan2.2 [27] have demonstrated a remarkable ability to generate high-fidelity, long-duration sequences with complex semantic alignment. The impressive visual quality has led to the conceptualization of these systems as “World Models”, implying an internal representation of physical laws. However, whether these models truly simulate 3D space or simply replicate 2D statistical patterns remains a subject of intense debate. Our work aims to provide a geometric yardstick to quantify this distinction. Video Quality Assessment and Prompt Following. The evaluation of generative videos has traditionally relied on distribution-based metrics such as Fréchet Video Distance (FVD) [24] and Inception Score (IS) [22], which prioritize frame-level textural and aesthetic quality over structural integrity. To assess text-video semantic alignment and prompt following, CLIP-based scores [20] have become the standard. Building on these foundations, comprehensive suites like VBench [12] and T2V-CompBench [23] have recently been proposed to provide multi-dimensional evaluations, encompassing dynamic quality, temporal consistency, and compositional prompt adherence. Physical Consistency and Scene Understanding. The transition from video synthesis to world simulation requires evaluating true physical realism rather than mere perceptual quality. While recent benchmarks assess macroscopic physical laws—such as VIDEOPHY [3] and WorldModelBench [15] using VLM judges, PhysBench [7] querying interaction dynamics, PhyGenBench [18] testing dynamic constraints, and WorldBench [25] fitting physical constants—they primarily probe high-level semantic plausibility. More recently, WorldScore [8] proposes a unified benchmark for world generation, focusing on long-range scene transitions and camera controllability via SLAM-based tracking and multi-modal assessment. Closer to our approach, MEt3R [2] assesses multi-view consistency by warping semantic features through latent 3D pointmaps, and TRAJAN [1] employs a trajectory autoencoder to detect anomalies in point tracks. However, these methods measure implicit consistency via learned representations. PDI-Bench fills a critical gap by enforcing explicit physical laws as quantitative constraints. Instead of evaluating high-level plausibility, our metric audits fundamental scale–depth alignment, 3D motion consistency, and structural rigidity to expose the spatial integrity of world models. Visual Perception and 3D Reconstruction. The feasibility of 3D physical auditing is underpinned by the surge in foundation models for visual perception. Video segmentation has been revolutionized by memory-based propagation models like SAM 2 [21], while temporal correspondence has reached pixellevel precision through dense trackers such as CoTracker3 [13]. Furthermore, robust 3D reconstruction from dynamic monocular sequences has become accessible via semantic-aware SfM frameworks like MegaSaM [16]. These perceptual backbones serve as the foundational sensors in our pipeline, projecting 2D video dynamics into a verifiable 3D world coordinate system.
Geometric Grounding
TRAJAN
Causal Predictive
VIDEOPHY/ PhysBench/ WorldModelBench/ PhyGenBench
Behavioral Logic
VBench/ T2V-CompBench
Visual Fidelity
PDI-Bench WorldScore MEt3R WorldBench
FVD/IS Pixel-level
Object-level
3D Metric-level
4
Step 0: Input Video
Step 1: Semantic Targeting (SAM 2)
prompts Step 2: 3D Geometric Uplifting (MegaSaM)
& camera poses World-space Pointmaps:
Step 3: 3D Structural Anchoring (CoTracker3)
Anchor seeding SIFT → Shi-Tomasi → uniform-grid
2D pixel-space trajectories
Figure 4 ∣ Overview of the PDI-Bench Pipeline. Our Target-Uplift-Anchor pipeline initiates with Semantic Targeting (SAM 2) for object isolation. It then performs Geometric Uplifting via MegaSaM to construct 3D world-space pointmaps P𝑤𝑜𝑟𝑙𝑑 , followed by Structural Anchoring with CoTracker3 to lift 2D pixel𝑛 space trajectories into 3D coordinates q𝑡 . The resulting spatial-temporal data are synthesized into the PDI metric for consistency auditing. See Sec. 3.1 for details.
3. Method The core objective of PDI-Bench is to bridge the gap between 2D pixel dynamics and 3D physical regularities. We operationalize this objective through a multi-stage Target-Uplift-Anchor pipeline, culminating in a three-dimensional geometric audit quantified by our proposed Perspective Distortion Index(PDI). 3.1. Target-Uplift-Anchor: Perceptual Pipeline As shown in Fig. 4, we orchestrate a hierarchical perception-to-reasoning stack through a collaborative Target-Uplift-Anchor workflow to extract the physical evidence required for geometric auditing. Semantic Targeting (SAM 2). We initiate the pipeline by identifying the auditing subject using Florence2 [28] for automated text-to-box prompting. These prompts are then fed into SAM 2 [21] to generate and 𝑇 propagate a temporal sequence of binary masks { 𝑀𝑡 }𝑡=1 . From these masks, we derive the instantaneous pixel height ℎ𝑡 and the 2D spatial boundaries. 3D Geometric Uplifting (MegaSaM). To recover the latent 3D physical environment, we employ 𝑇 MegaSaM [16] to obtain a coherent depth sequence {𝑍𝑡 }𝑡=1 , the estimated focal length 𝑓 , and camera poses. More importantly, MegaSaM projects every pixel into a unified 3D world coordinate system, 𝑇 × 𝐻 ×𝑊 ×3 yielding world-space pointmaps P𝑤𝑜𝑟𝑙𝑑 ∈ R . This critical step lifts 2D observations into pure 3D space, completely decoupling object kinematics from camera ego-motion.
5
3D Structural Anchoring (CoTracker3). With the 3D world-space constructed, we deploy CoTracker3 [13] to monitor the subject’s internal structural integrity. Within the region defined by the initial mask 𝑀1 , we seed anchor queries via a SIFT → Shi-Tomasi → uniform-grid cascade. After filtering by visibility 𝑛 𝑛 and displacement-jump thresholds, we obtain a set of reliable 2D pixel-space trajectories {(𝑢𝑡 , 𝑣𝑡 )}. 𝑛 𝑛 Using CoTracker3’s 2D pixel-space trajectories {(𝑢𝑡 , 𝑣𝑡 )} as spatial indices into the MegaSaM world-space 𝑛 𝑛 𝑛 pointmaps P𝑤𝑜𝑟𝑙𝑑 , we lift each tracked anchor to its 3D coordinate: q𝑡 = P𝑤𝑜𝑟𝑙𝑑 [𝑡, 𝑣𝑡 , 𝑢𝑡 ]. This operation effectively transforms 2D visual tracking into structurally meaningful 3D trajectories for subsequent rigidity auditing. Finally, the extracted multi-dimensional features from this Target-Uplift-Anchor workflow are synthesized into the Perspective Distortion Index (PDI) via three orthogonal weighted residuals: 𝜖𝑠𝑐𝑎𝑙𝑒 , 𝜖𝑡𝑟𝑎 𝑗 , and 𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 . 3.2. The Perspective Distortion Index (PDI) We synthesize the multi-dimensional geometric evidence into the Perspective Distortion Index (PDI), defined as a weighted sum of three orthogonal physical metrics: PDI = 𝑤1 ⋅ RMSE(𝜖𝑠𝑐𝑎𝑙𝑒 ) + 𝑤2 ⋅ RMSE(𝜖𝑡𝑟𝑎 𝑗 ) + 𝑤3 ⋅ 𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦
(1)
3
where ∑𝑖=1 𝑤𝑖 = 1. Each term is designed to be scale-invariant and to capture a distinct failure mode. We apply the Root Mean Square Error (RMSE) to the scale and trajectory residuals to sensitize the index to catastrophic, high-magnitude physical hallucinations. Conversely, 𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 is defined as the temporal mean of a robust dispersion statistic (MAD). We forgo RMSE for this term to avoid a redundant second-order penalty on an already aggregated measure of spatial incoherence, thereby preserving the direct physical interpretability of structural instability. 3.2.1. Scale-Depth Alignment (𝜖𝑠𝑐𝑎𝑙𝑒 ) To establish a physical baseline, we model the generation process via pinhole camera geometry. For an object with physical height 𝐻 and depth 𝑍 , its projected pixel height ℎ satisfies ℎ = 𝑓 ⋅ 𝐻 /𝑍 . Since 𝑓 and 𝐻 are constant for a rigid body, we derive the Scale-Depth Invariant: ℎ𝑡 ⋅ 𝑍𝑡 = 𝑓 ⋅ 𝐻 = Constant,
where ℎ𝑡 ∈ SAM 2, 𝑍𝑡 ∈ MegaSaM.
(2)
Any fluctuation in this product indicates non-physical scaling (e.g., “volume breathing”). We quantify (𝑡 )
this via a log-space residual 𝜖𝑠𝑐𝑎𝑙𝑒 to ensure symmetric penalties for expansions and contractions: (𝑡 )
𝜖𝑠𝑐𝑎𝑙𝑒 = ∣ln(ℎ𝑡 ⋅ 𝑍𝑡 ) − median𝑘∈[1,5] (ln(ℎ𝑘 ⋅ 𝑍 𝑘 ))∣
(3)
where the median of the first five frames establishes a stable baseline against initialization noise. The final RMSE(𝜖𝑠𝑐𝑎𝑙𝑒 ) serves as a robust measure of scaling severity, where ℎ𝑡 and 𝑍𝑡 are the per-frame pixel height and median object depth, respectively. 3.2.2. 3D Motion Consistency (𝜖𝑡𝑟𝑎 𝑗 ) Rather than relying on 2D projective cues, we audit kinematic plausibility directly in the MegaSaM world coordinate system, which fully decouples object motion from camera ego-motion. 6
Centroid Extraction & Kinematics. The per-frame 3D foreground centroid C𝑡 is computed as the coordinate-wise median of all foreground points in the world-space pointmaps P𝑤𝑜𝑟𝑙𝑑 masked by 𝑀𝑡 . After applying temporal median filtering (𝑘 = 3) to suppress depth flickering noise, we compute the frame-rate-normalized 3D velocity and acceleration using the frame interval Δ𝑡 = 1/fps: v𝑡 = (C𝑡+1 − C𝑡 )/Δ𝑡,
a𝑡 = (v𝑡+1 − v𝑡 )/Δ𝑡
(4)
To evaluate adherence to Newtonian inertia, we decompose kinematic anomalies into two orthogonal, parallel components: abnormal acceleration magnitude and unnatural directional shifts. 1. Acceleration Magnitude Penalty (˜ 𝑎𝑡 ). To quantify spatial jitter without bias from absolute scale, we define the relative acceleration ratio 𝑟𝑡 and its soft-saturated counterpart 𝑎˜𝑡 as: ∥a𝑡 ∥
𝑟𝑡 = 𝑣 , ref
𝑎 ˜𝑡 = 2 ⋅ tanh(𝑟𝑡 /5) ∈ [0, 2)
(5)
where 𝑣ref = max(median(∥v∥), 2 ⋅ median(∥a∥), 𝜀) is a robust speed reference to prevent noise amplifi−6 cation near-stall (𝜀 = 10 ). The tanh compression preserves linear scaling for minor deviations while bounding extreme outliers. 2. Directional Continuity Penalty (𝜑𝑡 ). Macroscopic objects cannot undergo instantaneous sharp turns without external forces. We penalize abrupt directional reversals via the cosine dissimilarity of consecutive velocity vectors. To prevent micro-tremors from triggering false penalties when the object is functionally stationary, this metric is activated only when the instantaneous speed exceeds a noise threshold (∥v∥ > 0.1 𝑣ref ): 𝜑𝑡 = {
1 − cos ∠(v𝑡−1 , v𝑡 ), 0,
if ∥v𝑡−1 ∥, ∥v𝑡 ∥ > 0.1 𝑣ref otherwise
(6)
Note that 𝜑𝑡 naturally falls within [0, 2], where 0 indicates identical directions and 2 indicates a complete reversal, aligning perfectly with the scale of the magnitude penalty. Final Residual. The magnitude and directional penalties are dimensionally aligned and combined with equal weights to form the holistic trajectory residual: (𝑡 )
𝜖𝑡𝑟𝑎 𝑗 = 0.5 ⋅ 𝑎 ˜𝑡 + 0.5 ⋅ 𝜑𝑡
(7)
3.2.3. Structural Rigidity (𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 ) To quantify non-physical internal deformations (e.g., the “jello effect” or “volume breathing”), we audit the structural cohesion of the object across time based on 3D pairwise distance consistency. 𝑛
We obtain the 3D world-space coordinates q𝑡 by sampling the MegaSaM pointmaps at the specific pixel locations defined by CoTracker3’s 2D trajectories. To guarantee high-fidelity correspondence and avoid boundary artifacts (such as depth bleeding), we select a set of optimal anchor pairs at the initial frame (𝑡 = 0) through a rigorous triple-filtering strategy: 1) Visibility: Anchors must maintain a CoTracker3 tracking confidence > 0.5. 2) Depth Smoothness: We compute the Sobel gradient magnitude of the pointmap’s Z-channel and discard points in the top-quartile (> 75th percentile) gradient regions 7
to preclude depth discontinuities. 3) Optimal Pair Scoring: Valid points are paired by maximizing a joint heuristic score that balances the signal-to-noise ratio (large spatial separation) and inland 𝑗 𝑗 𝑖 𝑖 𝑖 reliability:𝒮𝑖, 𝑗 = ∥q0 − q0 ∥2 × min( 𝐷𝑚𝑎𝑠𝑘 , 𝐷𝑚𝑎𝑠𝑘 ), where 𝐷𝑚𝑎𝑠𝑘 denotes the pixel distance from anchor 𝑖 to the nearest mask boundary. In classical kinematics, a rigid body requires the 3D Euclidean distance between internal points to remain 𝑗 𝑖 constant, i.e., ∥q𝑡 − q𝑡 ∥2 = Constant. We define the distance ratio 𝑟𝑖 𝑗 (𝑡 ) and the robust per-frame rigidity score as: 𝑗 𝑖 MAD({𝑟𝑖 𝑗 (𝑡 )}) ∥q𝑡 − q𝑡 ∥2 𝑟𝑖 𝑗 (𝑡 ) = (8) , score(𝑡 ) = 𝑗 median({𝑟𝑖 𝑗 (𝑡 )}) + 𝜀 ∥q0𝑖 − q0 ∥2 where MAD denotes the median absolute deviation, providing immunity against monocular global scale drift. Since 𝑡 = 0 inherently yields a zero score, it is excluded to prevent artificial suppression. The final rigidity metric is computed as the average over the active sequence: 𝑇 −1
𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 =
1 ∑ score(𝑡 ) 𝑇 −1
(9)
𝑡 =1
4. Experiments In this section, we conduct a systematic evaluation of state-of-the-art generative video models using the PDI-Bench framework to quantify the discrepancy between visual plausibility and geometric consistency. By auditing diverse scenarios, we expose vulnerabilities in latent 3D spatial representations and provide a diagnostic roadmap for physically grounded world simulators. 4.1. Experimental Setup PDI-Dataset and Evaluation Scenarios. To systematically audit these generative systems, we curate a comprehensive benchmark containing 183 video sequences derived from 28 diverse textual prompts. This dataset comprises 15 high-quality, real-world Ground Truth (GT) videos serving as a reliable physical baseline for calibration, alongside 168 synthetic videos. To provide a holistic view of the field’s current capabilities, we evaluate six representative state-of-the-art models categorized into two tiers: Open-source architectures (Wan 2.2 [27], HunyuanVideo [14]) and Closed-source systems (Sora (OpenAI) [19], Seedance 2.0 Fast (ByteDance via Doubao) [6], CogVideoX-3 (Zhipu AI via ChatGLM) [29], and Veo 3.1-Fast (Google via Flow) [9]). Furthermore, our evaluation is specifically designed to stresstest 3D spatial awareness by covering five critical geometric challenges: (1) Longitudinal Convergence (Longit. Conv.), (2) Dynamic Tracking (Dyn. Track.), (3) Biological Motion (Bio. Motion), (4) Curved Motion (Curved Mot.), and (5) Partial Occlusion (Part. Occl.). For the final PDI-score calculation, we empirically set the weights for scale-depth alignment, motion convergence, and structural rigidity to (𝑤𝑠𝑐𝑎𝑙𝑒 , 𝑤𝑡𝑟𝑎 𝑗 , 𝑤𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 ) = (0.4, 0.4, 0.2), respectively. 4.2. Perceptual and Geometric Fidelity Guard To ensure the integrity of the PDI-Bench pipeline, we implement a multi-stage Fidelity Guard that independently validates the outputs of the underlying perception models before PDI synthesis.
8
Original Video Frame
MegaSaM reconstruction
Reprojected Renderings
Figure 5 ∣ Reconstruction Audit. We validate the fidelity of MegaSaM 3D pointmaps by reprojecting them onto target frames to ensure geometric consistency. Semantic Segmentation Audit (SAM 2): To verify that the segmentation masks generated by SAM 2 consistently adhere to the intended target, we utilize a Vision-Language Model (VLM), specifically Doubao[5]. We generate the RGB-mask overlays and the corresponding binary masks and prompt the VLM to evaluate whether the highlighted region accurately covers the object described by the text_query. Point Tracking Audit (CoTracker3): We audit tracking quality using spatiotemporal trail maps with cyan lines denoting historical trajectories. The VLM judges whether these trails represent the target object’s motion trend. 3D Reconstruction Audit (MegaSaM): The fidelity of the 3D uplifted pointmaps is validated via a Cross-frame Reprojection Consistency check. For a sampled pair of frames { 𝐼 𝐴 , 𝐼 𝐵 }, we re-synthesize the view of frame 𝐵 by projecting the world-space pointmaps of frame 𝐴 onto the image plane of 𝐵 using the estimated camera poses: 𝐴 ˆ𝐼 𝐵 = ℛ(P𝑤𝑜𝑟𝑙𝑑 , 𝑅 𝐵 , T𝐵 , 𝐾 ) (10) where ℛ denotes an off-screen renderer utilizing Z-buffering and point splatting. The reconstruction is deemed successful only if it satisfies predefined thresholds for Coverage (density), MAE (photometric error), and L2 distance. A conceptual illustration of this validation scheme is provided in Fig. 5. 4.3. Quantitative Evaluation via PDI Table 1 presents a comprehensive ranking of state-of-the-art video generators based on our PDI Score. The results reveal a significant “physics gap” between perceived visual realism and underlying geometric consistency. Validation of Physical Baseline. Real-world Ground Truth (GT) videos anchor the benchmark with a PDI Score of 0.1206 and a remarkably low scale residual (𝜖𝑠 = 0.0660). This confirms that our 3D-uplifting pipeline accurately captures the near-perfect perspective laws of the physical world, setting a rigorous lower bound for generative auditing. Leading Generative Stability. Among generative models, Seedance 2.0 and CogVideoX-3 exhibit the highest fidelity to physical laws. Seedance achieves the best overall stability with 0.0% outliers and a leading MathPass rate of 89.3%, suggesting superior geometric self-consistency. CogVideoX-3 demonstrates near-GT performance in Motion Consistency (𝜖𝑡 = 0.2033) and Structural Rigidity (𝜖𝑟 = 0.2065), indicating highly smooth and non-deformable object synthesis.
9
Table 1 ∣ Quantitative Comparison of Physical Consistency on PDI-Bench. We report the PDI Score (mean of residuals) and its breakdown. Lower values indicate higher physical realism. GT represents real-world videos as the benchmark baseline. Best generative results are bolded. Rank
Model
PDI Score ↓
CI95
1 2 3 4 5 6 7
Scale 𝜖𝑠 ↓
Traj 𝜖𝑡 ↓
Rigid 𝜖𝑟 ↓
Std
Ground Truth (GT)
0.1206
Seedance 2.0 CogVideoX-3 Veo 3.1 Wan 2.2 Sora HunyuanVideo
0.2422 0.2480 0.4521 0.5595 0.8255 0.8825
Outlier
MathPass ↑
[0.1018, 0.1386]
0.0660
0.1764
0.1182
0.0378
0.0%
86.7%
[0.1954, 0.2920] [0.1656, 0.4093] [0.2611, 0.7247] [0.2572, 1.0766] [0.2652, 1.4847] [0.3094, 1.6018]
0.2295 0.3135 0.7507 0.9317 1.6753 1.8469
0.2064 0.2033 0.2271 0.2096 0.2711 0.2515
0.3392 0.2065 0.3049 0.5150 0.2345 0.2160
0.1315 0.3065 0.6980 1.2301 1.7312 1.7730
0.0% 3.6% 7.1% 7.1% 14.3% 14.3%
89.3% 85.7% 50.0% 67.9% 70.4% 57.1%
The Scale Hallucination Crisis. A critical finding is the high distortion observed in visually renowned models such as Sora and HunyuanVideo. Despite their aesthetic appeal, they suffer from severe scale hallucinations, with 𝜖𝑠 values exceeding 1.67 (a 25× increase over GT). This highlights a fundamental failure in current transformer-based architectures to maintain the ℎ ⋅ 𝑍 = const perspective invariant. Furthermore, their high standard deviations (> 1.7) and outlier ratios (14.3%) reflect a stochastic instability in physical modeling across different prompts. 4.4. Category-Specific Physical Consistency Analysis Building on Table 2 (a)–(e), we analyze how different motion patterns trigger specific physical hallucinations. Longitudinal Convergence. This scenario audits scaling during axial motion. HunyuanVideo ( 𝑃𝐷𝐼 = 0.10) and CogVideoX-3 ( 𝑃𝐷𝐼 = 0.15) approach the GT baseline, demonstrating a strong grasp of perspective laws. Conversely, Wan2.2 and Veo 3.1 exhibit high scale errors (𝜖𝑠 > 0.32), resulting in a “sliding” effect where object size mismatches relative 3D depth. Biological Motion. Evaluating articulated dynamics reveals the difficulty of maintaining consistency in non-rigid entities. Seedance 2.0 leads this category ( 𝑃𝐷𝐼 = 0.25), effectively preserving structural integrity. In contrast, Veo 3.1 and HunyuanVideo suffer from “volumetric breathing” hallucinations, where body mass fluctuates inconsistently (𝜖𝑠 > 1.97) during gait cycles. Curved Motion. Non-linear trajectories and rotations represent the most significant challenge. Sora exhibits catastrophic failure here ( 𝑃𝐷𝐼 = 2.13), driven by massive scale distortion (𝜖𝑠 = 4.87). This suggests that transformer-based generators often fail to preserve the ℎ ⋅ 𝑍 invariant during rotational transformations. CogVideoX-3 and Seedance 2.0 remain the most robust in this scenario. Partial Occlusion. This tests spatial memory and object permanence. HunyuanVideo shows severe degradation ( 𝑃𝐷𝐼 = 2.41, 𝜖𝑠 = 5.38), indicating the model “forgets” physical dimensions while the object is obscured. Conversely, Sora and Seedance 2.0 demonstrate superior resilience, maintaining structural cohesion (𝜖𝑟 < 0.45) upon the object’s re-emergence. Dynamic Tracking. This scenario assesses the decoupling of camera ego-motion from object kinematics. CogVideoX-3 ( 𝑃𝐷𝐼 = 0.16) and HunyuanVideo ( 𝑃𝐷𝐼 = 0.17) excel, nearing GT fidelity. However, Sora struggles significantly with scale (𝜖𝑠 = 2.84), as its world model tends to conflate camera proximity with non-physical object growth, leading to a collapse of perspective coupling.
10
(a) Longitudinal Convergence
(b) Dynamic Tracking
(c) Biological Motion
Model
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
Model
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
Model
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
GT (ref ) HunyuanVideo CogVideoX-3 Sora Seedance 2.0 Veo 3.1 Wan2.2
0.0715 0.1004 0.1540 0.2496 0.2623 0.2958 0.3305
0.0198 0.0580 0.0814 0.2656 0.2084 0.4344 0.3278
0.1406 0.1533 0.1909 0.2634 0.2351 0.2115 0.2273
0.0369 0.0795 0.2252 0.1903 0.4245 0.1869 0.5424
GT (ref ) CogVideoX-3 HunyuanVideo Veo 3.1 Seedance 2.0 Wan2.2 Sora
0.1167 0.1620 0.1722 0.2170 0.2334 0.2941 1.2824
0.0623 0.1039 0.1839 0.3233 0.2214 0.1661 2.8387
0.1734 0.2222 0.2097 0.1677 0.2039 0.2141 0.2719
0.1123 0.1577 0.0738 0.1029 0.3162 0.7098 0.1910
GT (ref ) Seedance 2.0 CogVideoX-3 Wan2.2 Sora HunyuanVideo Veo 3.1
0.1319 0.2536 0.2773 0.2827 0.3924 0.9760 1.0230
0.0887 0.3401 0.4031 0.3632 0.5968 2.1265 1.9738
0.1950 0.1816 0.1781 0.1748 0.2555 0.2310 0.2268
0.0921 0.2246 0.2239 0.3377 0.2571 0.1649 0.7135
(d) Curved Motion
(e) Partial Occlusion
(f ) Human Expert Study
Model
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
Model
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
Rank
GT (ref ) Seedance 2.0 CogVideoX-3 Wan2.2 Veo 3.1 HunyuanVideo Sora
0.1550 0.2558 0.2567 0.5223 0.6037 0.7467 2.1277
0.1136 0.3001 0.2944 0.8862 1.0484 1.4705 4.8660
0.1366 0.2122 0.2174 0.2042 0.2728 0.2321 0.3225
0.2745 0.2542 0.2600 0.4305 0.3759 0.3282 0.2617
GT (ref ) Seedance 2.0 Sora Veo 3.1 CogVideoX-3 Wan2.2 HunyuanVideo
0.1346 0.2101 0.2201 0.2414 0.3964 1.3157 2.4104
0.0626 0.1075 0.1617 0.2270 0.6964 2.8131 5.3793
0.2115 0.1960 0.2482 0.2640 0.2058 0.2208 0.4248
0.1250 0.4433 0.2807 0.2251 0.1774 0.5109 0.4436
1 2 3 4 5 6 7
Model Ground Truth (GT) Seedance 2.0 CogVideoX-3 Veo 3.1 Wan2.2 Sora HunyuanVideo
Mean Score ↓
Std. Dev.
1.57 2.96 2.97 3.36 3.37 3.60 3.66
1.00 1.86 1.47 1.56 2.26 1.99 2.29
Table 2 ∣ Component-wise Physical Consistency and Human Alignment. (a)–(e) Component-wise Error Analysis across Different Scenarios. Each sub-table evaluates models on a specific challenge. GT (ref) denotes the real-world baseline. Lower values indicate higher physical consistency. (f) Human Expert Evaluation Results. Scores range from 1 (Best) to 10 (Worst). The human consensus ranking exhibits a perfect alignment ( 𝜌 = 1.0) with our automated PDI Score. 4.5. Human Expert Study and Perceptual Alignment To validate the alignment between PDI-Bench and human physical intuition, we conducted a perceptual study involving seven computer vision experts. The evaluation set comprised 105 unique video clips: 15 real-world (GT) videos and 90 AI-generated clips (15 per model) sourced from six state-of-the-art generators. For each model, three clips were sampled for each of the five physical scenarios. Adhering to the PDI scoring protocol, experts assigned ratings ranging from 1 (Physical Realism) to 10 (Catastrophic Failure). With each clip reviewed by all seven experts, we obtained 𝑛 = 105 subjective ratings per model. Results and Alignment Analysis. As summarized in Table 2 (f), the expert consensus yields a model ranking identical to our automated results. Real-world GT consistently anchors the high-fidelity bound with the lowest mean score (1.57±1.00), confirming the experts’ ability to distinguish natural physics from synthetic artifacts. Among the AI models, Seedance 2.0 and CogVideoX-3 emerge as the top performers with nearly identical scores (2.96 and 2.97, respectively). Their low mean scores suggest a higher degree of temporal consistency and fewer visible violations of basic physical laws. Veo 3.1 and Wan 2.2 occupy the middle tier (3.36–3.37). While competitive, they exhibit more frequent physical inaccuracies compared to the top-tier models. In contrast, Sora and HunyuanVideo were penalised for the severe “scale hallucinations” and “motion jitter” previously identified by our 3D geometric audit. Notably, the high standard deviations in the lower-ranked models (e.g., 2.29 for HunyuanVideo) suggest that their failure modes are often dramatic and highly visible to human observers. This consistent alignment underscores the validity of PDI-Bench as an objective, automated proxy for human perceptual assessment of physical laws in video generation.
5. Case Study: Diagnosing Autoregressive Extrapolation Beyond standard evaluation, we conducted a stress test on autoregressive (AR) long-video generation. We evaluate the Self-Forcing paradigm [11] built on the Wan2.1-T2V-1.3B architecture [26]. By extrapolating 81-frame trained sequences to 129 frames across 28 prompts (see the Appendix A.2), we analyze the decay of 3D physical consistency beyond the training context window. 11
Table 3 ∣ Self-Forcing AR Analysis. Scale drift occurs despite stable trajectories. Category
PDI ↓
Scale ↓
Traj ↓
Rigid ↓
Longit. Conv. Dynamic Track. Biological Mot. Curved Motion Partial Occl.
1.3063 0.1994 2.0819 0.3111 2.7570
2.8462 0.1515 4.6326 0.3158 6.2172
0.3078 0.2604 0.3489 0.2530 0.4097
0.2237 0.1730 0.4464 0.4177 0.5311
Overall Mean
1.3407
2.8583
0.3170
0.3531
Kinematic Success vs. Geometric Collapse. Table 3 reveals a striking dichotomy in the AR model’s physical capabilities. Across all scenarios, the 3D Kinematic Trajectory error (𝜖𝑡 ) remains remarkably stable (overall mean: 0.3170). This indicates that the Self-Forcing training paradigm, combined with rolling KV caching, successfully mitigates high-frequency spatial jitter and unnatural reversals often seen in long AR generation.
However, this kinematic smoothness masks a catastrophic failure in 3D projective geometry. The overall Scale-Depth Alignment error (𝜖𝑠 ) surges to 2.8583, particularly in Longitudinal Convergence and Biological Motion. This demonstrates severe scale hallucination: as the model extrapolates beyond its 81-frame training horizon, it loses the “spatial memory” of the object’s original 3D volume, causing objects to expand or contract independently of their depth variations. Vulnerability to Spatial Disconnects. The data further reveal that AR generation is highly sensitive to continuous visual context. In Dynamic Tracking, where the subject remains centrally focused, the scale error remains low (0.1515). Conversely, in Partial Occlusion, the model exhibits its worst performance ( 𝑃𝐷𝐼 = 2.7570). When the object is temporarily obscured, the AR mechanism loses its structural anchor within the context window, causing a complete failure in object permanence when the subject re-emerges. In Curved Motion, while trajectory is maintained, the complexity of rotation leads to a higher rigidity residual (𝜖𝑟 = 0.4177), indicating a mild “jello effect” during non-linear displacement. These insights highlight PDI-Bench’s unique capability to provide fine-grained, frame-level diagnostic signals for emerging video generation architectures.
6. Conclusion In this paper, we introduced PDI-Bench, a novel quantitative framework for auditing the perspective and scale consistency of generative video world models. By constructing a collaborative Target-UpliftAnchor workflow, we successfully translated 2D pixel dynamics into verifiable 3D geometric reasoning. Our proposed Perspective Distortion Index (PDI) provides a multi-modal metric to identify subtle hallucinations such as “volume breathing” and “skating” effects. Through extensive benchmarking on our PDI-Dataset, we demonstrated that our framework offers a robust geometric yardstick that complements existing semantic-based metrics. We believe PDI-Bench will serve as a foundational tool for evaluating and improving the physical intelligence of future artificial world simulators. Limitations. While PDI-Bench establishes a rigorous geometric yardstick, it presents three primary limitations. First, its accuracy relies on off-the-shelf perception tools; in degraded videos where 3D uplifting fails, the framework must fall back to 2D proxies, reducing depth-aware precision. Second, our geometric invariants rely on a rigid-body assumption, making the metric less theoretically suited for highly non-rigid or amorphous subjects. Finally, disentangling complex 3D rotation from axial translation using purely monocular cues is fundamentally ill-posed, occasionally introducing measurement noise despite our consensus-based mitigations.
12
References [1] K. Allen, C. Doersch, G. Zhou, M. Suhail, D. Driess, I. Rocco, Y. Rubanova, T. Kipf, M. S. M. Sajjadi, K. Murphy, J. Carreira, and S. van Steenkiste. Direct motion models for assessing generated videos, 2025. URL https://arxiv.org/abs/2505.00209. [2] M. Asim, C. Wewer, T. Wimmer, B. Schiele, and J. E. Lenssen. Met3r: Measuring multi-view consistency in generated images, 2026. URL https://arxiv.org/abs/2501.06336. [3] H. Bansal, Z. Lin, T. Xie, Z. Zong, M. Yarom, Y. Bitton, C. Jiang, Y. Sun, K.-W. Chang, and A. Grover. Videophy: Evaluating physical commonsense for video generation, 2024. URL https://arxiv. org/abs/2406.03520. [4] A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023. URL https://arxiv.org/abs/2311.15127. [5] ByteDance. Doubao: A family of large language models. https://www.volcengine.com/ product/doubao, 2026. Accessed: 2026-05-06. [6] ByteDance. Seedance 2.0 fast: High-efficiency video generation foundation model. https://www. doubao.com/, 2026. Accessed: 2026-04-19. [7] W. Chow, J. Mao, B. Li, D. Seita, V. Guizilini, and Y. Wang. Physbench: Benchmarking and enhancing vision-language models for physical world understanding, 2025. URL https://arxiv.org/abs/ 2501.16411. [8] H. Duan, H.-X. Yu, S. Chen, L. Fei-Fei, and J. Wu. Worldscore: A unified evaluation benchmark for world generation, 2025. URL https://arxiv.org/abs/2504.00983. [9] Google. Flow: Where the next wave of storytelling happens. https://labs.google/fx/tools/ flow, 2026. Accessed: 2026-03-04. [10] J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models, 2020. URL https: //arxiv.org/abs/2006.11239. [11] X. Huang, Z. Li, G. He, M. Zhou, and E. Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion, 2025. URL https://arxiv.org/abs/2506.08009. [12] Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu. Vbench: Comprehensive benchmark suite for video generative models, 2023. URL https://arxiv.org/abs/2311.17982. [13] N. Karaev, I. Makarov, J. Wang, N. Neverova, A. Vedaldi, and C. Rupprecht. Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024. URL https://arxiv.org/abs/ 2410.11831. [14] W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y. Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P. Li, S. Li, W. Wang, W. Yu, X. Deng, Y. Li, Y. Chen, Y. Cui, Y. Peng,
13
Z. Yu, Z. He, Z. Xu, Z. Zhou, Z. Xu, Y. Tao, Q. Lu, S. Liu, D. Zhou, H. Wang, Y. Yang, D. Wang, Y. Liu, J. Jiang, and C. Zhong. Hunyuanvideo: A systematic framework for large video generative models, 2025. URL https://arxiv.org/abs/2412.03603. [15] D. Li, Y. Fang, Y. Chen, S. Yang, S. Cao, J. Wong, M. Luo, X. Wang, H. Yin, J. E. Gonzalez, I. Stoica, S. Han, and Y. Lu. Worldmodelbench: Judging video generation models as world models, 2025. URL https://arxiv.org/abs/2502.20694. [16] Z. Li, R. Tucker, F. Cole, Q. Wang, L. Jin, V. Ye, A. Kanazawa, A. Holynski, and N. Snavely. Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos, 2024. URL https: //arxiv.org/abs/2412.04463. [17] Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, L. He, and L. Sun. Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024. URL https://arxiv.org/abs/2402.17177. [18] F. Meng, J. Liao, X. Tan, W. Shao, Q. Lu, K. Zhang, Y. Cheng, D. Li, Y. Qiao, and P. Luo. Towards world simulator: Crafting physical commonsense-based benchmark for video generation, 2024. URL https://arxiv.org/abs/2410.05363. [19] OpenAI. Sora: Creating video from text. https://openai.com/sora, 2025. Accessed: 202603-20. [20] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020. [21] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer. Sam 2: Segment anything in images and videos, 2024. URL https://arxiv.org/abs/2408.00714. [22] T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen. Improved techniques for training gans, 2016. URL https://arxiv.org/abs/1606.03498. [23] K. Sun, K. Huang, X. Liu, Y. Wu, Z. Xu, Z. Li, and X. Liu. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation, 2025. URL https://arxiv.org/abs/ 2407.14505. [24] T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges, 2019. URL https://arxiv. org/abs/1812.01717. [25] R. Upadhyay, H. Zhang, J. Solomon, A. Agrawal, P. Boreddy, S. S. Narayana, Y. Ba, A. Wong, C. M. de Melo, and A. Kadambi. Worldbench: Disambiguating physics for diagnostic evaluation of world models, 2026. URL https://arxiv.org/abs/2601.21282. [26] T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. Feng, N. Zhang, P. Li, P. Wu, R. Chu, R. Feng, S. Zhang, S. Sun, T. Fang, T. Wang, T. Gui, T. Weng, T. Shen, W. Lin, W. Wang, W. Wang, W. Zhou, W. Wang, W. Shen, W. Yu, X. Shi, X. Huang, X. Xu, Y. Kou, Y. Lv, Y. Li, 14
Y. Liu, Y. Wang, Y. Zhang, Y. Huang, Y. Li, Y. Wu, Y. Liu, Y. Pan, Y. Zheng, Y. Hong, Y. Shi, Y. Feng, Z. Jiang, Z. Han, Z.-F. Wu, and Z. Liu. Wan: Open and advanced large-scale video generative models, 2025. URL https://arxiv.org/abs/2503.20314. [27] Wan-Video. Wan2.2: Wan: Open and advanced large-scale video generative models. https: //github.com/Wan-Video/Wan2.2, 2025. GitHub repository. [28] B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan. Florence-2: Advancing a unified representation for a variety of vision tasks, 2023. URL https://arxiv.org/abs/2311. 06242. [29] Zhipu AI. Cogvideox-3: Text-to-video diffusion models. https://chatglm.cn/video, 2026. Accessed: 2026-4-18.
15
A. Additional Experimental Details A.1. PDI-Dataset Construction The PDI-Dataset consists of 183 video sequences in total, partitioned into real-world and synthetic subsets. Real-world sequences. The real-world portion of PDI-Dataset contains 15 short clips collected from the public video platform Pexels (https://www.pexels.com/). We manually search for casual handheld or gimbal-assisted footage in ordinary environments and select clips that mirror the five core geometric scenarios introduced in the main paper: Longitudinal Convergence, Dynamic Tracking, Biological Motion, Curved Motion and Partial Occlusion. All videos are manually inspected to ensure (i) clear visibility of the auditing subject, (ii) sufficient parallax for 3D reconstruction, and (iii) minimal motion blur or rolling-shutter artifacts. Synthetic sequences. The synthetic portion comprises 168 videos generated from 28 text prompts, each instantiated once by six representative generative video models: Wan 2.2[27], HunyuanVideo[14], CogVideoX-3[29], Seedance 2.0-Fast[6], Sora[17], and Veo 3.1-Fast [9]. Note on Model Variants. To maintain transparency and reproducibility, we clarify the sources of the generative models evaluated in this benchmark. The CogVideoX-3 results reported herein were obtained via the official Zhipu Qingying web platform (accessed via ChatGLM). It is important to note that this represents a high-performance proprietary variant and should not be conflated with the open-weights CogVideoX-5B or other community-driven iterations. Similarly, Seedance 2.0-Fast, Sora (OpenAI) and Veo 3.1-Fast (Google via Flow) were accessed via their respective commercial web interfaces as of April 2026. All synthetic videos presented in our benchmark reflect the baseline commercial performance available to end-users at the time of evaluation. Note that the Sora samples in our dataset were generated using the $20 monthly consumer subscription rather than the enterprise API, representing the baseline commercial performance of the model. The 28 text prompts are grouped into the five stress-test scenarios listed above and correspond to the high-level descriptions used in our PDI-Dataset (e.g., a car on a straight road for Longitudinal Convergence, a bus entering a roundabout or a tractor turning at a field boundary for Curved Motion). For each model, we follow its official inference configuration (e.g., guidance scale, number of diffusion steps, native resolution) and, when the interface allows, fix the random seed to reduce stochastic variance; otherwise, we always use the first clip returned by the model to avoid cherry-picking. Temporal resolution and pre-processing. All clips are normalized to a fixed frame rate of 24 fps and trimmed to a duration of 4–12 seconds, leading to 96–300 frames per video. For models that natively output longer videos, we uniformly crop a central window to avoid initialization artifacts and abrupt endings. Each frame is resized to a resolution of 𝐻 × 𝑊 = 512 × 512 pixels while preserving the aspect ratio via zero-padding if necessary.
16
A.2. Complete PDI-Bench Prompt Gallery To ensure the reproducibility of our benchmark and facilitate future comparisons, we provide the complete list of text prompts used to generate the video sequences in PDI-Bench. As detailed in Table 4, the dataset comprises 28 meticulously engineered prompts distributed across five distinct projective geometric scenarios: Longitudinal Convergence, Dynamic Tracking, Biological Motion, Curved Motion, and Partial Occlusion. Each prompt is designed to isolate specific 3D-to-2D spatial transformations and camera-object interactions. Table 4 ∣ Complete List of Prompts in PDI-Bench. The prompts are carefully designed to evaluate specific geometric transformations and camera motions across six challenging categories. Category
Text Prompt
Longitudinal Convergence 1. A handheld following shot of a red vintage car driving away on a straight desert highway, harsh noon light and heat haze on the horizon, subtle shake and lateral drift. 2. A high-speed train moving toward the viewer on a straight track, low-angle handheld perspective, rails and gravel receding toward a clear vanishing point. 3. A yellow school bus driving away on a straight tree-lined suburban street, the shot tracking from a low position behind, morning light and clean asphalt. 4. A silver metallic sphere rolling away on a long reflective marble floor in a bright gallery, the shot following closely with slight sway. 5. A heavy cargo truck moving away on a straight bridge at night, tail lights glowing, subtle frame shake, city lights in the distance. 6. A large shipping container being pushed away on a straight industrial dock, cranes and water behind, moving viewpoint, overcast industrial light. Dynamic Tracking 1. A handheld following shot of a red sports car driving on a straight multi-lane highway, city skyline and roadside trees in the background receding rapidly with parallax. 2. A smooth following shot of an autonomous suitcase moving through a vast airport terminal, repeated columns and floor patterns rushing past in frame. 3. A close handheld shot following a large chrome sphere rolling along a straight, reflective museum corridor, exhibits and windows flowing past. 4. A following shot from a vehicle alongside, keeping pace with a large truck carrying a blue container on a long bridge, waves and bridge cables creating dynamic background motion. 5. A smooth following shot of a metal logistics crate moving along a straight automated conveyor, complex factory machinery in the background rushing past. 6. A handheld following shot of a large metal ball rolling through a straight modern art gallery, surrounding artworks and viewers receding rapidly with parallax. Biological Motion Continued on next page
17
Table 4 – Continued from previous page Category
Text Prompt 1. A smooth following shot of a large eagle flying at high speed parallel to a cliff, rock face and sea below, clear sky. 2. A following shot from a moving boat of a dolphin swimming and leaping in the waves alongside, spray and sunlight. 3. A handheld shot of a large octopus swimming away in a complex coral reef, tentacles waving, colorful fish and coral, blue water and light shafts. 4. A backward-moving shot following a snake slithering through dense colorful flowers on the ground, petals and stems, soft daylight. 5. A moving shot following a peacock walking and shaking its tail feathers in a palace garden, fountains and trimmed hedges, ornate tiles.
Curved Motion 1. A handheld tracking perspective follows a silver compact SUV navigating a sharp hairpin turn on a winding mountain road. The view orbits slightly to capture the vehicle transitioning from a front-view to a side-view against the pine forest background. 2. A low-angle shot follows a sports car drifting through a 90-degree corner on a professional race track. The car rotates intensely while the moving shot emphasizes the shifting vanishing lines of the curb and tire marks. 3. A cinematic tracking shot follows a city bus driving through a large, ornate stone roundabout. The view maintains a side perspective, showing the bus constantly changing its orientation relative to the central fountain and surrounding city traffic. 4. A ground-level perspective tracking a small delivery robot as it makes a sharp turn at a sidewalk corner. The shot stays close, highlighting the rotation of the robot’s boxy frame against the detailed brickwork. 5. A handheld shot follows a green tractor making a wide turn at the edge of a plowed field. The view moves with the vehicle, capturing the shifting angles of the heavy wheels and mechanical parts against the vast landscape. Partial Occlusion 1. A car driving along a street at night, wheels briefly obscured by a low roadside guardrail for under a second, handheld shot moving alongside, street lamps and storefronts. 2. A train passing behind a row of thin vertical power line poles, the shot tracking its movement from a moving platform, sky and industrial landscape. 3. A bus moving through a city street, briefly partially hidden by a thin traffic sign, the shot following from the sidewalk. 4. A vintage car driving past a row of thin trees, never fully leaving the moving view, autumn leaves and road. 5. A boat sailing behind a thin pier support, remaining partially visible throughout, handheld shot from the dock, sea and sky. Continued on next page 18
Table 4 – Continued from previous page Category
Text Prompt 6. A robot crate moving through a warehouse, passing behind a thin metal rack, the shot following alongside, shelves and boxes, industrial lighting.
Reconstruction-aware weighting. The final PDI score is synthesized as a weighted sum of three orthogonal physical residuals: PDI Score = 𝑤1 ⋅ RMSE(𝜖𝑠𝑐𝑎𝑙𝑒 ) + 𝑤2 ⋅ RMSE(𝜖𝑡𝑟𝑎 𝑗 ) + 𝑤3 ⋅ 𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 ,
(11)
where ∑𝑖 𝑤𝑖 = 1. When the 3D reconstruction from MegaSAM passes our quality check (i.e., satisfying the ground-plane SVD and reprojection constraints), we adopt a uniform prior with (𝑤1 , 𝑤2 , 𝑤3 ) = (0.4, 0.4, 0.2), assigning nearly equal importance to spatial scaling, temporal kinematics, and structural integrity. A.3. Evaluation Protocol For each model and scenario, we compute PDI scores on all valid sequences and report the median and the 95% Bootstrap Confidence Interval (CI) to ensure robustness against generative outliers. To guarantee statistical significance, each video is evaluated multiple times with different random seeds for anchor sampling, and the resulting scores are averaged. To facilitate cross-model comparison, we employ a GT-Anchored Normalization scheme. We first calculate robust statistics (Median and MAD) for each residual dimension across the real-world Ground Truth (GT) subset to define the "physics-perfect" baseline. Each raw residual is then standardized into a robust Z-score and mapped to a [0, 100] score using a scaled half-logistic function. This dual-track reporting—providing both raw physical residuals (PDI-Error) and normalized scores (PDI-Score)—allows for a transparent and interpretable assessment of the physical common sense embedded in state-of-the-art video generators.
B. Computing Resources All experiments and evaluations in this work were conducted on a Linux workstation equipped with NVIDIA RTX 3090 GPUs. The detailed hardware is as follows: • Hardware Configuration: – GPU: NVIDIA GeForce RTX 3090. – Video Memory: 24,576 MiB (24 GB) GDDR6X per GPU. – Compute Capability: 8.6 (Ampere architecture). • Resource Utilization: The PDI-Bench pipeline leverages multi-GPU parallelization to process the benchmark. Each video sequence is typically assigned to a single GPU worker to perform semantic segmentation (SAM2), point tracking (Co-Tracker), and 3D reconstruction (Mega-SAM) sequentially.
19
C. Proofs of Main Theorems In this section, we provide formal derivations for the geometric identities and metric properties that underlie the Perspective Distortion Index (PDI). C.1. Geometric Invariants of Perspective Projection Let a 3D point P = ( 𝑋, 𝑌 , 𝑍 ) be projected onto the image plane at p = (𝑥, 𝑦 ) with focal length 𝑓 . By similar triangles, we have 𝑋 𝑥 = , 𝑍 𝑓
𝑌 𝑦 = . 𝑍 𝑓
(12)
Consider an object of physical height 𝐻 aligned vertically in 3D. Let ( 𝑋, 𝑌1 , 𝑍 ) and ( 𝑋, 𝑌2 , 𝑍 ) denote its bottom and top endpoints. The projected pixel height ℎ is given by ℎ = ∣ 𝑦2 − 𝑦1 ∣. Using the relation above, »» 𝑓𝑌2 𝑓 ∣𝑌2 − 𝑌1 ∣ 𝑓 𝑌1 »» 𝑓𝐻 »= − = . (13) ℎ = »» »» 𝑍 𝑍 »»» 𝑍 𝑍 Similarly, the horizontal coordinate of the object’s centroid satisfies 𝑥 = 𝑓 𝑋 /𝑍 . Under a rigid-body assumption ( 𝐻 and 𝑋 constant), we obtain the parameter-free invariant 𝑓 𝐻 /𝑍 ℎ 𝐻 𝑥 = 𝑓 𝑋 /𝑍 = 𝑋 ,
(14)
establishing that the ratio of projected height to radial displacement is constant for any linear motion. C.2. Square-Inverse Scaling Law Starting from ℎ = 𝑓 𝐻 /𝑍 , treat 𝐻 and 𝑓 as constants and differentiate with respect to 𝑍 : 𝑑ℎ 𝑓𝐻 −2 = − 𝑓 𝐻𝑍 = − 2 . 𝑑𝑍 𝑍
(15)
For two nearby depths 𝑍1 and 𝑍2 = 𝑍1 + Δ𝑍 , a first-order Taylor expansion gives Δℎ ≈
𝑓𝐻 𝑑ℎ »»» »» ⋅ Δ𝑍 = − 2 Δ𝑍, 𝑑𝑍 »»𝑍=𝑍¯ 𝑍¯
(16)
2 where 𝑍¯ lies between 𝑍1 and 𝑍2 . Using the identity 𝑍¯ ≈ 𝑍1 𝑍2 for small relative changes, we obtain the square-inverse law reported in the main text:
Δℎ ≈ −
𝑓 𝐻 Δ𝑍 . 𝑍1 𝑍2
(17)
This shows that physically correct scaling must obey a quadratic suppression with depth; any linear-in-1/𝑍 behavior corresponds to a systematic velocity distortion.
20
C.3. Spatio-temporal Coupling in the Image Plane We next derive the depth-coupled suppression of pixel-wise motion used in Eq. (3) of the main text. Starting from the pinhole relation for the horizontal coordinate, 𝑥 = 𝑓𝑥
𝑋 , 𝑍
(18)
consider a 3D displacement (Δ𝑋, Δ𝑌 , Δ𝑍 ) that moves the point from ( 𝑋, 𝑌 , 𝑍 ) to ( 𝑋 + Δ𝑋, 𝑌 + Δ𝑌 , 𝑍 + Δ𝑍 ). The new image coordinate is 𝑋 + Δ𝑋 ′ 𝑥 = 𝑓𝑥 . (19) 𝑍 + Δ𝑍 The induced pixel displacement is therefore 𝑋 + Δ𝑋 𝑋 − ) 𝑍 + Δ𝑍 𝑍 𝑍 ( 𝑋 + Δ𝑋 ) − 𝑋 (𝑍 + Δ𝑍 )
Δ𝑥 = 𝑥 − 𝑥 = 𝑓𝑥 ( ′
= 𝑓𝑥
𝑍 (𝑍 + Δ𝑍 ) 𝑍Δ𝑋 − 𝑋 Δ𝑍 . = 𝑓𝑥 𝑍 (𝑍 + Δ𝑍 )
(20)
This recovers the expression used in the main paper and shows that pixel motion is quadratically suppressed with depth for fixed 3D velocities. C.4. Metric Implementation Details Scale-depth residual (Scale). Given per-frame object pixel height ℎ𝑡 (from SAM2 masks) and aligned depth 𝑧𝑡 (from Mega-SAM), we audit the perspective-scale invariant in log space: 𝑠𝑡 = ln(max(ℎ𝑡 , 𝜖)) + ln(max( 𝑧𝑡 , 𝜖)).
Using the first 𝑛ref = min(5, 𝑇 ) frames, we set 𝑠ref = median(𝑠1 , . . . , 𝑠𝑛ref ),
and compute framewise residuals
(𝑡 )
𝜖scale = ∣𝑠𝑡 − 𝑠ref ∣ ,
𝑡 ≥ 2.
The Scale component entering PDI is 𝐸scale = RMSE(𝜖scale ) .
This term directly measures violation of the ℎ𝑡 𝑧𝑡 ≈ const relation under perspective projection. Motion consistency (Traj). We use a 3D kinematic audit in world coordinates from Mega-SAM pointmaps. For each frame, we extract a robust foreground centroid by masked 3D median pooling; if valid foreground points are insufficient, we inherit the last valid centroid. The centroid trajectory is then temporally denoised by a 1D median filter (𝑘 = 3) per coordinate. With Δ𝑡 = 1/fps, velocity and acceleration are v𝑡+1 − v𝑡 x𝑡+1 − x𝑡 , a𝑡 = . v𝑡 = Δ𝑡
Δ𝑡
21
We define a robust speed reference −6
𝑠ref = max(median(∥v∥), 2 ⋅ median(∥a∥), 10
and acceleration penalty (𝑡 )
𝑝𝑎 = 2 tanh(
),
∥a𝑡 ∥/𝑠ref ). 5
Direction-change penalty is (𝑡 )
1 − cos 𝜃𝑡 , 0,
if both adjacent speeds are sufficiently large, otherwise,
𝑝𝜃 = {
where 𝜃𝑡 is the angle between adjacent velocity vectors. The trajectory residual is (𝑡 )
1 (𝑡 )
1 (𝑡 )
𝜖traj = 2 𝑝𝑎 + 2 𝑝𝜃 ,
and the Traj component is
𝐸traj = RMSE(𝜖traj ) .
Structural Rigidity (Rigidity). To quantify non-physical internal deformation (e.g., the “jello effect”), we evaluate object structural stability using a prioritized three-strategy hierarchy, and the active strategy output is directly used as the third PDI component. 𝑛
1) 3D Pairwise Rigidity (Primary). We sample world-space points q𝑡 from Mega-SAM pointmaps at CoTracker locations. Anchor pairs are selected at 𝑡 = 0 by triple filtering: (i) visibility filtering, (ii) depth-gradient reliability filtering, and (iii) pair scoring that favors both large 3D separation and interior-region reliability (distance to mask boundary). For each frame, ∥q𝑡 − q𝑡 ∥2
𝑟𝑖 𝑗 (𝑡 ) =
MAD({𝑟𝑖 𝑗 (𝑡 )})
𝑗
𝑖
∥q0 − q0 ∥2 𝑗
𝑖
,
𝜌𝑡 =
median({𝑟𝑖 𝑗 (𝑡 )}) + 𝜖
.
The strategy-1 rigidity residual is the temporal mean over 𝑡 ≥ 2: 𝑇
1 (1) 𝜖rigid = ∑ 𝜌𝑡 . 𝑇 −1 𝑡 =2 (Frames with insufficient visible pairs inherit the previous frame score.) 2) 3D Height Stability (Fallback when Strategy 1 is not entered). If 3D points are valid but strategy 1 is unavailable at the dispatcher level, we compute per-frame 3D object height from foreground 𝑦 -span: 3𝐷
ℎ𝑡
and use coefficient of variation: (2)
𝜖rigid =
= 𝑃95 ( 𝑦𝑡 ) − 𝑃5 ( 𝑦𝑡 ), std({ℎ𝑡 }𝑡=1 ) 3𝐷 𝑇
mean({ℎ3𝑡 𝐷 }𝑇𝑡=1 ) + 𝜖
.
22
3) 2D Pairwise Consistency (Degraded fallback). When 3D evidence is unavailable, we use 2D CoTracker pairwise distance ratios: 𝑟𝑖 𝑗 (𝑡 ) = 2𝐷
𝑑 𝑖 𝑗 (𝑡 ) 𝑑 𝑖 𝑗 (0)
std({𝑟𝑖 𝑗 (𝑡 )}) 2𝐷
2𝐷
,
𝜌𝑡
=
mean({𝑟𝑖2𝑗𝐷 (𝑡 )}) + 𝜖
,
and compute 𝑇
1 (3) 2𝐷 𝜖rigid = ∑ 𝜌𝑡 . 𝑇 𝑡 =1 Finally, the rigidity component used by PDI is (1) ⎧ ⎪ 𝜖rigid , ⎪ ⎪ ⎪ (2) ⎪ ⎪ ⎪𝜖 , 𝜖rigid = ⎨ rigid (3) ⎪ ⎪ 𝜖rigid , ⎪ ⎪ ⎪ ⎪ ⎪ ⎩0,
if Strategy 1 is selected, else if Strategy 2 is selected, else if Strategy 3 is selected, if no strategy is available.
D. Geometric Invariants and Perspective Coupling in Longitudinal Motion D.1. Diagnostic Perspective Analysis for Longitudinal Convergence For scenarios involving longitudinal or oblique motion (e.g., objects receding along a path), we introduce the Generalized H-VP Homogeneity as an auxiliary diagnostic tool. This constraint evaluates whether the projected centroid of a rigid object converges to the scene’s vanishing point (𝑉 𝑃 ) in synchronization with its scale reduction. Under the pinhole camera model, this yields a coupled inverse-depth law: Dist(p1 , 𝑉 𝑃 ) ℎ1 𝑍𝑡 , = = 𝑍1 ℎ𝑡 Dist(p𝑡 , 𝑉 𝑃 )
(21)
where ℎ𝑡 is the pixel height, 𝑍𝑡 is the 3D depth, and p𝑡 is the image centroid at frame 𝑡 . While our primary PDI-Score focuses on metrics applicable to arbitrary motion, Eq. (21) provides a deeper geometric probe specifically for the Longitudinal Convergence category. It ensures that the object’s “speed” of receding (trajectory) and its “rate” of shrinking (scale) are physically coupled. In practice, this distance-based formulation is used when a stable 𝑉 𝑃 can be estimated; for near-transverse motion, we transition to the angular formulation described below. D.2. Perspective Coupling via Angular Alignment To further quantify the coupling between foreground motion and background environment (detecting the “sticker-on-screen” effect), we calculate the angular divergence between two independent vanishing points: the motion-derived 𝑉 𝑃 𝑓 𝑔 and the geometry-derived 𝑉 𝑃𝑏𝑔 (extracted via LSD line clustering). Let (𝑐 𝑥 , 𝑐 𝑦 ) be the principal point; we define the respective direction vectors as: d 𝑓 𝑔 = 𝑉 𝑃 𝑓 𝑔 − (𝑐 𝑥 , 𝑐 𝑦 ),
d𝑏𝑔 = 𝑉 𝑃𝑏𝑔 − (𝑐 𝑥 , 𝑐 𝑦 ). 23
(a) Camera Model
(b) Perspective Convergence
Figure 6 ∣ Geometric principles of perspective. (a) The pinhole camera model mapping a 3D scene point P to a 2D image point Xp. (b) Perspective convergence illustrating the vanishing point (VP). Left: Original sequence; Middle: Geometric overlay; Right: Schematic diagram of ground plane projection. The perspective coupling residual, denoted as Δ𝜃 , is computed via cosine similarity: Δ𝜃 =
1 − cos ∠(d 𝑓 𝑔 , d𝑏𝑔 ) , 2
(22)
where Δ𝜃 ∈ [0, 1]. A value of 0 indicates that the object is moving perfectly toward the scene’s natural horizon, while 1 indicates a total contradiction in perspective spaces. We utilize this angular metric instead of Euclidean distance to remain robust against transverse motions where 𝑉 𝑃 coordinates tend toward infinity. D.3. Mathematical Derivation of H-VP Homogeneity To prove the identity in Eq. (21), assume a 3D point P𝑡 on a rigid object moving linearly in world space: P𝑡 = P0 + 𝑠𝑡 d,
where d = ( 𝐷𝑋 , 𝐷𝑌 , 𝐷𝑍 ) and 𝐷𝑍 ≠ 0.
The pinhole projection ( 𝑥𝑡 , 𝑦𝑡 ) is given by 𝑥𝑡 = 𝑓 𝑋𝑡 /𝑍𝑡 and 𝑦𝑡 = 𝑓 𝑌𝑡 /𝑍𝑡 . As the object recedes (𝑠𝑡 → ∞), its image coordinates converge to the vanishing point: 𝑉 𝑃 = ( lim
𝑠𝑡 →∞
𝑓 ( 𝑋0 + 𝑠𝑡 𝐷 𝑋 ) 𝑓 (𝑌0 + 𝑠𝑡 𝐷𝑌 ) 𝑓 𝐷 𝑋 𝑓 𝐷𝑌 )=( ). , lim , 𝑍0 + 𝑠𝑡 𝐷𝑍 𝑠𝑡 →∞ 𝑍0 + 𝑠𝑡 𝐷𝑍 𝐷𝑍 𝐷𝑍
The horizontal distance between the centroid and the 𝑉 𝑃 is: 𝑉 𝑃 𝑥 − 𝑥𝑡 =
𝑓 ( 𝐷 𝑋 𝑍 𝑡 − 𝐷 𝑍 𝑋𝑡 ) 𝑓 𝐷𝑋 𝑓 𝑋𝑡 − = . 𝐷𝑍 𝑍𝑡 𝐷𝑍 𝑍 𝑡
Substituting 𝑋𝑡 = 𝑋0 + 𝑠𝑡 𝐷𝑋 and 𝑍𝑡 = 𝑍0 + 𝑠𝑡 𝐷𝑍 , the numerator simplifies to: 𝐷 𝑋 (𝑍0 + 𝑠𝑡 𝐷𝑍 ) − 𝐷𝑍 ( 𝑋0 + 𝑠𝑡 𝐷 𝑋 ) = 𝐷 𝑋 𝑍0 − 𝐷𝑍 𝑋0 = Constant.
Thus, the 2D Euclidean distance to the 𝑉 𝑃 follows the relationship Dist(p𝑡 , 𝑉 𝑃 ) = 𝑍 , where 𝒞 is a 𝑡 constant determined by the initial geometry and motion direction. Given that a rigid body’s projected 𝑓 ⋅𝒞
24
𝑓𝐻
height is ℎ𝑡 = 𝑍 , it follows that: 𝑡
ℎ𝑡 ∝
1 𝑍𝑡
and
Dist(p𝑡 , 𝑉 𝑃 ) ∝
1 . 𝑍𝑡
Dividing the initial state by the state at time 𝑡 yields the homogeneity relation: Dist(p1 , 𝑉 𝑃 ) ℎ1 𝑍𝑡 . = = 𝑍1 ℎ𝑡 Dist(p𝑡 , 𝑉 𝑃 ) This concludes the derivation, proving that scale changes and radial convergence must be linearly coupled in 3D-consistent videos.
E. Qualitative Visualizations E.1. Visualizing the PDI Computation Pipeline To demonstrate the transparency, interpretability, and robust engineering of our proposed metric, we walk through the step-by-step execution of the PDI evaluation pipeline on a representative generated video (e.g., a vehicle driving forward). The pipeline systematically extracts multi-modal geometric evidence to compute the final physical realism score.
Figure 7 ∣ Input Video Sequence. Five uniformly sampled frames from a representative generated video featuring a moving vehicle. This raw temporal sequence serves as the initial input for our PDI computation pipeline. Step 1: Semantic Targeting (SAM 2). We initiate the pipeline by identifying the auditing subject using Florence-2 for automated text-to-box prompting. These prompts are then fed into SAM 2 to generate 𝑇 and propagate a temporal sequence of binary masks { 𝑀𝑡 }𝑡=1 . This step isolates the subject and provides frame-wise instantaneous pixel heights ℎ𝑡 and 2D spatial boundaries (Figure 8a). Step 2: 3D Geometric Uplifting (MegaSaM). To recover the latent 3D physical environment, we employ 𝑇 MegaSaM to obtain a coherent depth sequence {𝑍𝑡 }𝑡=1 , the estimated focal length 𝑓 , and camera poses. More importantly, MegaSaM projects every pixel into a unified 3D world coordinate system, yielding 𝑇 × 𝐻 ×𝑊 ×3 world-space pointmaps P𝑤𝑜𝑟𝑙𝑑 ∈ R (Figure 9). This critical step lifts 2D observations into pure 3D space, decoupling object kinematics from camera ego-motion. Step 3: 3D Structural Anchoring (CoTracker3). With the 3D world-space constructed, we deploy CoTracker3 to monitor the subject’s internal structural integrity. Within the region defined by 𝑀1 , we 𝑛 𝑛 seed anchor queries and obtain a set of reliable 2D pixel-space trajectories {(𝑢𝑡 , 𝑣𝑡 )} (Figure 8b). By using these as spatial indices into the MegaSaM pointmaps, we lift each tracked anchor to its 3D coordinate: 𝑛 𝑛 𝑛 q𝑡 = P𝑤𝑜𝑟𝑙𝑑 [𝑡, 𝑣𝑡 , 𝑢𝑡 ], transforming 2D visual tracking into structurally meaningful 3D trajectories.
25
(a) Target Isolation. SAM 2 accurately isolates the target vehicle.
(b) Dense Tracking. CoTracker3 produces long-term 2D trajectories inside the target mask.
Figure 8 ∣ Pipeline intermediate 2D processing. The system first segments the semantic target (a) and subsequently extracts dense kinematic trajectories (b) for downstream auditing. Step 4: Three-Dimensional Geometric Auditing. Utilizing the extracted 2D tracks and 3D geometries, the system audits the object’s physical rationale across three orthogonal dimensions: • Scale (𝜖𝑠𝑐𝑎𝑙𝑒 ): Audits whether the target’s projected height and depth satisfy the ℎ ⋅ 𝑍 = const invariant (pinhole camera model). • Trajectory (𝜖𝑡𝑟𝑎 𝑗 ): Evaluates whether the centroid motion in 3D world coordinates follows Newtonian inertia, penalizing abrupt acceleration jumps or non-physical reversals. • Rigidity (𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦 ): Quantifies the temporal instability of the 3D distances between anchor pairs via the Coefficient of Variation, penalizing “volume breathing” or “Jello effect” artifacts.
(b) Frame 𝑡2
(c) Frame 𝑡3
(a) Frame 𝑡1
Figure 9 ∣ Dense 3D Reconstruction via MegaSaM. Multi-view visualizations of the recovered 3D point clouds. Extracting these temporally consistent physical structures establishes a reliable geometric foundation for auditing spatial rigidity. Step 5: Score Synthesis & Final Report. The individual dimension errors are aggregated into the final Perspective Distortion Index (PDI) using the weights (𝑤1 = 0.4, 𝑤2 = 0.4, 𝑤3 = 0.2): PDI = 𝑤1 ⋅ RMSE(𝜖𝑠𝑐𝑎𝑙𝑒 ) + 𝑤2 ⋅ RMSE(𝜖𝑡𝑟𝑎 𝑗 ) + 𝑤3 ⋅ 𝜖𝑟𝑖𝑔𝑖𝑑𝑖𝑡 𝑦
(23) 26
A lower PDI score indicates superior physical rationale. The system ultimately generates a terminal audit report (shown below), providing fine-grained diagnostics to pinpoint specific failure modes.
F. Qualitative Analysis of Failure Modes While PDI-Bench provides robust quantitative scores, qualitative examination reveals distinct geometric failure modes that persist even in state-of-the-art generative models. Temporal Orientation and Structural Incoherence. A unique failure mode observed in models like Veo 3.1-Fast is the abrupt inversion or morphing of object features, where frontal and rear characteristics appear to swap or deform mid-sequence (Fig. 10). Crucially, these artifacts often manifest as instantaneous events. As the transition occurs within a duration shorter than the temporal window of our SfM-based reconstruction, the underlying pointmaps struggle to track the rapid topology change, leading to transient spikes in Structural Rigidity (𝜖𝑟 ) and Motion Consistency (𝜖𝑡 ). While PDI-Bench captures the resulting geometric inconsistency, these extreme behaviors highlight the difficulty of monocular 3D systems in maintaining a coherent object-centric frame of reference during non-physical transitions.
(a) Biological Motion: Articulated dynamics under rapid motion.
(b) Longitudinal Convergence: Geometric artifacts during camera-object interaction.
Figure 10 ∣ Qualitative Failure Modes. (a) A bird in flight exhibiting motion-induced blur and feature morphing. (b) A bus sequence showing temporal inconsistencies where the object’s appearance and 3D structure become spatially decoupled across consecutive frames.
G. Limitations and Future Directions While PDI-Bench establishes a rigorous geometric yardstick, several limitations remain. First, the framework inherits a dependency on off-the-shelf perception modules: when SAM 2, CoTracker3, or MegaSAM fail in low-texture or low-parallax regimes, the system must fall back to 2D proxies and reconstructionaware weighting, reducing the depth sensitivity of the final score. Second, our core invariants are derived under a rigid-body assumption; although the rigidity residual partially compensates for mild non-rigid motion, highly deformable or amorphous objects (e.g., fluids, cloth, or crowds) are not fully captured by the current formulation. Third, disentangling complex 3D rotations from axial translation using monocular cues is fundamentally ill-posed; the Model Consensus mechanism mitigates most false positives, but residual noise can persist in extreme curved-motion scenarios.
27
We see three main avenues for future work. One direction is to integrate learned priors by distilling PDIBench into a lightweight neural assessor that approximates our geometric residuals without explicit SfM, enabling larger-scale or near real-time auditing. Another is to extend the current Euclidean invariants to cover richer physical phenomena, such as fluid dynamics, articulated bodies, and multi-object interactions, potentially via hybrid Lagrangian–Eulerian formulations. Finally, incorporating multi-view or multi-sensor inputs (e.g., stereo, depth sensors, IMU traces) could reduce the inherent ambiguities of monocular geometry and allow PDI-Eval to serve as a calibration tool for future, more physically grounded world models.
H. Broader Impacts This work introduces PDI-Bench, a benchmark designed to evaluate the physical consistency and geometric realism of video generation models. We believe the proposed benchmark can provide several positive societal impacts. First, by enabling systematic evaluation of geometric and physical plausibility, PDI-Bench may help improve the reliability and interpretability of generative video systems. Such evaluation tools can benefit applications in robotics, embodied AI, simulation, digital content creation, and educational visualization, where physically coherent video generation is important. Second, the benchmark promotes transparency by exposing failure modes of current video generation models, which may encourage the development of safer and more trustworthy generative systems. At the same time, advances in realistic video generation may also introduce potential risks. Improving the physical realism of generated videos could contribute to the creation of increasingly convincing synthetic media, which may be misused for misinformation, deceptive content generation, or malicious manipulation. In addition, highly realistic generated videos could complicate the detection of synthetic media in sensitive contexts. By quantitatively measuring physical consistency and identifying model weaknesses, the benchmark may support future research on trustworthy generative modeling, robustness analysis, and synthetic media detection. We encourage future work to further investigate safeguards and responsible deployment practices for realistic video generation technologies.
28