Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos Anurag Ghosh1 Francesco Pittaluga2 Khiem Vuong1 Angela Chen1 Juan Alvarez-Padilla3 Manmohan Chandraker2,4 Srinivasa Narasimhan1 1
Carnegie Mellon University
2
NEC Labs America
Philadelphia
3
MIT∗ 4 UC San Diego
Cone
Barrels
Barrier
Person
Vehicle
arXiv:2606.07366v1 [cs.CV] 5 Jun 2026
140 m
In-the-wild Dashcam Video
Lifted Long-Tail Objects
Georeferenced Scene Reconstruction
Overlaid on Open-Street Map
Closed-Loop Simulation
t = 1 sec
Boston
San Francisco
t = 8 sec
t = 5 sec
New York
Chicago
Figure 1: Top row: Dash2Sim takes an in-the-wild monocular dashcam video (left) and recovers a georeferenced 3D reconstruction at metric scale (center). Long-tailed objects such as cones, barrels, and barricades are detected, tracked, and lifted to 3D (right). Middle row: The resulting 4D driving log supports closed-loop simulation with reactive agents. Bottom row: Applying Dash2Sim to a large dashcam dataset produces the ROADWork4D benchmark dataset, bringing the long-tail scenarios of in-the-wild driving to a range of tasks, including novel-view synthesis and closed-loop planning.
Abstract: Self-driving simulations typically rely on data collected in a small number of cities or on hand-authored synthetic scenarios. Dashcam videos cover a far broader range of locations and situations, including rare or long-tailed scenarios. They are considered less usable for simulation because it is difficult to recover accurate 4D scenes from monocular in-the-wild videos. Work zones are one such class of long-tailed situations that dashcams capture. We present Dash2Sim, a framework that turns in-the-wild monocular dashcam videos into metric, georeferenced 4D driving logs compatible with existing simulators, and verifies each one against an independently maintained map without annotations. We apply Dash2Sim to a large video corpus to create the ROADWork4D benchmark dataset, which spans 4,244 scenes with 2.7M 3D objects across 17 cities. On a verified subset ROADWork4D-CL (2,201 scenes), we study privileged closed-loop planners and find that work zone scenarios are difficult: while rule-based and hybrid planners generalize better than learning-based ones, all fall short, failing to make the lane changes that temporary work zone channels require. Beyond planning, dense depth recovered by Dash2Sim improves novel-view synthesis quality by up to 19% on perceptual metrics, suggesting its potential to provide rich conditioning for closed-loop sensor simulation from monocular videos. Keywords: Self-Driving Simulation, 4D Scene Recovery, Long-Tail Self-Driving ∗
Work done at Carnegie Mellon University
Table 1: Existing self-driving simulation paradigms. We compare scale and capabilities across existing driving simulation paradigms. Our framework recovers 4D driving logs from in-the-wild dashcam videos, supporting evaluation on realistic long-tail scenes. Rows are grouped by closed-loop support and data source. Scale Source
Long-Tail
Locales
Maps
Hours
Scenes
Closed-Loop
Reactive
Open-loop KITTI [1] nuScenes [2] Argoverse [3, 4] WOMD [5] WOD-E2E [6]
Fleet Fleet Fleet Fleet Fleet
✗ ✗ ✗ ✗ ✓
1 city 2 cities 6 cities 6 cities 6 cities
✗ ✓ ✓ ✓ ✗
6 5 — 570 12
— 1000 250k 100k 4,021
✗ ✗ ✗ ✗ ✗
✗ ✗ ✗ ✗ ✗
Pseudo Closed-loop, real fleet data Navsim [7, 8] Fleet
✗
4 cities
✓
—
12k
✓*
✗
Closed-loop, real fleet data nuPlan [9] Fleet
✗
4 cities
✓
1282
—
✓
✓
✓ ✓ ✓ ✓
— — — —
— 220 200 335
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
✓
35.2
4,244
✓
✓
Closed-loop, synthetic or synthetically-augmented data CARLA [10] Synthetic ✗ 8 Virtual towns Bench2Drive [11] Synthetic ✓ 12 Virtual towns Fail2Drive [12] Synthetic ✓ — InterPlan [13] Fleet-Synthetic ✓ 4 cities Ours
1
Eval
Benchmark
Dashcam
✓
17 cities
Introduction
Dashboard cameras have quietly become standard equipment on a large portion of consumer vehicles [14]. The footage they collectively produce, even just the fraction uploaded to public video platforms, spans more cities and unusual situations than any publicly available autonomous vehicle dataset. However, exploiting this data for learning how to drive autonomously is underexplored. There are two main reasons for this discrepancy. First, accurate 4D scene recovery from monocular videos is difficult and considered brittle [15]. For closed-loop simulation (we consider the privileged planning setup [9]), the 4D driving log needs metric, geo-referenced trajectories, localized objects with consistent identities, and a routable map. The monocular setting is especially hard, as any recovered geometry is scale-ambiguous [16], and, in driving specifically, forward motion does not provide enough parallax even if the scene is rigid (which it is not). Second, even when 4D scene recovery succeeds, the quality of the recovered data is difficult to determine at scale. Thus, prior work has used dashcam videos largely for perception [17, 18, 19], or abstracted scenes into behavioral scripts for synthetic engines [20, 21] that do not recover the entire scene. Instrumented fleets sidestep these issues by collecting data with cross-sensor and map agreement, but are expensive to scale, underrepresenting long-tailed settings such as work zones (Tab. 1). Work zones specifically are a persistent failure mode for commercial driving systems even in 2026 [22, 23] (See Fig. E.2). We address these challenges with two observations. First, existing public infrastructure such as co-located street imagery can act as anchor references to reliably recover scene geometry at scale. Second, replaying monocular 4D driving logs in a closed-loop simulator, against an independently maintained map, yields a scalable, annotation-free signal of data quality (Fig. A.1). This allows us to filter out driving logs without manual inspection. Our framework recovers ego trajectories and metric-scale dense depth, and lifts 3D objects from monocular dashcam videos to produce ROADWork4D, a benchmark of 4D driving logs from work zones. Crucially, 2,201 scenes pass non-reactive log-replay verification to produce ROADWork4DCL, for training and evaluating privileged closed-loop planners. On closed-loop privileged planning, hybrid Pluto [24] and rule-based PDM-Closed [25] outperform all learning-based planners by up to 40 points (§4.2) but still perform poorly overall, showing that rules generalize better even to a class of long-tailed scenarios, extending a known fleet-data finding in normal driving situations [25]. Our ablations indicate that work-zone layouts are the primary generalization challenge, suggesting that learned planners would benefit from training data with better coverage of rare work zones. Lastly, the recovered depth from ROADWork4D improves novel-view synthesis by up to 19% on perceptual metrics (§4.3), opening the door for closed-loop sensor simulation from monocular videos. Our results bring to light the potential of dashcam-derived 4D logs as a viable input to existing closed-loop simulators and a useful conditioning signal for closed-loop sensor simulation, serving as a complement to fleet-derived and synthetic data sources for advancing long-tail autonomous driving. 2
2
Related Work
Table 2: ROADWork4D at a glance. StatisScaling In-the-wild Data for AV and Robotics. Scal- tics for the entire dataset. The verified subset ing in machine learning and robotics increasingly comes used for closed-loop planning is summarized in from new, in-the-wild data sources rather than more in- Table A.1. domain collection: pooling heterogeneous robot logs Scale across embodiments [26], distilling web-scale video Scenarios 4,244 into manipulation policies [27], and curating in-theUS cities 17 wild web data, which drives the largest gains in mulFrames @ 5 Hz 633,404 timodal learning [28]. Even end-to-end driving planDriving time 35.2 h ners are now built on internet-scale foundation modReconstruction els [29]. On the data side, many in-the-wild dashcam SfM 3D points 276.3 M datasets [17, 18, 19] help cover the long tail but largely 3D annotations stop at two-dimensional perception. Earlier dashcamInstance tracks 142,863 to-simulation efforts [20, 21] replay extracted vehicle dynamic / static 65.0k / 77.9k trajectories as scripted scenarios inside a synthetic enPer-frame 3D boxes 2.74 M gine rather than recovering the full scene. By turning abundant dashcam video into simulation-ready logs with an annotation-free verification signal that filters quality at scale (§3), Dash2Sim brings this scaling perspective to autonomous driving, offering a viable path beyond expensive fleet collection toward solving the long-tail problem. Autonomous Driving and Simulation. Autonomous driving simulation falls into two categories. The first builds on data collected by instrumented fleets driving through a small number of cities [1, 2, 5, 9], with rare-scenario variants [6, 30] also drawn from fleets. The second relies on synthetic simulators [10], with rare and adversarial scenarios [11, 12] created by hand: an approach that inherently has larger visual and behavioral sim-to-real gaps than in-the-wild data sources. Another direction of work addresses sensor simulation, either via generative models [31, 32, 33, 34, 35] or via Gaussian splatting [36, 8, 37]. Both directions require rich conditioning signals and largely rely on fleet data sources. Dash2Sim provides 4D scenes at scale, along with conditioning for sensor simulation, layout [38], and behavior generation from monocular video. Closed-loop Planning for Autonomous Driving. Deployable planners are developed on realistic closed-loop simulators [9], with rule-based [25], learned, and hybrid planners [39, 40, 24, 41] all actively explored in the privileged planning setting. End-to-end VLAs [42, 43, 44, 45, 46] are also being heavily explored; they primarily target image-to-trajectory evaluation, are generally evaluated using fleet-derived benchmarks [6, 7], and assume multi-modal sensor inputs. Regardless, all of these methods are trained and evaluated on fleet logs, which have limited long-tail coverage. We propose the ROADWork4D-CL benchmark for training and evaluation on long-tail scenarios.
3
Dash2Sim Framework
Recovering a closed-loop driving log from in-the-wild dashcam video requires three properties that fleet vehicles sidestep through instrumentation. The ego trajectory should be metric and georeferenced at high accuracy. Surrounding agents and objects should carry open-set categories and consistent 3D identities across time. The log must be registered to a routable map. A dashcam provides a single uncalibrated video stream with, at best, consumer-grade GPS. Dash2Sim proposes the following approach. We anchor reconstruction to co-located street imagery (§3 and §D.1), detect and track agents and objects in 2D [17, 47], lift them to 3D [48] conditioned on the recovered depth (§D.4), and produce logs compatible with existing simulators (§3). We discuss key design decisions below, and provide detailed descriptions of the framework in §D. Metric Reconstruction from Dashcam Video via Geo-anchored SfM. A few issues make this challenging. Monocular localization is scale-ambiguous since the recovered trajectory and scene are 3
Denver
Philadelphia
Figure 2: Dense, geo-referenced point clouds recovered by Dash2Sim. ROADWork4D scenes after reconstruction. Roads, lane markings, sidewalks, vegetation, and small long-tailed work-zone objects (cones, tubular markers, signs) are reconstructed at metric scale from in-the-wild monocular videos. Best viewed zoomed in. Supplemental material, website and walkthrough video contain additional media and visualizations.
determined only up to a global similarity transform [16]. The dashcam’s GPS, if available, routinely places the vehicle in adjacent city blocks (See Fig. 4), and most in-the-wild videos lack GPS entirely. Our observation is that recovering a global metric reference is a retrieval problem, not a sensing problem: Google Street View panoramas provide dense, GPS-tagged anchors, building on work that uses street imagery for localization [49, 50, 51]. We retrieve nearby panoramas, render perspective views, and run SfM over the dashcam and Street View images together. A similarity transform from the SfM-recovered street imagery poses to their known GPS locations then transfers scale and geo-references the dashcam trajectory (Fig. 4, blue). Street imagery can provide high accuracy on dashcam videos (verified on nuScenes [2] in §D.2). We use dashcam GPS only to narrow down retrieval of co-located street imagery, and visual place recognition [52] can potentially replace it when no GPS is available (§D.3). A detailed description of our approach is provided in §D.1. Validating 4D Driving Logs for Driving Simulation at Scale. Closed-loop planning needs three components in addition to a 4D driving log: a simulator interface, a routable map, and a reactive-agent model. We employ nuPlan [9] as our closed-loop simulator interface, and we use OpenStreetMap [53] for our maps following prior work [54, 55]. A more detailed description is provided in §D.6. To determine which subsets are usable for closed-loop planning without expensive manual verification, we observe that the source videos contain no at-fault driving events: the driver completed each route without any incidents. In privileged planning, the simulator scores a trajectory against ground-truth perception and map, so the driving score evaluates the recovered trajectory, objects and agents, and the map jointly. Replaying the recorded ego trajectory as the planner’s output and rolling out with non-reactive agents therefore tests whether all three are mutually consistent: any penalty is likely to be an error in our recovered log or the map, as it is unlikely to be an error in the driving. Non-reactive agents ensure that we measure errors in our log, and map independence ensures that our logs are spatially and physically consistent. For example, a misaligned map would place the ego outside drivable areas. Because this score-based verifier needs no ground truth, it scales to all scenarios across 17 cities (§A). As a different simulator with alternative scoring mechanisms could yield a different subset, we will release all scenarios for further analysis. Data source. We use the ROADWork [17] dataset as our primary data source, as it provides dashcam video recordings of navigating through work zones across the U.S. We focus on work zones since they are a persistent hurdle for commercial autonomous deployment [22, 56] (See §E.2 for detailed discussion). We name the resulting corpus ROADWork4D (4,244 scenarios), which supports many autonomous-driving tasks. We use its automatically verified subset ROADWork4D-CL (2,201 scenarios), which passes non-reactive log-replay verification (§3), for closed-loop planning. ROADWork4D at a glance. It spans 17 US cities and 724.5 km of recovered ego trajectory, with 142,863 instance tracks providing 2.74 M per-frame 3D boxes across both common dynamic agents and long-tailed road objects (See Table 2). 4
Indianapolis
Boston
Columbus
Figure 3: Driving logs from ROADWork4D recovered by Dash2Sim. Our logs include long-tailed work zone scenarios with rare objects, layouts and complex lane change behaviors such as crossing the double yellow line in a work zone (top row). We believe our Dash2Sim framework and ROADWork4D dataset would spur research in a variety of self-driving tasks, including closed loop planning and sensor simulation. Supplemental material, website and walkthrough video contain additional media and visualizations.
4
Evaluation and Downstream Applications
First, we describe how we verify the 4D driving logs that Dash2Sim recovers for ROADWork4D, since manually verifying every annotation is infeasible at scale. Second, although ROADWork4D supports many uses, we highlight two downstream applications that motivate its utility for autonomous driving. 4.1
Recovering and Verifying 4D Driving Logs Coarse Traj.
We validate Dash2Sim components against Recovered Traj. ground truth or downstream proxies. For example, Street View anchoring recovers metric ego trajectories with median translation error within 10 cm and rotation error within 1◦ on nuScenesmini [2], with scale within 0.3% of unity (Figure 4), though a potential alignment error of 3–9 m remains in the global frame (§D.2). Sequences captured in Singapore were also reconBoston New York structed, suggesting the approach can generalize Philadelphia with co-located street imagery. We validate the value of our recovered depth via novel-view syn- Figure 4: Recovered ego trajectories. Red: coarse GPS tracks from the dashcam video. Blue: metric 6-DoF thesis (§4.3). We also show that dashcam GPS, poses recovered by Dash2Sim. Coarse GPS drifts into while useful, is not necessary for co-located adjacent blocks, while recovered poses align with the street imagery retrieval (§D.3). We validate our driven lanes. 3D object lifting pipeline on a multi-modal longtailed dataset, WorkZone3D [57] (§D.5). Overall, Dash2Sim recovers 4,244 scenes with ego trajectories, dense depth, and object tracks from 4,375 videos in the underlying source [17]. ROADWork4D is suitable for studying a variety of tasks, such as closed-loop planning, novel-view synthesis, generative simulation, or in-the-wild AV data curation. We will release the entire set to the community for advancing long-tailed autonomous driving research. 5
Table
3: Closed-loop Evaluation on ROADWork4D-CL. Learned planners [40, 24, 58] struggle to generalize, while hybrid planner [24] and rule-based planner [25] perform much better. CLS-NR Planner
Score↑
Coll.%↓
Table 4: Effect of channel sensitivity (CLS-NR). Channel-sensitive evaluation penalizes drivable-area deviations from the GT route.
CLS-R Score↑
Channel-Sensitive
Coll.%↓
Learned PlanTF [58] Diffusion Planner [40] Pluto [24]⋆
17.7 24.0 32.7
4.3 51.8 10.5
18.0 31.3 40.3
4.3 39.1 6.8
Rule-Based PDM-Closed
53.0
2.2
53.7
1.8
Hybrid Pluto [24]
57.5
4.7
59.7
1.6
Channel-Insensitive
Planner
Score↑
DAC Fail%↓
Score↑
DAC Fail%↓
Learned PlanTF [58] Diffusion Planner [40] Pluto [24]
17.7 24.0 32.7
18.0 30.9 23.4
22.9 30.9 43.0
5.3 9.2 4.3
Rule-Based PDM-Closed
53.0
27.3
57.4
9.4
Hybrid Pluto [24]
57.5
17.8
68.8
1.2
As an end-to-end test, we employ non-reactive log-replay for verification. Overall, 2,201 scenes from all 17 cities produce a high driving score (> 90) under verification. The verified subset approximately follows the geographical distribution of the entire set (See §A). This confirms that the recovered trajectory, objects, and independent map are mutually consistent and the logs are usable for closedloop planning. While a high score indicates log usability, a lower score does not preclude its validity. In nuPlan’s case, the scoring mechanism is conservative. For example, the nuPlan log-replay score on nuPlan Test14-Hard [58] (long-tailed driving situations from the fleet) is 85, below our conservative threshold. We discuss these limitations in §A. 4.2
Privileged Closed-Loop Planning with Dashcam Videos
Planners trained on normal driving data likely have limited exposure to long-tailed scenarios such as work zones (See Fig. 3 for some scenes, Fig. C.3 shows more examples), making ROADWork4D-CL a suitable long-tailed generalization benchmark for privileged closed-loop planning (See §E.2 for rationale). Setup. We evaluate PlanTF [58], Diffusion Planner [40], Pluto [24] (learned and hybrid variants), and PDM-Closed [25] in a privileged closed-loop setting using the nuPlan simulator [9]. All planners are trained on 1 million nuPlan scenario tokens without fine-tuning on ROADWork4D-CL, except for when we train PlanTF [58] to verify learnability and generalization across cities. We report non-reactive (CLS-NR) and IDM-reactive (CLS-R) closed-loop scores following nuPlan with one change in the definition of drivable area to include work-zone-specific channel sensitivity (§C.1). NC
1.0 Results and Insights. Our results show that a counterin0.8 tuitive finding from prior work [25], shown on regular driv0.6 0.4 PlanTF ing scenarios, extends even to a broad class of long-tailed EP DAC Diffusion Planner 0.2 Pluto (Learned) scenarios: rule-based and hybrid planners generalize betPDM-Closed ter than learned planners. The hybrid planner, Pluto [24], Pluto (Hybrid) achieves the highest overall score by combining a learned planner with a rule-based safety mechanism. A rule-based Comf. TTC planner is not far behind, outperforming all purely learned planners by 20 to 35 points (Table 3). Figure 5: Planner Error Analysis. Generalization Analysis of Planners. Among the learned PlanTF [58] shows low progress, Diffusion Planner [40] is aggressive and has a high planners, Pluto [24] outperforms PlanTF [58] and Diffu- crash rate, and Pluto [24] sacrifices comfort sion Planner [40] by a large margin, and each learned plan- for progress. PDM-Closed [25], while being ner fails differently (Fig. 5). We observe that the choice of safe, is very conservative. Pluto Hybrid [24] input representation plays a role, as Pluto’s reference-line balances all metrics. input [24] likely provides additional structure for planning, though this design choice also reduces channel sensitivity required in work zones (§C.1).
6
Table 5: Held-out novel-view synthesis on ROADWork4D scenes. Adding depth supervision from Dash2Sim improves both low-level image processing metrics and high-level perceptual metrics, with the largest improvements on the perceptual metrics that quantify human-observable visual fidelity. Method OmniRe [37] OmniRe [37] + Ours
PSNR↑
SSIM↑
LPIPS↓
FID↓
cFID↓
DINO-S↓
DSim↓
22.21 22.42 (+0.9%)
0.769 0.777 (+1.0%)
0.269 0.232 (-14%)
104.1 89.7 (-14%)
18.5 14.9 (-19%)
0.090 0.081 (-10%)
0.120 0.101 (-16%)
Role of Channel Sensitivity in Work Zones. In channel-sensitive evaluation, the drivable area is restricted to only the lane segments along the ground-truth route the human driver followed. In channel-insensitive evaluation, all mapped lane segments are considered drivable (details in §C.1). Lane closures, merges, and detours in work zones frequently require crossing into adjacent lanes or taking alternate paths, a failure mode commercial autonomous vehicles also exhibit [56]. All planners improve in channel-insensitive mode (Table 4). Curiously, Pluto [24] shows the largest improvement, with its DAC failure rate dropping drastically. We observe that compared to PlanTF [58] and Diffusion Planner [40], Pluto’s [24] input representation includes reference lines that encode lane centerlines, which likely provides additional structure for planning in normal driving but reduces channel sensitivity when work-zone closures require deviation from lane centerlines. PDMClosed [25] shows a similar pattern, indicating that centerline following, while helpful in normal driving, is likely not an optimal design choice for navigating work zones. Additional Insights. In §C.1, we further analyze the role of static work zone objects on planner performance. Removing dynamic agents has little effect on overall scores, indicating that work zone layouts are the primary generalization challenge for privileged motion planners. Also, to verify if motion planning is learnable from ROADWork4D, we train PlanTF [58] on increasing amounts of data (initially from one city, then three cities, then five, ten, and finally all cities), and observe that data from even one additional city improves performance on the full 17-city set. We observe that it eventually matches zero-shot Diffusion Planner [40] in performance, while recovering most of the gap with data from 10 cities only, suggesting that diversity across cities helps with generalization. Our results conclusively show that 4D driving logs recovered from dashcam video by Dash2Sim are sufficient to observe and study meaningful differences in closed-loop planning behavior. Our work additionally extends important insights from prior work that were extracted from fleet-derived data [25], supporting dashcam-based simulation as a complementary evaluation source for autonomous driving. 4.3
Towards Closed-Loop Sensor Simulation from Dashcam Videos
Privileged planning assumes access to ground-truth perception and cannot expose failures in detecting or interpreting the rare objects and temporary signage necessary for navigating work zones (§E.2). Closed-loop sensor simulation can test these perception-dependent failures by rendering sensor streams during simulation rollouts, motivating the need for novel-view synthesis to generate realistic sensor inputs. Setup. We reconstruct ROADWork4D scenes using OmniRe [37], a widely adopted urban driving reconstruction framework for novel-view synthesis. We consider the photometric-only mode as our baseline, and adding an inverse-depth ℓ1 loss using the per-frame depth in §3 (OmniRe + Ours). Every 10th frame is held out for novel-view evaluation following [37]. Reconstruction quality. Depth supervision improves every metric on the held-out test split (Table 5). LPIPS improves by 14%, CLIP-FID by 19%, DreamSim by 16%, and PSNR improves by 0.9%. Full-split results (§C.2) show depth also improves rendering of training views, implying that depth regularizes the underlying Gaussian Splat representations. To evaluate the rendering quality of small long-tail objects, we perform evaluation on object crops across 9 work zone categories in ROADWork4D. §C.2 shows that our depth supervision provides largest improvements on small objects like cones (20%). Figure 6 shows visual comparisons: depth 7
OmniRe
Ours
Boston (GT)
Houston (GT)
Figure 6: Novel View Synthesis of ROADWork4D Scenes with Our Depth Supervision. Dense depth from Dash2Sim improves novel-view synthesis quality, especially for small long-tailed objects such as vertical panels and temporary traffic control signs that contain textual details relevant for navigating work zones [17].
improves perceptual details of both common and long-tailed objects, including legibility of temporary signage that end-to-end models must interpret to navigate work zones [17]. More broadly, these results suggest that the dense depth recovered by Dash2Sim from monocular videos can serve as a useful conditioning signal for closed-loop sensor simulation of long-tailed driving scenarios.
5
Conclusion
Limitations and Future Work. Dash2Sim recovers 4D driving logs from monocular dashcam videos for closed-loop simulation, but has limitations when compared to fleet-derived simulation along several axes. First, a single forward-facing camera leaves most of the scene unobserved; second, our proposed framework has errors across stages; and third, some error sources such as mismatched maps lie outside the framework and are difficult to automate. The choice of simulator and scoring mechanism also affects our “yield”, and alternative simulators could produce a different subset of ROADWork4D-CL. Moreover, aspects of our planning evaluation rely on rule-based reactive agents [9, 59] that do not model complex interactions often required for long-tailed autonomous driving, especially while driving in work zones. We posit that improved 3D reconstruction [60, 61], learned traffic models [62], and real-world validation [63] are promising directions for closing many of the remaining gaps. We discuss each limitation and potential future work in detail in §B. Despite these limitations, Dash2Sim is, to our knowledge, the first framework to recover planningcompatible 4D driving logs from monocular in-the-wild videos at scale. Existing benchmarks require data from instrumented fleets, limiting coverage of long-tail scenarios. Dash2Sim improves accessibility by enabling learning from in-the-wild videos, a potentially unlimited data source to complement other sources. Non-reactive log-replay verification provides a scalable data quality signal for the recovered 4D logs, which could pave the way for expansion to new data sources. ROADWork4D extends the long-tail coverage established by ROADWork [17] from the 2D to the 4D regime, supporting a variety of tasks in self-driving simulation and planning. Our planning evaluation shows that planners trained and evaluated on fleet data do not generalize reliably to these long-tailed scenarios. And our novel-view synthesis evaluation shows the potential of closed-loop sensor simulation with dashcams. More broadly, our work points to in-the-wild dashcam-derived simulation as a scalable path to improving long-tail robustness of autonomous driving systems.
8
(a) Instrumented Collection
(b) Interactive Collection
(c) In-the-wild Observation Recovered agent & env. state
Ego sensor A Cross-modal agreement
Data quality
Agent
action obs.
Environment
task success
Spatial reference
Simulation consistency
Data quality
Physical constraints
Ego sensor B
Co-collected Sensor Agreement multi-modal sensors co-located on agent
Data quality
Task Success in Environment access to environment available
Cross-source Physical Consistency in Simulation physical realism provides quality signal
Figure A.1: Three paradigms for data verification. (a) Instrumented platforms can verify through cross-modal sensor agreement. (b) When the agent has environment access, task success provides the quality signal [64, 65]. (c) For in-the-wild data where neither is available, mutual consistency across independently maintained references serves as a quality signal.
A
Discussion
In-the-wild robotics data curation. Accurately recovering a scene from in-the-wild videos is difficult, but deciding which recovered logs are useful for training on, at scale and without labels, is also a difficult problem. Prior work uses task success as a label-free quality signal at scale in other settings [64, 65], and offline RL uses environment reward to assess which data is high-quality [66, 67] (Fig. A.1). In multimodal and language understanding, large-scale curation and filtering of in-the-wild web data has yielded the largest gains [28], suggesting that in-the-wild data curation could play a similar role for scaling autonomous driving and solving the long-tail. The same verification principles apply to dashcam videos that contain no at-fault events: the score checks whether the reconstructed environment is consistent with the video, and closed-loop metrics play the same role for driving. Scalable in-the-wild verification. Unlike fleets, which verify through co-collected sensors, and synthetic benchmarks, which are correct by construction, monocular video-derived benchmarks have neither, so assessing data quality at scale is far more difficult. Monocular 4D scene recovery in the wild is also brittle [15], which makes a scalable, annotation-free quality signal essential. We therefore employ non-reactive log-replay verification (§3), which uses the simulator’s driving score, grounded in physical constraints (drivable-area compliance, collision avoidance, etc.) and evaluated against an independently maintained map. Our verified scenarios approximately match the initial geographic distribution (Fig. A.2; Table 2 reports the entire recovered corpus), indicating that the verification paradigm is robust and not prone to systematic bias in Dash2Sim. However, while this mechanism can measure certain inconsistencies (e.g. map errors, errors in ego-pathway due to 3D reconstruction, phantom objects appearing in the driven pathway), it is far from complete. For example, semantic errors do not affect the log-replay score for the most part (nuPlan does distinguish between collisions with static objects and dynamic agents). It also requires that source videos contain no at-fault driving events, so extending it to, say, videos of accidents, requires distinguishing real collisions from reconstruction artifacts. Lastly, and more fundamentally, score-based verification inherits the simulator’s notion of “good driving”, so a faithfully recovered log can still be penalized when the real driving departs from such a notion. This is a property of any score-based simulator: many of these metrics are calibrated to nominal, comfortable and rule-following behavior, whereas the long-tail scenarios worth capturing often demand the opposite.
Table A.1: ROADWork4D-CL closedloop benchmark at a glance. The verified subset of ROADWork4D used for closedloop planning evaluation; full dataset statistics are in Table 2. Scale Scenarios US cities Frames @ 5 Hz Driving time
2,201 17 328,454 18.3 h
Reconstruction SfM 3D points
141.7 M
3D annotations Instance tracks 77,978 dynamic / static 36.7k / 41.3k Per-frame 3D boxes 1.50 M
Concretely, even when the recovered 4D log is correct, its score can be low because (a) comfort terms penalize the hard braking and sharp steering that are 9
800 Number of Scenarios
ROADWork4D Benchmark: Scenario Distribution by City
837
All (4244) nuPlan Compatible (2201)
654
600
519
464
400
463 371
356 244
225
200
209 111
0
ton
s
Bo
Los
An
ge
les
t
De
t roi
208 136
109
158
107
150 79
121 99
110
56
107
62
101
52
92
41
64 39
42
13
38
8
ver
s y o o a o e e e n is is ix oni isc lott phi attl Cit bu ag pol pol sto en vill n De n Ant Franc Char iladel Se York olum Chic nnea iana Hou Pho kson C w Jac Mi Ind Sa San Ph Ne
Figure A.2: ROADWork4D scenario distribution by city. All 4,244 scenarios (blue) and the 2,201 that pass non-reactive log-replay verification for privileged closed-loop planning (green), across 17 US cities. The scenario frequency distribution approximately matches before and after verification. routine in dense or evasive driving; (b) drivable-area and driving-direction checks penalize crossing into a closed or oncoming lane, even when cones, a flagger, or a contraflow channel make it the only legal path (when verifying with nuPlan [9], we use the proposed channel-insensitive mode for this reason); (c) a static map does not represent temporary work zones, so a correct trajectory may appear to leave the road entirely; and (d) near-misses and recoveries from another agent’s mistake score poorly (in nuPlan parlance, a low TTC score) precisely because the metrics reward uneventful driving, even though these safety-critical events are exactly what we most want to recover and study. A score-based verifier thus provides an essential but conservative estimate of data quality. Publicly available infrastructure as external reference. Non-reactive log-replay verification is one instance of a broader design principle in Dash2Sim: using publicly available, independently maintained sources as references. Street imagery resolves the scale ambiguity in monocular video and geo-references the trajectory. OpenStreetMap supplies the map and simultaneously serves as the independent verification reference. The simulator’s log-replay score translates implicit geometric and physical constraints into a computable quality signal without annotations. Thus, opportunistically leveraging public infrastructure can support physically-grounded simulation at scale.
10
B
Extended Limitations and Future Work
Scope of the Work. We scope our work to showing that dashcam videos are a useful data source for autonomous driving perception and planning, and to providing a framework that recovers 4D logs from such videos at scale. Numerous adjacent opportunities remain, such as building a truly closed-loop sensor simulation from dashcam videos, which we leave as future work, having shown the potential of the Dash2Sim framework and the corresponding ROADWork4D benchmark for such tasks. We hope our work encourages the broader community to make progress on the difficult, ongoing challenge of long-tailed driving. For privileged closed-loop evaluation, we acknowledge many limitations and note that the ROADWork4D-CL benchmark can be extended in several useful ways for better planning evaluation in long-tailed situations. We note a few such limitations: (a) We do not recover traffic light or signaling information, which would be useful for more realistic simulation. (b) Considered planners do not exploit the “open-set” nature of our data which may yield additional benefits, i.e. we map all the object categories to existing nuPlan categories [9] for compatibility with existing planners. (c) The dynamics of the recovered agent tracks could be improved beyond an EKF [68] with a constantvelocity motion model, which rarely holds in the real world (See §D.4). (d) We use IDM-based reactive agents provided by the simulator. These agents follow road-centerline paths and respond to the ego vehicle with car-following dynamics [59]. They do not capture complex interactive behaviors such as yielding negotiations, aggressive merges, or flagger compliance. Incorporating learned traffic models [62] would further improve fidelity of closed-loop evaluation. Dashcam video captures a single viewpoint. Fleet vehicles carry surround-view LiDAR and multiple cameras. A single dashcam observes a forward-facing view and cannot match that density. This sensing modality leaves many regions unobserved, introducing an implicit sim-to-real gap even for perfectly reconstructed logs. Sensor simulation methods can potentially synthesize plausible observations for unobserved viewpoints (See Fig. C.4 for some initial results). Dash2Sim provides the 4D driving logs that would act as conditioning for such methods. Noise in ROADWork4D. Our framework, Dash2Sim, composes monocular reconstruction, detection, tracking, and 3D-object lifting. Each stage contributes to downstream error. We do our best to quantify the error of every component and discuss mitigation strategies wherever possible (See §D). A few limitations are beyond our scope, for example, extending Dash2Sim beyond the daytime, goodweather captures in the ROADWork [17] dashcam dataset, since bad weather is a known failure mode for many of the underlying methods in the Dash2Sim framework. We also expect that advances in 3D reconstruction [60, 61, 69, 70] and vision foundation models [47] will directly improve recovered log quality without architectural changes to the framework. Some error sources, notably mismatched or incomplete map topology (See Fig. D.3), lie outside Dash2Sim and are difficult to automate. Scale and Scope of ROADWork4D. ROADWork4D is relatively modest in scale compared to other robotics datasets [26] spanning millions of videos. We also focus our demonstrations of Dash2Sim on a broad class of long-tailed situations, work zones. Navigating work zones is a persistent hurdle for commercial operators (Fig. E.2), and we believe this focus is a practical way to study a problem that commercial autonomous vehicles currently face. While ROADWork4D is an order of magnitude larger in some aspects than existing long-tailed driving benchmarks [6, 11, 12] and thus broadly useful, we expect that further scaling with dashcams would offer additional insights and research directions for the field. We hope that our work shows that in-the-wild data sources can complement expensive curated datasets. Correlation with real-world deployment. Our evaluation shows that planners exhibit measurable performance differences on ROADWork4D scenarios relative to fleet-derived benchmarks, but whether these differences predict on-road deployment outcomes remains an open problem. Validation frameworks that augment real-world testing with simulation [63] are necessary for bridging this gap.
11
Table C.1: Effect of channel sensitivity (CLS-NR). Table C.2: Effect of Static objects (CLSChannel-sensitive evaluation penalizes drivable-area NR). Scores are low even without traffic, deviations from the GT route. Restated to aid the which shows that construction zone layouts discussion here. are the primary generalization challenge. Channel-Sensitive Planner
Score↑
DAC Fail%↓
With Dynamics
Channel-Insensitive Score↑
DAC Fail%↓
Learned PlanTF [58] Diffusion Planner [40] Pluto [24]
17.7 24.0 32.7
18.0 30.9 23.4
22.9 30.9 43.0
5.3 9.2 4.3
Rule-Based PDM-Closed
53.0
27.3
57.4
9.4
Hybrid Pluto [24]
57.5
17.8
68.8
1.2
C
Additional Evaluation and Analysis
C.1
Privileged Closed-Loop Planning
Objects Only
Planner
Score↑
Coll.%↓
Score↑
Coll.%↓
Learned PlanTF [58] Diffusion Planner [40] Pluto [24]
17.7 24.0 32.7
4.3 51.8 10.5
18.3 32.2 31.0
4.9 35.7 9.6
Rule-Based PDM-Closed
53.0
2.2
56.4
0.6
Hybrid Pluto [24]
57.5
4.7
62.2
0.8
Evaluation protocol. We use the nuPlan simulator [9] with the same score metric NC 1.0 used in nuPlan evaluation [25]: no-collision 0.8 (NC), drivable-area compliance (DAC), time0.6 to-collision (TTC), comfort, ego progress (EP), 0.4 PlanTF EP DAC Diffusion Planner driving direction, and speed limit compliance 0.2 Pluto (Learned) and the metric aggregation mechanism. Each PDM-Closed scenario uses a 15 s simulation window, matchPluto (Hybrid) ing nuPlan’s closed-loop horizon. Our scenes are 30 s long, so one could create additional scenarios by sliding the temporal window, an exComf. TTC ploration we leave to future work. Traffic light status is absent from our logs, and we map all Figure C.1: CLS-R Breakdown. Results in the reobjects and agents to appropriate nuPlan cate- active case closely match those in the non-reactive gories for compatibility. The routable map is case. derived from OpenStreetMap [53] lane centerlines, which have known limitations discussed earlier in §B and later in §D.2. Channel Sensitivity While Computing Drivable Areas. We make one change in the definition of drivable area to account for channel sensitivity required in work zones. In normal driving, all mapped lane segments in the route are drivable, so the drivable area is defined as the union of all lane segments. In work zones, we consider only the lane segments close to the ground-truth route to be drivable. This is because in many work zones, the entire point of the temporary traffic control objects (such as cones, barricades, vertical panels, etc.) is to direct traffic into specific channels, which may or may not align with the lanes. The drivable area is thus only those lane segments and their corresponding lane polygons. This helps us understand whether planners can perform the lane changes in work zone scenarios, rather than simply following any valid lane to the goal point. In the main paper, we report channel-sensitive scores and base our discussion on them, since they are more relevant for work zones, but we also report channel-insensitive scores to isolate the effect of channel sensitivity on planner performance. Role of Static Work Zone Objects. We consider an additional simulation mode, where we analyze the effect of static work-zone objects by removing all dynamic agents from the simulation, leaving only the ego trajectory, static work-zone objects, and the map. This ablation primarily isolates the layout axis of the long-tail challenges work zones present (§E.2). The results show that work-zone layouts and channel sensitivity, not dynamic agents, drive the difficulty: removing dynamic agents entirely yields only modest improvements in scores and collision rates (Table C.2). Rule-based 12
Closed-loop score (CLS-NR)
0.25 Diffusion Planner (zero-shot) 0.24 0.23 0.22 0.211 0.208 0.21 0.20 0.190 0.19 0.18 0.17 PlanTF (zero-shot)
1
5
10
Number of training cities
NC
0.238
1.0 0.8 0.6 0.4
EP
Comf.
17
DAC
0.2
PlanTF (zero-shot) Diffusion Planner (zero-shot) PlanTF (trained, 17 cities)
TTC
Figure C.2: Motion planning is learnable from ROADWork4D. We train PlanTF [58] on city-wise subsets of ROADWork4D-CL and evaluate every model on the test set containing scenarios from all the cities. (Left) Closed-loop score increases with the number of training cities, with even data from one city improving performance over a zero-shot model and nearly reaching zero-shot Diffusion Planner [40] when trained on all cities. (Right) Zero-shot PlanTF (red) is conservative, with high nocollision (NC), drivable-area compliance (DAC), and time-to-collision (TTC) but little ego progress (EP), which is ultimately the point of driving. Training on all 17 cities (green) significantly improves ego progress and comfort, reaching the zero-shot Diffusion Planner [40] (blue). PDM-Closed [25] has dramatically fewer collisions but struggles to make progress, consistent with its known inability to change lanes [25]. Motion Planning is Learnable from ROADWork4D dataset. Training data from even one city improves closed-loop performance on the full 17-city test set, which shows that planners can learn planning behavior from 4D logs recovered from dashcams. We train PlanTF [58] on subsets of city logs of ROADWork4D, from starting from a subset that contains logs from single city up to the one that contains all of them. We evaluate every trained model on the same 17-city test set as earlier. Zero-shot PlanTF barely advances along the route, with very low ego progress (30), which lets the other metrics, such as drivable-area compliance, score high at the expense of actually making progress (Fig. C.2). Training on a single city considerably increases ego progress at the expense of drivablearea compliance, and adding data from more cities then improves drivable-area compliance, reducing drivable-area failures by 16.2%. The score matches zero-shot Diffusion Planner [40] (Fig. C.2), which has shown much better performance on existing long-tailed fleet-derived planning datasets [40] when compared to PlanTF [58].
13
Boston
Charlotte
Chicago
Columbus
Denver
Detroit
New York
Philadelphia
San Antonio
Indianapolis
Los Angeles
Minneapolis
Figure C.3: Dash2Sim recovers a geographically diverse dataset of work-zone driving logs. Bird’s-eye views of logs reconstructed from dashcam videos, one each from twelve cities across the U.S. These logs can be used to train and evaluate motion planners in closed-loop simulation.
14
Table C.3: Rendering quality on long-tailed work-zone objects across ROADWork4D scenes. Percentage improvement in parentheses. PSNR↑
SSIM↑
LPIPS↓
n
OmniRe
+ Depth
OmniRe
+ Depth
OmniRe
+ Depth
arrow board barricade barrier cone drum ttc message board ttc sign tubular marker vertical panel
229 942 828 717 718 163 1063 793 381
19.22 20.67 22.09 21.97 21.31 20.25 20.72 20.64 19.00
19.47 21.12 22.45 23.10 21.84 20.76 21.13 21.31 19.84
0.523 0.598 0.632 0.659 0.597 0.515 0.583 0.583 0.524
0.534 0.620 0.645 0.706 0.619 0.547 0.609 0.615 0.567
0.322 0.239 0.254 0.242 0.244 0.292 0.223 0.234 0.226
0.299 (-7%) 0.210 (-12%) 0.224 (-12%) 0.193 (-20%) 0.211 (-14%) 0.259 (-11%) 0.200 (-10%) 0.202 (-14%) 0.196 (-13%)
weighted mean
5834
20.94
21.51
0.595
0.622
0.242
0.211 (-13%)
Class
Table C.4: Full-split reconstruction quality across ROADWork4D scenes (train + test frames). Method OmniRe OmniRe + Depth
C.2
PSNR↑
SSIM↑
LPIPS↓
FID↓
cFID↓
DINO-S↓
DSim↓
26.14 26.53 (+1.5%)
0.851 0.870 (+2.2%)
0.208 0.170 (-18%)
49.3 43.7 (-11%)
7.2 5.9 (-18%)
0.071 0.064 (-10%)
0.069 0.056 (-19%)
Novel-View Synthesis
Overview. Sensor simulation from driving logs follows two approaches: generative models that synthesize sensor streams [71, 72, 73], or explicit reconstruction via Gaussian Splatting [36, 8]. In both cases, simulation fidelity depends on how well the conditioning signals are grounded in geometry, and the dense depth from §3 can serve as such a signal. We take the second route as it allows us to evaluate the depth produced by Dash2Sim, isolating its contribution without conflating it with a generative model’s ability to fill missing details. §4.3 showed that depth supervision from Dash2Sim improves held-out novel-view quality on every metric. This section provides the full-split and per-class breakdown. Details of our setup. We select 100 geographically diverse videos from ROADWork4D for our evaluation. Following OmniRe [37], we report PSNR and SSIM. These metrics average over the entire frame and are therefore insensitive to long-tailed objects that occupy only a small fraction of the image. Moreover, they penalize low-level misalignment that minimally affects the perceptual quality of the rendered scene [74, 75]. We additionally report LPIPS [74], FID [76], CLIP-FID [77] (cFID), DINO-Struct [78] and DreamSim [79] (DSim), which are increasingly adopted in driving novel view synthesis evaluations [80, 81, 8, 35] and capture higher-level visual fidelity. Full-split reconstruction quality. Table C.4 reports metrics on the full split. Depth supervision improves all metrics, including on training views where the photometric loss already provides direct pixel-level supervision: LPIPS improves by 18% and DreamSim by 19%. This shows that the depth regularizes the Gaussian representation itself. As splat positions and shapes are better constrained by the available depth, the representation generalizes to both seen and unseen views. Monocular video is the most difficult setting for OmniRe [37], which is primarily evaluated on multi-modal fleet data. This partially explains the drop in absolute scores from full split to held-out split. While a direct comparison is difficult due to differences in data source and sensor setup, the PSNR from monocular logs is in a comparable range to recent fleet-based sensor simulation [36, 8, 37], suggesting that dashcam-derived driving logs may become sufficient in the future for realistic sensor simulation with improvements in rendering. 15
Yaw = -20
Yaw = 0
Yaw = -20
Philadelphia
T = 3 seconds
Philadelphia
T = 5 seconds
Boston
Houston
Figure C.4: Potentially Simulating Other Sensors. Qualitative results of rendering additional camera viewpoints (yaw of ±20◦ ) from a single monocular dashcam video, conditioned on the dense depth recovered by Dash2Sim. A forward-facing dashcam leaves much of the scene unobserved; these renderings suggest that this depth could condition closed-loop sensor simulation beyond the original viewpoint. Rendering of small long-tail objects. Table C.3 evaluates rendering quality on 9 long-tailed workzone object categories across 5,834 crop-level instances in the held-out test split. These objects span only a few pixels in the dashcam image (§3) and the ego vehicle observes each one from a narrow range of directions over a short temporal window. LPIPS [74] improves for every long-tailed category, with the largest improvements on thin, vertical objects like cones (20%), tubular markers (14%), and vertical panels (13%). This shows that the dense depth from §3 improves object-level rendering of small work-zone objects, not just the overall scene. Potentially Rendering Multiple Sensor Views. We present some initial qualitative results of rendering multiple sensor streams from monocular videos in Fig. C.4.
16
Semantics
Geometry Monocular Dashcam Video
Promptable Tracker
Simulation
Long-Tailed Detector
Phoenix
Co-Located Street Imagery
2D Detection and Tracking
Geo-anchored SfM
Sparse Recon.
2D Object Tracks
Monocular Depth Prior
Ego, Objects, Agents on Open Street Map Geometry Aware Depth Refinement
Metric Reconstruction
Depth Conditioned 3D Lifting
3D Tracked Objects
4D Driving Log
Figure D.1: Dash2Sim Framework. Dash2Sim produces a metric, geo-referenced 4D driving log from in-the-wild monocular dashcam video in three stages, shown left to right. Geometry recovers a metric ego trajectory and dense per-image depth. Semantics detects, tracks, and lifts agents and long-tailed objects to 3D using promptable foundation models [47, 48]. For Simulation, we retrieve the OpenStreetMap [53] tile and place the ego, agents, and objects.
D
Dash2Sim Framework: Description and Validation
D.1
Description of the Geometry Framework
Several issues make metric reconstruction from dashcam video challenging. Monocular localization is scale-ambiguous, as the recovered trajectory and scene are determined only up to a global similarity transform [16]. The dashcam’s GPS, if available, routinely places the ego vehicle in adjacent city blocks instead of the driven lane, and most in-the-wild video lacks GPS entirely. Long-tailed objects (e.g., cones, tubular markers) also span only a few pixels in the dashcam image, so feature triangulation produces too few 3D points on them to support downstream 3D lifting and novel view synthesis, potentially, for closed-loop sensor simulation. Dash2Sim’s geometric framework (Fig. D.1, left) addresses all three. Sparse reconstruction. We compute image-level descriptors with EigenPlaces [82] and retrieve nearest neighbors with FAISS [83] to build the scene graph, then extract local features with SuperPoint [84] and match image pairs with LightGlue [85]. We run Global SfM [60, 86] jointly over the dashcam and Street View streams. The resulting reconstruction is aligned to global coordinates via a similarity transform fit to the Street View poses. We discard the noisy dashcam GPS (Fig. 4, red) during this alignment to avoid incorporating its biases. We fit a ground plane using a road-surface mask from a semantic segmentation model [87] for downstream simulator compatibility. Geometry-aware depth refinement. The joint SfM above is sparse and produces few 3D points. As long-tailed objects are small, sparse depth is insufficient for accurate localization [57] and view synthesis. We refine per-image depth with MP-SfM [88], which fuses a monocular depth-and-normal prior [89] with multi-view geometric constraints. We run this stage on the dashcam frames alone, excluding co-located street imagery since the two captures occurred at different times and disagree on dynamic and transient scene content. The output is per-image, metric, geometrically consistent depth. Feedforward depth estimators alone [89, 69, 70] did not produce sufficient accuracy in our experiments, but might improve performance when combined with the sparse reconstruction step [61], which we leave to future work. 17
Boston
Boston
San Francisco
Charlotte
Figure D.2: More Driving logs from ROADWork4D recovered by Dash2Sim. Our logs include long-tailed work zone scenarios with rare objects and dynamic agents. Our supplemental website and walkthrough video contain additional media and visualizations.
D.2
Street Imagery Anchoring Yields Accurate Metric Reconstruction
Overview. Street imagery anchors sparse reconstruction, recovers metric scale, and geo-references camera pose from monocular videos (Figure 4). We evaluate pose accuracy on nuScenes-mini [2] (daytime scenes), which provides dense per-frame ground-truth pose. Dashcam corpora lack such ground truth, so we use nuScenes as our proxy for evaluation. Setup. For each scene, Dash2Sim retrieves nearby Google Street View panoramas using the egovehicle coarse GPS, projects them into perspective crops, and reconstructs a sparse model anchored to the panoramas’ georeferenced coordinates. We evaluate the resulting dashcam camera poses against nuScenes ego-pose ground truth following the standard SLAM or odometry benchmark protocol [90]: we estimate the Sim(3) transform that best aligns the predicted trajectory to the ground truth via Umeyama alignment [91] and report the recovered scale ratio, absolute trajectory error (ATE) after alignment, and relative pose error (RPE). Findings. Table D.1 reports median results across nuScenes-mini. Anchoring to Street View imagery achieves sub-decimeter ATE under Sim(3) alignment, with every video sequence successfully registered. In several scenes, Dash2Sim achieves a median translation error within 10 cm and rotation RPE well below one degree. Moreover, the recovered scale ratio is within 0.3 % of unity, showing that panorama anchors are sufficient to recover absolute metric scale from monocular video without IMU, stereo cameras, or LiDAR. Note that several nuScenes-mini sequences are captured in Singapore, suggesting that Dash2Sim can generalize given co-located street imagery. 18
Condition
Recon.
Scale
ATE (m)
RPEt (m)
RPEr (°)
Street View anchored (Ours)
100%
1.0024
0.084
0.107
0.10
Coarse GPS anchored No Street View
100% 43%
1.6102 1.5708∗
0.084 0.070∗
4.010 3.311∗
0.10 0.48∗
Table D.1: Sparse Reconstruction via Dash2Sim. ATE and RPEt are translation RMSE in meters; RPEr is rotation RMSE in degrees. All metrics are computed after Sim(3) alignment of the predicted trajectory to ego-pose ground truth. Street View anchored is the Dash2Sim’s sparse reconstruction framework. Coarse GPS anchored replaces panorama alignment anchors with simulated consumergrade GPS while keeping Street View imagery in the reconstruction. No Street View removes Street View imagery and anchoring entirely and SfM fails on more than half the scenes due to insufficient horizontal parallax in dashcam-only frames. ∗ Median over reconstructed scenes only. San Francisco
Chicago
Figure D.3: Examples of Mismatch and Failures with SD Maps. Cases where the video and the SD map disagree. Top: A lane-count mismatch: OSM shows 2 lanes, but the video shows 4 lanes with construction on some of them (the map may predate a transition of that street to a narrower configuration). Since OSM is crowd-sourced and updated nightly, it is difficult to retrieve a map that matches the state of the road at the time of video capture. Bottom: A missing slip lane: OSM models the junction as two centerlines meeting at a node, omitting the curved connector and its traffic island, so trajectories following the actual curb appear to leave the drivable area.
We run two ablations to decompose the contribution of co-located street imagery. Coarse GPS anchored keeps Street View imagery as part of the SfM but replaces the panorama alignment anchors with simulated consumer-grade GPS positions along the ego trajectory, perturbed with iid Gaussian noise (σh =10 m, σv =15 m) modeling multipath-dominated error in urban areas. RPE rises from 0.11 m to 4.01 m and the recovered scale degrades from 1.002 to 1.610, while rotation RPE and Sim(3)-aligned ATE are unchanged as expected. No Street View removes Street View imagery entirely, running SfM on monocular frames alone with the same coarse GPS anchors. Reconstruction fails on more than half the scenes due to insufficient horizontal parallax in forward-facing monocular-only frames. The surviving scenes show comparable localization degradation to the GPS-only ablation. Co-located Street imagery thus contributes in two distinct ways: its geo-referenced coordinates provide reliable anchoring for metric-scale recovery, and its cross-view diversity enables SfM convergence on scenes where monocular video parallax alone is insufficient. 19
Limitations in Reconstruction and Map Sources. Reporting trajectory error after local rigidsimilarity alignment is the standard benchmarking choice [1, 90, 92], but these metrics do not capture many different types of errors and limitations, for example, in global coordinate alignment across different sources. Recent work on city-scale visual localization [93] confirms this limitation2 , showing that consumer-grade GNSS and associated products (Google Street View, Virtual 3D Assets in Google Earth, etc.) have minor errors in dense urban environments due to many potential factors [94, 95, 96]. Our anchors derive from Street View panorama coordinates, which are not survey-grade. We measured the residual absolute-frame errors between our reconstructions and nuScenes’ map-anchored ground truth at 3 to 9 meters horizontally, among other sources of error. The nuScenes [2] ground truth itself might carry similar errors, but we could not confirm this. Our map source itself (OpenStreetMap) has further limitations, owing to its Standard Definition (SD) quality compared to the true HD maps provided with fleet-collected benchmarks. For example, being crowd-sourced, OpenStreetMap does not always provide lane width, nor does it provide slip-lane or curb information at turns (i.e., many streets appear to intersect as road lines, with no drivable area around curbs), which is a problem for tight turns. We partly mitigate this during closed-loop evaluation by being more lenient when computing what constitutes a drivable area violation, but this is an unavoidable problem with SD Maps. Moreover, prior work [54, 55] has shown that studying autonomous driving with these limitations of SD Maps is still valuable. We partially address some of these issues via non-reactive log-replay simulation (§3), which filters out 4D driving logs where these issues cause at-fault events. Many modern approaches sidestep the map entirely: they either operate in a local ENU frame [6, 46] or build fully synthetic simulations [55] from real logs. Addressing these issues would improve the yield for privileged closed-loop planning [9], but we leave them out of scope of this work. D.3
Visual Place Recognition as an Alternative to GPS-Based Retrieval
Overview. Dash2Sim uses coarse GPS to retrieve nearby street view imagery, an assumption that holds for most dashcam data. To test whether GPS can be removed entirely, we evaluate visual place recognition as an alternative. Setup. We take the 227 annotated San Francisco images in ROADWork [17] dataset as queries and match them with MegaLoc [52] embeddings against ≈106k Mapillary [97] street-level references covering ≈20 km² of the city. Findings. Table D.2 shows that 60% of queries retrieve a reference within 50 m of ground truth at top-1, and that, conditioned on coverage, MegaLoc [52] reaches 76% at top-1 and 85% at top-5, well within what Dash2Sim needs as a spatial location prior. Most of the unconditional error therefore comes from sparse Mapillary coverage rather than the retrieval model. Note that R@∞ is a purely spatial bound: a reference within d meters need not visually match the query, since viewing direction, time of day, weather, and occlusions can all differ between dashcam captures and crowdsourced Mapillary images; thus, achievable recall is likely below R@∞. For video queries, redundancy across nearby frames opens an orthogonal opportunity to filter and reweight street view retrievals across the sequence [98, 99]. Figure D.4 shows a handful of retrievals. MegaLoc [52] retrieves nearby references despite substantial appearance variation between query and reference captures: different traffic, lighting, time of year, occlusions, and presence of long-tail objects such as work vehicles or blocked-off streets. We present this simple method as proof-of-concept for applying VPR at scale. We use Mapillary because it is openly queryable and allows retrieving city-scale imagery. Street View coverage is denser and more uniform, so we expect comparable or better retrieval. For further improvements in accuracy and size of the reference set, hybrid methods are applicable [100]. 2 https://www.youtube.com/watch?v=gLoMiIcCiCk&t=1098s; Quote: We’re essentially within a centimeter based on those ground control points... But Google Earth isn’t [as accurate]. Google Earth is actually several meters off in many places.
20
Query (ROADWork) Top-1 retrieval
San Francisco
1.8 m
2.5 m
2.0 m
Figure D.4: Street View Imagery recovery via Visual Place Recognition. Top-1 retrievals from a Mapillary [97] reference database for ROADWork [17] queries, with haversine distance between the query’s ground-truth GPS and the retrieved image’s GPS. Despite significant differences, retrieved images localize within a few meters of ground truth given coverage. Table D.2: Street View Imagery recovery via Visual Place Recognition. 227 San Francisco ROADWork [17] queries against ∼106k Mapillary [97] references (≈20 km²), retrieved with MegaLoc [52]. R@k is the fraction of queries with a top-k retrieval within d meters of ground truth; R@∞ is the coverage ceiling (any reference within d). Bracketed percentage scores are R@k/R@∞, i.e. the share of geometrically solvable queries correctly retrieved. d = 25 m d = 50 m d = 100 m R@1 R@5 Coverage (R@∞)
D.4
0.44 (65%) 0.50 (74%)
0.60 (76%) 0.67 (85%)
0.71 (77%) 0.76 (83%)
0.68
0.79
0.92
Description of the Semantics Framework
Promptable detection and tracking. Our observation is that the bottleneck for detecting long-tailed classes is generalization to the “long-tailed” text prompt, not the object tracker. SAM3 [47] accepts text and mask prompts, so we prompt with text for common objects and with exemplar masks from a domain-specific detector [17] for long-tailed objects. A claim-and-suppress strategy then uses the detector in tandem with the tracker to recover each instance, which we detail next. Iterative Detection and Tracking of Long-Tailed Objects. Segmentation Foundation Models [47] are already comparable to domain-specific detectors when prompted appropriately [17]. For longtailed work-zone categories, prior work [17] shows that SAM models recover 2D segmentations at up to 90% relative accuracy when paired with a domain-specific detector. We adopt a strategy to efficiently prompt SAM3 [47] across all instances, composing an image-level detector [17] in a detect-then-track loop. Tracking-by-detection via IoU-based matching is a common strategy for 2D multi-object tracking [101, 102], which we adapt for our setting. Please see our recovered 2D tracks in Fig. D.5. Algorithm 1 provides the sketch of the algorithm. For each long-tailed class c, we maintain a pool of active image-level detections, select the highest-confidence surviving detection as the exemplar, and prompt SAM3 [47] with its box together with positive and negative point hints sampled from the exemplar’s mask. To help SAM3 [47] disambiguate the exemplar from nearby same-class instances at prompt time, other same-frame detections of class c above a co-prompt confidence threshold τco are additionally passed as auxiliary prompts. SAM3 [47] propagates a tubelet across the video, and every detection whose 2D box overlaps the tubelet at any frame by IoU greater than τiou is considered “explained” and retired from the pool, together with the exemplar and the auxiliary detections that contributed to its prompt. The loop terminates when no detection remains above the exemplar 21
Boston
Charlotte
Chicago
Denver
Los Angeles
San Antonio
San Francisco
Figure D.5: 2D Object Tracks recovered by Dash2Sim. We show examples of 2D object tracks. Our approach accurately recovers both common and long-tailed objects (examples include “Arrow Board”, “Vertical Panels”, “Cones”, “TTC Sign Board”) with consistent IDs.
threshold τex or a per-class iteration budget Kc is reached. In practice, a small Kc suffices because a single template typically recovers most instances of its class within a scene. Depth-conditioned 4D lifting. We observe that the per-image depth recovered in §3 can act as the conditioning signal an open-vocabulary 3D detector needs for lifting small long-tailed objects. We prompt WildDet3D [48] with 2D tracks and condition it on this depth. Following prior work that uses novel-view synthesis as a proxy for depth quality [51, 103], we evaluate the recovered depth through its effect on rendering quality in §4.3. We validate 3D localization of long-tailed objects in §D.5. Please see our recovered 3D tracks in Fig. D.6 3D Track Smoothing and Class-Prior Rejection. We smooth each lifted 3D track with an Extended Kalman Filter [68] under a constant-velocity motion model to reduce per-frame jitter in WildDet3D [48]’s center and yaw estimates. We then reject detections that violate class-prior heuristics, namely predicted 3D dimensions outside class-specific bounds (e.g., a person is roughly 1.7 meters tall) and tracks that are too far from the fitted ground plane (§D.1). 22
Boston
Charlotte
Chicago
Denver
Indianapolis
New York City
San Francisco
Figure D.6: 3D Object Tracks recovered by Dash2Sim. We show examples of 3D object tracks on ROADWork4D logs. Our approach recovers common and long-tailed objects when given depth supervision (examples include “Arrow Board”, “Vertical Panels”, “Cones”, “TTC Sign Board”) with consistent IDs.
D.5
Evaluation of Lifting Long-Tailed Objects
Overview. WildDet3D [48] is already state-of-the-art when prompted on common driving objects [3]. The remaining question is whether depth conditioning improves 3D lifting of these objects, which span only a few pixels when considering work zone objects. ROADWork [17] provides only 2D annotations, so 3D detection cannot be quantified on it. We evaluate on WorkZone3D [57] instead, which provides ∼20,000 images with LiDAR and 3D ground truth over a similar set of long-tailed categories. Setup. To isolate 3D localization accuracy, we follow a prompt-conditioned protocol that matches the lifting procedure described earlier. For every class present in a frame, a frozen 2D detector [57] produces an instance mask from which we sample point prompts and query WildDet3D [48] in its geometric mode to return exactly one 3D box per prompt. We compare two regimes: Img (image only) and Img+LiDAR, where sparse LiDAR points serve as an upper-bound proxy for a geometrically derived depth prior. The depth recovered in §3 is dense but noisy, whereas LiDAR is sparser but 23
Algorithm 1 Iterative Exemplar Mining for Long-Tailed Object Tracking Require: Video V ; class-agnostic video propagator P; image-level detections D = {di } with class ci , score si , box bi , mask mi , frame fi ; long-tailed class set C; thresholds τex , τco , τiou ; per-class iteration budget Kc . Ensure: Set of object tracks T . 1: T ← ∅ 2: for each class c ∈ C do 3: Ac ← { d ∈ D | ci = c } ▷ active detections of class c 4: Tc ← ∅ 5: for k = 1, . . . , Kc do 6: e ← arg maxd∈Ac sd ▷ pick highest-confidence exemplar 7: if e = ∅ or se < τex then 8: break 9: end if 10: O ← { d ∈ Ac \ {e} | fd = fe , sd ≥ τco } ▷ co-class detections in same frame 11: pe ← S AMPLE P OINTS(me ) ▷ positive/negative point hints from exemplar mask 12: Q ← {(be , pe )} ∪ {(bd , S AMPLE P OINTS(md ))}d∈O ▷ multi-instance prompt set 13: X ← P.P ROPAGATE(V, fe , Q) ▷ tubelet X = {(f, id, m̂f,id )} 14: if X = ∅ then 15: Ac ← Ac \ ({e} ∪ O) ▷ retire exemplar; do not retry 16: continue 17: end if 18: for each (f, id, m̂) ∈ X do ▷ suppress detections this tubelet explains 19: for each d ∈ Ac with fd = f do 20: if I O U(bd , B BOX(m̂)) > τiou then 21: Ac ← Ac \ {d} 22: end if 23: end for 24: end for 25: Tc ← Tc ∪ R EID(X ) ▷ shift IDs for global uniqueness 26: Ac ← Ac \ ({e} ∪ O) 27: end for 28: T ← T ∪ Tc 29: end for 30: return T
Table D.3: Evaluating 3D-detection on long-tailed categories. mAP is averaged over centerdistance thresholds {0.5, 1, 2, 4} m. Depth conditioning gives a consistent improvement. AP-11
AP-40
Class
#GT
Img
Img+LiDAR
Img
Img+LiDAR
Barrel Channelizer Cone
1,735 5,544 1,113
0.325 0.171 0.129
0.325 0.180 0.130
0.293 0.127 0.124
0.294 0.127 0.126
Mean
—
0.208
0.212
0.181
0.182
accurate, so LiDAR bounds the benefit that geometric depth conditioning can provide. We report AP@11 and AP@40 averaged over center-distance thresholds {0.5, 1, 2, 4} m. Findings. Table D.3 shows that adding a depth prior gives a consistent improvement across all three long-tailed categories on both AP@11 and AP@40, with the largest gain on Channelizer (AP@11: 0.171 → 0.180). This validates the design choice of conditioning on a depth prior when lifting long-tailed objects in Dash2Sim. Qualitative results in Figure D.7 show that WildDet3D [48] can produce well-localized boxes for the long-tailed categories when provided with depth conditioning. 24
Pittsburgh
Figure D.7: Qualitative 3D Lifting Results on WorkZone3D [57]. Predicted 3D boxes from WildDet3D [48] for long-tailed work-zone categories (cones, barrels, channelizers). Boxes are well-localized when a depth prior is available. Image-only prompting produces more misses and depth errors on small distant instances. D.6
4D Driving Logs to Driving Simulation: Implementation Details
Closed-loop planning needs three components on top of a 4D driving log: a simulator interface, a routable map, and a reactive-agent model. The nuPlan [9] privileged-planning simulator accepts metric, map-aligned logs, provides rule-based IDM reactive agents, and lets existing planners run on ROADWork4D-CL. For the routable map, the nuPlan dataset assumes HD maps collected by the same fleet that captured the logs, which we do not have. We retrieve an OpenStreetMap [53] tile using the geo-referenced ego trajectory, following prior work [54, 55] that used OpenStreetMap lane graphs in place of HD maps. OpenStreetMap is maintained independently, so the map provides an external reference that the recovered log must be consistent with.
25
SLOW
STOP
SLOW
STOP
Directing Traffic Shoveling
Figure E.1: Work zones involve long-tailed objects, layouts and behaviors. Examples along the three long-tail axes present in work zones. (Left) Rare objects and entities: cones, barrels, arrow boards, construction vehicles, and a wide vocabulary of temporary signage (e.g., Road Closed, Burger King Closed During Construction, Delivery Vehicles Only). (Middle) Rare layouts: temporary lane closures and detours that contradict the map, forcing the ego vehicle off expected lanes. (Right) Rare behaviors: human flaggers holding Stop/Slow paddles, construction workers shoveling in active traffic, and traffic officials directing traffic with gestures. Left and Right Images courtesy ROADWork [17] dataset.
E
ROADWork4D Benchmark Details
E.1
Long-tail Driving
Long-tail driving refers to scenarios that occur rarely but determine system safety and deployment readiness. Within a fixed operational design domain (ODD), the scenario distribution is heavytailed: each additional fleet hour draws predominantly from common scenarios, making rare events increasingly expensive to collect. Rare cases span multiple axes that often co-occur [17]: rare objects (debris, animals), rare layouts (temporary lane closures, washed-out lane markings), rare behaviors (vehicles driving on the wrong side of the road, “illegal”-but-correct maneuvers), and rare environmental conditions (snow, fog, rain, glare). Prior work generally targets each axis in isolation, for instance through dedicated adverse-weather collection [104, 105]. E.2
Work zones as a representative long-tail autonomous driving testbed
Work zones as a category span all three long-tail axes. A typical work zone may contain rare objects (cones, barrels, arrow boards, and a vocabulary of temporary signage defined by federal standards [106]), rare layouts (lane closures that contradict the static map, contraflow setups, contradictory markings), and rare behaviors (flaggers holding stop and slow paddles, workers in active traffic, officers directing traffic with gestures). Fig. E.1 shows examples along each axis. Work zones are a significant source of traffic fatalities: NHTSA reports 898 fatalities in U.S. construction or maintenance zones in 2023 alone [108]. They also remain a persistent challenge for commercial self-driving systems. Fig. E.2 compiles publicly documented failures from commercial autonomous vehicles between 2022 and 2026, ranging from collisions and merges into oncoming traffic to regulatory action following fatalities [107] and operational halts as recently as 2026 [23]. These failures span multiple commercial platforms and multiple years. Work zones are also tractable to study. They occur frequently in routine driving, appear in disengagement and crash filings as a measurable category, and the ROADWork [17] dashcam corpus provides large-scale 2D annotations for work-zone perception, complementing ROADWork4D. Overall, ROADWork4D offers an additional setting for long-tail planning that complements existing planning benchmarks. E.3
Dashcams as a complementary data source
Fleet data covers the head of the scenario distribution well. Instrumented vehicles carry high-quality sensors but operate in a small number of cities, and each additional fleet hour is increasingly unlikely 26
Cruise
Cruise
Waymo
Waymo
2022
2023
2024
Waymo
Tesla
Xiaomi
2024
2024
2025
2023 Waymo
2025
Waymo
Waymo
Tesla
Waymo
2025
2026
2026
2026
Figure E.2: Work Zones are a persistent hurdle. A few work zone failures by commercial autonomous driving and driver-assisted vehicles, drawn from publicly available information (news reports, social media videos, etc.) between 2022 and 2026. Work zones are challenging, and contrary to prevailing expectations, remain a persistent hurdle for reliable self-driving. Failure modes include colliding with emergency vehicles, driving into wet concrete, not stopping for human flaggers, getting boxed in by construction equipment, merging into oncoming traffic, and fatal accidents that prompted regulatory action [107]. These recent incidents [23] suggest that solving autonomous driving in work zones is a current, practical, and critical challenge. Do note the visualization is not exhaustive or necessarily representative of the abilities of the commercial systems, as some commercial operators have a much larger fleet compared to others and thus a larger error surface area.
to encounter a novel rare event. Curating long-tail subsets from fleet logs [6, 30] helps, but can only surface what the fleet already drove through. Synthetic simulators [10, 11, 12] can construct rare scenarios by hand, but the range of rare objects, layouts, and behaviors they produce is bounded by the scenario designer’s creativity, and both visual and behavioral sim-to-real gaps remain. Dashcam video covers a different part of the distribution. Consumer dashcams are owned by an estimated 30% of U.S. drivers [14], span far more cities and road conditions than any single fleet, and civilian uploads tend to over-represent unusual driving events. The cost is noisier scene reconstruction: monocular video lacks the multi-modal sensing of fleet vehicles. Dash2Sim works within this constraint to produce metric 4D logs that complement fleet-derived benchmarks for studying the long tail of autonomous driving. Fig. A.2 shows the geographic distribution of ROADWork4D scenarios. ROADWork4D spans 17 US cities, exceeding existing fleet-derived benchmarks (nuPlan covers 4 cities, nuScenes 2, Argoverse 6). Many of these cities, including Chicago, Columbus, Indianapolis, Jacksonville, and Charlotte, are not represented in any existing planning benchmark and do not yet have a commercial autonomous ride-hailing service collecting fleet data. This geographic breadth matters because, although U.S. regulations [106] establish national minimum standards for temporary traffic control, work zone regulations are implemented and supplemented at the state and local level [109, 110], inducing domain gaps in both perception and planning for autonomous driving [17]. The non-reactive log-replay verified subset approximately follows the distribution of the full set, and all 4,244 scenarios remain available for open-loop evaluation and end-to-end training. §E.4 compares ROADWork4D against existing benchmarks. E.4
Comparison with existing datasets, benchmarks and simulators
Column definitions. Source: data origin (Fleet for calibrated autonomous-vehicle fleets, Synthetic for procedural generation in CARLA or its derivatives, Dashcam for in-the-wild consumer-grade video sources). Long-Tail: whether the benchmark is explicitly curated for rare scenarios. Locales: 27
distinct cities or synthetic towns. Maps: whether routable map data is provided. Hours/Scenes: total recorded driving time and number of evaluation scenes, segments, or routes (unit varies by benchmark and is clarified per row below). Closed-Loop: whether a policy’s actions change future observations, as opposed to scoring a predicted trajectory against a fixed log. Reactive: whether non-ego agents respond to the ego vehicle’s actions rather than replaying logged trajectories. Comments on planning benchmarks. nuScenes [2] provides 1,000 20-second scenes (5.5 hours annotated). Argoverse 2 [4] provides 1,000 sensor sequences and 250,000 motion-forecasting scenarios. Many multi-modal perception datasets can be repurposed as planning benchmarks [54], but no open-source simulator integration exists for them, so we omit them from planning comparisons. WOMD [5] contains ∼104,000 20-second segments. WOD-E2E [6] curates 4,021 long-tail segments at an occurrence frequency less than <0.03%. Navsim [7] builds on OpenScene [111, 112] and nuPlan [9] data, with 12,000 evaluation samples and a short-horizon non-reactive rollout that we count as pseudo closed-loop. Navsim V2 [8] extends this with sensor-grounded pseudo closed-loop evaluation via 3D Gaussian Splatting. nuPlan [9] provides the most complete open-source closedloop reactive simulator, with IDM-based reactive agents for privileged planning evaluation. Learned agent extensions [62] to nuPlan [9] have been proposed but are not yet open-sourced. InterPlan [13] augments nuPlan logs by spawning additional traffic agents and modifying routes to force lane changes. The maps and ego logs are realistic (inherited from nuPlan), but the added interactions are synthetic, so we group it with synthetic benchmarks. The full release contains 335 scenarios. Bench2Drive [11] provides 220 evaluation routes spanning 44 interactive scenarios. Fail2Drive [12] provides 200 paired routes across 17 scenario classes.
28
References [1] A. Geiger, P. Lenz, and R. Urtasun. Are we ready for autonomous driving? the KITTI vision benchmark suite. In CVPR, 2012. [2] H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom. nuScenes: A multimodal dataset for autonomous driving. In CVPR, 2020. [3] M.-F. Chang, J. Lambert, P. Sangkloy, J. Singh, S. Bak, A. Hartnett, D. Wang, P. Carr, S. Lucey, D. Ramanan, and J. Hays. Argoverse: 3D tracking and forecasting with rich maps. In CVPR, 2019. [4] B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, D. Ramanan, P. Carr, and J. Hays. Argoverse 2: Next generation datasets for self-driving perception and forecasting. In NeurIPS Datasets and Benchmarks Track, 2021. [5] S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, and D. Anguelov. Large scale interactive motion forecasting for autonomous driving: The Waymo open motion dataset. In ICCV, 2021. [6] R. Xu, H. Lin, W. Jeon, H. Feng, Y. Zou, L. Sun, J. Gorman, E. Tolstaya, S. Tang, B. White, B. Sapp, M. Tan, J.-J. Hwang, and D. Anguelov. WOD-E2E: Waymo open dataset for end-to-end driving in challenging long-tail scenarios. arXiv preprint arXiv:2510.26125, 2025. [7] D. Dauner, M. Hallgarten, T. Li, X. Weng, Z. Huang, Z. Yang, H. Li, I. Gilitschenski, B. Ivanovic, M. Pavone, A. Geiger, and K. Chitta. NAVSIM: Data-driven non-reactive autonomous vehicle simulation and benchmarking. In NeurIPS Datasets and Benchmarks Track, 2024. [8] W. Cao, M. Hallgarten, T. Li, D. Dauner, X. Gu, C. Wang, Y. Miron, M. Aiello, H. Li, I. Gilitschenski, B. Ivanovic, M. Pavone, A. Geiger, and K. Chitta. Pseudo-simulation for autonomous driving. In CoRL, 2025. [9] H. Caesar, J. Kabzan, K. S. Tan, W. K. Fong, E. M. Wolff, A. H. Lang, L. Fletcher, O. Beijbom, and S. Omari. nuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles. arXiv preprint arXiv:2106.11810, 2021. [10] A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun. CARLA: An open urban driving simulator. In CoRL, 2017. [11] X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan. Bench2Drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving. In NeurIPS Datasets and Benchmarks Track, 2024. [12] S. Gerstenecker, A. Geiger, and K. Renz. Fail2Drive: Benchmarking closed-loop driving generalization. arXiv preprint arXiv:2604.08535, 2026. [13] M. Hallgarten, J. Zapata, M. Stoll, K. Renz, and A. Zell. Can vehicle motion planning generalize to realistic long-tail scenarios? In IROS, 2024. 30 https://www.autoinsurance.com/research/ [14] autoinsurance. dash-cam-usage-report/, 2026. Accessed: 2026-04-29. [15] A. Allshire, H. Choi, J. Zhang, D. McAllister, A. Zhang, C. M. Kim, T. Darrell, P. Abbeel, J. Malik, and A. Kanazawa. Visual imitation enables contextual humanoid control. In CoRL, 2025. [16] R. Hartley and A. Zisserman. Multiple View Geometry in Computer Vision. 2003. 29
[17] A. Ghosh, S. Zheng, R. Tamburo, K. Vuong, J. Alvarez-Padilla, H. Zhu, M. Cardei, N. Dunn, C. Mertz, and S. G. Narasimhan. Roadwork: A dataset and benchmark for learning to recognize, observe, analyze and drive through work zones. In ICCV, 2025. [18] O. Zendel, K. Honauer, M. Murschitz, D. Steininger, and G. F. Dominguez. Wilddash-creating hazard-aware benchmarks. In ECCV, 2018. [19] O. Zendel, M. Schörghuber, B. Rainer, M. Murschitz, and C. Beleznai. Unifying panoptic segmentation for autonomous driving. In CVPR, 2022. [20] S. K. Bashetty, H. B. Amor, and G. Fainekos. Deepcrashtest: Turning dashcam videos into virtual crash tests for automated driving systems. In ICRA, 2020. [21] Y. Miao, G. Fainekos, B. Hoxha, H. Okamoto, D. Prokhorov, and S. Mitra. From dashcam videos to driving simulations: Stress testing automated vehicles against rare events. In AAAI Workshops, 2025. [22] J. Bote. Cruise vehicle gets stuck in wet concrete while driving in San Francisco. https:// www.sfgate.com/tech/article/cruise-stuck-wet-concrete-sf-18297946.php, 2023. Waymo halts freeway rides after robotaxis strug[23] TechCrunch. gle in construction zones. https://techcrunch.com/2026/05/21/ waymo-halts-freeway-rides-after-robotaxis-struggle-in-construction-zones/, 2026. [24] J. Cheng, Y. Chen, and Q. Chen. Pluto: Pushing the limit of imitation learning-based planning for autonomous driving. arXiv preprint arXiv:2404.14327, 2024. [25] D. Dauner, M. Hallgarten, A. Geiger, and K. Chitta. Parting with misconceptions about learning-based vehicle motion planning. In CoRL, 2023. [26] A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In ICRA, 2024. [27] C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, et al. Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158, 2024. [28] S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. NeurIPS, 2023. [29] J.-J. Hwang, R. Xu, H. Lin, W.-C. Hung, J. Ji, K. Choi, D. Huang, T. He, P. Covington, B. Sapp, et al. Emma: End-to-end multimodal model for autonomous driving. arXiv preprint arXiv:2410.23262, 2024. [30] R. Wagner, O. S. Tas, J. Villa, F. Hauser, Y. Shen, M. Steiner, D. Strutz, C. Fernandez, C. Kinzig, G. S. Guitierrez-Cabello, et al. Longtail driving scenarios with reasoning traces: The kitscenes longtail dataset. arXiv preprint arXiv:2603.23607, 2026. [31] A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado. GAIA-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023. [32] N. Agarwal, A. Ali, M. Bala, Y. Balaji, et al. Cosmos world foundation model platform for physical AI. arXiv preprint arXiv:2501.03575, 2025. 30
[33] X. Ren, Y. Lu, T. Cao, R. Gao, S. Huang, A. Sabour, T. Shen, T. Pfaff, J. Z. Wu, R. Chen, et al. Cosmos-drive-dreams: Scalable synthetic driving data generation with world foundation models. arXiv preprint arXiv:2506.09042, 2025. [34] X. Ren, T. Shen, J. Huang, H. Ling, Y. Lu, M. Nimier-David, T. Müller, A. Keller, S. Fidler, and J. Gao. Gen3c: 3d-informed world-consistent video generation with precise camera control. In CVPR, 2025. [35] J. Wang, B. Sun, Y. Bai, V. Casser, S. Peng, Z. Zhu, M.-L. Shih, X. Masotto, S.-Y. Su, K. V. Parvate, et al. Sensor2sensor: Cross-embodiment sensor conversion for autonomous driving. arXiv preprint arXiv:2605.22809, 2026. [36] H. Zhou, L. Lin, J. Wang, Y. Lu, D. Bai, B. Liu, Y. Wang, A. Geiger, and Y. Liao. Hugsim: A real-time, photo-realistic and closed-loop simulator for autonomous driving. TPAMI, 2025. [37] Z. Chen, J. Yang, J. Huang, R. De Lutio, J. Martinez Esturo, B. Ivanovic, O. Litany, Z. Gojcic, S. Fidler, M. Pavone, et al. Omnire: Omni urban scene reconstruction. In ICLR, 2025. [38] K. Chitta, D. Dauner, and A. Geiger. Sledge: Synthesizing driving environments with generative models and rule-based traffic. In ECCV, 2024. [39] Z. Li, K. Li, S. Wang, S. Lan, Z. Yu, Y. Ji, Z. Li, Z. Zhu, J. Kautz, Z. Wu, et al. Hydramdp: End-to-end multimodal planning with multi-target hydra-distillation. arXiv preprint arXiv:2406.06978, 2024. [40] Y. Zheng, R. Liang, K. Zheng, J. Zheng, L. Mao, J. Li, W. Gu, R. Ai, S. Li, X. Zhan, et al. Diffusion-based planning for autonomous driving with flexible guidance. In ICLR, 2025. [41] A. Ghosh, S. Narasimhan, M. Chandraker, and F. Pittaluga. Rad-lad: Rule and language grounded autonomous driving in real-time. arXiv preprint arXiv:2603.28522, 2026. [42] C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li. Drivelm: Driving with graph visual question answering. In ECCV, 2024. [43] X. Zhou, X. Han, F. Yang, Y. Ma, V. Tresp, and A. Knoll. Opendrivevla: Towards end-to-end autonomous driving with large vision language action model. In AAAI, 2026. [44] Z. Xu, Y. Bai, Y. Zhang, Z. Li, F. Xia, K.-Y. K. Wong, J. Wang, and H. Zhao. Drivegpt4-v2: Harnessing large language model capabilities for enhanced closed-loop autonomous driving. In CVPR, 2025. [45] B. Liao, S. Chen, H. Yin, B. Jiang, C. Wang, S. Yan, X. Zhang, X. Li, Y. Zhang, Q. Zhang, et al. Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving. In CVPR, 2025. [46] J. Lu, J. Guan, Z. Huang, J. Li, G. Li, L. Kong, Y. Li, H. Wang, S. Xu, Y. Luo, et al. Onevl: One-step latent reasoning and planning with vision-language explanation. arXiv preprint arXiv:2604.18486, 2026. [47] N. Carion, L. Gustafson, Y.-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V. Alwala, H. Khedr, A. Huang, et al. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719, 2025. [48] W. Huang, J. Zhang, S. Li, T. Jia, J. Duan, Y. Cheng, J. Cho, M. Wallingford, R. Soraki, C. D. Kim, et al. Wilddet3d: Scaling promptable 3d detection in the wild. arXiv preprint arXiv:2604.08626, 2026. [49] A. R. Zamir, T. Wekel, P. Agrawal, C. Wei, J. Malik, and S. Savarese. Generic 3d representation via pose estimation and matching. In ECCV, 2016. 31
[50] K. Vuong, R. Tamburo, and S. G. Narasimhan. Toward planet-wide traffic camera calibration. In WACV, 2024. [51] K. Vuong, A. Ghosh, D. Ramanan, S. Narasimhan, and S. Tulsiani. Aerialmegadepth: Learning aerial-ground reconstruction and view synthesis. In CVPR, 2025. [52] G. Berton and C. Masone. Megaloc: One retrieval to place them all. In CVPR Workshops, 2025. [53] OpenStreetMap contributors. OpenStreetMap. https://www.openstreetmap.org, 2017. Data retrieved from https://planet.openstreetmap.org. [54] N. Sriram, B. Liu, F. Pittaluga, and M. Chandraker. Smart: Simultaneous multi-agent recurrent trajectory prediction. In ECCV, 2020. [55] P. Cai, Y. Lee, Y. Luo, and D. Hsu. Summit: A simulator for urban driving in massive mixed traffic. In ICRA, 2020. Eskenazi. Waymo rolls toward San Francisco Airport. a [56] J. showdown is brewing. https://missionlocal.org/2024/12/ waymo-rolls-toward-san-francisco-airport-showdown-brewing/, 2024. [57] S. Sural, N. Sahu, and R. Rajkumar. Workzone3d: A multimodal dataset for 3d work zone perception in autonomous driving. In WACV, 2026. [58] J. Cheng, Y. Chen, X. Mei, B. Yang, B. Li, and M. Liu. Rethinking imitation-based planner for autonomous driving. In ICRA, 2024. [59] M. Treiber, A. Hennecke, and D. Helbing. Congested traffic states in empirical observations and microscopic simulations. Physical review E, 2000. [60] L. Pan, D. Baráth, M. Pollefeys, and J. L. Schönberger. Global structure-from-motion revisited. In ECCV, 2024. [61] L. Pan, J. L. Schönberger, and M. Pollefeys. Global structure-from-motion meets feedforward reconstruction. In CVPR, 2026. [62] S. Hagedorn, L. Donkov, A. Distelzweig, and A. P. Condurache. When planners meet reality: How learned, reactive traffic agents shift nuplan benchmarks. arXiv preprint arXiv:2510.14677, 2025. [63] R. Luo, H. Yang, M. Watson, A. Sharma, S. Veer, E. Schmerling, and M. Pavone. Sim2val: Leveraging correlation across test platforms for variance-reduced metric estimation. arXiv preprint arXiv:2506.20553, 2025. [64] L. Pinto and A. Gupta. Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. In ICRA, 2016. [65] S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen. Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection. IJRR, 2018. [66] L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch. Decision transformer: Reinforcement learning via sequence modeling. NeurIPS, 2021. [67] S. Emmons, B. Eysenbach, I. Kostrikov, and S. Levine. Rvs: What is essential for offline rl via supervised learning? arXiv preprint arXiv:2112.10751, 2021. [68] G. L. Smith, S. F. Schmidt, and L. A. McGee. Application of statistical filter theory to the optimal estimation of position and velocity on board a circumlunar vehicle. 1962. 32
[69] J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny. Vggt: Visual geometry grounded transformer. In CVPR, 2025. [70] N. V. Keetha, N. Müller, J. Schönberger, L. Porzi, Y. Zhang, T. Fischer, A. Knapitsch, D. Zauss, E. Weber, N. Antunes, et al. Mapanything: Universal feed-forward metric 3d reconstruction. In 3DV, 2025. [71] A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023. [72] S. Gao, J. Yang, L. Chen, K. Chitta, Y. Qiu, A. Geiger, J. Zhang, and H. Li. Vista: A generalizable driving world model with high fidelity and versatile controllability. NeurIPS, 2024. [73] X. Wang, Z. Zhu, G. Huang, X. Chen, J. Zhu, and J. Lu. Drivedreamer: Towards real-worlddrive world models for autonomous driving. In ECCV, 2024. [74] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. [75] A. Ghildyal and F. Liu. Shift-tolerant perceptual similarity metric. In ECCV, 2022. [76] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. NeurIPS, 2017. [77] T. Kynkäänniemi, T. Karras, M. Aittala, T. Aila, and J. Lehtinen. The role of imagenet classes in fr\’echet inception distance. arXiv preprint arXiv:2203.06026, 2022. [78] N. Tumanyan, O. Bar-Tal, S. Bagon, and T. Dekel. Splicing vit features for semantic appearance transfer. In CVPR, 2022. [79] S. Fu, N. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data. arXiv preprint arXiv:2306.09344, 2023. [80] C. Wickrema, S. Leary, S. Sarkar, M. Giglio, E. Bianchi, E. Mace, and M. Twardowski. Benchmarking image similarity metrics for novel view synthesis applications. arXiv preprint arXiv:2506.12563, 2025. [81] M. Khan, H. Fazlali, D. Sharma, T. Cao, D. Bai, Y. Ren, and B. Liu. Autosplat: Constrained gaussian splatting for autonomous driving scene reconstruction. In ICRA, 2025. [82] G. Berton, G. Trivigno, B. Caputo, and C. Masone. Eigenplaces: Training viewpoint robust models for visual place recognition. In ICCV, 2023. [83] M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou. The faiss library. arXiv preprint arXiv:2401.08281, 2024. [84] D. DeTone, T. Malisiewicz, and A. Rabinovich. Superpoint: Self-supervised interest point detection and description. In CVPR Workshops, 2018. [85] P. Lindenberger, P.-E. Sarlin, and M. Pollefeys. Lightglue: Local feature matching at light speed. In ICCV, 2023. [86] J. L. Schonberger and J.-M. Frahm. Structure-from-motion revisited. In CVPR, 2016. [87] B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022. 33
[88] Z. Pataki, P.-E. Sarlin, J. L. Schönberger, and M. Pollefeys. Mp-sfm: Monocular surface priors for robust structure-from-motion. In CVPR, 2025. [89] M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. TPAMI, 2024. [90] M. Grupp. evo: Python package for the evaluation of odometry and slam. https://github. com/MichaelGrupp/evo, 2017. [91] S. Umeyama. Least-squares estimation of transformation parameters between two point patterns. TPAMI, 1991. [92] J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers. A benchmark for the evaluation of rgb-d slam systems. In IROS, 2012. [93] A. Krishnan, S. Liu, P.-E. Sarlin, O. Gentilhomme, D. Caruso, M. Monge, R. Newcombe, J. Engel, and M. Pollefeys. Benchmarking egocentric visual-inertial slam at city scale. In ICCV, 2025. [94] J. J. Spilker Jr, P. Axelrad, B. W. Parkinson, and P. Enge. Global positioning system: theory and applications, volume I. American Institute of Aeronautics and Astronautics, 1996. [95] B. Klingner, D. Martin, and J. Roseborough. Street view motion-from-structure-from-motion. In ICCV, 2013. [96] L.-T. Hsu. Analysis and modeling gps nlos effect in highly urbanized area. GPS solutions, 2018. [97] Mapillary. https://www.mapillary.com. Accessed: 2026-04-29. [98] A. Torii, J. Sivic, and T. Pajdla. Visual localization by linear combination of image descriptors. In ICCV Workshops, 2011. [99] A. Ghosh, Y. Patel, M. Sukhwani, and C. Jawahar. Dynamic narratives for heritage tour. In ECCV Workshops, 2016. [100] P. Lindenberger, P.-E. Sarlin, J. Hosang, M. Balice, M. Pollefeys, S. Lynen, and E. Trulls. Scaling image geo-localization to continent level. In NeurIPS, 2025. [101] A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft. Simple online and realtime tracking. In ICIP, 2016. [102] Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu, and X. Wang. Bytetrack: Multi-object tracking by associating every detection box. In ECCV, 2022. [103] J. Tung, G. Chou, R. Cai, G. Yang, K. Zhang, G. Wetzstein, B. Hariharan, and N. Snavely. Megascenes: Scene-level view synthesis at scale. In ECCV, 2024. [104] C. Sakaridis, D. Dai, and L. Van Gool. Acdc: The adverse conditions dataset with correspondences for semantic driving scene understanding. In ICCV, 2021. [105] M. Bijelic, T. Gruber, F. Mannan, F. Kraus, W. Ritter, K. Dietmayer, and F. Heide. Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather. In CVPR, 2020. [106] Federal Highway Administration. Manual on Uniform Traffic Control Devices for Streets and Highways. U.S. Department of Transportation, 2026. URL https://mutcd.fhwa.dot. gov/. 34
[107] Reuters. China bans “smart” and “autonomous driving” terms in vehicle ads. https://www.reuters.com/business/autos-transportation/ china-bans-smart-autonomous-driving-terms-vehicle-ads-2025-04-17/, 2025. [108] U.S. Department of Transportation, National Highway Traffic Safety Administration. FARS Data: People Killed in Construction or Maintenance Zones. https://www-fars.nhtsa. dot.gov/People/PeopleAllVictims.aspx, 2026. [109] U.S. Code of Federal Regulations. Title 23 c.f.r. § 655.603 — standards. https://www. ecfr.gov/current/title-23/section-655.603, 2024. Establishes that States adopt the National MUTCD or publish State MUTCDs/Supplements in substantial conformance. [110] U.S. Code of Federal Regulations. Title 23 c.f.r. part 630, subpart j — work zone safety and mobility. https://www.ecfr.gov/current/title-23/part-630/subpart-J, 2024. Requires State-level processes and procedures for work zone management. [111] O. Contributors. Openscene: The largest up-to-date 3d occupancy prediction benchmark in autonomous driving. https://github.com/OpenDriveLab/OpenScene, 2023. [112] W. Tong, C. Sima, T. Wang, L. Chen, S. Wu, H. Deng, Y. Gu, L. Lu, P. Luo, D. Lin, et al. Scene as occupancy. In ICCV, 2023.
35