LTM: Large-scale Terrain Model for Landscapes Xiao Fu, Yue Hu, Meida Chen, Peter Anthony Beerel, Barath Raghavan University of Southern California
arXiv:2607.08711v1 [cs.CV] 9 Jul 2026
Abstract Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfireprone regions often span vast areas where conventional reconstruction methods underperform. Airborne LiDAR systems provide high-resolution terrain data, but they are expensive and infrequently updated. Image-based methods offer a lower-cost alternative, but struggle due to sparse visual features and limited image overlap. We propose a multimodal reconstruction framework leveraging outdated Digital Elevation Models (DEMs) as geometric priors for imagebased 3D reconstruction. Our key innovation is physics-based pixel-pixel alignment between images and DEM data, dramatically reducing computational complexity by eliminating expensive feature matching procedures. To validate our approach, we developed a large-terrain simulator based on a real wildfire-prone area, generating realistic images enabling a comprehensive evaluation. Given posed images and legacy DEMs, our method produces high-fidelity depth maps while maintaining real-time performance. We find significant improvements in reconstruction accuracy and computational efficiency over existing techniques, offering a scalable solution for wildfire response.
Introduction Environmental monitoring has yet to fully leverage the vast network of cameras already deployed in wildfireprone and other environmentally vulnerable regions. These camera systems capture rich spatio-temporal imagery that offers valuable insights into dynamic natural landscapes. Such imagery can help identify both gradual ecological changes, such as vegetation growth, and sudden environmental events, including avalanches (Barbolini et al. 2011), floods (Mudashiru et al. 2021), and wildfires (Fu et al. 2024). Accurate mapping of such landscapes not only enhances disaster preparedness but also aids automated semantic analysis, including assessments of disaster intensity. In particular, fuel maps, which provide vegetation-related spatial semantics, are widely used in conjunction with terrain models to support wildfire propagation simulators such as FARSITE (Finney 1998) and FlamMap (Finney 2006). In addition, 3D reconstruction techniques generate detailed repreCopyright © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Figure 1: A novel framework for large-scale dynamic terrain updates in end-to-end 3D semantic mapping.
sentations of natural terrain, which serve as the spatial foundation on which fuel maps are overlaid, enabling the extraction of timely and accurate insights that improve early warning systems and support more informed and responsive disaster management. Landscape, terrain, and vegetation are traditionally captured using remote sensing methods using spacecraft (Hirano, Welch, and Lang 2003) or aircraft-based platforms (Dobrowski et al. 2008; Rodriguez and Aggarwal 2002). However, the resulting 3D models are typically updated only on an annual basis, if that, due to the vast areas that must be scanned (Gorelick et al. 2017; Krishnan et al. 2011; Kervyn et al. 2007). This infrequent update cycle is insufficient for effective disaster mitigation, as significant changes in vegetation and surface conditions can occur over seasonal or even monthly timescales (Yang, Meng, and Zhang 2011). In contrast, ground-based camera networks offer higher temporal resolution, but lack the top-down, widearea coverage that is critical for comprehensive landscape monitoring. While prior ground image-based 3D reconstruction methods have shown strong performance in urban environments, their effectiveness in wildfire-prone natural landscapes remains limited (Iglhaut et al. 2019). These approaches typically focus on estimating camera pose and scene structure from image sequences, enabling spatial inference of objects and surfaces. However, their applicability to vegetated environments is constrained by several factors. Vegetation often lacks color contrast and is constantly changing struc-
turally due to wind and seasonal growth (Nguyen et al. 2015; Sidle 2017; Schmidt et al. 2025). At the same time, the large spatial extent covered by each ground-level camera often leads to sparse image overlap, reducing feature correspondence and increasing reconstruction errors (Deng et al. 2025). In addition, hand-crafted feature descriptors such as SIFT (Lowe 2004) and ORB (Rublee et al. 2011) struggle to detect stable, repeatable keypoints in vegetation-dominated scenes due to low texture, occlusions, and dynamic patterns (Qadri and Kantor 2021). Ground-level landscape imagery is often spatially sparse and temporally asynchronous, further complicating model-based reconstruction in largescale terrain. Moreover, unlike urban environments—which exhibit fine-grained, centimeter-level geometry and rapid second-level changes—natural landscapes require meterlevel spatial resolution (Hodgson et al. 2003) and evolve over weekly to monthly timescales, demanding fundamentally different reconstruction assumptions. Existing image-based methods use various 3D representations to achieve more accurate scene rendering and mapping, such as neural radiance fields (NeRF) to enhance the accuracy of reconstruction (Mildenhall et al. 2021). Inspired by this approach, we adopt a more suitable representation, DEM, for landscape applications instead of relying on voxel-based methods. However, existing image-based methods are predominantly designed for urban environments because of the scarcity in available large-scale vegetated datasets. Even when applying state-of-the-art largescale NeRF 3D reconstruction techniques (Turki, Ramanan, and Satyanarayanan 2022), performance degrades significantly in non-urban natural landscape scenes (Deng et al. 2025; Lu et al. 2023; Hermann et al. 2024). For example, satellite-based NeRF approaches such as CityNeRF (Turki, Ramanan, and Satyanarayanan 2022) incorporate digital surface models (DSM) into their evaluation pipeline but remain optimized for urban settings. We argue that existing stateof-the-art 3D representations are insufficient for accurate and efficient mapping and reconstruction in landscape environments. Therefore, we propose an improved method for mapping in landscape spatial monitoring that uses secondary scene information, particularly depth maps generated from ground-based RGB images. We present a cross-modality landscape mapping approach that accurately detects changes in scenes while updating scene-captured 3D models as presented in Figure 1. First, we adopt a vertical historical large-scale baseline model using a DEM in raster format. Second, we leverage existing learning-based monocular depth estimation tools to generate horizontal depth information from ground-based imagery. Third, by combining vertical and horizontal spatial information with automated identification of changing sections, we can accurately update the landscape by achieving consensus across multiple images. Through the identification of these shifting landscape features, we can dynamically update the baseline raster models to maintain current and accurate environmental representations. Contribution. • Monocular depth mapping. We develop a self-
supervised monocular depth estimation system specifically designed for large-scale vegetated outdoor environments, addressing the challenges of feature extraction in wildfire-prone landscapes. • Outdoor cross-modal dataset. We compile and evaluate existing cross-modality datasets and graphic-based simulators focused on wildfire-prone environments, providing comprehensive benchmarks for localization and mapping accuracy in large-scale wildfire-prone settings. • 3D representation and segmentation for large-scale outdoor landscapes. We introduce the strategic use of Digital Elevation Models (DEMs) as 2D arrays that effectively capture 3D environmental information. Our approach successfully achieves robust 2D-3D alignment and improves existing rendering and mapping frameworks. This method specifically provides alignment of each image pixel to the DEM raster pixel.
Related Work Spatial Representation and Reconstruction Traditional 3D representations, including point clouds (Rusu and Cousins 2011), meshes, and voxel grids (Wu et al. 2015), excel in urban reconstruction with high fidelity and flexible visualization. Combined with feature descriptors like ORB (Rublee et al. 2011) and SIFT (Lowe 2004), these representations increasingly adopt neural and statistical methods to enhance accuracy and visual quality. Traditional 3D reconstruction methods rely heavily on feature descriptors through approaches such as Multi-view Stereo (MVS) (Schönberger et al. 2016) and Simultaneous Localization and Mapping (SLAM) (Mur-Artal, Montiel, and Tardos 2015). Recent advances, particularly Neural Radiance Fields (NeRF) (Mildenhall et al. 2021) and 3D Gaussian Splatting methods like PixelSplat (Charatan et al. 2024), have shown promise but face significant challenges in nonurban scenarios (Mall, Hariharan, and Bala 2023).
2D-3D Alignment and Multi-modal Improvement Feature matching and 2D-3D alignment is essential for 3D reconstruction. Prior work has explored cross-modality registration between 2D images and 3D representations, including implicit (e.g., NeRF) and explicit (e.g., LiDAR point clouds) (Zhou et al. 2023; Wu et al. 2023; Li et al. 2022b; Chen et al. 2024; Bhunia, Li, and Bilen 2024; Cho, Kim, and Kim 2025). Line correspondences offer an alternative when visual features are sparse (Yu et al. 2020). Key challenges remain, including runtime-storage trade-offs (Brachmann, Cavallari, and Prisacariu 2023) and large search spaces (Li and Lee 2021). While neural feature matchers such as LoFTR (Sun et al. 2021), SuperGlue (Sarlin et al. 2020), and others (Panek, Kukelova, and Sattler 2022; Cheng et al. 2025) achieve robust results, they remain computationally intensive and less effective in large-scale, vegetated settings. DEM-supported vision tasks Prior work (Liu, Li, and He 2022) has described the advantages of DEMs in large-scale terrain representation, particularly their efficiency and storage benefits. Early work has explored using DEM models
for outdoor geo-localization (Baatz et al. 2012), which addresses the difficulty of feature recognition by using skylineshaped mountain ridges as matching points.
3D Reconstruction and Semantics Sparse images for vast area Large-scale environmental monitoring typically relies on satellite-based systems (Mall, Hariharan, and Bala 2023; Turki, Ramanan, and Satyanarayanan 2022) for broad coverage; however, their limited temporal and spatial resolution motivates the development of complementary image-based reconstruction methods. Recent neural approaches, including DUST3R (Wang et al. 2024), POW3R (Jang et al. 2025), and MAST3R (Leroy, Cabon, and Revaud 2024), demonstrate effective performance in sparse image scenarios by integrating image features and geometric priors for efficient 3D reconstruction. Depth estimation Early depth estimation approaches relied on multi-view stereo (Li and Snavely 2018; Godard et al. 2019), but demonstrated poor performance in outdoor vegetated environments. Recent advances in monocular depth estimation have achieved impressive results using only RGB images as input, with state-of-the-art methods including UniDepth (Piccinelli et al. 2024, 2025), Depth Anything (Yang et al. 2024a,b), DepthPro (Bochkovskii et al. 2024), Metric3D V2 (Hu et al. 2024), and Marigold (Ke et al. 2025). These models broaden the applicability of depth estimation across diverse settings while maintaining efficiency and scalability. 3D semantic segmentation Recent advances in end-toend 3D semantic segmentation have demonstrated significant benefits for scene understanding by integrating volumetric reconstruction with semantic labeling. Methods like LSM (Fan et al. 2024) show how 3D segmentation networks generate spatially-coherent semantic representations. These end-to-end 3D segmentation techniques that extend beyond open-vocabulary 2D methods like LSeg (Li et al. 2022a) offer particular promise for large-scale landscape semantic analysis.
Preliminaries Digital Elevation Models DEMs (Mukherjee et al. 2013) are typically represented using raster data structures. While this sacrifices geometric detail compared to mesh or point cloud formats, it offers significant computational advantages for mega-scale applications spanning cities or wildlands. For landscapes exceeding 1 square km in scale, raster-based DEMs as a 3D representation provide an ideal balance between representational fidelity and computational efficiency. This efficiency becomes particularly valuable when integrating outdated landscape information with real-time imagery for image-based landscape mapping, enabling scalable reconstruction pipelines that can process vast geographic areas while maintaining reasonable computational requirements. OpenTopography (OpenTopography; Krishnan et al. 2011) provides an annual DEM update based on the results of the LiDAR survey.
(a) Storage
(b) Visualization
Figure 2: DEM efficiently represents large-scale 3D terrain through per-pixel elevation storage. This 16 km2 area shows detailed ravines and ridges.
As presented in Figure 2, the 2D raster on the right presents a 4000 m × 4000 m area where each pixel of 1 m × 1 m is associated with the elevation of such a location. The left presents the same terrain projected into a 3D space, as you can see the gully and ridges in the mountainous area. The area shown covers all that was impacted by the 2019 Getty Fire.
Fuel Map Large-scale terrain semantic segmentation provides essential fuel information for fire behavior analysis and risk assessment. LANDFIRE (Rollins 2009) provides fuel maps with 30 m×30 m grid resolution covering wildfire-prone areas in the U.S., based on established classifications including Anderson’s 13 fuel models (Anderson 1982) and Scott and Burgan’s 40 fuel models (Andrews 2018), which categorize vegetation by height and flammability to define burnability characteristics. However, existing fuel classification systems rely on static, coarse-resolution data that cannot capture monthly or quarterly environmental changes critical for disaster response and fire behavior analysis (Finney 1998, 2006).
Problem Formulation Existing methods that adopt 3D representations and semantic segmentation fall short when providing solutions for large-area vegetated outdoor environments. It is essential to combine outdated Digital Elevation Models (DEMs) with current posed imagery to generate timely 3D terrain and fuel models. Terrain models are frequently outdated due to natural environmental changes, including prescribed burns, vegetation growth, and habitat shifts. Existing surveying frequencies remain insufficient for critical applications such as disaster mitigation. Landscape changes in wildfire-prone environments occur more gradually than urban transformations, but the temporal lag between actual terrain and fuel changes and available digital representations creates significant gaps in environmental monitoring capabilities. Meanwhile, the large areas impacted by wildfire require methods suitable for sparse and spatially imbalanced imagery. Existing infrastructure often provides minimal view overlap, and even when overlapping views exist, establishing reliable correspondences between
Method NeRF (Neural Radiance Fields) Gaussian Splatting Gaussian Splatting + Process Multi-view Stereo (COLMAP) SLAM Radiant Foam Monocular Depth Estimation 2D-3D Line Correspondence SuperGlue (Feature Matching) DEM-Supported Methods
Task Novel View Synthesis Novel View Synthesis Uncertainty-aware Rendering Dense 3D Reconstruction Real-time Mapping Real-time Ray Tracing Single-Image Depth Prediction Pose Estimation from 3D Map Feature Matching Across Images Large-Scale Terrain Mapping
3D Representation Volume (Implicit) Point Cloud (Implicit) Point Cloud (Implicit) Point Cloud/Mesh Point Cloud / Mesh Volume (Sparse Voxels) Depth Map 3D Lines (LiDAR Map) N/A (Keypoints) Elevation Map
Mapping Complexity (W/ Poses) O(NvMP) O(NpMP) O(Npˆ3+NpMP) O(NfMP) O(NfMP) O(NvMP) O(NfMP) O(NfMP) O(Nfˆ2MPˆ2) O(NvMP)
Mapping Complexity (W/o Poses) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MP) O(Nfˆ2MPˆ2) O(Nfˆ2MP)
Table 1: Compared to other 3D representations, DEM for landscape in 3D modeling can largely reduce the computation cost for cross-modality tasks for 3D reconstruction in creating 2D-3D correspondence. images remains challenging due to repetitive vegetation patterns and sparse, distinctive features.
Method Overview We propose an image-based pipeline for large-scale wildfire-prone landscape mapping. The pipeline integrates depth estimation, 2D-3D alignment, and image-based segmentation, leveraging the inherent stability of natural terrain. This stability enables historical DEMs to serve as reliable geometric priors for terrain reconstruction using current imagery. Our rasterization approach efficiently represents large-scale areas, compressing 10 km × 10 km landscapes into manageable 2D arrays. The raster-based structure supports multi-resolution spatial analysis, enabling both coarse semantic fuel mapping (30m resolution) and fine-grained terrain mapping (1m updated DEMs). We first examine existing spatial representations and image-based 3D reconstruction approaches, analyzing their limitations for wildfire landscapes. Learning-based and photogrammetry-based methods are most widely adopted among 3D reconstruction techniques. We then investigate how these methods can be extended and combined with landscape-focused representations to address the unique challenges of vegetated environment monitoring. Objectives Given outdated DEM and real-time images, our goal is to provide an update to the DEM and fuel maps. This will allow the 3D semantic map of a large-scale wildfire-prone area to be kept up-to-date.
Parameter Analysis for 3D Reconstruction We perform a computational complexity analysis that compares 3D representations based on raster, point, and volumetric data for large-scale terrain reconstruction. Table 1 presents the breakdown of the computational costs on different spatial scales and reconstruction methods. For raster-based DEM representations, computational complexity scales as the number of pixels, points, or other spatial registration methods. If the analysis is constructed to a square area of land where the width and length of the area are the same. Meanwhile, the precision (e.g., 1m gap between each data point) of 3D representation remains the same. As the coverage area increases, point-based methods suffer from cubic growth of width in point density requirements to main-
Algorithm 1: Pixel-based On-raster Ray Tracing Input: Outdated 3D Models, real-time posed images Parameter: DEM [][], cam info θ,ϕ, Camelev , img[][] Output: Depth map depth[][], img-DEM[][] 1: for u, v in img[][] do 2: ray ⃗ ← F (θ, ϕ, u, v) 3: x, y, z = Xcam , Ycam , Camelev 4: while DEM [x][y] < z do 5: x = x + ∆X (ray) ⃗ 6: y = y + ∆Y (ray) ⃗ 7: z = z + ∆elev (ray) ⃗ 8: end while 9: depth[u][v] = |(x, y) − (Xcam , Ycam )| 10: img-DEM[u][v] = (x, y) 11: end for 12: return depth[][], img-DEM[][]
tain surface fidelity, while raster-based approaches maintain quadratic scaling with respect to width. Our analysis reveals that raster-based representations achieve superior computational efficiency for large-scale vegetated terrain, with performance gains becoming more pronounced for larger scales. For the 10 km × 10 km wildland areas typical in disaster monitoring applications, rasterbased methods demonstrate a reduction in computational requirements compared to equivalent point cloud representations while maintaining sufficient geometric accuracy for large-scale outdoor vegetated terrain. It is noted that the raster 3D representation cannot capture hollow and concave features of landscape (e.g., overhangs, caves). But those scenarios are less relevant to wildfire propagation.
Photogrammetry-based DEM-integrated Method Raster-pixel alignment We designed ray tracing techniques to perform 2D-3D alignment between DEM raster pixels and image pixels. In Algorithm 1, we demonstrate a method to perform this pixel-to-pixel matching. Given camera poses, camera locations, camera field of view (FOV), and terrain DEM, the algorithm outputs the depth map of the image and the 2D image-DEM alignment. As a result, each entry of the image-resolution 2D array is associated with DEM coordinates. In a selected DEM area, coordinates are represented by the distance from the origin divided by the DEM precision (i.e., 1 m).
In line 3, the location initialization for ray traversal is established in this designated coordinate system. The initial elevation is either collected from the camera’s associated barometer or based on the DEM and common camera relative height offset (e.g., tripod height). In line 2, the camera ray vector is created based on camera poses, camera FOV, and image resolution. Given the image pixel location, the ray is generated based on the proportional FOV offset. The ray direction defines the path along which the testing location traverses through space. The variables (x, y) define the testing location used to query the terrain surface elevation, while elevation defines the ray elevation at the current testing location. In line 4, the stopping condition is defined as when the ray elevation falls below the terrain surface elevation. Thus, we identify the raster coordinate with which the image pixel aligns. We then calculate the depth of the pixel as the distance between the camera location and the found raster coordinate in line 9. Vegetated perturbation topographical simulator We present a novel topographical landscape simulator for vegetated terrains with vegetation perturbation and burned landscapes. Our framework combines high-resolution digital elevation models with procedurally generated vegetation using Unreal Engine, enabling photorealistic rendering while providing precise ground truth data for geometry, semantics, and temporal evolution. We preserve geospatial consistency through real-world DEM data adapted to game engine input heightmaps, facilitating meaningful sim-to-real benchmarking. Existing large-scale datasets and simulators, including KITTI (Geiger et al. 2013), 7-Scenes (Glocker et al. 2013), CARLA (Dosovitskiy et al. 2017), and AirSim (Shah et al. 2017), are predominantly designed for urban environments. While datasets for vegetated landscapes exist, they are typically aerial-based and lack the terrestrial-level imagery and topographical realism essential for evaluating 3D reconstruction or vegetation segmentation in large-scale outdoor environments. To address the lack of controlled vegetated terrain in remote wildland areas, we develop a custom simulation environment for outdoor vegetated scenes. Vegetation-shift images segmentation Building upon fuel type concepts (Anderson 1982), we focus on three primary vegetation categories critical for wildfire assessment: grass (herbaceous), shrub, and tree (timber). By incorporating semantic segmentation into our image-based 3D mapping framework, we provide both geometric reconstruction and semantic understanding of vegetated landscapes, enabling more comprehensive environmental monitoring than geometric reconstruction alone. Open vocabulary-based semantic segmentation methods (i.e., LSeg (Li et al. 2022a)) for vegetation segmentation is executed by choosing the primary vegetation categories. Then, we utilized the segmentation results to create a fuel map. End-to-end 3D semantic fuel map After aligning DEM raster cell with corresponding vegetation on LANDFIREscale fuel maps, each pixel represents vegetation within a 30 m × 30 m grid cell. While this contrasts with the 1 m
precision of DEM maps, fuel maps are typically overlaid with DEMs during wildfire behavior analysis. The 2D fuel map can be projected into 3D space using DEM-derived elevation information. Assume E(x, y) is the elevation map (e.g., from DEM), in meters. L(x, y) is the fuel map, categorical. Coordinate (x, y) ∈ Z2 are pixel indices. s = 30 m is the fuel map spatial resolution. We project each coordinate to 3D space using the elevation map and associate the corresponding fuel map cell label: P (x, y) = (x · s, y · s, E(x, y)) , L(x, y) Fuels are categorized into three primary types: grass, shrub, and tree. Additional image categories include sky and mountain, which represent two major non-fuel domains present in wildfire-prone landscape imagery. To resolve multiple image pixels projecting onto the same fuel map cell (each representing a 30 m×30 m area), we use majority voting over the label space in Equation 1. X ˆli,j = arg max 1L[u,v]=l (1) l∈L
i−D[u][v]∈N (i,j)
ˆli,j is the aggregated label assigned to fuel cell (i, j), L is the set of possible fuel labels, N (i, j) is the set of 30 m × 30 m area centered at (i, j). i-D[u][v] is the image-to-DEM alignment img-DEM[][] we acquire from Algorithms 1. Depth map with terrain priors Depth maps provide a efficient 3D representation where each pixel encodes the distance from camera to surface. We generate depth maps using Algorithm 1, enabling direct 3D projection when combined with known camera poses. Following FireLoc (Fu et al. 2024), we fuse neural depth estimation with DEM constraints through pixel-wise sampling and RANSAC regression. The approach can be enhanced using state-ofthe-art monocular depth frameworks (Piccinelli et al. 2025; Bochkovskii et al. 2024). This hybrid method leverages DEMs for stable terrain baselines while neural networks capture detailed surface variations, addressing training data limitations in outdoor environments. Performance degrades under significant occlusion conditions.
Experiments Datasets Site selection and terrain model We selected the Getty Fire site, focusing on the critical area spanning both sides of Interstate 405. Camera positions were established east of the highway with westward orientation to capture the region where the Getty Fire threatened to breach containment and cause extensive damage. Digital Elevation Model (DEM) data were obtained from OpenTopography. Simulated environment for large-scale vegetated scenes Our simulator generates realistic wildfire-prone landscapes with controllable vegetation parameters, enabling systematic evaluation of semantic mapping under varying environmental conditions. Built using Unreal Engine, the framework combines high-resolution digital elevation models with procedurally generated vegetation to produce photorealistic rendering while providing precise ground truth for
(a) Images
(b) Segmentation
(c) Images
(d) Segmentation
Figure 3: We present 2 pairs of real-world, simulated images, and their respective image-based segmentation results. The segmentation are vegetation types based on wildfire fuel models. geometry, semantics, and temporal changes. The simulator replicates the same topographical area as our smartphonecollected real-world data, maintaining identical camera configurations. This enables direct sim-to-real comparison for vegetation change scenarios, including plant growth, prescribed burns, and other environmental modifications that alter landscapes over time. We evaluate both simulator performance and compare between simulated and real environments. Real-world scenes collection and alignment We collected real-world imagery from wildland locations using an iPhone 14 Pro to validate our simulation environment. The alignment between simulated and real scenes enables environmental parameter calibration, while the simulator provides ground truth geometry unavailable in real deployments. This hybrid approach balances realism with quantitative evaluation requirements. Following established practices in location-based computer vision (Kendall, Grimes, and Cipolla 2015), we use smartphone sensors to capture images with pose and location data. Due to inherent drift in smartphone GPS, IMU, and barometer sensors, we manually filter results through visual inspection to ensure data quality. The simulator then generates corresponding images using identical pose and location information from the smartphone data.
Experiment Details Simulations were conducted on NVIDIA RTX 3090 GPUs, while depth estimation experiments used dual NVIDIA A6000 GPUs. The simulation environment enables systematic perturbation of landscape parameters, vegetation density, terrain consistency, and cloud coverage condition.
Evaluation Metrics Sim-to-real accuracy We evaluate 3D reconstruction quality through image-level accuracy metrics, as terrain reconstruction ultimately depends on accurate depth estimation from individual viewpoints. Primary metrics include the Structural Similarity Index Measure (SSIM)(Wang et al. 2004) for geometric consistency and the Learned Perceptual Image Patch Similarity (LPIPS)(Zhang et al. 2018) for the evaluation of perceptual quality.
(a) Real-world vegetation
(b) Real-world vegetation
(c) Simulated perturbation
(d) Simulated perturbation
Figure 4: Fuel maps comparing real-world and simulated camera setups over a 1500 m × 1500 m area. Simulated data features vegetation perturbations. Depth estimation and semantic segmentation The Depth estimation error can be calculated through errors, RMSE, meaning the distance between the actual location and the predicted location. While the semantic segmentation can be evaluated through the F-1 score or general fuel map cell classification errors.
Experimental Results Perception of the sim-to-real We present both simulated images and corresponding real-world captures to support the vegetation drift analysis. In Figure 3a and Figure 3c, two representative scenes are shown. Each pairing is a realworld image collected with a calibrated camera and a simulated rendering generated using the same camera poses and geolocation data. The simulated results accurately capture key terrain features, including mountain ridges, ravines, and general landscape morphology. In addition, the simulator successfully reflects vegetation transitions, such as the shift from shrubland and grass to coniferous forest.
(a) Image
(b) TopoDepth (ours)
(c) UniDepth
(d) DepthPro
Figure 5: Monocular depth estimation in mountainous terrain. Our method outperforms UniDepth and DepthPro, which struggle with long-range scenes. While baselines capture fine details like vegetation outlines, they fail to estimate depth for mountain sections larger than 100 m away. TopoDepth provides accurate depth estimation for distant mountains. Results use non-aligned real-world and simulated image pairs. Vegetation segmentation Image-based vegetation segmentation performs well in identifying the plant changes from the real-world shrubs to the simulated conifer terrain in the exact location. This is significant that the segmentation models can differentiate the shrub-like vegetation and conifer as presented in Figure 3. The camera locations are at grid coordinates (2218, 1101) for Figure 3a, (2216, 1349) for Figure 3c in meters from the origin. The camera origin can be observed in Figure 4, where the 3D semantic segmentation results are presented. Fuel map Figure 4 demonstrates successful fuel type classification, capturing the transition from conifer forest to shrubland. We evaluate the model in generic scenarios where vegetation transitions occur uniformly across the landscape (e.g., shrub/grass to conifer forest), and the model effectively identifies these broad vegetation shifts. The resulting fuel classifications are projected onto 3D semantic maps using our proposed methodology. Camera coverage remains stable throughout data collection, as evidenced by the correspondence between real and segmented imagery: Figure 4a (original shrub and grass), Figure 4c (simulated coniferous forest), and Figure 3a represent one viewpoint, while Figure 4b (original shrub and grass), Figure 4d (simulated coniferous forest), and Figure 3c represent another. This demonstrates the complete end-to-end pipeline success from initial image acquisition through fuel segmentation to final 3D fuel mapping. Monocular depth estimation We evaluate TopoDepth against state-of-the-art one-shot models, as shown in Figure 5. UniDepth demonstrates general distance inference but lacks cross-scene consistency, particularly failing to maintain scale coherence across varying viewpoints. DepthPro differentiates near and far objects but cannot capture large-scale topographic structures, incorrectly classifying sky regions and distant peaks as mid-range depths. While prior methods achieve reasonable performance in controlled synthetic datasets, they exhibit severe degradation on real
Sim-to-real
LPIPS 0.4437
SSIM 0.3614
FID 354.3967
Table 2: Sim-to-real evaluation. mountainous terrain where scale ambiguity and vegetation occlusion pose significant challenges. TopoDepth produces spatially consistent depth estimates that accurately follow complex terrain transitions across varying elevations, maintaining errors within tens of meters—critical for wildfire-prone landscape mapping. Baseline methods are omitted from quantitative comparison due to depth predictions exhibiting 1-2 orders of magnitude greater error, miscalculating mountain ranges at 1000 meters as tens of meters away, rendering them unsuitable for deployment. By integrating DEM priors, our method overcomes the fundamental scale ambiguity in monocular depth estimation, achieving robust 3D understanding across simulated, real-world aligned, and image-raster aligned scenarios. Sim-to-real evaluation We validate our FireLoc simulator against real-world fire imagery in a limited set of images (Table 2). As the first simulation to incorporate dynamic vegetation in wildfire modeling, our quantitative metrics demonstrate strong sim-to-real fidelity, establishing reliable benchmarks for fire prediction applications.
Conclusion We present a novel multi-modal 3D semantic mapping method for large-scale outdoor vegetated environments. Our approach combines raster-based 3D representations with temporal information integration to achieve efficient terrain mapping in wildfire-prone landscapes. By adapting reconstruction techniques to address sparse features, temporal dynamics, and extreme scale requirements in vegetated environments, we demonstrate significant improvements in computational efficiency and accuracy. This work advances com-
puter vision tools for environmental monitoring and disaster mitigation.
References Anderson, H. E. 1982. Aids to determining fuel models for estimating fire behavior. The Bark Beetles, Fuels, and Fire Bibliography, 143. Andrews, P. L. 2018. The Rothermel surface fire spread model and associated developments: A comprehensive explanation. Gen. Tech. Rep. RMRS-GTR-371. Fort Collins, CO: US Department of Agriculture, Forest Service, Rocky Mountain Research Station. 121 p., 371. Baatz, G.; Saurer, O.; Köser, K.; and Pollefeys, M. 2012. Large scale visual geo-localization of images in mountainous terrain. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part II 12, 517–530. Springer. Barbolini, M.; Pagliardi, M.; Ferro, F.; and Corradeghini, P. 2011. Avalanche hazard mapping over large undocumented areas. Natural hazards, 56(2): 451–464. Bhunia, A.; Li, C.; and Bilen, H. 2024. Looking 3D: Anomaly Detection with 2D-3D Alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 17263–17272. Bochkovskii, A.; Delaunoy, A.; Germain, H.; Santos, M.; Zhou, Y.; Richter, S. R.; and Koltun, V. 2024. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073. Brachmann, E.; Cavallari, T.; and Prisacariu, V. A. 2023. Accelerated coordinate encoding: Learning to relocalize in minutes using rgb and poses. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5044–5053. Charatan, D.; Li, S. L.; Tagliasacchi, A.; and Sitzmann, V. 2024. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 19457–19467. Chen, S.; Cavallari, T.; Prisacariu, V. A.; and Brachmann, E. 2024. Map-relative pose regression for visual relocalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20665–20674. Cheng, Z.; Deng, J.; Li, X.; Yin, B.; and Zhang, T. 2025. Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain Alignment. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 2491–2499. Cho, K.; Kim, D. Y.; and Kim, E. 2025. Zero-shot scene change detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 2509–2517. Deng, K.; Zhang, Y.; Yang, J.; and Xie, J. 2025. GigaSLAM: Large-Scale Monocular SLAM with Hierarchical Gaussian Splats. arXiv preprint arXiv:2503.08071. Dobrowski, S. Z.; Safford, H. D.; Cheng, Y. B.; and Ustin, S. L. 2008. Mapping mountain vegetation using species distribution modeling, image-based texture analysis, and
object-based classification. Applied Vegetation Science, 11(4): 499–508. Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; and Koltun, V. 2017. CARLA: An open urban driving simulator. In Conference on robot learning, 1–16. PMLR. Fan, Z.; Zhang, J.; Cong, W.; Wang, P.; Li, R.; Wen, K.; Zhou, S.; Kadambi, A.; Wang, Z.; Xu, D.; Ivanovic, B.; Pavone, M.; and Wang, Y. 2024. Large Spatial Model: Endto-end Unposed Images to Semantic 3D. arXiv:2410.18956. Finney, M. A. 1998. FARSITE, Fire Area Simulator–model development and evaluation. 4. The Station. Finney, M. A. 2006. An overview of FlamMap fire modeling capabilities. In In: Andrews, Patricia L.; Butler, Bret W., comps. 2006. Fuels Management-How to Measure Success: Conference Proceedings. 28-30 March 2006; Portland, OR. Proceedings RMRS-P-41. Fort Collins, CO: US Department of Agriculture, Forest Service, Rocky Mountain Research Station. p. 213-220, volume 41. Fu, X.; Hu, Y.; Sutrave, P.; Beerel, P. A.; and Raghavan, B. 2024. FireLoc: Low-latency Multi-modal Wildfire Geolocation. In Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, 1–14. Geiger, A.; Lenz, P.; Stiller, C.; and Urtasun, R. 2013. Vision meets robotics: The kitti dataset. The international journal of robotics research, 32(11): 1231–1237. Glocker, B.; Izadi, S.; Shotton, J.; and Criminisi, A. 2013. Real-time RGB-D camera relocalization. In 2013 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), 173–179. IEEE. Godard, C.; Mac Aodha, O.; Firman, M.; and Brostow, G. J. 2019. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE/CVF international conference on computer vision, 3828–3838. Gorelick, N.; Hancher, M.; Dixon, M.; Ilyushchenko, S.; Thau, D.; and Moore, R. 2017. Google Earth Engine: Planetary-scale geospatial analysis for everyone. Remote sensing of Environment, 202: 18–27. Hermann, M.; Kwak, H.; Ruf, B.; and Weinmann, M. 2024. Leveraging Neural Radiance Fields for Large-Scale 3D Reconstruction from Aerial Imagery. Remote Sensing, 16(24): 4655. Hirano, A.; Welch, R.; and Lang, H. 2003. Mapping from ASTER stereo image data: DEM validation and accuracy assessment. ISPRS Journal of Photogrammetry and remote sensing, 57(5-6): 356–370. Hodgson, M. E.; Jensen, J. R.; Schmidt, L.; Schill, S.; and Davis, B. 2003. An evaluation of LIDAR-and IFSARderived digital elevation models in leaf-on conditions with USGS Level 1 and Level 2 DEMs. Remote sensing of environment, 84(2): 295–308. Hu, M.; Yin, W.; Zhang, C.; Cai, Z.; Long, X.; Chen, H.; Wang, K.; Yu, G.; Shen, C.; and Shen, S. 2024. Metric3D v2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation. arXiv preprint arXiv:2404.15506.
Iglhaut, J.; Cabo, C.; Puliti, S.; Piermattei, L.; O’Connor, J.; and Rosette, J. 2019. Structure from motion photogrammetry in forestry: A review. Current Forestry Reports, 5(3): 155–168. Jang, W.; Weinzaepfel, P.; Leroy, V.; Agapito, L.; and Revaud, J. 2025. Pow3r: Empowering unconstrained 3d reconstruction with camera and scene priors. In Proceedings of the Computer Vision and Pattern Recognition Conference, 1071–1081. Ke, B.; Qu, K.; Wang, T.; Metzger, N.; Huang, S.; Li, B.; Obukhov, A.; and Schindler, K. 2025. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. arXiv preprint arXiv:2505.09358. Kendall, A.; Grimes, M.; and Cipolla, R. 2015. Posenet: A convolutional network for real-time 6-dof camera relocalization. In Proceedings of the IEEE international conference on computer vision, 2938–2946. Kervyn, M.; Kervyn, F.; Goossens, R.; Rowland, S. K.; and Ernst, G. 2007. Mapping volcanic terrain using highresolution and 3D satellite remote sensing. Krishnan, S.; Crosby, C.; Nandigam, V.; Phan, M.; Cowart, C.; Baru, C.; and Arrowsmith, R. 2011. OpenTopography: a services oriented architecture for community access to LIDAR topography. In Proceedings of the 2nd international conference on computing for Geospatial Research & Applications, 1–8. Leroy, V.; Cabon, Y.; and Revaud, J. 2024. Grounding Image Matching in 3D with MASt3R. Li, B.; Weinberger, K. Q.; Belongie, S.; Koltun, V.; and Ranftl, R. 2022a. Language-driven Semantic Segmentation. In International Conference on Learning Representations. Li, J.; and Lee, G. H. 2021. DeepI2P: Image-to-point cloud registration via deep classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15960–15969. Li, Y.; Yu, A. W.; Meng, T.; Caine, B.; Ngiam, J.; Peng, D.; Shen, J.; Lu, Y.; Zhou, D.; Le, Q. V.; et al. 2022b. Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 17182–17191. Li, Z.; and Snavely, N. 2018. Megadepth: Learning singleview depth prediction from internet photos. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2041–2050. Liu, X.; Li, D.; and He, Y. 2022. Random Mapping Method for Large-Scale Terrain Modeling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 5395–5403. Lowe, D. G. 2004. Distinctive image features from scaleinvariant keypoints. International journal of computer vision, 60: 91–110. Lu, C.; Yin, F.; Chen, X.; Liu, W.; Chen, T.; Yu, G.; and Fan, J. 2023. A large-scale outdoor multi-modal dataset and benchmark for novel view synthesis and implicit scene reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 7557–7567.
Mall, U.; Hariharan, B.; and Bala, K. 2023. Change-Aware Sampling and Contrastive Learning for Satellite Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5261–5270. Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1): 99–106. Mudashiru, R. B.; Sabtu, N.; Abustan, I.; and Balogun, W. 2021. Flood hazard mapping methods: A review. Journal of hydrology, 603: 126846. Mukherjee, S.; Joshi, P. K.; Mukherjee, S.; Ghosh, A.; Garg, R.; and Mukhopadhyay, A. 2013. Evaluation of vertical accuracy of open source Digital Elevation Model (DEM). International Journal of Applied Earth Observation and Geoinformation, 21: 205–217. Mur-Artal, R.; Montiel, J. M. M.; and Tardos, J. D. 2015. ORB-SLAM: a versatile and accurate monocular SLAM system. IEEE transactions on robotics, 31(5): 1147–1163. Nguyen, T. T.; Slaughter, D. C.; Max, N.; Maloof, J. N.; and Sinha, N. 2015. Structured light-based 3D reconstruction system for plants. Sensors, 15(8): 18587–18612. OpenTopography. ???? OpenTopography: High-Resolution Topography Data and Tools. https://opentopography.org/ start. Accessed: 2025-07-12. Panek, V.; Kukelova, Z.; and Sattler, T. 2022. Meshloc: Mesh-based visual localization. In European Conference on Computer Vision, 589–609. Springer. Piccinelli, L.; Sakaridis, C.; Yang, Y.-H.; Segu, M.; Li, S.; Abbeloos, W.; and Gool, L. V. 2025. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler. arXiv:2502.20110. Piccinelli, L.; Yang, Y.-H.; Sakaridis, C.; Segu, M.; Li, S.; Van Gool, L.; and Yu, F. 2024. UniDepth: Universal Monocular Metric Depth Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Qadri, M.; and Kantor, G. 2021. Semantic feature matching for robust mapping in agriculture. arXiv preprint arXiv:2107.04178. Rodriguez, J. J.; and Aggarwal, J. 2002. Matching aerial images to 3-D terrain maps. IEEE Transactions on pattern analysis and machine intelligence, 12(12): 1138–1149. Rollins, M. G. 2009. LANDFIRE: a nationally consistent vegetation, wildland fire, and fuel assessment. International Journal of Wildland Fire, 18(3): 235–249. Rublee, E.; Rabaud, V.; Konolige, K.; and Bradski, G. 2011. ORB: An efficient alternative to SIFT or SURF. In 2011 International conference on computer vision, 2564–2571. Ieee. Rusu, R. B.; and Cousins, S. 2011. 3d is here: Point cloud library (pcl). In 2011 IEEE international conference on robotics and automation, 1–4. IEEE. Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; and Rabinovich, A. 2020. Superglue: Learning feature matching with graph
neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4938– 4947. Schmidt, F.; Daubermann, J.; Mitschke, M.; Blessing, C.; Meyer, S.; Enzweiler, M.; and Valada, A. 2025. ROVER: A Multiseason Dataset for Visual SLAM. IEEE Transactions on Robotics, 41: 4005–4022. Schönberger, J. L.; Zheng, E.; Pollefeys, M.; and Frahm, J.M. 2016. Pixelwise View Selection for Unstructured MultiView Stereo. In European Conference on Computer Vision (ECCV). Shah, S.; Dey, D.; Lovett, C.; and Kapoor, A. 2017. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and service robotics: Results of the 11th international conference, 621–635. Springer. Sidle, R. 2017. Strengths and Limitations of UAV and Ground-based Structure from Motion Photogrammetry in a Gullied Savanna Catchment. Sun, J.; Shen, Z.; Wang, Y.; Bao, H.; and Zhou, X. 2021. LoFTR: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8922–8931. Turki, H.; Ramanan, D.; and Satyanarayanan, M. 2022. Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12922– 12931. Wang, S.; Leroy, V.; Cabon, Y.; Chidlovskii, B.; and Revaud, J. 2024. DUSt3R: Geometric 3D Vision Made Easy. In CVPR. Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4): 600–612. Wu, H.; Wen, C.; Shi, S.; Li, X.; and Wang, C. 2023. Virtual Sparse Convolution for Multimodal 3D Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 21653–21662. Wu, Z.; Song, S.; Khosla, A.; Yu, F.; Zhang, L.; Tang, X.; and Xiao, J. 2015. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1912–1920. Yang, L.; Kang, B.; Huang, Z.; Xu, X.; Feng, J.; and Zhao, H. 2024a. Depth Anything: Unleashing the Power of LargeScale Unlabeled Data. In CVPR. Yang, L.; Kang, B.; Huang, Z.; Zhao, Z.; Xu, X.; Feng, J.; and Zhao, H. 2024b. Depth Anything V2. arXiv:2406.09414. Yang, L.; Meng, X.; and Zhang, X. 2011. SRTM DEM and its application advances. International Journal of Remote Sensing, 32(14): 3875–3896. Yu, H.; Zhen, W.; Yang, W.; Zhang, J.; and Scherer, S. 2020. Monocular camera localization in prior lidar maps with 2d-3d line correspondences. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4588–4594. IEEE.
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, 586–595. Zhou, J.; Ma, B.; Zhang, W.; Fang, Y.; Liu, Y.-S.; and Han, Z. 2023. Differentiable registration of images and lidar point clouds with voxelpoint-to-pixel matching. Advances in Neural Information Processing Systems, 36: 51166–51177.