Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation Sri Hrushikesh Varma Bhupathiraju
arXiv:2609.16336v1 [cs.CR] 14 Sep 2026
University of Florida, Gainesville, USA [email protected]
Tetsu Ishizue
Nicholas U. Costagliola
Ozora Sako
The University of Electro-Communications, Tokyo, Japan [email protected]
University of Florida, Gainesville, USA [email protected]
Keio University, Tokyo, Japan [email protected]
Kentaro Yoshioka
Takeshi Sugawara
Sara Rampazzi
Keio University, Tokyo, Japan [email protected]
The University of Electro-Communications, Tokyo, Japan [email protected]
University of Florida, Gainesville, USA [email protected] Repeated pattern
Abstract Stereo cameras are integrated into autonomous systems such as self-driving cars, drones, and robots to offer precise depth estimation in a cost-effective manner compared to LiDAR technology. In this work, we reveal an intrinsic vulnerability in stereo cameras that stems from their pixel sampling and calibration processes, which can influence the outputs of stereo matching algorithms. Attackers can achieve fine-grained control over the estimated depth of real obstacles using simple repeating patterns, without relying on sophisticated adversarial machine learning techniques. Furthermore, deep learning-based depth estimation models exhibit a similar vulnerability. We evaluate the impact of this attack on two widely used stereo matching algorithms (BM and SGBM), three deep learning models (PSMNet, MoCha-Stereo, and UniMatch), a stereo-LiDAR fusion model (SGM-DDC), and two popular commercial stereo cameras, the ZED2 and Intel RealSense D435. For example, in the ZED2 camera, an attacker can displace obstacles up to 20 meters farther or 12 meters closer. In our real-world evaluation in a driving setting, a brief 0.5 second attack can trigger emergency braking in a popular autonomous driving framework. We further demonstrate the feasibility at driving speeds up to 40 km/h using CARLA. Finally, we confirm the ineffectiveness of state-of-the-art defenses, and we propose a novel strategy that leverages similarity scores to dynamically detect and suppress the depth discrepancies. Our work highlights vulnerabilities hidden in stereo matching and deep learning depth estimation models, addressing critical limitations in autonomous system deployments.
CCS Concepts • Security and privacy → Hardware attacks and countermeasures; • Computer systems organization → Embedded and cyber-physical systems.
This work is licensed under a Creative Commons Attribution 4.0 International License. CCS ’26, The Hague, Netherlands © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2871-6/2026/11 https://doi.org/10.1145/3830454.3832649
Induced fake depth
Repeated pattern
Victim stereo cameras mounted on UGV
Fake depth estimated by stereo camera Victim AV with stereo camera
Figure 1: Our attack induces controlled shifts in depth estimation of stereo cameras by exploiting an inherent vulnerability triggered by simple patterns with repeating elements.
Keywords Cyber-Physical Systems, Physical Attacks, Autonomous Systems, Stereo Cameras
1
Introduction
In the evolving landscape of autonomous systems, stereo vision has emerged as a passive, lightweight, and cost-effective alternative solution to LiDARs [24, 46]. These sensors mimic human depth perception using dual viewpoints and triangulation, delivering 3D data without sophisticated lasers. For instance, companies such as NODAR [46] and Dolunts [21] have shown their capabilities in drones and self-driving vehicles under challenging conditions, such as foggy environments, while reducing hardware costs by an order of magnitude [11]. At the heart of these systems lie Stereo Depth Estimation (SDE) algorithms and their deep-learning variants, which fuse image data into depth maps of the surrounding environment. Classical matching-based SDE algorithms, such as Block Matching [26] (BM) and Semi-Global Block Matching [51] (SGBM), estimate depth by comparing small blocks of pixels between two images. A block from the reference image is horizontally scanned across the other image, and the position with the best similarity is chosen as the match. The horizontal shift, called disparity, is then used to compute depth. Deep learning-based SDE models further extend this by learning complex pixel and feature relationships to achieve more accurate and robust disparity predictions [15, 62].
CCS ’26, November 15–19, 2026, The Hague, Netherlands
The literature has examined adversarial machine learning attacks designed to introduce depth estimation errors in deep learning SDE models [16, 18, 40, 57, 59]. Specifically, recent studies have shown that patterns containing repeated elements, such as stripes, waves, and concentric shapes, can increase attack success rates. These works exploit such periodic structures to generate adversarial examples on benchmark datasets [6, 39]. However, despite their ability to induce errors, these works neither explain the underlying cause of the phenomenon nor demonstrate control over the resulting depth, restricting their applicability to specific scenarios, complex optimization, and certian deep learning-based SDE models. This study aims to answer the following critical research questions: What vulnerability causes depth estimation errors in SDE block matching algorithms and deep-learning models in the presence of repeated elements? Can an adversary exploit such factors to induce consistent and controlled depth estimation manipulation? Here, we uncover a hidden vulnerability rooted in two fundamental properties of stereo cameras: sampling artifacts and calibration errors. Sampling artifacts are pixel-level distortions introduced during image discretization that alter pixel intensities and local features, while calibration errors are pixel-level inaccuracies and nonlinearities caused by lens distortions and manufacturing variations that warp features used for stereo matching. Our study demonstrates that these vulnerabilities can be exploited to induce incorrect depth estimations, allowing adversaries to systematically bias local matching scores by placing simple patterns with repeating elements in the scene and thereby controlling the resulting depth estimations, as illustrated in Figure 1. These artifacts alter the image pairs at the pixel level, ultimately affecting how SDE algorithms perform their matching. Moreover, we find that this phenomenon transfers to deep learning-based SDE models: they inadvertently learn and replicate the same vulnerability, making them equally susceptible. Based on this, we design a projection attack that controllably manipulates the depth estimation of both SDE algorithms and deep learning models without any sophisticated adversarial optimization or tracking. We systematically formalize and characterize how these vulnerabilities affect depth matching in both classical and deep learning SDE algorithms, with extensive evaluation in diverse scenarios. Specifically, we evaluate the attack on two classical algorithms (BM and SGBM [51]), three deep learning models (PSMNet [15], MoChaStereo [17], and UniMatch [62]) and one stereo-LiDAR fusion-based model (SGM-DDC [64]). For example, BM can be manipulated to shift obstacles up to 19m closer or 5m farther away from the victim stereo camera, with fine-grained control at 0.1m resolution. Similarly, the attack succeeds on two commercial stereo cameras widely used in autonomous systems (ZED2 [52] and RealSense D435 [32]), in realistic driving scenarios, with vehicle speeds up to 15 km/h. We further demonstrate the feasibility of the attack in high-speed scenarios of up to 40 km/h in the CARLA simulator [22]. The attack persists for at least 0.5 sec, sufficient to trigger emergency braking or unsafe maneuvers in state-of-the-art autonomous driving frameworks [55]. While prior work only suggests adversarial training as a potential defense against adversarial examples and impractical manual parameter tuning for reducing depth errors, we propose and evaluate a novel defense technique that detects potentially dangerous
Sri Hrushikesh Varma Bhupathiraju et al.
repeating patterns in the scene and mitigates the resulting depth shift. The approach analyzes the matching scores of local image regions, suppressing the vulnerability effects. In summary, our study makes the following contributions: • The discovery of a hidden vulnerability in stereo cameras arising from intrinsic characteristics of stereo vision systems, namely (i) sampling artifacts and (ii) calibration errors. The vulnerability can be exploited using simple patterns to control depth estimation on classical block matching and deep learning-based models, which unexpectedly learn and replicate the same vulnerability. • An extensive analysis of the attacker capability to manipulate two SDE algorithms (BM and SGBM), three deep learningbased SDE models (PSMNet, MoCha-Stereo, and UniMatch) and a stereo-LiDAR fusion model (SGM-DDC). For example, in synthesized scenarios, an attacker can cause depth estimation errors of up to 58.9m in deep-learning models, producing fake depth estimates as close as 3.7m from the victim camera. • Vulnerability evaluation in real-world scenarios of two popular stereo cameras, ZED2 and RealSense D435, achieving depth shifts > 12m. Our attack produces consistent depth estimation errors lasting ≥ 0.5 seconds in real-world driving scenarios at speeds up to 15 km/h, sufficient to trigger emergency braking in autonomous driving systems.1 • A novel defense methodology that overcomes the limitations of past approaches and detects structured repeated patterns in the scene, mitigating depth estimation errors with a success rate of 96.5% and 100% for classical and deeplearning-based SDE algorithms, respectively.
2 Background and Previous Work 2.1 Depth Estimation in Stereo Cameras Stereo-based depth estimation (SDE) reconstructs a 3D depth map by comparing a pair of 2D images captured from two horizontally aligned cameras, mimicking human binocular vision [54]. Unlike monocular depth estimation systems [7] which infer depth from scale, perspective, or motion, SDEs use disparity, the horizontal offset between corresponding regions in the left and right images. This disparity is then converted into depth using geometric triangulation. SDE algorithms are widely used in autonomous systems [20, 30, 46] for accurate obstacle detection and avoidance. Our study focuses on the two main categories of SDE: (i) Block matching methods [38, 51], which compute disparity by evaluating local pixel similarity across a predefined window, and (ii) Deep learning-based methods [15, 61, 62], which leverage high-level semantic features and learned representations via machine-learning algorithms, to establish these correspondences. While the first category is typically adopted in resource-constrained environments, such as drones and mobile robotics, the second achieves higher accuracy at the cost of increased computational demands, making it more suitable for autonomous driving applications. 1 Details and videos are available at https://cpseclab.github.io/stereocam/. Artifacts
available at https://zenodo.org/records/20767261. See the Open Policy section for details.
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
Real world object
Disparity (D)
Disparity range
CCS ’26, November 15–19, 2026, The Hague, Netherlands Victim AV with stereo camera
Victim AV with stereo camera Induced fake depth
de
de da
(xLi, yLi)
Left image
(xRi+D, yRi+D)
(xRi, yRi)
Right image
Figure 2: Depth estimation in stereo cameras. For each block in the left image (𝑥𝑖𝐿 , 𝑦𝑖𝐿 ), classical block matching algorithms start from the same position in the right image (𝑥𝑖𝑅 , 𝑦𝑖𝑅 ) and 𝑅 , 𝑦 𝑅 ) at disparity 𝐷, search for the most similar block (𝑥𝑖+𝐷 𝑖+𝐷 within the disparity range.
Block matching methods. In Block Matching (BM) algorithms for stereo cameras, an image (typically the left one) is divided into small rectangular regions, or blocks. For a block in the left image of coordinates (𝑥𝑖𝐿 , 𝑦𝑖𝐿 ), the algorithm examines the same coordinates (𝑥𝑖𝑅 , 𝑦𝑖𝑅 ) in the right image, searching leftward along the same horizontal (epipolar) line for the best block match. The disparity, as illustrated in Figure 2, is obtained by identifying the most similar block within a defined disparity range. The displacement between the reference block (𝑥𝑖𝑅 , 𝑦𝑖𝑅 ) and the best-matching 𝑅 , 𝑦 𝑅 ) defines the disparity value 𝐷, which is subseblock (𝑥𝑖+𝐷 𝑖+𝐷 quently used to estimate the depth. The similarity between two blocks is typically evaluated using metrics such as the Sum of Absolute Differences (SAD) or Normalized Cross-Correlation (NCC) [29]. Semi-Global Block Matching (SGBM) [38, 51] methods are enhanced variations of BM that aggregate matching costs along multiple 1D paths. BM, SGBM, and their variants are widely used in commercial stereo cameras for drones and robots, such as ZED2 [52] and Intel RealSense [32] examined in this work. Deep learning-based methods. Deep learning–based SDE models replace the classical matching algorithms, with feature extraction and cost-volume networks that infer disparity end-to-end from the stereo image pairs. Typical model architectures include encoderdecoder or convolutional neural networks (CNNs), and attention mechanisms to give more emphasis to certain correspondences (e.g., edges, textures) and less on ambiguous or noisy regions (e.g., flat surfaces or occlusions) [28]. In this work, we analyze three models as representative of three different architectures: PSMNet, MoCha-Stereo, and UniMatch.
2.2
Attacks against Stereo Depth Estimation
Researchers have long studied how to induce fake depths to SDEs by generating adversarial examples against their machine learning models. Typically, they apply structured pixel-level digital noise or adversarial patches to stereo image pairs to produce erroneous depth measurements [6, 27, 53, 57, 59]. Although these methodologies introduce uncontrolled depth errors, they are highly modelspecific, and their applicability in real-world scenarios remains limited because they require direct manipulation of the images or the autonomous system to implement pixel-level modifications.
Induced fake depth
(a) Displaying pattern on billboard
da (b) Projecting pattern on obstacle
Figure 3: Illustration of two attack scenarios explored in this work: the projected pattern on a billboard and the back of a van. The attacker controls the fake depth 𝑑𝑒 within the frustum of the attack pattern region placed at a distance 𝑑𝑎 . In contrast, Liu et al. [40] created physically realizable adversarial printed patches, leveraging full white-box access to the deep learning (DL) model under attack. Liu et al. [39] and Berger et al. [6] instead observed that the presence of repeated elements in the adversarial patches (e.g., stripe, wavy, and concentric designs) increased depth estimation errors. Building on this observation, they generate printed adversarial repeated elements still optimized against specific DL models. Both these works primarily attribute this phenomenon to the susceptibility of the models to adversarially optimized patterns, without examining why such phenomena occur in the first place. In addition, while all these works succeed on maximizing the depth errors in single stereo image pairs and DL models, they fail to ensure control over the magnitude and direction of the induced errors and their consistency over time and movement, forcing a careful patch crafting, and drastically limiting the attack success in real-world scenarios, like, for example, placing printed posters in front of the car trajectory to succeed. A separate branch of research focuses instead on stereo camera sensor attacks to induce depth errors. For instance, Zhou et al. [66] and Fu et al. [25] use light injection to create glare that drones interpret as obstacles. These attacks require one or more light sources to be perfectly aligned with the camera lenses and apply continuous injection to create the glare. While these attacks have been shown to succeed in real-world testing, the success depends on the lens setup of the victim stereo camera, and there is no control over the shape of the glare or the created obstacle. Contrary to these works, we demonstrate control over depth estimation of genuine objects under real-world moving conditions, and without the use of adversarial optimization.
3 Adversary Model and Vulnerability 3.1 Threat Model Attacker Goal. We consider stereo cameras deployed in autonomous vehicles (AVs) as a primary target of the attack, as shown in Figure 3. Different from previous work, which focuses only on maximizing the depth error for a given image pair and deep learning model, here the goal of the attacker is to mislead the victim’s AV’s stereobased perception, either based on classical block matching or deep learning, by consistently controlling the estimated depth of target object surfaces present in the driving scene. Underlying Principle and Scenario. Instead of generating physical optimized adversarial patches, the adversary can project using a commercial projector simple repeating geometric shapes, such as squares or stripes, onto object surfaces and locations (e.g., buildings,
CCS ’26, November 15–19, 2026, The Hague, Netherlands
traffic signs, backs of cars or vans) or by displaying them on digital billboards placed at the roadside or intersections (see Figure 3) for a short time duration (e.g., 0.5 seconds) to activate the attacker’s desired outcome. Such scenarios are in line with previous works targeting RGB cameras [43, 48, 67]. Simply by regulating the geometric structures’ shape and size, it is possible to trigger underlying non-linearities and artifacts, forcing disparity mismatch to specific values. In this way, the attacker can manipulate the depth estimation in a controlled and consistent manner without the need for complex depth-specific optimization per scene. This manipulation causes objects to be perceived as closer or farther than their actual distance from the victim AV. For example, a vehicle may fail to brake and collide by overestimating the distance to an obstacle, whereas underestimating the distance can trigger unnecessary emergency braking or abrupt maneuvers, compromising user safety. Assumptions. We assume that the adversary has knowledge of the type of stereo camera used by the victim AV and the type of the SDE algorithm (i.e., BM or deep learning-based model), the corresponding algorithmic parameters (e.g, disparity range, block size, stabilization, and confidence thresholds), and the calibration settings. This information can be inferred from publicly available camera manuals and data sheets, like the ones analyzed in this work [32, 52]. While proprietary algorithms are not fully disclosed, stereocamera vendors publish technical specifications (window size, disparity range, etc.) to comply with standards and interoperability requirements. Alternatively, the attacker can perform a black-box study by acquiring a stereo camera similar to the one used in the victim AV. Based on this assumption, we perform the attack on two widely used commercial stereo cameras, ZED2 [52] and RealSense D435 [32], both of which employ proprietary SDE algorithms. The attack is remote, not requiring any firmware or hardware access to the victim vehicle or stereo camera. We do assume oracle access to the depth estimation algorithm employed by the victim stereo camera, consistent with the assumptions from previous works [6, 39]. Finally, we focus on easy, realizable, low-effort repeated geometric shapes (e.g., simple stripes or chessboard patterns) that an unskilled attacker can project over flat or semi-flat surfaces without requiring sophisticated optimization, white-box model access, or the generation of adversarial examples. An adversary can select more sophisticated geometries to account for the specific structure of the object surface to project on. First we synthesize such scenarios on the KITTI dataset [26] and evaluate classical and deep learning-based SDE models in Section 5. Real-world dynamic evaluation in indoor and outdoor settings are presented in Section 6, using commercial RealSense and ZED2 stereo cameras.
3.2
Vulnerability in Block Matching Algorithms
This section describes the principles that enables attackers to precisely manipulate perceived depth by exploiting (i) sampling artifacts and (ii) calibration errors in stereo cameras, which are empirically validated on the BM and the ZED2 SGBM algorithms.
Sri Hrushikesh Varma Bhupathiraju et al.
Intermediate pixel value due to pixel sampling
Displacement in corresponding pixels due to calibration error
(a) Sampling Artifact
(b) Calibration Error
Figure 4: (a) Sampling artifacts produce intermediate pixel intensity values at the stripe boundaries due to averaging across the pixel grid in our Gazebo simulation2 . (b) A redcyan anaglyph of the stereo pair reveals calibration-induced geometric distortion: the cyan (left image) and red (right image) misalignments cause spatially varying shifts in disparity across a checkerboard pattern. Left Image Attack pattern synthesized in Gazebo
Right Image
dw
Simulated stereo camera
Incorrect blocks matched due to similar sampling artifacts
Ground truth block
Figure 5: Gazebo simulation scenario with stripe pattern in front of stereo camera (left). Illustration of the incorrect block matching triggered by sampling artifacts (right). 3.2.1 Sampling Artifacts. Image sensors in stereo cameras capture continuous physical scenes and represent them with discrete pixels [42]. The light intensity recorded at each pixel represents the average optical light intensity from that region of the scene. Given a physical point 𝑝 = (𝑥, 𝑦, 𝑧) and its optical light intensity 𝐼 opt (𝑝), the corresponding pixel intensity value 𝑃𝑖 is denoted by ∫ 𝑆𝑅𝑖+1 1 𝑃𝑖 = 𝐼 opt (𝑝) d𝐴 (1) 𝐴(𝑆𝑅𝑖 ) 𝑆𝑅𝑖 where 𝑆𝑅𝑖 is the sampling range of the 𝑖 𝑡ℎ pixel in the horizontal axis, defined as the minimum real-world distance represented by a single pixel 3 . Assuming square pixels, the physical area 𝐴 corresponding to the 𝑖 𝑡ℎ pixel is given by 𝐴(𝑆𝑅𝑖 ) = 𝑆𝑅𝑖2 . When the stereo camera captures images with high-contrast boundaries (e.g., object edges) in the physical space, each image sensor captures and discretizes the boundary, resulting in subtle pixel-level inconsistencies. The pixel intensity values in those particular regions consequently assumes an intermediate value corresponding to the average intensity of the region sampled by the stereo camera, as illustrated in Figure 4 (a). These intermediate values, referred to as sampling artifacts, are inherent to imaging processes [42] and can produce a measurable disparity offset in SDE algorithms and, consequently, changes in the estimated depth.
2 Due to grayscale conversion in Gazebo visualization rendering, the pattern appears
3 For this formalization, we restrict our analysis to the horizontal axis, consistent with
as gray and black stripes instead of black and white.
the matching direction employed in classical block matching algorithms [47]
SAD(𝐵(𝑖 Ref ), 𝐵(𝑖𝑇 + 𝑛 · 𝑑 𝑤 )) ∝ |𝑆𝐴(𝑖 Ref ) − 𝑆𝐴(𝑖𝑇 + 𝑛 · 𝑑 𝑤 )| (2) Where 𝑛 · 𝑑 𝑤 represents a horizontal shift by 𝑛 repetitions of the pattern, with 𝑛 ∈ {1, 2, · · · }. Furthermore, each repeated element in the pattern yields a similar SAD score because it reappears every 𝑑 𝑤 pixels. Thus, the BM algorithm selects the repeated block whose sampling artifact magnitude (and consequently SAD minima) most closely matches that of the reference block, rather than necessarily the true geometric match, producing an incorrect disparity estimate. Appendix A provides a detailed formulation of the vulnerability. Since 𝑆𝐴(𝑖) varies periodically with the repeating pattern, an adversary can exploit this predictable behavior to manipulate the location of the SAD minima and systematically bias stereo depth estimation. The adversary can simply craft repeated geometric elements in the pattern whose distances (e.g., adjusting 𝑑 𝑤 ) produce the sampling artifacts necessary to shift the estimated depth to the desired value and trigger controlled matching of selected regions. Real-World Validation. We further validate these findings using real-world images captured with a ZED2 stereo camera, with the repeated stripe pattern displayed on a monitor, placed at 𝑑𝑎 = 1m in controlled indoor conditions (𝑑 𝑤 = 12mm as in the simulation). Despite the pixel-level noise in the stereo image pairs, the minima in the normalized SAD scores vary according to the sampling artifact magnitude, consistent with the simulation results, as illustrated in Figure 6 (right). This experiment confirms that sampling artifacts can be used to control disparity estimation.
SAD score
SAD minima
Pixel Disparity
Normalized SAD Score
Formalization. In scenes containing repeated high-contrast boundaries, such as a black-white chessboard pattern, the sampling artifacts exhibit differences in both pixel intensity and spatial position. Although the physical pattern does not change, each repetition of the geometric shape is captured differently depending on the boundary’s alignment with the left and right cameras. To examine this effect, we employ a classical BM algorithm that uses the SAD score, as described in Section 2.1, to estimate disparities. The SAD score is computed by summing the absolute differences between corresponding pixel intensities in two image blocks (𝐵(𝑖)), providing a measure of their similarity. We synthesize a simple stripe pattern in the popular Gazebo simulation environment [36] and capture stereo images using an emulated ZED2 stereo camera, as shown in Figure 5 (left). Figure 5 (right) illustrates the left and right images of the stripe pattern synthesized in Gazebo. 𝐵(𝑖 Ref ) denotes the reference block in the left image, and 𝐵(𝑖𝑇 ) the true corresponding block in the right image, both containing the artifacts (gray pixels). The red block marks the region with a matching sampling artifact that is incorrectly paired with the left image. The BM algorithm selects the global minimum in the SAD score as the optimal match. Under natural conditions, the algorithm converges at the minimum matching score of SAD(𝐵(𝑖 Ref ), 𝐵(𝑖𝑇 )), finding the perfect match. However, the high-contrast alternate pattern and resulting artifacts produce oscillating SAD scores, resulting in multiple SAD minima, each of which is a potential match, as shown in Figure 6 (left). Let 𝑆𝐴(𝑖) denote the sampling artifact magnitude at the block position 𝑖, estimated from the pixel intensity. Based on this, the SAD score at the minima is proportional to the difference in pixel intensities, described as:
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Normalized SAD Score
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
Simulated
Real World
Pixel Disparity
Figure 6: Normalized SAD scores following an oscillating trend due to the black and white stripe pattern (left). The SAD minima in simulated and real-world scenarios (right). The normalized SAD scores are scaled in the (right) graph to better illustrate the SAD minima.
3.2.2 Calibration Error. SDE algorithms use calibrated image pairs for disparity estimation [54]. Stereo camera vendors typically perform a factory calibration process to determine the parameters to pre-process the raw camera images. This critical process establishes a precise mapping between a 3D point in the physical world and its corresponding 2D projection in the image, compensating for imperfections such as lens aberration or misalignment between the lens and the sensor plane [1]. However, unavoidable nonlinearities cause pixel correspondence errors that persist in the reconstructed depth map [60]. These distortions arise because corresponding pixels in the left and right images are displaced in different directions, resulting in mismatches along both the horizontal and vertical axes, as shown in Figure 4 (b). Although such calibration errors are inherent to stereo systems, they usually have little impact in benign cases. Yet, in the presence of high-contrast boundaries such as the ones present in repeated patterns, the calibration errors warp across each repeated element, causing the correct match to appear structurally different in the image pairs, and consequently trigger an incorrect match. The resulting mapping depends on the camera’s parameters (such as focal length, principal point, and translational and rotational matrices) and magnitude of the calibration error. Formalization. To examine how calibration errors cause depth estimation inaccuracies, we model the calibration error of our ZED2 camera using the fundamental parameterization of radial and tangential distortions [5], and the calibration models provided by the ZED2 [52] vendor. The parameterization is as follows: 𝑥𝑑 𝑥 = 𝑢 1 + 𝑘 1𝑟 2 + 𝑘 2𝑟 4 + 𝑘 3𝑟 6 𝑦𝑑 𝑦𝑢 2𝑎 1𝑥𝑢 𝑦𝑢 + 𝑎 2 (𝑟 2 + 2𝑥𝑢2 ) + 𝑎 1 (𝑟 2 + 2𝑦𝑢2 ) + 2𝑎 2𝑥𝑢 𝑦𝑢
(3)
Here, (𝑥𝑑 , 𝑦𝑑 ) and (𝑥𝑢 , 𝑦𝑢 ) denote the pixel coordinates of the distorted and undistorted images, respectively. The distortion vector 𝐶 = (𝑘 1, 𝑘 2, 𝑎 1, 𝑎 2, 𝑘 3 )𝑇 comprises the radial distortion parameters (𝑘 1, 𝑘 2, 𝑘 3 ) and tangential distortion parameters (𝑎 1, 𝑎 2 ), and 𝑟 represents the radial distance of the image pixel. Using the standard stereo calibration technique [1], we estimate the distortion parameters for the ZED2 cameras, which are then applied to Equation 3. Using this formulation, we quantify the residual calibration error that remains in the camera after factory calibration and rectification. The resulting distortion vectors are 𝐶𝐿 = (-0.0415, 0.0099, 0.00035, -0.00013, -0.00492)𝑇 and 𝐶𝑅 = (-0.0424,
0.0112, 0.00015, 9.69×10−5 , -0.00509)𝑇 , where 𝐶𝐿 and 𝐶𝑅 are the displacement vectors for the left and right cameras, respectively. This calibration error causes pixel shift as shown in Figure 4 (b). The calibration error magnitude (𝐸𝐶 ) is estimated 3.7 pixels in the ZED2 camera, derived by the maximum distance between points (𝑥𝑑 , 𝑦𝑑 ) and (𝑥𝑢 , 𝑦𝑢 ) in Equation 3. This parameterization captures semantic-level calibration errors but cannot account for pixel-level calibration inaccuracies. Calibration is based on finite mathematical approximations that inherently fail to model stochastic noise, non-planar surface imperfections, and higher-order lens irregularities. A 3 pixel calibration error is generally considered acceptable for a commercial stereo camera [1]. Nevertheless, this level of error is sufficient to cause depth estimation errors in the presence of repeated elements. To evaluate the effect of calibration error on BM, we use the same Gazebo simulation setup of the sampling artifact simulation. In this case, we employ a chessboard pattern, since calibration errors occur along both the horizontal and vertical axes. We define calibration error magnitude (𝐸𝐶 ) as the largest pixel shift induced by calibration error, ranging from 0-15 pixels at intervals of 3 (up to five times the estimated calibration error magnitude of ZED2, 3.7 pixels). This range in calibration error is used to describe the relationship between calibration error and depth estimation. Consistent with the methodology in Section 3.2.1, we set 𝑑 𝑤 = 12mm to ensure the pattern is perceptible to the ZED2 camera. Unlike sampling artifacts, which influence the minima of the SAD scores, calibration errors affect the global score matching of the BM algorithm. For this reason, we consider the mean of the estimated disparity within the chessboard pattern as our metric. The estimated mean disparity values at increasing calibration error magnitudes are shown in Figure 7. As hypothesized, even at 𝐸𝐶 = 3 pixels the estimated mean disparity doubles, increasing from 9.9 at zero error to 22.8 in our simulation. The disparity continues to rise until an error magnitude of 12, after which it plateaus. This plateau occurs because the BM algorithm’s disparity range (default value of 64) caps the estimation, preventing further increase and consequently flattening the average value. An adversary can select a pattern with 𝑑 𝑤 such that the perceived pixel size of its geometric shapes is close to 𝐸𝐶 and increase the estimated disparity. These results demonstrate how calibration errors can be exploited to induce incorrect block matching. Real-World Validation. We validate our simulation results with real-world experiments, using the same setup as in Section 3.2.1. The chessboard pattern displayed on a monitor at 𝑑𝑎 =1m is captured with the ZED2 camera, starting from 𝐸𝐶 = 3 up to 15 pixels. As illustrated in Figure 7, we observe a similar trend with the mean disparity increases until magnitude 12 and then plateaus following the disparity range value of the ZED2 SGBM algorithm. Unlike simulation, the mean disparity values reach higher values (up to 63), indicating that pixel imperfections in real stereo images further amplify the effect.
3.3
Vulnerability in Deep Learning Models
Deep learning-based SDE models employ convolutional, encoderdecoder, or attention-based architectures to learn and predict correspondences between high-level features extracted from the left and
Sri Hrushikesh Varma Bhupathiraju et al. Mean Estimated Disparity
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Simulated
Real World
Calibration Error Magnitude (pixels)
Figure 7: Estimated disparity variation at increasing calibration error in simulated and real-world conditions. Note that ZED2 has a calibration error of 3 pixels by design. right images, as described in Section 2.1. In contrast with classical block matching algorithms, they are generally trained to estimate the matching cost for each pixel in the full disparity range, followed by cost aggregation techniques that integrate local and global information to suppress noise in the disparity distribution [15, 39, 61]. Here, we experimentally verify the presence and effect of the discovered vulnerability in one of the most popular depth estimation models used in autonomous vehicles, PSMNet [15] (full characterization is detailed in Section 2.1). To achieve this, we train the model using the standard autonomous driving KITTI [26] dataset under default parameters. By following the Gazebo formalization described in Section 3.2, we synthesize stripe and chessboard patterns with sampling and calibration artifacts. The synthesized patterns are then embedded into 100 randomly chosen stereo image pairs from the KITTI dataset, emulating a pattern projected on the back surface of a van or other obstacle, 10m in front of the victim vehicle stereo camera (equivalent to ≈ 18% of the image pairs with a true disparity of 31). The patterns are luminance-matched for each KITTI image by computing the image’s average brightness over the whole scene (calculated as luma [33]) and shifting the pattern’s average luminance to match the KITTI scenes. 3.3.1 Sampling Artifacts. To demonstrate if sampling artifacts affect the depth estimation in PSMNet, we introduce the artifacts at selected stripes within the pattern, beginning from the ground truth location (Position 1), as shown in Figure 8. The sampling artifact magnitude (pixel intensity) in the left image at Position 1 is fixed, while in the right image, the artifact is iteratively shifted in the next stripe to the left. This shift causes the corresponding maximum disparity to increase following the artifact position shift, as seen in the corresponding depth maps. Figure 9 (left) shows that the average maximum predicted disparity increases over the 100 images, as the sampling artifact is shifted farther from the ground-truth position (from Position 1 to 4). These results demonstrate that the feature correspondences learned by the ML models are directly influenced by the location and magnitude of the sampling artifacts. Consequently, incorrect depth is estimated whenever a sampling artifact of similar magnitude appears. We further validate the influence of the sampling artifact pixel intensity on the disparity estimation of PSMNet in Appendix B. 3.3.2 Calibration Error. As with the sampling artifacts, we conduct a similar evaluation to demonstrate that calibration errors also influence the depth estimation in PSMNet. We consider a chessboard pattern with increasing chessboard-block widths 𝑑 𝑤 from 10 to
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation Position 1
Ground truth location
Position 2
Position 3
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Position 4
0
36
72
108
144
180
L
Estimated Mean Disparity
Figure 8: The synthesized patterns with sampling artifacts at different positions and the corresponding depth maps from PSMNet. The change in the maximum disparity corresponds to the sampling artifact position. Average Max Disparity Estimation
Smaller PRS
Corresponding maximum disparities
Sampling Artifact
Position of synthesized sampling artifact
dw (mm)
Calibration Error Magnitude (pixels)
Figure 9: The average maximum disparity estimation at increasing positions of the synthesized artifact (left). The mean disparity estimation for PSMNet at calibration error magnitudes at increasing 𝑑 𝑤 (right). The graphs show that feature matching in deep learning-based SDE models is influenced by sampling artifacts and calibration errors. 40mm. We then apply calibration errors ranging from 0 to 15 pixels to 100 randomly selected KITTI images. The resulting disparity estimations produced by PSMNet are shown in Figure 9 (right). The estimated disparity increases as the calibration-error magnitude grows, similar to the behavior observed for the BM algorithm in Section 3.2.2. At larger 𝑑 𝑤 (above 40mm), the effect of calibration on the structure of the chessboard pattern is reduced, limiting its influence on disparity estimation. This is consistent with the BM algorithm behavior and confirms that the vulnerability observed in the BM and SGBM algorithms also extends to PSMNet.
4
Smaller PG
W
Right image with sampling artifact at different positions
Attack Characterization
As defined in the threat model, the attack consists of displaying or projecting on a flat or semi-flat target surface with a simple repeated pattern crafted to trigger the desired sampling artifacts and calibration errors, leading to shift in depth. This section analyzes the stereo matching process and the attacker capabilities with respect to the following parameters: (i) Pattern design, (ii) Algorithmic block matching parameters (block size and search range), (iii) Distance from the victim stereo camera, (iv) Physical size of the pattern, and (v) Contrast of the pattern. An attacker can systematically adjust the above parameters to achieve the desired depth shift. For this study, all simulations are conducted in Gazebo, where a virtual stereo camera is configured to match the ZED2 camera specifications by the vendor. We then validate the analysis in controlled real-world indoor scenarios, projecting the patterns using a ViewSonic PA700W projector [56]. In these real-world experiments, we collect 10 images of each pattern and average them to minimize
PRS
PG
Figure 10: Parameterization of the pattern based on granularity of the chessboard pattern (𝑃𝐺 ), and size of the repeating black squares in the pattern (𝑃𝑅𝑆 ).
the pixel noise in the image pairs. For characterization, we consider the BM algorithm with the SAD score as our reference metric. Formalization. An adversary places a structured attack pattern of physical width 𝑊 and length 𝐿, at a distance 𝑑𝑎 from the victim stereo camera, as illustrated in Figure 1. The relative pixel size of the pattern (𝑤 𝑝 , 𝑙𝑝 ) in the image pair can be measured as 1 𝑓𝑥 0 𝑊 𝑤𝑝 = (4) 𝑙𝑝 𝑑𝑎 0 𝑓 𝑦 𝐿 where 𝑓𝑥 and 𝑓𝑦 denote the focal lengths of the camera in the x and y-directions, respectively. The objective is to induce a desired depth estimation value, defined as 𝑑𝑒 , by exploiting the sampling artifacts and calibration error. To achieve this, we parameterize a chessboard pattern, representing the simplest form of geometric shape repetition. The structure is characterized by granularity (𝑃𝐺 ), which defines the width of repetition, and the repeating element size (𝑃𝑅𝑆 ), as illustrated in Figure 10. The adversary can measure the 𝑑 𝑤 corresponding to desired 𝑃𝐺 and 𝑃𝑅𝑆 , using Equation 4, by setting 𝑊 = 𝐿 = 𝑑 𝑤 .
4.1
Pattern Design
To achieve control over depth estimation, an attacker varies 𝑃𝐺 and 𝑃𝑅𝑆 to assess the relationship between the pattern and the corresponding disparity measurement. We span the granularity 𝑃𝐺 from 1 to 50 pixels, beyond the typical block size used by SGBM/BM algorithms (i.e., 15–30) and the size of the chessboard black square 𝑃𝑅𝑆 ∈ [1, 𝑃𝐺 ] for each 𝑃𝐺 . For example, we set 𝑑𝑎 to 1.5m and the size of the chessboard pattern in Gazebo to 1.5×1m (𝑊 × 𝐿 ≈ 490 × 300 pixels), which emulates the rear surface of a typical car obstacle and within a standard size of a roadside billboard [49]. The BM algorithm is applied with a block size of 15 pixels and a disparity range of 64 (default). We then collect the relative change in the disparity achieved by each pattern. The attack is considered successful in changing the measured disparity if more than 10 pixels exhibit a disparity change, consistent with thresholds in popular autonomous driving frameworks for obstacle detection (e.g., Autoware [55]). The correct disparity in the given scenario is 21 pixels (relative disparity 0 with no attack). To confirm the consistency of our results also in real-world setting, we project the same patterns in front of the ZED2 camera,
(1) (3)
(2) Granularity (PG)
Sri Hrushikesh Varma Bhupathiraju et al.
Estimated Relative Disparity
Estimated Relative Disparity
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Distance (m)
Figure 11: Relative disparity errors in BM with respect to pattern design and contrast (left). The trend shows (1) a mismatch due to calibration error, (2) periodic drops in disparity depending on the position and magnitude of the closest sampling artifact, and (3) a steady, linear increase in disparity error. Disparity error at increasing distances (right).
setting the initial granularity to 𝑃𝐺 = 3 pixels (35mm), given the resolution of the ZED2. Results. The trend in the estimated relative disparity as 𝑃𝐺 increases in the BM algorithm for both simulation and real-world evaluation is illustrated in Figure 11 (left). In both cases, the disparity exhibits a cyclic pattern, increasing with 𝑃𝐺 until the disparity range value is reached, followed by a sharp drop and a subsequent repetition of the upward trend. These abrupt drops (e.g., toward disparity values of 60–70) occur because increasing 𝑃𝐺 alters the sampling artifact magnitude, shifting it away from the reference block and outside the disparity range associated with the current false match. As a result, the matching cost reaches a new temporary minimum, leading the algorithm to select a different correspondence based on the similarity caused by the sampling artifacts. The granularity 𝑃𝐺 can be adjusted to exploit the calibration error or to control specific depth estimations by shifting the sampling artifact position. For example, within the 𝑃𝐺 ranges of 14–23 pixels (4676mm) or 23–50 pixels (76-165mm), a linear increase in estimated disparity can be leveraged for control. The analysis also shows that changes in 𝑃𝑅𝑆 do not affect the BM algorithm’s results, since BM matches blocks strictly horizontally. In Section 5, we show that SGBM and deep learning–based SDE models, which rely on pixels and features across all directions, are influenced by the 𝑃𝑅𝑆 value. Domination of Calibration Error. At lower granularity (up to 6 pixels 𝑃𝐺 ), the calibration error dominates the effect of the sampling artifact, producing a large depth shift that reaches the disparity range limit of the BM algorithm (64). This occurs because, although sampling artifacts are formed on the image, they are displaced and warped by geometric distortions, altering the artifact pixels in a manner consistent with the modeled pattern structure (see Equation 3). Given a calibration error magnitude 𝐸𝐶 and a sampling range 𝑆𝑅, the projected pattern width 𝑑 𝑤 must satisfy 𝑑 𝑤 ≥ 𝐸𝐶 · 𝑆𝑅 in order to exceed the calibration error. Block Matching Algorithm Parameters. Maintaining the same 𝑑𝑎 , 𝑃𝐺 , and pattern size, we investigate how BM algorithm parameters, such as block size and disparity range, influence the disparity estimation. We increase the block size from 5 to 105 at intervals of 10 pixels (since blocks in BM algorithm need to have a well-defined center, they cannot be divisible by 2), until the stereo matching algorithm cannot predict any relevant depth information. We then
increase the disparity range from 16 to 160 (disparity range divisible by 16 is a standard practice [47]). We observe that increasing the block size expands the trend horizontally, as a larger pattern granularity (𝑃𝐺 ) is required to induce disparity errors. This reflects the fact that with larger blocks, the algorithm aggregates more pixel information, thereby shifting the point at which repeated structures begin to cause mismatches. In contrast, the disparity range expands the trend vertically by limiting the range of possible disparity values. While the vulnerability triggered by the chessboard pattern remains unchanged, the amplitude of the cyclic disparity errors is capped by this disparity range, as illustrated in Figure 11. Thus, the size of the block determines the threshold required to manipulate the disparity, whereas the range of the disparity limits the magnitude of the estimated disparity while keeping the influence of granularity unchanged.
4.2
Attack Design
In this section, we examine the basic geometry underlying depth estimation errors and analyze how the distance and the spatial region of the repeated pattern influence the resulting depth estimates. In stereo vision, the cameras observe the scene through a viewing volume known as the camera frustum, which defines the spatial region within which objects can be captured, and depth can be estimated. Figure 3 illustrates the typical frustum of a stereo camera, consisting of a truncated pyramid extending from the camera center, with its width and height at a given depth 𝑍 determined by the field of view and focal length 𝑓 [19]. For an object with physical size 𝑆 at an original depth 𝑑𝑎 , its projected size on the image plane (𝑆 P ) scales inversely with depth according to 𝑆 P = 𝑆 · 𝑑𝑎 /𝑑𝑒 , where 𝑑𝑒 is the induced depth after manipulating disparity, as illustrated in Figure 3. When the target object depth is altered to be perceived closer to the stereo camera (decreasing 𝑑𝑒 ), 𝑆 P increases, and as it is moved farther, 𝑆 P decreases, reflecting the perspective scaling within the frustum. All depth manipulations conducted by the attacker are therefore constrained to this frustum volume, with the size of the object varying according to its position within the frustum. Measuring the frustum, it is possible to characterize how the induced disparity changes with pattern distances and sizes. 4.2.1 Distance from the Stereo Camera. Using a pattern design with 𝑃𝐺 and 𝑃𝑅𝑆 of 23 pixels (maximum disparity error from experiments in Section 4.1), we vary the distance between the stereo camera and the projection surface, 𝑑𝑎 , from 3-20m in 1m increments, in the Gazebo environment. The projected pattern has a physical width of 1.5m (240 pixels in image space). As 𝑑𝑎 increases, the perceived size of the pattern decreases proportionally to 𝑓 /𝑑𝑎 , where 𝑓 is the focal length of the camera. The disparity estimated by the BM algorithm exhibits a similar decreasing trend with increasing 𝑑𝑎 , as shown in Figure 11 (right) in both simulated and real-world testing with the projected pattern. This behavior occurs because the number of potential matching candidates for stereo correspondence decreases as the projected pattern occupies fewer pixels in the images, thereby reducing the maximum achievable disparity error. We further show in Section 6.1.1 that this trend persists also in RealSense D435. 4.2.2 Physical Size of the Pattern. To analyze the relationship between pattern size and disparity error, we vary the width (𝑊 ) and
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
length (𝐿) of the projected pattern from 0.5 to 3 meters in increments of 0.5 meters, while fixing the projection distance at 𝑑𝑎 = 3 meters and keeping the pattern design constant. Both Gazebo-based simulations and real-world experiments show that the estimated disparity is directly influenced by the width of the pattern in pixel space: as the physical size of the pattern increases, its width in pixels also increases, leading to a corresponding rise in disparity estimation. This increase follows an approximately linear trend, starting from 44 pixels for 𝑊 = 0.5m, until the disparity reaches the disparity range. Beyond this point, the error saturates and remains fixed at the disparity range value. This behavior arises because stereo matching algorithms search for correspondences only within a predefined disparity range; once this range is exceeded, additional pattern size no longer reduces disparity error. These results indicate that, depending on the stereo depth estimation algorithm implemented in a target stereo camera, an adversary could estimate this limit and design pattern sizes that reliably induce erroneous depth perception. Further, based on the camera frustum formalization described earlier in this section, if an attack region of width 𝑊 is required at distance 𝑑𝑎 , then the width 𝑊 ′ needed at a different attack distance 𝑑𝑎′ can be estimated as 𝑊 ′ = 𝑊 · 𝑑𝑎′ /𝑑𝑎 . Similarly, for a given attack region of width 𝑊 , the corresponding distance 𝑑𝑎 required to achieve the attack can be obtained from the inverse relation. Based on these properties of disparity estimation, an ad𝑓 ·𝑊 versary can design the attack using the relation 𝑑𝑎 = 𝐷 , where 𝑓 is the camera’s focal length and 𝐷 is the estimated disparity. Here, 𝑑𝑎 indicates the distance to the pattern and 𝑊 its physical width. To induce the same depth estimation error at a greater distance, the attacker can simply scale the pattern by this relation.
4.2.3 Pattern Contrast. Sampling artifacts typically arise at boundaries between different pixels, where their intensity takes on an intermediate value between the intensities on either side. Even when the contrast in the pattern is small, these artifacts can still influence stereo matching by altering the perceived pixel structures. To demonstrate this, we select the pattern design corresponding to the maximum disparity error in Figure 11, and progressively reduce its contrast in both the Gazebo simulation and real-world scenario by 10% increments (≈25 pixel intensity levels) up to 90%, as a 100% reduction would remove the pattern. We find that the disparity error persists until the contrast is reduced to 80%, in both simulation and real-world conditions. Below this level, the magnitude of the sampling artifact becomes too small to cause disparity errors. Based on the above characterization, the adversary can analytically derive the pattern design for a target 𝑑𝑒 , camera distance 𝑑𝑎 , pattern size and contrast, without optimizing for a specific scene.
5
Evaluation on Driving Dataset
Using the characterization outlined in Section 4, we perform a comprehensive analysis of attack effectiveness and induced depth control for both BM and SGBM algorithms implemented via the OpenCV library [47], deep-learning-based SDE models (PSMNet [15], MoCha-Stereo [17] and UniMatch [62]), and stereo-LiDAR fusionbased model (SGM-DDC [64]) over the entire KITTI driving dataset.
CCS ’26, November 15–19, 2026, The Hague, Netherlands
5.1
Experimental Setup
For this evaluation, we collect real-world chessboard patterns displayed on a monitor placed at a distance of 𝑑𝑎 = 1m to account for sampling artifacts and calibration errors in the captured stereo pairs. We then synthesize these patterns using the methodology described in Section 3.3 to overlay them onto all 7518 stereo image pairs from the KITTI dataset. Figure 16 in the Appendix illustrates an example synthesized image. The patterns are placed in front of the victim stereo camera, with a fixed true discrepancy ranging between 5 and 35 pixels, representing an equivalent real-world physical distance between 62 and 9m, emulating potential vehicles and billboard distances and consistent with assumptions adopted in prior work [39, 40]. We evaluate 627 patterns across 7 true discrepancies, resulting in 4389 total combinations per stereo pair. We extract the relative disparity change and assume the attack is successful only when a minimum of a 10-pixel region exhibits the desired disparity. This is consistent with the 10-point threshold used for obstacle detection in state-ofthe-art autonomous driving frameworks such as Baidu Apollo and Autoware [3, 55]. Note that for this evaluation, the calibration error magnitude in the synthesized pattern is 3 pixels, as measured in ZED2. To ensure visibility of the pattern, we set 𝑃𝐺 and 𝑃𝑅𝑆 initial values to 3 pixels, as finer granularity cannot be resolved by our ZED2 stereo camera due to its resolution.
5.2
Stereo Matching Algorithms
For both BM and SGBM algorithms, we set the block size to 15 pixels and the disparity search range to 64 pixels, following the default configurations of the algorithms under analysis [47]. In BM, we observe that the maximum induced depth exhibits a similar trend, consistent with the pattern shown in Figure 11. In BM, a 𝑃𝑅𝑆 > 5 pixels is necessary to shift the predicted depth, since smaller values result in patterns that are too fine to be reliably captured by the victim stereo camera. Within a 𝑃𝐺 range of 23–63 pixels, which provides precise control over disparity estimation, the attack achieves fine-grained manipulation at the single-pixel disparity level. This translates to control over the fake-depth distance 𝑑𝑒 , allowing an adversary to move the pattern surface from 1m to 13m in front of the camera. In contrast, SGBM shows a different behavior, where the maximum disparity increases with simultaneous growth in 𝑃𝐺 and 𝑃𝑅𝑆 . This occurs because SGBM introduces smoothness constraints by aggregating costs across multiple directions. While this multi-path aggregation suppresses outliers and reduces vulnerability to locally ambiguous matches, it makes SGBM less sensitive to lower 𝑃𝐺 values but susceptible to higher ones. Similar to BM, the attack demonstrates a fine-grained control over 𝑑𝑒 between 11.5m and 2.5m (disparities 26-58). Furthermore, a larger 𝑃𝑅𝑆 increases the vertical similarity considered by the SGBM algorithm, resulting in a higher calculated disparity. Natural Motion. In natural driving scenarios, the camera’s finite exposure duration produces motion-blur artifacts in regions where scene elements move relative to the camera. This blur smears object boundaries and alters pixel intensities through sampling artifacts and region-specific calibration errors. To evaluate the attack under realistic motion, we synthesize motion blur on KITTI images using
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Sri Hrushikesh Varma Bhupathiraju et al.
Table 1: Maximum disparity estimation error induced by the attack across the tested ground truth disparities. Ground Truth Disparity (pixels) Model PSMNet MoCha-Stereo UniMatch
5
10
15
20
25
30
35
79 30 187
77 69 132
51 34 12
55 70 10
51 51 8
72 69 15
49 73 20
the experimental setup described in Section 5.1. We test the attack on the BM and SGBM algorithms. Motion blur is generated using the point spread function [54]. The details of the motion blur simulation are provided in Appendix D. We apply the point spread function assuming a vehicle moving at a constant speed of 40 km/h, resulting in an average displacement of approximately four pixels for the projected pattern. We observe that the attack exhibits the same overall behavior for BM and SGBM, consistent with Section 5.2. For BM, the results follow the trend shown in Figure 11, indicating precise control of depth estimation within a 𝑃𝐺 range of 23–63 pixels, with a maximum 𝑑𝑒 of 12m. For SGBM, the maximum disparity increases jointly with 𝑃𝐺 and 𝑃𝑅𝑆 , yielding a maximum 𝑑𝑒 of 9m. This aligns with Section 5.2, where multi-path aggregation in SGBM suppresses larger depth errors. Overall, these results show that the sampling artifacts and calibration errors introduced by the patterns persist and enable controlled depth manipulation despite motion-blur at speeds of 40 km/h.
5.3
Deep Learning SDE algorithms
Here, we evaluate the shift in estimated depth on the state-of-theart PSMNet [15], MoCha-Stereo [17] and UniMatch [62] models, trained on the KITTI dataset [54], using their default training parameters. Specifically, we consider the median calculated disparity of 80% of the pattern region (from the center), and measure the shift in desired depth due to 𝑃𝐺 and 𝑃𝑅𝑆 variation. Table 1 lists the maximum disparity estimation error induced by the attack at increasing ground truth disparities. PSMNet Analysis. In PSMNet, the attacker can increase the calculated disparity by up to 79 pixels relative to a pattern placed at a 5 pixel disparity (ground truth). This corresponds to a 59m shift in estimated depth toward the victim camera from a pattern located at 𝑑𝑎 =62m. Similar to BM and SGBM, 𝑃𝑅𝑆 > 5 pixels is necessary to induce a substantial shift in the calculated disparity, as the block must be large enough for the model to reliably detect and match. We observe a cyclic change in estimated depth with increasing 𝑃𝐺 of the pattern geometry, as seen in Figure 12. Fine-grained control over the estimated depth can be achieved with an average step size of approximately 1 pixel. As the true disparity increases, the achievable range of depth shift decreases as expected. Additionally, we find that the calculated disparity can increase up to 22 pixels, corresponding to a controllable increase in depth estimation of up to 15m. MoCha-Stereo Analysis. The attack causes a controlled depth estimation error of up to 52m, corresponding to an estimated depth of 10m in front of the victim camera at a true distance of 62m. We observe a similar cyclic change in estimated depth with increasing
𝑃𝐺 as seen in PSMNet, with a higher sensitivity to 𝑃𝑅𝑆 variation (Figure 12). Control over the estimated depth is achieved with a step size between 0.9 and 1.55 pixels, consistent across all disparities. UniMatch Analysis. In UniMatch, we observe a shift in calculated disparity of up to 187 pixels, equivalent to an estimated depth of 60.5m towards the victim camera at a true distance of 62m. In contrast with PSMNet, UniMatch presents a higher effect with a smaller 𝑃𝑅𝑆 . We observe similar peak granularities to PSMNet at select 𝑃𝐺 (e.g., 33-34) as shown in Figure 12, indicating transferability of selected patterns between PSMNet and UniMatch. We also achieve controllable depth estimation between 1 and 187 pixels, with an interval between 0.5 and 1.2 pixels, except at a true disparity of 10, where a single outlier increases the interval to 19 pixels. Furthermore, we find that calculated disparity can increase up to 20 pixels, corresponding to an increase in depth estimation of 12m. In contrast with PSMNet, between 10 and 20m, we achieve a maximum depth control of just 6 pixels (∼2.4m).
5.4
SDE Stereo - LiDAR Fusion Model
We further extend our evaluation to a Stereo-LiDAR Fusion Model, SGM-DDC [64]. This model augments the semi-global matching (SGM) algorithm with a discrete disparity-matching cost (DDC) to insert sparse LiDAR disparities into the SGM cost volume. First, sparse LiDAR disparities are propagated to neighboring pixels to form a semi-dense LiDAR prior. The SGM cost volume is then compared against this LiDAR prior, and the DDC term favors disparities where both the stereocamera and LiDAR measures agree with each other. A final consistency check is used to ensure cohesion between the sensor views. To evaluate the fusion model against our patterns, we synthesize a LiDAR point cloud to represent a realistic flat wall, such as a billboard displaying the chessboard pattern. The results show a maximum change in depth from 62m to 16m from the target vehicle, a shift in depth of 46m. This estimated depth error decreases quickly within 31m, with a shift of just 4.4m possible with a more restricted range of control, varying from 4 and 11 pixels depending on the true depth of the pattern, with an interval between 0.64 and 1.43 pixels. When the pattern is placed between 63m and 21m away from the camera, the range of control in depth is only 4 pixels. This reduction in range of control and overall error corresponds with a reduced effect of 𝑃𝐺 variation, particularly at farther distances. SGM-DDC explicitly favors disparities closer to the LiDAR prior, which is unaffected by the pattern 𝑃𝐺 and 𝑃𝑅𝑆 . This constrains the attacker to induce a shift in depth estimation consistent with the LiDAR prior. SGM-DDC is therefore more resilient to pattern variation than the non-fusion models, but it remains affected by repeated patterns.
6
Real World Evaluation
We examine the effectiveness of the attack in static indoor scenarios and dynamic outdoor settings with the ZED2 [52] and Intel RealSense D435 [32] stereo cameras, using the ViewSonic PA700W to project the patterns over a wall and the back of a van. To assess end-to-end consequences on the autonomous driving framework we use the Euclidean clustering within Autoware [55] to perform
Granularity (PG)
CCS ’26, November 15–19, 2026, The Hague, Netherlands
MoCha-Stereo
UniMatch
Size of repeating element (PRS)
Size of repeating element (PRS)
PSMNet
Size of repeating element (PRS)
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
Granularity (PG)
Granularity (PG)
Figure 12: Example of heatmaps showing the maximum predicted disparity for PSMNet, Mocha, and UniMatch models, reported as a delta relative to the ground-truth disparity.
6.1
(a) Illustration of the Experimental Setup
Static Indoor Scenario
To analyze the shift in estimated depth and controllability of the attack under real-world conditions, we set 𝑑𝑎 = 3m, with the chessboard patterns projected onto a flat surface, consistent with Section 4.2. The distance of 3m is chosen to ensure that the entire projected pattern is captured within both the left and right images of the tested stereo cameras, and the ambient illumination is maintained at a constant value of 500 lux (typical of an indoor laboratory environment). At 𝑑𝑎 = 3m, a minimum pattern width 𝑊 of 0.9m and length 𝐿 of 0.6cm is required to induce false depth in ZED2, while 𝑊 of 60 cm suffices for RealSense. These widths correspond to ≤ 1.4% of the image pair area in both ZED2 and RealSense cameras. The attack produces a shift in depth of up to 2.4m in RealSense and 2.2m in ZED2, meaning the surface is perceived at 0.6m and 0.8m distance to the camera, as illustrated in Figure 13 (a). Following the characterization of the stereo matching algorithms, we evaluate the controllability of the attack on both cameras using chessboard patterns with increasing granularity (𝑃𝐺 ). Please note that we apply this for ZED2 and RealSense without reverse engineering or knowledge of their proprietary SGBM, using the same pattern against classical BMs. Results and Observations. For the RealSense camera, the pattern enables fine-grained control, with depth adjustable in increments of 0.1m between 0.4 and 2.6m. The depth increases linearly as 𝑃𝐺 varies from 14 to 27 pixels. Since the RealSense camera has a depth range of 3m, the attack does not displace the pattern farther from its original distance. In contrast, the ZED2 camera exhibits coarser control, with changes in depth of approximately 2.2, 1.4, and 0.5m observed at 𝑃𝐺 = 27, 33, and 46 pixels, respectively. At lower values of 𝑃𝐺 (510), the attack shifts the estimated depth farther by up to 4m, while maintaining fine-grained control at approximately 1m resolution. Moreover, the RealSense camera requires a repeating geometric shape size to satisfy 𝑃𝑅𝑆 = 𝑃𝐺 (a complete chessboard pattern) to induce false depth, whereas the ZED2 camera requires 𝑃𝑅𝑆 such that 𝑃𝐺 − 7 < 𝑃𝑅𝑆 < 𝑃𝐺 . We hypothesize that this difference arises because ZED2 employs a more advanced stereo matching algorithm that incorporates features such as edges and contours, reducing finegrained control over depth estimation. Nonetheless, the evaluation
ZED2 Camera
Projector
Estimated point cloud from ZED2
(b) Induced fake depth at increasing distance Fake depth distance (de) (m)
obstacle detection, and consider the attack successful if the false depth is detected as an obstacle.
Pattern distance (da) (m)
Figure 13: (a) Indoor scenario setup. (b) Induced depth error (𝑑𝑒 ) at increasing pattern distances (𝑑𝑎 ) for ZED2 and RealSense cameras in indoor scenario . demonstrates that it is possible to induce consistent false depth as close as 0.8m from the victim camera in both ZED2 and RealSense. Note that, for each experiment, a depth shift is considered successful if the resulting point-cloud points are detected as a genuine obstacle by Autoware’s Euclidean clustering. 6.1.1 Controllability at Increasing Distances. The distance between the stereo camera and the projection surface is increased from 𝑑𝑎 = 3m to 20m at increments of 1m. Building on the characterization of the attack in Section 4.2.1, we project a pattern of 𝑊 × 𝐿 = 2 × 1.5m, as the prior results demonstrate that this dimension is necessary to induce false depth at extended distances. We select pattern with 𝑃𝐺 and 𝑃𝑅𝑆 that resulted in the highest depth shifts for ZED2 and RealSense, respectively. For each distance, the experiment is repeated 10 times. Figure 13 (b) presents the minimum induced false depth (𝑑𝑒 ) in RealSense and ZED2 stereo cameras at increasing projection distances. The attack can shift depth estimation from up to 𝑑𝑎 = 15m in ZED2 and 16m in RealSense. For example, the attack can shift the surface depth as close as 2m from both cameras, when 𝑑𝑎 = 10m. The trend in 𝑑𝑒 initially increases exponentially and then transitions to a near-linear saturation when 𝑑𝑎 > 12m for RealSense and 𝑑𝑎 > 8m for ZED2. This behavior is consistent with the experimental results in Figure 11 (right), where the incorrect disparity values decrease exponentially at first and then approach saturation at larger 𝑑𝑎 . These findings demonstrate the controllability of the attack from distances of up to 15m. 6.1.2 Scene Illumination. To evaluate the attack under varying illumination, we use the same experimental setup as in Section 6.1
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Outdoor Evaluation
For our outdoor realistic scenarios, we collect estimated depth maps from the ZED2 and RealSense cameras mounted on an AgileX Hunter 2.0 [2] unmanned ground vehicle (UGV). The surface is positioned 1.5m to the right of the UGV, emulating a roadside billboard as shown in Figure 3, with 𝑊 × 𝐿 = 2 × 1.5m. Under nighttime conditions with an ambient illumination of 50 lux, the attack produces similar trends to the indoor setting. For example, the estimated depth plateau at 2.2m for ZED2 and 1.8m for RealSense when 𝑑𝑎 = 10m. Similarly, the attack induces false depth up to 𝑑𝑎 = 16m for ZED2 and 18m for RealSense. The slightly extended attack range observed at night is attributed to lower ambient illumination, which enhances the projected pattern sharpness and contrast. Under daytime conditions, with ambient illumination ranging from 1000 to 1500 lux, the attack induces false depth, causing obstacle detection at 𝑑𝑒 = 2.5m in front of the camera when 𝑑𝑎 = 8m for ZED2, and at 𝑑𝑒 = 1.8m, when 𝑑𝑎 = 10m for RealSense. The reduced attack range is due to the higher illumination, which diminishes the brightness of the projected pattern captured by the sensor, consistent with the contrast results of Section 4.2.3. This reduction in brightness lowers the effective contrast of the pattern and decreases the magnitude of sampling artifacts. Nevertheless, these results show the practicality of the attack in real-world conditions. Using the same experimental setup, we collect stereo image pairs with the victim cameras mounted on a UGV traveling at 10 km/h (refer to Section 12 for details), at distances to the projected pattern ranging from 𝑑𝑎 = 20 to 3m. We consider the attack successful, if consistent depth shift for ≥ 0.5 sec duration is achieved, sufficient to trigger emergency braking in AD systems such as Autoware [55]. The consistent attack durations achieved in these experiments are listed in Table 2. In nighttime conditions, the attack with 𝑃𝐺 = 49 and 𝑃𝑅𝑆 = 39 produces a stable false depth of 2.8m in ZED2 for 4.6 sec and approximately 1.3m in RealSense for 3.2 sec. Under daytime illumination, the attack with 𝑃𝐺 = 30 and 𝑃𝑅𝑆 = 25 achieves consistent false depth for up to 2.1 sec in ZED2 and 1.8 sec in RealSense. These results are similar to those in static scenarios with a 0.3m variance. Furthermore, the attack requires higher 𝑃𝐺 and 𝑃𝑅𝑆 than in the indoor experiments in Section 6.1, since a successful attack requires consistent depth shift from further distances. These results show the possibility for attackers to consistently induce false depth for a sufficient amount of time to trigger automatic responses in the autonomous system, such as unexpected and sudden braking or dangerous maneuvers. 6.2.1 Driving Scenario. We extend the evaluation to a real-world scenario where an attacker projects a pattern onto a van obstacle
Projector Camera attached to the vehicle Nighttime Conditions
Time (s)
Error in Estimated Depth (m)
6.2
Pattern projected on Van Obstacle
Error in Estimated Depth (m)
with the patterns corresponding to highest shift in depth. The ambient illumination in the room increases from 0 to 1000 lux, in steps of 100 lux. For each illumination level, the evaluation is repeated 10 times, and the estimated depth maps from the ZED2 and RealSense cameras are collected. Across all the tested illumination conditions, the estimated false depth remains consistently at 0.8m for ZED2 and 0.6m for RealSense. These findings align with the results in Section 4.2.3, confirming that sampling artifacts affect stereo matching even under higher illumination conditions.
Sri Hrushikesh Varma Bhupathiraju et al.
Daytime Conditions
Time (s)
Figure 14: Experimental setup in outdoor scenarios, with van obstacle (top). The induced depth by the ZED2 camera in night (bottom-left) and daytime (bottom-right) conditions. Table 2: Attack duration (in seconds) for both RealSense D435 and ZED2 cameras in night and daytime conditions. Condition Daytime Nighttime
10 km/h (UGV) RealSense D435 ZED2 1.8s 2.1s 3.2s 4.6s
15 km/h (car) RealSense D435 ZED2 1.0s 0.7s 1.2s 1.0s
positioned either on the roadside or in an adjacent lane to the victim vehicle, as illustrated in Figure 3. The projector, placed 4m behind the van, casts a 1.8m wide attack pattern onto its rear surface as shown in Figure 14. The van is located 1.5m to the side of the victim car, and the depth maps of the ZED2 and RealSense are collected following the setup in Section 6.2. The victim car is traveling at ≈15 km/h, moving from 𝑑𝑎 = 20m to 3m distances from the van. The experiments are conducted under nighttime (50 lux) and daytime (850–1100 lux) conditions. We select patterns to induce false depth between 4-7m from the van. Figure 14 shows consistent incorrect depth estimation lasting more than 0.5 sec for both ZED2 and RealSense cameras under both lighting conditions. Specifically, false obstacles persist for 1.06 sec and 1.23 sec in daytime and nighttime, respectively, for RealSense and for 1 sec and 0.7 sec for ZED2. As in the case of the billboard pattern, the van is not in the victim’s trajectory; however, the attack succeeds in shifting the van’s depth into the victim’s path along the frustum. These findings confirm the possibility to conduct the attack by projecting onto real-world obstacles, with the potential to trigger unsafe and unexpected behaviors in the victim vehicles.
7
High Speed Evaluation
The attack under high-speed driving conditions is evaluated on CARLA [22], a widely used AV testing simulator. We synthesize a scenario in which the victim vehicle, equipped with a ZED2 stereo camera, drives at a constant speed, with a proxy billboard displaying the attack pattern placed 2m to the right of its trajectory, following the setup shown in Figure 3(a). The vehicle approaches the billboard
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
CCS ’26, November 15–19, 2026, The Hague, Netherlands
from 50m away and passes it at speeds of 10, 20, 30, and 40 km/h. We collect stereo images throughout the vehicle trajectory and apply the BM algorithm to estimate depth from the captured frames. 𝑃𝐺 and 𝑃𝑅𝑆 are set to 26 pixels and 10 pixels, respectively, as they demonstrated the highest depth estimation error in static scenarios. The results show consistent depth manipulation for durations of 5.3, 2.6, 2.3, and 2.1 s at vehicle speeds of 10, 20, 30, and 40 km/h, respectively. The shorter attack durations observed at higher speeds are because the vehicle traverses the fixed 50m distance in less time. These durations all exceed the 0.5 s chattering threshold, the minimum detection time required for the vehicle to react in autonomous driving frameworks such as Autoware [55]. These results indicate that the attack can induce consistent depth shifts at typical driving speeds of 40 km/h, potentially triggering unsafe driving behaviors, including sudden braking and unexpected maneuvers.
they occur at regular intervals and exhibit consistent magnitude and gradient. For each block, the disparity range containing such detected oscillatory peaks is labeled as a repeated element, and applying this process across all blocks yields a mask of the attacked region. When identified, we draw a corresponding 2D bounding box and perform block matching using the entire bounding box as a single unit. Standard matching is then applied to the stereo image pair, and the minimum SAD score obtained for the bounding box is taken as the correct disparity estimation within the repeated pattern region. The proposed defense on 200 stereo image pairs collected under real-world conditions achieves 96.5% success rate in mitigating the attack, with the resulting post-mitigation depth error remaining below 0.5m. This corresponds to a 98% success rate in nighttime and 94.5% in daytime conditions, demonstrating consistent performance across varying environments. We further extend our methodology to deep learning models. To achieve this, we extract the PSMNet model’s intermediate results from the cost-volume matrix. The cost-volume matrix of deep learning-based SDE algorithms measures the feature-level disparity between stereo images, analogous to SAD scores in classical SDE algorithms. The feature magnitudes from the matrix are normalized, then equidistant minima are detected in the feature magnitudes, following the methodology described above for the SAD scores. To evaluate this, we apply the attack to the 200 image pairs from KITTI, using the same experimental setup described in Section 5.1. This methodology achieves 100% success rate in suppressing the depth estimation error to below 0.1m. These results demonstrate the effectiveness of the proposed defense in mitigating the attack against both classical and deep learning-based SDE algorithms. Furthermore, the defense can also be used to mitigate natural repeated patterns, which can unintentionally trigger the same effects.
8
Defenses
While patterns with repeating elements are recognized as a challenging case for stereo depth estimation, stereo camera vendors recommend parameter adjustments to mitigate depth errors [35]. For example, RealSense suggests changing the census register diameter and the peak threshold to account for depth errors. Such tuning is impractical in dynamic real-world environments where autonomous systems operate, because parameters would require continuous, real-time adaptation to the spectral properties of incoming stereo images. Moreover, such tuning can only mitigate attacks at a single pattern granularity, whereas an adversary can vary across multiple granularities. Barrois et al. [4] propose detecting repetitive regions using FFT and to estimate the true distance using a stereo model-free optimization. To verify the effectiveness of this method, we implement and evaluate it on 200 stereo image pairs collected from the ZED2 camera using the setup described in Section 6.2, comprising daytime and nighttime scenarios. Since our patterns are placed at 3–20m away, we consider the attack mitigated when the depth estimation error is reduced to less than 0.5m, an acceptable relative error [41]. As a result, the rate of successful mitigation is only 13% of the evaluated stereo image regions. This lower success rate is primarily caused by the FFT-based detector that fails to capture the repeated elements. As real-world patterns do not remain perfectly uniform, this violates the assumptions in the original method, especially when 𝑃𝑅𝑆 < 𝑃𝐺 . This is consistent with the design limitations discussed in the work, which require near-perfectly repeated structures.
8.1
Proposed Defense
To overcome the limitations of previous approaches, we propose a new defense that leverages the effect of repeated elements on matching scores (SAD) in BM algorithms. In the presence of repeated elements, the trace of the matching scores over the pixel offsets includes periodic peaks corresponding to the pattern, as shown in Figure 6. The proposed method uses such repeated peaks to detect patterns placed by an attacker. To achieve this, we normalize the SAD scores for each block and identify minima produced by oscillations induced by repeated structures in the pattern. To ensure that these minima correspond to structured repetition rather than noise, we detect peaks only when
9
Discussion
Training Dataset. Deep learning based models are affected by their training methodology. For example, when analyzing the UniMatch model pretrained on the SceneFlow dataset [45] instead of KITTI, we observe a shift in the range of possible depth errors and the patterns that are most effective at a given distance. The maximum predicted disparity within the pattern region decreases by over 50 pixels at a true distance of 60 m, but increases by over 50 pixels at a true distance of 10 m, with a completely different selection of effective patterns. This shows how multiple factors concur to trigger the vulnerability in deep-learning models, including the dataset used, specific model architecture, and training methodology. Additionally, convolutional aggregation in deep learning-based models can blend features extracted from the pattern and the background at the pattern edges, as shown in Figure 17. Further details are in Appendix C. Changing Illuminance and Texture. The attacker can adjust the projection brightness depending on distance and environmental conditions, using the formula 𝐵𝑝 = 𝐵 + 𝑘 · 𝐼 , where 𝐵 is the average projected pixel brightness at a distance, 𝑘 is the flat surface reflectance factor, and 𝐼 is the average ambient illumination. We use this during our daytime outdoor evaluation (natural light ranging from 1000-1500 lux) with a conventional projector at 4 meters.
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Further, projection patterns can be adjusted based on the texture of the surface used for the projection. For example, on a surface with curvature 𝐶𝑈 , adversaries can use standard radial image contortions based on Δ𝑖 = 𝑖 − (1/𝐶𝑈 ) · sin(𝑖 · 𝐶𝑈 ), where 𝑖 is the radial pixel distance, to maintain perspective across image frames. Stealthiness. The adversary can further deploy a stealthy attack by disguising the repeated pattern with naturalistic textures by leveraging surrounding objects [65] or by exploiting the infrared domain, as demonstrated in prior work [8, 50] to project patterns. We additionally verify that adversaries can leverage infrared projections against the RealSense D435 camera, as some stereo systems do not deploy infrared filters to support time-of-flight depth estimation or to enhance low-light performance. Limitations. In our real-world dynamic experiments, a single attack pattern is used to manipulate depth estimation on the ZED2 and RealSense cameras. A more sophisticated adversary, however, could employ dynamic patterns that adapt to the victim vehicle’s position, enabling longer-lasting depth estimation errors while reducing the necessary attack surface. We further limit our experiments to patterns displayed or projected onto flat or semi-flat surfaces. In theory, adversaries could exploit obstacles with arbitrary shapes to project structured patterns, which make the precise obstacle depth estimation more challenging. Although our analysis is limited to four state-of-the-art stereo matching algorithms and deep learning models, current stereo processing techniques are derived from these approaches, which can potentially extend vulnerability risks to other architectures and stereo cameras within the same family.
Sri Hrushikesh Varma Bhupathiraju et al.
11
12
Related Works
Physical attacks for spoofing and modifying sensor readings are a growing concern. Prior works have demonstrated that cameras are susceptible to a range of spoofing attacks, including laser injection [37, 44, 50, 63] and electromagnetic interference attacks [37]. Similarly, LiDARs and radars are also susceptible to such physical attacks [10, 13, 14, 31, 34]. These attacks have been shown to be highly effective within the ADAS domain, where sensor perturbations can affect downstream decision-making. Bhupathiraju et al. [9] demonstrate that thermal cameras are susceptible to attacks on their image processing by either naturally occurring or maliciously placed heat sources, causing failures in downstream object detection. Chen et al. [16] leverage repeating patterns and build optimized adversarial patterns to attack camera-based SLAM algorithms for monocular depth estimation. Nassi et al. [43] exploit a vulnerability in object detectors that do not consider the depth of the obstacle. This work specifically considers stereo camera depth estimation, and the ambiguity caused by repeated elements within the scene. Deep Neural Networks also pose their own unique security concerns within the ADAS domain. In particular, prior works have demonstrated that Deep Neural Networks are highly susceptible to adversarial perturbations, including patch-based attacks [58] [12] [23]. In contrast, this work specifically uncovers a hidden vulnerability in stereo cameras that can be used to achieve controlled depth manipulation, without the need for adversarial optimization.
Ethical Considerations
All experiments in this work use publicly available datasets and software. The use of publicly available datasets ensures transparency and reproducibility. The attack evaluation in our indoor and outdoor scenarios was conducted in controlled environments mimicking real-world driving while adhering to the safety policy of our institutions and the speed limits, with the vehicles operating at up to 15 km/h. In line with ethical guidelines, we have disclosed our findings to the relevant stakeholders (ZED and RealSense) and are awaiting their responses. No human studies were involved in this research.
13 10
Conclusion
Our work identifies a dangerous vulnerability in stereo cameras that can be exploited using structured repeated patterns to induce depth estimation errors. These vulnerabilities also affect deep learning SDE algorithms without requiring the creation of adversarial examples. Our evaluation demonstrates that an adversary can take advantage of this vulnerability to control classical block matching algorithms and ML models, including two widely used stereo cameras deployed in autonomous systems. We evaluate attacks in high-speed simulated and controlled real-world scenarios, showing their practicality in driving conditions. Finally, we design a defense method that detects repeated elements and mitigates the resulting depth estimation.
Acknowledgment
This paper was edited for grammar using Grammarly and Microsoft Copilot. We disclosed the vulnerability and our findings to the vendor. This research was supported in part by the JST CREST JPMJCR23M4, JST FOREST JPMJFR2531, JST Next-generation Edge AI Semiconductor JPM-JES2515, and JSPS KAKENHI 24K02940 and 24K14943.
References [1] Steffen Abraham and Wolfgang Förstner. 2005. Fish-Eye-Stereo Calibration and Epipolar Rectification. ISPRS Journal of photogrammetry and remote sensing 59, 5 (2005), 278–288. [2] AgileX Robotics Team. 2023. Hunter2.0 User Manual. https://cdn.shopify.com/s/ files/1/0551/0630/6141/files/HUNTER2.0_USER_MANUAL2023.12_50805.pdf [3] Baidu Inc. 2017. Apollo. http://apollo.auto. Accessed: 2021-10-08. [4] Björn Barrois, Marcus Konrad, Christian Wöhler, and Horst-Michael Groß. 2010. Resolving Stereo Matching Errors due to Repetitive Structures Using Model Information. Pattern recognition letters 31, 12 (2010), 1683–1692. [5] Seven S Beauchemin and Ruzena Bajcsy. 2001. Modelling and Removing Radial and Tangential Distortions in Spherical Lenses. In Multi-Image Analysis: 10th International Workshop on Theoretical Foundations of Computer Vision Dagstuhl Castle, Germany, March 12–17, 2000 Revised Papers. Springer, 1–21. [6] Zachary Berger, Parth Agrawal, Tian Yu Liu, Stefano Soatto, and Alex Wong. 2022. Stereoscopic Universal Perturbations across Different Architectures and Datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15180–15190. [7] Amlaan Bhoi. 2019. Monocular Depth Estimation: A Survey. arXiv preprint arXiv:1901.09402 (2019). [8] Sri Hrushikesh Bhupathiraju, Takami Sato, Michael Clifford, Takeshi Sugawara, Qi Alfred Chen, and Sara Rampazzi. 2024. On the vulnerability of traffic light recognition systems to laser illumination attacks. In ISOC Symposium on Vehicle Security and Privacy (VehicleSec). ISOC. [9] S Hrushikesh Bhupathiraju, Shaoyuan Xie, Michael Clifford, Qi Alfred Chen, Takeshi Sugawara, and Sara Rampazzi. 2026. The Heat is On: Understanding and Mitigating Vulnerabilities of Thermal Image Perception in Autonomous Systems. In Network and Distributed System Security Symposium (NDSS).
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
CCS ’26, November 15–19, 2026, The Hague, Netherlands
[10] Sri Hrushikesh Varma Bhupathiraju, Jennifer Sheldon, Luke A Bauer, Vincent Bindschaedler, Takeshi Sugawara, and Sara Rampazzi. 2023. EMI-LiDAR: Uncovering Vulnerabilities of LiDAR Sensors in Autonomous Driving Setting Using Electromagnetic Interference. In Proceedings of the 16th ACM Conference on Security and Privacy in Wireless and Mobile Networks. 329–340. [11] Brad Rosen. 2025. The Future of Autonomous Vehicle Sensors: A Shift Toward Advanced Stereo Vision. https://www.nodarsensor.com/blog/the-future-ofautonomous-vehicle-sensors-a-shift-toward-advanced-stereo-vision [12] Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer. 2017. Adversarial Patch. arXiv preprint arXiv:1712.09665 (2017). [13] Yulong Cao, S Hrushikesh Bhupathiraju, Pirouz Naghavi, Takeshi Sugawara, Z Morley Mao, and Sara Rampazzi. 2023. You Can’t See Me: Physical Removal Attacks on LiDAR-Based Autonomous Vehicles Driving Frameworks. In 32nd USENIX security symposium (USENIX Security 23). 2993–3010. [14] Yulong Cao, Chaowei Xiao, Benjamin Cyr, Yimeng Zhou, Won Park, Sara Rampazzi, Qi Alfred Chen, Kevin Fu, and Z Morley Mao. 2019. Adversarial Sensor Attack on LiDAR-Based Perception in Autonomous Driving. In Proceedings of the 2019 ACM SIGSAC conference on computer and communications security. 2267– 2281. [15] Jia-Ren Chang and Yong-Sheng Chen. 2018. Pyramid Stereo Matching Network. In Proceedings of the IEEE conference on computer vision and pattern recognition. 5410–5418. [16] Baodong Chen, Wei Wang, Pascal Sikorski, and Ting Zhu. 2024. Adversary is on the Road: Attacks on Visual SLAM Using Unnoticeable Adversarial Patch. In 33rd USENIX Security Symposium (USENIX Security 24). 6345–6362. [17] Ziyang Chen, Wei Long, He Yao, Yongjun Zhang, Bingshu Wang, Yongbin Qin, and Jia Wu. 2024. Mocha-stereo: Motif channel attention network for stereo matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 27768–27777. [18] Kelvin Cheng, Tianfu Wu, and Christopher Healey. 2022. Revisiting NonParametric Matching Cost Volumes for Robust and Generalizable Stereo Matching. Advances in Neural Information Processing Systems 35 (2022), 16305–16318. [19] Arzu Çöltekin. 2006. Foveation for 3D Visualization and Stereo Imaging. Helsinki University of Technology. Introduction to the Aircraft Obstacle Avoidance System. [20] DJI. 2025. https://support.dji.com/help/content?customId=01700006547&spaceId= 17&re=US&lang=en&documentType&paperDocType=ARTICLE [21] Dolunts. 2025. Discover the World with a New Level of Technology. https: //dolunts.com/ [22] Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. 2017. CARLA: An open urban driving simulator. In Conference on robot learning. PMLR, 1–16. [23] Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. 2018. Robust PhysicalWorld Attacks on Deep Learning Visual Classification. In Proceedings of the IEEE conference on computer vision and pattern recognition. 1625–1634. [24] Foresight . 2023. 3D Stereo Vision for the Autonomous Vehicle Industry. https://www.foresightauto.com/3d-stereo-vision-for-the-autonomousvehicle-industry/ [25] Zhangjie Fu, Yueyan Zhi, Shouling Ji, and Xingming Sun. 2021. Remote Attacks on Drones Vision Sensors: An Empirical Study. IEEE Transactions on Dependable and Secure Computing 19, 5 (2021), 3125–3135. [26] Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. 2013. Vision Meets Robotics: The KITTI Dataset. The international journal of robotics research 32, 11 (2013), 1231–1237. [27] Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and Harnessing Adversarial Examples. arXiv preprint arXiv:1412.6572 (2014). [28] Mohd Saad Hamid, NurulFajar Abd Manap, Rostam Affendi Hamzah, and Ahmad Fauzan Kadmin. 2022. Stereo Matching Algorithm Based on Deep Learning: A Survey. Journal of King Saud University-Computer and Information Sciences 34, 5 (2022), 1663–1673. [29] Heiko Hirschmuller and Daniel Scharstein. 2008. Evaluation of Stereo Matching Costs on Images with Radiometric Differences. IEEE transactions on pattern analysis and machine intelligence 31, 9 (2008), 1582–1599. [30] Honda Motor Co. 2025. Meet the Autonomous Work Vehicle. https://www. honda.com/mobility/Autonomous-Work-Vehicle [31] David Hunt, Kristen Angell, Zhenzhou Qi, Tingjun Chen, and Miroslav Pajic. 2024. MadRadar: A Black-Box Physical Layer Attack Framework on mmWave Automotive FMCW Radars. In NDSS. [32] Intel Corporation. 2025. RealSense D435i. https://www.intel.com/ content/www/us/en/products/sku/190004/intel-realsense-depth-camerad435i/specifications.html [33] International Telecommunication Union, Radiocommunication Sector (ITU-R). 2011. Recommendation ITU-R BT.601-7: Studio encoding parameters of digital television for standard 4:3 and wide-screen 16:9 aspect ratios. Recommendation BT.601-7. ITU. https://www.itu.int/dms_pubrec/itu-r/rec/bt/r-rec-bt.601-7-201103-i!!pdfe.pdf
[34] Zizhi Jin, Qinhong Jiang, Xuancun Lu, Chen Yan, Xiaoyu Ji, and Wenyuan Xu. 2024. PhantomLiDAR: Cross-Modality Signal Injection Attacks against LiDAR. arXiv preprint arXiv:2409.17907 (2024). [35] Kevin Zhao, Dan Nie, Yusheng Hsu, Edward Tsuchiya. 2025. Mitigation of Repetitive Pattern Effect of Intel RealSense Depth Cameras D400 Series. https://realsenseai.com/wp-content/uploads/2025/06/Mitigate-RepetitivePattern-Effect-Whitepaper_v03.pdf [36] Nathan Koenig and Andrew Howard. 2004. Design and Use Paradigms for Gazebo, An Open-Source Multi-Robot Simulator. In IEEE/RSJ International Conference on Intelligent Robots and Systems. 2149–2154. [37] Sebastian Köhler, Richard Baker, and Ivan Martinovic. 2022. Signal Injection Attacks against CCD Image Sensors. In Proceedings of the 2022 ACM on Asia Conference on Computer and Communications Security. 294–308. [38] MJPM Lemmens. 1988. A Survey on Stereo Matching Techniques. International Archives of Photogrammetry and Remote Sensing 27, B8 (1988), 11–23. [39] Hangcheng Liu, Xu Kuang, Xingshuo Han, Xingwan Wu, Haoran Ou, Shangwei Guo, Xingyi Huang, Tao Xiang, and Tianwei Zhang. 2025. Optimization-Free Patch Attack on Stereo Depth Estimation. arXiv preprint arXiv:2506.17632 (2025). [40] Yang Liu, Jucai Zhai, Chihao Ma, Pengcheng Zeng, Xinan Wang, and Yong Zhao. 2024. Physical Attack for Stereo Matching. In Proceedings of the International Conference on Computer Vision and Deep Learning. 1–5. [41] David F Llorca, Miguel A Sotelo, Ignacio Parra, Manuel Ocaña, and Luis M Bergasa. 2010. Error Analysis in a Stereo Vision-Based Pedestrian Detection Sensor for Collision Avoidance Applications. Sensors 10, 4 (2010), 3741–3758. [42] Jean J Lorre and Alan R Gillespie. 1980. Artifacts in Digital Images. In Applications of Digital Image Processing to Astronomy, Vol. 264. SPIE, 123–135. [43] Ben Nassi, Yisroel Mirsky, Dudi Nassi, Raz Ben-Netanel, Oleg Drokin, and Yuval Elovici. 2020. Phantom of the ADAS: Securing Advanced Driver-Assistance Systems from Split-Second Phantom Attacks. In ACM CCS 2020. 293–308. [44] Dudi Nassi, Raz Ben-Netanel, Yuval Elovici, and Ben Nassi. 2019. MobilBye: Attacking ADAS with Camera Spoofing. arXiv preprint arXiv:1906.09765 (2019). [45] N.Mayer, E.Ilg, P.Häusser, P.Fischer, D.Cremers, A.Dosovitskiy, and T.Brox. 2016. A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation. In IEEE International Conference on Computer Vision and Pattern Recognition (CVPR). http://lmb.informatik.uni-freiburg.de/Publications/ 2016/MIFDB16 arXiv:1512.02134. [46] NODAR Inc. 2025. Long-Range 3D Vision for Automotive. https://www. nodarsensor.com/industries/automotive/ [47] OpenCV. 2025. OpenCV: Open Source Computer Vision. [48] Naman Patel, Prashanth Krishnamurthy, Siddharth Garg, and Farshad Khorrami. 2019. Adaptive Adversarial Videos on Roadside Billboards: Dynamically Modifying Trajectories of Autonomous Vehicles. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 5916–5921. [49] Naman Patel, Prashanth Krishnamurthy, Siddharth Garg, and Farshad Khorrami. 2021. Overriding Autonomous Driving Systems Using Adaptive Adversarial Billboards. IEEE Transactions on Intelligent Transportation Systems 23, 8 (2021), 11386–11396. [50] Takami Sato, Sri Hrushikesh Varma Bhupathiraju, Michael Clifford, Takeshi Sugawara, Qi Alfred Chen, and Sara Rampazzi. 2024. Invisible Reflections: Leveraging Infrared Laser Reflections to Target Traffic Sign Perception. arXiv preprint arXiv:2401.03582 (2024). [51] Daniel Scharstein and Richard Szeliski. 2002. A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms. International journal of computer vision 47, 1 (2002), 7–42. [52] StereoLabs. 2025. ZED 2. https://www.stereolabs.com/products/zed-2 [53] Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013. Intriguing Properties of Neural Networks. arXiv:1312.6199 (2013). [54] Richard Szeliski. 2022. Computer Vision: Algorithms and Applications. Springer Nature. [55] The Autoware Foundation. [n. d.]. Autoware: Open-Source To Self-Driving. https://github.com/CPFL/Autoware. Accessed: 2025-11-11. [56] ViewSonic . [n. d.]. Projector PA700W Datasheet. https://github.com/CPFL/ Autoware. Accessed: 2025-11-11. [57] Pengfei Wang, Xiaofei Hui, Beijia Lu, Nimrod Lilith, Jun Liu, and Sameer Alam. 2024. Left-Right Discrepancy for Adversarial Attack on Stereo Networks. arXiv preprint arXiv:2401.07188 (2024). [58] Hui Wei, Hao Tang, Xuemei Jia, Zhixiang Wang, Hanxun Yu, Zhubo Li, Shin’ichi Satoh, Luc Van Gool, and Zheng Wang. 2024. Physical Adversarial Attack Meets Computer Vision: A Decade Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 9797–9817. [59] Alex Wong, Mukund Mundhra, and Stefano Soatto. 2021. Stereopagnosia: Fooling Stereo Networks with Adversarial Perturbations. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 2879–2888. [60] Yalin Xiong and Larry Matthies. 1997. Error Analysis of a Real-Time Stereo System. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, 1087–1093.
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Sri Hrushikesh Varma Bhupathiraju et al.
[61] Haofei Xu and Juyong Zhang. 2020. AANet: Adaptive Aggregation Network for Efficient Stereo Matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 1959–1968. [62] Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. 2023. Unifying Flow, Stereo and Depth Estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023). [63] Chen Yan, Wenyuan Xu, and Jianhao Liu. 2016. Can You Trust Autonomous Vehicles: Contactless Attacks against Sensors of Self-Driving Vehicle. Def Con 24, 8 (2016), 109. [64] Yasuhiro Yao, Ryoichi Ishikawa, and Takeshi Oishi. 2025. Stereo-LiDAR fusion by semi-global matching with discrete disparity-matching cost and semidensification. IEEE Robotics and Automation Letters (2025). [65] Tianyue Zheng, Jingzhi Hu, Rui Tan, Yinqian Zhang, Ying He, and Jun Luo. 2024. {𝜋 -Jack } : { Physical-World } Adversarial Attack on Monocular Depth Estimation with Perspective Hijacking. In 33rd USENIX Security Symposium (USENIX Security 24). 7321–7338. [66] Ce Zhou, Qiben Yan, Yan Shi, and Lichao Sun. 2022. DoubleStar: Long-Range Attack towards Depth Estimation Based Obstacle Avoidance in Autonomous Systems. In 31st USENIX security symposium (USENIX Security 22). 1885–1902. [67] Husheng Zhou, Wei Li, Zelun Kong, Junfeng Guo, Yuqun Zhang, Bei Yu, Lingming Zhang, and Cong Liu. 2020. DeepBillboard: Systematic Physical-World Testing of Autonomous Driving Systems. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering. 347–358.
Appendix A Influence of Sampling Artifact on SAD Score Here, we formalize how sampling artifacts affect pixel intensities and, in turn, bias the minima of the block matching SAD cost, shifting predicted disparity. For a physical point 𝑝 = (𝑥, 𝑦, 𝑧), the pixel intensity can be estimated using Equation 1. The corresponding pixel intensity 𝑃 𝑅 of the 𝑖 𝑡ℎ pixel in the right image is given by: ∫ 𝑆𝑅𝑖+𝐷+1 1 𝑃 𝑅 (𝑖 + 𝐷) = 𝐼 opt (𝑝)d𝐴 (5) 𝐴(𝑆𝑅𝑖 ) 𝑆𝑅𝑖+𝐷 Here, 𝐷 denotes the ground truth pixel disparity between left 𝑓 ·𝑏 and right images, computed as 𝐷 = 𝑑𝑎 , where f is the focal length and b is the baseline distance between the two cameras (12 cm for ZED2). To model the variations, we first consider a simple black and white stripe pattern, placed at a distance 𝑑𝑎 from a stereo camera. We assume that each stripe in the pattern has a width 𝑑 𝑤 . The pixel intensity of the sampling artifact varies based on the optical intensity received from point 𝑝 in the repeated pattern, defined as: ( 𝐼 opt (𝑝) =
1 if 0 ≤ (𝑝 (𝑥) mod 𝑑 𝑤 ) < 𝑑 𝑤 /2 0 Otherwise
(6)
Here, 𝑝 (𝑥) defines the x-coordinate of the physical point 𝑝. This equation models pixel-level inconsistencies in sampling artifacts that ultimately introduce measurable disparity offsets. These offsets arise from variations in the sampling artifacts, which differ in magnitude between the left and right images. Such subtle differences in the sampling artifacts ultimately influence the final matching in SDE algorithms. Influence on SAD Score. In BM algorithms, a block 𝐵𝐿 in the left image is compared with the a candidate block 𝐵𝑅 on the right image within the disparity search range. The algorithm finally selects a corresponding block 𝐵𝐷 in the right image that minimizes the SAD score, as given by 𝐵𝐷 = arg min (SAD(𝐵𝐿 , 𝐵𝑅 )) 𝐵𝑅
.
(7)
When repeated elements are in the scene, multiple blocks in the right image (𝐵𝑅 ) have similar small SAD scores due to their inherent similarity. This creates an oscillating trend in the estimated SAD scores, as illustrated in Figure 6 (blue oscillating pattern). Without sampling artifacts, these blocks demonstrate the same SAD score at the minima of the oscillating pattern, as shown in Figure 6 (a). In such cases, the block with the highest disparity (i.e., the closest obstacle) is typically selected. In reality, however, each 𝐵𝑅 has a different SAD due to sampling artifacts: a particular 𝐵𝑅 with an artifact similar to that of 𝐵𝐿 has a lower SAD score than the others. The relative position between a pixel and the repeated pattern determines the pixel intensity, as suggested by Equation 1, and the SAD scores over a disparity become periodically high and low due to different intervals of pixels and the repeated pattern, as illustrated in Figure 6 (b). Figure 6 annotates the difference in the minima of the SAD scores with and without the presence of the SAD score, illustrating the influence on the minima of the oscillations. As we discuss further, this minimum ultimately influences the matching produced by the BM algorithm. To further analyze the influence of sampling artifacts on disparity matching, we formalize the vulnerability. We define an image block at position 𝑖 = (𝑖𝑥 , 𝑖 𝑦 ) in the image as 𝐵(𝑖). Since block matching in stereo vision is performed along the horizontal axis, we focus on the 𝑥-direction. Let 𝐵(𝑖 Ref ) denote the reference block in the left image, 𝐵(𝑖𝑇 ) the true corresponding block in the right image, and 𝐵(𝑖 𝐹 ) a falsely matched block in the right image. The SAD score 𝐵(𝑖 Ref ) and an arbitrary block 𝐵(𝑖) in the right image is expressed as 𝑆𝐴𝐷 (𝐵(𝑖 Ref ), 𝐵(𝑖)). As shown in Figure 6 (left), the SAD scores become periodically high and low, reflecting the width 𝑑 𝑤 of the repeating element width, which is expressed as SAD(𝐵(𝑖 Ref ), 𝐵(𝑖)) ≈ SAD(𝐵(𝑖 Ref ), 𝐵(𝑖 T + 𝑛 · 𝑑 𝑤 )),
(8)
where 𝑇 + 𝑛 · 𝑑 𝑤 represents a horizontal shift by 𝑛 repetitions of the pattern, with 𝑛 ∈ {1, 2, · · · }. This shows that the SAD score for 𝐵(𝑖 Ref ) is approximately the same at each repeated element in the right image, since the pattern reappears every 𝑑 𝑤 pixels. Such periodicity produces multiple candidate regions with similar SAD values, each of which can lead to a false match. This behavior is typical of block matching algorithms, which rely on local pixel intensities. We define the average pixel intensity of the sampling artifact at contrasting boundaries (Figure 4 (a)) as the sampling artifact magnitude. Let the magnitude at block position 𝑖 = (𝑖 (𝑥), 𝑖 (𝑦)) be denoted as 𝑃 (𝑖), where 𝑖 (𝑥) = 𝑛 · 𝑑 𝑤 (at every repetition of the pattern). Here, 𝑃 (𝑖) is estimated from the pixel intensity using Equation 5. The SAD score depends on the similarity between the sampling artifacts of two blocks: SAD(𝐵(𝑖 Ref ), 𝐵(𝑖)) ∝ |𝑃 (𝑖 Ref ) − 𝑃 (𝑖)|.
(9)
Thus, for repeating patterns, the SAD score between a reference block and a potential match (𝑖 (𝑥) = 𝑖𝑇 (𝑥) +𝑛 ·𝑑𝑤) is proportional to the difference in sampling artifact magnitude. An incorrect match occurs when SAD(𝐵(𝑖 Ref ), 𝐵(𝑖 𝐹 )) < SAD(𝐵(𝑖 Ref ), 𝐵(𝑖𝑇 ))),
(10)
which implies that the sampling artifact magnitude of the reference block is closer to that of the false match than to the true match. This condition can be expressed as: |𝑃 (𝑖 Ref ) − 𝑃 (𝑖 F )| < |𝑃 (𝑖 Ref ) − 𝑃 (𝑖 T )|.
(11)
The minima of the SAD scores, which determine the optimal match in BM algorithms, become dominated by sampling artifacts (Figure 6). Consequently, the BM algorithm tends to select the offset 𝑛 that minimizes the difference: arg min |𝑃 (𝑖 Ref ) − 𝑃 (𝑖𝑇 + 𝑛 · 𝑑 𝑤 )| .
(12)
𝑛
This leads to a lower matching cost but an incorrect disparity estimate. Because the sampling artifact magnitude 𝑃 (𝑖) varies periodically with 𝑑 𝑤 , as described in Equation 6, the minima of the SAD score are directly influenced. Since disparity estimation relies on these minima, the sampling artifact ultimately affects which block is chosen as the correct match when repeated elements create multiple candidates. An adversary could exploit this by estimating the periodic trend of sampling artifact intensity from Equation 6 and adjusting the pattern according to Equation 12 to manipulate depth estimation.
B
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Average Max Disparity Estimation
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
Normalized pixel similarity
Figure 15: The average maximum disparity increases with respect to higher normalized pixel similarity (higher similarity in sampling artifact pixel). The graph further validates that feature matching in deep learning-based SDE models is influenced by the sampling artifacts.
Sampling Artifact Consistency in PSMNet
To further validate the effect of sampling artifacts on the PSMNet model, we compute the pixel similarity between the left and right images within the artifact region using normalized pixel similarity, defined as the absolute difference between the sampling artifact intensities in each image pair. A high normalized similarity indicates that the corresponding artifact magnitudes are closely matched. The maximum disparity is measured with increments of 0.2, corresponding to these similarity values. As shown in Figure 15, artifacts with higher normalized similarities (similar sampling artifact magnitudes) are consistently matched, mirroring the behavior of classical block-matching algorithms. This further confirms that the vulnerability observed in BM algorithms also extends to deep learning models such as PSMNet.
C
Global Properties of Deep-learning Models.
Deep learning–based SDE models typically rely on convolutional layers to estimate disparity by analyzing local similarity between the left and right images. Consequently, at the boundaries of the attack pattern, features from both the pattern and the surrounding foreground are blended through convolutional aggregation. This interaction can create a perspective hijack, leading to incorrect depth estimation for objects located at the edges of the attack region. Examples of such edge-based errors are shown in Figure 17. We do not include these edge-induced disparity errors in our evaluation, as they arise specifically from the convolutional feature extraction of deep learning models. Nonetheless, adversaries could exploit this behavior to mislead SDE models and extend the affected region of a scene.
D
Motion Blur Simulation
The motion-blurred image can be modeled by integrating the continuously changing image during the exposure interval 𝑇 . If the
Figure 16: Image synthesis with an attack pattern. Each pattern is cropped from a real-world image of a computer monitor, then luminance matched to the background image, and placed on top of the image with the corresponding disparity.
camera moves toward the obstacle with constant velocity 𝑣, the blurred image is given by 𝐼𝑏 (𝑥, 𝑦) =
1 𝑇
∫ 𝑇
𝐼 (𝑥 − Δ𝑥 (𝑡), 𝑦 − Δ𝑦(𝑡)) 𝑑𝑡,
(13)
0
where (Δ𝑥 (𝑡), Δ𝑦 (𝑡)) represents the image displacement caused by camera motion at time 𝑡. For forward motion, the displacement is radial and increases with distance from the image center:
CCS ’26, November 15–19, 2026, The Hague, Netherlands
Sri Hrushikesh Varma Bhupathiraju et al.
equivalent point spread function corresponds to the normalized trajectory traced by the pixel over the exposure duration 𝑇 .
E
Figure 17: Edge artifacts caused by the pattern on three different scenes.
𝑓 𝑣𝑡 , (14) 𝑍 (𝑥, 𝑦) where 𝑓 is the focal length and 𝑍 (𝑥, 𝑦) is the scene depth. Thus, each pixel accumulates intensity along a radial trajectory during exposure, producing a spatially varying radial motion blur. The Δ𝑟 (𝑡) =
Open Science Policy
We are dedicated to promoting the principles of open science and ensuring the reproducibility of research in the field of autonomous system security. To uphold this commitment, we provide comprehensive access to all relevant research artifacts and materials as detailed below. Details and demo videos recordings are available on our website: https://sites.google.com/view/stereocam/home. Our research artifacts, including the scripts to synthesize and evaluate the attack on SDE models, the resulting depth estimation point clouds from real-world experiments, and the script to test and evaluate the defense methodology, are available at: https://zenodo.org/ records/20767261.