Lights, Camera, Attack: Exploiting Temporal HDR Fusion with Pulsed Light Alkim Domeke Clemson University
Michael Kühr Technical University of Munich
arXiv:2609.37742v1 [cs.CR] 29 Sep 2026
Mohammad Hamad Technical University of Munich
Roman Gilliatt Clemson University
Sebastian Steinhorst Technical University of Munich
Long Cheng Clemson University
Bing Li Clemson University
Mert D. Pesé Clemson University
Abstract
2. Adversarial Pulsed Light
1. Benign Scene
3. Camera ISP with Temporal HDR
4. Example Corrupted Fused Output
HDR
Modern cameras widely use temporal High Dynamic Range (HDR) to improve visibility by capturing a sequence of exposures with different integration times and fusing them into a single image. This process implicitly assumes that scene illumination remains sufficiently stable during capture. We introduce FLASH (Fusion-Level Attack by Saturating HDR), an external pulsed-light attack that deliberately attacks this assumption by creating cross-exposure inconsistency before downstream perception. FLASH exploits an algorithmic assumption rather than relying on sensor damage or hardware failure, and requires neither physical camera access, access to raw exposure brackets, knowledge of the fusion algorithm, nor exact phase lock to the camera. Across eight physical camera platforms spanning embedded, surveillance, photography, smartphone, and automotive use cases, and matched optical controls, FLASH causes pipeline-dependent darkening, overexposure, and visibility loss. This includes extreme-darkening rates of 50.0% on an iPhone 16 Pro and 33.7% on a Wyze Battery Cam Pro. On the Wyze camera, FLASH triggers the system-level low-visibility response in 10/10 trials, compared with 0/10 continuous-light and randomized-frequency flashing controls. In a controlled stationary OpenPilot case study, 23.0% of frames exhibit severe darkening in the traffic-cone target region, with target-background CNR decreasing by up to 90.8%. Under FLASH, the OpenPilot interface also fails to display the system-level path state observed in the corresponding control trials. In a controlled night-only HDR reconstruction stress test, a proof-of-concept exposure-rejection defense reduces median output-brightness deviation by 79.16%. These results show that temporal HDR fusion itself requires security-aware validation of exposure evidence.
Benign Temporal HDR Fusion
FLASH Attack on Temporal HDR Fusion time
Exposure 1
Exposure 2
…
Exposure n
time Exposure 1
Exposure 2 … Exposure b … Exposure n
…
…
Assumption: Consistent Scene Illumination Across Exposures
Clean:
…
Fusion
Fusion
Normal Fused Output Continuous:
Example Corrupted Fused Output Randomized Flashing:
FLASH creates inconsistent exposure evidence during capture. FLASH: Darker than Clean !
Figure 1: Overview of FLASH on temporal HDR fusion. Temporal HDR fuses exposures of different integration times under an assumption of stable scene illumination. FLASH creates cross-exposure inconsistency by perturbing one or more exposures, which can corrupt the fused output. The bottom row shows a Wyze Battery Cam Pro example, where continuous and randomized-frequency flashing brighten the scene while FLASH produces a darker output than the clean baseline.
1
Introduction
High Dynamic Range (HDR) imaging allows cameras to preserve visual detail across scenes containing both bright and dark regions. A common approach is temporal multi-exposure High Dynamic Range (HDR), in which the camera captures multiple observations of the same scene using different integration times and combines them into one output frame. Shorter exposures preserve information in bright regions, while longer exposures collect more light from darker regions; the camera’s Image Signal Processing (ISP) then aligns and fuses these observations into the final image [16, 36]. Tempo-
DISTRIBUTION STATEMENT A. Approved for public release; distribution is unlimited. OPSEC11092
1
ral multi-exposure fusion is identified in prior imaging literature as a widely applied HDR technique [54] and is used in mobile computational photography [23], commercial surveillance cameras [9], automotive imaging [49], and embedded vision systems [27, 46]. Despite differences in implementation, these pipelines rely on a common assumption: the scene and its illumination remain sufficiently consistent while the sequence of exposures is captured [16, 23]. This assumption creates a physical attack surface. If illumination changes rapidly during the exposure sequence, different exposures can record inconsistent versions of the same scene. Prior imaging work has shown that ordinary LED flicker can interfere with temporal HDR capture [12, 17, 48]. We show that this behavior can be induced deliberately. In particular, an external light source can inject different amounts of optical energy into different exposure intervals, causing the resulting exposure evidence to become mutually inconsistent. Rather than simply making the captured scene brighter, this inconsistency can interact with exposure fusion and subsequent ISP processing to produce darkening, overexposure, contrast loss, or other corrupted output. The resulting failure therefore occurs during image formation, before any downstream perception model receives the frame. We introduce Fusion-Level Attack by Saturating HDR (FLASH), a non-contact physical attack that deliberately creates this cross-exposure inconsistency using pulsed light. Fusion-Level Attack by Saturating HDR (FLASH) specifically targets sequential temporal HDR pipelines and does not target spatial or single-shot HDR architectures that capture dynamicrange information simultaneously. The attacker requires neither physical access to the camera, access to raw exposure brackets, knowledge of the proprietary fusion algorithm, nor exact phase synchronization with the camera. Instead, periodic open-loop illumination creates unequal pulse overlap across exposure intervals. Longer integration windows generally provide greater opportunity for pulse overlap, but FLASH does not require a particular number of exposures or a specific exposure index to be corrupted. This distinguishes FLASH from continuous strong-light attacks that primarily disrupt acquisition through persistent sensor saturation [32, 40], as well as attacks that optimize perturbations against downstream perception models [19, 22, 28]. FLASH instead exploits an algorithmic assumption inside temporal HDR fusion itself. We evaluate FLASH through analytical modeling, controlled simulation, matched optical controls, physical camera experiments, and a controlled stationary comma 3X/OpenPilot case study. Simulation characterizes how geometry, source strength, and ambient illumination affect the attack and shows that the effect is strongest under low ambient illumination, short distance, and near-axis aiming. Across physical embedded, surveillance, consumer, and smartphone camera pipelines, FLASH produces device-dependent responses rather than a uniform brightness change. The strongest framelevel collapse occurs on the iPhone 16 Pro and Wyze Battery
Cam Pro, with extreme-darkening rates of 50.0% and 33.7%, respectively, relative to continuous-light controls. On the Wyze camera, FLASH also triggers the built-in low-visibility response in 10/10 trials, compared with 0/10 continuous-light and randomized-frequency flashing controls. In the controlled stationary comma 3X case study, 23.0% of frames exhibit severe darkening in the traffic-cone target region, and targetbackground separability decreases by up to 90.8%. Under FLASH, the OpenPilot interface also displays neither the green planned path nor the red blocked-path state observed in the corresponding control conditions. We treat these as camera-stream degradation and displayed-interface observations and do not attribute them to a particular internal perception or planning component. Because the vulnerability arises before exposure fusion, we also evaluate whether inconsistent exposure evidence can be rejected before reconstruction. In a controlled three-exposure night-only HDR stress test, a proof-of-concept exposure-level rejection method detects the corrupted exposure in 756/756 cases with a 0.0% clean false rejection rate and reduces median absolute luma deviation from 60.46% to 12.60%, a 79.16% relative reduction. This defense evaluation is performed in a controllable reconstruction pipeline rather than inside proprietary commercial ISP firmware. This paper makes the following main contributions: • We identify temporal HDR fusion as a physical attack surface and introduce FLASH, which creates crossexposure illumination inconsistency using externally generated pulsed light. The attack operates before downstream perception and does not require physical camera access, raw exposure access, knowledge of the fusion algorithm, or exact phase synchronization. • We characterize FLASH through analysis, simulation, matched optical controls, and physical evaluation across embedded, surveillance, consumer, smartphone, and automotive camera systems. The results show that different camera pipelines map cross-exposure inconsistency into different output failure modes, including frame-level collapse, brightening, exposure compensation, glare, and local visibility degradation. We further evaluate comma 3X camera-stream degradation and displayed OpenPilot path state in a controlled outdoor stationary setting. • We develop and evaluate a proof-of-concept exposurelevel rejection method that removes anomalous exposure evidence before fusion. In a controlled three-exposure night-only HDR stress test, it rejects the corrupted exposure in 756/756 cases with a 0.0% false rejection rate and reduces median absolute luma deviation by 79.16%. To facilitate reproducibility and support future research, we make our HDR-Lab framework, defense method, and simulation files available as open-source code at https: 2
or downstream perception rather than deliberately creating inconsistent exposure evidence inside temporal HDR fusion. Timing-based camera attacks. Other attacks exploit temporal properties of camera capture. Rolling-shutter attacks modulate illumination timing to introduce adversarial stripes, erase objects, or spoof traffic-light signals by exploiting the sequential readout of image rows [14, 25, 29, 45, 57]. Scenewide flashing attacks similarly vary illumination over time, but generally target downstream video-recognition models rather than the internal relationship among HDR exposure brackets [42]. In contrast, FLASH uses temporal illumination specifically to create unequal optical contamination across the exposure observations used for temporal HDR fusion. It therefore targets exposure-fusion consistency rather than rolling-shutter readout or model-specific temporal sensitivity. ISP and HDR failures. Prior work has also studied attacks and non-adversarial failures within camera-processing pipelines. Pipeline attacks inject faults or manipulate stages such as sensor interfaces, ISP parameters, and image-scaling operations [32, 33, 39, 41, 55]. Separately, the imaging literature documents failures of temporal HDR under changing illumination. LED flicker can cause different exposure intervals to capture different illumination states [12, 17, 48]. When exposure timing does not align with an LED’s illumination cycle, the source may be captured inconsistently across observations [17, 48]. Longer integration times and multi-exposure capture have been used to reduce such flashing artifacts by increasing the likelihood that illumination is observed during capture [17, 48]. Other known HDR failures include ghosting caused by scene motion [56] and halos or posterization introduced during tone mapping [24]. Emerging spatial HDR architectures, including split-diode, sub-pixel, and multi-tap designs, reduce the temporal exposure gap by capturing dynamic-range information concurrently rather than through sequential exposure brackets [20, 26, 53]; these architectures are therefore outside the scope of this work. Prior work thus establishes that temporal illumination variation can degrade HDR imaging, but does not formulate deliberate cross-exposure illumination inconsistency as a physical attack on temporal multi-exposure fusion. FLASH targets this fusion assumption inside the camera ISP, rather than relying on persistent sensor saturation, rolling-shutter readout, or optimization against a downstream perception model.
//anonymous.4open.science/r/Lights-Camera-Attac k-HDR-Manipulation-with-FLASH-Attacks-CB90/.
2
Background
Modern camera systems use ISPs to convert raw sensor readings into images for downstream perception models. As reported by Kühr et al. [31], ISPs usually include both hardware and software stages. Many of these pipelines are proprietary and not standardized, so they behave as black-box preprocessing before data reaches application-level algorithms. HDR and wide-dynamic-range processing are common components of modern camera pipelines, including automotive ISP designs and image sensors with on-chip HDR preprocessing [8, 37]. These stages extend dynamic range in scenes with strong brightness variation, such as indoor-outdoor transitions, night photography, and backlit surveillance views. To achieve this, temporal HDR-enabled cameras capture multiple exposures of the same scene, with different integration times, and combine them into one image. This principle follows early multi-exposure dynamic-range reconstruction work, where differently exposed pictures of the same scene are combined to recover a higher-dynamic-range representation [16, 35, 36]. In such pipelines, shorter exposures preserve highlight details, while longer exposures preserve shadow details. However, temporal HDR fusion assumes consistent illumination across bracketed exposures [16, 23]. In practice, flashing and other rapid lighting changes violate this condition and can degrade fusion quality because multi-exposure pipelines are designed for stable illumination [12,17,48,53]. Abrupt pulsed light can cause one or more exposures, particularly those with longer integration times, to receive substantially more injected light than others, creating cross-bracket inconsistency and potentially producing darkened fusion output. HDR controls vary across devices. Consumer cameras may not expose HDR as a user setting, surveillance cameras often place it in administrative interfaces, and automotive systems may integrate it into sensor or ISP configurations managed by the platform [7, 10, 38, 44, 47, 50]. Operational users may therefore be unable to disable HDR or disabling it can reduce visibility in low-light, backlit, and high-contrast scenes [23].
3
Related Work 4
Optical sensor attacks. Prior physical optical attacks commonly interfere with image acquisition or downstream perception. Light-based attacks use lasers, strobes, or infrared illumination to saturate sensors, introduce artifacts, or mislead perception models [11,18,21,52,58]. Continuous strong-light attacks cause broad frame-level overexposure through persistent illumination [32,40], while spatial optical attacks place or project structured perturbations into selected scene or image regions [18, 34]. These attacks primarily target sensor capture
Threat Model
System Model. We model the victim as a camera that forms temporal HDR output by capturing n ≥ 2 exposures with different integration times and combining them through an ISP fusion and tone-mapping pipeline [13]. Temporal HDR assumes that scene illumination remains sufficiently consistent across these exposure observations, as detailed in Appendix A. FLASH violates this assumption by injecting pulsed light during capture so that one or more exposures receive unequal 3
Table 1: Comparison of FLASH with prior physical optical attacks. Attack Class Camera blinding [21, 40] Infrared injection [52] Scene-wide flashing [42] Rolling-shutter attacks [30, 45, 57] Spatial optical attacks [18, 34] FLASH
External Only
No CameraSpecific Profiling
Model Agnostic
No Object Alignment
No Exact Phase Lock
Targets HDR Failure
✓ ✓ ✓ ✓ ✓ ✓
✓ × ✓ × (✓) ✓
✓ ✓ × (✓) × ✓
✓ (✓) ✓ (✓) × ✓
✓ ✓ ✓ (✓) ✓ (✓)†
× × × × × ✓
Terms. A checkmark indicates that the property holds, a cross that it does not, and a parenthesized checkmark that it varies across the cited attacks. "External only" means no access to the camera, ISP, or host system. "No camera-specific profiling" means no target-specific camera characterization. "Model agnostic" means no downstream model access or optimization. "No object alignment" means no pattern placement on a scene object or image region. "No phase lock" means no exact synchronization with sensor readout or exposure timing. Line of sight and sufficient optical power remain necessary. † FLASH still requires a pulse cadence that overlaps vulnerable exposure brackets.
injected illumination. Longer exposures generally provide greater opportunity for pulse overlap, but the attack does not require a particular exposure index to be corrupted. The resulting cross-exposure inconsistency can produce darkened, overexposed, clipped, or otherwise distorted HDR output. Attacker Goal. The attacker aims to corrupt the final camera output by inducing inconsistent exposure evidence during temporal HDR capture. The goal is to produce abnormal brightness, contrast, or visibility before the frame reaches downstream human or machine perception. We treat FLASH as a camera-output attack and do not assume control over, or knowledge of, any downstream perception model. Attacker Capabilities. The attacker controls only an external pulsed-light source and can approximately position and aim it toward the target camera. The attack is non-contact, openloop, and black-box. The attacker does not require physical access to the camera or host system, access to RAW exposure buffers, exposure registers, ISP firmware or settings, sensor trigger lines, internal telemetry, or the downstream perception model. Exact knowledge of the exposure schedule, fusion algorithm, or camera frame rate is not assumed, and exact phase synchronization with sensor capture is not required. Public or externally observable information, such as approximate camera placement and the use of temporal multi-exposure HDR, may be used to select a feasible attack configuration. Operational Constraints. FLASH requires line of sight between the optical source and the target camera and sufficient received illumination to produce unequal contamination across exposure observations. Attack feasibility therefore depends on source strength, camera-to-source distance, aiming angle, relative source height, ambient illumination, and temporal pulse overlap. These factors constrain the effective operating region of the attack and are evaluated experimentally in Section 6. The attack does not require continuous illumination or permanent modification of the environment. Successful execution requires only that the pulse train overlap the camera’s exposure sequence unevenly enough to create cross-exposure inconsistency; the specific timing relationship
is described in Section 5. Evaluation Scope. We evaluate FLASH at three levels. First, controlled simulation isolates how spatial, illumination, and temporal conditions affect HDR output degradation. Second, component-level physical experiments characterize the response of black-box embedded, surveillance, consumer, and smartphone camera pipelines using matched optical controls. Third, a controlled outdoor stationary comma 3X/OpenPilot case study evaluates camera-stream degradation and displayed path states using a speed-limit sign, pedestrian, traffic cone, and general road scene. These are evaluation settings rather than distinct attack mechanisms; all use the same external pulsed-light mechanism to create cross-exposure inconsistency before downstream perception. We additionally use the Raspberry Pi Camera Module 3, AI Camera, and HQ Camera as controllable component-level platforms for characterization, parameter sweeps, and baseline comparisons.
5 5.1
FLASH Attack Design Overview
FLASH targets cameras that form each temporal HDR output by fusing n ≥ 2 exposure observations captured with different integration times. Let B = {1, . . . , n} denote the set of exposure indices, and let Tb denote the integration time of the bth exposure, for b ∈ B . The attacker directs periodic bright optical pulses toward the camera during capture. Because exposure windows differ in duration and temporal offset, the pulses can overlap different exposures by different amounts. Sufficiently unequal injected energy creates crossexposure inconsistency that enters the proprietary fusion and tone-mapping pipeline and can produce darkening, overexposure, or visibility loss. Longer exposure windows generally provide greater opportunity for pulse overlap, but FLASH does not require a 4
A
Side View
Flashing Light
Top View
HDR Exposures Corrupted Output Frames
OFF
ON
…
ON
OFF
1 2 …
b
T1 T2
Tb
ON
… ON
…
OFF
ON
OFF
ON
OFF
ON
… ON
OFF
n
1 2 …
b
Tn
T1 T2
Tb
N
ON
… OFF
…
ON
OFF
ON
OFF
n Tn
N+1
time
Figure 3: Representative overview of FLASH timing across consecutive HDR frames. Each output frame N is formed by fusing n exposures with increasing integration times, T1 < T2 < · · · < Tn . Periodic light pulses overlap different subsets of these exposures, creating cross-exposure inconsistencies that corrupt the fused HDR output.
Figure 2: Physical setup and variables used in the FLASH evaluation. The side view shows the flashlight height z f , camera height zc , and horizontal camera-to-flashlight distance d f c . The flashlight emits pulses with strength Pf and frequency fflash , while the camera records at frame rate ffps under ambient-light condition A. The top view defines the flashlight yaw angle θ between its optical axis and the line from the flashlight to the camera. At θ = 0◦ , the flashlight points directly at the camera.
camera is: Pf g(θ) E f (aa, s ) ≈ κ 2 , d f c + (z f − zc )2 where g(θ) represents the directional beam response and κ captures fixed source and optical factors. The actual camera response is measured empirically because the source optics, sensor response, and ISP are treated as black boxes. Flashlight Height. Let z f ∈ R≥0 denote the flashlight height above the reference plane, with z f ∈ [zmin , zmax ]. Camera-to-Flashlight Horizontal Distance. Let d f c ∈ R>0 denote the horizontal distance between the flashlight and camera, with d f c ∈ [dmin , dmax ]. Flashlight Angle. Let θ ∈ R denote the yaw angle between the flashlight optical axis and the direction from the flashlight to the camera. Received irradiance is affected through g(θ), with θ ∈ [θmin , θmax ]. Ambient Illumination. We distinguish the ambient-light quantity used in simulation from that measured in physical experiments. In simulation, ambient lighting is controlled by area-light source power Pasim , while physical ambient illumiphys nance is measured at the camera plane as Ev,a in lux. These quantities are not treated as directly equivalent. Flash Emission Strength. Let Pf ∈ R≥0 denote the flashlight emission strength. In simulation, Pf denotes the nominal attack-source power in watts; in physical experiments, it denotes the corresponding source setting used by the hardware. It determines received flash irradiance together with distance, relative height, and aiming angle, with Pf ∈ [Pf ,min , Pf ,max ]. Temporal Interaction. The temporal interaction depends on camera frame rate ffps , flash frequency fflash , pulse width τ p , and relative phase φ. In our evaluation, fflash is attackerselected, ffps is an operating condition, the pulse width is fixed by the source configuration, and no exact phase lock is used. Temporal HDR implementations acquire n ≥ 2 observations during distinct exposure windows. These may be sequential frames or staggered intervals within a frame cycle. For exposure b ∈ B , Tb > 0 denotes its integration time. The number,
specific number of exposures or a particular exposure to be corrupted. The affected subset depends on the exposure schedule, pulse width, pulse frequency, and relative phase. The attack proceeds by placing and aiming the source from a feasible line-of-sight location, selecting a source strength and pulse cadence, and inducing unequal contamination across exposure observations. FLASH does not assume access to the exposure schedule, sensor readout, ISP, or fusion algorithm.
5.2
Attack Parameters
We distinguish attacker-selected settings from camera and environmental conditions. The attacker-selected configuration is a = [z f , d f c , θ, Pf , fflash ]⊤ , while s = [zc , A, ffps ]⊤ describes the operating condition. Camera height zc is fixed by the platform, while A denotes the ambient-light condition and ffps the frame rate. We represent A as Pasim in simulation and phys Ev,a in physical experiments. For a given s , the attacker selects a ∈ C (ss), where C (ss) contains configurations satisfying source-range, line-of-sight, hardware, and safety constraints. Appendix C formalizes these constraints. FLASH does not solve a white-box optimization problem; we instead evaluate controlled sweeps over feasible configurations. Spatial and Optical Parameters. We use a right-handed coordinate system with a horizontal reference plane and vertical axis z. In driving case studies, the reference plane is the road surface; in other settings, it is the local mounting plane. Figure 2 depicts this geometry. The target camera is mounted at height zc . We define the effective vertical offset and source-to-camera distance as: q zeff = z − z , r = d 2f c + (z f − zc )2 . f c f c f A first-order model of the flash irradiance received by the 5
Our physical evaluation considers Feval = {24, 30, 60, 120} fps. Their least common multiple gives fflash = 120 Hz, producing integer ratios of 5, 4, 2, and 1, respectively. The commercial flashlight controller reports cadence in RPM; under its cycles-per-minute convention, 120 Hz corresponds to 7200 RPM. For fractional or fluctuating rates such as 29.97 or 59.94 fps, relative phase drifts over time. We therefore treat the common-multiple cadence as a nominal selection rather than a guarantee of exact synchronization. Its effectiveness is established empirically in the evaluation.
duration, and temporal offsets of the exposures depend on the camera implementation. For staggered sensor HDR, exposure intervals for a sensor row may be separated by readout offsets: Willassen et al. report long and short exposures of 43 ms and 0.043 ms [53], while other commercial CMOS architectures report representative timings of TS ≈ 0.8 ms, TM ≈ 6.25 ms, and TL ≈ 50.0 ms [15]. These examples motivate modeling the sequence generically. For b, c ∈ B with Tb > Tc , the longer exposure provides greater opportunity for overlap under an unknown relative phase. Let ffps ∈ R>0 be the nominal camera frame rate, with output-frame start time: tk =
k , ffps
6
ffps ∈ [ ffps,min , ffps,max ].
We evaluate FLASH in three stages. First, simulation characterizes the physical and environmental conditions under which FLASH causes HDR output degradation. Second, physical camera experiments characterize how black-box camera pipelines respond relative to matched optical controls. Third, a controlled stationary comma 3X/OpenPilot case study evaluates camera-stream degradation and displayed path states. Evaluation Questions. In this section, we evaluate the following EQs: • EQ1: Under what physical and environmental conditions does FLASH cause camera-output degradation? We evaluate the effects of source distance, angle, height, strength, ambient-light condition, and capture frame rate through controlled simulation and physical parameter sweeps. • EQ2: How do black-box physical camera pipelines respond to FLASH relative to matched optical controls? We compare FLASH with clean, continuous-light, and randomized-frequency conditions across physical camera platforms and characterize the resulting response types using output-level luma and target-background contrast measurements. • EQ3: In controlled stationary comma 3X/OpenPilot trials, how do the camera stream and displayed path state differ under FLASH? We separately evaluate offline camera-stream degradation and the path state displayed by the OpenPilot’s own interface. We do not observe OpenPilot’s internal perception or planning states.
Let fflash ∈ R>0 denote pulse frequency, with Tflash =
1 , fflash
fflash ∈ [ fflash,min , fflash,max ].
For exposure b ∈ B associated with output frame k, let Wk,b = [tk + δb , tk + δb + Tb ] denote its exposure interval, where δb is its temporal offset. Let p(t; fflash , τ p , φ) ∈ [0, 1] denote the normalized pulse train. The injected flash energy is proportional to Z
Qk,b = E f (aa, s )
Wk,b
Experiments
p(t; fflash , τ p , φ) dt.
The spatial and optical settings determine E f , while frame rate, flash frequency, pulse width, relative phase, exposure duration, and exposure offset determine pulse overlap. FLASH seeks configurations for which Qk,b ̸= Qk,c for at least some b, c ∈ B , with sufficient difference to disturb temporal HDR fusion. Let Xk,b denote the clean measurement for exposure b, and let H denote the unknown temporal HDR fusion and tonemapping pipeline. The attacked and clean outputs are IkFLASH = H Xk,b + Qk,b b∈B , Ikclean = H Xk,b b∈B . Because H , the exposure offsets, exact exposure durations, and number of exposure observations are generally unavailable, FLASH does not compute an internal gradient or exact phase schedule. Section 6 instead measures the output effect through controlled parameter sweeps. Open-Loop Cadence Selection. For a stable nominal frame rate, selecting ffflash ∈ Z>0 causes the pulse pattern to repeat at fps the same nominal offset across output frames. This provides a repeatable cadence relationship but does not guarantee selective overlap with a particular exposure; actual overlap also depends on pulse width, relative phase, exposure durations, and offsets.
6.1
Evaluation Setup
Simulation Environment. We implement the simulation in Blender within a controlled indoor environment under daytime and nighttime illumination baselines. The evaluated attack-source power Pf ranges from 2 W to 20 W. Complete scene-construction details and the representative source classes associated with these power levels are provided in Appendix E. Component-Level Hardware Setup. We evaluate seven physical cameras as black-box imaging pipelines in a fixed in6
door setup. For the default trials, the camera and flash source are mounted at zc = z f = 1.5 m, the camera-to-source distance is d f c = 2 m, and the target board is placed 2 m from the camera. The manufacturer-rated Pf = 2 W source is aimed toward the camera and target region. Ambient condition A and control-light illuminance are measured at the camera plane, and floor markings maintain consistent geometry across devices. We record the final image or video output using the available or default capture mode. The physical default uses 2 W at 2 m, while the simulation uses 12 W at 4 m to isolate parameter effects without trivially saturating the closest condition. Table 2 summarizes the evaluated platforms and their roles. Complete setup details and a photograph are provided in Appendix F. System-Level comma 3X OpenPilot Setup. For the systemlevel case study, the comma 3X operates through its normal camera and OpenPilot pipeline in a controlled outdoor setting while the vehicle remains stationary. Matched target-present control and FLASH trials keep the experimenter, source, target, and vehicle in the same positions; the control keeps the source inactive, while the FLASH condition differs only in the emitted pulses. The source uses Pf = 2 W and is placed at d f c = 3 m, θ = 0◦ , and z f = 1.5 m. We evaluate a speed-limit sign, traffic cone, pedestrian, and general road scene. For each trial, we record whether OpenPilot displays a green planned path, a red blocked-path indication, or no path, and we analyze full-frame and target-region degradation from the exported camera stream.
6.2
consistent illumination, while the 4.0 m distance avoids the frequent sensor saturation observed at 2.0 m. Together, these settings produce a strong but non-saturated response, allowing the effect of optical power to be isolated. We then vary optical power across 2 W, 4 W, 8 W, 12 W and 20 W. Table 3 shows that nighttime attack strength increases monotonically with power, from −15.08% at 2 W to −80.20% at 20 W. Under daytime lighting, the same sweep produces only small changes, from 0.65% to 3.48%. Height sweep. Using the same non-saturating distance of 4.0 m, we fix the angle at 0◦ and power at 12 W, then vary source height across 0.5 m, 1.5 m and 3.0 m. Figure 5 shows that camera-level placement at 1.5 m gives the strongest response under both ambient conditions. At night, the score is −54.57% at 1.5 m, compared with −46.60% at 0.5 m and −45.33% at 3.0 m. Daytime results follow the same pattern with smaller changes. Simulation key findings. FLASH is strongest under low ambient illumination, at short d f c , near θ = 0◦ , at cameralevel z f , and with high Pf . The spatial sweep reaches extreme darkening near 2 m and 0◦ ; degradation rises from −15.08% at 2 W to −80.20% at 20 W and peaks at −54.57% for z f = 1.5 m. Ambient illumination, distance, angle, and power dominate the response.
6.3
Physical Camera Evaluation
The physical camera evaluation addresses EQ2 through matched clean, continuous-light, randomized-frequency, and FLASH conditions across black-box camera pipelines. We then organize the observed results by dominant response type. The parameter sweeps on the Raspberry Pi Camera Module 3 and iPhone 16 Pro further address EQ1 across distance, angle, source setting, measured ambient illuminance, and frame rate.
Simulation-Based Evaluation
The simulation evaluation addresses EQ1 by varying attack distance, horizontal angle, source height, optical power, and simulated ambient-light condition and measuring the resulting HDR output degradation after ISP processing. For each condition, we report a normalized output-degradation score relative to the matched no-attack baseline. Negative values indicate output darkening or compression, while positive values indicate brightening relative to the clean baseline. Scores at or below −100% indicate catastrophic HDR failure, where the simulated ISP output collapses to an extremely darkened frame. We interpret these cases as defined in Section 6. Spatial sweep. We first sweep distance (d f c ) and horizontal angle (θ). The flashlight is fixed at 12 W and 1.5 m, while distance varies from 2 m to 20 m and aim angle varies from −45◦ to 45◦ . Figure 4 shows that the strongest response occurs at short range and near 0◦ aim. The effect drops sharply after approximately 10 m and weakens as distance increases or the beam moves off axis. Nighttime results are substantially stronger, with extreme-darkening failures appearing only under low ambient illumination. Daytime results remain below about 8% even at the closest settings. Power sweep. We fix the geometry at 4.0 m, 0◦ , and 1.5 m. The 0◦ angle and camera-level source height provide direct,
6.3.1
Evaluation Design and Cross-Device Comparison
Black-box evaluation and controls. Because proprietary cameras do not expose RAW brackets or ISP internals, we treat physical devices as black-box pipelines and evaluate only their final image or video output. Each camera is tested under four matched default conditions: (1) clean baseline, (2) continuous light (matched to the measured illuminance of the FLASH source at the camera plane), (3) randomizedfrequency flashing light, and (4) FLASH. These controls separate timing-sensitive FLASH behavior from ordinary brightlight exposure and randomized-frequency flashing. Default cross-device comparison. We first run the four matched conditions across all seven physical cameras using the default setup. The default setup uses a camera-tosource distance of 2 m, attack angle of 0◦ , source height of 1.5 m, low ambient illuminance, and the high flashlight setting. The Raspberry Pi HQ Camera was operated in its default single-exposure mode and serves as the non-HDR baseline. Its 7
Table 2: Physical platforms and their roles in the FLASH evaluation. The first seven platforms are evaluated at the component physical level, while comma 3X is used for the system-level OpenPilot case study.
Raspberry Pi Camera Module 3 Platform Class Embedded
Raspberry Pi AI Camera Embedded
Raspberry Pi HQ Camera Embedded
Evaluation Use Mechanism tests; parameter sweeps
Cross-device comparison
Non-HDR baseline
Device
onn Indoor Camera Surveillance Surveillance camera evaluation
Wyze Battery Cam Pro Surveillance Surveillance camera evaluation
Kodak PixPro FZ45 Consumer Photography evaluation
iPhone 16 Pro
comma 3X
Smartphone AV system Smartphone Controlled outdoor evaluation; case study parameter sweeps
Figure 4: Simulation spatial sweep results for the flashlight attack. The source is fixed at 12 W and 1.5 m, while distance and horizontal aim angle are varied. The left plot shows the daytime baseline and the right plot shows the nighttime baseline. Table 3: Simulation power sweep results at 4.0 m distance, 0◦ angle, and 1.5 m height. Cond. Night Day
2W
4W
8W
12W
3.0m
1.97%
1.5m
2.36%
0.5m
2.02%
3.0m
-45.33%
20W
−15.1% −24.4% −40.3% −54.6% −80.2% 0.7% 1.1% 1.8% 2.4% 3.5%
IMX477 sensor supports DOL-HDR, but this mode is not implemented by the stock Raspberry Pi camera stack. Raspberry Pi 5 separately provides an optional ISP-based computational HDR mode that accumulates frames captured with the same exposure setting. This mode differs from the temporal multiexposure fusion targeted in this work and was not enabled during capture. We do not expect it to be immune to bright light because non-HDR ISP stages can still react to illumination changes. Instead, we use it to test whether the observed response is specific to HDR or HDR-relevant pipelines. Metric application and normalization. We use the signed luma-change, severe-darkening, extreme-darkening, and CNR metrics defined in Section 6 for the default cross-device physical-camera comparison. For the physical camera recordings, each non-clean condition is compared with a matched reference recording from the same camera and setup. The
1.5m
1.5m
0.5m
-54.57% 1.5m
-46.60%
Figure 5: Simulation height sweep results at 4.0 m distance, 0◦ angle, and 12 W flashlight power under daytime (left) and nighttime (right) ambient conditions. clean-normalized comparison uses the no-light baseline as the reference and is reported for completeness. To isolate timing-sensitive FLASH behavior, our main default crossdevice analysis uses control-normalized comparisons. In this analysis, FLASH is compared against continuous light and randomized-frequency flashing under the same default setup. This separates timing-dependent degradation from ordinary bright-light exposure and randomized-frequency flashing. For the parameter sweeps, we use clean versus FLASH pairs and report median signed decoded luma-unit change, because the 8
Table 4: Control-normalized full-frame response and dominant response under FLASH in the default setup.
Continuous: median Y' = 68.2
Randomized Flashing: median Y' = 81.5
FLASH: median Y' = 9.0
Frames Med. δY ′ Sev. Ext. Dominant response (%) (%) (%)
Camera FLASH vs. continuous light iPhone 16 Pro Kodak PixPro FZ45 onn Indoor Camera Raspberry Pi AI Camera Wyze Battery Cam Pro
1299 331 292 143 172
Raspberry Pi Module 3 Raspberry Pi HQ (non-HDR)
589 300
FLASH vs. randomized flashing iPhone 16 Pro 1299 Kodak PixPro FZ45 351 onn Indoor Camera 438 Raspberry Pi AI Camera 143 Wyze Battery Cam Pro 155 Raspberry Pi Module 3 Raspberry Pi HQ (non-HDR)
589 300
213.9 50.0 50.0 Intermittent collapse 159.1 0.0 0.0 Brightening/compensation -4.6 0.0 0.0 Localized degradation 350.9 0.0 0.0 Brightening/compensation -42.6 33.7 33.7 Full-frame dark collapse 317.6 0.0 0.0 Brightening/compensation -41.5 27.0 1.0 Mixed response
Figure 6: Four-condition Wyze Battery Cam Pro comparison. Continuous light and randomized-frequency flashing raise median Y ′ from 11.2 to 68.2 and 81.5, while FLASH reduces it to 9.0, which is darker than the clean scene (11.2), supporting timing-dependent collapse rather than bright-light interference.
10.4 25.1 24.9 Intermittent collapse -6.8 0.0 0.0 Mild darkening -5.2 0.0 0.0 Localized degradation 4.3 0.0 0.0 Brightening/compensation -46.8 33.5 33.5 Full-frame dark collapse 162.3 19.7 8.0 Intermittent collapse 269.4 30.3 6.3 Mixed response
Note: Severe darkening denotes frames with δY ′ ≤ −50%. Extreme darkening denotes frames with δY ′ ≤ −80%. Highlighted cells indicate nonzero scene-wide severedarkening or extreme-darkening failure rates.
dian luma, while FLASH reduces the output to a level close to the clean low-light baseline. This behavior indicates that the collapse is associated with pulse timing rather than bright illumination alone.
goal of the sweep is to compare trends across distance, angle, source setting, ambient illuminance, and FPS rather than to summarize collapse rates. Cross-device results. Table 4 shows that FLASH produces pipeline-dependent responses rather than uniform darkening. The strongest full-frame collapse appears on the iPhone 16 Pro and Wyze Battery Cam Pro. The Raspberry Pi Camera Module 3 shows timing-sensitive degradation relative to randomized-frequency flashing, while the Kodak PixPro FZ45 and Raspberry Pi AI Camera primarily show brightening or exposure-compensation behavior. The onn Indoor Camera shows localized degradation without full-frame collapse. We therefore organize the results below by response type. 6.3.2
Clean: median Y' = 11.2
The Wyze camera also exhibits a built-in response that is not captured by frame-level luma metrics alone. Across 10 repeated FLASH trials, all 10 trials activated the camera’s automatic spotlight or infrared behavior. The same response occurred in 0 out of 10 continuous-light trials and 0 out of 10 randomized-frequency trials. We retain this automatic behavior because spotlight and infrared modes are part of normal surveillance-camera operation. This result indicates that the Wyze pipeline interpreted the FLASH-affected scene as a low-visibility condition, while continuous and randomizedfrequency flashes did not achieve the same treatment. Raspberry Pi Camera Module 3. The Raspberry Pi Camera Module 3 shows timing-sensitive degradation mainly relative to randomized-frequency flashing. In that comparison, 19.7% of frames meet the severe-darkening threshold and 8.0% meet the extreme-darkening threshold.
Dark-Collapse and Timing-Sensitive Responses
iPhone 16 Pro. The iPhone 16 Pro exhibits intermittent framelevel collapse. Relative to continuous light, FLASH produces severe-darkening and extreme-darkening rates of 50.0%. Relative to randomized-frequency flashing, the corresponding rates are 25.1% and 24.9%. This indicates intermittent framelevel collapse, even though the median signed luma change is positive in both comparisons. Wyze Battery Cam Pro. The Wyze Battery Cam Pro exhibits strong full-frame collapse. Relative to continuous light, FLASH produces a median signed luma reduction of 42.6%, with severe-darkening and extreme-darkening rates of 33.7%. Relative to randomized-frequency flashing, the median signed luma reduction is 46.8%, with severe-darkening and extremedarkening rates of 33.5%. Figure 6 shows that continuous and randomized-frequency flashing light increase the decoded me-
Non-HDR interpretation boundary. The Raspberry Pi HQ Camera is treated as an interpretation boundary rather than as evidence of HDR-specific failure. It is a non-HDR camera, yet it shows severe-darkening rates of 27.0% and 30.3% in the two control-normalized comparisons. Its extreme-darkening rates remain much lower, at 1.0% and 6.3%, and its median response is inconsistent across controls: −41.5% relative to continuous light but +269.4% relative to randomized-frequency flashing. We therefore use it as evidence that strong optical stimulation can also disturb non-HDR camera output, while the strongest extreme-darkening collapse appears on HDR or HDR-relevant black-box pipelines in our evaluated set. 9
6.3.3
Brightening, Compensation, and Localized Responses
with proprietary exposure, tone-mapping, and capture-mode processing. Although response magnitude varies across capture modes, ∆Y ′ remains nonzero at every tested frame rate.
Raspberry Pi AI Camera. On the Raspberry Pi AI Camera, FLASH produces a brightening response. The full-frame median change is +350.9% relative to continuous light and +4.3% relative to randomized-frequency flashing. Its targetregion median changes are +398.1% and +9.4% under the same two comparisons. These results indicate exposurecompensation behavior rather than full-frame collapse. Kodak PixPro FZ45. On the Kodak PixPro FZ45, FLASH is much brighter than continuous light in the full-frame metric, with a median δY ′ of +159.1%, and it also increases targetregion luma by +185.6% relative to continuous light. Relative to randomized-frequency flashing, Kodak shows a small fullframe reduction of −6.8% and a stronger background-region reduction of −26.1%, although these changes do not reach the severe-darkening threshold. onn Indoor Camera. The onn Indoor Camera’s regional measurements show stronger localized degradation. Relative to continuous light, the gray-card region decreases by 38.8%, the target region decreases by 22.0%, and target-background CNR decreases by 48.8%. This ISP does not exhibit full-frame extreme-darkening under the tested conditions, although its median full-frame luma change is −4.6% relative to continuous light and −5.2% relative to randomized-frequency flashing. The onn results therefore indicate localized visibility and contrast loss rather than full-frame collapse.
6.4 System Case Study: comma 3X CameraStream Degradation and OpenPilot Path State The system case study addresses EQ3 through controlled stationary comma 3X/OpenPilot trials. The vehicle remains parked throughout each trial. We evaluate three conditions: an unobstructed baseline with no target or attack equipment, a target-present control without FLASH, and a target-present FLASH condition. During the FLASH condition, the source is 3 m directly in front of the forward-facing camera at 0◦ and a height of 1.5 m. This favorable stress-testing geometry is selected based on the preceding simulation and componentlevel evaluations. The evaluated targets are a speed-limit sign, an orange traffic cone, and a pedestrian, together with general road-scene visibility. Displayed path state. For each condition, we record whether OpenPilot displays a green planned path, a red blocked path, or no path. This is an interface-level observation; we do not observe OpenPilot’s internal perception or planning state. Offline camera-stream processing. Separately, we export the comma 3X camera streams and compare each FLASH video with its corresponding no-FLASH reference. Target-present trials use the target-present control as the reference, while the general road-scene trial uses the unobstructed baseline. We compute full-frame signed luma change, target-region signed luma change, target-background signed CNR change, severedarkening rate, extreme-darkening rate, and representative reference-versus-FLASH frames. The traffic-sign and cone targets are analyzed from the same video pair, so their fullframe metrics are reported once. Their target-region and CNR metrics are reported separately because they use different object ROIs. Table 6 summarizes these measurements. OpenPilot Displayed Path State. In the unobstructed baseline, OpenPilot displays a green planned path. When an object of interest is present without FLASH, OpenPilot displays a red blocked path. Under FLASH, it displays neither the green planned path nor the red blocked path. The missing displayed path is therefore distinct from the red blocked-path state observed in the target-present control. Cone Visibility. The cone target provides the strongest quantitative system-level degradation. Although the shared fullframe sign/cone scene does not reach the severe-darkening threshold, the cone ROI shows intermittent target-level collapse. The cone target ROI reaches a minimum δY ′ of −61.8%, and 23.0% of compared frames meet the severedarkening threshold. The target-background CNR also shows strong intermittent loss, with a worst-frame δCNR of −90.8% and 23.0% of frames below the −80% CNR-change threshold. This indicates that FLASH can substantially reduce the visibil-
Physical camera key findings. FLASH causes pipelinedependent responses. Extreme-darkening reaches 50.0%/24.9% on the iPhone and 33.7%/33.5% on the Wyze camera relative to continuous/randomized flashing. Wyze triggers its low-visibility response in 10/10 FLASH trials versus 0/10 controls. Other devices show collapse, brightening, compensation, or localized degradation shaped by geometry, ambient illumination, and capture mode. 6.3.4
Operational Robustness Evaluation
To further address EQ1, we perform matched clean-versusFLASH physical parameter sweeps on the iPhone 16 Pro and Raspberry Pi Camera Module 3, the two evaluated devices supporting selectable Frame Per Second (fps) modes. We vary camera-to-source distance d f c , aiming angle θ, source setting Pf , measured ambient illuminance A, and camera frame rate ffps . We report median signed decoded luma-unit change ∆Y ′ because percent-normalized changes become unstable when low-light clean baselines approach zero. Table 5 summarizes the results; the dataset composition, complete sweep plot, and per-setting analysis are provided in Appendix G. Distance and off-axis aiming attenuate the response on both cameras, while increased ambient illuminance strongly suppresses the low-light response. Source setting and frame rate produce non-monotonic, device-dependent results, consistent 10
Table 5: Physical operational robustness under matched clean-versus-FLASH trials. Entries report median signed decoded luma-unit change ∆Y ′ for the settings listed in order. Parameter
Settings
iPhone 16 Pro
Module 3
Trend
Distance d f c Aiming angle θ Source setting Pf Ambient illuminance A Frame rate ffps
2, 4, 6, 8 m 0◦ , 10◦ , 20◦ , 30◦ 25%, 50%, 100% 0, 27.3, 239 lux 24, 30, 60, 120 fps
60.60, 28.85, 16.90, 12.62 60.60, 41.15, 17.53, 14.66 39.19, 52.59, 32.94 60.44, −6.12, −4.51 67.39, 65.87, 45.43, 52.61
2.82, 0.74, 0.29, 0.42 2.82, 1.92, 0.46, 0.50 2.18, 0.71, 2.60 5.98, 1.00, −0.12 N/A, 12.03, 6.56, 11.83
Decreases with distance. Strongest near axis. Non-monotonic. Suppressed by ambient illuminance. Capture-mode dependent.
Table 6: Offline comma 3X camera-stream degradation under FLASH relative to the corresponding no-FLASH references. Full-frame metrics capture scene-level effects, while target ROI and CNR metrics capture object-level visibility. Highlighted cells indicate the strongest observed effects. Target/Condition
Meas.
Frames Med. δY ′ (%) Min. δY ′ (%) Sev. (%) Min. δCNR (%)
Shared sign/cone scene Full frame 100 Speed-limit sign Target ROI 100 Cone Target ROI 100 Pedestrian Target ROI 560 Road-scene visibility Full frame 120–560
86.6 101.9 234.2 305.3 63.5–73.0
ity and separability of a road object in the exported comma 3X stream, even when the full-frame average does not collapse. Figure 7a shows the representative clean-versus-FLASH sign and cone example used for this system-level visualization. Road-Scene Visibility. The general road-scene trials show that FLASH introduces visible glare, flare, and local visibility obstruction in the comma 3X stream. Figure 8 shows the corresponding in-car OpenPilot view during the controlled evaluation, where the displayed comma 3X camera stream contains a dark horizontal band under FLASH. Across three FLASH trials, full-frame median δY ′ remains positive, ranging from +63.5% to +73.0%, because the attack source adds substantial light to the scene. The worst full-frame values also remain positive, ranging from +56.6% to +65.2%. Pedestrian Visibility. The pedestrian trial evaluates FLASHinduced artifacts around a person in the deployed comma 3X front-camera stream. Here, the dominant effect is not target-region darkening but glare and local saturation near the injected light source. The pedestrian ROI becomes brighter under the luma metric, with a median δY ′ of +305.3%, a minimum δY ′ of +204.4%, and 0.0% severe-darkening or extreme-darkening frames. The target-background CNR also increases, with a median δCNR of +313.0%. We therefore report this trial as a glare and saturation artifact case rather than an extreme-darkening failure case. Figure 7b shows the representative pedestrian example, where FLASH introduces glare and a dark horizontal band that reduces local road visibility. Traffic Sign Visibility. The traffic-sign trial evaluates a speedlimit sign in the exported comma 3X front-camera stream. In this case, the dominant effect is not target-region darken-
-38.7 21.4 -61.8 204.4 56.6–65.2
0.0 0.0 23.0 0.0 0.0
– 44.2 -90.8 210.6 –
ing but optical interference in the surrounding camera stream. The representative frames show glare, lens flare, and localized visibility obstruction near the injected light source. The selected speed-limit sign ROI becomes brighter under the luma metric, with a median δY ′ of +101.9%, a minimum δY ′ of +21.4%, and 0.0% severe-darkening or extreme-darkening frames. We therefore treat this trial as a scene-visibility artifact case rather than an extreme-darkening failure case. The traffic-sign portion of Figure 7a shows this effect qualitatively, with reduced sign-region visibility under FLASH. System-level key findings. OpenPilot displays a green path in the unobstructed baseline and a red blocked path in the target-present control, but neither under FLASH. Offline analysis shows scene-dependent artifacts, including 23.0% severe cone-region darkening and a worst-case CNR loss of 90.8%. The sign and pedestrian trials instead show glare, lens flare, and local visibility obstruction.
7
Proof-of-Concept Mitigation
FLASH creates cross-bracket inconsistency by causing one or more exposures in a temporal HDR sequence to receive illumination that is inconsistent with the rest of the bracket. Because this inconsistency is before exposure fusion, mitigation can act on the exposure evidence before the final frame is formed. We evaluate a proof-of-concept exposure-level rejection method in a controlled three-exposure setting in which corruption is injected into the long exposure. This allows us to test whether removing anomalous exposure evidence before 11
Cone Region Obscured
Reduced Road Visibility CLEAN (Headlights ON)
FLASH (Headlights ON)
CLEAN (Headlights ON)
(a) Traffic-sign and cone trial.
}
Dark Band and Glare Reduces Road Visibility.
Sign Region Affected by FLASH
Glare and Dark Band on Pedestrian Region, Pedestrian Obscured
FLASH (Headlights ON)
(b) Pedestrian trial.
Figure 7: Representative comma 3X unobstructed baseline reference and FLASH frames. Dashed boxes mark the evaluated target regions. In the sign and cone trial, FLASH reduces traffic-sign visibility and obscures the cone region. In the pedestrian trial, FLASH introduces glare and a dark horizontal band that reduces local road visibility. Quantitative results summarized in Table 6.
Dark
} Horizontal Band.
FLASH reduces Comma 3X’s road visibility with a dark horizontal band.
Figure 8: In-car OpenPilot view during the controlled evaluation. Under FLASH, the comma 3X stream contains a dark horizontal band that reduces local road visibility, and the interface displays neither a green planned path nor a red blocked path. Quantitative camera-stream results reported in Table 6.
Corrupted long exposure fused Visual Artifacts Present Median |δY'| = 60.46%
Corrupted long exposure rejected Visual Artifacts Improved Median |δY'| = 12.60%
(a) Undefended FLASH
(b) Defended against FLASH
Figure 9: Exposure-level rejection in the controlled nightonly HDR stress test. Rejecting the corrupted long exposure before fusion succeeds in all 756 cases with a 0.0% clean false rejection rate and reduces median absolute output-luma deviation from 60.46% to 12.60% (79.16%).
fusion can reduce the resulting reconstruction deviation. To support controlled analysis of internal HDR behavior, we also developed HDR-Lab, a modular open-source computational photography framework. Unlike the proprietary black-box camera pipelines used in the physical evaluation, HDR-Lab exposes intermediate stages and allows controlled manipulation of individual exposure inputs. Implementation details and the access link are provided in Appendix B. Temporal exposure rejection. We propose an exposure-level anomaly filter that operates inside the HDR reconstruction pipeline before exposure fusion. Conceptually, the filter tests the exposures in a bracket for lighting inconsistencies. In our evaluated implementation, each bracket contains short, medium, and long exposures. For each exposure, the filter computes a luminance statistic and compares it with the corresponding clean or recent reference behavior. If one exposure produces an abnormal luminance deviation, the filter flags that exposure as corrupted. A flagged exposure is excluded from fusion or assigned a near-zero fusion weight, and the output frame is reconstructed from the remaining exposures. In our three-exposure implementation, the defense rejects at most one corrupted exposure and preserves the rest of the bracket. This differs
from whole-frame rejection because the output frame can still be reconstructed from the remaining exposure evidence. The defense targets the cross-bracket inconsistency exploited by FLASH. If an anomalous exposure remains in the bracket, it can influence reconstruction before subsequent ISP stages render the final frame. Exposure rejection instead removes the detected inconsistent exposure before fusion. Our evaluation tests this principle specifically for corruption of the long exposure in a three-exposure bracket. Controlled defense evaluation. Commercial cameras do not expose their internal exposure buffers, so we evaluate the defense in a controllable HDR reconstruction pipeline rather than inside proprietary on-chip ISP firmware. We use the night EXR simulation scenes as radiance inputs and construct synthetic three-exposure HDR brackets containing short, medium, and long exposures. The stress test injects corruption only into the long exposure. Then the defense system applies exposure-level rejection before fusion, and we compare the reconstructed FLASH output before and after defense. Across 756 night-only stress-test cases, the detector rejects the corrupted long exposure in 756/756 cases, with 12
a 0.0% clean false rejection rate. The defense reduces the median absolute luminance deviation from 60.46% before defense to 12.60% after defense, a 47.86 percentage-point reduction and a 79.16% relative reduction. The unoptimized Python/OpenCV prototype adds 108.85 ms median latency per HDR bracket, so this timing result is implementationspecific. This evaluation does not claim modification of closed commercial ISP implementations or robustness to arbitrary exposure counts or corruption patterns. It evaluates whether exposure-level rejection can reduce FLASH-induced reconstruction deviation in the tested three-exposure, singlecorruption setting. Additional mitigation directions. Two additional systemlevel approaches may reduce the reliability of FLASH, although we do not evaluate them experimentally. First, capturetiming dither could introduce small randomized shifts in frame or bracket timing, making it harder for an open-loop pulsed source to repeatedly produce the same overlap pattern across exposure intervals. Such a method would require camera-driver or firmware support and may affect systems that expect stable capture timing. Second, cameras equipped with infrared illumination or automatic spotlight modes may reduce the relative effect of injected visible light by increasing scene illumination. This approach is hardware- and scenedependent and should be treated as deployment-specific hardening rather than a general defense. Scope of the mitigation evaluation. Our defense evaluation considers three-exposure brackets in which corruption is injected only into the long exposure and at most one exposure is rejected. Although the rejection principle is intended to detect inconsistent exposure evidence more generally, we do not evaluate arbitrary numbers of exposures, corruption of other exposure positions, or simultaneous corruption of multiple exposures. The method also requires the corrupted exposure to exceed the anomaly threshold and at least one remaining exposure to preserve scene information. Abrupt benign lighting changes may also produce cross-bracket inconsistencies, and rejecting an exposure can reduce HDR reconstruction quality because fewer measurements remain for fusion.
ities may preserve evidence of pedestrians or obstacles and reduce risk, especially with redundant coverage, uncertainty estimation, sensor-health monitoring, and outlier rejection. However, fusion does not guarantee immunity: cameras remain important for lane markings, sign content, traffic-light state, object appearance, and other semantic information that radar or LiDAR may not recover. A system may also continue weighting a corrupted camera stream if it remains syntactically valid and triggers no failure indicator. Vehicle-level risk therefore depends on sensor coverage, fusion architecture, modality weighting, and failure handling, with cameradominant systems and cases where secondary sensors provide incomplete evidence remaining more exposed. Our systemlevel evaluation demonstrates degradation in the exported comma 3X camera stream and does not establish failure of object detection, sensor fusion, planning, or vehicle control. Evaluating whether radar or LiDAR fusion suppresses or propagates the effects of FLASH requires an end-to-end multimodal vehicle platform and is left for future work. Attack Scope and Evaluation Limitations. FLASH targets temporal HDR fusion and is not expected to transfer directly to single-exposure spatial HDR. Architectures such as split-diode pixels, sub-pixel sensors, and programmable multi-tap sensors reduce the timing gap exploited by the attack [20, 26, 53]. Our evaluation assumes non-contact lineof-sight deployment, low-to-moderate ambient lighting, and feasible emitter placement within safety and regulatory constraints. The tested devices and environments are limited, and commercial cameras are evaluated as black-box pipelines because RAW brackets, fusion weights, and proprietary ISP decisions are not exposed. We therefore rely on matched controls and a non-HDR baseline to support the timing-dependent interpretation. The defense is evaluated in a controllable HDR pipeline rather than inside proprietary camera firmware.
9
We introduced FLASH, a physical attack that exploits the stable-illumination assumption of temporal HDR fusion by creating inconsistent illumination across sequential exposures before downstream perception. Through simulation, matched optical controls, and physical evaluation across camera platforms, we show that this cross-bracket inconsistency produces pipeline-dependent failure modes, including darkening, overexposure, and local visibility degradation. The strongest physical results include extreme-darkening output rates of 50.0% on the iPhone 16 Pro and 33.7% on the Wyze Battery Cam Pro, while the controlled comma 3X case study shows up to a 90.8% reduction in target-background CNR in the trafficcone target region. In a controlled three-exposure night-only HDR stress test, proof-of-concept exposure rejection reduces median absolute output-luma deviation by 79.16%. These findings identify temporal HDR fusion itself as a physical attack surface and show that camera pipelines should vali-
Defense Finding. In the controlled three-exposure nightonly HDR stress test, exposure-level rejection detects the corrupted long exposure in 756/756 cases with a 0.0% clean false rejection rate and reduces median absolute luma deviation from 60.46% to 12.60%, a 79.16% relative reduction. The result applies to the evaluated single-corruption setting; broader exposure configurations and corruption patterns remain unevaluated.
8
Conclusion
Discussion and Limitations
Sensor Fusion and Vehicle-Level Safety. FLASH directly corrupts the camera stream without affecting independent radar or LiDAR sensing. In multi-sensor systems, these modal13
date cross-exposure consistency before accepting exposure evidence for fusion.
[5] Texas penal code § 28.03: Criminal mischief. Texas Legislature Online. Accessed 2025-06-26. URL: https: //statutes.capitol.texas.gov/Docs/PE/htm/P E.28.htm.
Ethical Considerations
[6] Texas penal code § 28.03, criminal mischief. FindLaw. Accessed 2025-06-26. URL: https://codes.findla w.com/tx/penal-code/penal-sect-28-03/.
This research was conducted solely to assess theoretical vulnerabilities and improve camera imaging-pipeline safety; no actual interference in open traffic, private property, photography sessions, or public roadways was performed. Pedestrian examples were recorded with consent, and identifying facial regions were concealed and not published for anonymity.
[7] Apple Inc. Adjust HDR camera settings on iPhone. Apple Support, iPhone User Guide, n.d. Accessed July 11, 2026. URL: https://support.apple.com/guid e/iphone/adjust-hdr-camera-settings-iph2c afe2ebc/ios.
Open Science
[8] Arm. Arm Mali-C71AE: High Performance Image Signal Processing with Advanced Safety. https: //developer.arm.com/community/arm-community -blogs/b/embedded-and-microcontrollers-blo g/posts/arm-mali-c71ae-image-signal-proce ssing-advanced-safety, 2020. Accessed: 2026-0506.
To facilitate reproducibility and encourage further exploration of algorithmic vulnerabilities in ISP pipelines, we provide the complete source code for our experimental framework. This release includes the full HDR-Lab implementation and the Blender-based physical simulation parameter optimization environment presented in this work. https://anonymous. 4open.science/r/Lights-Camera-Attack-HDR-Manip ulation-with-FLASH-Attacks-CB90/.
[9] Axis Communications AB. Wide dynamic range. White paper. Accessed July 28, 2026. URL: https://whit epapers.axis.com/en-us/wide-dynamic-range.
Acknowledgment
[10] Axis Communications AB. AXIS Q6325-LE PTZ Camera user manual. Axis Documentation, n.d. Section “Handle scenes with strong backlight.” Accessed July 11, 2026. URL: https://help.axis.com/en-us/ax is-q6325-le.
This work was supported by Clemson University’s Virtual Prototyping of Autonomy Enabled Ground Systems (VIPRGS), under Cooperative Agreement W56HZV-21-2-0001 with the US Army DEVCOM Ground Vehicle Systems Center (GVSC).
[11] Sri Hrushikesh Bhupathiraju, Takeshi Sugawara, Takami Sato, Qi Alfred Chen, Michael Clifford, and Sara Rampazzi. On the vulnerability of traffic light recognition systems to laser illumination attacks. In ISOC Symposium on Vehicle Security and Privacy (VehicleSec). ISOC, 2024. doi:10.14722/vehiclesec.2024.230 24.
References [1] 18 u.s.c. § 33: Destruction of motor vehicles or motor vehicle facilities. Cornell Law School, Legal Information Institute. Accessed 2025-06-26. URL: https: //www.law.cornell.edu/uscode/text/18/33.
[12] Michael Brading, Brian Keelan, and Hieu Tran. Image sensors for camera monitor systems, 2016. doi:10.1 007/978-3-319-29611-1_5.
[2] California penal code § 417.25: Aiming or pointing a laser scope. California Legislative Information. Accessed 2025-06-26. URL: https://leginfo.legisl ature.ca.gov/faces/codes_displaySection.xh tml?sectionNum=417.25&lawCode=PEN.
[13] Prashant Chaudhari, Franziska Schirrmacher, Andreas Maier, Christian Riess, and Thomas Köhler. Mergingisp: Multi-exposure high dynamic range image signal processing. In Christian Bauckhage, Juergen Gall, and Alexander Schwing, editors, Pattern Recognition, page 328–342. Springer International Publishing, 2021.
[3] California penal code § 417.27: Laser pointers at moving vehicles. California Legislative Information. Accessed 2025-06-26. URL: https://leginfo.legislature. ca.gov/faces/codes_displaySection.xhtml?se ctionNum=417.27&lawCode=PEN.
[14] Zhaojie Chen, Puxi Lin, Zoe Lin Jiang, Zhanhang Wei, Sichen Yuan, and Junbin Fang. An illumination modulation-based adversarial attack against automated face recognition system. In Information Security and Cryptology, volume 12612 of Lecture Notes
[4] New york penal law § 145.14: Criminal tampering in the third degree. New York State Senate. Accessed 2025-06-26. URL: https://www.nysenate.gov/leg islation/laws/PEN/145.14. 14
in Computer Science, pages 53–69. Springer, 2021. doi:10.1007/978-3-030-71852-7_4.
range and low-light imaging on mobile cameras. ACM Transactions on Graphics (TOG), 35(6):1–12, 2016.
[15] Yiheng Chi, Xingguang Zhang, and Stanley H. Chan. Hdr imaging with spatially varying signal-to-noise ratios, 2023. URL: https://arxiv.org/abs/2303.1 7253, arXiv:2303.17253.
[24] Alain Horé and Orly Yadid-Pecht. A new filter for reducing halo artifacts in tone mapped images. In 2014 22nd International Conference on Pattern Recognition, pages 889–894, 2014. doi:10.1109/ICPR.2014.163.
[16] Paul E. Debevec and Jitendra Malik. Recovering high dynamic range radiance maps from photographs. In Proceedings of the 24th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’97, page 369–378, USA, 1997. ACM Press/Addison-Wesley Publishing Co. doi:10.1145/258734.258884.
[25] Keqi Huang and Fanqiang Lin. Lights flicker: Furtive attack of facial recognition systems. In Proceedings of the 2022 5th International Conference on Telecommunications and Communication Engineering, pages 208–213. ACM, 2022. doi:10.1145/3577065.3577103.
[17] Brian Deegan. Led flicker: Root cause, impact and measurement for automotive imaging applications. Electronic Imaging, 30(17):146–1–146–1, 2018. URL: https://library.imaging.org/ei/articles/ 30/17/art00003, doi:10.2352/ISSN.2470-1173. 2018.17.AVM-146.
[26] S. Iida, Y. Sakano, T. Asatsuma, M. Takami, I. Yoshiba, N. Ohba, H. Mizuno, T. Oka, K. Yamaguchi, A. Suzuki, K. Suzuki, M. Yamada, M. Takizawa, Y. Tateshita, and K. Ohno. A 0.68e-rms random-noise 121db dynamicrange sub-pixel architecture cmos image sensor with led flicker mitigation. In 2018 IEEE International Electron Devices Meeting (IEDM), pages 10.2.1–10.2.4, 2018. doi:10.1109/IEDM.2018.8614565.
[18] Ranjie Duan, Xiaofeng Mao, A. K. Qin, Yuefeng Chen, Shaokai Ye, Yuan He, and Yun Yang. Adversarial laser beam: Effective physical-world attack to dnns in a blink. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16057–16066. IEEE, 2021. doi:10.1109/CVPR46437.2021.01580.
[27] Prabu Kumar. Key embedded vision applications of hdr cameras, 2021. Accessed July 28, 2026. URL: https://www.e-consystems.com/blog/camera/ technology/key-embedded-vision-application s-of-hdr-cameras/.
[19] Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physicalworld attacks on deep learning models, 2018. URL: https://arxiv.org/abs/1707.08945, arXiv: 1707.08945.
[28] Alexey Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world, 2017. URL: https://arxiv.org/abs/1607.02533, arXiv: 1607.02533. [29] Sebastian Köhler, Richard Baker, and Ivan Martinovic. Signal injection attacks against ccd image sensors. In Proceedings of the 2022 ACM on Asia Conference on Computer and Communications Security, pages 294– 308. ACM, 2022. doi:10.1145/3488932.3497771.
[20] Yu Feng, Keiichiro Kagawa, Kamel Mars, Keita Yasutomi, and Shoji Kawahito. Programmable dynamic range hdr imaging with led-flicker and motion artifact mitigation using a four-tap cmos image sensor. IEEE Sensors Journal, 25(10):18302–18311, 2025. doi: 10.1109/JSEN.2025.3557801.
[30] Sebastian Köhler, Giulio Lovisotto, Simon Birnbach, Richard Baker, and Ivan Martinovic. They see me rollin’: Inherent vulnerability of the rolling shutter in cmos image sensors. In Annual Computer Security Applications Conference, pages 399–413. ACM, 2021. doi:10.1145/3485832.3488016.
[21] Zhangjie Fu, Yueyan Zhi, Shouling Ji, and Xingming Sun. Remote attacks on drones vision sensors: An empirical study. IEEE Transactions on Dependable and Secure Computing, 19(5):3125–3135, 2022. doi: 10.1109/TDSC.2021.3085412.
[31] Michael Kühr, Mohammad Hamad, Pedram MohajerAnsari, Mert Pesé, and Sebastian Steinhorst. Sok: Security of the image processing pipeline in autonomous vehicles, 09 2024. doi:10.48550/arXiv.2409.01234.
[22] Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples, 2015. URL: https://arxiv.org/abs/1412.6 572, arXiv:1412.6572.
[32] Junjian Li and Honglong Chen. Adversarial raw: Imagescaling attack against imaging pipeline. arXiv preprint arXiv:2206.01733, 2022. URL: http://arxiv.org/ abs/2206.01733.
[23] Samuel W Hasinoff, Dillon Sharlet, Ryan Geiss, Andrew Adams, Jonathan T Barron, Florian Kainz, Jiawen Chen, and Marc Levoy. Burst photography for high dynamic 15
[33] Junjian Li, Honglong Chen, Zhichen Ni, Yudong Gao, Weifeng Liu, and Nan Jiang. Image-scaling attack on image signal processing pipelines in deep neural networks based outdoor vision applications. IEEE Transactions on Consumer Electronics, pages 1–1, 2024. doi:10.1109/TCE.2024.3413720.
[41] Buu Phan, Fahim Mannan, and Felix Heide. Adversarial imaging pipelines. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16046–16056. IEEE, 2021. doi:10.1109/CVPR 46437.2021.01579. [42] Roi Pony, Itay Naeh, and Shie Mannor. Over-the-air adversarial flickering attacks against video recognition networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 515–524, June 2021.
[34] Yanmao Man, Ming Li, and Ryan Gerdes. GhostImage: Remote perception attacks against camera-based image classification systems. In 23rd International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2020), pages 317–332, San Sebastian, October 2020. USENIX Association. URL: https://www.usenix.o rg/conference/raid2020/presentation/man.
[43] M.A. Robertson, S. Borman, and R.L. Stevenson. Dynamic range improvement through multiple exposures. In Proceedings 1999 International Conference on Image Processing (Cat. 99CH36348), volume 3, pages 159– 163 vol.3, 1999. doi:10.1109/ICIP.1999.817091.
[35] Steve Mann and Rosalind W. Picard. Being undigital with digital cameras: Extending dynamic range by combining differently exposed pictures. Technical Report 323, M.I.T. Media Lab Perceptual Computing Section, Cambridge, MA, USA, 1994. Also appeared in IS&T’s 48th Annual Conference, Cambridge, Massachusetts, May 1995, pp. 422–428.
[44] Samsung Electronics. New galaxy S22 camera updates let you capture the stars like a pro. Samsung Global Newsroom, October 2022. Published October 26, 2022. Accessed July 11, 2026. URL: https://news.samsu ng.com/global/new-galaxy-s22-camera-updat es-let-you-capture-the-stars-like-a-pro.
[36] T. Mertens, J. Kautz, and F. Van Reeth. Exposure fusion: A simple and practical alternative to high dynamic range photography. Computer Graphics Forum, 28(1):161– 171, 2009. URL: https://onlinelibrary.wiley. com/doi/abs/10.1111/j.1467-8659.2008.01171 .x, arXiv:https://onlinelibrary.wiley.com/ doi/pdf/10.1111/j.1467-8659.2008.01171.x, doi:10.1111/j.1467-8659.2008.01171.x.
[45] Athena Sayles, Ashish Hooda, Mohit Gupta, Rahul Chatterjee, and Earlence Fernandes. Invisible perturbations: Physical adversarial examples exploiting the rolling shutter effect. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14661–14670. IEEE, 2021. doi:10.1109/CVPR46437. 2021.01443.
[37] OMNIVISION. OX08B40 8.3-Megapixel Automotive Image Sensor Product Brief. https://www.ovt.co m/wp-content/uploads/2024/07/OX08B40-PB-v1. 1-WEB.pdf, 2024. Accessed: 2026-05-06.
[46] Aru Ranjan Singh, Thomas Bashford-Rogers, Demetris Marnerides, Kurt Debattista, and Sumit Hazra. HDR image-based deep learning approach for automatic detection of split defects on sheet metal stamping parts. The International Journal of Advanced Manufacturing Technology, 125:2393–2408, 2023. doi:10.1007/s0 0170-022-10763-6.
[38] OMNIVISION Technologies, Inc. OX08B40: 8.3megapixel automotive image sensor with LED flicker mitigation and 140 db high dynamic range. Product Brief Version 1.1, OMNIVISION Technologies, Inc., July 2024. URL: https://www.ovt.com/wp-conte nt/uploads/2024/07/OX08B40-PB-v1.1-WEB.pdf.
[47] Sony Corporation. Auto HDR. ILCE-7M3 α7 III Help Guide, 2018. Accessed July 11, 2026. URL: https: //helpguide.sony.net/ilc/1720/v1/en/conten ts/TP0001653148.html.
[39] Tatsuya Oyama, Kota Yoshida, Shunsuke Okura, and Takeshi Fujino. Adversarial examples created by fault injection attack on image sensor interface. IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, E107.A(3):344–354, 2024. doi:10.1587/transfun.2023CIP0025.
[48] Sony Semiconductor Solutions Group. LED Flicker Mitigation. https://www.sony-semicon.com/en/te chnology/automotive/lfm.html, 2024. Accessed: 2026-05-06.
[40] Jonathan Petit, Bas Stottelaar, and Michael Feiri. Remote attacks on automated vehicles sensors : Experiments on camera and lidar. 2015. URL: https://api. semanticscholar.org/CorpusID:39608826.
[49] Isao Takayanagi and Rihito Kuroda. Hdr cmos image sensors for automotive applications. IEEE Transactions on Electron Devices, 69(6):2815–2823, 2022. doi:10 .1109/TED.2022.3164370. 16
[50] Texas Instruments. Automotive ADAS reference design for four camera hub with integrated ISP and DVP outputs. Technical Report TIDUCB9, Texas Instruments, September 2016. URL: https://www.ti.com/lit/u g/tiducb9/tiducb9.pdf.
[58] Zhe Zhou, Di Tang, Xiaofeng Wang, Weili Han, Xiangyu Liu, and Kehuan Zhang. Invisible mask: Practical attacks on face recognition with infrared. arXiv preprint arXiv:1803.04683, 2018. URL: http://arxiv.org/ abs/1803.04683.
[51] U.S. Department of Justice. Criminal resource manual #1426: Destruction of motor vehicles. U.S. Department of Justice. Accessed 2025-06-26. URL: https://ww w.justice.gov/archives/jm/criminal-resourc e-manual-1426-destruction-motor-vehicles.
A
Vulnerability Analysis of the HDR Pipeline
Global assumptions. We assume the camera operates in the linear response regime with fixed gain; RAW values are available before tone curves or clipping unless stated otherwise; the flash illuminates the same pixel set R in every exposure; and the flash adds radiance E f > 0 without hard clipping except where noted. These assumptions isolate FLASH from unrelated ISP effects. Temporal HDR fusion model. Temporal HDR captures a sequence of exposures with different integration times, typically short, medium, and long [16, 36, 43]. For exposure j, the observed pixel value can be written as I j (x) = ∆t j E(x), where E(x) is scene radiance and ∆t j is exposure time. Under FLASH, one exposure receives an additional radiance term, Ik′ (x) = ∆tk (E(x) + E f (x)). Because longer exposures have larger integration windows, they generally have a greater probability of overlapping a flash pulse. This can create crossbracket inconsistency when one or more exposures receive substantially more injected energy than others, causing the exposure sequence to represent inconsistent scene illumination. Radiance bias. In weighted radiance reconstruction, per-pixel radiance is estimated from multiple exposures as
[52] Wei Wang, Yao Yao, Xin Liu, Xiang Li, Pei Hao, and Ting Zhu. I can see the light: Attacks on autonomous vehicles using invisible lights. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pages 1930–1944. ACM, 2021. doi:10.1145/3460120.3484766. [53] Trygve Willassen, Johannes Solhusvik, Robert Johansson, Sohrab Yaghmai, Howard E. Rhodes, Sohei Manabe, Duli Mao, Zhiqiang Lin, Dajiang Yang, Orkun Cellek, Eric A. G. Webster, Siguang Ma, and Bowei Zhang. A 1280 x 1080 4 . 2 µ m split-diode pixel hdr sensor in 110 nm bsi cmos process. 2015. URL: https://api.semanticscholar.org/CorpusID: 13190491. [54] Yuliang Wu, Ganchao Tan, Jinze Chen, Wei Zhai, Yang Cao, and Zheng-Jun Zha. Event-based asynchronous hdr imaging by temporal incident light modulation. Opt. Express, 32(11):18527–18538, May 2024. URL: https: //opg.optica.org/oe/abstract.cfm?URI=oe-3 2-11-18527, doi:10.1364/OE.520808.
Ê(x) =
∑ j w j (x)I j (x)/∆t j . ∑ j w j (x)
If exposure k is flashed and not rejected, the estimate becomes
[55] Qixue Xiao, Yufei Chen, Chao Shen, Yu Chen, and Kang Li. Seeing is not believing: Camouflage attacks on image scaling algorithms. In 28th USENIX Security Symposium (USENIX Security 19), pages 443–460. USENIX Association, 2019. URL: https://www.usenix.org /conference/usenixsecurity19/presentation/ xiao.
Ê ′ (x) = Ê(x) +
wk (x) E f (x), ∑ j w j (x)
which biases the reconstructed HDR map upward. The later tone-mapping stage may then compress the output as if the whole scene were brighter, producing darkened or extremedarkening frames. Saturation and missing-data effect. If the flashed exposure saturates, its weight may be reduced or set to zero by saturation-aware fusion. In that case, the corrupted bracket becomes missing data, with wk (x) ≈ 0. If the remaining exposures are underexposed, the fusion step has insufficient reliable information for the affected pixel region. This can reduce local detail, suppress target contrast, or force a dark fused value. Exposure-fusion failure. For exposure-fusion methods such as Mertens fusion, the per-pixel weight can be written as W j (x) = C j (x)wc S j (x)ws E j (x)we , where C j , S j , and E j measure contrast, saturation, and well-exposedness [36]. FLASH
[56] Yifan Xiao, Peter Veelaert, and Wilfried Philips. Deep hdr deghosting by motion-attention fusion network, 2022. URL: https://www.mdpi.com/1424-822 0/22/20/7853, doi:10.3390/s22207853. [57] Chen Yan, Zhijian Xu, Zhanyuan Yin, Xiaoyu Ji, and Wenyuan Xu. Rolling colors: Adversarial laser exploits against traffic light recognition. In 31st USENIX Security Symposium (USENIX Security 22), pages 1957– 1974. USENIX Association, 2022. URL: https: //www.usenix.org/conference/usenixsecuri ty22/presentation/yan. 17
Table 7: HDR-Lab Algorithm Comparison Matrix. The framework implements canonical algorithms at each stage of the pipeline to enable combinatorial evaluation of vulnerability to adversarial illumination. Pipeline Stage
Implemented rithms
Alignment
MTB, Homography
Merging
Tone Mapping
Algo- Function / Characteristics
Corrects for camera motion between exposure brackets. MTB is sensitive to global binary threshold shifts. Robertson, Debevec, Recovers response curves and fuses expoMertens sures. Mertens performs exposure fusion directly, bypassing radiance-map construction used by Robertson/Debevec. Mantiuk, Drago, Rein- Compresses dynamic range for display. hard, Local Global operators preserve saturation; local operators may introduce artifacts.
Figure 10: HDR-Lab graphical interface used to configure and execute modular HDR processing pipelines. The tool exposes swappable alignment, merging, and tone mapping stages and provides side-by-side visualization of source inputs, processed output, and per-run configuration details.
can either overexpose the flashed bracket or make the nonflashed brackets appear inconsistent with it. In both cases, the normalized weights become unstable, and the fused image can lose target-background contrast.
B
HDR-Lab Implementation Details Here, ℓLoS indicates whether the flashlight has an unobstructed line of sight to the camera, and δclear denotes the minimum clearance from deployment-specific obstacles or protected regions. The condition csrc = 1 requires the source to support the selected power and cadence under the fixed pulse-width configuration. The condition csafe = 1 represents applicable deployment-specific safety restrictions. Physical feasibility and attack effectiveness. Membership in C (ss) establishes only that a configuration satisfies the placement, hardware, and safety constraints. It does not require selective bracket contamination or a particular amount of output degradation. These properties characterize attack effectiveness rather than physical feasibility. Because the exact exposure windows and relative phase are unavailable to the attacker, effectiveness is measured from the resulting camera output in Section 6. Motion-dependent feasibility. When the camera or source moves, let a k and s k denote their values at frame k. A configuration remains feasible over an attack interval K only if a k ∈ C (ssk ), ∀k ∈ K .
The HDR-Lab framework allows for a combinatorial evaluation of pipeline configurations under identical illumination conditions. It implements multiple canonical algorithms at each stage of the pipeline to enable controlled experimentation across alignment, merging, and tone mapping strategies. Table 7 details the specific algorithms implemented within the framework. This modular design facilitates the evaluation of how specific standard techniques (such as shift-based alignment or global versus local tone mapping) react to adversarial illumination. Figure 10 illustrates the graphical interface used to configure these modular pipelines and visualize the impact of adversarial illumination on intermediate processing steps.
C
Parameter Constraints and Interactions
The main methodology defines the attacker-selected configuration a , operating condition s , component-wise parameter bounds, received flash irradiance, and bracket-level injected energy. This appendix specifies only the additional deployment constraints used to construct the feasible set. Feasible set. Let B ⊂ R5 denote the Cartesian product of the admissible intervals for the attacker-selected parameters defined in the main methodology. For an operating condition s , we define ℓLoS (aa, s ) = 1, δclear (aa, s ) ≥ ∆safe , C (ss) = a ∈ B . csrc (aa) = 1, csafe (aa, s ) = 1
Motion can change the source-to-camera geometry, received irradiance, line-of-sight availability, and clearance. Motion itself does not change the nominal cadence relationship unless the camera frame rate or the source-camera clock relationship also changes.
D
Metric Definitions
For each frame, we compute a Rec.709 luma proxy from decoded RGB values: Y ′ = 0.2126R′ + 0.7152G′ + 0.0722B′ , 18
where R′ , G′ , and B′ denote gamma-encoded color channels following the Rec.709 luma coefficients. We use Y ′ as an output-level brightness metric rather than as a calibrated photometric luminance measurement. Let Ω denote an image region, such as the full frame, gray-card region, target region, or background region. For frame i, the mean luma in region ′ ′ Ω under condition c is Ȳi,Ω,c = meanx∈Ω Yi,c (x) . For a test condition and its matched reference condition, the per-frame signed luma change is: ′ δYi,Ω (%) = 100 ×
′ ′ Ȳi,Ω,test − Ȳi,Ω,ref ′ Ȳi,Ω,ref
Table 8: Simulated optical power levels and representative commercial source classes. Representative Source Class Low-power strobe Everyday-carry flashlight Security flashlight High-output flashlight Searchlight-class source
1 N ′ ∑ ⊮ δYi,Ω ≤ −50% , N i=1
rext,Ω =
1 N ′ ∑ ⊮ δYi,Ω ≤ −80% . N i=1
This metric captures changes in target-background separability under the same decoded-output measurement model used for the luma-change metrics.
E
Simulation Environment Details
The simulation is conducted in Blender within a closed indoor cube of size 50 m × 50 m × 50 m. The cube origin is placed at z = 25 m so that the floor lies on the world z = 0 plane. The interior face normals are flipped to support internal ray tracing, and the ceiling face is removed to admit overhead ambient illumination. The interior surfaces use a checker texture with scale 10.0 and a noise texture with scale 60.0 and detail 15.0, multiplied into the base color to introduce fine spatial variation. All surfaces use a Principled BSDF with specular IOR level 0.0 and roughness 1.0, which suppresses mirrorlike reflections and keeps the measurements dominated by diffuse irradiance. Atmospheric effects are modeled with a world volume that applies volume scatter with density 0.01 and anisotropy −1.0. Ambient illumination is provided by a 50 m × 50 m area light positioned at a height of 50 m. We use two ambient source-power settings, Pasim = 500 kW for daytime and Pasim = 20 kW for nighttime. These are simulation source-power settings and are not treated as equivalent to the lux measurements used in the physical evaluation. The attack source is modeled as a circular emitter with diameter 0.1 m, and shadow visibility is disabled to avoid self-occlusion artifacts from the source geometry. The evaluated optical power levels are 2 W, 4 W, 8 W, 12 W and 20 W, corresponding to the representative commercial source classes summarized in Table 8.
Here, N is the number of matched frames used in the comparison. Thus, severe darkening is the percentage of frames with at least a 50% luma reduction relative to the matched reference, and extreme-darkening rate is the percentage of frames with at least an 80% luma reduction. We also compute target-background contrast-to-noise ratio using the same luma proxy. Let ΩT and ΩB denote the target and background regions. For frame i under condition c, the per-frame CNR is: ′ ′ Ȳi,Ω − Ȳi,Ω T ,c B ,c CNRi,c = q , σ2i,ΩT ,c + σ2i,ΩB ,c ′ where Ȳi,Ω,c is the mean Rec.709 luma proxy in region Ω, and 2 σi,Ω,c is the corresponding regional luma variance. For a test condition and its matched reference condition, the per-frame signed CNR change is
δCNRi (%) = 100 ×
2W 4W 8W 12 W 20 W
.
For parameter-sweep trend analysis, we also use the de′ = Ȳ ′ ′ coded luma-unit change ∆Yi,Ω i,Ω,test − Ȳi,Ω,ref . We report g′ Ω over matched the median signed luma-unit change ∆Y frames for the physical parameter sweeps. This luma-unit metric is used only to summarize sweep trends; the default cross-device and system-level failure metrics remain percentnormalized. For the default cross-device and system-level comparisons, we report the mean, median, and minimum signed luma f′ Ω , δY ′ change over frames δY ′ Ω , δY min,Ω , along with the severe-darkening and extreme-darkening rates: rsev,Ω =
Optical Power
The spatial sweep evaluates distances from 2 m to 20 m in 2 m increments and horizontal angles from −45◦ to 45◦ . The height sweep fixes the distance at 4.0 m and the angle at 0◦ , then evaluates source heights of 0.5 m, 1.5 m and 3.0 m. The power sweep uses the same 4.0 m distance, 0◦ angle, and 1.5 m source height while varying source power across the five evaluated levels.
CNRi,test − CNRi,ref . CNRi,ref
We report summary statistics of δCNRi over frames for comparisons where target and background regions are defined. 19
F
Component-Level Physical Setup
Distance sweep (d f c ). The distance sweep shows the strongest response at the closest tested range. On the iPhone 16 Pro, the median ∆Y ′ decreases from +60.60 units at 2 m to +28.85, +16.90, and +12.62 units at 4 m, 6 m and 8 m. The Raspberry Pi Camera Module 3 shows the same attenuation with smaller magnitude, decreasing from +2.82 units at 2 m to +0.74, +0.29, and +0.42 units at 4 m, 6 m and 8 m. This confirms that the received flash contribution weakens as d f c increases. Angle sweep (θ). The angle sweep peaks near the optical axis. On the iPhone 16 Pro, the median ∆Y ′ decreases from +60.60 units at 0◦ to +41.15, +17.53, and +14.66 units at 10◦ , 20◦ , and 30◦ . The Raspberry Pi Camera Module 3 decreases from +2.82 units at 0◦ to +1.92, +0.46, and +0.50 units. This matches the simulation results showing that off-axis placement reduces the injected light reaching the camera. Source-setting sweep (Pf ). The source-setting sweep is nonmonotonic, which is expected for black-box camera pipelines with automatic exposure and tone mapping. On the iPhone 16 Pro, the median ∆Y ′ is +39.19, +52.59, and +32.94 units at 25%, 50%, and 100% settings. On the Raspberry Pi Camera Module 3, the corresponding values are +2.18, +0.71, and +2.60 units. Thus, a higher Pf setting does not always produce a larger final luma shift because the camera pipeline can compensate differently across settings. Ambient-illuminance sweep (A). Increasing measured ambient illuminance suppresses the low-light response. On the iPhone 16 Pro, the median ∆Y ′ changes from +60.44 units at 0 lux to −6.12 units at 27.3 lux and −4.51 units at 239 lux. On the Raspberry Pi Camera Module 3, the response decreases from +5.98 units at 0 lux to +1.00 and −0.12 units at 27.3 lux and 239 lux. These results show that increasing A reduces the relative contribution of the injected flash. Frame-rate sensitivity ( ffps ). The frame-rate sweep shows that capture mode affects the response, but the effect is not confined to one frame rate. On the iPhone 16 Pro, the median ∆Y ′ is +67.39, +65.87, +45.43, and +52.61 units at 24 fps, 30 fps, 60 fps and 120 fps. On the Raspberry Pi Camera Module 3, the response is +12.03, +6.56, and +11.83 units at 30 fps, 60 fps and 120 fps. These results show that timing behavior and capture mode shape the FLASH response, while the attack effect remains measurable across the tested ffps settings.
Figure 11 shows the fixed component-level setup used across the physical camera experiments.
Flashlight
Checkerboard Example Target
Camera
Gray Reference Card
Figure 11: Component-level physical setup used for the FLASH evaluation. The setup includes the target board, checkerboard target, gray reference card, flash source, and camera under test. We conduct the component-level physical evaluation in a fixed indoor setup. The camera under test is mounted on a tripod at zc = 1.5 m, and the target board is placed 2 m from the camera. The flash source is mounted on a separate tripod at z f = 1.5 m and aimed toward the camera and target region. The scene includes a checkerboard target and a gray reference card. Ambient condition A and control-light illuminance are measured at the camera plane in lux, and floor markings maintain consistent camera, target, distance, and angle positions across devices. All physical cameras are treated as black-box imaging pipelines, and we record the final image or video output using the available or default capture mode. The default physical setup uses a manufacturer-rated Pf = 2 W commodity flashlight at d f c = 2 m to provide sufficient injected illumination. The simulation default instead uses Pf = 12 W at d f c = 4 m. This provides a controlled midstrength simulation condition without trivially saturating the closest setting with the strongest 20 W source.
G
H Legal Implications of Blinding Autonomous Vehicle (AV) and Surveillance Cameras
Physical Operational Robustness Details Deliberate optical interference with AV sensors, including high-intensity flashlights or lasers aimed at cameras or LiDAR, is likely covered by existing criminal statutes. Federal law treats willful acts that “damage, disable, or tamper with” a motor vehicle in a way that endangers human life as a felony punishable by up to twenty years’ imprisonment (18 U.S.C. § 33) [1, 51]. State laws add related prohibitions: Texas may
The saved physical-camera dataset contains 94 videos: 28 default videos, 28 distance and angle sweep videos, 12 sourcesetting sweep videos, 12 ambient-illuminance sweep videos, and 14 frame-rate sensitivity videos. Figure 12 reports the median signed decoded luma-unit response across the physical sweep settings. 20
iPhone 16 Pro
iPhone 16 Pro median signed Y (units)
70
Distance
60
Raspberry Pi Camera Module 3
Angle
Power
Ambient
FPS
Raspberry Pi Camera Module 3 median signed Y (units)
Physical parameter sweeps, median signed luma response
12.5 10.0 7.5
40
5.0 20 2.5 0
0.0
60 12 0
30
24
lx 9l x
x
.3
0l
23
Sweep setting
27
%
0%
50
10
% 25
0° 10 ° 20 ° 30 °
8m
6m
4m
2m
10
Figure 12: Complete physical parameter sweep results for the iPhone 16 Pro and Raspberry Pi Camera Module 3. Values report median signed decoded luma-unit change for matched clean-versus-FLASH trials. Separate y-axis scales preserve the within-device trends. classify intentional sensor blinding as criminal mischief [5,6], while California forbids directing a laser beam into a moving vehicle with intent to harass or annoy [2, 3]. Related civil claims may include negligence, trespass to chattels, vandalism, repair costs, medical costs, and punitive damages. Similar rules apply to interference with security or traffic cameras, including vandalism, criminal trespass, and public nuisance theories [4].
I
Publication Disclaimer
Reference herein to any specific commercial company, product, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or the Department of the Army (DoA). The opinions of the authors expressed herein do not necessarily state or reflect those of the United States Government or the DoA and shall not be used for advertising or product endorsement purposes.
21