Conceptio › Archive › arXiv CS
arXiv CSopen access

Learning Responsibility-Attributed Adversarial Scenarios for Testing Autonomous Vehicles

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

arXiv:2605.13751v1 [cs.RO] 13 May 2026

Learning Responsibility-Attributed Adversarial Scenarios for Testing Autonomous Vehicles

Yizhuo Xiao1 , Haotian Yan2 , Ying Wang3 , Zhongpan Zhu2,4 , Yuxin Zhang5 , Xintao Yan6 , Mustafa Suphi Erden1 , Cheng Wang1,∗ 1 School of Engineering and Physical Sciences, Heriot-Watt University, Edinburgh, U.K. 2 State Key Laboratory of Autonomous Intelligent Unmanned Systems, Tongji University, Shanghai, China 3 College of Computer Science and Technology, Jilin University, Changchun, China 4 University of Shanghai for Science and Technology, Shanghai, China 5 National Key Laboratory of Automotive Chassis Integration and Bionics, Jilin University, Changchun, China 6 Department of Civil Engineering, The University of Hongkong, Hongkong, China ∗ Corresponding author: [email protected]

Abstract Establishing trustworthy safety assurance for autonomous driving systems (ADSs) requires evidence that failures arise from avoidable system deficiencies rather than unavoidable traffic conflicts. Current adversarial simulation methods can efficiently expose collisions, but generally lack mechanisms to distinguish these fundamentally different failure modes. Here we present CARS (Context-Aware, Responsibility-attributed Scenario generation), a framework that integrates responsibility attribution directly into adversarial scenario generation. CARS combines context-aware adversary selection with a generative adversarial policy optimized in closed-loop simulation to construct collision scenarios that are both physically feasible and diagnostically attributable. Across benchmark datasets spanning heterogeneous national traffic environments, CARS consistently discovers feasible collision scenarios with high attribution rates under multiple regulation-prescribed careful and competent driver models. By coupling adversarial generation with normative responsibility assessment, CARS moves simulation testing beyond collision discovery toward the construction of interpretable, regulation-aligned safety evidence for scalable ADS validation.

Introduction Autonomous driving systems (ADSs) are moving toward broader public-road testing and deployment. Regulators, developers, and the public therefore need credible evidence that these systems can handle rare, high-risk traffic interactions 1 . Road testing alone cannot provide that evidence at practical scale. Safety-critical events are rare in naturalistic traffic, and statistical confidence from public-road exposure would require extremely large driving distances 2–4 . Closed-loop simulation and scenariobased testing have therefore become central to ADS safety validation 1,5–7 . These methods are also increasingly connected to regulatory practice and to evaluation-efficient use of test evidence 8–11 . Recent work has made generated safety-critical scenarios more frequent, severe, diverse, controllable, or reactive to the ADS 12–19 . These advances are important. However, most adversarial generators still optimize or report collision occurrence as the primary outcome. A collision count alone cannot distinguish ADS failures from unrealistic or infeasible adversarial motion, or from encounters that no human driver could reasonably avoid. For instance, a crash caused by an unavoidable sudden intrusion Preprint.

cannot meaningfully be attributed to deficient ADS decision-making. We argue that inducing the generation of scenarios with elevated collision rates is a necessary condition for adversarial scenario generation, but the core criterion for evaluating test validity is whether a collision can be attributed to the system under test. Existing adversarial scenario generation methods treat collision itself as the optimization objective, conflating collisions that the system should have avoided but did not with collisions that no reasonable driving behavior could have prevented, which fundamentally undermining the credibility of test conclusions.

Figure 1: Conceptual challenges in responsibility-attributed adversarial scenario generation. a, Adversary selection instability. Static assignment or one-shot geometric heuristics fix the attacker (red) at simulation start; as traffic evolves the original adversary drifts out of the conflict region while other agents (gray) that develop a credible threat are never reconsidered, leaving the ADS (blue) unopposed. b, Trajectory-capacity limitation. A policy that predicts a single-Gaussian distribution can still generate severe collisions, but the resulting severity may not survive attribution and feasibility checks. c, Responsibility ambiguity. Without a responsibility-attribution standard, scenarios in which the ADS could have braked in time and scenarios that any reasonable response would also have suffered are reported under the same failure-rate metric, conflating ADS shortcomings with the encounter’s inherent unavoidability. Resolving responsibility attribution requires a behavioral reference standard that is independent of the system under test, used to delineate the encounters an ADS should reasonably be expected to avoid. Without such a reference, the same collision count mixes failures of the ADS with collisions that no reasonable driver could have prevented, and the resulting test evidence cannot be interpreted as evidence of ADS deficiency. The international regulatory framework UN ECE R157 8 already provides the conceptual basis for such a standard through the careful and competent driver model (CCDM), a normative reference that defines the response a careful and competent driver (CCD) would produce in the same encounter. A collision is attributable to the ADS under test if the CCDM would have avoided it under the same scenario; otherwise the encounter lies outside the scope of meaningful ADS attribution. The CCD concept has been instantiated as several computational models with different structural assumptions, including the Fuzzy Safety Model (FSM) 20,21 prescribed by UN ECE R157, the Japanese CCD model (CC-JP) 22 . A related formalism, Responsibility-Sensitive Safety (RSS) 23 instead defines safety envelopes. These models share the same attribution rule and differ mainly in how they bound the reference driver’s braking response and reaction delay. 2

However, existing adversarial scenario generators are not organized around this attribution rule. Their optimization remains at the level of inducing collisions. This leaves adversarial scenario generation with a structural gap, which spans adversary selection, trajectory generation, and collision evaluation, as summarized in Figure 1. A responsibility-attributed scenario must begin with an adversary that can create an avoidable conflict from a plausible traffic position. Its generated motion must remain feasible while exerting collision pressure. Most importantly, the generation target should be an ADS-attributable collision: one that occurs for the ADS under test but would be avoided by a CCD reference. These requirements link adversary selection and trajectory generation to the CCD attribution reference.

Figure 2: CARS responsibility-attributed adversarial scenario generation framework. a, Contextaware adversary selection. At each simulation step, CARS scores surrounding agents by their kinematic relationship to the ADS and stabilizes the selected agent over time, allowing the threat assignment to track the evolving conflict. b, Adversarial trajectory generation. The selected agent is initialized from a Gaussian-mixture diffusion prior. Closed-loop reinforcement learning fine-tunes the agent to approach the ADS and create safety-critical interactions. c, Attribution-aware objective. CARS targets ADS-attributable collisions under a CCD reference. A collision is attributable when it occurs for the ADS under test but would have been avoided by the CCD in the same encounter. To address this gap, we present CARS (Context-Aware, Responsibility-attributed Scenario generation), a framework for generating ADS-attributable adversarial scenarios under a CCDM reference (Figure 2). CARS is built around three design properties. First, context-aware selection chooses the adversarial counterpart from agents that can create a genuine avoidable conflict from their current traffic position. Second, multi-component generation uses a Gaussian-mixture diffusion policy to represent multiple plausible action-sequence hypotheses under the same traffic context 18,24,25 , which is then fine-tuned by closed-loop reinforcement learning toward safety-critical interactions 26 . Third, attribution-aware generation optimizes the adversarial policy for collisions whose responsibility can be assigned to the ADS under the CCD standard. We validate the effectiveness and generalization of CARS across multiple experiments. With regard to effectiveness, the majority of generated collisions are correctly attributed to the ADS under FSM, and the attribution remains consistent under the CC-JP and RSS cross-checks, indicating that the result is not specific to a single CCDM specification. With regard to generalization, the same generation policy transfers without retraining across three datasets spanning heterogeneous national traffic environments: nuScenes urban driving 27 , AD4CHE Chinese highway traffic 28 , and RounD German roundabout interaction 29 . 3

In summary, these results show that adversarial simulation can construct responsibility-attributed test evidence, not only produce collisions. This work moves adversarial scenario generation from collision discovery to responsibility-attributed scenario generation, providing a methodology aligned with the regulatory governance framework for scalable ADS safety testing.

Results Evaluation design We evaluate CARS by asking whether generated collisions are attributable to the ADS, physically feasible, severity-diverse, and transferable beyond the training setting. The primary evaluation is conducted on the nuScenes urban-driving dataset 27 . Cross-domain transfer is then tested on highway traffic from AD4CHE 28 and roundabout interactions from RounD 29 , without retraining the adversary policy (dataset details in Supplementary Section S14; map in Supplementary Fig. S3). The primary nuScenes evaluation uses a Gaussian-mixture diffusion planner as the ADS under test; ADS-planner robustness experiments replace it with alternative planners. Generated collisions are evaluated under three reference models with different modeling assumptions. FSM serves as the primary attribution reference because it operationalizes the UN ECE R157 CCD standard with a fuzzy braking response that produces graduated decelerations rather than binary threshold decisions. This continuous response also provides a natural severity scale for distinguishing Easy, Medium, and Hard encounters 8 . CC-JP and RSS are retained as cross-checks under different structural assumptions. Table 1 reports responsibility validity under the three reference models, severity diversity (Hcrit ), positive braking-deficit exposure (BD+ %, the fraction of scenarios with BD > 0), and trajectory infeasibility (IP%); definitions and per-axis kinematic checks are given in Methods and Supplementary Tables S1, S5, and S6. Adversary selection Context-aware selection keeps the active adversary aligned with the evolving threat rather than fixing the attacker at the start of a scenario. We evaluate the selector on validation target–adversary (tgt–adv) pairs and report the full classification and ranking statistics in Supplementary Table S4. Because several surrounding agents can be plausible threats in the same scene, ranking is the relevant test: the labeled adv appears among the three highest-ranked agents in 97.8% of validation tgt–adv pairs. Candidate adversaries are rescored at every simulation step. A temporal confirmation gate promotes a new adv only after its score remains highest for several consecutive frames, preventing momentary score spikes from switching the policy away before the current adv has shaped the interaction. The selector is not intended to infer a unique adversarial intent from a single frame. Its role is to keep the adversarial policy attached to the surrounding agents that currently forms the most relevant conflict with the ADS. Figure 3 shows how the selected adversary changes as the conflict develops. In the nuScenes example, the initial adv is replaced only after a closer agent sustains a higher adversarial score over the confirmation window; this newly selected adv then closes on the tgt. The same selection rule identifies a lane-change threat in AD4CHE and a roundabout-entry yielding conflict in RounD, indicating that the mechanism tracks the evolving conflict rather than a fixed scene template. Most rollouts retain the initially selected adv, but 12–16% require a switch. These switches occur when the originally assigned adv no longer dominates the interaction, showing why adversary selection must remain active during rollout. Responsibility attribution on nuScenes In the primary nuScenes evaluation, CARS retains 88.7% attribution under the primary FSM reference (Table 1). The auxiliary checks give 79.7% attribution under CC-JP and 97.1% under RSS, indicating that the attribution result is not specific to a single reference specification. These auxiliary checks are interpreted as sensitivity tests rather than as a replacement for the primary FSM criterion. As a stricter robustness subset, 73.8% of the collisions are attributable under all three reference models. Only 2.2% are unpreventable under all three, so 97.8% remain preventable by at least one reference model. This distinction matters for validation: a generator that simply forces the adversary into unrecoverable 4

Figure 3: Context-aware adv re-selection across datasets. a, nuScenes. b, AD4CHE. c, RounD. Each row shows three BEV snapshots from one rollout scenario at t=0, the adv switch frame, and the collision frame, with a heatmap strip beneath giving the per-step adversarial probability Padv (t) for exemplary candidates A1 (upper) and A2 (lower); darker red indicates higher probability. The tgt is shown in blue, the currently active adv in red, and other agents in gray. geometry would raise unpreventable rates under all references, whereas CARS mostly produces encounters that remain avoidable under the primary reference or at least one auxiliary check. Figure 4a–c illustrates representative disagreements among the reference models. Panels a–c show cases in which FSM, CC-JP, and RSS, respectively, would still collide while the other references recover. These cases are not treated as a replacement for the primary FSM criterion; they show how auxiliary references expose sensitivity to model-specific avoidability assumptions. If CARS were mainly producing geometrically impossible collisions, all three references would fail together. Instead, the all-reference-unpreventable set remains small. This pattern supports the interpretation that CARS produces challenging but still evaluable encounters, rather than collisions whose attribution is already lost by construction. Severity diversity and kinematic feasibility CARS-generated collisions span a broad range of FSM braking demand while remaining within the feasibility bounds used for IP. Severity is summarized by the FSM Easy, Medium, and Hard criticality tiers, which describe the braking demand imposed on the FSM reference: Easy cases remain within a comfort-level response, whereas Hard cases approach the emergency-braking regime. On nuScenes, generated scenarios populate all three tiers, yielding Hcrit =0.798 (Table 1; per-tier and per-method breakdowns in Supplementary Table S1 and Supplementary Table S2). Figure 4d–f characterizes the collision kinematics behind this spread. The BD distributions separate the FSM criticality tiers, confirming that the tier labels reflect differences in braking demand. The adv 5

speed and longitudinal closing-speed panels show that collisions occur during sustained approaches, rather than from stationary or marginal contacts. The Easy and Medium tiers remain populated, indicating that the generated set does not collapse into a single emergency-braking regime. Additional nuScenes collision geometries are provided in Supplementary Section S15. The same scenarios remain physically plausible under the feasibility checks used for IP. CARS has IP% equal to 0.04% in Table 1, and the percentile-level acceleration, jerk, and lateral-acceleration checks remain below the feasibility bounds (Supplementary Table S6). Together, the severity and feasibility results show that CARS preserves a spread of attributable interaction regimes without relying on kinematic artifacts. a. FSM fails

b. CC-JP fails

Adversary

d. Physical-margin distribution

Target

c. RSS fails

Other agents

e. Adv speed at collision

FSM

CC-JP

RSS

f. Approach urgency

Figure 4: Responsibility attribution and collision kinematics on CARS-generated nuScenes scenarios. a–c, Representative attribution disagreements in which FSM fails (a), CC-JP fails (b), and RSS fails (c). adv trajectory in red; tgt trajectory in blue; the tgt’s stop position under each reference model is shown as a front-edge line (purple, FSM; orange, CC-JP; green, RSS), as keyed in the legend. d, Max braking deficit BD, box-and-whisker per FSM criticality tier (Easy / Medium / Hard); box spans the interquartile range, whiskers extend to the 5th–95th percentiles, points are outliers, and median is annotated. BD > 0 indicates that required stopping distance exceeded the available gap. e, Adv speed at collision, raincloud plot per FSM criticality tier (median annotated). f, Longitudinal closing speed at collision, kernel-density estimate per criticality tier (median annotated). Comparison with adversarial generators Under the same responsibility-attribution pipeline, collision-oriented generators do not automatically produce responsibility-attributable scenarios. We compare CARS with STRIVE 12 , SafeSim 13 , and a rule-based Bezier-CAT interception baseline inspired by CAT 15 , evaluating each method’s generated collisions with the same FSM, CC-JP, and RSS attribution checks. An internal CARS (K=1 adv) ablation isolates the effect of replacing the adversarial policy’s Gaussian-mixture output head with a one-component Gaussian output head. Figure 5 visualizes the comparison, and Table 1 reports the full metric matrix. Under the same FSM criterion, CARS retains 88.7% attribution, approximately twice the strongest external baseline, while also achieving the highest severity diversity and the lowest infeasibility rate among the compared methods (Figure 5, Table 1). The main metrics are stable across seven 6

rollout seeds (Supplementary Fig. S2). The baselines occupy different failure regimes. STRIVE finds impacts that rarely remain attributable under any reference model. SafeSim improves attribution relative to STRIVE, but still shows lower severity diversity and higher infeasibility than CARS. The rule-based Bezier-CAT interception baseline has the highest BD+ and IP rates, showing that geometric collision pressure can come at the cost of feasibility. These failure modes differ, but they share the same limitation: collision occurrence alone does not imply usable validation evidence. The internal CARS (K=1 adv) ablation shows why raw severity diversity is not sufficient. Replacing the adversarial policy’s Gaussian-mixture output head with a one-component Gaussian output head leaves Hcrit almost unchanged, but nearly halves FSM attribution and sharply increases infeasibility. After requiring zero kinematic violation, only 25.0% of FSM-preventable K=1 scenarios remain, whereas CARS retains 97.2% with essentially unchanged severity diversity (Supplementary Table S3). Thus, the added trajectory capacity of the Gaussian-mixture adversary is expressed not as a broader severity histogram alone, but as severity diversity that survives attribution and feasibility checks. a. Severity diversity

b. Three-CCDM validity

c. Validity-feasibility trade-off

d. Severity tier coverage

Figure 5: Baseline comparison on nuScenes. a, Severity-diversity entropy Hcrit per method (raw point estimate on the FSM-preventable subset; same values as Table 1). b, Responsibility validity under FSM, CC-JP, and RSS per method, grouped bar. c, Validity-feasibility trade-off: FSM validity (x) versus infeasibility percentage IP (y); bubble area scales with the square root of the FSM-preventable scenario count. Lower-right is better. d, Per-method criticality tier coverage on the FSM-preventable subset, horizontal stacked bar. Generalization across planners and domains The learned adversarial policy remains effective when the ADS-under-test planner is changed without retraining the adversarial policy. On nuScenes, replacing the Gaussian-mixture diffusion planner used during training with a one-component diffusion planner leaves FSM attribution at 87.8%, and replacing it with the architecturally distinct CTG planner 30 retains 86.0% FSM attribution (Table 1). Both planner-robustness tests preserve low infeasibility, with IP% at 0.04% and 0.02%, respectively. These results indicate that CARS does not depend on the particular ADS planner architecture used during training. We next deploy the same adversarial policy on AD4CHE and RounD without retraining. AD4CHE tests highway interactions and RounD tests roundabout interactions, giving two traffic geometries outside the urban nuScenes training setting (Figure 6). FSM attribution remains 76.4% on AD4CHE 7

and 57.5% on RounD, while severity diversity remains close to the nuScenes result (Table 1). Physical feasibility transfers cleanly to AD4CHE (IP = 0.00%) and remains low on RounD (IP = 2.32%), where curved approaches and laterally dominated interactions differ most from the training distribution. Together, these results show that the adversarial policy retains measurable attribution and low infeasibility across ADS planners and traffic geometries, while the lower RounD attribution rate reflects the added difficulty of roundabout interactions. Figure 6 provides the geometric and metric context for the cross-domain tests. Panel a overlays generated conflicts in a tgt-centered relative-position frame, giving a qualitative domain map. Attribution is quantified in panel b and Table 1: all three reference models avoid a smaller fraction of RounD collisions than nuScenes or AD4CHE collisions. This drop is not specific to the FSM reference alone. Panel d shows that RounD contains a larger proportion of Hard FSM-preventable cases, indicating that roundabout interactions shift the FSM-based attribution profile toward higher braking demand. This pattern is consistent with the curved and laterally dominated geometry of RounD, where a longitudinal braking reference has fewer ways to resolve conflicts than in more aligned approach settings. Together, Figure 6 and Table 1 show that cross-domain deployment preserves measurable attribution and low infeasibility, but the roundabout domain shifts the generated evidence toward harder and less easily attributable encounters. Table 1 | Unified evaluation of CARS across adversarial-method comparisons, ADS-planner robustness tests, and cross-dataset generalization. In the ADS-planner rows, the adversarial policy is fixed and only the ADS-under-test planner is replaced. Method / Setting

FSM % ↑

CC-JP % ↑

RSS % ↑

Hcrit ↑

BD+ % ↓

IP % ↓

Adversarial methods on nuScenes STRIVE

7.3

5.8

6.1

0.528

53.8

36.39

SafeSim

44.8

44.8

44.8

0.628

48.3

12.94

Bezier-CAT

21.1

36.4

15.0

0.260

84.1

73.08

CARS (K=1 adv )

45.2

35.5

53.2

0.797

66.1

27.40

CARS

88.7

79.7

97.1

0.798

22.5

0.04

ADS-planner robustness (fixed CARS adv) One-component diffusion planner

87.8

80.0

96.7

0.834

21.7

0.04

CTG planner 30

86.0

80.0

92.4

0.745

29.2

0.02

CARS on AD4CHE

76.4

63.8

80.9

0.751

36.6

0.00

CARS on RounD

57.5

52.0

70.7

0.759

50.9

2.32

Cross-dataset generalization

Discussion CARS shows that adversarial simulation can generate collisions whose responsibility is attributable, not merely collisions whose occurrence can be counted. Collision counts alone can mix ADS failures with infeasible adversarial motion or encounters that no reasonable reference response would avoid. CARS addresses this by linking three generation-side choices: online adv selection, multi-component trajectory generation, and closed-loop reward shaping against a CCD reference. The resulting scenarios remain largely attributable under the primary FSM reference, retain severity diversity after feasibility checks, and remain measurable when the ADS planner or traffic domain changes. The main implication is that adversarial scenario generation should be evaluated as evidence construction. For ADS validation, a generated collision is useful only if it supports a claim about the ADS under test. That claim requires an independent reference response. If the reference also collides, the scenario is a weak failure case for the ADS. If the reference avoids the collision, the scenario becomes interpretable as an attributable failure. This changes the objective from inducing impacts to constructing collisions that survive attribution, severity, and feasibility checks. This framing also clarifies the role of reference models. UN R157 defines ADS performance relative to a CCD, and FSM provides one operational longitudinal form of that standard. CC-JP and RSS expose how the attribution decision changes under different reference assumptions. Their disagreements are 8

a. Approach geometry across datasets

b. CCDM validity (%)

c. Kinematic urgency at impact

d. Severity tier shift

Figure 6: Cross-dataset generalization across nuScenes, AD4CHE, and RounD. a, Approach geometry overlaid across all three datasets, with shape encoding dataset (⃝ nuScenes; □ AD4CHE; △ RounD), color encoding FSM criticality tier (Hard/Medium/Easy), and 1D kernel-density marginals on top (s) and right (l). b, Responsibility validity under FSM, CC-JP, and RSS across the three datasets, color encoding validity %. c, Adv speed at collision, ridge plot per dataset. d, Severity tier shift across datasets, horizontal stacked bar on the FSM-preventable subset.

not a nuisance to hide. They identify boundary cases where responsibility depends on the chosen reference model. Reporting this structure is more informative than collapsing all generated collisions into one failure rate. The present study has several limits. The reference models are longitudinal, so the responsibility estimate is strongest for longitudinal conflicts. Roundabout and laterally dominated interactions can shift the FSM-based criticality profile because a braking-only reference has fewer ways to resolve the encounter. CARS also selects one active adv at a time, leaving coordinated multi-agent attacks outside the present evaluation. All evaluated agents are vehicles; vulnerable road users would require class-specific dynamics and responsibility assumptions. Future work should extend responsibility-attributed generation beyond these boundaries. Multiagent adversarial policies could generate coordinated threats. Reference models for lateral, crossing, and multi-agent conflicts would broaden the set of attributable scenarios. Vehicle–pedestrian and vehicle–cyclist interactions will require different feasibility envelopes and responsibility rules. As ADS regulation moves beyond longitudinal Automated Lane Keeping Systems (ALKS) settings, adversarial generators should be coupled to explicit reference standards so that generated scenarios remain useful as validation evidence. 9

Methods CARS is a closed-loop adversarial scenario generation framework with online adv selection, diffusionbased adversarial generation, and CCD-based attribution verification. First, a context-aware selector identifies the surrounding agent most likely to form a responsibility-relevant conflict with the ADS. Second, CARS trains a Gaussian-mixture diffusion model to generate adv action sequences conditioned on traffic context. The model is then fine-tuned with reinforcement learning in closed-loop simulation to create sustained approach conflicts rather than isolated contacts. Finally, attribution is verified with a CCD reference by holding the generated adv trajectory fixed and replacing the tgt response. This section describes the main components; detailed adversary-selection features, diffusion losses, reinforcement learning fine-tuning, and the identifiability argument are provided in the Supplementary Information. Scenario formulation We consider a multi-agent traffic scene with agents A = {1, . . . , N }. Let itgt denote the agent controlled by the ADS under test; we refer to this agent as tgt. Each agent i has state sit = [xit , yti , vti , θti ], containing position, speed, and heading, and action ait = [v̇ti , θ̇ti ], containing longitudinal acceleration and yaw rate. State transitions follow unicycle dynamics, sit+1 = f (sit , ait ).

(1)

At each simulation step, agents are partitioned into the tgt, the adv, and background agents Abg (t). The policies for tgt and Abg (t) remain fixed during adversarial generation; only the policy controlling the active adv is optimized. Let C(ξ) ∈ {0, 1} indicate whether rollout ξ contains a collision. For a generated adv trajectory Tadv , let ξADS (Tadv ) denote the original rollout with the ADS under test controlling tgt, and let ξCCD (Tadv ) denote the counterfactual rollout in which the tgt response is replaced by a CCD reference while Tadv is fixed. The collision is classified as ADS-attributable when Iattr (Tadv ) = I[C(ξADS (Tadv )) = 1 ∧ C(ξCCD (Tadv )) = 0] .

(2)

Context-aware adversary selection The active adv is selected online rather than fixed at simulation start. For each surrounding agent i, CARS computes a 13-dimensional feature vector ϕi,t describing its relative geometry and kinematics with respect to tgt, including distance, relative speed, heading alignment, longitudinal and lateral offsets, closing rates, bumper-to-bumper gap, time-to-collision, bearing angle, and short-horizon approach trends (definitions in Supplementary Section S2). A histogram-based gradient boosting classifier estimates the adversarial likelihood Padv (i | ϕi,t ), and the highest-ranked candidate at step t is ît = arg max Padv (i | ϕi,t ). (3) i∈A\{itgt }

Candidate scores are recomputed at every control step so that the active adv can follow the evolving threat. To avoid oscillation caused by transient geometric changes, a new candidate is promoted only after Kconf consecutive steps of agreement:  î , if ît−j = ît ∀j ∈ {0, . . . , Kconf − 1}, i∗t+1 = ∗t (4) it , otherwise, where i∗t is the active adv index after temporal confirmation, with i∗0 = î0 at rollout start. Gaussian-mixture adversarial generation Diffusion models (DMs) are used across ADS stress testing, from controllable perception attacks to traffic-scenario generation 14,31 . CARS first trains a DM to generate adv action sequences conditioned on map, tgt history, and neighbor history 14,32–34 . Given context c, the generator represents the future action sequence τa = {at+1 , . . . , at+T }; the forward diffusion process is given in Supplementary Section S3. Instead of predicting a single Gaussian denoising distribution, CARS uses a Gaussianmixture diffusion model (GMDM) to represent a richer conditional action-sequence distribution, 10

following recent mixture-parameterized diffusion and planning models for trajectory generation 24,25 . For noisy action sequence τar at diffusion step r, the denoising model predicts "K # T Y X 0 r 0 pθ (τa | τa , c, r) = πh,k N (τa,h ; µh,k , Σh,k ) , (5) h=1

k=1

where h indexes the planning horizon and πh,k , µh,k , and Σh,k are the mixture weight, mean, and covariance of Gaussian component k. Training minimizes a negative log-likelihood objective with mixture-usage and component-separation regularizers, Ltotal = LNLL − λH Hπ + λrepel Lrepel ,

(6)

with detailed definitions in Supplementary Section S4. The pretrained GMDM is then fine-tuned by reinforcement learning in closed-loop simulation, shifting its output distribution toward action sequences that create collision pressure. This refinement is implemented with proximal policy optimization (PPO) 35,36 . The reward combines longitudinal progress toward tgt, lateral proximity, lateral closing, a terminal collision bonus, and a time penalty: Rtotal = wprog Rprog + wℓ Rℓ + wℓc Rℓc + Rhit − λtime .

(7)

These dense terms encourage adv to enter and remain within the tgt-relative conflict geometry before impact, so that generated collisions arise from sustained approach dynamics. The reinforcementlearning objective and reward components are given in Supplementary Section S5. Responsibility attribution Attribution is verified by holding the generated adv trajectory fixed and replacing the tgt response with an independent reference model. We use FSM as the primary attribution reference for two reasons. First, it gives an operational form of the UN R157 CCD standard for longitudinal ALKS conflicts. Second, its fuzzy surrogate safety measures produce graduated braking responses, which are more consistent with graded driver braking behavior in ADS testing than a binary safe/unsafe switch 8,20,21 . CC-JP and RSS are used as auxiliary checks with different structural assumptions: CC-JP provides an alternative CCD reference from the same regulatory context, while RSS provides a formal safety envelope. Reporting all three references tests whether attribution results depend on a single reference specification. For FSM attribution, tgt follows its original recorded path while its speed along that path is controlled by the FSM braking response (Supplementary Section S7). The generated adv trajectory remains fixed. FSM combines a Proactive Fuzzy Surrogate Safety Metric (PFS), which assesses whether tgt can stop if adv brakes hard, and a Critical Fuzzy Surrogate Safety Metric (CFS), which assesses imminent collision risk under current accelerations. The saturated membership functions and safe/unsafe distance definitions for PFS and CFS are given in Supplementary Section S6. FSM then uses the Sugeno reaction rule, ( CFS (btgt,max − btgt,comf ) + btgt,comf , CFS > 0, btgt = (8) PFS btgt,comf , CFS = 0, where btgt is the commanded positive deceleration magnitude, and btgt,comf and btgt,max denote the comfort and maximum deceleration limits. The command is then followed by the regulatory reaction delay and jerk-limited ramp. The same PFS/CFS values define the FSM criticality tiers used in the Results: Easy (PFS≤ 0.85), Medium (PFS> 0.85 and CFS< 0.9), and Hard (CFS≥ 0.9). A generated collision is attributed to the ADS under a reference model when the reference-controlled tgt avoids the collision in the counterfactual rollout. FSM attribution should therefore be read as a longitudinal attribution claim. It assumes that the generated adv trajectory respects the relevant kinematic envelope, that the path-constrained counterfactual rollout is valid for the encounter, and that the conflict belongs to the longitudinal class addressed by UN R157. Under these conditions, an FSM-attributed collision provides evidence that the ADS failed in an encounter that a CCD-style longitudinal reference could avoid; the formal lower-bound statement is provided in Supplementary Section S1. 11

Implementation details The GMDM backbone uses a ResNet-18 map encoder and a temporal convolutional denoising network conditioned on map context, tgt history, and neighbor history. The Gaussian-mixture head uses K=8 components, selected by the mixture-component sweep in Supplementary Fig. S1, and the planning horizon is T =20 steps at 10 Hz. The adv selector is implemented with a histogram-based gradient boosting classifier and a five-step temporal confirmation gate. Additional adv-selection features, diffusion losses, and reinforcement learning details are provided in the Supplementary Information. Trajectory feasibility is measured from the adv trajectory, with details provided in the Supplementary Section S13. Let a denote longitudinal acceleration and j = da/dt denote longitudinal jerk. IP% is computed per scenario as the fraction of time steps that exceed at least one physical feasibility bound: |a| > 7 m/s2 8 , |j| > 12.65 m/s3 8 , or |alat | > 3.0 m/s2 37 . The reported IP% is averaged across scenarios, enabling cross-method comparison when rollout lengths differ. Full percentile-level kinematic checks are reported in Supplementary Table S6.

Data Availability The nuScenes dataset that we used to train CARS is publicly available at https://www.nuscenes. org. The RounD and AD4CHE dataset that we used for cross-dataset generalization experiment is publicly available at https://levelxdata.com/round-dataset/ and can be requested at https://lab.autozyx.com/ad4che.html.

Code Availability The project website and the source code of the CARS framework is publicly available at https: //robosafe-lab.github.io/CARS.

Acknowledgements This work was funded by UK Research and Innovation (UKRI) under the UK government’s Horizon Europe funding guarantee [grant number EP/Z533464/1].

Author Contributions Y.X. developed the methodology, implemented the framework, conducted experiments, and wrote the manuscript. H.Y. contributed to the visualization and experimental design. Y.W. and Z.Z. oversaw the research project and guided the team. Y.Z. provided data and cleaned the data. X.Y and M.S.E. contributed to writing – review & editing. C.W. focused on conceptualization, supervision, funding, and editing. All authors reviewed and approved the final version.

Competing Interests The authors declare no competing interests.

References 1. Qian, C., Xu, J., Xing, X. & Guo, F. Test case sampling optimization for safety validation of automated driving systems. Nat. Commun. 17, 3114 (2026). 2. Kalra, N. & Paddock, S. M. Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability? Transp. Res. Part A Policy Pract. 94, 182–193 (2016). 3. Feng, S. et al. Dense reinforcement learning for safety validation of autonomous vehicles. Nature 615, 620–627 (2023). 12

4. Liu, H. X. & Feng, S. Curse of rarity for autonomous vehicles. Nat. Commun. 15, 4808 (2024). 5. Feng, S., Yan, X., Sun, H., Feng, Y. & Liu, H. X. Intelligent driving intelligence test for autonomous vehicles with naturalistic and adversarial environment. Nat. Commun. 12, 748 (2021). 6. Yan, X. et al. Learning naturalistic driving environment with statistical realism. Nat. Commun. 14, 2037 (2023). 7. Ding, W. et al. A survey on safety-critical driving scenario generation—a methodological perspective. IEEE Trans. Intell. Transp. Syst. 24, 6971–6988 (2023). 8. United Nations Economic Commission for Europe (UNECE). UN regulation no. 157 — uniform provisions concerning the approval of vehicles with regard to Automated Lane Keeping Systems (ALKS). Tech. Rep., UNECE World Forum for Harmonization of Vehicle Regulations (WP.29) (2021). URL https://unece.org/transport/documents/2021/03/standards/ un-regulation-no-157-automated-lane-keeping-systems-alks. 9. Tang, L. et al. Scenario-based accelerated testing for SOTIF in autonomous driving: a review. IEEE Internet Things J. 12, 1453–1470 (2024). 10. Song, Q., Bensoussan, A. & Mousavi, M. R. Synthetic vs. real: an analysis of critical scenarios for autonomous vehicle testing. Autom. Softw. Eng. 32, 37 (2025). 11. Wu, X. et al. Make full use of testing information: An integrated accelerated testing and evaluation method for autonomous driving systems. Accid. Anal. Prev. 224, 108280 (2026). 12. Rempe, D., Philion, J., Guibas, L. J., Fidler, S. & Litany, O. Generating useful accident-prone driving scenarios via a learned traffic prior. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 17305–17315 (2022). 13. Chang, W.-J., Pittaluga, F., Tomizuka, M., Zhan, W. & Chandraker, M. Safe-Sim: Safety-critical closed-loop traffic simulation with diffusion-controllable adversaries. In Proc. Eur. Conf. Comput. Vis. (ECCV), 242–258 (Springer, 2024). 14. Xu, C., Petiushko, A., Zhao, D. & Li, B. DiffScene: Diffusion-based safety-critical scenario generation for autonomous vehicles. In Proc. AAAI Conf. Artif. Intell., vol. 39, 8797–8805 (2025). 15. Zhang, L., Peng, Z., Li, Q. & Zhou, B. CAT: Closed-loop adversarial training for safe end-toend driving. In Conference on Robot Learning, vol. 229 of Proceedings of Machine Learning Research, 2357–2372 (PMLR, 2023). 16. Xie, Y., Zhang, Y., Dai, K. & Yin, C. A real-time critical-scenario-generation framework for defect detection of autonomous driving system. IET Intell. Transp. Syst. 18, 114–128 (2024). 17. Zhu, B. et al. Critical scenarios adversarial generation method for intelligent vehicles testing based on hierarchical reinforcement architecture. Accid. Anal. Prev. 215, 108013 (2025). 18. Xiao, Y., Erden, M. S. & Wang, C. Controllable latent diffusion for traffic simulation. Preprint at https://arxiv.org/abs/2503.11771 (2025). 19. Wang, C., Kong, L., Tamborski, M. & Albrecht, S. V. HAD-Gen: Human-like and diverse driving behavior modeling for controllable scenario generation. Accid. Anal. Prev. 223, 108270 (2025). 20. Mattas, K. et al. Fuzzy surrogate safety metrics for real-time assessment of rear-end collision risk: a study based on empirical observations. Accid. Anal. Prev. 148, 105794 (2020). 21. Mattas, K. et al. Driver models for the definition of safety requirements of automated vehicles in international regulations. Application to motorway driving conditions. Accid. Anal. Prev. 174, 106743 (2022). 22. Japan Automobile Manufacturers Association. Automated driving safety evaluation framework ver. 1.0. Tech. Rep., Japan Automobile Manufacturers Association (JAMA) (2020). URL https://www.jama.or.jp/english/reports/framework.html. 13

23. Shalev-Shwartz, S., Shammah, S. & Shashua, A. On a formal model of safe and scalable self-driving cars. Preprint at https://arxiv.org/abs/1708.06374 (2017). 24. Chen, H. et al. Gaussian mixture flow matching models. In Proceedings of the 42nd International Conference on Machine Learning, vol. 267 of Proceedings of Machine Learning Research, 9783– 9802 (PMLR, 2025). 25. Kim, H. J., Yin, Z., Lai, L., Lee, J. & Ohn-Bar, E. Branchout: Capturing realistic multimodality in autonomous driving decisions. In Proceedings of the 9th Conference on Robot Learning, vol. 305 of Proceedings of Machine Learning Research, 1940–1952 (PMLR, 2025). 26. Black, K., Janner, M., Du, Y., Kostrikov, I. & Levine, S. Training diffusion models with reinforcement learning. In International Conference on Learning Representations (2024). 27. Caesar, H. et al. nuScenes: A multimodal dataset for autonomous driving. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11621–11631 (2020). 28. Zhang, Y. et al. The AD4CHE dataset and its application in typical congestion scenarios of traffic jam pilot systems. IEEE Trans. Intell. Veh. 8, 3312–3323 (2023). 29. Krajewski, R., Moers, T., Bock, J., Vater, L. & Eckstein, L. The rounD dataset: A drone dataset of road user trajectories at roundabouts in Germany. In 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), 1–6 (IEEE, 2020). 30. Zhong, Z. et al. Guided conditional diffusion for controllable traffic simulation. In Proc. IEEE Int. Conf. Robot. Autom. (ICRA), 3560–3566 (IEEE, 2023). 31. Zheng, Y., Xiao, Y., Zhu, Z., Erden, M. S. & Wang, C. CADiffusion: Controllable adversarial diffusion for attacking lane detection of autonomous vehicles. In Proc. IEEE Int. Conf. Intell. Transp. Syst. (ITSC), 4516–4522 (IEEE, 2025). 32. Ho, J., Jain, A. & Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, vol. 33, 6840–6851 (2020). 33. Janner, M., Du, Y., Tenenbaum, J. B. & Levine, S. Planning with diffusion for flexible behavior synthesis. In Proceedings of the 39th International Conference on Machine Learning, vol. 162 of Proceedings of Machine Learning Research, 9902–9915 (PMLR, 2022). 34. Jiang, C. et al. MotionDiffuser: Controllable multi-agent motion prediction using diffusion. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9644–9653 (2023). 35. Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal policy optimization algorithms. Preprint at https://arxiv.org/abs/1707.06347 (2017). 36. Schulman, J., Moritz, P., Levine, S., Jordan, M. I. & Abbeel, P. High-dimensional continuous control using generalized advantage estimation. Preprint at https://arxiv.org/abs/1506. 02438 (2015). 37. United Nations Economic Commission for Europe (UNECE). Proposal for a supplement to the 03 series of amendments to UN regulation no. 79 (steering equipment). Tech. Rep. ECE/TRANS/WP.29/GRVA/2020/7, UNECE World Forum for Harmonization of Vehicle Regulations (WP.29), Working Party on Automated/Autonomous and Connected Vehicles (GRVA) (2020). URL https://unece.org/DAM/trans/doc/2020/wp29grva/ ECE-TRANS-WP29-GRVA-2020-07e.pdf.

14

Supplementary Information S1

Lower-bound interpretation of FSM-attributed responsibility

This note formalizes the identifiability statement summarized in the main-text Methods (Responsibility attribution). For a generated collision scenario ξ, let ρtgt denote the recorded geometric path of the target vehicle and let Tadv denote the generated adv trajectory. For any target-control policy π, let Cπ (ξ) be the replay outcome obtained by holding Tadv fixed and applying π to the target, and let Collision(Cπ (ξ)) ∈ {0, 1} indicate whether that replay still collides. Let Π be the set of admissible target-control policies satisfying the kinematic envelope, and let Πlon ⊆ Π be the subset of policies that control only the target’s longitudinal speed or braking along the fixed path ρtgt . Define Prev(ξ) = 1[∃ π ∈ Π : Collision(Cπ (ξ)) = 0] , Prevlon (ξ) = 1[∃ π ∈ Πlon : Collision(Cπ (ξ)) = 0] ,

(S1) (S2)

[ FSM (ξ) = 1[Collision(CπFSM (ξ)) = 0] , Prev

(S3)

where πFSM is the FSM braking policy used in the path-constrained replay. The claim is [ FSM (ξ) ≤ Prevlon (ξ) ≤ Prev(ξ). Prev

(S4)

The argument proceeds in two parts: the chain of inequalities and the equality conditions. We use three assumptions only when discussing equality, not for the lower-bound inequality itself: • A1 (adversary-envelope validity). The fixed adv trajectory Tadv respects the kinematic envelope assumed by the reference model. • A2 (path-constrained replay validity). Evaluating longitudinal target control along ρtgt is an admissible approximation for the conflict class under study; i.e., the replay operator Cπ represents a longitudinal response on the fixed encounter geometry. • A3 (FSM-completeness on the longitudinal envelope). On the longitudinal conflict class addressed by UN R157, if any admissible longitudinal-only policy in Πlon can avoid the collision, then the FSM policy also avoids it. [ FSM (ξ) ≤ Prevlon (ξ). The FSM controller πFSM is by construction a member of Proof of Prev [ FSM (ξ) = 1, i.e. Collision(CπFSM (ξ)) = 0, Πlon (it brakes longitudinally and does not steer). If Prev then πFSM ∈ Πlon is a witness for the existential statement defining Prevlon (ξ), so Prevlon (ξ) = 1. [ FSM (ξ) = 0) holds because if no Πlon controller can The contrapositive (Prevlon (ξ) = 0 ⇒ Prev avoid the collision, then πFSM ∈ Πlon in particular cannot. Pointwise dominance of indicators yields the inequality. Proof of Prevlon (ξ) ≤ Prev(ξ).

Πlon ⊆ Π by definition. Hence

{∃ π ∈ Πlon : Collision(Cπ (ξ)) = 0} ⊆ {∃ π ∈ Π : Collision(Cπ (ξ)) = 0}, so the indicator of the former is bounded above by that of the latter. Equality of the first inequality under A1–A3. Assumption A3 (FSM-completeness) is precisely [ FSM (ξ) = 1 on the longitudinal envelope addressed by the implication Prevlon (ξ) = 1 ⇒ Prev [ UN R157. Combined with PrevFSM (ξ) ≤ Prevlon (ξ) above, this yields equality on that subset. A1 ensures that FSM’s evaluation is conducted on an adv trajectory inside the kinematic envelope FSM itself assumes; A2 ensures that the path-constrained replay operator is the same one used to define Πlon . Without A1 or A2, the regulatory interpretation and equality claim should not be invoked, although the lower-bound inequality remains valid whenever CπFSM is still a member of the replay class represented by Πlon . Equality of the second inequality. Prevlon (ξ) = Prev(ξ) holds whenever Prev(ξ) = 1 ⇒ Prevlon (ξ) = 1, i.e. whenever the scenario’s preventability is realizable through longitudinal control alone. Scenarios that are preventable only through steering or combined longitudinal–lateral control fall in the strict-inequality regime. 15

Aggregate consequence. Summing pointwise inequalities over the corpus and dividing by |Scol | preserves the inequalities. Hence 1 X[ 1 X 1 X r̂FSM = PrevFSM (ξ) ≤ Prevlon (ξ) ≤ Prev(ξ). (S5) |Scol | |Scol | |Scol | ξ

ξ

ξ

The reported r̂FSM = 88.7% is therefore a deterministic lower bound on the true preventability fraction across the corpus, regardless of which subset of scenarios is Πlon -realizable, and is exact on the longitudinal-realizability subset under A1–A3. Bias under relaxation. The argument relies on no probabilistic structure (no scenario sampling distribution, no expectation over noise); the bound is deterministic and pointwise. The lower-bound conclusion in (S4) requires only that CπFSM corresponds to a longitudinal controller in Πlon replayed on ρtgt ; it remains valid under wider conditions than the equality. If A1 or A2 fails, the regulatory interpretation and equality claim should not be invoked; the lower-bound claim still holds as long as the FSM replay remains a longitudinal controller in Πlon . Relaxing A3 alone changes only the equality on the Πlon -realizable subset, leaving the lower-bound conclusion intact.

S2

Adversarial Agent Selection: Feature Vector

The 13-dimensional feature vector ϕi,t for each surrounding agent i relative to tgt is defined at control step t as follows. Let tgt have position ptgt,t , velocity vtgt,t , speed vtgt,t , heading θtgt,t , and length ∥ ⊤ Ltgt . Let etgt,t = (cos θtgt,t , sin θtgt,t )⊤ and e⊥ tgt,t = (− sin θtgt,t , cos θtgt,t ) denote its longitudinal and lateral unit vectors. For each surrounding agent i with state (pi,t , vi,t , vi,t , θi,t , Li ), define ∥ ei,t = (cos θi,t , sin θi,t )⊤ and ∆pi,t = pi,t − ptgt,t . di,t = ∥∆pi,t ∥,

∆vi,t = vi,t − vtgt,t ,

ci,t = cos(θi,t − θtgt,t ),

si,t = ∆p⊤ i,t etgt,t ,

⊥ li,t = ∆p⊤ i,t etgt,t , l˙i,t = (vtgt,t − vi,t )⊤ e⊥ tgt,t ,

ṡi,t = (vtgt,t − vi,t )⊤ etgt,t , gi,t = si,t −

TTCi,t = gi,t / max(ṡi,t , ϵ),

βi,t =

∥

∥

Ltgt +Li − b, 2 ∥ (ptgt,t − pi,t )⊤ ei,t

(S6)

, max(∥ptgt,t − pi,t ∥, ϵ) (5) (5) d˙i,t = (di,t − di,t−5 )/(5∆t), ∆|l|i,t = |li,t | − |li,t−5 |, where di,t is Euclidean distance, ∆vi,t relative speed, ci,t heading alignment, si,t /li,t longitudinal/lateral offset, ṡi,t /l˙i,t longitudinal/lateral closing speed, gi,t bumper-to-bumper gap (b is a bumper (5) margin), TTCi,t time to collision, βi,t the cosine of the bearing angle from agent i toward tgt, d˙ i,t

(5)

the distance change rate over the most recent 5 steps, and ∆|l|i,t the change in absolute lateral offset. Together with the agent speed vi,t , this yields a 13-dimensional feature vector ϕi,t .

S3

Diffusion Model: Forward Process

The forward diffusion process defines a Markov chain that gradually corrupts the clean action sequence τa0 into Gaussian noise over R diffusion steps. Each forward transition q(τar | τar−1 ) is a fixed Gaussian kernel:  √ q(τar | τar−1 ) = N αr τar−1 , βr I , (S7) Qr R where {βr }r=1 is a predefined variance schedule, αr = 1 − βr , and ᾱr = u=1 αu . The resulting marginal distribution admits the closed-form: √ √ τar = ᾱr τa0 + 1 − ᾱr ϵ, ϵ ∼ N (0, I). (S8)

S4

GMDM Loss Components

The total loss in (6) comprises three terms: Ltotal = LNLL − λH Hπ + λrepel Lrepel . 16

(a) Negative log-likelihood.

The negative log-likelihood term is LNLL (θ) = −

T X

 0 log pθ τa,h | τar , c, r .

(S9)

h=1

(b) Entropy regularization.

The mixture-weight entropy term is ! T K X X Hπ = − πh,k log(πh,k ) , h=1

(S10)

k=1

where K denotes the number of Gaussian components and πh,k is the weight of component k at horizon step h. (c) Component repulsion.

The component-repulsion term is   T X 1 ∥µh,k − µh,ℓ ∥22 1 X , exp − Lrepel = T K(K − 1) s2 h=1

(S11)

k̸=ℓ

where s controls the distance scale for component separation.

S5

PPO Fine-Tuning: Detailed Formulation

The reverse denoising process is formulated as a multi-step Markov decision process. At diffusion step r, the state and policy are defined as sr = (τar , r, c), The objective is to maximize

πθ (τar−1 | sr ) = pθ (τar−1 | τar , c).

  J(θ) = Eπθ R(τa0 , c) .

The policy gradient across R denoising steps is " R # X r−1 r 0 ∇θ J = E ∇θ log pθ (τa | τa , c) R(τa , c) .

(S12) (S13)

(S14)

r=1

The clipped PPO surrogate loss is " R #  X  (r)  (r) (r) (r) Lπ (θ) = −E min ρθ Â , clip ρθ , 1−ϵ, 1+ϵ Â ,

(S15)

r=1 (r)

where ρθ = pθ (τar−1 | τar , c) / pθold (τar−1 | τar , c) is the importance ratio, Â(r) is the generalized advantage estimate at denoising step r, and ϵ is the clipping coefficient. lon Reward components. Let dlon 0 and d1 denote the longitudinal component of the adv–tgt relative position projected onto the tgt heading at the start and end of an n-step shaping window, and ℓ0 , ℓ1 the corresponding lateral components. The individual reward terms in Rtotal ((7)) are:  lon  d − dlon Rprog = clip 0 lon 1 , −0.5, 1 , d0 + ν  exp(−ℓ1 /σℓ ) if ℓ0 > ℓ1 , Rℓ = −0.5 exp(−ℓ1 /σℓ ) otherwise,   ℓ0 − ℓ 1 Rℓc = clip , −0.5, 1 , (S16) ∆teff (ℓ0 + ν)

where σℓ is a lateral decay scale, ∆teff is the shaping-window duration, and ν > 0 is a small constant for numerical stability. The clip at −0.5 allows the reward to penalize diverging motion up to a bounded magnitude, so the adv avoids pathological retreat without being driven off-manifold by a sharp negative gradient. 17

S6

FSM: Full Specification

The implementation follows UN Regulation No. 157 8 . The regulatory parameters are the reaction time τ =0.75 s, the tgt comfort deceleration btgt,comf =4 m/s2 , the tgt maximum deceleration btgt,max =6 m/s2 , the assumed adv maximum deceleration badv,max =7 m/s2 , the minimum standstill distance dmin =2 m, and the jerk limit J=12.65 m/s3 . Here btgt denotes the commanded positive deceleration magnitude before the reaction-delay and jerk-ramp implementation. vadv and vtgt denote longitudinal speeds; atgt is the signed longitudinal acceleration of tgt. dlon and dlat denote the longitudinal and lateral separations in the tgt-aligned frame, vadv,lat is the adv lateral approach speed toward tgt, and Ladv and Ltgt denote vehicle lengths. Lateral safety pre-check. Lateral risk is identified only when all four conditions hold; failure of any condition causes FSM to skip the longitudinal fuzzy stage and apply no braking: (a) dlon > 12 (Ltgt + Ladv ) (adv rear ahead of tgt front), (b) vadv,lat > 0 (adv moving toward tgt laterally), (c) vtgt > vadv (tgt closing on adv longitudinally), dlon + Ltgt + Ladv dlat (d) < + 0.1. vadv,lat vtgt − vadv Fuzzy membership.

(S17)

PFS and CFS share a saturated linear membership µ ∈ [0, 1]:   xsafe − x µ(x; xsafe , xunsafe ) = clamp , 0, 1 . xsafe − xunsafe

(S18)

PFS. dPFS safe = vtgt τ + dPFS unsafe = vtgt τ +

2 vtgt

2btgt,comf 2 vtgt

2btgt,max

− −

2 vadv

2badv,max 2 vadv

2badv,max

+ dmin , (S19)

,

PFS PFS = µ(dlon − dmin ; dPFS safe , dunsafe ).

The dmin safety distance appears both as a constant offset in dPFS safe and as a shift on the membership input. CFS. CFS is computed only when vtgt > vadv ; otherwise CFS=0. ∗ max(atgt , −btgt,comf ) and vtgt = vtgt + a′tgt τ . Two branches:

Define a′tgt

=

 (vtgt − vadv )2 (vtgt − vadv )2   ∗  , , vtgt ≤ vadv ,   2|a′tgt | 2|a′tgt | CFS CFS (S20) dsafe , dunsafe =  ∗ (v ∗ − vadv )2 (vtgt − vadv )2    ∗  dnew + tgt , dnew + , vtgt > vadv , 2btgt,comf 2btgt,max  ∗ CFS where dnew = (vtgt + vtgt )/2 − vadv τ . Then CFS = µ(dlon ; dCFS safe , dunsafe ). The first branch handles the case where the tgt would already be slower than the adv after the reaction window, in which case dsafe = dunsafe degenerates the membership to a step function under UN R157. Braking response. ( btgt =

CFS (btgt,max − btgt,comf ) + btgt,comf ,

CFS > 0,

PFS btgt,comf ,

CFS = 0.

(S21)

When CFS is non-zero, the reaction is at least btgt,comf and rises linearly to btgt,max as CFS approaches one. When CFS vanishes, PFS scales the comfort deceleration only. 18

Reaction-time and jerk-limited ramp. The deceleration is implemented after a delay τ , then increases at the constant jerk rate J = 12.65 m/s3 until it reaches btgt : ( 0, t < ttrig + τ,  bactual (t) = (S22) min bactual (t−∆t) + J∆t, btgt (t) , t ≥ ttrig + τ, where ttrig is the first step at which btgt > 0. The path-constrained replay in Section S7 integrates bactual as the FSM-commanded deceleration aFSM (t).

S7

Path-Constrained Re-Simulation

The tgt follows its original recorded path, parameterized by arc length s; only its speed is modified by the reference models. Let vorig (t) denote the tgt’s original speed at step t. Let aref (t) ≥ 0 denote the reference-model deceleration command along this fixed path; for FSM, aref (t) = aFSM (t) from Equation (S22), which already includes the reaction delay and jerk-limited ramp. A cumulative speed deficit ∆v then accumulates this commanded deceleration:  ∆v(0) = 0,     ∆v(t) = ∆v(t−1) + a (t) · ∆t, ref  (S23)  v(t) = max 0, vorig (t) − ∆v(t) ,    s(t+1) = s(t) + v(t) · ∆t, where ∆t is the simulation time step. When ∆v = 0, the re-simulated trajectory reproduces the original recording exactly.

S8

Gaussian Mixture Component Selection

We select the number of Gaussian mixture components in the GMDM by training models with K ∈ {2, 3, 4, 5, 6, 7, 8} and measuring the normalized mixture-weight entropy Hnorm at convergence. Figure S1 reports the sweep across five independent seeds. We use K=8 in the paper because it achieves the highest single-seed utilization on the PPO-training seed (Hnorm =0.82) and the narrowest seed-to-seed spread among the tested values, indicating reliable mixture-component usage.

Figure S1: Normalized entropy Hnorm of the GMDM mixture weights at convergence, as a function of the number of Gaussian components K. The dark line is the seed used for PPO training (seed 42) and the light envelope is the range across five independent seeds. K=8 achieves the highest single-seed utilization (Hnorm =0.82, highlighted in red) and is used throughout the paper.

S9

Statistical Robustness across Random Seeds

We repeat the closed-loop rollout with six additional seeds (seven seeds in total) using the same PPO-fine-tuned adv policy and the same nuScenes scene–agent pairs. Figure S2 reports collision 19

count, FSM attribution, Hcrit , and median TTCmin across seeds. Hcrit is computed on each seed’s FSM-preventable subset, matching the definition used in main-text Table 1. a. Collision Episodes

b. FSM attribution (%)

c. Hcrit

d. TTCmin median (s)

Figure S2: Statistical robustness across seven random seeds. Each panel shows the distribution of a key metric: box plot (median and interquartile range), individual seed values (dots), and mean (diamond).

S10

CARS Scenario Quality Stratified by FSM Criticality

Supplementary Tables S1–S3 provide the tier-level statistics underlying the severity-diversity and feasibility analyses in the main text. Supplementary Table S1 reports per-tier scenario quality for FSM-preventable CARS collisions. Supplementary Table S2 reports the Hard/Medium/Easy distribution used to compute Hcrit across methods and transfer settings. Supplementary Table S3 recomputes Hcrit after applying stricter kinematic-feasibility requirements.

S11

Adversary Classifier Evaluation

Supplementary Table S4 reports binary-classification and scene-level ranking metrics for the contextaware adv selector on the validation split.

S12

Cross-Reference Comparison

Supplementary Table S5 compares FSM, CC-JP, and RSS on the same 408 CARS collision scenarios on nuScenes. Quantities restricted to unpreventable cases are computed separately for each reference model’s unpreventable subset.

S13

Full adv Trajectory Kinematics

Supplementary Table S6 reports the adv trajectory kinematics underlying the IP% column in main-text Table 1. Higher-order kinematic estimation from discretely sampled position is intrinsically noise-amplifying: triple finite differencing of cm-precision position at ∆t = 0.1 s amplifies measurement noise by ∼ 103 , producing a physically meaningless jerk floor on the order of 20 m/s3 even for smooth motion. This is an estimator-level requirement and is independent of any subsequent feasibility evaluation. We therefore apply a Savitzky-Golay filter (window 7 frames at 10 Hz, cubic polynomial) to the position trace before differencing for all baselines; the table reports the mean of per-scenario maxima and 95th-percentile values. The RounD entry additionally subtracts the geometric centripetal floor v 2 /rlocal from |alat | so that normal cornering on the fixed roundabout geometry is not counted as a lateral violation. We write j = da/dt for longitudinal jerk. IP% uses the bounds |a| ≤ 7 m/s2 8 , |j| ≤ 12.65 m/s3 8 , and |alat | ≤ 3.0 m/s2 37 .

20

Supplementary Tables

Table S1 | CARS scenario quality by FSM criticality tier on nuScenes. Metric

Hard

Medium

Easy

Overall

Episodes (n, %) 26 (7.2) TTCmin median (s) ↓ 0.61 BD+ % ↑ 19.2 d¯min (m) ↓ 1.49

207 (57.2) 2.39 15.9 3.49

129 (35.6) 3.25 10.1 4.64

362 2.72 14.1 3.76

Table S2 | FSM criticality tier distribution across methods. Method / Configuration

N

Hard%

Medium%

Easy%

Hcrit

Baselines (nuScenes) STRIVE 12 SafeSim 13 Bezier-CAT 15

30 13 121

73.3 46.2 91.7

26.7 53.8 0.0

0.0 0.0 8.3

0.528 0.628 0.260

Ours (nuScenes) CARS CARS (K=1 adv )

362 28

7.2 35.7

57.2 57.1

35.6 7.1

0.798 0.797

ADS-planner robustness (fixed CARS adv) One-component diffusion planner 158 8.9 CTG planner 30 271 6.6

53.8 64.9

37.3 28.4

0.834 0.745

Cross-dataset generalization CARS on AD4CHE 28 CARS on RounD 29

66.9 69.0

24.5 15.2

0.751 0.759

359 533

8.6 15.8

Table S3 | FSM-preventable severity diversity after applying kinematic feasibility requirements. N0 and H0 additionally require zero per-step kinematic violation (IP = 0). N0.10 and H0.10 allow at most 10% violating time steps. Method / Configuration

Nprev

Hcrit

N0

H0

N0.10

H0.10

CARS CARS (K=1 adv )

362 28

0.798 0.797

352 7

0.796 0.373

362 9

0.798 0.622

STRIVE 12 SafeSim 13 Bezier-CAT 15

30 13 121

0.528 0.628 0.260

3 0 0

0.000 NA NA

4 8 0

0.000 0.512 NA

Table S4 | Adversary classifier evaluation on the validation split (62 scenes, 2,131 agent-level samples, 135 tgt–adv pairs). Metric

Value

Binary classification (per class) Precision / Recall / F1 (adversary) Precision / Recall / F1 (non-adversary)

0.75 / 0.79 / 0.77 0.99 / 0.98 / 0.98

Scene-level ranking of the labeled adversary Top-1 / Top-3 / Top-5 accuracy 83.0% / 97.8% / 99.3%

21

Table S5 | Reference-model comparison over 408 nuScenes scenarios. Metrics defined in UP Supplementary Table S1; v̄col : mean collision speed (m/s) for unpreventable scenarios only. Reference model

UP%

TTCmin ↓

d¯min ↑

BD+ %

UP ↓ v̄col

FSM CC-JP RSS

11.3 20.3 2.9

2.46 0.72 2.19

3.32 0.83 4.41

22.5 56.4 15.0

6.01 5.06 5.47

Table S6 | Full adv trajectory kinematics. Here j = da/dt denotes longitudinal jerk. Subscript max denotes the per-trajectory maximum magnitude; subscript 95 denotes the 95th percentile of the magnitude. Method

Longitudinal acceleration

Jerk

Lateral / yaw

amax (m/s2 )

a95 (m/s2 )

jmax j95 alat,95 θ̇max (m/s3 ) (m/s3 ) (m/s2 ) (rad/s)

Baselines STRIVE 12 SafeSim 13 Bezier-CAT 15

6.12 6.61 28.93

5.16 4.68 24.60

41.52 37.38 24.90 16.30 57.43 56.20

3.28 0.376 1.07 0.317 18.70 10.337

Ours CARS CARS (K=1 adv )

2.26 3.83

1.90 3.35

8.49 6.36 18.07 16.14

0.42 2.13

0.105 0.517

ADS-planner robustness (fixed CARS adv) One-component diffusion planner 2.24 CTG planner 30 2.14

1.88 1.83

8.18 7.94

6.20 6.08

0.44 0.41

0.109 0.097

Cross-dataset generalization CARS on AD4CHE 28 CARS on RounD 29

0.38 2.71

1.64 1.22 22.59 12.14

0.16 1.84

0.039 0.458

0.45 3.29

22

S14

Three Evaluation Datasets in Detail

The evaluation suite spans three continents, three road topologies, and contrasting driving cultures (Figure S3). nuScenes provides 1,000 urban scenes recorded in Boston (US) and Singapore. RounD provides 22 drone-recorded sessions of a four-arm signalised roundabout in Neuweiler, Germany. AD4CHE provides 68 drone recordings of multi-lane highway traffic across four Chinese cities. Together the three datasets stress-test CARS under signalised intersections, circular merging, and high-speed lane changing respectively.

a

b

c

a) nuScenes

b) RounD

c) AD4CHE

Car-mounted cameras + LiDAR + HD map Boston, USA + Singapore

Overhead drone trajectories

Overhead drone trajectories

Neuweiler, Germany

Xi'an, Changchun, Hefei, Shenzhen

Urban intersections, lane changes

Four-arm signalised roundabout

Multi-lane highways, high speed

Scenes

1000

Scenes

22

Duration

4.65 h

Duration

1.66 m/s

Mean Speed

Mean Speed

d

Scenes

68

6.03 h

Duration

5.12 h

6.18 m/s

Mean Speed

8.23 m/s

Figure S3: Three evaluation datasets spanning three continents and three road topologies. a, nuScenes front-facing camera frame at a Boston (US) intersection (1,000 urban scenes across Boston and Singapore). b, RounD overhead drone frame at the Neuweiler (Germany) roundabout (22 recordings of four-arm signalised geometry). c, AD4CHE overhead drone frame on a multi-lane highway (68 recordings across four Chinese cities). d, Geographic locations of the four data-collection countries; pins are colored by dataset.

S15

Gallery of CARS Collision Geometries on nuScenes

In addition to the severity diversity quantified in main-text Table 1 and the per-tier kinematic statistics in main-text Fig. 4, CARS also produces collisions across a wide variety of geometric interaction types. Figure S4 illustrates six representative cases drawn from the 408 nuScenes collision scenarios: rear-end approach, cut-in from ahead, lane drift across the boundary, two angled-approach geometries, and near-perpendicular side impact. The variety of approach geometries shows that CARS is not biased toward a single failure mode of the tgt policy; rather, it discovers heterogeneous configurations under which the same policy fails.

23

Figure S4: Gallery of collision geometries generated by CARS on nuScenes. a, Rear-end approach. b, Cut-in from ahead. c, Lane drift across the boundary. d, Angled approach from the front-left. e, Angled approach from the front-right. f, Near-perpendicular side impact. The tgt is shown in blue and the adv in red; filled boxes mark the start position and outlined boxes the end position. Other agents are shown in gray.

24

Record · ID 180715 · SHA-256 bdb85468a005fb1b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.