Conceptio › Archive › arXiv CS
arXiv CSopen access

A distribution-free certification framework for trustworthy crash-severity prediction

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A distribution-free certification framework for trustworthy crash-severity prediction Amir Rafea,∗ , Subasish Dasa a Civil Engineering, Texas State University, 601 University Drive, San Marcos, 78666 TX, USA

ABSTRACT

Keywords: conformal prediction distribution-free inference partial identification uncertainty quantification crash injury severity trustworthy artificial intelligence

Crash-severity models inform screening, dispatch and site prioritization, yet are deployed without a finite-sample statement of what one prediction means. Off-the-shelf guarantees fail here, because the features that make crash severity distinctive defeat them: the KABCO outcome is ordinal, the recorded label is a field assessment agreeing with medical severity about half the time, erring in a structured way, and deployment crosses jurisdictions and years calibration never saw. We develop a certification layer that wraps any severity model unmodified, with distributionfree guarantees using this structure: contiguous ordinal sets that read as “B or worse”; perclass validity for any pre-declared partition, with an oracle efficiency characterization; transfer of coverage to unobserved true severity through a declared reporting band, with a worst-case sharpness result; a one-sided certificate under deployment shift; and severity-weighted risk control. The guarantees compose with an attributable slack budget. The same analysis bounds what certification can achieve. A certified set’s informativeness is governed by a functional of the true law that no base model can evade and that cannot be lower-bounded distribution-free; given a declared misreporting channel identified from record-linkage data, a nonvacuous lower bound on that floor becomes computable. On 5.2 million Texas records across seven base models spanning four decades, the layer attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats, separating it from a remainder that stays bounded but distribution-free unidentifiable. The framework is released as an open-source package with theorem-level tests.

arXiv:2609.11592v1 [stat.ML] 10 Sep 2026

ARTICLE INFO

1. Introduction

A crash-severity model is a claim about an injury that has already happened and been written down by someone at the roadside. Such models are increasingly used to decide things: which corridors receive countermeasure funding, how a dispatcher weighs a call, which sites a safety office screens first. The models themselves have improved steadily for two decades, from ordered response specifications through random-parameters and latent-class formulations to gradient boosting and tabular foundation models. What has not arrived alongside them is any statement about what a single prediction means. A deployed severity model returns a category, or a probability vector, and the agency acting on it has no finite-sample guarantee of any kind attached to that output. The gap between the sophistication of the estimator and the emptiness of the accompanying guarantee is the subject of this paper. The gap matters because the three structural features that make crash severity statistically interesting are exactly the features that make a naive guarantee wrong. The outcome is ordinal. KABCO severity is ordered, from no injury through possible, non-incapacitating and incapacitating injury to fatality, and the ordering carries the operational meaning. A safety office does not want a set of severities scattered across the scale; it wants a statement of the form “B or worse”, which is an interval. A prediction set that omits the middle of the scale while including both ends is mathematically permissible and operationally meaningless. The label is noisy in a structured way. The severity recorded on a crash report is a police officer’s field assessment, not a medical diagnosis, and linkage studies that match crash reports to medical records find that the two agree roughly half the time. The disagreement is not random. It concentrates in adjacent categories, runs mostly in the under-reporting direction, varies by reporting agency, and behaves differently at the two ends of the scale, with fatality recording close ∗ Corresponding author

[email protected] (A. Rafe); [email protected] (S. Das) ORCID (s): 0000-0002-4089-2088 (A. Rafe); 0000-0002-1671-2753 (S. Das)

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 1 of 45

Certified crash-severity prediction

to exact. Every model in the literature is therefore trained and evaluated against a label that is not the quantity anyone cares about, and a coverage guarantee on the reported label is a guarantee about the report rather than about the injury. The deployment is hierarchical and shifting. Calibration happens on the data an agency has, and deployment happens in a county, a district, or a year that the calibration did not see. The severity literature has documented that the underlying relationships are themselves unstable across time. A guarantee established on one distribution is not automatically a guarantee anywhere else, and the honest question is not whether transfer happens but whether one can tell when it has failed. Layered on these is the field’s own central doctrine. The heterogeneity program in severity analysis has argued for two decades that unobserved heterogeneity is not a nuisance to be averaged away but the dominant structural feature of crash data. That claim has a consequence for guarantees: one holding on average over a heterogeneous population is a guarantee about the majority of that population, and the minority subpopulations, which in crash data are the severe ones, are where an averaged guarantee is least trustworthy and most needed. The consequence is not ours to claim as an observation. It has been documented in other domains, in drug discovery, where marginal conformal prediction meets its global target while minority-class coverage collapses (Tursunbadalov and Tursunbadalov, 2026), and in survey measurement, where marginal validity coexists with large subgroup gaps (Rafe and Das, 2026). What those studies report empirically, and in one case with a discouraging verdict on the repair, this paper derives: the undercoverage is a structural consequence of pooling a quantile across heterogeneous score distributions, not a defect of any base model, and it therefore cannot be fitted away. Conformal prediction supplies the natural starting point, since it converts any model’s output into a set with finitesample, distribution-free coverage under exchangeability alone. But the guarantee it supplies off the shelf is precisely the one the three features above defeat. It covers the label it was calibrated on, which is the report and not the injury. It covers marginally, which averages over exactly the heterogeneity the field says dominates. It assumes exchangeability, which deployment across jurisdictions and years breaks. Applying the recipe unchanged to a severity model inherits the recipe’s guarantees, not the ones the problem requires. This paper takes the opposite approach and asks what can be proved if the structure is used rather than ignored. Ordinality is what turns a bounded reporting error into a bounded set expansion, because it is what makes the set of severities compatible with a report a small part of the scale rather than all of it; contiguity is then what keeps that expansion additive rather than multiplicative, and what makes the output a statement a dispatcher can act on. The band structure of KABCO misreporting is stable even where its magnitudes are not, which makes a declared support band a more defensible premise than an estimated confusion kernel. The jurisdictional hierarchy is a natural set of predeclared groups, which is the strongest conditioning that the impossibility results permit. And severity itself supplies the cost asymmetry: omitting a fatality from a prediction set is not exchangeable with omitting a no-injury record, and a framework that treats them alike has failed at the thing it was built for. The result is a certification layer that wraps any severity model without modifying it, and attaches guarantees that respect what the model is predicting. Two of these results are as much impossibility statements as guarantees, and together they are the paper’s methodological center. The informativeness of a certified set, as distinct from its validity, is governed by a functional of the true severity law that no base model can evade and that no procedure can lower-bound without assumptions; the structured reporting channel is the assumption that makes that floor computable from record-linkage data. What the layer certifies on the strata a screening system exists to serve is therefore not only coverage, but a quantitative account of how much of the residual uncertainty the measurement process imposes and no model can remove. A quantity that is impossible to certify without a measurement model is identified with one, which is the sense in which the paper is an identification result rather than an application of conformal prediction.

Non-goals. This paper makes no causal claims, and the restriction is deliberate rather than a limitation to be excused later. Every statement is about the coverage of a prediction set or the expectation of a risk functional, conditional on the crash-reporting process that generated the data. Where a covariate is described as associated with severe outcomes, the statement is about reporting frequencies in the records and not about what would happen under an intervention. No result licenses a claim that a road feature, a vehicle technology, or a driver attribute produces an injury outcome. The decoded vehicle fields in particular record equipment availability rather than equipment engagement, so nothing here is a statement about the effectiveness of any safety system. Identification is a separate and legitimate research program with its own requirements, and certification neither substitutes for it nor competes with it. The two compose: a certified prediction set answers what is plausible given what was recorded, and a causal analysis answers what would change under an intervention, and an agency needs both. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 2 of 45

Certified crash-severity prediction

The remainder of the paper is organized as follows. Section 2 positions the work against the severity, transportationuncertainty and conformal literatures. Section 3 develops the framework and its guarantees, with the proofs of the three main theorems deferred to A. Section 4 describes the Texas records and the experimental design. Section 5 reports the experiments, each of which tests a theorem. Section 6 sets out what the guarantees do not mean and how they would be deployed, and the concluding section states the contributions, the limitations that survive them, and what follows.

2. Related work Three literatures meet in this paper: the crash-severity modeling tradition that supplies the problem and the base models, the uncertainty-quantification work that has begun to bring coverage language into transportation, and the conformal foundations that supply the mathematical ingredients. Each is mature. None of them, separately or in combination, supplies what a deployed severity model needs.

2.1. Severity modeling and the measurement of severity Crash-severity analysis has been an ordered-response problem since McCullagh (1980), and the methodological arc since then has been a sustained effort to make the conditional distribution more flexible. Savolainen, Mannering, Lord and Quddus (2011) survey the alternatives and their trade-offs. The heterogeneity program associated with Mannering and colleagues is the central development: Mannering, Shankar and Bhat (2016) argue that unobserved heterogeneity is not a nuisance to be averaged away but the dominant structural feature of crash data, and Mannering and Bhat (2014) set the methodological frontier that random-parameters and latent-class specifications were built to address. Mannering (2018) adds the temporal dimension, showing that the relationships themselves are unstable across time, which is a statement about deployment as much as about estimation. More recently the field has absorbed machine learning and deep models, and Seyfi, Karimi Mamaghan, Behnood and Mannering (2025) assess what that substitution does and does not buy. Our relationship to this literature is deliberately not competitive. The heterogeneity program identified the right structure, and this paper takes that structure as an input rather than relitigating it: the latent classes that a latentclass ordered logit produces are exactly the partition that Section 3.3 conditions on, and Section 5.1 wraps a randomparameters ordered logit inside the certification layer rather than arguing against it. What the tradition does not provide, and does not claim to provide, is a finite-sample statement about what a fitted model’s output means for one crash. A more flexible conditional distribution is a better estimate; it is not a guarantee. The second thread in this stream is measurement, and it is the one that motivates Section 3.6. Police-reported KABCO severity is not medically assessed severity. Burdett, Li, Bill and Noyce (2015) and the underlying linkage study (Burdett, 2014) match Wisconsin crash reports to medical records and find that agreement between the police rating and the medically referenced category is close to half, with the disagreement concentrated in the under-reporting direction and heavily structured by category. Burdett, Bill and Noyce (2022) show that the discrepancy varies systematically by reporting agency, which is precisely why this paper declares a band rather than estimating a confusion matrix: a kernel estimated in one jurisdiction and era does not transfer, while the band structure does. Taylor, Fliss, Schiro and Harmon (2024) reach a similar conclusion from linked North Carolina trauma-registry data, reporting sensitivity near 50% for police identification of serious injury. Farmer (2003) and Compton (2005) provide the national comparisons, and both support treating fatality recording as exact while leaving the nonfatal categories uncertain, which is the asymmetry Remark 3.38 exploits. This literature establishes that the label a severity model is trained and evaluated on is noisy in a specific, ordered, category-dependent way. It does not provide a way to state what a prediction means for the underlying injury.

2.2. Uncertainty quantification in transportation Transportation has begun to take uncertainty seriously, and the movement is real. Qian, Zhao, Zhang, Chen, Zheng and Zhou (2024) survey uncertainty quantification for traffic forecasting across Bayesian, ensemble, bootstrap and calibration approaches. Conformal methods specifically have arrived: Patil, Ahmed and Midlam-Mohler (2024) wrap a graph neural network for travel-time forecasting, Yang, Huang, Qiu and Cheng (2026) give coverage-guaranteed intervals for traffic demand, and Bohlouli, Varghese, Gentile and Eldafrawi (2025) apply Mondrian conformal prediction to mode choice. In crash analysis, Islam, Wang and Abdel-Aty (2024) pursue calibrated confidence for real-time crash and severity prediction. Closest to this paper is Wei, Miyazaki and Sato (2026), who apply generic and class-conditional conformal prediction to pre-crash injury risk and obtain a guaranteed confidence level. That work demonstrates both the demand A. Rafe and S. Das: Preprint submitted to Elsevier

Page 3 of 45

Certified crash-severity prediction

for coverage guarantees in injury prediction and the readiness of the field to use them, and we regard it as a predecessor rather than a competitor. The distinction we draw is between application and foundations. Applying split conformal prediction to a transportation problem inherits exactly the guarantees split conformal prediction already has: marginal coverage of the label the model was calibrated on, under exchangeability, with no statement about subpopulations, about the true injury behind a noisy code, or about a jurisdiction the calibration never saw. Those are not oversights in that work. They are the boundary of the tool it applies. The contribution available here is to move the boundary, by proving new statements that use the structure of the severity problem rather than importing a recipe that ignores it.

2.3. Conformal foundations The conformal framework originates with Vovk, Gammerman and Shafer (2005), and split conformal prediction in the form used here is standard (Lei, G’Sell, Rinaldo, Tibshirani and Wasserman, 2018). Its limits are equally well established: Vovk (2012) and Foygel Barber, Candès, Ramdas and Tibshirani (2021) prove that exact covariateconditional coverage is unattainable without trivial sets, which is why every conditional statement in this paper is made with respect to finitely many pre-declared groups. Work on conditional guarantees continues along that line, through Ding, Angelopoulos, Bates, Jordan and Tibshirani (2023) for many-class settings, Gibbs, Cherian and Candès (2025) for covariate-shift-indexed classes, Kiyani, Pappas and Hassani (2024) for learned partitions, and Dunn, Wasserman and Ramdas (2023) for hierarchical structure. Ordinal conformal prediction is an active subfield. Lu, Angelopoulos and Pomerantz (2022) and Chakraborty, Tyagi, Qiao and Guo (2024) construct CDF-based ordinal scores of the kind Definition 3.3 uses, and Xu, Guo and Wei (2023) formulate conformal risk control for ordinal losses. Set size in the ordinal setting has its own literature: Zhang, Chen, Shi, Ma, Xu and Yan (2025) construct minimum-length ordinal sets that are optimal per instance given the base model’s scores, which is a reminder that width has a set-construction component and is not simply a property of the estimator underneath. Our sets are not built for minimal length, and where this paper reports width it reports the width of the construction it certifies rather than the best achievable. The undercoverage of minority subpopulations under marginal calibration has recently been documented outside transportation, and the two nearest reports appear to disagree about the remedy. Tursunbadalov and Tursunbadalov (2026) find, in molecular property prediction, that marginal conformal prediction meets its global target while minorityclass coverage falls far below it across three model families, and that class-conditional calibration restores it at a modest cost in set size. Rafe and Das (2026) find, in survey-based social measurement, the first half but not the second: marginal validity coexists with subgroup gaps near thirteen percentage points, and group-specific calibration there worsened the efficiency and fairness trade-off rather than repairing it. The disagreement is not about whether conditioning works. It is about how much calibration data each cell receives, and both studies say so. The survey study’s own failure analysis attributes the result to calibration-cell fragmentation: its thinnest cell holds 32 calibration observations and its largest holds 355, the per-group quantile is estimated from those handfuls, and it overreacts to local idiosyncrasies in the score distribution, adding variance faster than it removes bias. That is Theorem 3.13(ii) read at the wrong end of its rate, since the efficiency of a class-conditional threshold −1∕2 degrades as 𝑂𝑝 (𝑛𝑐 ) and 𝑛𝑐 of order 30 is where the guarantee is formally intact and practically useless. The same paper’s regularized comparator, which shrinks thin-cell thresholds toward the global quantile, is the natural response and is exactly the partial-pooling instinct that the rollup rule of (38) encodes structurally. Our setting sits at the other end of that rate rather than on the other side of an argument. The thinnest cell used anywhere in this paper holds 4,125 calibration records, which is more than a hundred times the survey study’s thinnest and an order of magnitude above its largest, and the stratum map declares a minimum cell size and rolls up to a coarser geography whenever a cell would fall below it. The repair works here because the cells are thick, not because the groups are of a different kind. Our contribution to this exchange is therefore not another empirical report but the statement of when the repair must work: Theorem 3.11 restores per-cell coverage for any partition fixed in advance, Theorem 3.13 says what that costs and at what rate, and the rollup rule is what keeps a deployment on the right side of it. Label noise has been attacked from several directions: Einbinder, Feldman, Bates, Angelopoulos, Gendler and Romano (2024) study robustness and correction under noisy labels, Sesia, Wang and Tong (2025) model random contamination, Bortolotti, Wang, Tong, Menafoglio, Vantini and Sesia (2025) pursue noise-adaptive classification, Penso, Goldberger and Fetaya (2025) estimate a clean threshold under uniform noise, Xi, Liu, Zeng, Sun and Wei (2025) handle known uniform noise online, and Cohen, Goldberger and Tirer (2025) address noisy regression. Distribution shift is served by weighted conformal prediction (Tibshirani, Foygel Barber, Candès and Ramdas, 2019), the nonexchangeable extensions of Foygel Barber, Candès, Ramdas and Tibshirani (2023), group-weighted constructions A. Rafe and S. Das: Preprint submitted to Elsevier

Page 4 of 45

Certified crash-severity prediction

(Bhattacharyya and Foygel Barber, 2026), and recent work on likelihood-ratio regularization (Joshi, Kiyani, Pappas, Dobriban and Hassani, 2025), optimal transport with unlabeled targets (Correia and Louizos, 2025), and clipped weights (Wang and Goel, 2026). Risk beyond coverage is handled by conformal risk control (Angelopoulos, Bates, Fisch, Lei and Schuster, 2024), whose closest safety-critical application is the wildfire-evacuation mapping of Dayan (2026). Sequential monitoring of exchangeability by test martingales is due to Vovk, Gammerman and Shafer (2022). Each of these supplies an ingredient this paper uses and labels as such. The label-noise literature is the instructive case for what remains missing. Those methods correct or calibrate using a stochastic noise model, a noise rate, or a uniformity premise. Crash severity offers none of those reliably, because the confusion structure varies by agency and era (Burdett et al., 2022), but it does offer something the general case lacks, which is a stable ordinal band: a report of B is wrong by one category far more often than by three, and a report of K is essentially never wrong upward. Assuming a kernel one cannot transfer is a stronger and less defensible premise than declaring a support band one can.

2.4. The gap The three streams do not meet. Severity modeling supplies flexible conditional distributions and the heterogeneity structure, but no finite-sample guarantee. Transportation uncertainty work supplies applications of generic conformal prediction, which inherit generic guarantees. The conformal literature supplies validity, subgroup conditioning, noise robustness, shift correction and risk control as separate results, each proved in a setting that does not know what KABCO is. No existing framework gives contiguous ordinal prediction sets that carry a guarantee on the unobserved true severity under a declared reporting band, that remain valid on the heterogeneity classes the severity literature has spent two decades identifying, that report an honest and estimable certificate when transfer to a new jurisdiction cannot be validated, that control a severity-weighted omission risk, and that compose these guarantees with an accountable slack budget. Assembling that composition, proving what it costs, and shipping it as an artifact an agency can audit is the contribution of this paper.

3. Framework and guarantees Table 1 gathers the notation used below, and Figure 1 outlines the certification workflow. Table 1 Notation. Symbol

Meaning

Defined

Data and outcome  = {1, … , 5} 𝑌 , 𝑌̃ 𝑋∈ tr , cal 𝑛

KABCO severity, O≺C≺B≺A≺K true, police-reported severity covariates training, calibration index sets calibration size; test point is 𝑛+1

§3.1 §3.1 §3.1 §3.1 §3.1

Fitted tier (arbitrary, never trusted) 𝐹̂ (𝑘 ∣ 𝑥) 𝑐(𝑥) ̂ 𝑔(𝑥) ̂ 𝑤(𝑥)

base conditional CDF latent-class assignment, 𝐶 classes stratum map, 𝛾 ∈  density ratio for transfer

(1) (55) (38) (39)

Certification machinery 𝑠(𝑥, 𝑦) 𝐶𝜆 (𝑥) 𝑞, ̂ 𝑞̂𝑐 , 𝑞̂𝑐𝛾 𝛼  (𝑌̃ ) 𝛿 𝐶̃ ⊕ (𝑥) 𝜅(⋅), 𝛽 𝜆̂ 𝑑TV 𝑀𝑡 , 𝜀

ordinal cumulative score threshold set, a contiguous interval marginal, class, product-cell quantile miscoverage level declared compatibility band declared beyond-band mass (an input) band-expanded set severity cost, risk budget conformal risk-control threshold total-variation slack; bounded below only drift test martingale and its parameter

(2) (3) (6), (8) (6) (26) (30) (28) (47) (46) (53) (44)

Every result in this section carries a [NEW], [ADAPTED] or [KNOWN] attribution, and no adapted result is presented as new.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 5 of 45

Certified crash-severity prediction

Figure 1: CHOIR as a graph. Circles are random or fitted objects, the rectangle an exported artifact, shaded nodes observed data; plates denote replication over calibration records (𝑖 ∈ cal ), latent classes (𝑐), declared strata (𝛾) and the deployment ̂ is never trusted: stream (𝑡). The fitted tier, the base CDF 𝐹̂ , the class gate 𝑐, ̂ the stratum map 𝑔 and the density ratio 𝑤, no guarantee depends on it. Guarantees are created by the calibration quantile 𝑞̂𝑐𝛾 , transferred to true severity by the declared band ( , 𝛿), cost-weighted through (𝜅, 𝛽), and exported as a per-stratum certificate with additive, attributable slacks. The dashed loop is detection, not correction: an alarm mandates re-calibration and never edits an issued certificate.

Throughout, each result is labeled [NEW] (new statement and proof), [ADAPTED] (known technique, new setting or packaging; source cited), or [KNOWN] (stated for completeness with citation). No adapted result is presented as new. Proofs of Theorems 3.13, 3.33, and 3.45 are deferred to A; all other proofs appear in line.

3.1. Setup, notation, and standing assumptions Severity categories  = {1, … , 𝐾} with 𝐾 = 5, ordered 1(O) ≺ 2(C) ≺ 3(B) ≺ 4(A) ≺ 5(K) (KABCO: no injury, possible, non-incapacitating, incapacitating, fatal). Covariates 𝑋 ∈ . 𝑌 denotes true severity; 𝑌̃ the police-reported severity. Data are split into a training set tr , on which every fitted object lives (base model, partition, density-ratio model), and a calibration set cal = {1, … , 𝑛}; the test point is indexed 𝑛 + 1. All fitted objects are treated as fixed functions when reasoning about cal ∪ {𝑛 + 1}; this is what sample splitting buys and is used without further comment. Assumption 3.1 (A1: exchangeability). (𝑋1 , 𝑌̃1 ), … , (𝑋𝑛 , 𝑌̃𝑛 ), (𝑋𝑛+1 , 𝑌̃𝑛+1 ) are exchangeable. Remark 3.2 (Clustered crash data). Persons in one crash are dependent, so rows from the same crash violate A1 if they straddle calibration and test. Protocol: (i) all splits are by crash identifier; (ii) the primary analysis calibrates and tests on one randomly sampled driver row per crash, restoring the i.i.d.-across-units structure (crashes i.i.d. ⇒ sampled rows i.i.d. ⇒ A1). Using all rows under crash-level splitting is approximately valid and is reported as a robustness check; for a fully rigorous multi-row treatment see the grouped conformal literature (Dunn et al., 2023).

Base model. Any estimator fit on tr exporting an estimated conditional CDF 𝐹̂ (𝑘 ∣ 𝑥) = 𝑃̂ (𝑌 ≤ 𝑘 ∣ 𝑋 = 𝑥),

𝐹̂ (0 ∣ 𝑥) ≡ 0,

A. Rafe and S. Das: Preprint submitted to Elsevier

𝐹̂ (𝐾 ∣ 𝑥) ≡ 1,

(1) Page 6 of 45

Certified crash-severity prediction

non-decreasing in 𝑘. Ordered logit and probit, random-parameters logit, and latent-class logit export 𝐹̂ natively; any probabilistic classifier exports it by cumulating 𝑝(⋅ ̂ ∣ 𝑥). No assumption is ever made that 𝐹̂ is correct. Every guarantee below is distribution-free in 𝐹̂ ; model quality affects only set size, never validity.

3.2. Ordinal scores, interval sets, marginal validity Definition 3.3 (Ordinal cumulative score). { } 𝑠(𝑥, 𝑦) = max 𝐹̂ (𝑦 − 1 ∣ 𝑥), 1 − 𝐹̂ (𝑦 ∣ 𝑥) .

(2)

𝐹̂ (𝑦 − 1 ∣ 𝑥) is the predicted mass strictly below 𝑦; 1 − 𝐹̂ (𝑦 ∣ 𝑥) the mass strictly above: 𝑠 measures how deep into a tail the candidate label sits (cf. CDF-based ordinal scores, Lu et al., 2022; Chakraborty et al., 2024). Definition 3.4 (Threshold sets). 𝐶𝜆 (𝑥) = {𝑘 ∈  ∶ 𝑠(𝑥, 𝑘) ≤ 𝜆}.

(3)

Lemma 3.5 (Contiguity and non-triviality; [ADAPTED] — proof ours). For every 𝑥 and 𝜆 ≥ 0: (i) 𝐶𝜆 (𝑥) is a contiguous (possibly empty) interval, 𝐶𝜆 (𝑥) = {𝑎𝜆 (𝑥), … , 𝑏𝜆 (𝑥)} ∩ ;

(4)

(ii) if 𝜆 ≥ 1∕2 then 𝐶𝜆 (𝑥) ≠ ∅ and contains the predictive median 𝑘∗ (𝑥) = min{𝑘 ∶ 𝐹̂ (𝑘 ∣ 𝑥) ≥ 1∕2};

(5)

(iii) 𝜆 ↦ 𝐶𝜆 (𝑥) is nested non-decreasing. Proof. (i) 𝐶𝜆 (𝑥) = 𝐿 ∩ 𝑈 with 𝐿 = {𝑘 ∶ 𝐹̂ (𝑘 − 1 ∣ 𝑥) ≤ 𝜆}, 𝑈 = {𝑘 ∶ 𝐹̂ (𝑘 ∣ 𝑥) ≥ 1 − 𝜆}. Since 𝐹̂ (⋅ ∣ 𝑥) is non-decreasing, 𝐿 is a down-set and 𝑈 an up-set; 1 ∈ 𝐿 (as 𝐹̂ (0) = 0) and 𝐾 ∈ 𝑈 (as 𝐹̂ (𝐾) = 1). The intersection of a down-set {1, … , 𝑏𝜆 } and an up-set {𝑎𝜆 , … , 𝐾} is the interval {𝑎𝜆 , … , 𝑏𝜆 }, empty iff 𝑎𝜆 > 𝑏𝜆 . (ii) At 𝑘∗ : 𝐹̂ (𝑘∗ − 1 ∣ 𝑥) < 1∕2 ≤ 𝜆 and 1 − 𝐹̂ (𝑘∗ ∣ 𝑥) ≤ 1∕2 ≤ 𝜆, so 𝑠(𝑥, 𝑘∗ ) ≤ 1∕2 ≤ 𝜆. (iii) Immediate from Definition 3.4. Convention 3.6 (Non-emptiness). If 𝐶𝜆 (𝑥) = ∅, output the singleton {arg min𝑘 𝑠(𝑥, 𝑘)}. This can only enlarge sets, so every lower coverage bound below holds verbatim; we suppress it from the notation. Proposition 3.7 (Marginal validity of split conformal; [KNOWN] — Vovk et al., 2005; Lei et al., 2018). Let 𝑆𝑖 = 𝑠(𝑋𝑖 , 𝑌̃𝑖 ) for 𝑖 ∈ cal , with order statistics 𝑆(1) ≤ ⋯ ≤ 𝑆(𝑛) , and set 𝑞̂ = 𝑆(⌈(1−𝛼)(𝑛+1)⌉) (with 𝑞̂ = +∞, i.e. 𝐶 = , if ⌈(1 − 𝛼)(𝑛 + 1)⌉ > 𝑛). Under A1, ( ) ℙ 𝑌̃𝑛+1 ∈ 𝐶𝑞̂ (𝑋𝑛+1 ) ≥ 1 − 𝛼,

(6)

(7)

1 and if the 𝑛 + 1 scores are a.s. distinct, also ≤ 1 − 𝛼 + 𝑛+1 .

Proof. Exchangeability makes the rank of 𝑆𝑛+1 among {𝑆1 , … , 𝑆𝑛+1 } (sub-)uniform on {1, … , 𝑛 + 1}. The event 𝑌̃𝑛+1 ∈ 𝐶𝑞̂ (𝑋𝑛+1 ) equals {𝑆𝑛+1 ≤ 𝑞}, ̂ which contains the event that this rank is at most ⌈(1 − 𝛼)(𝑛 + 1)⌉, of probability ≥ 1 − 𝛼. The upper bound is the standard no-ties argument. Remark 3.8 (Ties). 𝑠 takes values in the finite set of 𝐹̂ -values, so ties are possible; the lower bound is unaffected, and exactness can be restored by uniform tie-breaking jitter. At 𝑛 ∼ 106 the tie correction is numerically irrelevant. Remark 3.9 (Why intervals matter). Contiguity is not cosmetic. First, “B or worse” is the operationally meaningful output for screening; a set omitting the middle of the scale while holding both ends is a permissible answer that no dispatcher can act on. Second, contiguity is what keeps the noise expansion of Theorem 3.31 additive: an interval ̃ admits only the multiplicative expands to an interval at a cost of 𝑏− + 𝑏+ categories, whereas a general set of size |𝐶| − + ̃ bound |𝐶|(𝑏 + 𝑏 + 1). The guarantee itself does not require it (Remark 3.32); the efficiency of the guarantee does. The ordinal structure of  is what does the mathematical work in Theorem 3.31, by making  (𝑦) ̃ a proper subset of . A. Rafe and S. Das: Preprint submitted to Elsevier

Page 7 of 45

Certified crash-severity prediction

Proposition 3.10 (Impossibility of exact conditional coverage; [KNOWN] — Vovk, 2012; Foygel Barber et al., 2021). Suppose 𝑃𝑋 is non-atomic. Then any procedure achieving ℙ(𝑌 ∈ 𝐶(𝑋) ∣ 𝑋 = 𝑥) ≥ 1 − 𝛼 for a.e. 𝑥 under all distributions returns trivial sets (sets that cannot adapt). The strongest achievable target is therefore coverage conditional on finitely many pre-declared groups, the mathematical justification for Sections 3.3 and 3.8, where the groups are the transportation-relevant ones: heterogeneity classes, jurisdictions, time. The non-atomicity of 𝑃𝑋 is the operative condition and is not a technicality: the obstruction is driven by the covariate law, not by || < ∞. Were 𝑋 finitely supported the statement would fail, and Theorem 3.11 with 𝑐̂ the identity on that support would be the counterexample. Crash covariates include continuous components (speed limit, age, traffic volume), so the condition holds here.

3.3. Heterogeneity-conditional validity: the econometrics bridge Let 𝑐̂ ∶  → {1, … , 𝐶} be any measurable class-assignment function fit on tr ; in CHOIR it is the MAP class of a latent-class model (deep gate or classical latent-class ordered logit). Let 𝑛𝑐 = #{𝑖 ≤ 𝑛 ∶ 𝑐(𝑋 ̂ 𝑖 ) = 𝑐} and let 𝑞̂𝑐 be the class-wise analogue of (6), (𝑐) 𝑞̂𝑐 = 𝑆(⌈(1−𝛼)(𝑛 , +1)⌉) 𝑐

(𝑐) (𝑐) 𝑆(1) ≤ ⋯ ≤ 𝑆(𝑛 the ordered scores {𝑆𝑖 ∶ 𝑐(𝑋 ̂ 𝑖 ) = 𝑐}, )

(8)

𝑐

and define the class-conditional predictor 𝐶 het (𝑥) = 𝐶𝑞̂𝑐(𝑥) (𝑥). ̂

(9)

Theorem 3.11 (Latent-class-conditional coverage; [ADAPTED] — Mondrian argument, Vovk et al., 2005; framing new). Under A1, for every class 𝑐 ∈ {1, … , 𝐶}, ( ) | ℙ 𝑌̃𝑛+1 ∈ 𝐶 het (𝑋𝑛+1 ) | 𝑐(𝑋 ̂ 𝑛+1 ) = 𝑐 ≥ 1 − 𝛼, |

(10)

and 𝐶 het (𝑥) is a contiguous interval for every 𝑥. Proof. Fix 𝑐; condition on 𝐸𝑐 = {𝑐(𝑋 ̂ 𝑛+1 ) = 𝑐} and on the index set 𝐼𝑐 ⊆ {1, … , 𝑛} of calibration points in class 𝑐. Because 𝑐̂ is a fixed measurable function, membership in class 𝑐 is a deterministic function of the data point; exchangeability of the full collection implies exchangeability of {(𝑋𝑖 , 𝑌̃𝑖 ) ∶ 𝑖 ∈ 𝐼𝑐 } ∪ {(𝑋𝑛+1 , 𝑌̃𝑛+1 )} conditionally on 𝐸𝑐 and |𝐼𝑐 | = 𝑛𝑐 : any permutation 𝜋 of the class-𝑐 points extends to a permutation of all 𝑛 + 1 points fixing the complement, and the conditioning event is permutation-symmetric within the class. Apply Proposition 3.7’s rank argument to these 𝑛𝑐 + 1 exchangeable scores with quantile index ⌈(1 − 𝛼)(𝑛𝑐 + 1)⌉; average over 𝑛𝑐 . Contiguity is Lemma 3.5 at 𝜆 = 𝑞̂𝑐 . Remark 3.12 (Arbitrariness is harmless). Theorem 3.11 requires nothing about 𝑐̂ being “correct”: even a nonsensical partition yields per-cell validity. Partition quality affects efficiency (set size), quantified next. The latent-class machinery of the random-parameters tradition cannot break validity; it can only earn efficiency. The guarantee is per cell and marginal: each cell covers at 1 − 𝛼 over the calibration draw, not simultaneously across cells at a family-wise level, so a deployment auditing many cells at once should read the coverages as marginal and budget accordingly. Theorem 3.13 (Oracle efficiency among conditionally valid predictors; [NEW]). Define the class-conditional score CDFs, their (1 − 𝛼)-quantiles, and the class proportions ( ) | 𝐺𝑐 (𝑡) = ℙ 𝑠(𝑋, 𝑌̃ ) ≤ 𝑡 | 𝑐(𝑋) ̂ =𝑐 , |

𝑞𝑐 = inf {𝑡 ∶ 𝐺𝑐 (𝑡) ≥ 1 − 𝛼},

𝑝𝑐 = ℙ(𝑐(𝑋) ̂ = 𝑐) > 0.

(11)

Assume: (R1) each 𝐺𝑐 is continuous and strictly increasing in a neighborhood of 𝑞𝑐 , and the mixture and strictly increasing in a neighborhood of its (1 − 𝛼)-quantile 𝑞mix ;

∑

𝑐 𝑝𝑐 𝐺𝑐 is continuous

(R2) for every 𝑘 ∈  and 𝑐: ℙ(𝑠(𝑋, 𝑘) = 𝑞𝑐 ∣ 𝑐(𝑋) ̂ = 𝑐) = 0. Consider the family  of predictors 𝐶𝝀 (𝑥) = 𝐶𝜆𝑐(𝑥) (𝑥) indexed by class-wise thresholds 𝝀 = (𝜆1 , … , 𝜆𝐶 ). Then: ̂ A. Rafe and S. Das: Preprint submitted to Elsevier

Page 8 of 45

Certified crash-severity prediction

(i) (Oracle characterization.) Within  , class-𝑐 coverage ≥ 1 − 𝛼 holds iff 𝜆𝑐 ≥ 𝑞𝑐 ; the unique expected-sizeminimal member of  valid in every class is 𝝀⋆ = (𝑞1 , … , 𝑞𝐶 ), with minimal size ∑ [ ] | ⋆ = 𝑝𝑐 𝔼 |𝐶𝑞𝑐 (𝑋)| | 𝑐(𝑋) ̂ =𝑐 . (12) | 𝑐

𝑝

(ii) (Attainment.) Under i.i.d. sampling within each class, as min𝑐 𝑛𝑐 → ∞: 𝑞̂𝑐 ←←→ ← 𝑞𝑐 for every 𝑐, with |𝑞̂𝑐 − 𝑞𝑐 | = −1∕2 𝑂𝑝 (𝑛𝑐 ) under (R1) with positive density at 𝑞𝑐 (van der Vaart, 1998, Cor. 21.5); and the expected size of 𝐶 het converges to  ⋆ . (iii) (Diagnosis of marginal conformal.) Marginal split conformal has class-𝑐 coverage converging to 𝐺𝑐 (𝑞mix ), which is < 1−𝛼 for every class with 𝑞𝑐 > 𝑞mix , and > 1−𝛼 for every class with 𝑞𝑐 < 𝑞mix . It is class-conditionally valid for all 𝑐 iff 𝑞1 = ⋯ = 𝑞𝐶 . Remark 3.14 (What (iii) does and does not identify). Part (iii) is elementary, and its content is that a mixture quantile is a mixture: classes whose score distributions sit above the pooled quantile must pay for those below it. It predicts a direction and nothing more. It supplies no magnitude, and in particular it does not identify which classes are the hard ones, since whether 𝑞𝑐 > 𝑞mix for a given subpopulation is an empirical property of the data and not a consequence of this theorem. That the hard classes in Texas crash records turn out to be unrestrained drivers, motorcyclist-involved crashes and rural high-speed records is an empirical finding, reported in Section 5 and not claimed here. The theorem’s value is that it makes the failure structural rather than incidental: it is not a defect of any particular base model and cannot be fitted away.

Honesty box (scope of Theorem 3.13). Theorem 3.13 does not claim class-conditional (Mondrian) sets are smaller on average than marginal sets; in general they are larger, because marginal conformal silently borrows coverage from hard classes to shrink easy ones. The correct claim, proved in A, is that the class-conditional predictor is the smallest one that does not do that. For safety applications, undercoverage of severe strata is precisely the failure mode that matters; the empirical section reports average sizes honestly. A second honest reading concerns (R1)–(R2): a discrete base model exports finitely many 𝐹̂ -values, so the score distributions are atomic and both conditions can fail exactly; the tie-breaking jitter of Remark 3.8 makes the scores continuous and restores them.

3.4. When certification is informative: a floor no model reaches below Theorem 3.13 compares members of the threshold family  at a fixed 𝐹̂ . It is therefore silent on the question a practitioner asks on being handed a certificate reading “somewhere between no injury and fatal”: would a better base model have done better? This subsection answers that question in its negative form, by exhibiting a floor on expected set size that no predictor breaches and in which no base model appears. Throughout, 𝑍 denotes the label being certified. Everything below is stated for 𝑍 = 𝑌̃ , the reported label on which the sets of Sections 3.2–3.3 are calibrated, and holds verbatim for 𝑍 = 𝑌 . Let 𝑝(𝑧 ∣ 𝑥) = ℙ(𝑍 = 𝑧 ∣ 𝑋 = 𝑥)

(13)

be the true conditional mass function, which is nowhere assumed known and nowhere estimated. Fix a stratum 𝐴 ∈ 𝜎(𝑋) with ℙ(𝑋 ∈ 𝐴) > 0. Definition 3.15 (Level sets and the level-set functional). For 𝑡 ≥ 0, 𝐿𝑡 (𝑥) = {𝑧 ∈  ∶ 𝑝(𝑧 ∣ 𝑥) > 𝑡}, 𝑁(𝑥, 𝑡) = |𝐿𝑡 (𝑥)|, (14) ( ) | and write 𝛾𝐴 (𝑡) = ℙ 𝑝(𝑍 ∣ 𝑋) > 𝑡 | 𝑋 ∈ 𝐴 for the coverage of the level-set rule at 𝑡, non-increasing with 𝛾𝐴 (0) = 1. | Define the floor as the value of the size-minimization problem itself, { [ } ] ( ) | | 𝐴 (𝛼) = inf 𝔼 |𝐶(𝑋)| | 𝑋 ∈ 𝐴 ∶ 𝐶 𝜎(𝑋)-measurable, ℙ 𝑍 ∈ 𝐶(𝑋) | 𝑋 ∈ 𝐴 ≥ 1 − 𝛼 , (15) | | which exists for every 𝛼 ∈ (0, 1) because 𝐶 ≡  is feasible. Theorem 3.17 identifies it with a level set: writing 𝑡𝐴 (𝛼) = sup{𝑡 ≥ 0 ∶ 𝛾𝐴 (𝑡) ≥ 1 − 𝛼}, [ ] ( ) | | 𝐴 (𝛼) = 𝔼 𝑁(𝑋, 𝑡𝐴 (𝛼)) | 𝑋 ∈ 𝐴 whenever ℙ 𝑝(𝑍 ∣ 𝑋) = 𝑡𝐴 (𝛼) | 𝑋 ∈ 𝐴 = 0. (16) | | A. Rafe and S. Das: Preprint submitted to Elsevier

Page 9 of 45

Certified crash-severity prediction

𝛼 ↦ 𝐴 (𝛼) is non-increasing. Remark 3.16 (Why the floor is defined as a value and not as a level set). Equation (16) can fail at an atom, and the degenerate case is the instructive one: if 𝑍 = 𝑓 (𝑋) for a measurable 𝑓 , then 𝛾𝐴 (𝑡) = 1 for every 𝑡 < 1, so 𝑡𝐴 (𝛼) = 1 and 𝑁(𝑥, 1) = |{𝑧 ∶ 𝑝(𝑧 ∣ 𝑥) > 1}| = 0, whereas the floor of (15) is 1−𝛼, attained by a singleton on a (1−𝛼)-mass covariate set and the empty set elsewhere (1 once the non-emptiness convention of Convention 3.6 is imposed). Defining 𝐴 (𝛼) by (15) keeps it positive at such laws, which matters here because those are exactly the laws Theorem 3.25 builds from. 𝐴 (𝛼) is a functional of the joint law of (𝑋, 𝑍) restricted to 𝐴 and of nothing else. It does not depend on 𝐹̂ , on the calibration split, or on any modeling choice. Theorem 3.17 (Level-set floor on expected set size; [ADAPTED] — oracle level-set optimality, Sadinle, Lei and Wasserman, 2019; ordinal and vacuity corollaries [NEW]). Let 𝐶 ∶  → 2 be any 𝜎(𝑋)-measurable set-valued predictor satisfying ( ) | ℙ 𝑍 ∈ 𝐶(𝑋) | 𝑋 ∈ 𝐴 ≥ 1 − 𝛼. (17) | Then, by (15), ] [ | 𝔼 |𝐶(𝑋)| | 𝑋 ∈ 𝐴 ≥ 𝐴 (𝛼), |

(18)

and the floor is a level set: the infimum in (15) is attained by 𝐶(𝑥) = 𝐿𝑡𝐴 (𝛼) (𝑥), with value 𝔼[𝑁(𝑋, 𝑡𝐴 (𝛼)) ∣ 𝑋 ∈ 𝐴] ( ) | as in (16), whenever ℙ 𝑝(𝑍 ∣ 𝑋) = 𝑡𝐴 (𝛼) | 𝑋 ∈ 𝐴 = 0, and attained after boundary randomization otherwise. The | substance is the second statement; (18) restates (15). Proof. Relax set membership to 𝑢 ∶  ×  → [0, 1] and consider [∑ ] [∑ ] | | min 𝔼 𝑢(𝑋, 𝑧) | 𝑋 ∈ 𝐴 s.t. 𝔼 𝑝(𝑧 ∣ 𝑋) 𝑢(𝑋, 𝑧) | 𝑋 ∈ 𝐴 ≥ 1 − 𝛼. | | 𝑢 𝑧∈

(19)

𝑧∈

Any 𝐶 obeying (17) yields the feasible point 𝑢(𝑥, 𝑧) = 𝟏{𝑧 ∈ 𝐶(𝑥)} with objective 𝔼[|𝐶(𝑋)| ∣ 𝑋 ∈ 𝐴], so the value of (19) lower-bounds 𝐴 (𝛼). For 𝜏 ≥ 0 the Lagrangian is [∑ ] ( )| 𝔼 𝑢(𝑋, 𝑧) 1 − 𝜏 𝑝(𝑧 ∣ 𝑋) | 𝑋 ∈ 𝐴 + 𝜏(1 − 𝛼), (20) | 𝑧

which decouples across (𝑥, 𝑧): the coefficient of 𝑢(𝑥, 𝑧) is negative exactly when 𝑝(𝑧 ∣ 𝑥) > 1∕𝜏, so a pointwise minimizer sets 𝑢 = 1 on 𝐿1∕𝜏 (𝑥) and 𝑢 = 0 off its closure. Choosing 𝜏 with 1∕𝜏 = 𝑡𝐴 (𝛼) makes the constraint bind, after boundary randomization when 𝑝(𝑍 ∣ 𝑋) has an atom at 𝑡𝐴 (𝛼), so complementary slackness holds and this level set solves (19). Away from an atom the solution is integral, hence of the form 𝑢 = 𝟏{𝑧 ∈ 𝐿𝑡𝐴 (𝛼) (𝑥)} and feasible for (15); so the relaxed value and 𝐴 (𝛼) coincide and both equal 𝔼[𝑁(𝑋, 𝑡𝐴 (𝛼)) ∣ 𝑋 ∈ 𝐴]. Remark 3.18 (Variational form). Equation (18) is often written as a budget allocation: spend miscoverage 𝑢(𝑥) ≥ 0 at 𝑥 subject to 𝔼[𝑢(𝑋) ∣ 𝑋 ∈ 𝐴] ≤ 𝛼, and pay the smallest set capturing conditional mass 1 − 𝑢(𝑥). Theorem 3.17 is that infimum solved. Its content is that the optimal budget is not free to vary arbitrarily with 𝑥: it is the one induced by a single threshold on 𝑝(⋅ ∣ 𝑥), which is the Neyman–Pearson structure and is what makes 𝐴 (𝛼) computable in principle from the law alone. Corollary 3.19 (The floor binds interval predictors). Every contiguous-interval predictor is in particular a set-valued predictor, so (18) applies verbatim to 𝐶 het , to 𝐶̃ ⊕ , and to every composed set of Theorem 3.50. Ordinal structure cannot evade the floor; it can only fail to reach it. Corollary 3.20 (Vacuity is a property of the data, not of the model). Fix 𝜀 ∈ (0, 1). If 𝐴 (𝛼) ≥ 𝐾 − 𝜀 then every predictor valid at level 1 − 𝛼 on 𝐴 in the sense of (17) has mean set size at least 𝐾 − 𝜀 on 𝐴. Anticipating the banded map N(𝑏− , 𝑏+ , 𝛿) of Assumption 3.29 below, the composed statement is sharper still: a base interval of width at least 𝐾 − 𝑏− − 𝑏+ satisfies 𝐶̃ ⊕ (𝑥) =  after expansion, so whenever 𝑁(𝑥, 𝑡𝐴 (𝛼)) ≥ 𝐾 − 𝑏− − 𝑏+ on all but an 𝜀-fraction of 𝐴, every valid composed predictor is vacuous off that fraction. No choice of base model, no volume of training data, and no calibration scheme alters this. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 10 of 45

Certified crash-severity prediction

Remark 3.21 (What this licenses, and the direction of the logic). Corollary 3.20 is a conditional statement whose condition is an unknown functional of the joint law. Its content is that it converts the question “is our base model inadequate?” into the question “is 𝐴 (𝛼) large?”, which is a question about Texas crash records rather than about CHOIR. Honesty requires stating the converse too, because it limits what any experiment can show: Theorem 3.17 equally implies that the observed size of any valid predictor is an upper bound on 𝐴 (𝛼), so no agreement across any library of base models can establish that the floor is high. A plug-in estimate of 𝐴 (𝛼) is evidence about the data regime and is reported as such in Section 5; it is not, and cannot be turned into, a proof. Nonparametric estimation of 𝑝(⋅ ∣ 𝑥) would settle the matter and is unavailable here for a concrete reason: the recorded covariates are close to unique per record, so exactly replicated covariate patterns, on which the empirical label frequency would estimate 𝑝(⋅ ∣ 𝑥) with no model, essentially do not occur. Coarsening the covariates does not help, since a coarser 𝜎-field raises the floor and so bounds it from the wrong side. Finally, the estimand is defined relative to 𝜎(𝑋) for the recorded fields 𝑋: a field the police form does not carry cannot lower a floor it does not enter. Remark 3.22 (Relation to Theorem 3.13). The two results quantify over different objects and neither implies the other. Theorem 3.13 ranges over the threshold family  at a fixed 𝐹̂ and identifies the smallest class-conditionally valid member, a statement about calibration given a model. Theorem 3.17 ranges over all 𝜎(𝑋)-measurable predictors and identifies a floor none reaches below, a statement about the data. Read together they bracket the achievable, 𝐴 (𝛼) ≤ 𝐴⋆ , with the gap attributable to 𝐹̂ and the floor attributable to 𝑃 . Corollary 3.23 (Informativeness frontier). For a declared informativeness target 𝑚 ∈ {1, … , 𝐾 − 1} define 𝛼𝐴⋆ (𝑚) = inf {𝛼 ∈ (0, 1) ∶ 𝐴 (𝛼) ≤ 𝑚}.

(21)

For 𝛼 < 𝛼𝐴⋆ (𝑚) no valid predictor on 𝐴 attains mean size at most 𝑚. Since 𝛼 ↦ 𝐴 (𝛼) is non-increasing, 𝛼𝐴⋆ is well defined and non-increasing in 𝑚. The map 𝑚 ↦ 𝛼𝐴⋆ (𝑚) is the object an agency needs in order to choose a screening level it can actually be served at, and it is what the empirical frontier of Section 5 estimates.

3.5. The floor cannot be certified from below Theorem 3.17 localizes informativeness in a single functional of the law, and Remark 3.21 observed that valid predictors bound 𝐴 (𝛼) only from above. It is natural to ask whether some other procedure could bound it from below, and so certify that a stratum is beyond certification. It cannot, and the obstruction is the same non-atomicity that drives Proposition 3.10. Let  denote the set of joint laws of (𝑋, 𝑍) on  ×  whose covariate marginal is non-atomic, and let 𝐷𝑛 = ((𝑋1 , 𝑍1 ), … , (𝑋𝑛 , 𝑍𝑛 )) be i.i.d. from 𝑃 ∈ . Definition 3.24 (Lower certificate). A (1−𝛾) lower certificate for the floor is a measurable map 𝐿̂ 𝑛 ∶ ( ×)𝑛 → [0, 𝐾] with ( ) ℙ𝑃 𝐿̂ 𝑛 (𝐷𝑛 ) ≤ 𝐴 (𝛼; 𝑃 ) ≥ 1 − 𝛾 for every 𝑃 ∈ . (22) It is distribution-free: no 𝑃 in the class may be excluded, which is precisely the standard the rest of this paper holds itself to. Theorem 3.25 (Non-certifiability of the floor; [ADAPTED] — distribution-free impossibility via indistinguishability, Foygel Barber, 2020; the level-set target and the ordinal reading [NEW]). Fix 𝐾 ≥ 2, 𝛼 ∈ (0, 1), 𝑛 ∈ ℕ, 𝛾 ∈ (0, 1) and 𝜀 > 0. Let 𝐿̂ 𝑛 be any (1 − 𝛾) lower certificate in the sense of Definition 3.24, and let 𝑃0 ∈  be the law with 𝑍 ⟂ 𝑋 and 𝑍 uniform on , for which 𝐴 (𝛼; 𝑃0 ) = (1 − 𝛼)𝐾. Then ( ) ℙ𝑃0 𝐿̂ 𝑛 (𝐷𝑛 ) ≤ 1 ≥ 1 − 𝛾 − 𝜀. (23) That is, on a law whose floor is (1 − 𝛼)𝐾, every distribution-free lower certificate returns the trivial bound with probability at least 1 − 𝛾 − 𝜀. No non-trivial lower bound on 𝐴 (𝛼) is certifiable from finite data without assumptions on 𝑃 . ⨆ Proof. Because 𝑃0,𝑋 is non-atomic, for any 𝑚 there is a partition  = 𝑚 𝑗=1 𝐵𝑗 with 𝑃0,𝑋 (𝐵𝑗 ) = 1∕𝑚 for every 𝑗. For 𝑣 = (𝑣1 , … , 𝑣𝑚 ) ∈  𝑚 let 𝑓𝑣 (𝑥) = 𝑣𝑗 for 𝑥 ∈ 𝐵𝑗 , a measurable function, and let 𝑄𝑣 ∈  be the law with the A. Rafe and S. Das: Preprint submitted to Elsevier

Page 11 of 45

Certified crash-severity prediction

same covariate marginal 𝑃0,𝑋 and 𝑍 = 𝑓𝑣 (𝑋) almost surely. Under 𝑄𝑣 the singleton predictor 𝐶(𝑥) = {𝑓𝑣 (𝑥)} has conditional coverage 1 ≥ 1 − 𝛼, so 𝐴 (𝛼; 𝑄𝑣 ) ≤ 1 by (15); note this is where Remark 3.16 is needed, since (16) degenerates at exactly these laws. Applying (22) to each 𝑄𝑣 and averaging over 𝑉 uniform on  𝑚 , [ ( )] 𝔼𝑉 ℙ𝑄𝑉 𝐿̂ 𝑛 (𝐷𝑛 ) ≤ 1 ≥ 1 − 𝛾. (24) Let 𝑄̄ 𝑚 be the 𝑉 -mixture of 𝑄⊗𝑛 , i.e. the law of 𝐷𝑛 when 𝑉 is drawn once and 𝑛 points are then drawn from 𝑄𝑉 . 𝑉 Let 𝐸 be the event that the 𝑛 sampled covariates fall in 𝑛 distinct cells. On 𝐸 the labels 𝑍𝑖 = 𝑉𝑗(𝑖) involve 𝑛 distinct coordinates of 𝑉 , which are i.i.d. uniform on  and independent of the covariates; the covariates themselves have marginal 𝑃0,𝑋 under both laws. Hence 𝑄̄ 𝑚 and 𝑃0⊗𝑛 agree on 𝐸, and by a birthday bound ℙ(𝐸 𝑐 ) ≤ 𝑛2 ∕(2𝑚) under either, so ( ) 𝑛2 . 𝑑TV 𝑄̄ 𝑚 , 𝑃0⊗𝑛 ≤ 2𝑚

(25)

Choosing 𝑚 ≥ 𝑛2 ∕(2𝜀) and combining (24) with (25) gives (23). Finally 𝐴 (𝛼; 𝑃0 ) = (1 − 𝛼)𝐾: under 𝑃0 the conditional law is uniform for every 𝑥, so covering reported-label mass 1 − 𝛼 costs expected size (1 − 𝛼)𝐾, attained by including a (1 − 𝛼) fraction of the 𝐾 labels and unimprovable by any rule reaching mass 1 − 𝛼. Either way 𝐴 (𝛼; 𝑃0 ) ≫ 1. Remark 3.26 (Attribution). The mechanism of Theorem 3.25 is not ours. Foygel Barber (2020) proves that any distribution-free confidence interval for the conditional label probability in binary regression obeys a length lower bound that is independent of the sample size and holds for every distribution without point masses, so that no distribution-free procedure can adapt to structure in the law. That is the same obstruction, under the same non-atomicity condition, established by the same style of indistinguishability argument, and the honest description of Theorem 3.25 is that it transports that argument to a different target. What is ours is the target and its consequence: the estimand here is the level-set functional 𝐴 (𝛼), a property of the whole conditional law on a stratum rather than of a single conditional probability, and the consequence is that the floor of Theorem 3.17 is not certifiable from below, which is what closes the account opened in Remark 3.21. A reader who knows Foygel Barber (2020) and Foygel Barber et al. (2021) should find Theorem 3.25 unsurprising; we state and prove it because the paper needs the specific statement, not because the technique is new. Remark 3.27 (What is and is not being claimed). The theorem does not say the floor is unknowable in principle; it says it is not certifiable distribution-free. The mechanism is exact and worth stating in words: when 𝑃𝑋 is non-atomic the sample never revisits a covariate value, so a diffuse conditional law and a deterministic labeling 𝑍 = 𝑓 (𝑋) generate identical data, and any 𝑓 agreeing with the observed labels is consistent with everything seen. Assumptions that tie 𝑝(⋅ ∣ 𝑥) across nearby 𝑥, such as smoothness or a parametric form, evade the construction and permit estimation; they are exactly the assumptions this paper declines to make elsewhere, and taking them here to rescue a lower bound would be inconsistent. The result is the natural companion of Proposition 3.10: non-atomicity of the covariate law obstructs conditional coverage, and it obstructs certifying the floor. Together with Theorem 3.17 the statement is complete. There is a floor, it governs informativeness, valid predictors bound it from above, and nobody, ourselves included, can bound it from below without leaving the distribution-free setting. Remark 3.28 (The empirical face of the obstruction). The construction is not a pathology invented for the proof; it describes the data. Estimating 𝑝(⋅ ∣ 𝑥) without model assumptions requires covariate patterns that repeat, and in the Texas records they do not: on the motorcyclist stratum, 54,872 records carry 54,868 distinct patterns across the 38 recorded fields, with four patterns occurring twice and none more often (Section 5). The sample never revisits an 𝑥, which is the empirical statement of the hypothesis Theorem 3.25 runs on. Coarsening the covariates restores repeats but changes the estimand, raising the floor of a smaller 𝜎-field, and so bounds from the wrong side.

3.6. Coverage on true severity under reporting noise Calibration uses reported labels 𝑌̃ ; the guarantee owed to the practitioner is on true severity 𝑌 . We never estimate a misclassification model; we declare a support (band) condition and prove what it buys.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 12 of 45

Certified crash-severity prediction

Assumption 3.29 (N( , 𝛿): compatibility). There is a set-valued map  ∶  → 2 (“plausible truths for a report”) satisfying 𝑦̃ ∈  (𝑦) ̃ for every 𝑦̃ ∈ , and 𝛿 ∈ [0, 1) with ( ) ℙ 𝑌𝑛+1 ∈  (𝑌̃𝑛+1 ) ≥ 1 − 𝛿. (26) Reflexivity is a normalization, not a restriction: a report is always compatible with itself, and every map instantiated in this paper satisfies it. It is stated because it is what gives 𝐶̃ ⊆ 𝐶̃ ⊕ , which Corollary 3.48 uses. Banded special case N(𝑏− , 𝑏+ , 𝛿):  (𝑦) ̃ = {𝑦 ∈  ∶ 𝑦̃ − 𝑏+ ≤ 𝑦 ≤ 𝑦̃ + 𝑏− }.

(27)

Here 𝑏+ guards over-reporting (true severity below the report; the set reaches downward) and 𝑏− guards under-reporting (true severity above the report; the set reaches upward). Under-reporting is the safety-critical direction, so 𝑏− is the load-bearing constant. KABCO default: 𝑏− = 𝑏+ = 1 with 𝛿 the beyond-adjacent mass, elicited from KABCO–MAIS linkage studies and swept in sensitivity curves (Section 3.6.1); refined category-dependent maps in Remark 3.38. The assumption concerns only the test point’s joint law (𝑌𝑛+1 , 𝑌̃𝑛+1 ); calibration labels are never de-noised because calibration only ever touches 𝑌̃ . ̃ ̃ Definition 3.30 (Compatibility expansion). For an interval 𝐶(𝑥) = [𝑎(𝑥), ̃ 𝑏(𝑥)] produced by any procedure of Sections 3.2–3.3, let ⋃  (𝑦). ̃ (28) 𝐶̃ ⊕ (𝑥) = ̃ 𝑦∈ ̃ 𝐶(𝑥)

In the banded case ] [ ̃ 𝐶̃ ⊕ (𝑥) = 𝑎(𝑥) ̃ − 𝑏+ , 𝑏(𝑥) + 𝑏− ∩ ,

(29)

again an interval. Theorem 3.31 (Noise transfer of coverage; [ADAPTED] — coverage transfer under label noise, Cauchois, Gupta, Ali and Duchi, 2022; Stutz, Roy, Matejovicova, Strachan, Cemgil and Doucet, 2023; the banded ordinal map and additive ̃ 𝑛+1 )) ≥ 1 − 𝛼 (Proposition 3.7), or ≥ 1 − 𝛼 expansion [NEW]). Let 𝐶̃ be any set-valued predictor with ℙ(𝑌̃𝑛+1 ∈ 𝐶(𝑋 conditionally on a class or stratum (Theorem 3.11) with N( , 𝛿) holding conditionally. Then under N( , 𝛿), ( ) ℙ 𝑌𝑛+1 ∈ 𝐶̃ ⊕ (𝑋𝑛+1 ) ≥ 1 − 𝛼 − 𝛿. (30) ̃ 𝑛+1 )} ∩ {𝑌𝑛+1 ∈  (𝑌̃𝑛+1 )} we have 𝑌𝑛+1 ∈  (𝑦) ̃ 𝑛+1 ), hence Proof. On 𝐴 = {𝑌̃𝑛+1 ∈ 𝐶(𝑋 ̃ for some 𝑦̃ ∈ 𝐶(𝑋 ⊕ ̃ ̃ ̃ ̃ 𝑌𝑛+1 ∈ 𝐶 (𝑋𝑛+1 ). And ℙ(𝐴) ≥ ℙ(𝑌𝑛+1 ∈ 𝐶) − ℙ(𝑌𝑛+1 ∉  (𝑌𝑛+1 )) ≥ 1 − 𝛼 − 𝛿. The conditional version is identical with all probabilities conditional. The proof is two lines by design: the difficulty was moved into having a banded compatibility map, which is transportation structure. Without ordinal structure,  (𝑦) ̃ is all of  and the expansion is vacuous. Transferring conformal coverage from a noisy or weakly supervised label to the clean one is not itself new (Cauchois et al., 2022; Stutz et al., 2023; Einbinder et al., 2024); what the ordinal band contributes is that the transfer costs an additive 𝑏− + 𝑏+ categories rather than degrading multiplicatively or by the full error rate, a distinction made precise against the total-variation baseline in Proposition 3.34 and Remark 3.35. ̃ its proof never Remark 3.32 (What contiguity does and does not buy). Theorem 3.31 holds for any set-valued 𝐶; ⋃ ̃ uses contiguity, and 𝐶̃ ⊕ = 𝑦∈  ( 𝑦) ̃ is defined for arbitrary 𝐶. What ordinality supplies is that  (𝑦) ̃ is a proper ̃ 𝐶̃ subset of , without which the expansion is vacuous; that is the load-bearing structure. What contiguity of 𝐶̃ supplies is separate and is about cost, not validity: for an interval, 𝐶̃ ⊕ is again an interval and the expansion costs 𝑏− + 𝑏+ ̃ − + 𝑏+ + 1), i.e. multiplicatively. Contiguity also additively, whereas for a general 𝐶̃ the bound degrades to |𝐶|(𝑏 carries the operational reading (“𝐵 or worse”) that a screening deployment acts on. Both are reasons to want interval sets; neither is a condition of Theorem 3.31. Theorem 3.33 (Sharpness; [NEW]). Fix the banded map  (𝑦) ̃ = [𝑦̃ − 𝑏+ , 𝑦̃ + 𝑏− ]. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 13 of 45

Certified crash-severity prediction

(i) (Tightness of the constant.) For every 𝛼 ∈ (0, 1), 𝛿 ∈ [0, 𝛼 ∧ (1 − 𝛼)), 𝑏± with 𝐾 ≥ 𝑏− + 2, and every 𝜂 > 0, there exists a joint law of (𝑋, 𝑌 , 𝑌̃ ) satisfying A1 and N(𝑏− , 𝑏+ , 𝛿) under which ( ) lim ℙ 𝑌𝑛+1 ∈ 𝐶̃ ⊕ ≤ 1 − 𝛼 − 𝛿 + 𝜂; (31) 𝑛→∞

hence no constant larger than 1 − 𝛼 − 𝛿 can replace the bound of Theorem 3.31. (ii) (The expansion cannot be trimmed on either side.) Fix 𝛼 ∈ (0, 1∕2) and assume 𝑏− ≥ 1 (respectively 𝑏+ ≥ 1 for the symmetric statement), so the trimmed rule is well defined. For the rule that expands upward only by 𝑏− − 1 (i.e., outputs [𝑎̃ − 𝑏+ , 𝑏̃ + 𝑏− − 1]) there is a law satisfying N(𝑏− , 𝑏+ , 0) with asymptotic true-label coverage ≤ 𝛼 < 1 − 𝛼; symmetrically, for the rule that expands downward only by 𝑏+ − 1 there is a law satisfying N(𝑏− , 𝑏+ , 0) with the same failure. Hence no one-sided reduction of the expansion preserves the guarantee. Proposition 3.34 (Structure buys tightness; [ADAPTED] — (i) cf. Einbinder et al., 2024, Cor. 3). Suppose only ℙ(𝑌̃ ≠ 𝑌 ) ≤ 𝜀tot is known (no band). Then: (i) the best transfer without expansion is ̃ ≥ 1 − 𝛼 − 𝜀tot , ℙ(𝑌 ∈ 𝐶)

(32)

⋃ and no label-space expansion whose output 𝑦∈ ̃ remains a proper subset of  improves the constant, since ̃ 𝐶̃ 𝐷(𝑦) without support restrictions an adversarial coupling puts the off-event mass outside that union; (ii) under the banded assumption, Theorem 3.31 attains 1 − 𝛼 − 𝛿 at the cost of at most 𝑏− + 𝑏+ added categories. In the KABCO regime elicited from linkage studies (𝜀tot of order 0.3–0.5, dominated by adjacent A/B/C confusion; beyond-band mass 𝛿 of order 0.01–0.05), the guarantee improves from 1 − 𝛼 − 𝜀tot (weak, and at the upper end near-uninformative at 𝛼 = 0.1) to 1 − 𝛼 − 𝛿 (near-nominal). Remark 3.35 (Attribution of Proposition 3.34). Part (i) is not new and is not claimed as new. Einbinder et al. (2024) obtain the same degradation from a purely structural premise, assuming only a total-variation bound between the clean and noisy label laws and correcting the nominal level accordingly; up to the constant, their Corollary 3 is (32). We state it here because it is the baseline Theorem 3.31 must beat, and the comparison in (ii) is the point: their premise and ours are both structural, but theirs is indexed by the total error rate 𝜀tot and ours by the beyond-band mass 𝛿. In the KABCO regime those differ by an order of magnitude, which is what the band buys. The adversarial-coupling clause in (i) is our addition and is a one-line observation, not a contribution. Proof. (i) is the union bound of Theorem 3.31 with  (𝑦) ̃ = {𝑦} ̃ on a 1−𝜀tot event, plus the stated adversarial coupling. (ii) is Theorem 3.31 plus arithmetic; the KABCO magnitudes are cited (Section 3.6.1), not proved.

3.6.1. Numeric grounding of the band KABCO–MAIS linkage studies quantify police misclassification against medically assessed severity. In the Wisconsin CODES linkage (2008–2012; 120,534 linked crash victims), overall agreement between the police KABCO rating and the MAIS-based reference category is 51%, so 𝜀tot ≈ 0.49 in the linked sample; essentially all disagreement is concentrated within one adjacent reference category in the under-reporting direction: victims rated B, C, or O whose reference severity is serious (MAIS 3+) constitute 2.9% of all cases, and per report category the beyond-adjacent underreporting mass is 0.4% (O), 1.7% (C), and 2.0% (B) (Burdett et al., 2015; Burdett, 2014). Fatality recording is exact in the over-reporting direction in both the Wisconsin linkage (100% of K-coded victims MAIS 6) and NASS/CDS (Farmer, 2003). Over-reporting from A is heavier and reaches two categories down (38.2% of A-coded victims at MAIS 1), which motivates the category-dependent map of Remark 3.38 rather than a constant symmetric band. Two caveats are declared: linked samples over-represent medically attended victims, and MAIS is a distinct scale mapped to KABCO by a declared correspondence; both are reasons the band is a sensitivity input swept over 𝛿 (and over candidate maps  ), never an estimated parameter. Remark 3.36 (Is the band minimal, or is the saturation self-inflicted?). A band of total width 𝑏− + 𝑏+ = 2 saturates a five-point scale from any base interval of width 3, so a reader is entitled to ask whether the uninformativeness reported in Section 5 is a property of the data or of our own declaration. The two directions answer differently, and both answers are needed. Upward, the band is not optional. Setting 𝑏− = 0 places every adjacent under-reporting event outside the band, and the linkage evidence puts essentially all of the 49% disagreement there, so 𝛿 would rise to roughly that value and the A. Rafe and S. Das: Preprint submitted to Elsevier

Page 14 of 45

Certified crash-severity prediction

floor 1 − 𝛼 − 𝛿 to roughly 0.41. Against this, 𝑏− = 1 leaves only the beyond-adjacent mass, which weighted by the reported KABCO composition of the test split is 0.822(0.004) + 0.099(0.017) + 0.062(0.020) ≈ 0.0062. One category upward buys about two orders of magnitude in 𝛿, and no certificate worth issuing survives without it. Downward, honesty requires conceding that the cited evidence does not force 𝑏+ = 1. The Wisconsin linkage quantifies over-reporting only for category A, at 38.2% reaching two categories down, which contributes 0.014(0.382) ≈ 0.0053 of beyond-band mass under a constant unit band and brings the total to ≈ 0.0115, comfortably inside the declared 𝛿 = 0.02. It reports no adjacent over-reporting rate for O, C, or B, so a deployment could defensibly declare 𝑏+ = 0 and absorb the over-reporting into a slightly larger 𝛿. That trim is the sharpest one the evidence permits, and it does not change the finding. With 𝑏− = 1 and 𝑏+ = 0 the expansion adds one category rather than two, and the class-conditional base widths on the motorcyclist stratum, which lie between 3.55 and 4.22 across all seven base models of Section 5, compose to between 4.55 and 5. Every one remains at or above width 4 on a scale of five. The rural high-speed stratum sits at 1.95 and composes to 2.95, and is not vacuous under either band, which is consistent with what Section 5 reports there. The saturation on the strata this paper is about is therefore inherited from the base sets and not manufactured by the declaration; the band is the mechanism by which it becomes visible, not its origin. Remark 3.37 (Known-Π alternatives). If a full misclassification kernel Π(𝑦̃ ∣ 𝑦) were known and stable, quantilecorrection approaches (Einbinder et al., 2024; Sesia et al., 2025; Bortolotti et al., 2025) could yield smaller sets. We deliberately require only the support/band of Π: in crash data Π’s magnitudes vary by jurisdiction and era while its band structure is stable; robustness to Π-misspecification is the design goal. The known-Π route is complementary. Remark 3.38 (Category-dependent bands; fatals exact). Replace the constant band by a category-dependent map, e.g.  (5) = {5} and 5 ∉  (𝑦) ̃ for 𝑦̃ ≤ 4 when fatality recording is treated as exact (death-record verification), and  (4) ⊇ {2, 3, 4} where two-level over-reporting from A is material (Section 3.6.1). Theorem 3.31 holds verbatim; its proof never used constancy of the band. Consequence: with fatals exact, the K boundary never inflates, which materially reduces the practical cost of the expansion; stated as a corollary in Section 3.9.

3.7. The reporting channel imposes an identifiable floor Theorem 3.25 closed the distribution-free account: the level-set floor 𝐴 (𝛼) governs informativeness and cannot be certified from below without assumptions. This subsection reopens it with the one assumption the domain supplies for free, a declared misreporting channel, and shows that the floor then admits a nonvacuous lower bound computable from linkage constants. The pairing is the point: what is impossible distribution-free (Theorem 3.25) becomes possible given the channel, which is an identifiability statement about the data regime rather than about any model. The floor itself is a genie-aided converse, the argument style of list-size lower bounds in coding (Guruswami and Vadhan, 2010), and the recovery of its inputs from a partially known channel places it in the partial-identification tradition (Molinari, 2008); the nearest predictive-set neighbor, oracle prediction sets under interval censoring (Liu, de Paula and Tamer, 2025), is a different problem, set construction under a censored outcome rather than a floor imposed by a misclassification channel. Write 𝑀(𝑡 ∣ 𝑦) = ℙ(𝑌̃ = 𝑡 ∣ 𝑌 = 𝑦) for the channel taking true severity 𝑌 to the reported label 𝑌̃ . On a stratum 𝐴 ∈ 𝜎(𝑋) with ℙ(𝑋 ∈ 𝐴) > 0, let 𝑝𝐴 (𝑦) = ℙ(𝑌 = 𝑦 ∣ 𝑋 ∈ 𝐴) be the true composition and ∑ 𝑝̃𝐴 (𝑡) = 𝑦 𝑀(𝑡 ∣ 𝑦) 𝑝𝐴 (𝑦) the reported composition, the latter observed. We require nondifferential reporting within the stratum: 𝑀(𝑡 ∣ 𝑦, 𝑥) = 𝑀𝐴 (𝑡 ∣ 𝑦) for 𝑥 ∈ 𝐴, that is, given true severity the reporting process does not vary with covariates inside 𝐴. This is weaker than global nondifferentiality and is the honest form, since agreement rates differ across road-user types (Remark 3.43). Definition 3.39 (Channel floor). For inclusion weights 𝑐(𝑡 ∣ 𝑦) ∈ [0, 1], ∑ ∑ ∑ ∑ 𝐴∗ (𝛼) = min 𝑝𝐴 (𝑦) 𝑐(𝑡 ∣ 𝑦) s.t. 𝑝𝐴 (𝑦) 𝑐(𝑡 ∣ 𝑦) 𝑀𝐴 (𝑡 ∣ 𝑦) ≥ 1 − 𝛼. 𝑐

𝑦

𝑡

𝑦

(33)

𝑡

This is a fractional knapsack: order the cells (𝑦, 𝑡) by 𝑀𝐴 (𝑡 ∣ 𝑦) and include greedily until the coverage constraint binds, with one fractional boundary cell. It is the smallest expected set size for covering the reported label at level 1 − 𝛼 by a predictor that knows the true label and may randomize; no contiguity is imposed. A value 𝐴∗ (𝛼) < 1 is attained only by fractional rules that occasionally return the empty set; under the non-emptiness convention (Convention 3.6) the deployed floor is max{1, 𝐴∗ (𝛼)}, so a computed value below one certifies that the channel imposes no excess width on 𝐴, not that the bound is violated. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 15 of 45

Certified crash-severity prediction

Theorem 3.40 (Channel floor on expected set size; [ADAPTED] — level-set optimality (Sadinle et al., 2019) composed with a genie reduction under within-stratum nondifferential reporting). Under within-stratum nondifferential reporting, any 𝜎(𝑋)-measurable, possibly randomized set-valued predictor 𝐶 with reported-label coverage ℙ(𝑌̃ ∈ 𝐶(𝑋) ∣ 𝑋 ∈ 𝐴) ≥ 1 − 𝛼 satisfies [ ] | 𝔼 |𝐶(𝑋)| | 𝑋 ∈ 𝐴 ≥ 𝐴∗ (𝛼). |

(34)

𝐴∗ (𝛼) depends only on the stratum channel 𝑀𝐴 and the true composition 𝑝𝐴 ; no base model, feature set, or calibration scheme enters. Proof. Write 𝑁( ) for the value of (33)’s minimization over  -measurable inclusion rules 𝑐 with the same marginal coverage constraint. Any valid 𝐶 gives a feasible rule 𝑐(𝑡 ∣ 𝑥) = 𝟏{𝑡 ∈ 𝐶(𝑥)}, so 𝔼[|𝐶(𝑋)| ∣ 𝐴] ≥ 𝑁(𝜎(𝑋)). Since 𝜎(𝑋) ⊆ 𝜎(𝑋, 𝑌 ), every 𝜎(𝑋)-measurable rule is 𝜎(𝑋, 𝑌 )-measurable, so the larger feasible set gives 𝑁(𝜎(𝑋)) ≥ 𝑁(𝜎(𝑋, 𝑌 )). Under within-stratum nondifferentiality the reported law given (𝑋, 𝑌 ) = (𝑥, 𝑦) is 𝑀𝐴 (⋅ ∣ 𝑦). For any 𝜎(𝑋, 𝑌 )-measurable 𝑐(𝑡 ∣ 𝑥, 𝑦) form its conditional average 𝑐(𝑡 ̄ ∣ 𝑦) = 𝔼[𝑐(𝑡 ∣ 𝑋, 𝑌 ) ∣ 𝑌 = 𝑦, 𝑋 ∈ 𝐴], which is ∑ ∑ 𝜎(𝑌 )-measurable; because the objective 𝑡 𝑐 and the coverage integrand 𝑡 𝑐 𝑀𝐴 (𝑡 ∣ 𝑦) are both linear in 𝑐, averaging preserves the objective and the constraint, so 𝑐̄ is feasible with the same value. The minimum is therefore attained on 𝑦-measurable weights, and 𝑁(𝜎(𝑋, 𝑌 )) = 𝑁(𝜎(𝑌 )) = 𝐴∗ (𝛼), the value of (33). The chain gives (34). For tightness, take 𝑋 = (𝑌 , 𝑈 ) with 𝑈 independent uniform: the rule that includes 𝑡 with probability equal to the optimal fractional weight 𝑐 ⋆ (𝑡 ∣ 𝑌 ), using 𝑈 as the randomization device, attains 𝐴∗ (𝛼); for a fixed covariate process that carries no such device the bound is generally strict. ∑ Corollary 3.41 (Channel-imposed vacuity). 𝐴∗ (𝛼) > 1 only if 𝑦 𝑝𝐴 (𝑦) max𝑡 𝑀𝐴 (𝑡 ∣ 𝑦) < 1 − 𝛼, that is, only if no single reported category covers the stratum’s reported label at 1 − 𝛼 even knowing the true label. The converse can fail: a rule may spend width two on a peaked column and width zero on another, so that width-one coverage falls short yet the LP value is below one (the empty-set budget of Definition 3.39). When the channel does force 𝐴∗ (𝛼) > 1, the observed width on 𝐴 decomposes as 𝐴∗ (𝛼) ⏟⏟⏟

≤ irreducible 𝜎(𝑋)-floor . ⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟

channel floor

≤ observed width

(35)

The middle term is bounded but not point-identified, by Theorem 3.25; 𝐴∗ (𝛼) is its certifiable lower end. On strata whose true severity concentrates on the smeared interior categories, the channel alone forces wide sets, independent of the base model. This is the mechanism behind the empirical vacuity of Section 5: it is a property of the reporting channel on those strata, not a defect of any predictor. ( Remark 3.42 (Graceful degradation off nondifferentiality). If reporting is differential within 𝐴 with sup𝑥∈𝐴, 𝑦 𝑑TV 𝑀(⋅ ∣ ) 𝑦, 𝑥), 𝑀𝐴 (⋅ ∣ 𝑦) ≤ 𝛾, then any predictor feasible under the true process has surrogate-channel coverage at least 1−𝛼−𝛾, so 𝔼[|𝐶(𝑋)| ∣ 𝐴] ≥ 𝐴∗ (𝛼 + 𝛾). The floor moves continuously in the differential. One consequence is diagnostic: a valid predictor achieving width below 𝐴∗ (𝛼) using scene fields coded by the same officer who assigns the label is evidence of rater-mediated dependence, not a contradiction. Remark 3.43 (Identification and the specification test; [NEW]). 𝐴∗ requires 𝑀𝐴 and 𝑝𝐴 , neither observed directly. The linkage constants of Section 3.6.1 are marginal rates, not a full channel, so 𝑀𝐴 is only partially identified; following the direct-misclassification approach of Molinari (2008), the identification region for 𝐴∗ is the set of channels that meet the declared per-category constraints and reproduce the observed reported composition, 𝑝̃𝐴 = 𝑀𝐴⊤ 𝑝𝐴 on the simplex, with 𝑝𝐴 recovered through each admitted channel in turn. The channel is a declared policy input swept over this set, never estimated, consistent with the band of Assumption 3.29. The certified lower end is the admitted channel that minimizes 𝐴∗ , attained by concentrating each column on its diagonal and nearest under-reporting cell, which the linkage evidence supports; it is computed directly rather than by grid search. Selection is handled by an explicit model. Writing 𝑠(𝑦, 𝑡) = ℙ(linked ∣ 𝑌 = 𝑦, 𝑌̃ = 𝑡), the linked and population channels satisfy 𝑀𝐴link (𝑡 ∣ 𝑦) ∝ 𝑀𝐴 (𝑡 ∣ 𝑦) 𝑠(𝑦, 𝑡), so 𝑀𝐴 (𝑡 ∣ 𝑦) ∝ 𝑀𝐴link (𝑡 ∣ 𝑦)∕𝑠(𝑦, 𝑡). When linkage tracks true severity alone, 𝑠 is constant in 𝑡 and the forward channel is invariant, the optimistic case; in this domain the on-scene indicators that drive the officer’s KABCO code also drive entry into the linkable medical stream, so 𝑠 depends on the reported label given the truth. We bound that A. Rafe and S. Das: Preprint submitted to Elsevier

Page 16 of 45

Certified crash-severity prediction

dependence by max𝑡 𝑠(𝑦, 𝑡)∕ min𝑡 𝑠(𝑦, 𝑡) ≤ 𝑅 for each 𝑦 and report the certified floor over 𝑅 ∈ {1, 2, 5} (Section 5); 𝑅 = 1 recovers invariance. This yields a testable prediction rather than a fitted quantity: since 𝐴∗ (𝛼) lower-bounds the width of every valid predictor, an achieved valid width below the identified interval refutes the declared channel on 𝐴, an overidentification test on the declared set. The identified set is accordingly the declared-consistent set intersected with {𝑀𝐴 ∶ 𝐴∗ (𝛼) ≤ 𝑤ach for all 𝐴}, where 𝑤ach is an achieved valid width; on the Texas records this intersection 𝐴 𝐴 removes nothing, but it is the general form. Section 5 reports the test on every stratum, with the reliable low-severity stratum first because it binds most tightly.

3.8. Spatiotemporal deployment shift A deployment cell is (time bin) × (spatial unit). Two regimes: strata observed in calibration, and genuinely new strata. Theorem 3.44 (Observed strata: group-conditional validity and rollup; [ADAPTED]). Let 𝑔 ∶  →  assign each record its stratum (e.g. county×year; 𝑔 is a fixed function of record fields). Calibrating per stratum as in Theorem 3.11 yields, for every stratum 𝛾 with calibration points, ( ) ℙ 𝑌̃𝑛+1 ∈ 𝐶(𝑋𝑛+1 ) ∣ 𝑔(𝑋𝑛+1 ) = 𝛾 ≥ 1 − 𝛼, (36) hence for any deployment mixture over observed strata with unknown weights 𝜔𝛾 , ∑ 𝑃test = 𝜔𝛾 𝑃𝛾 ⟹ 𝑃test (𝑌̃ ∈ 𝐶) ≥ 1 − 𝛼.

(37)

𝛾

With the hierarchical rollup 𝛾 ↦ parent(𝛾) applied whenever 𝑛𝛾 < 𝑛min (county → CBSA → district → statewide), the guarantee holds conditional on the coarsest cell actually used for calibration. Proof. The per-stratum statement is Theorem 3.11 with partition 𝑔; strata are defined by record fields, so genuinely new strata are empty in calibration and are exactly the case deferred to Theorem 3.45. The mixture statement is the tower property. Rollup: the union of leaf strata sharing a parent is a cell of a coarser fixed partition; apply Theorem 3.11 to that partition. Transport reading: redistribution of traffic and exposure across already-observed county-year strata, the dominant mode of medium-term drift, can never break the certificate. What it cannot cover is a new stratum.

Choice of 𝑛min . We set 𝑛min = ⌈2∕𝛼⌉ ⋅ 50

(38)

(e.g. 𝑛min = 1000 at 𝛼 = 0.1). Rationale: with 𝑛𝛾 ≥ 𝑛min , the per-cell quantile index granularity 1∕(𝑛𝛾 + 1) ≤ 𝛼∕100, so discreteness inflates per-cell coverage by at most one percent of the nominal miscoverage budget, and the binomial √ standard error of empirically auditing a cell’s coverage at level 1 − 𝛼 is at most 𝛼(1 − 𝛼)∕𝑛min ≈ 0.0095 at 𝛼 = 0.1. The floor is a granularity-and-auditability choice, not a validity requirement: Theorem 3.44 is valid at any 𝑛𝛾 ≥ 1. Theorem 3.45 (New stratum: covariate-shift transfer with estimated weights; composition [NEW], ingredients [ADAPTED] — Tibshirani et al., 2019; Lei and Candès, 2021; Foygel Barber et al., 2023). Let the target stratum 𝑔 ∗ (a new county, or a future year) satisfy the covariate-shift condition: 𝑃𝑔∗ has 𝑋-marginal 𝑃𝑋∗ and the same conditional 𝑃cal (𝑌̃ ∣ 𝑋). Let 𝑤(𝑥) = d𝑃𝑋∗ ∕d𝑃cal,𝑋 (𝑥) exist, and let 𝑤̂ ≥ 0 be an estimate fit on tr plus unlabeled covariates from ̂ 𝑔 ∗ . Run weighted split conformal with weights 𝑤: 𝑛 { } ∑ ̂ 𝑞̂𝑤̂ = inf 𝑡 ∶ 𝑝𝑤 𝟏{𝑆 ≤ 𝑡} ≥ 1 − 𝛼 , 𝑖 𝑖 𝑖=1

̂ 𝑝𝑤 𝑖 = ∑𝑛

̂ 𝑖) 𝑤(𝑋

̂ 𝑗 ) + 𝑤(𝑋 ̂ 𝑛+1 ) 𝑗=1 𝑤(𝑋

,

(39)

(∑ ) ̂ ̂ 𝑛+1 )∕ 𝑗 𝑤(𝑋 ̂ 𝑗 ) + 𝑤(𝑋 ̂ 𝑛+1 ) is placed at +∞ (Tibshirani et al., 2019); in where the remaining mass 𝑝𝑤 = 𝑤(𝑋 𝑛+1 particular 𝑞̂𝑤̂ = +∞ (i.e. 𝐶 = ) whenever the calibration mass cannot reach 1 − 𝛼. Define the estimated tilt 𝑄𝑤̂ on  by d𝑄𝑤̂ ∝ 𝑤̂ d𝑃cal,𝑋 .

(40)

Then: A. Rafe and S. Das: Preprint submitted to Elsevier

Page 17 of 45

Certified crash-severity prediction

Algorithm 1 New-stratum transfer certificate (Theorem 3.45): weighted calibration plus the exported three-number tuple Require: calibration scores {𝑆𝑖 }𝑖∈cal and covariates {𝑋𝑖 }𝑖∈cal ; training-split covariates; unlabeled target covariates from 𝑔 ∗ ; level 𝛼; one-sided confidence level 1 − 𝑎 for the slack bound; weight cap 𝑤max 1: split the unlabeled target sample into a ratio-fitting half f it and a diagnostic half diag } {̂ 𝑛 ̂ ⋅ 𝑛0 , 𝑤max ⊳ 2: fit a probabilistic classifier to (training-split covariates = 0) vs (f it = 1); set 𝑤(𝑥) ← min 𝑝(1∣𝑥) 𝑝(0∣𝑥) ̂ 1 any 𝑤̂ ≥ 0 is admissible: estimation error, incl. clipping, is absorbed by the slack in (41) 3: for each test point 𝑥 from 𝑔 ∗ do ̂ 4: 𝑞̂𝑤̂ (𝑥) ← weighted quantile (39) with test weight 𝑤(𝑥); predict 𝐶𝑞̂𝑤̂ (𝑥) (𝑥) ⊳ Thm. 3.45(i) 5: end for ̂ 6: draw a 𝑤-weighted resample of the calibration covariates ⊳ a sample from 𝑄𝑤̂ of (40) 7: train a discriminator ℎ on one half of diag vs one half of the resample; compute the balanced accuracy ba(ℎ) on the two held-out halves ⊳ sample-split: ℎ never sees its evaluation data { } ̂ 8: 𝑑 ← max 0, 2 ba (ℎ) − 1 , with ba a one-sided level-(1 − 𝑎) lower confidence bound formed by a TVLCB LCB LCB Bonferroni split of 𝑎 across exact (Clopper–Pearson) bounds on the two class-wise binomial accuracies ⊳ valid lower bound on the slack by (42) ̂ is declared realizable then 9: if the parametric family of 𝑤 10: compute the plug-in slack estimate of Theorem 3.45(iii) 11: end if ( ) ̂ 12: emit the certificate tuple 1 − 𝛼, 𝑑 ⊳ honesty box: the diagnostic certifies TVLCB , parametric slack estimate weakness, never strength (i) (Slack identity.) ) ( ( ) 𝑃𝑔∗ 𝑌̃𝑛+1 ∈ 𝐶𝑞̂ (𝑋𝑛+1 ) ≥ 1 − 𝛼 − 𝑑TV 𝑃 ∗ , 𝑄𝑤̂ , 𝑤̂

𝑋

(41)

and the slack is a purely covariate-space quantity between two sampleable distributions: it involves no labels and no unknown 𝑤. (ii) (Assumption-free diagnostic.) For any classifier ℎ trained to distinguish samples of 𝑃𝑋∗ (an unlabeled target ̂ sample drawn independently of the scored calibration points) from samples of 𝑄𝑤̂ (𝑤-weighted resampling of calibration covariates), the balanced accuracy satisfies 2 ba(ℎ) − 1 ≤ 𝑑TV (𝑃𝑋∗ , 𝑄𝑤̂ );

(42)

a sample-split estimate of 2 ba(ℎ) − 1 whose two class-wise accuracies are each given an exact (Clopper– Pearson) lower bound on their binomial proportion, combined by a Bonferroni split of the level, is a valid lower confidence bound on the slack. A large value certifies the transfer certificate is weak; a small value is necessary, not sufficient. (iii) (Upper bound under realizability.) If 𝑤 lies in a parametric family {𝑤𝜃 } (e.g. logistic density ratio in engineered covariates) and 𝑤̂ = 𝑤𝜃̂ with 𝜃̂ the MLE of the associated probabilistic classifier on 𝑚 unlabeled target and 𝑛′ ( ) calibration-split samples, then under standard regularity 𝑑TV (𝑃𝑋∗ , 𝑄𝑤̂ ) = 𝑂𝑝 (𝑚 ∧ 𝑛′ )−1∕2 .

Honesty box (scope of Theorem 3.45). (a) The theorem assumes covariate shift: the conditional severity-reporting process transfers. This is declared and partially falsifiable: with even a small labeled batch from 𝑔 ∗ , conditional coverage can be audited; without labels, concept drift is detectable only prospectively (conformal test martingales; Vovk et al., 2022), implemented as monitoring, not correction (Algorithm 2). (b) The classifier diagnostic lower-bounds the slack; an assumption-free finite-sample upper bound on total variation from unlabeled samples alone is impossible. The exported certificate is therefore the three-number tuple (nominal level 1 − 𝛼, slack lower confidence bound, parametric slack estimate under (iii)), and nothing more. Algorithm 1 states the full certified procedure.

Drift monitoring (detection, never correction). The covariate-shift condition of Theorem 3.45 is not label-free falsifiable, so deployment pairs every certificate with a sequential monitor. For each deployed record 𝑡 with score 𝑠𝑡 , A. Rafe and S. Das: Preprint submitted to Elsevier

Page 18 of 45

Certified crash-severity prediction

Algorithm 2 Deployment drift monitor (conformal test martingale; detection only) Require: frozen calibration scores {𝑆𝑖 }𝑖∈cal ; betting parameter 𝜀 ∈ (0, 1); alarm budget 𝛼mon 1: 𝑀 ← 1 2: for each deployed record 𝑡 = 1, 2, … do 3: compute 𝑠𝑡 and the smoothed p-value 𝑝𝑡 by (43) 4: 𝑀 ← 𝑀 ⋅ 𝜀 𝑝𝜀−1 ⊳ running product (44) 𝑡 5: if 𝑀 ≥ 1∕𝛼mon then 6: alarm: freeze new certificates; re-calibrate on post-drift data ⊳ false-alarm rate ≤ 𝛼mon by Ville’s inequality 7: end if 8: end for compute the smoothed conformal p-value against the frozen calibration scores, ( ) #{𝑖 ∈ cal ∶ 𝑆𝑖 > 𝑠𝑡 } + 𝑈𝑡 #{𝑖 ∈ cal ∶ 𝑆𝑖 = 𝑠𝑡 } + 1 𝑝𝑡 = , 𝑈𝑡 ∼ Unif (0, 1) i.i.d., 𝑛+1

(43)

and the power test martingale ∏ 𝜀 ∈ (0, 1). 𝜀 𝑝𝜀−1 𝑀𝑡 = 𝑠 ,

(44)

𝑠≤𝑡

Under the online protocol, in which each scored record is appended to the calibration bag before the next is scored, the 𝑝𝑡 are independent uniform under exchangeability, 𝑀𝑡 is a non-negative martingale with 𝑀0 = 1, and Ville’s inequality bounds the false-alarm rate exactly: ℙ(sup𝑡 𝑀𝑡 ≥ 1∕𝛼mon ) ≤ 𝛼mon ([KNOWN] — Vovk et al., 2022). Deployed against a frozen calibration set (Algorithm 2), the 𝑝𝑡 share one calibration sample and are exchangeable but not independent (Bates, Candès, Lei, Romano and Sesia, 2023), so the martingale property holds only up to that dependence; the dependence is 𝑂(1∕|cal |) and vanishes as the calibration set grows, negligible at the |cal | ∼ 106 scale here, and an exact deployed bound is recovered by the online update at the cost of a growing calibration bag. A triggered monitor mandates re-calibration; it never adjusts an issued certificate retroactively. Algorithm 2 is the deployed loop.

3.9. Severity-cost risk control Let 𝜅 ∶  → [0, ∞) be non-decreasing with 𝜅max = 𝜅(𝐾) > 0 (USDOT/FHWA comprehensive crash-cost scale; the domain admits the indicator cost of Corollary 3.48). Theorem 3.46 (Cost-weighted false-omission control; [ADAPTED] — conformal risk control, Angelopoulos et al., 2024). Define the loss 𝓁𝜆 (𝑥, 𝑦) ̃ =

𝜅(𝑦) ̃ 𝟏{𝑦̃ ∉ 𝐶𝜆 (𝑥)} ∈ [0, 1]. 𝜅max

(45)

Then 𝜆 ↦ 𝓁𝜆 is non-increasing and right-continuous with 𝓁1 ≡ 0 (as 𝑠 ≤ 1 pointwise implies 𝐶1 = ). With { } 1 𝑛 ̂ 𝜆̂ = inf 𝜆 ∶ 𝑛+1 𝑅𝑛 (𝜆) + 𝑛+1 ≤ 𝛽 ,

𝑅̂ 𝑛 (𝜆) = 1𝑛

𝑛 ∑

𝓁𝜆 (𝑋𝑖 , 𝑌̃𝑖 ),

(46)

𝑖=1

under A1: [ ] 𝔼 𝜅(𝑌̃𝑛+1 )𝟏{𝑌̃𝑛+1 ∉ 𝐶 ̂ (𝑋𝑛+1 )} ≤ 𝛽 𝜅max . 𝜆

(47)

Proof. Monotonicity: 𝐶𝜆 is nested non-decreasing (Lemma 3.5(iii)), so 𝟏{𝑦̃ ∉ 𝐶𝜆 } is non-increasing in 𝜆. Rightcontinuity: 𝜆 ↦ 𝟏{𝑠(𝑥, 𝑦) ̃ > 𝜆} is right-continuous. Boundedness with 𝐵 = 1 after normalization; and 𝓁𝜆max = 𝓁1 ≡ 0 ≤ 𝛽. These are exactly the hypotheses of the conformal risk control theorem (Angelopoulos et al., 2024), whose 𝐵 𝑛 ̂ threshold is inf {𝜆 ∶ 𝑛+1 𝑅𝑛 (𝜆) + 𝑛+1 ≤ 𝛽} with 𝜆̂ = 𝜆max if the set is empty; the conclusion is theirs, rescaled by 𝜅max . A. Rafe and S. Das: Preprint submitted to Elsevier

Page 19 of 45

Certified crash-severity prediction

Operational reading: expected severity-weighted cost of screening misses is provably at most 𝛽𝜅max . Sets remain intervals, so decision rules such as “escalate when 𝐶𝜆̂ (𝑥) ⊆ {4, 5} or min 𝐶𝜆̂ (𝑥) ≥ 3” are well-defined and monotone. Theorem 3.47 (Risk control on true severity under banded noise; [NEW]). Assume N(𝑏− , 𝑏+ , 𝛿) and let ( ) 𝜅 + (𝑦) ̃ = 𝜅 min(𝑦̃ + 𝑏− , 𝐾) ,

(48)

the largest cost compatible with report 𝑦̃ within the band. Run Theorem 3.46 with 𝜅 + in place of 𝜅 (same machinery; loss bounded by 1 after normalizing by 𝜅max ) and output expanded sets 𝐶 ⊕ . Then ̂ 𝜆

[

] 𝔼 𝜅(𝑌𝑛+1 )𝟏{𝑌𝑛+1 ∉ 𝐶 ⊕ (𝑋𝑛+1 )} ≤ 𝛽 𝜅max + 𝛿 𝜅max . ̂ 𝜆

(49)

𝑐 ) ≤ 𝛿: Proof. Split on the band event 𝐵𝑛+1 = {𝑌𝑛+1 ∈  (𝑌̃𝑛+1 )}, ℙ(𝐵𝑛+1

[ ] 𝔼[𝜅(𝑌 )𝟏{𝑌 ∉ 𝐶 ⊕ }] ≤ 𝔼 𝜅(𝑌 )𝟏{𝑌 ∉ 𝐶 ⊕ }𝟏𝐵 + 𝜅max 𝛿. On 𝐵: (a) 𝑌 ∉ 𝐶 ⊕ ⇒ 𝑌̃ ∉ 𝐶̃ (contrapositive of Definition 3.30); (b) 𝑌 ≤ 𝑌̃ + 𝑏− , so 𝜅(𝑌 ) ≤ 𝜅 + (𝑌̃ ) by monotonicity. Hence the first term is at most 𝔼[𝜅 + (𝑌̃ )𝟏{𝑌̃ ∉ 𝐶̃𝜆̂ }] ≤ 𝛽𝜅max by Theorem 3.46 applied to the 𝜅 + -loss (bounded by 𝜅(𝐾) = 𝜅max ). Corollary 3.48 (Fatal-omission control under a declared exact-fatality premise; [NEW]). Take 𝜅(𝑦) = 𝟏{𝑦 = 5} and suppose fatality recording is exact (𝑌̃ = 5 ⟺ 𝑌 = 5; Remark 3.38, and see Remark 3.49 on what this premise requires). Then, with no 𝛿 inflation, ( ) ℙ 𝑌𝑛+1 = 5, 𝑌𝑛+1 ∉ 𝐶 ⊕ (𝑋𝑛+1 ) ≤ 𝛽. (50) ̂ 𝜆

Proof. On {𝑌 = 5}: 𝑌̃ = 5 = 𝑌 ; then 𝑌 ∉ 𝐶 ⊕ ⇒ 𝑌 ∉ 𝐶̃ (as 𝐶̃ ⊆ 𝐶 ⊕ by the reflexivity of  , Assumption 3.29) ̃ So the risk is at most 𝔼[𝟏{𝑌̃ = 5}𝟏{𝑌̃ ∉ 𝐶̃ ̂ }] ≤ 𝛽 by Theorem 3.46 with this 𝜅; exactness replaces the ⇒ 𝑌̃ ∉ 𝐶. 𝜆 band event, so no 𝛿 enters. Under the exact-fatal map of Remark 3.38 the inflated cost 𝜅 + of (48) coincides with 𝜅, so 𝜆̂ is the Theorem 3.46 threshold and 𝐶̃ ⊕ is its expansion. ̂ 𝜆

Remark 3.49 (Which direction of exactness the fatal bound needs). The corollary is stated with 𝑌̃ = 5 ⟺ 𝑌 = 5, and the two directions are not equally supported, so the asymmetry is declared rather than glossed. The proof uses only the under-reporting direction, 𝑌 = 5 ⇒ 𝑌̃ = 5: a true fatality must not be recorded as anything milder. The evidence assembled in Section 3.6.1 establishes the over-reporting direction, 𝑌̃ = 5 ⇒ 𝑌 = 5 (in the Wisconsin linkage every K-coded victim is MAIS 6; Burdett et al., 2015; Farmer, 2003), which the proof does not need. The direction the proof does need is a real assumption about the reporting process, not a fact the linkage studies deliver. Its known failure mode is timing: a report is often filed before a victim dies, which is exactly why fatality counting conventions impose a post-crash window. In CRIS the severity field is subject to later amendment, so the premise is that the record consulted is the amended one. Deployments that score reports at filing time, before amendment, do not satisfy it, and for them the honest instrument is Theorem 3.47 with a declared 𝛿 at the K boundary rather than this corollary. We state the exact-fatal case as a declared premise, on the same footing as the band itself, and not as a measured property of the data. The corollary in words: the probability that a true fatality falls outside the flagged set is provably at most 𝛽, regardless of the model, the covariate distribution, and B/C reporting noise.

3.10. The composition theorem: certification calculus Theorem 3.50 (Composition; [NEW] as packaged, proof elementary by design). Let 𝑐̂ be a class partition (Section 3.3), 𝑔 a stratum map with observed strata obs and new strata handled by weights 𝑤̂ (Section 3.8), and let N( , 𝛿𝑐,𝛾 ) hold conditionally on each product cell (𝑐, 𝛾). Calibrate on the product partition {(𝑐, 𝛾)} with the rollup rule at floor 𝑛min ; apply the compatibility expansion. Then: (i) observed cell (𝑐, 𝛾): ( ) ℙ 𝑌 ∈ 𝐶 ⊕ ∣ 𝑐̂ = 𝑐, 𝑔 = 𝛾 ≥ 1 − 𝛼 − 𝛿𝑐,𝛾 ; A. Rafe and S. Das: Preprint submitted to Elsevier

(51) Page 20 of 45

Certified crash-severity prediction

Algorithm 3 CHOIR certification (condition → weight → calibrate → expand → risk-adjust) Require: calibration scores 𝑆𝑖 = 𝑠(𝑋𝑖 , 𝑌̃𝑖 ), 𝑖 ∈ cal ; class map 𝑐; ̂ stratum map 𝑔 with rollup hierarchy and floor 𝑛min of (38); compatibility map  with beyond-band mass 𝛿𝑐,𝛾 ; level 𝛼 (or cost 𝜅 and budget 𝛽); unlabeled target covariates for any new stratum 𝑔 ∗ 1: assign each 𝑖 ∈ cal its product cell (𝑐(𝑋 ̂ 𝑖 ), 𝑔(𝑋𝑖 )) 2: while any cell has 𝑛𝑐𝛾 < 𝑛min do 3: merge the cell into its parent stratum (county → CBSA → district → statewide) ⊳ Thm. 3.44 4: end while 5: for each cell (𝑐, 𝛾) do 6: if 𝛾 observed in calibration then 7: 𝑞̂𝑐𝛾 ← the ⌈(1 − 𝛼)(𝑛𝑐𝛾 + 1)⌉-th smallest score in the cell, as in (8) ⊳ Thm. 3.11 8: else (𝛾 = 𝑔 ∗ new) ̂ take the weighted quantile (39), record the slack lower confidence 9: run Algorithm 1 within the cell: fit 𝑤, bound ⊳ Thm. 3.45 10: end if 11: optionally replace the quantile by 𝜆̂ of (46) with cost 𝜅 + of (48) ⊳ Thms. 3.46–3.47 12: end for ̃ 13: at prediction: 𝐶(𝑥) = 𝐶𝑞̂𝑐(𝑥),𝑔(𝑥) (𝑥); expand to 𝐶̃ ⊕ (𝑥) via (28) ⊳ Def. 3.30 ̂ 14: emit the per-cell certificate (53), each slack with its provenance ⊳ Thm. 3.50 (ii) new stratum 𝑔 ∗ , class 𝑐: ) ( ∗ ( ) , , 𝑄𝑤,𝑐 𝑃𝑔∗ 𝑌 ∈ 𝐶 ⊕ ∣ 𝑐̂ = 𝑐 ≥ 1 − 𝛼 − 𝛿𝑐,𝑔∗ − 𝑑TV 𝑃𝑋∣𝑐 ̂

(52)

̂ where 𝑄𝑤,𝑐 of the class-𝑐 calibration covariate law; ̂ is the 𝑤-tilt (iii) replacing the quantile step by the risk-control step of Section 3.9 in any cell yields the corresponding bound 𝛽𝜅max + 𝛿𝑐,⋅ 𝜅max in that cell. All slacks are additive, and each slack term is attributable to exactly one declared assumption (N for 𝛿; covariate shift plus estimation for 𝑑TV ). Proof. (i) Theorem 3.11 applied to the product partition (a fixed measurable function of the record) gives cellconditional 1 − 𝛼 on reported labels; Theorem 3.31’s argument, run conditionally on the cell, subtracts 𝛿𝑐,𝛾 . (ii) Within class 𝑐, Theorem 3.45(i) applies with all distributions conditioned on 𝑐̂ = 𝑐 (the class is a covariate-measurable event, so the covariate-shift condition and the tilt are well defined class-wise); then subtract 𝛿 as in (i). (iii) Theorems 3.46 and 3.47 run cell-wise. Additivity is the union bound / telescoping of the two-step arguments already given. Remark 3.51 (The one delicate point). Mondrian (product-partition) calibration and weighted quantiles interact correctly because the weighting of Theorem 3.45 is applied within a cell of a fixed partition; per-cell quantiles are never mixed with global weights. The canonical order is: condition → weight → calibrate → expand → risk-adjust. The package enforces this order by construction; Algorithm 3 states it. Remark 3.52 (Slack budget as an artifact, and what it is not). The deployment certificate is the table of per-cell values ̂ 𝛾) = (1 − 𝛼) − Cert(𝑐,

∑

⏟⏟⏟

LCB 𝑑̂ TV 𝑐,𝛾 ⏟⏟⏟

declared; subtractable

estimated; not subtractable

(𝑗) 𝑗 𝛿𝑐,𝛾

−

,

(53)

each term carried with its provenance. The two kinds of term are not interchangeable, and conflating them would undo the honesty box of Theorem 3.45. The declared terms 𝛿 (𝑗) are inputs. Subtracting them is exactly Theorem 3.50, and the result is a genuine lower bound. The transfer term is not: Theorem 3.50(ii) subtracts the true 𝑑TV , whereas the exported quantity is only a lower confidence bound on 𝑑TV (Theorem 3.45(ii), which is all that unlabeled data permits). Subtracting a lower bound on ̂ is not a coverage guarantee and must never be a penalty yields an upper bound on the floor. Consequently Cert A. Rafe and S. Das: Preprint submitted to Elsevier

Page 21 of 45

Certified crash-severity prediction

∑ reported as one: it is the best case the evidence allows, and the true floor (1 − 𝛼) − 𝑗 𝛿 (𝑗) − 𝑑TV can be lower. Empirically it is: Section 3’s spatial experiment finds certified cells whose realized weighted coverage falls below ̂ which is permitted and expected. Cert, What the artifact delivers is therefore diagnostic rather than probative, and that is the honest claim: it is auditable, LCB its terms are attributable, and a large 𝑑̂ is a proof of weakness that no deployment should ignore. A small one is TV not a proof of strength. This is the “verifying, validating, and certifying” deliverable in the only form unlabeled target data can support.

3.11. Default base model and its non-load-bearing identifiability DLCON (deep latent-class ordinal network) is the default base model: a mixture of ordinal experts, 𝑃̂ (𝑌 ≤ 𝑘 ∣ 𝑥) =

𝐶 ∑

( ) 𝜋𝑐 (𝑥𝑊 ; 𝜙) 𝜎 𝜏𝑐,𝑘 − 𝑓𝑐 (𝑥; 𝜃𝑐 ) ,

𝑘 = 1, … , 𝐾 − 1,

(54)

𝑐=1

with 𝜎 the logistic function, per-class MLP experts 𝑓𝑐 (⋅; 𝜃𝑐 ), and a gating network on a designated covariate sub-vector 𝑥𝑊 of 𝑥 (slow, planning-level covariates when the latent classes are to be interpreted; 𝑥𝑊 = 𝑥 in our implementation), 𝜋𝑐 (𝑥𝑊 ; 𝜙) = ∑𝐶

exp 𝑔𝑐 (𝑥𝑊 ; 𝜙)

(55)

,

𝑐 ′ =1 exp 𝑔𝑐 ′ (𝑥𝑊 ; 𝜙)

where 𝑔(⋅; 𝜙) is an MLP. Thresholds are strictly ordered by construction via cumulative softplus increments, 𝜏𝑐,1 ∈ ℝ free,

𝜏𝑐,𝑘 = 𝜏𝑐,1 +

𝑘−1 ∑

(

) log(1 + 𝑒𝑢𝑐,𝑗 ) + 𝜖 ,

𝑘 = 2, … , 𝐾 − 1,

(56)

𝑗=1

with free parameters 𝑢𝑐,𝑗 and a fixed jitter 𝜖 = 10−3 guaranteeing 𝜏𝑐,1 < ⋯ < 𝜏𝑐,𝐾−1 . Training minimizes the ordinal negative log-likelihood plus the ranked probability score on the training labels, 1 ∑ NLL = − log 𝑝( ̂ 𝑌̃𝑖 ∣ 𝑥𝑖 ), 𝑝(𝑘 ̂ ∣ 𝑥) = 𝐹̂ (𝑘 ∣ 𝑥) − 𝐹̂ (𝑘 − 1 ∣ 𝑥), (57) 𝑛tr 𝑖∈ tr

RPS =

∑ 𝐾−1 ∑

1 𝑛tr 𝑖∈

(

)2 𝐹̂ (𝑘 ∣ 𝑥𝑖 ) − 𝟏{𝑌̃𝑖 ≤ 𝑘} ,

(58)

tr 𝑘=1

 = NLL + 21 RPS ,

(59)

optionally augmented by an entropy load-balancing penalty on the batch-mean gate distribution, which discourages gate collapse onto a single expert (an efficiency and interpretability aid; no guarantee depends on it). The MAP class 𝑐(𝑥) ̂ = arg max 𝜋𝑐 (𝑥𝑊 ; 𝜙) 𝑐

(60)

feeds Theorem 3.11. Proposition 3.53 (Identifiability up to permutation; [ADAPTED]). Consider the population model with linear experts 𝑓𝑐 = 𝜃𝑐⊤ 𝜓(𝑥) and multinomial-logistic gating in 𝑥𝑊 . If (a) classes are distinct, (𝜃𝑐 , 𝜏𝑐,⋅ ) ≠ (𝜃𝑐 ′ , 𝜏𝑐 ′ ,⋅ ) for 𝑐 ≠ 𝑐 ′ ; (b) the joint support of (𝜓(𝑋), 𝑥𝑊 ) is not contained in a proper affine subspace; and (c) 𝐶 equals the true number of components, then the mixture representation is identifiable up to label permutation. This follows from identifiability of finite mixtures of GLMs / mixtures of experts (Follmann and Lambert, 1991; Jiang and Tanner, 1999; Grün and Leisch, 2008), with ordered-logit experts handled through the binary-collapse argument applied at each cumulative threshold.

Scope note. No guarantee in Sections 3.2–3.10 depends on Proposition 3.53. Validity holds for arbitrary, even misspecified, 𝐹̂ and 𝑐. ̂ Proposition 3.53 answers the econometric question of whether the latent classes are well-defined objects: under (a)–(c), yes, in the linear-expert case; with deep experts the classes are used only as a partition, where Theorem 3.11 needs nothing. Architecture optional, guarantees unconditional: CHOIR’s central design principle. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 22 of 45

Certified crash-severity prediction

4. Data and experimental design 4.1. Source, unit of analysis, and attrition The empirical work uses the Texas Crash Records Information System (CRIS), the statutory repository of policereported crashes maintained by the Texas Department of Transportation, for calendar years 2017 through 2025. CRIS is distributed as linked crash, unit, and person tables. The analysis joins these to two external sources: a vehicle table decoded from the reported vehicle identification number through the National Highway Traffic Safety Administration vPIC service, and a geographic file mapping each crash identifier to a census block group and, where one applies, a core-based statistical area. The unit of analysis is the driver. A record enters the sample when its unit is a motor vehicle and the person is that unit’s driver, and the reported severity 𝑌̃ is the person’s KABCO injury code mapped to the ordinal scale of Section 3.1. Records whose severity is coded unknown are dropped rather than imputed, because an imputed severity is a fabricated label and every guarantee in Section 3 is a statement about the label actually recorded. Table 2 reports the attrition. Of 11,250,255 raw crash-unit-person rows, 10,678,038 belong to motor-vehicle units, 10,047,354 are driver rows, and 9,078,415 carry a valid severity code. Geocoding to a block group succeeds for 8,585,232 of these. Geocoding is recorded as an indicator and never used as a filter, so a record that fails to geocode still contributes to every analysis that does not condition on geography. Table 2 Attrition from raw CRIS records to the analysis sample. Stage

Rows

Dropped

CRIS crash-unit-person rows, 2017–2025 Motor-vehicle units Driver rows Valid KABCO severity Geocoded to block groupa

11,250,255 10,678,038 10,047,354 9,078,415 8,585,232

572,217 630,684 968,939

Primary analysis rows, one driver per crashb

5,213,921

a An indicator, not a filter; records that fail to geocode remain in the sample. b Sampled under Remark 3.2; see Section 4.2.

The last row of Table 2 follows from the exchangeability requirement rather than from any data-quality consideration, and it is the one that matters for validity.

4.2. Crash-level splitting and the one-driver-per-crash protocol Two people in the same crash do not carry independent severities. They share a speed, a road, a weather condition, and an impact. Assumption 3.1 is therefore violated by any split that places one occupant of a crash in the calibration set and another in the test set, and the violation is not a technicality. It inflates the effective calibration sample, and the quantile that results is calibrated against a dependence structure the theory does not admit. This is the substance of Remark 3.2, and the protocol it prescribes is followed without exception. Every split in this paper is taken at the crash identifier, so all rows from one crash fall on the same side of every partition. The primary analysis goes further and samples one driver row per crash uniformly at random under seed 20260704, yielding 5,213,921 records from the 9,078,415 valid-severity driver rows. This restores the structure the theory assumes: crashes are independent, so sampled rows are independent, so Assumption 3.1 holds. Every result in Section 5 is computed on these primary rows. Retaining all driver rows under crash-level splitting remains approximately valid and is reported as a robustness check; a fully rigorous treatment of multiple rows per crash belongs to the grouped conformal literature (Dunn et al., 2023). Three split designs carry the experiments. The random design S1 assigns 60% of crashes to training, 20% to calibration and 20% to test, and verifies the results that assume exchangeability. The temporal design S2 trains on 2017 through 2021, calibrates on 2022 and 2023, and tests on 2024 and 2025; this is the deployment geometry an agency actually faces, and the setting in which the covariates of the test period are observable while its labels are not. The spatial design S3 calibrates on 200 counties and tests on 54 held out counties stratified by rural and urban status, simulating transfer to jurisdictions the calibration never saw.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 23 of 45

Certified crash-severity prediction

4.3. Features, exclusions, and missingness Features are declared in two configurations, and the distinction states when a prediction is being made rather than expressing a modeling preference. Configuration P contains only what is knowable before a crash is investigated: crash context such as posted speed limit, weather, light and surface condition, road type and alignment, traffic control, functional system, annual daily traffic, day of week and hour; person attributes such as driver age, gender, restraint use and license class; vehicle attributes decoded from the vehicle identification number, including vehicle age, body class, gross weight class, drive type, door count, engine displacement and power, and the presence of electronic stability control, collision warning, lane departure warning, blind spot monitoring and curtain or side airbags; and geography. Configuration T adds what an investigator observes on arrival, namely manner of collision, harmful event, object struck, vehicle damage severity, airbag deployment and ejection. Configuration T is the emergency response and triage setting, where those fields are available at decision time, and it is reported as that rather than as a better model. Every reported guarantee in this paper uses Configuration P, so no certified coverage or risk statement depends on information a deployment could not have; Configuration T enters once, in the width-budget decomposition of Section 5, where it measures how much of the residual set width a richer scene-time record could recover and is labeled as such. Configuration T is defined here because it is the setting in which the restraint and collision fields are unambiguous, and because naming it is what keeps the leakage question from becoming an argument, and no certified guarantee below is computed under it. Restraint use sits at the boundary and the boundary is worth naming precisely, because it is sharper than a leakage concern. The field describes a pre-crash state but is recorded after the crash, by the same officer, on the same form, at the same moment as the severity code. This paper’s premise is that the officer’s severity assessment disagrees with medical severity roughly half the time and errs in a structured, agency-varying way. It would be incoherent to treat that officer’s severity field as a measurement problem requiring a theorem while treating the same officer’s restraint field as a clean covariate. The plausible failure is not leakage but correlated recording: an officer coding a fatality may attend to restraint status differently than one coding property damage. We retain the field in Configuration P, because a deployment would have it, and we flag that any coefficient or partition resting on it inherits the reporting process rather than describing the world. No guarantee in this paper depends on that field being accurate; validity holds for any partition and any base model, whatever they are built from. Certain fields are excluded from every configuration. Injury and death counts, suspected serious injury counts, and the crash-level severity code are excluded because they encode the outcome. Investigator narratives are excluded, as are blood alcohol results, which are present for a small and non-random minority of records and are recorded after the fact. All personally identifying material, including the vehicle identification number itself, license and registration identifiers, dates of birth, and free-text narratives, is dropped at ingest and appears in no released artifact. Two properties of the decoded vehicle fields constrain what may be said with them. The vPIC decode returns equipment availability rather than equipment engagement, so a record indicating that a vehicle line offers automatic emergency braking does not establish that the system was present, enabled, or active in a given crash. These fields are treated as vehicle-technology covariates, and nothing in this paper is a statement about what any such system does. Decode coverage is also systematically sparser for older vehicles, so an undecoded indicator enters the design matrix rather than the record being silently dropped. Missingness is handled explicitly and no record is discarded for incompleteness. Categorical fields receive an explicit missing level. Continuous fields are median imputed with an accompanying missingness indicator, so a model can condition on the fact of absence. Two continuous fields are substantially incomplete, annual daily traffic at 48.65% and engine power at 50.99%, and both are reported in Table 3 with their indicators noted. Incompleteness of traffic volume is also why the descriptive map below uses hour of day rather than volume as its second axis. Separately, two fields carry sentinel values rather than nulls: posted speed limit is recorded as −1 on 244,965 primary records, and vehicle age is negative on 18,335 records whose model year post-dates the crash year. Neither is a null, so neither registers in a missingness count, and both enter the design matrix as stored. The descriptive statistics in Table 3 exclude them and say so.

4.4. Composition of the sample Table 3 reports the composition of the analysis sample in three panels: records by year and reported outcome, the covariate profile across the severity ladder, and the feature blocks with their completeness. Panel A shows a stable annual volume near 580,000 records, with the lowest total of the series in 2020. Panel B is the panel that motivates the framework. Reading across the ladder from no injury to fatal, mean posted speed limit rises from 44.6 to 55.5 mph, the A. Rafe and S. Das: Preprint submitted to Elsevier

Page 24 of 45

Certified crash-severity prediction

rural share rises from 0.255 to 0.571, the motorcycle-driver share rises from 0.002 to 0.196, and median annual daily traffic falls from 34,378 to 12,336. The severe end of the KABCO scale is not a random subsample of the mild end. It is a structurally distinct population, and it is small: 15,774 fatal records, or 0.30% of the sample. Table 3 Composition of the analysis sample. Panel A. Driver records by year and reported outcome Year

O

C

B

A

K

Total

2017 2018 2019 2020a 2021 2022 2023 2024 2025

474,174 483,410 498,230 413,477 483,262 485,983 489,323 488,537 474,709

60,193 61,957 64,451 52,313 55,335 53,150 52,352 50,787 48,087

34,815 32,384 32,108 28,680 36,570 40,110 42,525 44,236 45,282

7,965 6,891 7,123 6,939 9,056 8,759 8,644 8,319 8,011

1,680 1,650 1,599 1,733 1,975 1,871 1,823 1,806 1,637

578,827 586,292 603,511 503,142 586,198 589,873 594,667 593,685 577,726

Total

4,291,105

498,625

336,710

71,707

15,774

5,213,921

Panel B. Covariate profile across the severity ladder

𝑛 mean speed limit (mph) rural share motorcycle-driver share median ADT (known)b mean driver age female share median vehicle age (yr) VIN undecoded geocoded

O

C

B

A

K

4,291,105 44.6 .255 .002 34,378 38.7 .396 8 .030 .926

498,625 46.0 .210 .020 35,376 39.9 .515 8 .025 .970

336,710 47.9 .304 .060 27,321 38.9 .444 9 .031 .972

71,707 51.3 .457 .190 16,842 38.9 .305 10 .043 .976

15,774 55.5 .571 .196 12,336 41.9 .207 11 .030 .989

Panel C. Feature blocks and completeness (Configuration P) Block

cont.

categ.

ind.

levels

crash context person vehicle (VIN)

3 1 4

10 5 11

0 0 1

80 33 148

geography

0

2

1

11

Total

8

28

2

272

continuous features (% missing) speed limit 0.00, ADT 48.65, hour 0.00 driver age 0.85 vehicle age 0.48, doors 26.07, displacement 4.63, engine power 50.99 – one-hot encoded at a 0.005 minimum frequency; rarer levels pooled

a The 2020 total is the lowest in the series. b ADT is 48.65% missing and engine power 50.99% missing; both enter the design matrix with median imputation and a missingness

indicator, so no record is dropped for incompleteness. c Panel B statistics for posted speed limit exclude the 244,965 records (4.70%) carrying the CRIS −1 sentinel, and those for vehicle

age exclude 18,335 records with a model year later than the crash year. Both values are sentinels rather than nulls, so Panel C’s missingness counts do not register them; the fitted models see the columns as stored. Motorcycle share is the CRIS-reported body style, which is populated for every record, rather than the VIN decode.

The same structure is visible in the raw data before any model is fit. Figure 2 maps the share of severe or fatal outcomes over the plane of hour of day and posted speed limit. Severe outcomes concentrate in a high-speed late-night region: cells at or above 55 mph between midnight and 04:00 pool to 5.8% against 0.9% on the daytime low-speed plateau, a sixfold contrast, with a peak cell of 7.4%. The figure describes reporting frequencies and carries no claim about why the contrast exists. It matters here for a narrower reason. A single calibration threshold computed over this plane is dominated by the plateau, because that is where almost all of the records are, and the ridge is where almost all of the severe outcomes are. Figure 3 shows the same separation jointly rather than marginally. The no-injury and severe-or-fatal groups occupy overlapping but visibly displaced regions of covariate space: the severe group sits ten miles per hour higher in posted A. Rafe and S. Das: Preprint submitted to Elsevier

Page 25 of 45

Certified crash-severity prediction

Figure 2: The severity landscape of Texas driver records, 2017–2025. Color gives the share of severe or fatal (K+A) outcomes in each hour by speed-limit cell, for the 4,877,108 records with posted limits between 15 and 85 mph; contours mark the 2% and 5% levels, and the colorbar tick marks the 1.68% K+A share of the full sample. Cells at or above 55 mph between 00:00 and 03:59 pool to 5.8%, against 0.9% for cells at or below 45 mph between 06:00 and 21:59. Cells with fewer than 500 records are masked.

limit and three years older in vehicle age, and carries a late-night tail nearly three times heavier. Driver age barely moves, which is worth stating plainly, because it shows the displacement is a property of where and in what people crash rather than of who they are. The two distributions overlap substantially, and that overlap is the difficulty. They are neither separable, which would make the problem easy, nor identical, which would make one threshold sufficient.

5. Results The experiments follow the order of the guarantees in Section 3, and each one tests a theorem rather than a leaderboard position. Every number traces to a stored result file computed on the frozen snapshot under seed 20260704. The acceptance rule, fixed before the experiments were run, is that empirical coverage must reach the target less three binomial standard errors; at these sample sizes that standard error is about 0.0004, so the rule is strict. One framing point governs the section. The contribution is the certification layer, not the base model. Where a base model wins on set size, that is a fact about the model and not evidence for the framework; where the guarantee holds regardless of which model sits underneath, that is the theorem doing its work.

5.1. Validity and model agnosticism Table 4 wraps seven base models in one unchanged certification layer: the ordered logit that has anchored severity analysis since McCullagh (1980), a multinomial logit, a latent-class ordered logit, a random-parameters ordered logit, gradient boosting, the DLCON network of Section 3, and the TabPFN tabular foundation model. Every set A. Rafe and S. Das: Preprint submitted to Elsevier

Page 26 of 45

Certified crash-severity prediction

Figure 3: Joint covariate geometry of no-injury (teal) and severe or fatal (coral) driver records. Filled contours hold 50% and 90% of each group, dashed lines mark the no-injury medians, and panel titles give the severe group’s median and interquartile range. The severe group sits at higher posted speed limits (median 55 against 45 mph) and in older vehicles (10 against 7 years), and carries a heavier late-night tail (20.8% of severe records fall between 00:00 and 05:59, against 7.8% of no-injury records), while driver age is nearly unchanged. All 84,419 complete-case severe records are shown against 150,000 no-injury records sampled at seed 20260704.

produced by every model at every level was contiguous, without exception. Twenty of the 21 model-by-𝛼 cells at 𝛼 ∈ {0.05, 0.10, 0.20} met the pre-registered acceptance criterion. The exception is worth reporting precisely, because it is the kind of result that is easier to absorb than to state. Gradient boosting at 𝛼 = 0.10 covers 0.8989 against a criterion of 0.8990, missing by 8 × 10−5 . The criterion is nominal coverage less three binomial standard errors of the test set, and it was fixed before the experiments were run, so the cell fails it and we do not revise the rule after seeing the number. What the rule does not budget for is the variability of the calibration draw itself. Split conformal’s guarantee is marginal over that draw, and for a calibration set of 808,594 records the resulting spread in realized coverage is of the same order as the test-set standard error the rule uses. The shortfall is a small fraction of that unmodeled term, so this is a miss against a criterion that is tighter than the theory requires, not evidence against Proposition 3.7. It is reported because a framework whose subject is honest certification cannot round its own acceptance test in its favor. The table separates into two columns that behave very differently. Coverage is indistinguishable across all seven models, which is what Proposition 3.7 asserts: validity is a property of the calibration step and cannot be improved by a better base model. Efficiency is where the models separate, and the separation deserves to be stated plainly rather than dressed up. Mean set width ranges from 1.51 to 3.41 categories, and a linear ordered logit of the kind specified A. Rafe and S. Das: Preprint submitted to Elsevier

Page 27 of 45

Certified crash-severity prediction Table 4 One certification layer over seven base models (E1, 𝛼 = 0.10). Base model

Coverage

Mean width

Fit

Predict

Calibrate

Ordered logit Multinomial logit Latent-class ordered logit Random-parameters ordered logit Gradient boosting (balanced) DLCON (ours) TabPFN v3

0.8998 0.8997 0.8998 0.8999 0.8989 0.9000 0.8992

1.54 1.53 1.53 1.54 3.41 1.52 1.55

6s 24 min 3 min 19 s 43 s 22 s 9s

2s 1s 3s 7s 14 s 0s 16 min

0.11 s 0.11 s 0.12 s 0.12 s 0.07 s 0.11 s 0.12 s

Held-out test records 𝑛 = 810, 082; nominal coverage 0.90. Empirical coverage spans 0.8989–0.9000 across all seven models, and every set is contiguous. 20 of 21 model-by-𝛼 cells at 𝛼 ∈ {0.05, 0.10, 0.20} met the pre-registered acceptance criterion of nominal less three binomial standard errors. The exception is Gradient boosting (balanced) at 𝛼 = 0.1 (coverage 0.8989 against a floor of 0.8990, short by 7.6e-05). The criterion is computed from the test-set standard error alone and therefore does not budget for the variability of the calibration draw, which for split conformal is of comparable size; the shortfall is well inside that unmodeled term. It is reported rather than absorbed. Calibration is the certification layer itself; fit and predict belong to the base model and are reported to show what the layer costs by comparison.

decades ago produces sets whose width is indistinguishable from a 2025 tabular foundation model. The balanced gradient-boosting model is the widest of the seven, because class balancing flattens its predicted distribution and the score of Definition 3.3 then sits deeper in both tails. None of this affects validity, and all of it affects usefulness. One qualification belongs with that table. The widths reported are those of the construction certified here, which fixes a threshold on the score of Definition 3.3, and they are not the narrowest sets obtainable from the same base model. Width in the ordinal setting has a set-construction component that is separable from the estimator, and Zhang et al. (2025) give sets of minimal length per instance given the base scores. That machinery is complementary to this paper rather than competing with it: any construction exporting contiguous sets can be certified by the layer developed here, and a narrower one would carry the same guarantees at less cost. What Table 4 establishes is that validity is invariant across base models, not that these widths are optimal.

5.2. Heterogeneity, and where marginal validity fails Marginal coverage is an average, and averages conceal. At nominal 0.90 the marginal predictor covers baseline drivers at 0.947, which is close to target and unremarkable. On the strata a safety agency exists to serve, it fails: unrestrained drivers are covered at 0.374, motorcyclists at 0.583, and rural high-speed records at 0.624 under the gradient boosting anchor. That anchor is the most extreme of the seven base models, and the size of the shortfall is a property of it rather than of the guarantee: the other six undercover the same two strata between 0.822 and 0.896. These coverages are binomial proportions on test-split strata of several thousand records or more, so their standard errors are below 0.01 and the shortfalls are structural rather than sampling noise. Read per reported KABCO category, the same pattern runs along the severity ladder, with the fatal category covered at 0.822 and the incapacitating category at 0.869 while the mild categories are overcovered at 0.952 and 0.958. The unrestrained stratum is defined narrowly and deliberately. CRIS records restraint use in several codes, and only one of them, “None,” states that a restraint was available and not used; the neighboring “Not Applicable” code is predominantly the motorcycle code, since 8,530 of its 10,518 test-split records are motorcyclists and are counted in that stratum instead. Defining the stratum as “None” alone leaves 10,335 records and is what the numbers above report. A looser definition folding in the inapplicable and residual codes would raise the coverage to 0.417, so the narrow definition is both the honest one and the less flattering one. This is not a defect of the base model and fitting a better one will not remove it. It is the behavior Theorem 3.13(iii) predicts. A single threshold calibrated on a population that is 82% uninjured is calibrated for that population, and the severe strata are a small minority paying for the majority’s threshold. The practical reading is uncomfortable and worth stating directly: a marginally valid severity model deployed for screening would be least reliable on exactly the motorcyclists and unrestrained drivers that a screening system is built to find. Conditioning repairs it, and repairs it regardless of where the partition came from. Class-conditional calibration on the declared partition restores every cell to 0.8982–0.9002, and the eight cells of a generic KMeans partition all land within 0.8980–0.9016. That the two are indistinguishable on validity is not a disappointment; it is Theorem 3.11 behaving as stated, since the theorem holds for any partition fixed in advance and arbitrariness therefore cannot break it. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 28 of 45

Certified crash-severity prediction

The framework is indifferent to a partition’s provenance, then, but the deployer should not be, because efficiency is where partitions separate and the separation favors the econometric construct. At matched validity, the latent-class ordered logit’s own gate gives the narrowest sets of any partition tested, at a mean width of 3.236 categories against 3.356 for the declared safety strata, 3.402 for a generic KMeans clustering, and 3.406 for the deep network’s gate. The classes it recovers are also legible rather than incidental: their widths run from 2.44 to 4.04 categories while each holds coverage within a few thousandths of nominal, which is to say the model has separated the population into subpopulations that genuinely differ in how predictable their severity is, and the certification layer then prices each one honestly. This is the sense in which the heterogeneity program earns its place here. Its classes are not required for validity, and the paper does not pretend otherwise. They are the most efficient partition on offer, and they are the only one that arrives with an interpretation attached. A clustering fit for geometric compactness and a neural gate fit for predictive loss both land near the marginal width, because neither was built to separate the score distribution; the latent-class model was, and two decades of severity econometrics is the reason it knows how. Efficiency in the aggregate is not the same as repair where it matters, and Section 5.3 shows the two can point in different directions. The partition that closes the deficit on the severity ridge is the one declared on the ridge’s own coordinates, not the one that minimizes width overall. A deployer therefore has two distinct questions, and they have two distinct answers: latent classes for the narrowest sets across the board, and declared bands for coverage in a region one has decided to protect. The repair is not free, and the price is the honest part of the result. The motorcyclist stratum’s mean width rises from 2.67 to 4.22 categories and the unrestrained stratum’s from 2.80 to 4.59. Those sets were narrow because they were wrong. Covering the truth at the stated rate requires them to be wide, and a set that says “B or worse” honestly is worth more than a narrow set that is silently wrong four times in ten. Class-conditional calibration redistributes width toward the records that need it rather than inflating it everywhere: it changes the width of 18.7% of records, by one or two categories, for a mean change of −0.006, and validity never depends on the partition geometry (Theorem 3.11).

5.3. The deficit, and what closes it Figure 4 carries the heterogeneity result. It maps the deviation of empirical coverage from nominal over the plane of Figure 2 in three maps. Under marginal calibration the deficit is not diffuse: it sits on the severity ridge, reaching 0.338 in the deepest cell with adequate support, the 75-mph bin at 03:00. A generic declared partition narrows it, lifting coverage at 75 mph from 0.611 to 0.692, but it cannot close a slice it does not encode, and exact conditional coverage everywhere is unattainable by Proposition 3.10. A partition declared on the plane’s own coordinates closes the ridge, taking the deepest cell to 0.847 and pooling to 0.8979. Two features of that figure qualify the claim rather than support it, and both are worth the space. First, the declared partition’s residual weakness has structure. Seven cells with adequate support still fall below 0.80, and they sit exactly where the declaration is coarse: the 50-mph night cells, and the day-band edge hours 06:00 and 21:00, where the hour-wall minimum is 0.825. The partition guarantees its bands, and only its bands. Second, the generic KMeans-8 partition does nothing at all for the night deficit. Along the hour axis its curve lies on top of the marginal curve, and its entire benefit is realized on the speed axis. A partition helps where it encodes the structure that matters and nowhere else, which is the practical content of Theorem 3.11.

5.4. Coverage of the true label under reporting noise Police-reported severity is not medically assessed severity, and Section 3.6.1 sets out what the linkage literature reports about the gap. Because true severity is unobserved in CRIS, the guarantee of Theorem 3.31 is audited semisynthetically: a known banded misclassification process is injected into held-out labels, calibration sees only the corrupted labels, and coverage is evaluated against the labels actually recorded. This is the standard design for noisecorrection work, and it is stated plainly rather than presented as field evidence. Table 5 reports the sweep. The band-expanded sets cover the true label at or above the guaranteed floor 1 − 𝛼 − 𝛿 at every 𝛿, with a comfortable margin. Three readings of that table matter, and two of them cut against the result. The first is the comparison the theory predicts. The floor available without band structure is 1−𝛼 −𝜀tot , about 0.767 for this injected process, against a banded floor running from 0.90 down to 0.80. The gap is real but it narrows as 𝛿 grows, from thirteen points at 𝛿 = 0 to under four at 𝛿 = 0.10. The vacuousness that motivates Theorem 3.31 is a claim about field magnitudes, where the linkage literature puts 𝜀tot near 0.49 and the beyond-band mass near 0.02, giving a A. Rafe and S. Das: Preprint submitted to Elsevier

Page 29 of 45

Certified crash-severity prediction

Figure 4: The coverage deficit and what closes it. Per-cell deviation of empirical coverage from nominal 0.90 over the plane of Fig. 2 (random-split test fold, 754,186 records, gradient-boosting base; 82 of 336 low-support cells hatched). (a) Marginal conformal fails on the severity ridge, reaching 0.338 in the 75-mph, 03:00 cell. (b) A generic declared partition narrows the 75-mph deficit to 0.692 but cannot close a slice it does not encode (Proposition 3.10). (c) A partition on the plane’s own coordinates closes the ridge (0.847 in that cell, 0.8979 pooled); the residue moves to the coarse cells, with an hour-wall minimum of 0.825. (d,e) The same comparison along each axis.

Table 5 Guaranteed and empirical true-label coverage under banded reporting noise (E4, 𝛼 = 0.10). Guaranteed floor 𝛿 0.00 0.01 0.02 0.05 0.10

Empirical coverage on 𝑌

Banded

Generica

Constant 𝑏=1

Category-dep.

No expansionb

0.90 0.89 0.88 0.85 0.80

0.7673 0.7670 0.7659 0.7644 0.7617

0.9870 0.9870 0.9870 0.9871 0.9872

0.9882 0.9882 0.9883 0.9883 0.9884

0.8970 0.8973 0.8974 0.8980 0.8990

a 1−𝛼−𝜀

tot , the floor available without band structure, where 𝜀tot = 0.1327–0.1383 is the realized error rate of the injected process. This is a property of the injection design and not the quantity Theorem 3.31 tracks, which is 𝛿. b Sets left unexpanded. This arm carries no guarantee; its coverage of 0.8970–0.8990 at mean width 3.40 is a property of this noise level, not a property the theory provides.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 30 of 45

Certified crash-severity prediction

generic floor near 0.41 against a banded floor of 0.88. It is not a claim about this injection, whose realized error rate is 0.133 and whose generic floor is therefore weak rather than vacuous. The distinction matters because Theorem 3.31’s floor is indexed by 𝛿 and not by 𝜀tot , and the swept grid brackets the literature’s 𝛿, which is the quantity the theorem tracks. The second is the price. Expanded coverage sits near 0.987, some nine points above nominal, at a mean width of 4.42 of five categories. That is conservative by construction: the constant band pays for worst-case noise within the band while the injected process is not adversarial, and Theorem 3.33 is tight only in the worst case over laws satisfying the band, not on any particular law. The category-dependent map of Remark 3.38 is the natural candidate for recovering that width, and Section 5.8 reports that on these data it does not: it is narrower only on the exact-fatal singleton, and wider wherever the A category reaches two categories down, so the two maps are empirically indistinguishable here. The width is not the band’s doing, and no choice of band recovers it. The third reading concerns the sweep itself, and it is a limitation of the experiment rather than of the theorem. The declared 𝛿 is the beyond-band mass, but the injection can place beyond-band errors only on the B and A categories, which are 7.5% of records, so a declared 𝛿 of 0.10 realizes a true beyond-band mass near 0.0075. We therefore also report a corrected arm in which the realized mass equals the declared 𝛿 exactly. In both arms the bound holds at every 𝛿, and in both the margin widens as 𝛿 grows rather than tightening, because injecting more noise inflates the calibration quantile and widens the base sets faster than the floor falls. The honest conclusion is that this sweep confirms the theorem’s arithmetic and cannot falsify it: Theorem 3.31 is a union bound, and no configuration of a process satisfying its assumption will violate it. A falsification with teeth is nonetheless available, but it targets the declared channel rather than the union bound, and it is the width-against-floor specification test reported later in this section. Note also that a declared 𝛿 above 0.075 is not achievable under a two-category-down design at all, which the corrected arm flags rather than silently saturating. The third is a result a reader will find, so we state it first. The arm that ignores the noise entirely, applying no expansion at all, still covers the true label at 0.897–0.899 at a mean width of 3.40. At this noise level the expansion is not needed to achieve coverage. What it is needed for is to know that coverage holds. The unexpanded arm’s success is a property of this configuration, in which noise enters only the calibration labels and inflates the quantile in a way that happens to compensate. Nothing in the theory protects it, and nothing protects it at the field magnitudes above. The expansion converts an empirical accident into a floor that can be stated in advance, and the extra width is the visible price of that certificate. That trade is the subject of the paper, and it is a trade a deployment is entitled to refuse.

5.5. Transfer across time and jurisdiction Table 6 deploys a calibration frozen on 2022–23 onto two subsequent years. Unweighted split conformal holds at state scale, covering 0.8993 in 2024 and 0.9008 in 2025, and that is itself a finding: across these years the covariate drift is mild enough at the scale of a whole state that marginal validity survives without correction. Both recovery methods hold as well. The instructive column is the last. The density-ratio certificates report a total-variation slack bound near 0.35, which says honestly that the parametric weight family does not capture the drift from 2017–21 to 2024–25. The certificate is not failing. The certificate is the mechanism that says so. Table 6 Temporal transfer to two unseen deployment years (E5, 𝛼 = 0.10). Coverage

Mean width

Method

2024

2025

2024

2025

TV-slack LCB

Unweighted split conformal County-Mondrian (Thm. 3.44) Density-ratio weighted (Thm. 3.45)

0.8993 0.9003 0.9024

0.9008 0.9023 0.9048

3.439 3.412 3.453

3.438 3.415 3.462

– – 0.348 / 0.353

Calibrated on 2022–23, deployed on 593,685 records in 2024 and 577,726 in 2025; binomial standard error ≈ 0.0004. The slack bound is one-sided: it certifies weakness, never strength.

Space is where the absence of a certificate becomes dangerous, and Figure 5 makes the case. Across 54 held-out counties, per-county coverage under the frozen state calibration ranges from 0.372 to 0.936. A state-level model that is valid on average is, county by county, anywhere between useless and conservative, and nothing in the marginal guarantee tells a county which one it has. Twenty-nine of the 54 met the sample-size requirement to earn a transfer A. Rafe and S. Das: Preprint submitted to Elsevier

Page 31 of 45

Certified crash-severity prediction

certificate; for those, density-ratio weighting restores weighted coverage to 0.872–0.922, and the exported tuple bounds the slack from below.

Figure 5: Spatial transfer certificates. Each of the 54 held-out counties is filled by its raw unweighted coverage under the frozen state calibration (0.372 to 0.936 at nominal 0.90); gray counties lie outside the hold-out design. The 29 ringed counties earned a certificate: density-ratio weighting restores weighted coverage to 0.872–0.922, with the total-variation slack bounded from below (0.000 to 0.097). Fill is raw coverage while the certificate is on the weighted transfer (Ward, inset: raw 0.741, weighted 0.911). Eight of the 29 fall below their best-case floors, which the one-sided certificate permits (Theorem 3.45).

The honesty of that figure is the point of it. A certified county can still render coral, because the fill is raw unweighted coverage while the certificate is issued on the weighted transfer. Eight of the 29 certified counties fall below their bestcase floors empirically, which the one-sided certificate explicitly permits. The floor is the best case the evidence allows, not a performance bound. It certifies weakness and never strength, and a framework claiming otherwise from unlabeled data would be claiming something provably unavailable. Two of those eight sit within 0.006 of their floors on roughly 1,200 rows apiece, which is inside sampling error; the certificate is a one-sided diagnostic and a county falling below it is evidence to investigate rather than proof of failure. One point of experimental discipline belongs here rather than in a footnote, because it constrains what the weighted arm may claim. Theorem 3.45 requires the density ratio to be fixed given the training split and the unlabeled target covariates, which means it may not be a function of the point being scored. The weighted arm therefore fits its ratio on one half of each target jurisdiction’s unlabeled records and predicts and evaluates only on the other half, and the A. Rafe and S. Das: Preprint submitted to Elsevier

Page 32 of 45

Certified crash-severity prediction

reported weighted coverages rest on those held out rows alone. The unweighted and group-weighted arms fit no ratio and carry no such constraint, so they are evaluated on the full population; because the halves are drawn uniformly at random, the arms estimate the same population coverage and only the weighted arm pays in variance.

5.6. Certificates in deployment A certificate issued once is not a deployment practice. Figure 6 shows the shipped monitoring apparatus over two years of held-out stream. Monthly coverage stays within 0.009 of nominal with a two-year mean of 0.900, and the quantiles of the conformal p-values depart from uniformity by at most 0.021.

Figure 6: Certificates in deployment. (a) Monthly volume and severe share over 2017–2025 with the frozen temporal splits. (b) Monthly quantiles of the conformal p-values depart from uniformity by at most 0.021 (±2 SE shading). (c1) The monthly-restarted test martingale (nominal per-month false-alarm rate 0.01; exact under the online protocol, approximate under the frozen calibration used here) alarms in three isolated months, against pooled and month-matched calibration. (c2) Monthly coverage stays within 0.009 of nominal (two-year mean 0.900). (d) The 2025 state-level certificate as a waterfall: nominal, less the declared band 𝛿, less the total-variation slack bound, a best-case floor of 0.527 against observed 0.905.

The monitor’s own behavior is the most useful part of the figure, and it is not the reassuring result. The monthlyrestarted test martingale crosses its alarm in three isolated months. The natural explanation is seasonality, since the calibration pools all months, so the alarms were re-tested against a month-matched calibration drawn from the same calendar months of the calibration years. They fire again, and in November 2024 more strongly. The seasonality reading is refuted by its own pre-declared test, and the honest description is that three isolated months carry real distribution shifts, detected and reported. Coverage in those months deviates by at most 0.006, so the shifts were A. Rafe and S. Das: Preprint submitted to Elsevier

Page 33 of 45

Certified crash-severity prediction

detected without material damage. The instrument’s job is detection, not correction: an alarm mandates re-calibration and never retroactively edits a certificate already issued.

5.7. Risk control and the price of validity Coverage is not the only quantity a screening deployment must control, because omitting a fatality is not exchangeable with omitting a no-injury record. Figure 7 reports both risk dials against their nominal bounds. The fatal-omission rate saturates at 0.0019 across 2,475 fatalities in the test fold, far under any budget a deployer would set, and the severity-weighted cost risk saturates at 6.92 against bounds an order of magnitude larger. Over the range originally tested, 𝛽 ≥ 0.01, the risk constraint is slack: the threshold sits at zero and the escalation workload is constant at 0.281. That saturation is reported rather than hidden, and it is why the operating curves extend to much smaller 𝛽, where the constraint actually binds.

Figure 7: Operating CHOIR. (a,b) Cost-weighted false-omission risk and fatal-omission rate (𝑛fatal = 2,475) against their nominal bounds over a log-spaced 𝛽 grid, with the four originally tested points marked. Both guarantees saturate well below the bound for 𝛽 above roughly 0.01, the range a deployer would use. At the extreme low-𝛽 edge, where the threshold reaches 0.85 to 0.89, the empirical value briefly exceeds the bound on this single test-set realization by at most 5% relative, which is the finite-sample noise expected around a marginal guarantee (Theorem 3.46, Corollary 3.48) and not a violation. (c) The escalation workload implied by each 𝛽, saturating at 0.281 once the constraint goes slack. (d,e) Mean set width over the plane of Fig. 2 under marginal and declared calibration: certification moves width into the cells where Fig. 4(a) showed coverage falling short, and takes width away where coverage was already adequate.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 34 of 45

Certified crash-severity prediction

At that extreme low-𝛽 edge the empirical value briefly exceeds its bound on this single test-set realization, once for fatal omission and six times for cost risk, by at most 5% relative. This is expected and is not a violation. Theorem 3.46 and Corollary 3.48 control an expectation over the calibration draw rather than a deterministic per-realization quantity, and these excursions sit within sampling error at the achievability edge, where the threshold saturates near one. The alternative would have been to plot only the range in which the bound never crosses, which is range shopping. Panels (d) and (e) price the guarantee. Certification does not inflate width uniformly; it moves width to where coverage was missing. In cells where marginal coverage fell below 0.80 the declared partition spends an additional 0.61 categories, and in cells where marginal coverage was already adequate it spends 0.18 categories fewer. The blocky structure of panel (e) is the declared partition itself, legible in the price surface.

5.8. Composition and cost The guarantees are stated separately and deployed together, so their composition must be checked rather than assumed. Table 7 stacks class-conditional calibration, the band expansion and the stratum structure on the product cells of the class by rural-urban grid. Every cell clears the additive floor 1 − 𝛼 − 𝛿 = 0.88 that Theorem 3.50 predicts. It clears it by so wide a margin that the check is worth reading carefully rather than as a success. Table 7 Composed guarantees on product cells (E8, 𝛼 = 0.10, 𝛿 = 0.02). Class

Stratum

𝑛

Coverage on 𝑌

Mean width

% at full scale

Floor

baseline baseline motorcycle motorcycle rural highspeed unrestrained unrestrained

rural urban rural urban rural rural urban

106,778 592,634 2,924 5,606 91,805 4,594 5,741

0.9950 0.9976 1.0000 1.0000 0.9963 0.9998 0.9998

4.317 4.240 5.000 5.000 4.992 4.999 4.999

39.7 29.4 100.0 100.0 99.3 99.9 99.9

0.88 0.88 0.88 0.88 0.88 0.88 0.88

810,082

0.9971

4.353

40.3

0.88

All

Every cell clears the additive floor 1 − 𝛼 − 𝛿 = 0.88. The rural high-speed stratum is rural by definition, so of the 4 × 2 possible cells 7 are non-empty. The last column is the share of records whose set is the whole scale, and it is the column to read: coverage near 1.000 on the severe strata is the arithmetic of returning every category, which mean width understates. The category-dependent map of Remark 3.38 was tested as the efficient alternative and is indistinguishable here to six decimal places, because the base sets on these strata are already near-vacuous before any expansion (Section 5.8); the constraint is the base model, not the band.

On the severe strata the composed set is the entire severity scale. For rural motorcyclists, 99.97% of records receive all five categories; for unrestrained drivers in either setting the figure is above 99.9%. Coverage of 1.000 on those cells is not evidence, it is the arithmetic of returning everything, and the certified statement handed to an agency for a motorcyclist is that the injury lies somewhere between none and fatal. That is true and it is useless, and reporting mean width conceals it, which is why Table 7 reports the share of records receiving the full scale instead. The natural explanation is that the constant band is too conservative, and it is wrong. We tested it directly: the category-dependent map produces an identical full-scale share in every cell, to six decimal places, because it is narrower only on the exact-fatal singleton and wider wherever the A category reaches two down. The cause is upstream of the band. Run without any expansion, the base sets on these strata are already near-vacuous, at mean width 4.28 of five for rural motorcyclists with 28% at full scale, and 4.65 with 65% at full scale for urban unrestrained drivers. Any expansion of one category saturates them. The honest reading is therefore a dilemma rather than a triumph, and it is the paper’s sharpest empirical finding. On the strata that matter, one may have the marginally calibrated sets, which are narrower but undercover, or the composed sets of Table 7, which cover at 1.000 and say nothing. The two horns are not equally sensitive to the base model. The second is not: every one of the seven composes to width four or more on these strata. The first is: the gradient boosting anchor undercovers unrestrained drivers at 0.374 while carrying the widest sets of the seven at mean width 3.41, so it is not an instance of narrow and wrong but of wide and wrong, whereas the six models with narrow sets near width 1.51 undercover the same stratum between 0.822 and 0.896. The dilemma that survives all seven is therefore between validity and informativeness, not between narrowness and correctness. The composition theorem holds, and what it composes on these strata is uninformative. The constraint is not the certification layer, and no band and no partition repairs it; the next paragraph shows that a large part of it is not the base model either. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 35 of 45

Certified crash-severity prediction

How much of the width is unavoidable is a quantitative question, and Theorem 3.40 answers it without reference to any model. The channel floor 𝐴∗ (𝛼) lower-bounds the mean width of every valid predictor of the reported label on stratum 𝐴, and it depends only on the declared reporting channel and the stratum’s true severity composition. We compute it as an identification region in the sense of Molinari (2008), ranging the within-stratum channel over the constraints of Section 3.6.1 and admitting only channels that reproduce the observed reported composition. Table 8 reports the resulting certified lower bound against the smallest base width any of the seven models achieves, and Figure 8 shows the same decomposition. Two readings follow. First, the specification test: on no stratum does an achieved valid width fall below the certified floor at any 𝑅, so the declared channel is not refuted, and the low-severity baseline stratum, where reporting is most reliable and the floor sits near one category, binds most tightly and passes. The test’s power is itself stratum-dependent, and the pass means more where power is higher. On baseline a channel whose floor exceeded the achieved 1.40, only about 0.4 categories above the invariance value, would be refuted, so the pass there is informative; on the motorcyclist stratum only a channel far noisier than any the linkage evidence admits could be caught, so the pass there is weak. Achieved per-cell coverage runs between 0.8997 and 0.9004, a hair off nominal; recomputing the floor at the realized level rather than at 1 − 𝛼 moves it by about 10−3 of a category and changes no verdict. Second, the decomposition of Corollary 3.41 under an explicit selection model. Linkage may depend on the reported label given the truth, and we bound that dependence by a ratio 𝑅 of the linkage probability across reported labels at fixed true severity (Remark 3.43); Table 8 reports the certified floor at 𝑅 ∈ {1, 2, 5}. Because the certified floor 𝐴∗ is a lower bound on the width of any valid predictor, it splits the observed width into three parts with the quantifiers pointing the way the theorem licenses: one category that any nonempty set must carry, a channel-imposed excess of at − 𝐴∗ that a better model or richer record least 𝐴∗ − 1 that no base model removes, and a remainder of at most 𝑤ach 𝐴 could in principle recover but Theorem 3.25 shows cannot be certified distribution-free. On the motorcyclist stratum of roughly 3.55 categories, the channel-imposed excess is at least 0.64 under selection-invariance and at least 0.03 under a fivefold dependence, so the recoverable remainder is at most 1.91 to 2.52; on unrestrained drivers the channel excess is at least 0.36 under invariance. Where the truth sits between the certified floor and the observed width is bounded but not identified, which is the Molinari framing of Remark 3.43: a declared channel set gives a floor no model beats, not a point estimate of how much the channel imposes. One premise underwrites the whole computation and belongs where it is used: Theorem 3.40 requires the reporting channel to be nondifferential within the declared stratum. That premise is cleanest on the motorcyclist stratum, a homogeneous vehicle class whose riders plausibly share a single reporting regime, and most delicate on unrestrained drivers, where the defining covariate is itself a candidate to shift the channel and within-stratum nondifferentiality is a stronger assumption. The decomposition above leads with motorcyclists for that reason, and the unrestrained floors should be read as the more assumption-laden of the two. Table 8 Channel floor 𝐴∗ against achieved width, per stratum, 𝛼 = 0.10, at three strengths of reported-label linkage-selection dependence 𝑅 (Remark 3.43); 𝑅 = 1 is selection-invariance. Each floor is the certified column-concentration lower bound over the declared channel set. The achieved width, the smallest base (pre-expansion) Mondrian width among the seven models, exceeds every floor, so the declared channel is not refuted at any 𝑅. Stratum

Achieved base width

𝑅=1

𝑅=2

𝑅=5

1.40 1.95 3.58 3.55

0.97 1.00 1.36 1.64

0.93 0.95 1.20 1.48

0.91 0.92 0.96 1.03

Baseline Rural high-speed Unrestrained Motorcyclist

A deployment wanting both validity and information would need a lower level, a richer covariate record, or a base model that separates these strata, and the channel floor caps what the last two can buy. The framework will certify whatever it is given, including the fact that what it was given is not enough, and now also how much of that shortfall no model can remove. The part of the shortfall that a richer record could recover, as opposed to the channel floor that none can, is itself measurable. A second declared feature configuration adds the fields an emergency responder observes at the scene but a dispatch-time model does not, ejection, airbag deployment, vehicle damage severity, the first harmful event, and the manner of collision; re-running the class-conditional calibration on it narrows the certified sets by a stratum-dependent amount while holding per-cell coverage at the nominal level. On the unrestrained stratum the triage fields recover about A. Rafe and S. Das: Preprint submitted to Elsevier

Page 36 of 45

Certified crash-severity prediction

Figure 8: Certified width decomposed per stratum at 𝛼 = 0.10, against the smallest base width the seven models achieve. Gray is the single category any nonempty set carries; coral is the model-independent channel floor, a certified lower bound on the excess the reporting channel imposes (solid to the heavy-selection floor 𝑅 = 5, hatched to the invariance floor 𝑅 = 1); the teal remainder is a certified upper bound on what a better model or record could recover, itself not certifiable distribution-free (Theorem 3.25).

half a category, from 3.77 to 3.19 for the ordered-logit base, because ejection is recorded there and separates severity strongly. On the motorcyclist stratum they recover nothing, from 3.61 to 3.64, because the fields that carry the signal for an enclosed occupant, restraint and airbag and ejection relative to a compartment, do not apply to a rider and the record does not hold the missing information. The width above the channel floor on that stratum is therefore not an artifact of a thin feature set that a triage deployment would repair; it is a covariate gap that the richest scene-time record available does not close, bounded above by Theorem 3.25 and not closed from below by the data. One qualification belongs here: the CRIS snapshot does not record helmet use, a dominant determinant of rider severity, so the claim is that the fields present in the record do not recover the width, not that no conceivable field could; a deployment that captured helmet status might narrow the recoverable remainder, though never below the channel floor. The certification layer reports which of the two situations a stratum is in, which is the operational content an agency deciding what to record would want. Cost is the last question, and the layer is close to free. Certification is a sort and a quantile lookup, so it is 𝑂(𝑛 log 𝑛) with a small constant. Measured on synthetic conditional distributions at increasing scale, on a 24-core desktop processor (Intel Core i9-13900K, 3.0 GHz base, 32 GB) with the machine otherwise idle, the layer costs 0.21 s at one million rows and 1.09 s at five million, and the measured times scale linearly to within a few percent across the decade from 105 to 5 × 106 . Set against base models that take minutes to hours to fit, and against a prediction pass of roughly fifteen minutes for the foundation model, the certification layer costs about a second. The timings in Table 4 span two machines and should be read as such: the five statistical and gradient-boosting models were fit on that processor, while DLCON and TabPFN were fit and applied on a GPU. That asymmetry does not affect the point, which is one of order of magnitude rather than of benchmark. Whatever else is true of the framework, its cost is not a reason to decline it.

6. Discussion 6.1. What the guarantees do not mean The value of a guarantee lies as much in its boundary as in its content, and the boundaries here are sharp enough to state without hedging. Nothing in this paper is a causal statement. Every result is a statement about the coverage of a prediction set or the expectation of a risk functional, conditional on the crash-reporting process that generated the data. When Section 4.4 records that the fatal category is associated with higher posted speed limits, that is a description of reporting frequencies in CRIS and not a claim that speed limits produce fatalities. The distinction is not pedantic. A certified prediction set answers “given what was recorded about this crash, which severities are plausible at level 1 − 𝛼”, and it is silent on A. Rafe and S. Das: Preprint submitted to Elsevier

Page 37 of 45

Certified crash-severity prediction

what would have happened under an intervention. The severity literature has spent considerable effort on identification (Mannering and Bhat, 2014; Mannering et al., 2016), and that effort is orthogonal to this one rather than superseded by it. A causal analysis and a certification layer compose; neither substitutes for the other. Marginal validity is not conditional validity, and the gap is not a technicality that better estimation closes. Proposition 3.10 makes the impossibility formal (Vovk, 2012; Foygel Barber et al., 2021): no procedure achieves exact covariate-conditional coverage over finitely many categories without returning trivial sets. Every conditional claim here is therefore made with respect to a partition declared in advance, and Section 5.3 shows what that costs in practice. The declared partition closes the ridge it encodes and leaves residual pockets exactly where its resolution runs out. A reader should take from this that declaring the right groups is a modeling responsibility the framework does not discharge, and that a partition is a hypothesis about where heterogeneity lives, not a guarantee that it has been found. The transfer certificate is the boundary most likely to be misread, so it is worth being blunt. The exported floor is not a lower bound on realized coverage in the target jurisdiction. It is the best case the available evidence permits, and Section 5.5 shows eight of 29 certified counties falling below their floors empirically, which the one-sided construction explicitly allows. This is not a defect of the estimator. An assumption-free finite-sample upper bound on total variation from unlabeled data alone is unavailable, so a certificate that claimed to lower-bound performance would be claiming something that cannot be delivered. What the certificate delivers is the ability to detect and quantify weakness, and a large reported slack, as in the temporal experiment’s 0.35, is the mechanism functioning rather than failing. Finally, the drift monitor detects and does not correct. Section 5.6 reports three genuine alarms, and the correct response to an alarm is re-calibration, not a retroactive edit to a certificate already issued. Concept drift, as opposed to covariate shift, is not correctable without labels, and labels arrive with the delay of the reporting process itself.

6.2. A deployment playbook The framework is designed to be operated, and the experiments map onto a concrete sequence an agency can follow. The first decision is the reporting configuration, and it is a decision about when the prediction is made rather than about accuracy. Configuration P uses only pre-crash knowable fields and is the setting for screening and countermeasure prioritization. Configuration T adds what an investigator observes on arrival and is the setting for emergency response, where those fields exist at decision time. Reporting the two separately is what keeps the leakage question from becoming an argument: the fields are not leakage in the setting where they are observed, and they are unavailable in the setting where they would be. The second decision is the partition, and Section 5.3 is the guidance. Validity holds for any partition fixed in advance, so the choice is an efficiency and relevance decision rather than a validity one. The lesson from the repair sequence is that a partition protects what it encodes. A generic clustering bought nothing along the axis it did not represent, while bands declared on the operationally meaningful coordinates closed the deficit where severity actually lives. An agency that cares about motorcyclists should declare motorcyclists, and should expect wider sets there rather than treating the widening as a defect. The third decision is the band, and it is a sensitivity input rather than an estimate. The linkage literature bounds the plausible range (Burdett et al., 2015, 2022; Taylor et al., 2024), and the discipline is to sweep 𝛿 and report the family of guarantees rather than to estimate a confusion kernel that will not transfer across agencies or eras. Where fatality recording can be treated as exact (Farmer, 2003; Compton, 2005), the category-dependent map of Remark 3.38 keeps the K boundary from inflating, which is where most of the practical cost of expansion would otherwise fall. The fourth decision is the operating point. Figure 7 is the dial: a budget 𝛽 implies an escalation workload, and the workload saturates at 0.281 for any 𝛽 above roughly 0.01. A deployer choosing between screening intensities is choosing a point on that curve, and the curve, not an accuracy metric, is the quantity that translates into staffing. The last step is the one that distinguishes an audited deployment from an unaudited one. Certificates are exported per stratum, the monitor runs on the deployment stream, and both are artifacts rather than reports. The composition experiment shows the slack budget is additive and attributable, so when a certified floor is low the deployer can read which term consumed it. That accountability, rather than any coverage number, is what a procurement process can actually check.

6.3. Relation to generic conformal tooling A natural question is what this framework offers over a mature general-purpose conformal library. Table 9 answers it on the terms that make the comparison meaningful, and the answer requires reading the table correctly.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 38 of 45

Certified crash-severity prediction

Coverage is a constraint to satisfy, not a score to maximize. All three methods target 0.90 and all three attain it, and a method reporting 0.899 is not better than one reporting 0.895 when the standard error is near 0.007. Set size is where a comparison could bite, and it does not: the sets are the same size. The differentiators are the last three columns, and they are the research questions of this paper. Generic toolkits give neither contiguity by construction, nor coverage on the unobserved true label, nor a fatal-omission bound, because they were never designed to. They are excellent at what they do, and what they do is the marginal guarantee that Section 5.1 shows is insufficient on exactly the strata that matter. One caveat belongs in the open. This comparison runs on the small synthetic dataset shipped with the package so that the examples execute without the restricted CRIS data. It is a package sanity check, not evidence for any research question. The evidence is Section 5, on 5.2 million real records. Table 9 Guarantees available from generic conformal tooling. Method

Coverage

Set size

Contiguous

True-label

Fatal-omission

CHOIR MAPIE crepes

0.895 0.899 0.899

2.19 2.22 2.22

by construction 99% 99%

yes no no

yes no no

Computed on the synthetic demonstration dataset shipped with the package, not on CRIS: this is a sanity check on the implementation, not evidence for any research question. All three methods target 0.90 and all three attain it; the binomial standard error at this sample size is about 0.007, so the coverage column separates nothing. The non-contiguous sets the generic tools return about 1% of the time are not operationally meaningful as “B or worse” statements.

6.4. Anticipated objections Four objections are worth answering directly rather than leaving to inference. Conformal prediction is not new. Correct, and the paper never claims otherwise. Split conformal prediction, Mondrian conditioning, weighted conformal prediction under covariate shift and conformal risk control are all established (Vovk et al., 2005; Lei et al., 2018; Tibshirani et al., 2019; Angelopoulos et al., 2024), and each is labeled [KNOWN] or [ADAPTED] where it appears. What is new is the set of statements that use the structure of the severity problem: transfer of coverage to the unobserved true label from a declared support band, the oracle and diagnosis results for heterogeneity classes, the label-free one-sided transfer certificate, and cost-inflated risk control under banded reporting noise. The special issue asks for mathematical foundations for transportation, and results that exploit ordinality, KABCO reporting structure and jurisdictional hierarchy are transportation mathematics rather than an application of a recipe. The latent classes are arbitrary. They are, and that is the point. Theorem 3.11 holds for any partition fixed on the training split, so arbitrariness cannot break validity; it can only cost efficiency. Section 5.2 reports both a declared partition and a learned KMeans partition, and both restore per-cell coverage. The classical latent-class specification is included precisely so that the comparison is against the severity literature’s own construct rather than a strawman. Texas only, so what generalizes? The theory of when transfer holds is the generalization, and it is the content of Section 3.8 rather than an aspiration. The spatial experiment does not claim that Texas counties resemble other states’ jurisdictions; it holds out 54 counties precisely to simulate the jurisdictions a calibration has never seen, and reports that uncertified transfer ranges from 0.372 to 0.936. The transportable object is the certificate, which any jurisdiction can compute on its own unlabeled covariates. Fatalities are only 0.3% of records. They are, which is why the fatal-omission guarantee is stated separately from coverage and why 2,475 test fatalities support it. Rarity is an argument for per-class calibration and explicit risk control, not against them, and it is exactly the regime in which a marginal guarantee is least informative.

6.5. Scope and open problems Three matters are open in ways that shape what should be done next rather than qualifying what was done here. The label-noise evidence is semi-synthetic and can be nothing else without linked medical records. Section 5.4 injects a known process because true severity is unobserved in CRIS, and the honest consequence is that the empirical check validates the theorem’s arithmetic rather than the field’s noise magnitudes. Linking CRIS to Texas trauma-registry records, as Taylor et al. (2024) did for North Carolina, would convert the sensitivity program into a measurement. A. Rafe and S. Das: Preprint submitted to Elsevier

Page 39 of 45

Certified crash-severity prediction

The composed sets are uninformative on the severe strata, and Section 5.8 shows why that is not the band’s fault. We had expected the constant band to be the culprit and the category-dependent map to be the remedy; the data refuted it, the two maps being indistinguishable because the base sets on those strata are already near-vacuous before any expansion. The open problem this exposes is not a better band but a base model that discriminates severity among rare, heterogeneous road users well enough that a valid set is also a narrow one. That is a modeling problem the certification layer cannot solve and, to its credit, cannot hide: the layer reports exactly how uninformative the underlying model is on the populations that matter. Separately, an elicitation protocol that a state agency could run to declare its own band, rather than importing magnitudes from a Wisconsin linkage, remains worth building. The covariate-shift premise is declared and only partially falsifiable. With even a small labeled batch from a target jurisdiction, conditional coverage becomes auditable; without labels, concept drift is detectable prospectively but not correctable. The monitor in Section 5.6 is the shipped response to that asymmetry, and its three alarms demonstrate both that it works and that a deployer must choose its granularity in advance. Choosing the monitor’s granularity after seeing which months alarm would be tuning to conclusion, which is why the month-matched comparison was declared before it was run and reported after it refuted the hypothesis it was built to test.

Conclusion Crash-severity models are deployed without guarantees, and the three features that make crash severity statistically distinctive are the same three that make an off-the-shelf guarantee wrong. This paper has developed a certification layer that wraps any severity model without modifying it and attaches guarantees that use that structure rather than ignoring it. The guarantees are distribution-free in the base model, which is the claim that matters for deployment: no result below depends on the fitted model being correct. They are not assumption-free, and the assumptions are named where they are used. Marginal and class-conditional validity hold in finite samples under exchangeability alone. The oracle characterization is asymptotic and requires regularity. The noise and transfer results are free of a noise kernel and of a shift model respectively, which is their point, but each rests on a declared structural premise. The contribution is a characterization of what certification can and cannot deliver for this outcome, assembled as a calculus rather than resting on any single theorem. On the side of what is possible are five families of guarantees that hold at once with an additive and attributable slack budget, so a deployer can read which term consumed a certified floor. Ordinality yields prediction sets that are contiguous by construction, so the output is the operationally meaningful “B or worse” rather than an arbitrary subset of the scale, and the noise expansion costs a fixed number of categories rather than a multiple of the set’s size. Conditioning on any partition fixed in advance restores per-class coverage, with an oracle characterization of the efficiency this buys and a diagnosis of why marginal calibration undercovers precisely the severe strata. A declared ordinal compatibility band transfers coverage to the unobserved true severity at a floor that degrades only by the declared beyond-band mass, and a sharpness result shows the additive loss is unimprovable from support information alone. A declared covariate-shift premise transfers coverage across jurisdictions and years, exporting an honest one-sided certificate where it cannot be validated. Severity-weighted omission risk is controlled, with no band slack where fatality recording is exact. On the side of what is not possible are two results that bound any procedure rather than any model. The informativeness of a certified set is governed by a level-set functional of the true law that no base model can evade, and that functional cannot be lower-bounded from finite data without assumptions. The reporting channel supplies the one assumption the domain affords for free: given a declared misreporting channel identified from record-linkage constants, the floor admits a computable lower bound, so what is impossible distribution-free becomes possible given a declared measurement model. This pairing is the paper’s identification contribution, and it converts part of the empirical vacuity of the vulnerable strata from an apparent failure of the base model into a derived quantity: on the motorcyclist stratum a model-independent share of the certified width is imposed by the reporting channel, removable by no base model and by no covariate uninformative about the reporting process. The honest summary is that no single result here is deep. Each rests on a weak and declarable premise rather than a heroic one, and several are short by design; the value is in the calculus and the identification, in assembling the premises an ordinal safety-critical outcome actually supports and pairing each possibility with its impossibility counterpart, not in the depth of any one theorem. Empirically, one unchanged layer wrapped seven base models spanning four decades of methodology, from the classical ordered logit to a tabular foundation model, and validity was identical across all of them while efficiency was not, which is the theorem’s prediction rather than a surprise. On 5.2 million Texas driver records the layer exposed a failure that a marginal guarantee conceals, namely that six of the seven models undercover unrestrained drivers and A. Rafe and S. Das: Preprint submitted to Elsevier

Page 40 of 45

Certified crash-severity prediction

motorcyclists while covering the population on average, that the size of the shortfall is governed by the base model and ranges from slight to severe across the seven, and that conditioning repairs every cell at a width cost the paper reports rather than hides. Certification cost seconds on top of models that cost minutes to hours to fit. The results carry limitations that are worth stating plainly, because each is a boundary of the evidence rather than of the theory. The label-noise guarantee is proved but its empirical check is semi-synthetic: true severity is unobserved in the crash records, so a known banded process was injected, and the experiment therefore validates the theorem’s arithmetic rather than the field’s noise magnitudes. Related to this, the constant band is conservative by construction and buys its guarantee at a width the category-dependent map would reduce. The transfer theorem assumes a covariateshift premise that unlabeled data cannot validate; concept drift is detectable prospectively but not correctable without labels, and the exported floor certifies weakness rather than strength, as eight of twenty-nine certified counties falling below their floors empirically makes concrete. The conditional guarantees hold for declared partitions only, and the repair sequence shows residual weakness surviving exactly where a declaration’s resolution runs out, so choosing the groups remains a modeling responsibility the framework does not discharge. Finally, the empirical work is one state’s records under one reporting regime, which is why the transportable object is the certificate rather than the coverage number. Each limitation points at its own next step. The semi-synthetic noise evidence would become a measurement by linking the crash records to a state trauma registry, which would convert the sensitivity sweep over the declared band into an elicited band with a defensible provenance, and would let the category-dependent map be calibrated to the jurisdiction that uses it rather than imported from another state’s linkage. The covariate-shift premise becomes auditable with even a small labeled batch from a target jurisdiction, which suggests a sampling design question worth its own treatment: how few labeled records from a new county are needed to turn a one-sided certificate into a two-sided one. The declared partition’s residual pockets invite the question of how to declare well, and the natural formulation is a design problem in which resolution is spent where the deficit is deepest subject to a minimum cell count. Beyond severity, nothing in the framework is specific to KABCO except the band and the cost vector, so the same layer applies wherever an ordered, safety-critical outcome is recorded by a human observer under a stable misreporting structure, and injury scales in adjacent transportation settings are the obvious next test.

Software. The framework is released as choircert on PyPI (imported as choir) under the MIT license, with documentation, the demonstration dataset used in Section 6.3, and theorem-level tests that reproduce each guarantee on simulated data. The certification layer is model-agnostic by construction: any estimator exporting a conditional CDF can be wrapped without modification.

Appendix A. Deferred proofs Proof of Theorem 3.13. (i) Class-𝑐 coverage at threshold 𝜆𝑐 is 𝐺𝑐 (𝜆𝑐 ); by (R1), 𝐺𝑐 (𝜆𝑐 ) ≥ 1 − 𝛼 ⟺ 𝜆𝑐 ≥ 𝑞𝑐 . ∑ Expected size 𝔼[|𝐶𝜆 (𝑋)| ∣ 𝑐̂ = 𝑐] = 𝑘 ℙ(𝑠(𝑋, 𝑘) ≤ 𝜆 ∣ 𝑐̂ = 𝑐) is non-decreasing in 𝜆; minimizing size subject to validity forces 𝜆𝑐 = 𝑞𝑐 classwise, and strict monotonicity of 𝐺𝑐 at 𝑞𝑐 gives uniqueness up to null sets. (ii) 𝑞̂𝑐 is the ⌈(1 − 𝛼)(𝑛𝑐 + 1)⌉-th order statistic of the within-class sample (i.i.d. in the i.i.d. case we invoke); convergence and the rate are classical (van der Vaart, 1998, Cor. 21.5). Size convergence: ∑ ( ) | 𝔼[|𝐶𝑞̂𝑐 (𝑋)| ∣ 𝑐̂ = 𝑐] − 𝔼[|𝐶𝑞𝑐 (𝑋)| ∣ 𝑐̂ = 𝑐] = ℙ 𝑞𝑐 ∧ 𝑞̂𝑐 < 𝑠(𝑋, 𝑘) ≤ 𝑞𝑐 ∨ 𝑞̂𝑐 | 𝑐̂ = 𝑐 → 0 (A.1) | 𝑘

by (R2) and 𝑞̂𝑐 → 𝑞𝑐 , then dominated convergence (|𝐶| ≤ 𝐾). (iii) The marginal empirical quantile converges to 𝑞mix by (R1); class-𝑐 coverage of the fixed threshold 𝑞mix is 𝐺𝑐 (𝑞mix ); monotonicity gives the dichotomy, and ∑ 𝑐 𝑝𝑐 𝐺𝑐 (𝑞mix ) = 1 − 𝛼 with all summands ≥ 1 − 𝛼 forces equality iff all 𝑞𝑐 coincide. Proof of Theorem 3.33. (i) Take 𝑋 degenerate and 𝑚 = 1, so that the band around 𝑚 is [1, 1 + 𝑏− ] and 𝐾 ≥ 𝑏− + 2 leaves at least one category above it. For small 𝜀 > 0 let 𝑌̃ put mass 1 − 𝛼 + 𝜀 on 𝑚 and the remaining mass 𝛼 − 𝜀 on 𝐾 (outside the band [1, 1 + 𝑏− ]), with 𝐹̂ chosen so that 𝑠(𝑥, 𝑚) < 𝑠(𝑥, 𝑘) for all 𝑘 ≠ 𝑚. Then ℙ(𝑆 = 𝑠(𝑥, 𝑚)) = 1 − 𝛼 + 𝜀, so for large 𝑛 the calibration quantile 𝑞̂ equals 𝑠(𝑥, 𝑚) with probability → 1 and 𝐶̃ = {𝑚}, whence 𝐶̃ ⊕ = [1, 𝑚 + 𝑏− ]. Couple 𝑌 to 𝑌̃ by: 𝑌 = 𝑌̃ except on an event of probability 𝛿 contained in {𝑌̃ = 𝑚}, where 𝑌 = 𝑚 + 𝑏− + 1 (possible since 𝛿 < 1 − 𝛼 < 1 − 𝛼 + 𝜀 and 𝑚 + 𝑏− + 1 ≤ 𝐾). N(𝑏− , 𝑏+ , 𝛿) holds by construction. Coverage decomposes as: the A. Rafe and S. Das: Preprint submitted to Elsevier

Page 41 of 45

Certified crash-severity prediction

unperturbed part of {𝑌̃ = 𝑚} (mass 1 − 𝛼 + 𝜀 − 𝛿) is covered; the perturbed event (mass 𝛿) has 𝑌 = 𝑚 + 𝑏− + 1 ∉ 𝐶̃ ⊕ ; and {𝑌̃ = 𝐾} (mass 𝛼 − 𝜀) has 𝑌 = 𝐾 ∉ 𝐶̃ ⊕ . Hence ℙ(𝑌 ∈ 𝐶̃ ⊕ ) → 1 − 𝛼 + 𝜀 − 𝛿; taking 𝜀 = 𝜂 gives the claim, and since 𝜂 > 0 is arbitrary no constant larger than 1 − 𝛼 − 𝛿 can hold uniformly over laws satisfying the assumptions. (ii) Upper trim. Take the construction of (i) with 𝑚 = 1 and 𝛿 = 0, and set 𝑌 = 𝑌̃ + 𝑏− a.s. on {𝑌̃ = 𝑚} (maximal under-reporting, allowed by N since 𝑌 − 𝑌̃ = 𝑏− ; feasible as 𝑚 + 𝑏− ≤ 𝐾). Whenever 𝐶̃ = {𝑚} the trimmed set is [1, 𝑚 + 𝑏− − 1], and 𝑌 = 𝑚 + 𝑏− ∉ [1, 𝑚 + 𝑏− − 1]. So on {𝑌̃ = 𝑚} (probability → 1 − 𝛼) the trimmed rule misses the truth; its true-label coverage is at most 𝛼 + 𝑜(1). Lower trim. Symmetrically anchor at 𝑚 = 𝐾 (feasible when 𝐾 ≥ 𝑏+ + 1, so that 𝑚 − 𝑏+ ≥ 1) and set 𝑌 = 𝑌̃ − 𝑏+ a.s. on {𝑌̃ = 𝑚} (maximal over-reporting); the trimmed set [𝑚 − 𝑏+ + 1, 𝐾] never contains 𝑌 = 𝑚 − 𝑏+ . In both cases only the miss on {𝑌̃ = 𝑚} is needed for the bound, so the placement of the residual mass is immaterial. Proof of Theorem 3.45. (i) Two steps. Step 1 (exactness under the estimated tilt). By weighted conformal (Tibshirani et al., 2019), calibration with weight function 𝑤̂ (fixed given tr ) is exactly valid if the test covariate is drawn from 𝑄𝑤̂ and the conditional of 𝑌̃ ∣ 𝑋 matches calibration: 𝑃𝑋∼𝑄𝑤̂ (𝑌̃ ∈ 𝐶𝑞̂𝑤̂ (𝑋)) ≥ 1 − 𝛼. Step 2 (change of test measure). The coverage event is 𝐸 = {𝑌̃𝑛+1 ∈ 𝐶𝑞̂𝑤̂ (𝑋𝑛+1 )}. Conditionally on calibration data, its probability under the two test-point laws 𝑃 ∗ = 𝑃𝑋∗ ⊗ 𝑃cal (𝑌̃ ∣ 𝑋) and 𝑄 = 𝑄𝑤̂ ⊗ 𝑃cal (𝑌̃ ∣ 𝑋) differs by at most 𝑑TV (𝑃 ∗ , 𝑄); since the joints share the conditional, 𝑑TV (𝑃 ∗ , 𝑄) = 12

| | ∗ |d𝑃 − d𝑄𝑤̂ | = 𝑑TV (𝑃𝑋∗ , 𝑄𝑤̂ ). | ∫ | 𝑋

(A.2)

Combine and average over calibration data. (ii) 2 ba(ℎ) − 1 = 𝑃𝑋∗ (ℎ=1) − 𝑄𝑤̂ (ℎ=1) ≤ sup𝐴 |𝑃𝑋∗ (𝐴) − 𝑄𝑤̂ (𝐴)| = 𝑑TV ; the binomial LCB is standard. (iii) Under realizability 𝑄𝑤𝜃 = 𝑃𝑋∗ and 0

(

𝑑TV 𝑄𝑤 ̂ , 𝑄𝑤𝜃 𝜃

0

)

𝑤𝜃 | |𝑤 ≤ 12 𝔼𝑃cal | 𝔼𝑤𝜃̂ − 𝔼𝑤 0 |, | 𝜃̂ 𝜃0 |

(A.3)

a Lipschitz functional of 𝜃̂ − 𝜃0 near 𝜃0 under dominated derivatives, with 𝜃̂ − 𝜃0 = 𝑂𝑝 ((𝑚 ∧ 𝑛′ )−1∕2 ) by standard M-estimation.

Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability The crash records analyzed in this study are not public. They were obtained from the Texas Crash Records Information System (CRIS), maintained by the Texas Department of Transportation, which distributes record-level extracts to researchers on request through the CRIS Query portal at https://cris.dot.state.tx.us/. Access requires a data request and acceptance of the department’s terms of use; the authors hold no redistribution right and therefore cannot supply the records directly. The vehicle attributes were decoded from reported vehicle identification numbers using the National Highway Traffic Safety Administration’s public vPIC service (https://vpic.nhtsa. dot.gov/api/), and the geographic reference files are public Census products. All personally identifying material, including vehicle identification numbers, license and registration identifiers, dates of birth, and free-text investigator narratives, was excluded at ingest and appears in no released artifact. Every step from the raw departmental files to the analysis snapshot is scripted and released, so a reader holding an equivalent CRIS extract can rebuild the snapshot and reproduce the results with a single command. A synthetic demonstration dataset carrying the same schema ships with the software package so that all examples, tests, and the comparison of Section 6.3 run without restricted data.

Code availability The certification framework is released as an open-source package under the MIT license. It is distributed as choircert on the Python Package Index and imported as choir; the source, documentation, and the full experimental A. Rafe and S. Das: Preprint submitted to Elsevier

Page 42 of 45

Certified crash-severity prediction

pipeline are at https://github.com/pozapas/choircert, and a citable archived release is deposited at Zenodo (DOI: 10.5281/zenodo.21434172). The release includes the theorem-level test suite, which reproduces each guarantee on simulated data, the data-construction pipeline, the experiment harness, and the scripts that generate every figure and table in this paper from stored result files. All randomness derives from a single declared seed (20260704).

Reproducibility statement Results were computed on a frozen analysis snapshot built by the released pipeline, whose row order is deterministic. Every cached model artifact carries a fingerprint of the split it was built from, and the loader refuses any artifact whose fingerprint does not match the current snapshot, so a stale cache raises an error rather than silently misaligning with the data. Each number reported in this paper traces to a stored result file rather than to a transcribed value, and the tables are generated from those files by script. Base models were fit on a sixteen-core laptop-class processor, with the exception of the DLCON network and the tabular foundation model, which were fit and applied on a GPU; the certification layer itself is CPU-only and its cost is reported in Section 5.8.

CRediT authorship contribution statement Amir Rafe: Conceptualization, Methodology, Software, Formal analysis, Writing – original draft, Visualization. Subasish Das: Conceptualization, Methodology, Supervision, Writing – review & editing.

References Angelopoulos, A.N., Bates, S., Fisch, A., Lei, L., Schuster, T., 2024. Conformal risk control, in: The Twelfth International Conference on Learning Representations. Bates, S., Candès, E., Lei, L., Romano, Y., Sesia, M., 2023. Testing for outliers with conformal p-values. The Annals of Statistics 51, 149–178. doi:10.1214/22-AOS2244. Bhattacharyya, A., Foygel Barber, R., 2026. Group-weighted conformal prediction. Electronic Journal of Statistics 20. doi:10.1214/26-EJS2506, arXiv:2401.17452. Bohlouli, R., Varghese, K.K., Gentile, G., Eldafrawi, M., 2025. Enhancing mode-choice models with conformal prediction: Uncertainty quantification and decision support using tree-based machine learning. Transport and Telecommunication Journal 26, 352–361. doi:10.2478/ ttj-2025-0027. Bortolotti, T., Wang, Y.X.R., Tong, X., Menafoglio, A., Vantini, S., Sesia, M., 2025. Noise-adaptive conformal classification with marginal coverage. arXiv preprint doi:10.48550/arXiv.2501.18060. Burdett, B., Bill, A., Noyce, D., 2022. Evaluation of law enforcement agency injury severity assessments. Transportation Research Record: Journal of the Transportation Research Board 2676, 246–255. doi:10.1177/03611981221086628. Burdett, B., Li, Z., Bill, A.R., Noyce, D.A., 2015. Accuracy of injury severity ratings on police crash reports. Transportation Research Record: Journal of the Transportation Research Board 2516, 58–67. doi:10.3141/2516-09. Burdett, B.A., 2014. Improving Accuracy of KABCO Injury Severity Assessment by Law Enforcement Officers. Master’s thesis. University of Wisconsin–Madison. Department of Civil and Environmental Engineering. Cauchois, M., Gupta, S., Ali, A., Duchi, J.C., 2022. Predictive inference with weak supervision. arXiv:2201.08315. journal of Machine Learning Research (2024). Chakraborty, S., Tyagi, C., Qiao, H., Guo, W., 2024. Distribution-free conformal prediction for ordinal classification, in: Proceedings of the Thirteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR. pp. 120–139. Cohen, Y., Goldberger, J., Tirer, T., 2025. Efficient conformal prediction for regression models under label noise. arXiv preprint doi:10.48550/ arXiv.2509.15120. Compton, C.P., 2005. Injury severity codes: A comparison of police injury codes and medical outcomes as determined by NASS CDS investigators. Journal of Safety Research 36, 483–484. doi:10.1016/j.jsr.2005.10.008. Correia, A.H.C., Louizos, C., 2025. Non-exchangeable conformal prediction with optimal transport: Tackling distribution shifts with unlabeled data, in: Advances in Neural Information Processing Systems. arXiv:2507.10425. Dayan, B., 2026. Conformal risk control for safety-critical wildfire evacuation mapping: A comparative study of tabular, spatial, and graph-based models. arXiv preprint doi:10.48550/arXiv.2603.22331. Ding, T., Angelopoulos, A.N., Bates, S., Jordan, M.I., Tibshirani, R.J., 2023. Class-conditional conformal prediction with many classes, in: Advances in Neural Information Processing Systems, pp. 64555–64576. doi:10.52202/075280-2817. Dunn, R., Wasserman, L., Ramdas, A., 2023. Distribution-free prediction sets for two-layer hierarchical models. Journal of the American Statistical Association 118, 2491–2502. doi:10.1080/01621459.2022.2060112. Einbinder, B.S., Feldman, S., Bates, S., Angelopoulos, A.N., Gendler, A., Romano, Y., 2024. Label noise robustness of conformal prediction. Journal of Machine Learning Research 25, 1–66. Farmer, C.M., 2003. Reliability of police-reported information for determining crash and injury severity. Traffic Injury Prevention 4, 38–44. doi:10.1080/15389580309855.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 43 of 45

Certified crash-severity prediction Follmann, D.A., Lambert, D., 1991. Identifiability of finite mixtures of logistic regression models. Journal of Statistical Planning and Inference 27, 375–381. doi:10.1016/0378-3758(91)90050-O. Foygel Barber, R., 2020. Is distribution-free inference possible for binary regression? Electronic Journal of Statistics 14, 3487–3524. doi:10.1214/ 20-EJS1749. Foygel Barber, R., Candès, E.J., Ramdas, A., Tibshirani, R.J., 2021. The limits of distribution-free conditional predictive inference. Information and Inference: A Journal of the IMA 10, 455–482. doi:10.1093/imaiai/iaaa017. Foygel Barber, R., Candès, E.J., Ramdas, A., Tibshirani, R.J., 2023. Conformal prediction beyond exchangeability. The Annals of Statistics 51, 816–845. doi:10.1214/23-AOS2276. Gibbs, I., Cherian, J.J., Candès, E.J., 2025. Conformal prediction with conditional guarantees. Journal of the Royal Statistical Society Series B: Statistical Methodology 87, 1100–1126. doi:10.1093/jrsssb/qkaf008. Grün, B., Leisch, F., 2008. Finite mixtures of generalized linear regression models, in: Shalabh, Heumann, C. (Eds.), Recent Advances in Linear Models and Related Areas. Springer, pp. 205–230. doi:10.1007/978-3-7908-2064-5_11. Guruswami, V., Vadhan, S., 2010. A lower bound on list size for list decoding. IEEE Transactions on Information Theory 56, 5681–5688. doi:10.1109/TIT.2010.2070170. Islam, M.R., Wang, D., Abdel-Aty, M., 2024. Calibrated confidence learning for large-scale real-time crash and severity prediction. npj Sustainable Mobility and Transport 1, 1. doi:10.1038/s44333-024-00001-9. Jiang, W., Tanner, M.A., 1999. On the identifiability of mixtures-of-experts. Neural Networks 12, 1253–1258. doi:10.1016/S0893-6080(99) 00066-0. Joshi, S., Kiyani, S., Pappas, G., Dobriban, E., Hassani, H., 2025. Conformal inference under high-dimensional covariate shifts via likelihood-ratio regularization. Advances in Neural Information Processing Systems 38. arXiv:2502.13030. Kiyani, S., Pappas, G.J., Hassani, H., 2024. Conformal prediction with learned features, in: Proceedings of the 41st International Conference on Machine Learning, PMLR. pp. 24749–24769. Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R.J., Wasserman, L., 2018. Distribution-free predictive inference for regression. Journal of the American Statistical Association 113, 1094–1111. doi:10.1080/01621459.2017.1307116. Lei, L., Candès, E.J., 2021. Conformal inference of counterfactuals and individual treatment effects. Journal of the Royal Statistical Society Series B: Statistical Methodology 83, 911–938. doi:10.1111/rssb.12445. Liu, W., de Paula, A., Tamer, E., 2025. Prediction sets and conformal inference with interval outcomes. arXiv:2501.10117. Lu, C., Angelopoulos, A.N., Pomerantz, S., 2022. Improving trustworthiness of AI disease severity rating in medical imaging with ordinal conformal prediction sets, in: Medical Image Computing and Computer Assisted Intervention – MICCAI 2022, Springer. pp. 545–554. doi:10.1007/978-3-031-16452-1_52. Mannering, F., 2018. Temporal instability and the analysis of highway accident data. Analytic Methods in Accident Research 17, 1–13. doi:10.1016/j.amar.2017.10.002. Mannering, F.L., Bhat, C.R., 2014. Analytic methods in accident research: Methodological frontier and future directions. Analytic Methods in Accident Research 1, 1–22. doi:10.1016/j.amar.2013.09.001. Mannering, F.L., Shankar, V., Bhat, C.R., 2016. Unobserved heterogeneity and the statistical analysis of highway accident data. Analytic Methods in Accident Research 11, 1–16. doi:10.1016/j.amar.2016.04.001. McCullagh, P., 1980. Regression models for ordinal data. Journal of the Royal Statistical Society: Series B (Methodological) 42, 109–127. doi:10.1111/j.2517-6161.1980.tb01109.x. Molinari, F., 2008. Partial identification of probability distributions with misclassified data. Journal of Econometrics 144, 81–117. doi:10.1016/ j.jeconom.2007.12.003. Patil, M., Ahmed, Q., Midlam-Mohler, S., 2024. Urban traffic forecasting with integrated travel time and data availability in a conformal graph neural network framework, in: 2024 IEEE 27th International Conference on Intelligent Transportation Systems, pp. 2482–2487. doi:10.1109/ ITSC58415.2024.10920016. Penso, C., Goldberger, J., Fetaya, E., 2025. Conformal prediction of classifiers with many classes based on noisy labels, in: Proceedings of the Fourteenth Symposium on Conformal and Probabilistic Prediction with Applications, PMLR. pp. 82–95. Qian, W., Zhao, Y., Zhang, D., Chen, B., Zheng, K., Zhou, X., 2024. Towards a unified understanding of uncertainty quantification in traffic flow forecasting. IEEE Transactions on Knowledge and Data Engineering 36, 2239–2256. doi:10.1109/TKDE.2023.3312261. Rafe, A., Das, S., 2026. Socio-conformal calibration in complex survey data: Marginal validity is not enough for subgroup reliability. arXiv:2605.05562. Sadinle, M., Lei, J., Wasserman, L., 2019. Least ambiguous set-valued classifiers with bounded error levels. Journal of the American Statistical Association 114, 223–234. doi:10.1080/01621459.2017.1395341. Savolainen, P.T., Mannering, F.L., Lord, D., Quddus, M.A., 2011. The statistical analysis of highway crash-injury severities: A review and assessment of methodological alternatives. Accident Analysis & Prevention 43, 1666–1676. doi:10.1016/j.aap.2011.03.025. Sesia, M., Wang, Y.X.R., Tong, X., 2025. Adaptive conformal classification with noisy labels. Journal of the Royal Statistical Society Series B: Statistical Methodology 87, 796–815. doi:10.1093/jrsssb/qkae114. Seyfi, M., Karimi Mamaghan, A.M., Behnood, A., Mannering, F., 2025. Analyzing crash injury severities with deep learning and advanced statistical models: An assessment of methodological challenges. Analytic Methods in Accident Research 48, 100405. doi:10.1016/j.amar.2025. 100405. Stutz, D., Roy, A.G., Matejovicova, T., Strachan, P., Cemgil, A.T., Doucet, A., 2023. Conformal prediction under ambiguous ground truth. Transactions on Machine Learning Research . Taylor, N.L., Fliss, M.D., Schiro, S.E., Harmon, K.J., 2024. Comparative analysis of injury identification using KABCO and ISS in linked north carolina trauma registry and crash data. Traffic Injury Prevention 25, 912–918. doi:10.1080/15389588.2024.2361052.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 44 of 45

Certified crash-severity prediction Tibshirani, R.J., Foygel Barber, R., Candès, E., Ramdas, A., 2019. Conformal prediction under covariate shift, in: Advances in Neural Information Processing Systems. Tursunbadalov, M., Tursunbadalov, M., 2026. A quiet failure in calibrated virtual screening: Marginal conformal prediction under-covers the minority class, and a class-conditional fix recovers it. arXiv:2607.06605. van der Vaart, A.W., 1998. Asymptotic Statistics. Cambridge University Press. doi:10.1017/CBO9780511802256. Vovk, V., 2012. Conditional validity of inductive conformal predictors, in: Proceedings of the Asian Conference on Machine Learning, PMLR. pp. 475–490. Vovk, V., Gammerman, A., Shafer, G., 2005. Algorithmic Learning in a Random World. Springer, New York, NY. doi:10.1007/b106715. Vovk, V., Gammerman, A., Shafer, G., 2022. Testing exchangeability, in: Algorithmic Learning in a Random World. Springer, pp. 227–263. doi:10.1007/978-3-031-06649-8_8. Wang, J., Goel, S., 2026. Weight clipping for robust conformal inference under unbounded covariate shifts. arXiv preprint doi:10.48550/arXiv. 2605.02072. Wei, J., Miyazaki, Y., Sato, F., 2026. Pre-crash injury risk prediction with guaranteed confidence level: A conformal and interpretable framework. Traffic Injury Prevention 27, 583–593. doi:10.1080/15389588.2025.2538725. Xi, H., Liu, K., Zeng, H., Sun, W., Wei, H., 2025. Exploring the noise robustness of online conformal prediction, in: Advances in Neural Information Processing Systems. Xu, Y., Guo, W., Wei, Z., 2023. Conformal risk control for ordinal classification, in: Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, PMLR. pp. 2346–2355. Yang, C., Huang, X., Qiu, S., Cheng, Y., 2026. CONTINA: Confidence interval for traffic demand prediction with coverage guarantee. Transportation Research Part C: Emerging Technologies 184, 105502. doi:10.1016/j.trc.2025.105502. Zhang, Z., Chen, X., Shi, Y., Ma, L.L., Xu, Z., Yan, Y., 2025. Provably minimum-length conformal prediction sets for ordinal classification. arXiv:2511.16845.

A. Rafe and S. Das: Preprint submitted to Elsevier

Page 45 of 45

Record · ID 673527 · SHA-256 fd4209bb6f0b3cba
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.