arXiv:2605.23663v1 [cs.HC] 22 May 2026
Detecting Drunk Driving Using Off-the-Shelf Smartwatches ROBIN DEUBER, ETH Zürich, Switzerland LANLAN YANG, University of Zürich, Switzerland MICHAL BECHNY, ETH Zürich, Switzerland CHRISTOPH HECK, ETH Zürich, Switzerland MATTHIAS PFÄFFLI, University of Bern, Switzerland MATTHIAS BANTLE, University of Bern, Switzerland FLORIAN VON WANGENHEIM, ETH Zürich, Switzerland ELGAR FLEISCH, ETH Zürich, Switzerland and University of St. Gallen, Switzerland WOLFGANG WEINMANN, University of Bern, Switzerland MANUEL GÜNTHER, University of Zürich, Switzerland FELIX WORTMANN∗ , University of St. Gallen, Switzerland VARUN MISHRA∗ , Northeastern University, USA Alcohol-impaired driving remains a major yet preventable cause of road traffic injury and death, with many drivers underestimating their level of intoxication. Compared to in-vehicle systems, mobile drunk-driving detection using consumer smartwatches offers a scalable way to trigger preventive interventions and increase awareness without additional in-vehicle hardware. We introduce a system that leverages wrist accelerometer data and heart rate variability-derived physiological signals to detect alcohol-related driving impairment. We collected data in a randomized, controlled three-arm test-track study (𝑛 = 54) and trained both logistic regression models with window-aggregated features and a two-tower 1D convolutional neural network (CNN), to detect alcohol-impaired driving. The CNN achieved a participant-averaged area under the receiver operating characteristic (AUROC) of 0.88 for detecting any alcohol intoxication and 0.86 for detecting driving above the WHO-recommended limit of 0.05 g/dL. To the best of our knowledge, this is the first work to (1) demonstrate drunk-driving detection using consumer smartwatches, (2) develop and evaluate such a system in a real vehicle on a closed test track, and (3) rigorously assess generalization to unseen participants. Together, these findings highlight the potential of wearable-based sensing to support scalable, measurement-driven prevention of alcohol-related traffic harm. CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing; Empirical studies in ubiquitous and mobile computing; • Applied computing → Consumer health. Additional Key Words and Phrases: mobile health, safety, driving, wearable sensing, physiological arousal ∗ Shared last authorship.
Authors’ Contact Information: Robin Deuber, [email protected], ETH Zürich, Zürich, Zürich, Switzerland; Lanlan Yang, [email protected], University of Zürich, Zürich, Zürich, Switzerland; Michal Bechny, [email protected], ETH Zürich, Zürich, Zürich, Switzerland; Christoph Heck, [email protected], ETH Zürich, Zürich, Zürich, Switzerland; Matthias Pfäffli, [email protected], University of Bern, Bern, Switzerland; Matthias Bantle, [email protected], University of Bern, Bern, Switzerland; Florian von Wangenheim, fwangenheim@ ethz.ch, ETH Zürich, Zürich, Switzerland; Elgar Fleisch, [email protected], ETH Zürich, Zürich, Switzerland and University of St. Gallen, St. Gallen, Switzerland; Wolfgang Weinmann, [email protected], University of Bern, Bern, Switzerland; Manuel Günther, [email protected], University of Zürich, Zürich, Zürich, Switzerland; Felix Wortmann, [email protected], University of St. Gallen, St. Gallen, Switzerland; Varun Mishra, [email protected], Northeastern University, Boston, Massachusetts, USA. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM.
The manuscript has been submitted to ACM and is currently under review.
0:2
•
Deuber et al.
1
Introduction Sensors
Heart rate and interbeat intervals
Classification
Feature generation
Pretrained physiological arousal detection
Validation
no alcohol
moderate
severe
Early Warning
-
+
+
Above Limit
-
-
+
Leave-one-subject-out cross-validation Fold 1
… test
Arousal Acc.
train
Fold 2
test
…
Fold 54
magnitude =
𝑎𝑥2 + 𝑎𝑦2 + 𝑎𝑧2
1D convolutional neural network
train
…
train
…
Z-score normalization
…
𝑡𝑖 + 180𝑠
…
𝑡𝑖
train
…
train train
test
Accelerometer
Fig. 1. Overview of the 1D CNN pipeline designed to perform two classification tasks (Early Warning and Above Limit).
According to a 2024 report by the World Health Organization (WHO), alcohol consumption caused approximately 2.6 million deaths worldwide in 2019 and represents a major modifiable risk factor contributing substantially to the global burden of disease and injury, with impacts extending across health, safety, and social domains [70]. In the context of road safety, alcohol-related road crashes caused approximately 298,000 deaths worldwide in 2019, of which 156,000 fatalities involved individuals who had not consumed alcohol themselves, highlighting that alcohol consumption poses substantial risks not only to one’s own health but also to the safety of others [70]. Drunk driving generally refers to operating a motor vehicle while alcohol consumption has impaired the cognitive, perceptual, or motor abilities required for safe driving. Empirical evidence shows that alcohol-related crash risk increases nonlinearly with intoxication level: a recent meta-analysis pooling over 100 effect estimates found an exponential relationship between blood alcohol concentration (BAC) and the risk of being killed or seriously injured, with odds ratios reaching up to approximately 240 for BAC levels between 0.12 and 0.20 g/dL [27]. Despite this continuous risk increase, legal BAC thresholds vary substantially across regions, including zero tolerance policies as well as limits such as 0.05 and 0.08 g/dL [20]. In practice, drunk driving most frequently occurs in everyday situations following social drinking, such as evenings out or celebratory events, where individuals may underestimate their level of impairment or overestimate their ability to compensate. While alcohol-impaired driving occurs across all demographic groups, it is disproportionately prevalent among young adults, particularly those aged 21–24, and among men, consistent with broader patterns of social drinking and risk-taking behavior shaped by situational constraints and social norms [11, 50]. Alcohol consumption induces pronounced physiological changes that affect how the body regulates stress, attention, and motor control. In particular, acute alcohol intake disrupts autonomic nervous system regulation, leading to reduced heart rate variability (HRV) [9, 33, 53, 54, 68]. Beyond autonomic effects, alcohol also impairs sensorimotor integration, disrupting gross motor control and movement smoothness [17, 46]. These physiological and motor disturbances compromise cognitive, visual, and fine motor functions essential for safe driving, including sustained attention, decision-making, hazard perception, and precise motor coordination [41, 47]. Consequently, intoxicated drivers respond more slowly to unexpected events, exhibit degraded lane keeping (e.g., increased lateral variability), show increased steering variability, and display impaired visual-motor coordination [10, 23, 29]. The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:3
Together, this evidence demonstrates that alcohol-related physiological dysregulation manifests as measurable deficits in visual control, movement smoothness, and driving performance that substantially increase crash risk [11, 24, 27, 50]. Beyond these physiological and behavioral impairments, alcohol intoxication can compromise self-awareness, and many drinkers struggle to estimate their level of intoxication and fitness to drive accurately. Experimental and field studies similarly show substantial miscalibration between subjective and objective intoxication, and BAC underestimation has been linked to riskier driving behavior [32, 34]. Acute alcohol tolerance can further reduce perceived danger and increase the willingness to drive [2]. At the same time, impaired driving also persists due to a compliance gap: some individuals drive despite being aware of impairment or legal risk. The present work focuses on the former mechanism (miscalibration and reduced situational awareness) rather than on deliberate rule violations. Objective impairment detection could therefore support accident prevention in two complementary ways. First, timely feedback could increase awareness where individuals might otherwise decide to continue driving despite impairment. Second, modern vehicles increasingly incorporate driver-state-aware assistance systems that could account for the driver’s condition into their control logic, for example, by initiating emergency braking earlier [65] and thereby increasing safety. Despite these opportunities, closing the awareness gap at scale requires sensing that is both objective and broadly accessible. Many existing approaches require in-vehicle or specialized sensors, or proprietary integration that is not available across the broader vehicle fleet (e.g. [19, 31]), which can limit coverage and raise practical deployment barriers. These considerations motivate wearable-based drunk-driving detection as a more accessible, non-invasive, and potentially rapid solution. Wrist-worn wearables are widely adopted and continue to ship at scale [21], and driving is a highly structured activity characterized by repetitive control actions that are observable at the wrist. Alcohol is known to affect autonomic regulation and motor behavior (e.g., reduced HRV, elevated physiological arousal, and altered movement patterns) in ways that are directly measurable using consumer-grade wearable devices via HRV features and accelerometry [9, 33, 54]. In summary, the rationale for smartwatch-based detection is not that all intoxicated drivers would necessarily refrain from driving once notified, but rather that a meaningful subset of impaired-driving episodes arises from miscalibration of one’s own intoxication level and fitness to drive. In such cases, objective feedback delivered through a widely available wrist-worn device may help reduce this awareness gap and support preventive decision-making; while the same signal could also serve as an input to driver-state-aware vehicle safety systems. Despite these factors and the potential to reduce the awareness gap among drivers, it remains unclear whether alcohol-impaired driving can be detected using off-the-shelf smartwatches in a real driving context. To address this gap, we introduce a smartwatch-based drunk-driving detection approach that leverages only physiology and accelerometer signals to identify alcohol-related impairment during real driving. In particular, we investigate to what extent alcohol-impaired driving can be detected in a real vehicle using off-the-shelf smartwatches, and assess the extent to which the resulting models generalize to unseen drivers. In summary, this work makes three main contributions. First, we demonstrate that alcohol-impaired driving can be detected using off-the-shelf smartwatches that sense accelerometer and HRV data, without requiring in-vehicle telemetry or driver-facing cameras. Second, we develop and evaluate such a system in a real vehicle on a closed test track, thereby moving wearable-based intoxication detection from unconstrained or gait-based settings to a structured driving context in which gait cues are unavailable and wrist movements are shaped by vehicle control. Third, we rigorously assess generalization to unseen drivers via leave-one-subject-out (LOSO) validation and provide a structured comparison of a logistic-regression baseline and a two-tower 1D convolutional neural network (CNN), including analyses of (temporal context) window length, modality ablations, and perdriving-phase normalization (see Figure 1).
The manuscript has been submitted to ACM and is currently under review.
0:4
•
Deuber et al.
In the spirit of open science and reproducibility, we release the source code accompanying this paper at https: //anonymous.4open.science/r/Detecting-Drunk-Driving-Using-Off-the-Shelf-Smartwatches/. In accordance with local ethics requirements, de-identified study data may be made available for non-commercial research upon reasonable request to the authors, subject to approval by the scientific study board and a data transfer agreement.
2 Related Work 2.1 Wearable-Based Detection of Alcohol Intoxication Alcohol intoxication can be assessed through a wide range of sensing modalities [52]. These include approaches that estimate BAC from reported consumption and physiological characteristics; breath analyzers; bodily fluid analysis (e.g., blood analysis via gas chromatography); transdermal sensing of ethanol or its metabolites; optical spectroscopy that detects ethanol signatures in tissue; and indirect inference based on physiological or behavioral signals (e.g., photoplethysmography (PPG), body temperature, or driving behavior). Focusing on wearable sensing, Davis-Martin et al. [18] categorize biosensor-based intoxication detection into gait-based assessment (e.g., accelerometer and gyroscope) and transdermal alcohol concentration sensing. Although transdermal alcohol concentration sensors directly quantify ethanol diffused through the skin, they are not integrated into consumer-grade smartwatches. Brobbin et al. identified 32 transdermal alcohol concentrationfocused wearable studies but reported significant methodological variability [7]. Recent work by Fairbairn et al. achieved up to 0.94 area under the receiver operating characteristic curve (AUROC) in intoxication detection using a wrist-worn transdermal alcohol biosensor evaluated across both controlled laboratory sessions and real-world field use [22]. In contrast, Panneer Selvam et al. demonstrated a wearable sweat-based biochemical sensor for detecting an ethanol metabolite [51]. Because metabolites can remain detectable substantially longer than contemporaneous intoxication, sweat-based systems are primarily suited to monitoring alcohol consumption over longer time horizons rather than real-time impairment at the moment of driving [51]. Several studies explored physiological sensing for intoxication detection outside of the driving context. Chen et al. proposed a non-invasive sobriety-test system based on finger PPG, and reported an accuracy of up to 85% for detecting alcohol-related changes in PPG signals [14]. Wang et al. developed an intoxication-identification system using electrocardiogram (ECG) and PPG sensors, and trained SVM classifiers, reporting an average identification performance of 95% in accuracy/F1-score; notably, they found PPG-based features to perform comparably to ECG while being more convenient to acquire [66]. Fingertip PPG has also been explored in small-scale prototypes using reflected red/infrared optical sensing and breath-alcohol reference measurements, achieving up to 87.5% classification accuracy between alcohol-consumed vs. non-consumed conditions [56]. Beyond optical sensing, Czaplik et al. evaluated bioimpedance spectroscopy and impedance cardiography in a controlled drinking trial (ethanol vs. water control), using repeated breath-alcohol measurements and showing that bioimpedance-derived parameters track increasing alcohol levels [16]. Wearable and mobile inertial sensing has also been used to estimate intoxication outside of the driving context, typically by leveraging gait- and motion-derived features from commodity accelerometers or gyroscopes. For example, Chawathe used smartphone accelerometer signals to classify intoxication/BAC-related states from gait-derived features and reported discriminative performance around AUROC ≈ 0.85 [12]. In a field setting, Killian et al. collected smartphone accelerometer data during a one-day event with 13 students and used an ankle-worn transdermal alcohol sensor to provide objective ground-truth labels; their best-performing classifier achieved 77.5% accuracy for detecting intoxicated versus sober windows (≥ 0.08 g/dL vs. < 0.08 g/dL) on 10-second segments [30]. McAfee et al. proposed AlcoWear, which infers intoxication from gait using inertial features extracted from smartphone and smartwatch accelerometer and gyroscope sensors, reporting 79.8% accuracy using smartwatch data and 89.5% using smartphone data for classifying BAC ranges [43]. More recently, Segura The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:5
et al. collected smartwatch inertial signals (accelerometer and gyroscope) together with heart rate (HR) over three weeks from 14 participants, using transdermal alcohol concentration as ground truth, and achieved an AUROC of 0.75 for detecting intoxication above BAC 0.05 g/dL using time-series machine learning (ML) models [59]. Taken together, prior wearable-based alcohol detection research demonstrates that alcohol-related signals can be captured using body-worn sensors, but the existing evidence stems predominantly from non-driving settings, specialized sensing modalities, or movement patterns such as gait. In contrast, our study investigates whether off-the-shelf smartwatches can detect alcohol-impaired driving in a real vehicle using only physiological and motion signals in a driving context.
2.2
In-Vehicle Vital Signs and Drunk-Driving Detection
Automotive environments are increasingly equipped with sensors capable of monitoring driver vital signs. Prior work has demonstrated ECG acquisition from the steering wheel [26] and hybrid sensing solutions combining ECG, PPG, and camera-based monitoring [67]. Remote PPG extraction from in-cabin cameras has also advanced significantly [6, 28]. Vital signs have also been measured using seat-integrated sensors in vehicles [60]. Commercial offerings integrating such capabilities are emerging, including from BMW [4] and Bosch [5]. Prior studies have investigated drunk-driving detection using in-vehicle sensing and driver monitoring to infer impairment from deviations in driving behavior. Early work primarily relied on vehicle-state and drivercontrol signals, and many studies relied on simulator-based experiments to classify intoxication from steering patterns, lane positioning, and other behavioral metrics [37, 62]. More recent work typically combines highresolution vehicle telemetry (e.g., steering, lane position, pedal inputs) with driver monitoring cameras (DMCs) capturing eye movements, gaze stability, and head pose, analyzed using ML to distinguish intoxicated from sober driving in simulators [13, 31, 35, 37, 38]. Extending this line of work to real-vehicle settings, a recent test-track study combining driver-vehicle interaction and DMC data reported an AUROC of 0.84 [19]. Complementary evidence from non-detection-oriented studies shows that camera-based indicators of impaired oculomotor control and disrupted visual scanning reliably reflect alcohol intoxication in both controlled and on-road driving contexts [1, 42, 71]. Despite high discriminative performance, this line of work relies on additional in-vehicle sensing hardware, proprietary vehicle integration, and careful calibration, which constrains accessibility due to hardware costs and limits applicability largely to recent vehicle generations and research-grade settings [49, 61]. Accordingly, our work complements this line of research by shifting the sensing locus from the vehicle to the driver: rather than relying on onboard telemetry, cameras, or proprietary vehicle integration, we study whether alcohol-impaired driving can be detected using only off-the-shelf smartwatch signals.
2.3
Research Gap
Wearable and mobile intoxication detection has primarily been studied in non-driving contexts, often relying on gait and general movement patterns. Only recent work has begun to combine motion- and physiologicalsignal–based detection. Driving, however, is a highly structured activity with constrained movements and reduced variability compared with general human motion. This structure may facilitate classification, but many informative cues available in unconstrained settings, such as walking-related patterns, are absent. Because the underlying movement distribution differs substantially, it remains unclear whether motion-based approaches transfer to driving contexts. In addition, some wearable-based approaches rely on sensors that are not available in consumer smartwatches, such as transdermal alcohol concentration sensors. It therefore remains unclear whether off-the-shelf smartwatches can detect alcohol-impaired driving. Beyond smartwatches, some wearable approaches, such as fingertip PPG sensors, are not feasible in driving contexts, while many other wearable-based intoxication detection systems have so far been evaluated only in laboratory settings. In contrast to these approaches, invehicle intoxication detection predominantly focuses on driving behavior and onboard sensing. As a result, The manuscript has been submitted to ACM and is currently under review.
0:6
•
Deuber et al.
wearable-based intoxication detection in driving contexts remains underexplored. This gap matters because detection approaches that depend on vehicle-integrated sensing or proprietary telemetry cannot scale beyond newer vehicles, whereas wearable-based methods can be deployed independently of vehicle infrastructure. To the best of our knowledge, no prior work has investigated drunk-driving detection using wearable physiological and accelerometer signals in real driving scenarios.
3
Data Collection
We conducted a randomized, controlled study (ClinicalTrials.gov NCT05796609) to collect physiological and behavioral data from drivers across varying levels of alcohol intoxication as they were driving. Data collection took place on a closed test track to maximize safety while closely approximating real-world driving conditions. The study protocol was approved by the local ethics committee (IRB) in Bern, Switzerland (ID 2022-02245). It was conducted between April 2023 and July 2023.
3.1
Participants
We recruited 72 participants through public advertisements and social media. Of these, 55 met the inclusion criteria, and 1 was excluded from the analysis due to missing data, resulting in a final sample of 𝑛 = 54. Eligible individuals were at least 21 years old, held a valid driving license for a minimum of three years, and reported regular driving activity. We excluded individuals with a history of substance abuse, pregnancy, or medical conditions contraindicating alcohol consumption. To strictly control for habitual alcohol misuse, we employed a two-stage screening process. First, candidates completed the alcohol use disorders identification test (AUDIT) [57], and those with scores of 15 or higher were excluded. Second, eligible candidates attended a physical screening visit where capillary blood samples were analyzed for phosphatidylethanol (PEth), an objective biomarker of alcohol consumption [40, 58]. Candidates with PEth levels indicating excessive chronic consumption (greater than 200 ng/mL) were excluded. The resulting sample consisted of 54 participants (age 37.2 ± 14.9 years), with a balanced gender distribution (28 female, 26 male). All participants provided written informed consent prior to the study and received financial compensation for their participation (200 CHF, equivalent to 222 USD).
3.2
Study Design
The study employed a three-group design to isolate the effects of alcohol from potential confounding factors such as fatigue, learning effects, and placebo responses. We assigned 31 participants to the treatment group, and they received alcoholic beverages tailored to induce specific BAC trajectories. They completed driving sessions at three intoxication levels: a sober baseline, a severe intoxication phase (target peak BAC 0.08 g/dL; complete session above 0.05 g/dL), and a moderate intoxication phase on the descending limb of the alcohol curve (below 0.05 g/dL, i.e., below the WHO-recommended legal limit). Treatment participants were blinded to both alcohol presence and dosage. Another 12 participants were assigned to the placebo group [64]. To control for expectancy effects associated with consuming a beverage believed to contain alcohol, these participants received a non-alcoholic placebo drink that mimicked the appearance of the treatment beverage and were blinded to their group assignment. The remaining 11 participants were assigned to the open-label reference group. This group received no beverages and completed all driving sessions in a sober state. Participants in the reference group were fully informed about their condition. This group served to control for circadian influences (e.g., increasing drowsiness over the course of the day) and practice effects related to repeated driving of the test track. We selected group sizes consistent with those commonly used in the driver state detection literature (in particular, [31, 36]).
The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
3.3
•
0:7
Apparatus and Sensors
We collected physiological signals and accelerometer data using a consumer-grade smartwatch (Garmin vivoactive 4S) that participants wore on their right wrist. Participants fitted the smartwatch securely, and the study team then verified the fit and adjusted it if needed to minimize motion artifacts during steering maneuvers while maintaining comfort. The smartwatch continuously recorded interbeat intervals (IBIs) in milliseconds (derived from PPG), HR in beats per minute (based on PPG), and triaxial accelerometer data in 𝑔. Breath alcohol concentration served as the ground truth for intoxication levels. Measurements were taken using a police-grade breathalyzer (Dräger 6820, Drägerwerk AG & Co. KGaA), which is certified for law-enforcement use. Breath alcohol concentration values were converted to BAC using a standard breath-to-blood conversion factor of 0.2 (i.e., 0.25 mg/L breath alcohol concentration corresponds to 0.05 g/dL BAC).
3.4
Procedure Driving phase 1
Driving phase 2
Blood alcohol concentration [g/dL]
no alcohol
Driving phase 3
severe
moderate
0.10 0.08 0.06
WHO limit
0.04 0.02 0.00 Treatment n = 31
Placebo n = 12
Reference n = 11
Treatment n = 31
Placebo n = 12
Reference n = 11
Treatment n = 31
Placebo n = 12
Reference n = 11
Fig. 2. Blood alcohol concentration across driving phases. Mean blood alcohol concentration (BAC; g/dL) for the treatment (n=31), placebo (n=12), and reference (n=11) groups shown separately for driving phases 1–3. Points indicate group means across participants. The dashed horizontal line marks the WHO-recommended BAC limit (0.05 g/dL).
Each study day corresponded to the duration of a full workday for a participant. Participants were requested to arrive in a fasted state (no food intake within four hours prior to arrival) to ensure comparable metabolic conditions. Upon arrival, participants underwent a urine test to screen for other substances and, where applicable, pregnancy. Participants first completed a familiarization drive to acclimatize to the vehicle handling and the test-track layout. Subsequently, all participants, independent of group assignment, completed a baseline driving session in a confirmed sober state (BAC = 0.00 g/dL). After the baseline drive, the alcohol administration phase commenced for the treatment group. Individual alcohol doses were calculated using the Widmark formula [69], adjusting for age, sex, body weight, and height to set a dose-calculation target BAC of 0.08 g/dL. The alcoholic beverage was vodka mixed with bitter orange juice to mask its taste. To maintain blinding, the placebo group received an identical volume of bitter orange juice as placebo drink. Both groups received their beverages in neutral bottles with narrow outlets to reduce exposure to smell. Beverages were consumed in a controlled environment over a The manuscript has been submitted to ACM and is currently under review.
0:8
•
Deuber et al.
fixed 20-minute period. Usually, only one participant was present during administration; if two participants were present, they were instructed not to discuss their perceived alcohol level or the drinks they received. A visualization of participant BACs is provided in Figure 2. For the treatment group, the second driving session (severe phase) typically began once BAC peaked and subsequently fell below 0.075 g/dL. Due to procedural variability (e.g., measurement fluctuations), some participants partially drove above this threshold. However, all observed values remained between 0.054 g/dL and 0.086 g/dL. A waiting period of at least 20 minutes after the last sip was enforced to prevent mouth-alcohol contamination of breathalyzer measurements. After the severe phase drive, participants rested while their BAC naturally decreased. The third driving session (moderate phase) was initiated when BAC generally had descended below 0.035 g/dL, with observed values ranging from 0.014 g/dL to 0.044 g/dL. Breathalyzer tests were administered immediately before and after each driving scenario to obtain precise ground truth labels for all data segments. Participants in the placebo and reference groups followed the same temporal schedule as the treatment group. Their second and third drives were scheduled at time points matched to the treatment group to ensure comparable fatigue and circadian states across all conditions.
(a)
(b)
(c)
Fig. 3. Illustrations of the experimental setup: (a) off-the-shelf smartwatch worn by participants; (b) temporary crossroads on the test track where participants were required to stop; and (c) temporarily marked crossing area on the track delineated with cones.
3.5
Driving Tasks and Environment
The driving tasks took place on a closed-circuit proving ground in Switzerland (see Figure 3). The study vehicle was a VW Touran with automatic transmission. The track featured a mix of road geometries, including long straight sections, wide curves, and a complex inner segment with tight turns and intersections. Road widths ranged from 6 to 10 meters. To approximate realistic driving demands, we designed three distinct scenarios. On average, each driving phase (i.e., one of the three) lasted 40 minutes. The highway scenario involved high-speed driving (up to 80 km/h) on the track’s outer oval. The rural scenario was a mixed-speed segment (up to 60 km/h) that included moderate curves, stop signs, and obstacles to maneuver around. The urban scenario was a lowerspeed segment (up to 50 km/h) involving tight turns, a pedestrian crosswalk requiring a full stop, and navigation around artificial obstacles (traffic cones representing roadworks or parked vehicles). Each driving session (no alcohol, severe, moderate) consisted of completing all three scenarios. To prevent order effects, the sequence of scenarios (e.g., urban–highway–rural) and the direction of travel were randomized for each participant and each The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:9
session. A licensed driving instructor was present in the passenger seat at all times, equipped with dual control pedals to intervene in the event of safety-critical errors.
4 Modeling and Evaluation 4.1 Data Preparation Two time series were extracted for each participant: a physiological arousal signal and an accelerometer-based motion signal. As a basis for the arousal signal, the smartwatch provides IBI measurements, which are event-based and irregularly sampled, making direct application of standard time-series modeling techniques challenging. We therefore transform IBI and heart-rate data into an equally spaced physiological representation using a pretrained arousal-estimation model. This choice is grounded in established physiological theory. Acute alcohol intake acts as a systemic stressor that activates the sympathetic nervous system and suppresses parasympathetic activity, leading to reduced heart rate variability and elevated heart rate, which are hallmarks of physiological arousal. Prior work has shown that alcohol-induced autonomic responses resemble those elicited by other physical and cognitive stressors. We leverage the arousal-estimation model developed by Mishra et al., which was trained on diverse arousal paradigms including mental, startle-based, and physical stressors [45]. This diversity enables robust characterization of autonomic activation across heterogeneous conditions. Our approach thus treats alcohol-induced impairment as a form of physiological arousal, grounded in shared autonomic mechanisms, while enabling stable and interpretable time-series modeling. However, this signal should not be interpreted as an alcohol-specific biomarker. Elevated physiological arousal can also arise from other sources of arousal, such as physical activity, task demands, and fatigue. Before inferring physiological arousal probabilities, we removed outliers and applied participant-specific z-score normalization based on each participant’s data distribution. After cleaning both streams, short-window features were computed for IBI (mean, standard deviation, median, minimum, maximum, 20th percentile, 80th percentile, and root mean square of successive differences) and for HR (mean, standard deviation, median, 20th percentile, and 80th percentile) and passed to the pretrained model (provided by Mishra et al. [45]) to obtain a continuous physiological arousal probability over time. Predicted arousal probabilities range from 0 to 1, with values closer to 0 indicating lower physiological arousal and values closer to 1 indicating higher physiological arousal. Accelerometer data was processed independently by loading tri-axial acceleration and computing a motion-intensity signal as the magnitude of the acceleration vector. For the logistic regression approach (in contrast to the CNN pipeline), we perform an additional featureextraction step. From the two time series (physiological arousal and acceleration), we constructed windowed segments that served as the basic modeling units. A sliding window with a fixed length (180 s) and a step size (45 s) produced overlapping segments that captured short-term dynamics in both physiological arousal and motion signals. We selected this window size as a compromise between faster detection with smaller windows and improved detection performance with larger windows, consistent with the general window-length trade-off discussed by Deuber et al. for drunk-driving detection using in-vehicle sensors [19]. For each window, we extracted the corresponding sequences of arousal probability and motion intensity and, in addition, computed a broad set of statistical and temporal features using the tsfresh library [15], including distributional, autocorrelation, and complexity-based measures. Feature extraction was applied separately to the physiological arousal and accelerometer streams, and only windows with sufficient data coverage (> 50 %) and complete overlap with driving intervals were retained. In total, we extracted 783 features per modality; Table 1 summarizes the feature families. All samples were aligned with the study protocol by mapping timestamps to the three driving phases. Treatment participants completed a no alcohol phase (phase 1), a severe intoxication phase (BAC > 0.05 g/dL, phase 2), and a moderate intoxication phase starting once BAC had typically fallen below 0.035 g/dL (phase 3). Placebo The manuscript has been submitted to ACM and is currently under review.
0:10
•
Deuber et al.
Table 1. Overview of extracted tsfresh features per modality.
feature family
tsfresh name
frequency-domain / spectral features
fft_coefficient, fft_aggregated, fourier_entropy, spkt_welch_density, energy_ratio_by_chunks distribution / quantile features quantile, index_mass_quantile, change_quantiles, binned_entropy wavelet / time–frequency features cwt_coefficients, number_cwt_peaks trend / regression features linear_trend, agg_linear_trend autocorrelation / dependence features autocorrelation, partial_autocorrelation, agg_autocorrelation counts / thresholding / crossings value_count, range_count, count_above, count_below, number_crossing_m, ratio_beyond_r_sigma entropy / complexity features approximate_entropy, sample_entropy, permutation_entropy, lempel_ziv_complexity, cid_ce summary statistics mean, median, variance, standard_deviation, skewness, kurtosis, minimum, maximum, . . . autoregressive features ar_coefficient nonlinear dynamics features friedrich_coefficients, max_langevin_fixed_point peak-related features number_peaks boolean / duplicate indicators has_duplicate, has_duplicate_max, has_duplicate_min stationarity / unit-root test features augmented_dickey_fuller similarity / query-matching features query_similarity_count other / uncategorized features
# features 422
77 62 53 23 21
18
17
11 5 5 3 3 1 62
and reference participants followed the same schedule but remained sober throughout. These phases defined the labels used in downstream modeling. We consider two binary detection tasks that reflect complementary goals: Early Warning captures any alcohol exposure (BAC > 0.00 g/dL) as an early-warning signal, whereas Above Limit targets episodes above the WHO-recommended limit (BAC > 0.05 g/dL). Figure 1 visualizes the two classification tasks. For the Early Warning task, all windows with BAC > 0.00 g/dL (treatment phases 2 and 3) were labeled positive; for the Above Limit task, only windows with BAC > 0.05 g/dL (treatment phase 2) were labeled positive. All placebo and reference samples were labeled negative. Beyond the Early Warning and Above Limit binary tasks, we also considered categorical classification tasks trained separately for treatment and control participants aiming to distinguish driving phases 1–3, to evaluate whether the models captured alcohol-related impairment patterns rather than the effects attributable to circadian rhythms, fatigue, or cumulative physiological arousal over the course of the day, thus disambiguating potential ordering effects. The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
4.2
•
0:11
Models and Evaluation
We evaluated two modeling approaches: a logistic regression model with LASSO regularization and a two-tower (late-fusion) 1D CNN trained directly on windowed time-series inputs. Both models were applied to the two binary alcohol-detection tasks. In the categorical setting, models were trained separately for treatment and control participants (placebo and reference groups). All evaluations followed a LOSO validation scheme, in which we iteratively hold out one participant for testing while training on the remaining participants. As a baseline approach, we implemented logistic regression with LASSO regularization. Our pipeline is conceptually similar to the in-vehicle sensor-based drunk-driving detection approach of Koch et al. but differs in sensing modality and feature construction [31]. Specifically, our baseline uses a substantially larger feature set computed via the tsfresh package [15]. As an additional robustness analysis, we evaluated multiple window configurations, combining window sizes from 30 s to 600 s with step sizes equal to one quarter of the respective window length and a minimum window coverage of 50% (i.e., at least 50% of expected samples present). The model was trained with class-balanced weights using the liblinear solver. Probabilistic predictions for the held-out participant were used to assess the performance metrics as described above. Beyond the feature-based logistic-regression baseline, we trained a two-tower 1D CNN, operating directly on time-series segments from raw physiological arousal and accelerometer signals. This CNN constitutes a step toward end-to-end learning, while still relying on the pretrained physiological-arousal-estimation model to transform irregular physiological measurements into an equally spaced input sequence. We use the CNN model to examine whether learned temporal representations can capture impairment-related structure beyond what is available from hand-crafted features alone. We adopted a two-tower late-fusion architecture because the physiological arousal and accelerometer streams differ substantially in temporal resolution and signal characteristics (1 Hz vs. 25 Hz), including temporal granularity, smoothness, and likely noise structure. Moreover, they reflect complementary dimensions of impairment, namely autonomic arousal and motor behavior. Accordingly, separate modality-specific convolutional towers first encode each stream before their representations are fused in a shared classification head. To operationalize this architecture, both modalities were segmented into fixed-length windows (180 s) with a 15 s step, using the same window length as in the logistic regression approach. Within each window, both modalities were resampled to regular temporal grids: the physiological arousal signal was mapped to a 1 Hz grid, and the accelerometer magnitude to a 25 Hz grid (40 ms). Windows were retained only when at least one-third of the expected samples were present. We used a more permissive coverage threshold than in the logistic-regression pipeline because the 1D CNN requires a fixed, exact number of input samples per window. Accordingly, missing values were imputed via time-based linear interpolation; any remaining leading or trailing gaps were filled via forward and backward fill. This procedure produced two synchronized sequences per window (one per modality) that were passed to the respective convolutional towers of the 1D CNN (see Figure 4). The physiological arousal tower of the CNN comprises 3 repeated 1D convolution blocks (Conv1d + BatchNorm + ReLU with temporal downsampling with 16, 32, and 64 channels and kernel size 5), whereas the accelerometer-magnitude tower comprises 4 such blocks (32, 64, 128, and 128 channels and kernel size 7). Each tower applies adaptive global average pooling to obtain a fixed-length embedding; embeddings are concatenated and passed to a shared fully connected head that outputs logits, with dropout (rate 0.3) applied during training. In each LOSO fold, one participant was held out for testing, and the remaining participants were split into training and validation sets by randomly selecting 10 (≈20%) for validation. Before training, each modality was initially standardized using z-score normalization fitted on the inner training data. The model was optimized with AdamW, using a class-weighted binary cross-entropy loss to address imbalance, together with a ReduceLROnPlateau scheduler. For the categorical phase classification, we trained the models using categorical cross-entropy loss. Training employed early stopping based on validation AUROC, and the best-performing model per fold was
The manuscript has been submitted to ACM and is currently under review.
0:12
•
Deuber et al.
Arousal
Acceleration
Input [B, C, L]
Input [B, C, L]
Conv Block
3 times
Conv Block Conv1d
Conv1d
BatchNorm1d
BatchNorm1d
ReLU
ReLU
Downsample
Downsample
Global Pooling
Global Pooling
4 times
Linear ReLU Linear FC Head
Logits Output [B, num_classes]
Fig. 4. Two-tower (late-fusion) 1D CNN architecture. Here, B denotes batch size, C channels, and L window length.
subsequently evaluated on the held-out participant. As in the logistic regression baseline, macro- and micro-level AUROC and area under the precision-recall curve (AUPRC) were computed for comparability across experiments. As an additional test, we report CNN results under a per-phase normalization scheme. Specifically, we compute z-scores within each participant and phase and apply them to the corresponding phase data. This per-phase normalization represents a deliberately conservative evaluation setting. By independently z-normalizing each participant’s signals within each driving phase, we explicitly remove phase-level mean and variance shifts, setting each phase to zero mean and unit variance. This procedure suppresses baseline differences between sober and intoxicated driving that may arise from alcohol-induced physiological changes (e.g., elevated heart rate and reduced HRV). As a result, the model must rely primarily on finer-grained temporal and structural patterns in physiological and motion signals, rather than global level shifts. While this normalization likely removes informative signal and is therefore expected to reduce performance, it provides a stringent test of whether discrimination is driven by dynamic impairment-related patterns rather than coarse baseline differences. As an additional analysis, we evaluated whether the proposed model can predict continuous BAC values rather than solving binary classification tasks. While 0.05 g/dL is the WHO-recommended limit and one of the most commonly applied operational thresholds [70], BAC thresholds vary across jurisdictions [20]. Modeling BAC as a continuous outcome therefore provides a more comprehensive view of performance across the full intoxication spectrum. We used the same two-tower CNN architecture, input modalities, training procedure, and The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:13
LOSO evaluation framework as in the classification setting. The main differences concerned target construction, model output, loss function, and evaluation. Specifically, instead of assigning each time window a binary label, we derived a continuous BAC target for each timestamp by linearly interpolating between time-stamped BAC measurements. The final output layer was adapted to produce a single continuous prediction rather than a binary logit, and the model was optimized using smooth L1 loss. We summarize regression performance using mean absolute error and Pearson correlation between predicted and reference BAC values. To maintain comparability with the Above Limit classification task, we additionally report the AUROC obtained when thresholding the regression outputs to identify whether participants were above 0.05 g/dL.
5
Results
We report performance for two binary classification tasks: Early Warning and Above Limit. Only participants in the treatment group exhibit both positive and negative labels, whereas placebo and reference participants contain exclusively negative samples. Consequently, per-participant AUROC/AUPRC can only be computed for treatment participants, and macro-averages (per-participant averages; mean ± std across LOSO-held-out participants) therefore characterize participant-level discrimination within this group. In contrast, to evaluate performance in a more deployment-relevant setting that includes both intoxicated and consistently sober drivers, we additionally report pooled predictions that incorporate control participants (micro-averages). Because these participants lack positive labels, a pooled evaluation is required to include them. For completeness and comparability, we also report pooled predictions for the treatment group alone. To contextualize AUPRC under varying class imbalance, we report a chance-level baseline defined by the positive-class prevalence of the respective evaluation split: in precision–recall space, an uninformative classifier achieves an AUPRC equal to this prevalence [55].
5.1
Logistic Regression Model
Table 2 summarizes the logistic regression baseline. Within the treatment group, the model achieved solid discrimination for Early Warning (AUROC 0.80 ± 0.11; AUPRC 0.89 ± 0.08), while performance was slightly lower for Above Limit (AUROC 0.75 ± 0.10; AUPRC 0.60 ± 0.13). The AUPRC gap is consistent with the stronger class imbalance in Above Limit (random AUPRC ≈ 0.33 vs. ≈ 0.67 for Early Warning in the treatment-only evaluation). Across both tasks, pooled (micro) performance was slightly lower than the corresponding perparticipant averages (e.g., Early Warning: pooled AUROC 0.77 vs. macro AUROC 0.80 ± 0.11; Above Limit: pooled AUROC 0.71 vs. macro AUROC 0.75 ± 0.10). When pooling treatment and control participants, AUPRC decreased substantially, reflecting the lower prevalence of positive labels in that combined evaluation (random AUPRC 0.38 for Early Warning and 0.19 for Above Limit), whereas pooled AUROC remained comparatively stable (0.73 and 0.72, respectively). Figure 5 shows the receiver operating characteristic (ROC) curves for the logistic regression models. Table 2 also compares models trained on physiological-arousal-only vs. accelerometer-only features. Using a single modality (and keeping the standard window size of 180 s) reduced performance relative to the combined model (Table 2), but both modalities retained measurable predictive value. In the treatment-only evaluation for Early Warning, physiological-arousal-only and accelerometer-only models achieved similar AUPRC (both around 0.83 ± 0.10–0.84 ± 0.10), while the accelerometer-only model yielded higher AUROC than the physiologicalarousal-only model. For Above Limit, accelerometer-only clearly outperformed physiological-arousal-only (macro AUROC 0.73 ± 0.11 vs. 0.64 ± 0.12; macro AUPRC 0.59 ± 0.15 vs. 0.46 ± 0.13). Table 3 analyzes the impact of window size. For Early Warning (treatment-only pooled evaluation), longer aggregation windows consistently improved performance: AUROC increased monotonically from 0.73 (30 s) to 0.80 (600 s), and AUPRC increased from 0.84 to 0.89. When pooling both treatment and control groups, the The manuscript has been submitted to ACM and is currently under review.
0:14
•
Deuber et al.
Table 2. Performance of LASSO-regularized logistic regression (180 s windows) for the Early Warning and Above Limit binary alcohol-detection tasks; as well as ablations of physiological-arousal-only and accelerometer-only (acc). We report macro-averaged (per-participant) and micro-averaged (pooled) AUROC and AUPRC, together with random AUPRC baselines, for treatment participants and for the combined treatment+control group.
per-participant average
pooled predictions
per-participant average
pooled predictions
AUROC treatment AUPRC rand. AUPRC AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
AUROC treatment AUPRC rand. AUPRC AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
Early Warning
Above Limit
0.80 ± 0.11 0.89 ± 0.08 0.67 ± 0.03 0.77 0.87 0.67 0.73 0.61 0.38
0.75 ± 0.10 0.60 ± 0.13 0.33 ± 0.02 0.71 0.58 0.34 0.72 0.40 0.19
arousal only
acc only
arousal only
acc only
0.70 ± 0.14 0.83 ± 0.10 0.67 ± 0.03 0.70 0.83 0.67 0.62 0.50 0.38
0.74 ± 0.14 0.84 ± 0.10 0.67 ± 0.03 0.70 0.81 0.67 0.70 0.57 0.38
0.64 ± 0.12 0.46 ± 0.13 0.33 ± 0.02 0.64 0.44 0.34 0.61 0.24 0.19
0.73 ± 0.11 0.59 ± 0.15 0.33 ± 0.02 0.69 0.55 0.34 0.72 0.39 0.19
influence of window size was rather smaller and notable particularly in AUPRC, which increased from 0.60 (30 s) to 0.64 (600 s). For Above Limit, the effect of window size was also limited: AUROC remained within a relatively narrow range (0.67–0.71), and AUPRC varied only modestly (0.51–0.58). A similar pattern was observed when pooling treatment and control participants, albeit at different AUPRC levels. Table 4 summarizes the mean absolute coefficients of the two logistic regression models across tsfresh feature families, together with the overall mean absolute coefficient values. Coefficients were computed within the same LOSO evaluation framework and then averaged across folds. Overall, acceleration features received substantially larger coefficients than physiological arousal features in both classification tasks (0.409 vs. 0.109 for Early Warning and 0.454 vs. 0.186 for Above Limit), indicating that the logistic regression models relied more strongly on acceleration-derived features. Within the acceleration modality, entropy / complexity features showed the largest mean absolute coefficients for both tasks. For Above Limit, peak-related features ranked second, whereas the pattern was more mixed for Early Warning. Within the physiological arousal modality, the largest coefficients in Early Warning were observed for peak-related and entropy / complexity features. For Above Limit, the most prominent physiological arousal feature families were entropy / complexity, peak-related, autoregressive, and trend / regression features. The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
Early Warning
0.8
0.8
0.6
0.6
0.4 0.2 0.0
0.0
0.2
0.4 0.6 1 Specificity
0.8
0.4 0.2
AUROC Logistic Regression: 0.73 CNN: 0.71
0.0
1.0
0:15
Above Limit
1.0
Sensitivity
Sensitivity
1.0
•
AUROC Logistic Regression: 0.72 CNN: 0.79 0.0
0.2
0.4 0.6 1 Specificity
0.8
1.0
Fig. 5. Receiver operating characteristic (ROC) curves for the Early Warning and Above Limit tasks, comparing logistic regression and CNN models on pooled predictions with treatment and control groups included. The dashed gray line indicates the performance of a random classifier.
Table 3. Effect of window length on logistic regression performance for the Early Warning and Above Limit tasks. Pooled AUROC and AUPRC are shown for treatment-only and treatment+control evaluations across window lengths from 30 s to 600 s, along with random AUPRC baselines. The 180 s window (marked with ∗ ) corresponds to the default configuration used in subsequent analyses.
5.2
30
60
120
180∗
300
450
600
Early Warning
AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
0.73 0.84 0.67 0.73 0.60 0.38
0.75 0.85 0.67 0.73 0.61 0.38
0.76 0.86 0.67 0.73 0.60 0.38
0.77 0.87 0.67 0.73 0.61 0.38
0.77 0.87 0.67 0.72 0.61 0.38
0.80 0.88 0.67 0.76 0.65 0.38
0.80 0.89 0.67 0.74 0.64 0.38
Above Limit
AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
0.67 0.51 0.33 0.71 0.36 0.19
0.70 0.55 0.33 0.72 0.40 0.19
0.71 0.57 0.33 0.73 0.41 0.19
0.71 0.58 0.34 0.72 0.40 0.19
0.70 0.55 0.33 0.72 0.39 0.19
0.71 0.56 0.33 0.72 0.39 0.19
0.71 0.58 0.33 0.71 0.39 0.19
1D CNN Model
The CNN preprocessing pipeline yielded 14,433 samples. Table 5 reports the CNN results. For treatment participants, the CNN improved per-participant average performance over logistic regression for both tasks, reaching AUROC 0.88 ± 0.09 and AUPRC 0.93 ± 0.05 for Early Warning, and AUROC 0.86 ± 0.11 and AUPRC 0.78 ± 0.17 for Above Limit. At the pooled (micro) level, improvements were not uniform across tasks: for Early Warning, The manuscript has been submitted to ACM and is currently under review.
0:16
•
Deuber et al.
Table 4. Mean absolute coefficients by tsfresh feature family, presented separately by classification task and modality. Values are reported as mean ± standard deviation. The bottom row shows the overall mean and standard deviation across feature families. Missing entries denote feature families that were excluded because their features contained too many missing values after tsfresh extraction (specifically, 100.00% missing values for similarity / query-matching features and 3.07% missing values for nonlinear dynamics features in the physiological arousal modality). Early Warning feature family
Above Limit
arousal
acc
arousal
acc
frequency-domain / spectral features distribution / quantile features wavelet / time–frequency features trend / regression features autocorrelation / dependence features counts / thresholding / crossings entropy / complexity features summary statistics autoregressive features nonlinear dynamics features peak-related features boolean / duplicate indicators stationarity / unit-root test features similarity / query-matching features other / uncategorized features
0.069 ± 0.002 0.146 ± 0.006 0.045 ± 0.006 0.145 ± 0.011 0.104 ± 0.006 0.035 ± 0.003 0.172 ± 0.021 0.072 ± 0.010 0.141 ± 0.010 – 0.191 ± 0.012 0.065 ± 0.009 0.135 ± 0.021 – 0.093 ± 0.007
0.065 ± 0.001 0.559 ± 0.031 0.024 ± 0.002 0.540 ± 0.031 0.604 ± 0.048 0.147 ± 0.007 0.976 ± 0.065 0.619 ± 0.066 0.534 ± 0.052 0.599 ± 0.131 0.589 ± 0.077 0.042 ± 0.006 0.274 ± 0.034 – 0.147 ± 0.011
0.125 ± 0.004 0.244 ± 0.011 0.043 ± 0.005 0.354 ± 0.023 0.209 ± 0.020 0.036 ± 0.007 0.362 ± 0.028 0.215 ± 0.022 0.303 ± 0.017 – 0.302 ± 0.025 0.099 ± 0.010 0.064 ± 0.014 – 0.069 ± 0.007
0.087 ± 0.003 0.364 ± 0.021 0.049 ± 0.003 0.604 ± 0.041 0.589 ± 0.034 0.132 ± 0.015 1.806 ± 0.094 0.492 ± 0.047 0.517 ± 0.054 0.137 ± 0.124 1.149 ± 0.111 0.040 ± 0.007 0.259 ± 0.058 – 0.123 ± 0.011
overall mean and standard deviation
0.109 ± 0.010
0.409 ± 0.040
0.186 ± 0.015
0.454 ± 0.045
pooled AUROC slightly decreased (0.75 vs. 0.77), whereas for Above Limit it slightly increased (0.74 vs. 0.71). However, these differences are small and likely within estimation variability, so we refrain from over-interpreting the trends. When pooling treatment and control participants, Early Warning metrics were slightly lower than the corresponding logistic regression baselines (e.g., pooled AUROC 0.75 vs. 0.77), while for Above Limit both pooled AUROC and AUPRC increased (0.79 vs. 0.72 and 0.51 vs. 0.40, respectively). Figure 5 shows the ROC curves for the CNN models. We additionally evaluated a preprocessing variant that normalizes each individual’s signals per corresponding driving phase, to assess whether the models rely on phase-level distribution shifts across the study day (e.g., baseline drift or session-order effects) rather than impairment-related patterns. For both tasks, phase-wise normalization reduced performance compared to the standard preprocessing (e.g., Early Warning macro AUROC 0.82 ± 0.15 vs. 0.88 ± 0.09; Above Limit macro 0.81 ± 0.13 vs. 0.86 ± 0.11). This result suggests that phasespecific level shifts and longer-term temporal context carry an informative signal for intoxication detection, which is attenuated when each phase is normalized independently. We further report the ablation results for the CNN. Consistent with the logistic regression baseline, accelerometeronly models outperformed physiological-arousal-only models, particularly for Above Limit (macro AUROC 0.84 ± 0.10 for accelerometer-only vs. 0.61 ± 0.17 for physiological-arousal-only). However, the combined model achieved the best overall performance, indicating that physiological arousal estimates and wrist motion provide complementary information during driving. Figure 6 visualizes the results. Furthermore, to approximate a deployment-relevant scenario, we evaluated temporally smoothed predictions by aggregating window-level predicted probabilities into 15 s bins and computing a per-driving-segment cumulative The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:17
Table 5. Performance of the two-tower 1D CNN for the Early Warning and Above Limit binary alcohol-detection tasks under standard and per-phase normalization; as well as ablations of physiological-arousal-only and accelerometer-only (acc). We report macro-averaged (per-participant) and micro-averaged (pooled) AUROC and AUPRC, together with random AUPRC baselines, for treatment participants and for the combined treatment+control sample.
Early Warning
per-participant average
pooled predictions
per-participant average
pooled predictions
AUROC treatment AUPRC rand. AUPRC AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
AUROC treatment AUPRC rand. AUPRC AUROC treatment AUPRC rand. AUPRC AUROC treatment + control AUPRC rand. AUPRC
Above Limit
standard
normalized
standard
normalized
0.88 ± 0.09 0.93 ± 0.05 0.67 ± 0.04 0.75 0.85 0.67 0.71 0.54 0.38
0.82 ± 0.15 0.90 ± 0.09 0.67 ± 0.04 0.69 0.82 0.67 0.65 0.54 0.38
0.86 ± 0.11 0.78 ± 0.17 0.33 ± 0.04 0.74 0.63 0.33 0.79 0.51 0.19
0.81 ± 0.13 0.73 ± 0.16 0.33 ± 0.04 0.68 0.53 0.33 0.65 0.30 0.19
arousal only
acc only
arousal only
acc only
0.74 ± 0.17 0.86 ± 0.11 0.67 ± 0.04 0.71 0.82 0.67 0.65 0.50 0.38
0.82 ± 0.14 0.89 ± 0.09 0.67 ± 0.04 0.64 0.77 0.67 0.65 0.58 0.38
0.61 ± 0.17 0.44 ± 0.20 0.33 ± 0.04 0.60 0.38 0.33 0.57 0.21 0.19
0.84 ± 0.10 0.73 ± 0.16 0.33 ± 0.04 0.70 0.53 0.33 0.72 0.39 0.19
moving average. Figure 7 reports pooled AUROC as a function of elapsed time since segment start, computed cumulatively over windows observed up to time 𝑡. The resulting curves stabilize quickly, suggesting that reliable, continuously updated predictions can be obtained within the first ≈ 5 minutes of driving. When reformulating the task as continuous BAC regression rather than binary classification, the model achieved a mean absolute error of 0.019 g/dL. The Pearson correlation coefficient between predicted and reference BAC values was 0.433, and the AUROC obtained when using the regression outputs for the Above Limit classification task was 0.746. Although this performance is lower than that of the dedicated binary CNN classifier for Above Limit detection (AUROC of 0.79), it remains within a comparable range. These results suggest that the model captures BAC-associated signal variation and retains comparable discrimination for identifying driving above the WHO-recommended limit. At the same time, regression performance was likely constrained by the zero-inflated distribution of the training data.
5.3
Categorical Phase Classification
To further probe whether the learned representations capture alcohol-related impairment patterns rather than generic time-of-day or repeated-drive effects, we trained separate categorical CNN models for the treatment and The manuscript has been submitted to ACM and is currently under review.
0:18
•
Deuber et al.
Early Warning
0.8
0.8
0.6
0.6
0.4 0.2 0.0
AUROC Combined: Arousal only: Acc. only: 0.0
0.2
0.4 0.6 1 Specificity
Above Limit
1.0
Sensitivity
Sensitivity
1.0
0.2
0.71 0.65 0.65 0.8
0.4
1.0
0.0
AUROC Combined: Arousal only: Acc. only: 0.0
0.2
0.4 0.6 1 Specificity
0.79 0.57 0.72 0.8
1.0
Fig. 6. Receiver operating characteristic (ROC) curves for the Early Warning and Above Limit tasks, comparing CNN performance using physiological-arousal-only, accelerometer-only, and combined (physiological arousal and accelerometer) inputs. The dashed gray line indicates the performance of a random classifier.
Early Warning
1.0
Above Limit
Cumulative AUROC
0.8 0.6 0.4 0.2 0.0
500
1000 1500 2000 2500 Time since segment start [s]
3000
500
1000 1500 2000 2500 Time since segment start [s]
3000
Fig. 7. We temporally smooth model outputs by aggregating window-level predicted probabilities into 15 s bins and computing a per-segment cumulative moving average (CMA) within each driving segment (i.e., each participant × phase). At each elapsed time 𝑡, we report the pooled AUROC computed over all windows observed up to 𝑡. Shaded regions indicate 99% DeLong confidence intervals. The dashed vertical line at 180 s marks the earliest time at which a CMA-based prediction is available (one full window).
control groups to classify phases 1–3 (Table 6). For treatment participants, phase classification was substantially above random guessing (macro AUROC 0.87 ± 0.08; macro AUPRC 0.78 ± 0.13 under standard preprocessing), whereas performance for control participants was markedly lower (macro AUROC 0.61 ± 0.13; macro AUPRC 0.48 ± 0.12) and close to the by-chance level of performance. These results are consistent with phase separability The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:19
being driven primarily by alcohol-induced changes in the treatment group, with only limited phase-related (i.e., contemporal) patterns in the control group (e.g., due to fatigue, habituation, or other time-varying factors). When we applied per-phase normalization, it further reduced categorical classification performance in the control group, bringing it closer to random predictions, with a per-participant average AUROC of 0.55. In contrast, the treatment group performance remained strong, albeit slightly lower than the standard processing. Table 6. Categorical phase-classification performance of the 1D CNN within the treatment group (𝑛 = 31) and the control group (𝑛 = 23). We report macro-averaged (per-participant) and micro-averaged (pooled) AUROC and AUPRC under standard and per-phase normalization, together with random AUPRC baselines.
per-participant average pooled predictions
AUROC AUPRC rand. AUPRC AUROC AUPRC rand. AUPRC
treatment (𝑛 = 31) standard normalized
control (𝑛 = 23) standard normalized
0.87 ± 0.08 0.78 ± 0.13 0.33 ± 0.00 0.80 0.66 0.33
0.61 ± 0.13 0.48 ± 0.12 0.33 ± 0.00 0.56 0.38 0.33
0.82 ± 0.09 0.72 ± 0.12 0.33 ± 0.00 0.78 0.65 0.33
0.55 ± 0.12 0.42 ± 0.09 0.33 ± 0.00 0.54 0.37 0.33
6 Discussion 6.1 Summary of Findings Our results show that mobile, wearable-based drunk-driving detection based on wrist motion and physiological arousal estimation is feasible. In the following, we discuss several key aspects of the findings. Detection of any alcohol level versus sober (Early Warning) works better than detection of whether participants are above or below the WHO-recommended limit of 0.05 g/dL. This is consistent with previous work that reports the same pattern [19, 31]. The results indicate that human response to alcohol is more individual at higher BAC levels (around 0.05 g/dL) than for the binary distinction sober vs. non-sober. From a technical perspective, class imbalance is also more pronounced for the Above Limit task. When comparing the logistic regression approach with the CNN approach, the picture is more nuanced. For Early Warning (the better-performing task), both model classes yield similar performance, suggesting that logistic regression already exploits most of the available separability. For Above Limit, however, the CNN improves on the logistic regression baseline. The window-length analysis further indicates that shorter windows may be sufficient for practical detection deployment, with only limited benefits from longer temporal context, especially compared to prior work [31]. This is consistent with Fig. 7, which suggests that temporally smoothed predictions reach stable discrimination after a short elapsed time. The additional regression analysis further suggests that the model captures meaningful continuous BAC-related variation. However, because the present work focuses on classification, these results should primarily be viewed as complementary evidence and as indicating an avenue for future work on continuous intoxication estimation. Across conditions, pooled predictions (micro-averages) consistently underperform macro-averages over participants. A plausible explanation is that macro AUROC/AUPRC only requires correct ranking within each participant (i.e., participant-specific score scales and operating points), whereas pooled evaluation additionally requires scores to be comparable across participants. In the latter, performance is determined by a single global ranking (and, for any fixed operating point, an implicit shared threshold) despite subject-specific differences. The manuscript has been submitted to ACM and is currently under review.
0:20
•
Deuber et al.
Consequently, these metrics are not directly comparable and highlight the need for improved cross-participant calibration. A design consideration is that the intoxication phases in the treatment group necessarily followed a fixed temporal order. Consequently, BAC level, session index, accumulated driving experience, time-on-task, and circadian state are partially confounded by design. The placebo and reference groups were included to partially mitigate this limitation by matching the temporal structure without alcohol exposure, thereby reducing the risk that models rely on trivial differences between early and late sessions. As a further robustness analysis, we evaluated per-phase normalization, in which each participant’s signals were z-normalized separately within each driving phase. If classification performance were primarily driven by coarse phase-level distribution shifts, performance would be expected to collapse under this normalization. Instead, per-phase normalization consistently reduces performance, albeit only slightly. This suggests that the models indeed exploit phase-level drifts over the study day, yet remain predictive even after such shifts are removed. To further probe this issue, we added a categorical phase-classification analysis in which we trained separate CNN models to classify phases 1–3 within the treatment and control groups. If generic repeated-drive, fatigue, or circadian effects dominated the learned representations, phase classification should also be strong among sober control participants. Instead, the findings further reinforce our model’s ability to capture alcohol-related effects. The driving phase classification for control participants was slightly above chance, but separability between phases was significantly higher for the treatment group. This supports the interpretation that phase separability in the treatment group is not merely a generic artifact of session order. Regarding modalities, the ablation analyses show that accelerometer features contribute more strongly than physiological arousal features, particularly for Above Limit, while performance differences are smaller for Early Warning. This pattern is also consistent with the coefficient analysis of the logistic regression models, where mean absolute coefficients were generally larger for accelerometer-derived than for physiological arousal features, while differences between Early Warning and Above Limit remained less pronounced. One interpretation is that higher levels of impairment are more strongly expressed in movement dynamics that are relevant for driving, which is consistent with reports that more severe intoxication is typically associated with greater impairment [8]. This interpretation should be considered in light of unobserved variability in steering behavior. Participants were not constrained in their driving posture (e.g., one- vs. two-handed steering), and we did not record hand dominance or hand usage. While this preserves naturalistic driving behavior, steering style may influence wristworn accelerometer signals and contribute to inter-individual variability. Future work should therefore explicitly account for hand dominance, steering style, and device placement when interpreting motion-based features. At the same time, accelerometer and physiological arousal fusion yields the best overall performance, indicating that the two signals provide complementary evidence. A further interpretation is that we rely on a pretrained physiological arousal detection model and do not perform full end-to-end learning for physiological signals; this may slightly limit performance compared to accelerometer features. Conversely, the weaker physiologicalarousal-only performance should not be interpreted as evidence that physiological sensing is uninformative in principle. Using a device-agnostic physiological arousal representation may improve portability across wearables that differ in access to raw PPG versus only derived signals (e.g., IBI). Future work could therefore evaluate end-to-end learning directly from raw PPG, which may yield additional gains while retaining the option of physiological-arousal-based representations for cross-device generalization. While direct comparisons are limited by differences in datasets, experimental protocols, and evaluation settings, our results are broadly in line with, and in some cases slightly higher than, reported performance in prior invehicle drunk-driving detection work using onboard sensors. The two most state-of-the-art studies in simulator and real-vehicle settings [19, 31] apply the same two binary detection tasks and report the same or somewhat lower performance using in-vehicle driver monitoring cameras and vehicle-based sensors. We achieved 0.88 ± 0.09 and 0.86 ± 0.11 for Early Warning and Above Limit, respectively, while Koch et al. reported 0.88 ± 0.09 and The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:21
0.79 ± 0.10 in a simulator setting [31], and Deuber et al. achieved 0.84 ± 0.11 and 0.80 ± 0.10 in a real vehicle using in-vehicle sensing [19]. Although direct comparison is limited (e.g., due to lack of LOSO evaluation, different task formulations, or other methodological differences), our results are broadly comparable to mobile drunk-driving detection outside of driving contexts and highlight wearable-based sensing as a viable alternative to in-vehicle approaches, with the advantage of scaling to drivers in older vehicle fleets without dedicated in-cabin sensing hardware.
6.2
Practical Relevance
Wearable-based intoxication detection is practically relevant because it directly targets a key failure mode: many drivers miscalibrate their own impairment and underestimate their objective intoxication [2, 32, 34, 39]. The primary value of a smartwatch-based detector is not to replace legal enforcement or breath testing, but to reduce this awareness gap by providing objective feedback in situations where subjective judgment is unreliable. Because the proposed approach relies only on a widely available wrist-worn device, it is inherently scalable and can be deployed without additional vehicle hardware, proprietary telemetry access, or in-cabin cameras. This makes it a pragmatic alternative to vehicle-centric systems: it can reach drivers in older vehicle fleets and across heterogeneous mobility contexts, while remaining non-invasive and comparatively low-cost. A second practical implication follows from driver-state-aware assistance systems. If impairment sensing is available, safety functions can, in principle, adapt their behavior to the driver’s condition. In this context, a wearable-derived signal can act as an auxiliary indicator that the driver’s state has changed in a manner consistent with alcohol exposure, and thus that impairment is likely. The results also suggest a deployment-relevant design direction: while mean performance improves only slightly with longer temporal aggregation, the performance confidence interval narrows, indicating more stable predictions, with reliable, continuously updated estimates available within ≈5 minutes. Because nuisance alarms would be unacceptable, the most realistic operationalization is a temporally smoothed risk score that integrates evidence over longer durations rather than triggering on single windows. Such a score could support the two complementary intervention pathways discussed above: (1) user-facing feedback and (2) vehicle-side safety system adaptation. Both pathways emphasize prevention and harm reduction rather than adjudicating legal impairment.
6.3
Contributions
This work makes three core contributions. First, it demonstrates the feasibility of detecting alcohol-impaired driving using only consumer-grade smartwatch sensing in a realistic driving context, without relying on specialized transdermal sensors or proprietary in-vehicle integration. This lowers cost and access barriers and positions wearable-based sensing as a pragmatic complement to vehicle-centric approaches. Importantly, our goal is risk-aware, symptom-based impairment detection: we identify wearable-derived signals associated with alcohol-impaired driving that can support preventive feedback and driver-state-aware assistance systems, rather than attempting causal inference or diagnostic attribution of intoxication. Second, it shows the development and evaluation of an off-the-shelf smartwatch-based drunk-driving detection system in a real vehicle on a closed test track. The study uses a three-arm study design (treatment, placebo, open-label reference) not as a methodological contribution in itself, but to strengthen the interpretation of wearable-derived impairment signals by accounting for expectancy, fatigue, and time-on-task confounds. Third, we show generalization to unseen drivers by consistently applying LOSO cross-validation. In addition, we provide a structured evaluation across model families and experimental factors, including a state-of-the-art baseline and a two-tower 1D CNN, alongside analyses of window size, modality contributions (accelerometry versus physiological arousal), normalization, and categorical phase classification to probe alcohol-specific structure. Taken together, these contributions help The manuscript has been submitted to ACM and is currently under review.
0:22
•
Deuber et al.
close the gap between prior wearable intoxication studies that largely occur outside the driving context and prior drunk-driving detection work, which relies primarily on in-vehicle sensors, and they establish a foundation for future on-road validation and deployment-oriented research.
6.4
Limitations and Outlook
While our findings are promising, several limitations could constrain generalization and highlight opportunities for future work. First, while the sample size is substantial for a controlled alcohol administration study, it does not capture the full diversity of real-world drivers in terms of demographics, health status, driving styles, and wearable usage habits. Larger, more heterogeneous studies are needed to assess generalizability to the broader driving population. Second, the current labels reflect alcohol exposure, but the specificity of the learned signal is not yet established: other driver states (e.g., fatigue, distraction, non-alcohol-related physiological arousal, illness, or medication effects) may induce partially similar physiological or movement changes and could give rise to false positives. Third, for legal and ethical reasons, this study deliberately uses a closed-track design, which represents the highest-fidelity setting currently permissible for controlled alcohol exposure. Although the protocol was designed to approximate real driving, open-road conditions would likely introduce additional variability (e.g., route heterogeneity) and a corresponding distribution shift. Thus, while the test-track design provides high internal validity, its ecological validity remains limited. Driving took place during daytime, without surrounding traffic, secondary non-driving tasks, or externally induced distractions. In addition, the presence of a licensed driving instructor with dual pedals may have created a supervised driving context, which could both increase rule-compliant driving and reduce perceived accident risk. Finally, all data were collected in the same vehicle model, which limits conclusions about generalization across vehicle types, seating positions, steering dynamics, and cabin layouts. Accordingly, future work should prioritize on-road validation under naturalistic conditions, for example, via observational studies generating a sober validation dataset. Such data would make it possible to quantify distribution shift, assess false-positive rates in everyday driving, and evaluate whether calibration or personalization strategies are needed before deployment. These studies would increase ecological validity, although at the cost of reduced experimental control and the absence of controlled intoxication labels. In the longer term, wearable sensing could also be integrated into large-scale naturalistic driving studies, where extensive real-world driving data is collected and rare naturally occurring impaired-driving episodes may be observed under appropriate ethical and legal safeguards. Moreover, while the CNN improves per-participant discrimination, the gap between macro and micro performance indicates that cross-participant calibration is a deployment-relevant consideration. Macro-averaged metrics primarily reflect whether the model can rank intoxicated and sober windows within the same participant, whereas pooled metrics more closely approximate a cold-start deployment scenario in which a common model and threshold are applied to previously unseen drivers. This distinction is important because wearable sensing models often suffer from inter-individual variability in physiology, movement patterns, device placement, and baseline signal distributions. Similar challenges have been discussed in prior mobile and wearable sensing work, where personalization and generalization are treated as key requirements for robust real-world deployment [3, 25, 44, 63]. For practical smartwatch-based drunk-driving detection, this suggests that a purely population-level model may be most appropriate as an initial cold-start model, but that performance could likely benefit from personalization. Promising strategies include collecting a short sober driving baseline, normalizing risk scores relative to typical physiological and motion profile of an individual, adapting decision thresholds per user, or incrementally updating calibration parameters as more user-specific data become available. Future work should therefore evaluate explicit calibration and personalization strategies and quantify how much user-specific data is needed to close the macro–micro performance gap. The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:23
More broadly, future studies could further build on this foundation by exploring additional time-series ML approaches under the same LOSO scheme and assess whether they yield incremental robustness gains. Likewise, temporal postprocessing (e.g., rolling aggregation, hysteresis thresholds, Bayesian filters, or sequence models) may provide a pragmatic way to reduce jitter and improve decision stability in real time. Another promising direction is to leverage wearable foundation models (e.g., [48]) pretrained on large-scale unlabeled or weakly labeled data (potentially lower quality) and then fine-tuned on high-quality experimental datasets. A practical limitation of the present approach concerns detection latency. Our current pipeline uses 180 second windows, and when combined with temporal smoothing, stable predictions emerge after approximately the first 5 minutes of driving. The 180 second window length was selected a priori based on prior literature rather than optimized on the present data, to reduce the risk of overfitting. At the same time, our a posteriori analysis across multiple window lengths suggested that longer windows yielded only limited performance gains. While such window lengths may be acceptable for longer trips, they may reduce applicability in short-distance driving scenarios. Future work should therefore investigate whether shorter windows can retain sufficient predictive performance while enabling earlier detection. Moreover, a limitation is the domain shift introduced by applying the physiological arousal model of Mishra et al. [45] outside the context in which it was developed. The model was trained and evaluated on controlled physiological arousal paradigms, including mental, startle-based, and physical stressors, whereas our setting combines alcohol exposure, active driving, task demands, fatigue, and safety supervision. We therefore interpret this signal as a proxy for autonomic arousal rather than an alcohol-specific biomarker. Our experimental design was intended to partially address this specificity challenge. The placebo and reference groups completed the same three-session schedule while remaining sober, which helps separate alcohol-related effects from generic time-on-task, repeated-driving, and circadian effects. In addition, the categorical phase-classification analysis and the per-phase normalization analysis suggest that phase separability was stronger under alcohol administration than under the sober control schedules. Nevertheless, these analyses cannot rule out all non-alcohol sources of physiological arousal. Future work should therefore investigate end-to-end models trained directly on raw or minimally processed physiological signals, rather than relying on an intermediate arousal-estimation model developed in a different context. A further direction for future work is multimodal fusion between wearable and in-vehicle sensing. The present paper focuses on smartwatch-only detection, as this setting is the most scalable and does not depend on proprietary vehicle integration or additional onboard hardware. Nevertheless, combining wearable signals with vehicle-based measurements may further improve detection performance and robustness, particularly if the two modality groups capture complementary aspects of alcohol-related impairment. Future work should therefore compare smartwatch-only, vehicle-only, and fused models to assess the added value of multimodal integration. Finally, privacy and ethical considerations are central for wearable-based impairment sensing, as continuous physiological monitoring is inherently sensitive. Future work should therefore reinforce data-minimization strategies, prioritize on-device processing, ensure transparency to users, and define clear governance around when and how risk scores are shared with vehicles or third parties, alongside safeguards to prevent misuse.
7
Conclusion
Alcohol-impaired driving remains a major, yet preventable, cause of road traffic harm, and many drivers misjudge their own intoxication, motivating accessible, non-invasive tools that close this awareness gap without requiring additional in-vehicle hardware. In this work, we use consumer-grade smartwatches in a randomized, controlled test-track study to collect wrist motion and physiological arousal signals across sober and intoxicated driving phases, and train both LASSO-regularized logistic regression and two-tower 1D CNN models under LOSO
The manuscript has been submitted to ACM and is currently under review.
0:24
•
Deuber et al.
validation for two binary detection tasks (Early Warning and Above Limit). Our results show that smartwatchbased sensing can reliably distinguish sober from alcohol-impaired driving, with performance clearly above chance and broadly comparable to in-vehicle approaches, thereby supporting the premise that the motivation articulated in the introduction can be addressed using widely available wearable devices. Overall, this work establishes wearable-based drunk-driving detection as a viable and scalable complement to vehicle-centric systems and legal enforcement, while highlighting the need for further research on calibration, real-world deployment, and ethical safeguards.
References [1] Christer Ahlström, Raimondas Zemblys, Svitlana Finér, and Katja Kircher. 2023. Alcohol impairs driver attention and prevents compensatory strategies. Accident Analysis & Prevention 184 (May 2023), 9 pages. doi:10.1016/j.aap.2023.107010 [2] Michael T Amlung, David H Morris, and Denis M McCarthy. 2014. Effects of acute alcohol tolerance on perceptions of danger and willingness to drive after drinking. Psychopharmacology 231, 22 (2014), 4271–4279. doi:10.1007/s00213-014-3579-1 [3] Nikola Banovic and John Krumm. 2018. Warming Up to Cold Start Personalization. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 1, 4, Article 124 (Jan. 2018), 13 pages. doi:10.1145/3161175 [4] BMW Group. 2025. In-car tracking of vital signs: BMW Group takes medicine to the road. BMW Group. Retrieved December 3, 2025 from https://www.bmwgroup.com/en/news/general/2025/automotive-health.html [5] Bosch Mobility. 2025. Systems for interior sensing in commercial vehicles. Bosch Mobility. Retrieved December 3, 2025 from https: //www.bosch-mobility.com/en/solutions/interior/interior-sensing-cv/ [6] Tayssir Bouraffa, Dimitrios Koutsakis, and Salvija Zelvyte. 2025. Deep Learning-based rPPG Models Towards Automotive Applications: A Benchmark Study. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW). IEEE Computer Society, Los Alamitos, CA, USA, 1081–1090. doi:10.1109/WACVW65960.2025.00130 [7] Eileen Brobbin, Paolo Deluca, Sofia Hemrage, and Colin Drummond. 2022. Accuracy of wearable transdermal alcohol sensors: systematic review. Journal of Medical Internet Research 24, 4 (2022), e35178. doi:10.2196/35178 [8] Ty Brumback, Dingcai Cao, and Andrea King. 2007. Effects of alcohol on psychomotor performance and perceived impairment in heavy binge social drinkers. Drug and alcohol dependence 91, 1 (2007), 10–17. doi:10.1016/j.drugalcdep.2007.04.013 [9] Stefan Brunner, Raphaela Winter, Christina Werzer, Lukas von Stülpnagel, Ina Clasen, Annika Hameder, Andreas Stöver, Matthias Graw, Axel Bauer, and Moritz F Sinner. 2021. Impact of acute ethanol intake on cardiac autonomic regulation. Scientific reports 11, 1 (2021), 13255. doi:10.1038/s41598-021-92767-y [10] Vince D Calhoun, David Altschul, Vince McGinty, Regina Shih, David Scott, Edie Sears, and Godfrey D Pearlson. 2004. Alcohol intoxication effects on visual perception: an fMRI study. Human brain mapping 21, 1 (2004), 15–26. doi:10.1002/hbm.10145 [11] Centers for Disease Control and Prevention. 2025. Risk Factors for Impaired Driving. Centers for Disease Control and Prevention. Retrieved December 12, 2025 from https://www.cdc.gov/impaired-driving/risk-factors/index.html [12] Sudarshan S Chawathe. 2020. Using accelerometers in mobile phones to estimate blood alcohol levels. In 2020 IEEE International Smart Cities Conference (ISC2). IEEE, IEEE, Piscataway, NJ, USA, 1–8. doi:10.1109/isc251055.2020.9239049 [13] Huiqin Chen and Lei Chen. 2017. Support Vector Machine Classification of Drunk Driving Behaviour. International Journal of Environmental Research and Public Health 14, 1 (Jan. 2017), 108. doi:10.3390/ijerph14010108 [14] Yang-Yi Chen, Chun-Liang Lin, Yu-Cheng Lin, and Changchen Zhao. 2018. Non-invasive detection of alcohol concentration based on photoplethysmogram signals. IET Image Processing 12, 2 (2018), 188–193. doi:10.1049/iet-ipr.2017.0625 [15] Maximilian Christ, Nils Braun, Julius Neuffer, and Andreas W Kempa-Liehr. 2018. Time series feature extraction on basis of scalable hypothesis tests (tsfresh–a python package). Neurocomputing 307 (2018), 72–77. doi:10.1016/j.neucom.2018.03.067 [16] Michael Czaplik, Mark Ulbrich, Nadine Hochhausen, Rolf Rossaint, and Steffen Leonhardt. 2019. Evaluation of a new non-invasive measurement technique based on bioimpedance spectroscopy to estimate blood alcohol content: A pilot study. Biomedical Engineering/Biomedizinische Technik 64, 3 (2019), 365–371. doi:10.1515/bmt-2018-0070 [17] John C Dalrymple-Alford, P Anne Kerr, and Richard D Jones. 2003. The effects of alcohol on driving-related sensorimotor performance across four times of day. Journal of studies on alcohol 64, 1 (2003), 93–97. doi:10.15288/jsa.2003.64.93 [18] Rachel E Davis-Martin, Sheila M Alessi, and Edwin D Boudreaux. 2021. Alcohol use disorder in the age of technology: a review of wearable biosensors in alcohol use disorder treatment. Frontiers in psychiatry 12 (2021), 642813. doi:10.3389/fpsyt.2021.642813 [19] Robin Deuber, Patrick Langer, Mathias Kraus, Matthias Pfäffli, Matthias Bantle, Filipe Barata, Florian von Wangenheim, Elgar Fleisch, Wolfgang Weinmann, and Felix Wortmann. 2025. Moving Beyond the Simulator: Interaction-Based Drunk Driving Detection in a Real Vehicle Using Driver Monitoring Cameras and Real-Time Vehicle Data. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, Article 84, 25 pages. doi:10.1145/3706598. 3714007
The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:25
[20] Drinkdriving.org. 2025. Drink Driving Limits Worldwide. Drinkdriving.org. Retrieved December 12, 2025 from https://www.drinkdriving. org/worldwide_drink_driving_limits.php [21] eurostat. 2023. Growing importance of internet-connected devices. eurostat, Luxembourg. Retrieved December 19, 2025 from https: //ec.europa.eu/eurostat/web/products-eurostat-news/w/ddn-20230829-1 [22] Catharine E. Fairbairn, Jiaxu Han, Eddie P. Caumiant, Aaron S. Benjamin, and Nigel Bosch. 2025. A wearable alcohol biosensor: Exploring the accuracy of transdermal drinking detection. Drug and Alcohol Dependence 266 (2025), 112519. doi:10.1016/j.drugalcdep.2024.112519 [23] Harriet Garrisson, Andrew Scholey, Joris C Verster, Brook Shiferaw, and Sarah Benson. 2022. Effects of alcohol intoxication on driving performance, confidence in driving ability, and psychomotor function: a randomized, double-blind, placebo-controlled study. Psychopharmacology 239, 12 (2022), 3893–3902. doi:10.1007/s00213-022-06260-z [24] Charles Goldenbeld. 2024. Road Safety Thematic Report – Alcohol and Drugs. Technical Report. European Road Safety Observatory, Brussels, Belgium. https://road-safety.transport.ec.europa.eu/document/download/bd2408b2-64ce-44a8-a4ca-d7820c7c91ba_en?filename= ERSO-TR-alcohol_drugs_2023.pdf [25] Andreas Grammenos, Cecilia Mascolo, and Jon Crowcroft. 2018. You Are Sensing, but Are You Biased? A User Unaided Sensor Calibration Approach for Mobile Sensing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2, 1, Article 11 (March 2018), 26 pages. doi:10.1145/3191743 [26] Fayssal Hamza Cherif, Lotfi Hamza Cherif, Mohammed Benabdellah, and Georges Nassar. 2020. Monitoring driver health status in real time. Review of scientific instruments 91, 3 (2020), 035110. doi:10.1063/1.5098308 [27] Alena Katharina Høye and Ingeborg Storesund Hesjevoll. 2023. Alcohol and driving—how bad is the combination? A meta-analysis. Traffic injury prevention 24, 5 (2023), 373–378. doi:10.1080/15389588.2023.2204984 [28] Po-Wei Huang, Bing-Jhang Wu, and Bing-Fei Wu. 2020. A heart rate monitoring framework for real-world drivers using remote photoplethysmography. IEEE journal of biomedical and health informatics 25, 5 (2020), 1397–1408. doi:10.1109/jbhi.2020.3026481 [29] Christopher Irwin, Elizaveta Iudakhina, Ben Desbrow, and Danielle McCartney. 2017. Effects of acute alcohol consumption on measures of simulated driving: a systematic review and meta-analysis. Accident Analysis & Prevention 102 (May 2017), 248–266. doi:10.1016/j.aap.2017.03.001 [30] Jackson A Killian, Kevin M Passino, Arnab Nandi, Danielle R Madden, John D Clapp, Nirmalie Wiratunga, Frans Coenen, and Sadiq Sani. 2019. Learning to Detect Heavy Drinking Episodes Using Smartphone Accelerometer Data.. In Proceedings of the 4th International Workshop on Knowledge Discovery in Healthcare Data co-located with the 28th International Joint Conference on Artificial Intelligence (IJCAI 2019) (CEUR Workshop Proceedings, Vol. 2429). CEUR-WS.org, Aachen, Germany, 35–42. [31] Kevin Koch, Martin Maritsch, Eva Van Weenen, Stefan Feuerriegel, Matthias Pfäffli, Elgar Fleisch, Wolfgang Weinmann, and Felix Wortmann. 2023. Leveraging driver vehicle and environment interaction: Machine learning using driver monitoring cameras to detect drunk driving. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). ACM, New York, NY, USA, 1–32. doi:10.1145/3544548.3580975 [32] Jöran Köchling, Berit Geis, Cho-Ming Chao, Jana-K Dieks, Stefan Wirth, and Kai O Hensel. 2021. The hazardous (mis) perception of Self-estimated Alcohol intoxication and Fitness to drivE—an avoidable health risk: the SAFE randomised trial. Harm reduction journal 18, 1 (2021), 122. doi:10.1186/s12954-021-00567-4 [33] Pekka Koskinen, Juha Virolainen, and Markku Kupari. 1994. Acute alcohol intake decreases short-term heart rate variability in healthy subjects. Clinical Science 87, 2 (1994), 225–230. doi:10.1042/cs0870225 [34] Jennifer R. Laude and Mark T. Fillmore. 2016. Drivers who self-estimate lower blood alcohol concentrations are riskier drivers after drinking. Psychopharmacology 233, 8 (Feb. 2016), 1387–1394. doi:10.1007/s00213-016-4233-x [35] John D. Lee, Dary Fiorentino, Michelle L. Reyes, Timothy L. Brown, Omar Ahmad, James Fell, Nic Ward, and Robert Dufour. 2010. Assessing the Feasibility of Vehicle-Based Sensors to Detect Alcohol Impairment. Technical Report. National Highway Traffic Safety Administration, Washington, DC, USA. 98 pages. https://www.nhtsa.gov/sites/nhtsa.gov/files/811358_0.pdf [36] Vera Lehmann, Thomas Zueger, Martin Maritsch, Michael Notter, Simon Schallmoser, Caterina Bérubé, Caroline Albrecht, Mathias Kraus, Stefan Feuerriegel, Elgar Fleisch, Tobias Kowatsch, Sophie Lagger, Markus Laimer, Felix Wortmann, and Christoph Stettler. 2024. Machine Learning to Infer a Health State Using Biomedical Signals — Detection of Hypoglycemia in People with Diabetes while Driving Real Cars. NEJM AI 1, 3 (Feb. 2024), 10 pages. doi:10.1056/aioa2300013 [37] Zhenlong Li, Xue Jin, and Xiaohua Zhao. 2015. Drunk driving detection based on classification of multivariate time series. Journal of Safety Research 54 (Sept. 2015), 61.e29–64. doi:10.1016/j.jsr.2015.06.007 [38] ZhenLong Li, HaoXin Wang, YaoWei Zhang, and XiaoHua Zhao. 2020. Random forest–based feature selection and detection method for drunk driving recognition. International Journal of Distributed Sensor Networks 16, 2 (Feb. 2020), 13 pages. doi:10.1177/1550147720905234 [39] Steven Love and Gregoire S Larue. 2025. A systematic review on the factors associated with the accuracy of self-estimated alcohol intoxication: Implications for drink driving. Journal of Safety Research 94 (2025), 425–435. doi:10.1016/j.jsr.2025.08.005 [40] Marc Luginbühl, Friedrich M. Wurst, Frederike Stöth, Wolfgang Weinmann, Christophe P. Stove, and Katleen Van Uytfanghe. 2022. Consensus for the use of the alcohol biomarker phosphatidylethanol (PEth) for the assessment of abstinence and alcohol consumption in clinical and forensic practice (2022 Consensus of Basel). Drug Testing and Analysis 14, 10 (July 2022), 1800–1802. doi:10.1002/dta.3340 The manuscript has been submitted to ACM and is currently under review.
0:26
•
Deuber et al.
[41] Teri L Martin, Patricia AM Solbeck, Daryl J Mayers, Robert M Langille, Yvona Buczek, and Marc R Pelletier. 2013. A review of alcoholimpaired driving: The role of blood alcohol concentration and complexity of the driving task. Journal of forensic sciences 58, 5 (2013), 1238–1250. doi:10.1111/1556-4029.12227 [42] Pierre Maurage, Nicolas Masson, Zoé Bollen, and Fabien D’Hondt. 2020. Eye tracking correlates of acute alcohol consumption: A systematic and critical review. Neuroscience & Biobehavioral Reviews 108 (Jan. 2020), 400–422. doi:10.1016/j.neubiorev.2019.10.001 [43] Andrew McAfee, Jacob Watson, Ben Bianchi, Christina Aiello, and Emmanuel Agu. 2017. AlcoWear: Detecting blood alcohol levels from wearables. In 2017 IEEE SmartWorld, Ubiquitous Intelligence & Computing, Advanced & Trusted Computed, Scalable Computing & Communications, Cloud & Big Data Computing, Internet of People and Smart City Innovation (SmartWorld/SCALCOM/UIC/ATC/CBDCom/IOP/SCI). IEEE, IEEE, Piscataway, NJ, USA, 1–8. doi:10.1109/uic-atc.2017.8397486 [44] Lakmal Meegahapola, William Droz, Peter Kun, Amalia de Götzen, Chaitanya Nutakki, Shyam Diwakar, Salvador Ruiz Correa, Donglei Song, Hao Xu, Miriam Bidoglia, George Gaskell, Altangerel Chagnaa, Amarsanaa Ganbold, Tsolmon Zundui, Carlo Caprini, Daniele Miorandi, Alethia Hume, Jose Luis Zarza, Luca Cernuzzi, Ivano Bison, Marcelo Rodas Britez, Matteo Busso, Ronald Chenu-Abente, Can Günel, Fausto Giunchiglia, Laura Schelenz, and Daniel Gatica-Perez. 2023. Generalization and Personalization of Mobile Sensing-Based Mood Inference Models: An Analysis of College Students in Eight Countries. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 6, 4, Article 176 (Jan. 2023), 32 pages. doi:10.1145/3569483 [45] Varun Mishra, Sougata Sen, Grace Chen, Tian Hao, Jeffrey Rogers, Ching-Hua Chen, and David Kotz. 2020. Evaluating the reproducibility of physiological stress detection models. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies 4, 4, Article 147 (Dec. 2020), 29 pages. doi:10.1145/3432220 [46] Fredrik Modig, Per-Anders Fransson, Måns Magnusson, and Mitesh Patel. 2012. Blood alcohol concentration at 0.06 and 0.10% causes a complex multifaceted deterioration of body movement control. Alcohol 46, 1 (2012), 75–88. doi:10.1016/j.alcohol.2011.06.001 [47] Herbert Moskowitz and Dary Fiorentino. 2000. A Review of the Literature on the Effects of Low Doses of Alcohol on Driving-Related Skills. Technical Report. National Highway Traffic Safety Administration, Washington, DC, USA. doi:10.21949/1525468 [48] Girish Narayanswamy, Xin Liu, Kumar Ayush, Yuzhe Yang, Xuhai Xu, Shun Liao, Jake Garrison, Shyam Tailor, Jake Sunshine, Yun Liu, Tim Althoff, Shrikanth Narayanan, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak Patel, Samy Abdel-Ghaffar, and Daniel McDuff. 2024. Scaling Wearable Foundation Models. arXiv:2410.13638 [cs.LG] doi:10.48550/arXiv.2410.13638 [49] National Highway Traffic Safety Administration. 2023. Advanced Impaired Driving Prevention Technology. Technical Report. National Highway Traffic Safety Administration, Washington, DC, USA. https://www.nhtsa.gov/sites/nhtsa.gov/files/2023-12/anprm-advancedimpaired-driving-prevention-technology-2127-AM50-web-version-12-12-23.pdf [50] National Highway Traffic Safety Administration. 2025. Traffic Safety Facts - Alcohol-Impaired Driving - 2023 Data (DOT HS 813 713). Technical Report. National Highway Traffic Safety Administration, Washington, DC, USA. https://crashstats.nhtsa.dot.gov/Api/Public/ ViewPublication/813713 [51] Anjan Panneer Selvam, Sriram Muthukumar, Vikramshankar Kamakoti, and Shalini Prasad. 2016. A wearable biochemical sensor for monitoring alcohol consumption lifestyle through Ethyl glucuronide (EtG) detection in human sweat. Scientific reports 6, 1 (2016), 23111. doi:10.1038/srep23111 [52] Szymon Paprocki, Meha Qassem, and Panicos A Kyriacou. 2022. Review of ethanol intoxication sensing technologies and techniques. Sensors 22, 18 (2022), 6819. doi:10.3390/s22186819 [53] Shawn F Reed, Stephen W Porges, and David B Newlin. 1999. Effect of alcohol on vagal regulation of cardiovascular function: contributions of the polyvagal theory to the psychophysiology of alcohol. Experimental and clinical psychopharmacology 7, 4 (1999), 484. doi:10.1037//1064-1297.7.4.484 [54] Magdalena Romanowicz, John E Schmidt, John M Bostwick, David A Mrazek, and Victor M Karpyak. 2011. Changes in heart rate variability associated with acute alcohol consumption: current knowledge and implications for practice and research. Alcoholism: Clinical and Experimental Research 35, 6 (2011), 1092–1105. doi:10.1111/j.1530-0277.2011.01442.x [55] Takaya Saito and Marc Rehmsmeier. 2015. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one 10, 3 (2015), e0118432. doi:10.1371/journal.pone.0118432 [56] Pornnapa Sanguansri, Nattapat Apiwong-Ngam, Athipong Ngamjarurojana, and Supab Choopun. 2021. Development of Non-Invasive Alcohol Analyzer Using Photoplethysmography. Journal of Physics: Conference Series 2145, 1 (2021), 012059. doi:10.1088/17426596/2145/1/012059 [57] John B Saunders, Olaf G Aasland, Thomas F Babor, Juan R De la Fuente, and Marcus Grant. 1993. Development of the Alcohol Use Disorders Identification Test (AUDIT): WHO Collaborative Project on Early Detection of Persons with Harmful Alcohol Consumption-II. Addiction 88, 6 (June 1993), 791–804. doi:10.1111/j.1360-0443.1993.tb02093.x [58] Alexandra Schröck, Annette Thierauf-Emberger, Stefan Schürch, and Wolfgang Weinmann. 2017. Phosphatidylethanol (PEth) detected in blood for 3 to 12 days after single consumption of alcohol—a drinking study with 16 volunteers. International Journal of Legal Medicine 131, 1 (Sept. 2017), 153–160. doi:10.1007/s00414-016-1445-x [59] Manuel Segura, Pere Vergés, Richard Ky, Ramesh Arangott, Angela Kristine Garcia, Thang Dihn Trong, Makoto Hyodo, Alexandru Nicolau, Tony Givargis, and Sergio Gago-Masague. 2025. Advancing Intoxication Detection: A Smartwatch-Based Approach.
The manuscript has been submitted to ACM and is currently under review.
Detecting Drunk Driving Using Off-the-Shelf Smartwatches
•
0:27
arXiv:2510.09916 [cs.LG] doi:10.48550/arXiv.2510.09916 [60] Michaela Sidikova, Radek Martinek, Aleksandra Kawala-Sterniuk, Martina Ladrova, Rene Jaros, Lukas Danys, and Petr Simonik. 2020. Vital sign monitoring in car seats based on electrocardiography, ballistocardiography and seismocardiography: A review. Sensors 20, 19 (2020), 5699. doi:10.3390/s20195699 [61] Smart Eye AB. 2023. Smart Eye’s Market-Leading Driver Monitoring Software Included in New Volvo EX90. Smart Eye AB. Retrieved February 7, 2025 from https://www.smarteye.se/news/smart-eyes-market-leading-driver-monitoring-software-included-in-new-volvo-ex90/ [62] Yifan Sun, Jinglei Zhang, Xiaoyuan Wang, Zhangu Wang, and Jie Yu. 2018. Recognition Method of Drinking-driving Behaviors Based on PCA and RBF Neural Network. Promet-Traffic&Transportation 30, 4 (Aug. 2018), 407–417. doi:10.7307/ptt.v30i4.2657 [63] Timo Sztyler and Heiner Stuckenschmidt. 2017. Online personalization of cross-subjects based activity recognition models on wearable devices. In 2017 IEEE International Conference on Pervasive Computing and Communications (PerCom). IEEE, Piscataway, NJ, USA, 180–189. doi:10.1109/PERCOM.2017.7917864 [64] Maria Testa, Mark T. Fillmore, Jeanette Norris, Antonia Abbey, John J. Curtin, Kenneth E. Leonard, Kristin A. Mariano, Margaret C. Thomas, Kim J. Nomensen, William H. George, Carol VanZile-Tamsen, Jennifer A. Livingston, Christopher Saenz, Philip O. Buck, Tina Zawacki, Michele R. Parkhill, Angela J. Jacques, and Lenwood W. Hayman. 2006. Understanding Alcohol Expectancy Effects: Revisiting the Placebo Condition. Alcoholism: Clinical and Experimental Research 30, 2 (Jan. 2006), 339–348. doi:10.1111/j.1530-0277.2006.00039.x [65] Jianqiang Wang, Chenfei Yu, Shengbo Eben Li, and Likun Wang. 2015. A forward collision warning algorithm with adaptation to driver behaviors. IEEE Transactions on Intelligent Transportation Systems 17, 4 (2015), 1157–1167. doi:10.1109/tits.2015.2499838 [66] Wen-Fong Wang, Ching-Yu Yang, and Yan-Fu Wu. 2018. SVM-based classification method to identify alcohol consumption using ECG and PPG monitoring. Personal and Ubiquitous Computing 22, 2 (2018), 275–287. doi:10.1007/s00779-017-1042-0 [67] Joana M Warnecke, Joan Lasenby, and Thomas M Deserno. 2023. Robust in-vehicle heartbeat detection using multimodal signal fusion. Scientific Reports 13, 1 (2023), 20864. doi:10.1038/s41598-023-47484-z [68] Frank Weise, Dieter Krell, and Norbert Brinkhoff. 1986. Acute alcohol ingestion reduces heart rate variability. Drug and alcohol dependence 17, 1 (1986), 89–91. doi:10.1016/0376-8716(86)90040-2 [69] Erik Matteo Prochet Widmark. 1932. Die theoretischen Grundlagen und die praktische Verwendbarkeit der gerichtlich-medizinischen Alkoholbestimmung. Number 11 in Fortschritte der naturwissenschaftlichen Forschung. Berlin : Urban & Schwarzenberg, Berlin, Germany. [70] World Health Organization (WHO). 2024. Global status report on alcohol and health and treatment of substance use disorders. Technical Report. World Health Organization (WHO), Geneva, Switzerland. https://www.who.int/publications/i/item/9789240096745 [71] Raimondas Zemblys, Christer Ahlström, Katja Kircher, and Svitlana Finér. 2024. Practical aspects of measuring camera-based indicators of alcohol intoxication in manual and automated driving. IET Intelligent Transport Systems 18, 8 (June 2024), 1408–1427. doi:10.1049/itr2.12520
The manuscript has been submitted to ACM and is currently under review.