Conceptio › Archive › arXiv CS
arXiv CSopen access

Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

D ETECTING AGITATION B EFORE B EHAVIORAL E SCALATION IN AUTISTIC YOUTH T HROUGH M ULTIMODAL W EARABLE S ENSING

arXiv:2609.24791v1 [cs.HC] 21 Sep 2026

Nibraas Khan1∗ Abigale Plunk2 John Staubitz3 Ingrid Shragge3 Jordan Brooks3 Suzanne Wright3 Alec Brewer4 James Dieffenderfer4 Alper Bozkurt4 Amy Weitlauf3 Nilanjan Sarkar1,2,5 1

Department of Computer Science, Vanderbilt University, Nashville, TN, USA Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, USA 3 Treatment and Research Institute for Autism Spectrum Disorders (TRIAD), Vanderbilt University Medical Center, Nashville, TN, USA 4 Department of Electrical and Computer Engineering, North Carolina State University, Raleigh, NC, USA 5 Department of Mechanical Engineering, Vanderbilt University, Nashville, TN, USA 2

A BSTRACT Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn device, and vocalizations from lapel microphones across 30 clinician-led sessions with 15 autistic youth, paired with expert behavioral annotations. We adapt four pretrained foundation models, one per modality, project each to a shared 128-dimensional space, and fuse them into a single group model. The model detected agitation with an area under the ROC curve of 0.724 at the clinician-annotated onset (within-participant permutation p = 0.0005), declining to 0.608 at 30 s before onset. Thirteen of fifteen participants were above chance. A from-scratch configuration reached only 0.58, while frozen and fine-tuned features performed comparably (0.71 and 0.72). Audio contributed most of the signal, and a watch-only configuration stayed near chance. Individualized agitation is therefore detectable, including in unannotated windows preceding the annotated onset, using foundation-model transfer with one shared model rather than one per child. Keywords Wearable Sensing · Physiological Computing · Autism Spectrum Disorder · Machine Learning · Early Warning Systems

1

Introduction

Autism spectrum disorder (ASD) is a heterogeneous neurodevelopmental condition characterized by differences in social communication and social interaction, along with restricted, repetitive patterns of behavior, interests, and activities [1]. Globally, the World Health Organization estimates that 1 in 100 children have ASD [2], while in the United States, a 2022 CDC survey identified approximately 1 in 31 eight-year-old children with ASD [3]. Among the most pressing clinical concerns in ASD are challenging behaviors including aggression, self-injury, and property destruction. Across childhood and adolescence, as many as 68% of autistic youth exhibit such behaviors [4]. Persistent episodes pose risks to youth and caregivers alike [5], impede skill acquisition [6], and can lead to exclusion from school services and community opportunities, diminishing quality of life [7]. These behaviors often require intensive interventions including specialized behavioral therapy, crisis response, caregiver training, and school-based supports [8]. Understanding why and under which circumstances these behaviors occur is therefore critical to effective and safe intervention. ∗

Corresponding author: [email protected]

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Because these behaviors can create unsafe situations, clinicians seek to understand their function, the pattern of environmental variables that evoke and reinforce them, and intervene preventively. Applied Behavior Analysis (ABA) provides an evidence-based framework for this process. It is implemented by Board Certified Behavior Analysts (BCBAs) in collaboration with caregivers and service systems [9]. BCBAs draw from a continuum of assessment methods, ranging from indirect tools such as structured interviews and rating scales, to direct observation, to experimental analyses that test causal hypotheses about behavioral function [10, 11, 12, 13, 14]. A common technique is the Interview-Informed Synthesized Contingency Analysis (IISCA; [12]), a brief analytic procedure within the broader category of Practical Functional Assessment (PFA; [15]). In an IISCA, a clinician first interviews caregivers to identify the conditions under which challenging behavior is likely and the individualized signs that precede it, then arranges those conditions in a controlled session and withdraws them as soon as those signs appear. The session ends the escalation rather than allowing it to run to a high-intensity episode. Those signs are the target of this work. We refer to the escalating state they express as agitation: a condition of rising distress expressed through changes in various modalities such as movement, vocalization, and autonomic arousal [16]. Agitation is not a single discrete act; rather, it builds over seconds and unfolds across multiple expressive and physiological channels that differ in how readily a person can observe them. The expressive channels comprise the visible and audible manifestations of distress, such as an abrupt shift in posture, the tightening of muscles in the jaw or shoulders, a sudden change in vocal pitch, or vocal protests like heavy breathing or pacing. Caregivers and clinicians become highly skilled at recognizing these external, idiosyncratic behavioral signs. However, agitation is simultaneously driven by internal autonomic channels. These unobservable manifestations include sympathetic nervous system responses like a surging heart rate, diminishing heart rate variability, and spikes in electrodermal activity. Because these physiological components occur entirely beneath the skin, they remain invisible to even the most attentive observer without specialized instrumentation [17, 18]. However, noting and recording these behaviors manually at scale can be technically challenging and resource-intensive. BCBAs are in high demand, services are costly, and dedicated human observation and data collection may compete with providing intervention, instruction, or care in schools, clinics, and homes. Many early signs of agitation are subtle (posture shifts, micro-gestures, brief vocalizations) and easy to miss without continuous attention, and its autonomic components cannot be seen at all. Coverage is inherently incomplete across settings and times of day, observer drift and reactivity can erode data quality, and privacy constraints on recording minors further limit where and how long continuous observation can occur. These constraints motivate technologies that render the signs of behavioral escalation measurable and scalable. Video analysis can quantify posture and repetitive movement [19], but camera-based approaches are constrained by occlusion, lighting, and privacy [20]. Body-worn Inertial Measurement Units (IMUs) extend coverage beyond the camera’s view [21, 22, 23], yet both video and IMU-based tools operate only once something is externally visible. Escalation often begins earlier in the autonomic nervous system, where Heart Rate (HR) increases, heart-rate variability decreases, and electrodermal activity rises [17, 24, 18]. Across sensing modalities, however, most studies train models to predict the challenging behavior itself rather than the agitation that precedes it. Targeting agitation directly creates a safer window for least-restrictive, function-based supports, because the state is present and measurable before the episode it leads to. It also improves data quality: teams can label reliable, lower-intensity signals without waiting for high-intensity episodes, increasing the number of positive training windows [25]. Prior work established the feasibility of treating agitation as the supervised target for real-time early warning in clinical settings, using a lightweight multimodal sensor suite [16]. In that pilot study with three participants, a multimodal wearable stack spanning movement, peripheral physiology, and audio with AdaBoost achieved an average recall of 69.97% at labeled onset and 55.81% at 25 s prior, demonstrating that meaningful early warning is achievable in practice [16]. However, while the pilot study demonstrated feasibility, it highlighted critical bottlenecks for clinical scaling. These included a constrained sensor suite and a reliance on models trained entirely from scratch on a small number of labeled events. Because agitation is highly idiosyncratic, capturing it effectively requires robust and comfortable sensing deployed across a larger population, alongside methods that overcome severe data scarcity without losing sensitivity to individual differences. Furthermore, attempting to predict highly specific behavioral signs spreads limited data too thinly as the diversity of those signs increases. We address that scarcity by moving the representation outside the cohort. Foundation models pretrained on movement, audio, photoplethysmography, and electrodermal corpora supply general structure for each signal type, learned from populations and tasks unrelated to agitation. Our labeled events are then spent on the decision boundary, and not on learning what a vocalization or an arm movement looks like in the first place. The present work targets these barriers with four primary contributions: • Clinically grounded multimodal dataset: We scaled collection to 15 participants across 30 sessions using an improved sensor array informed by a formative wearability study. The dataset captures upper body movement, physiology, and vocalizations during clinical ABA sessions, paired with BCBA verified annotations. 2

Detecting Agitation Before Behavioral Escalation in Autistic Youth

• Agitation modeling via foundation models: We overcome data scarcity and label fragmentation by pooling individualized signs into a single agitation class and fusing pretrained foundation models into a shared representation. • Evaluation: We evaluate encoder transfer (frozen, fine tuned, and from scratch), quantify modality contributions via ablation, and demonstrate that a shared group model performs on par with individual per participant models. We report all results using area under the ROC curve with bootstrap confidence intervals and within participant permutation tests. • Analysis of cohort heterogeneity: We analyze why detection succeeds or fails across different children. By mapping model performance to the specific phenotypic expression of a child’s agitation, such as the difference between vocal signs and diffuse postural signs, we provide interpretable boundaries for future clinical deployment. The rest of the paper is structured as follows: Section 2 reviews related work on sensing for challenging behaviors and for the agitation that precedes them. Section 3 describes the dataset, sensing modalities, feature design, and model. Section 4 presents quantitative results and model interpretability analyses. Section 5 offers discussion of deployment, ethics, and limitations. Section 6 concludes and outlines future work.

2

Related Work

Sensing has been applied to challenging behaviors and to the earlier, lower intensity states that precede them; yet much of the literature still targets the challenging behaviors themselves rather than explicitly modeling agitation. To prepare the reader for systems that act early in real settings, we frame the related work by the channels through which agitation is expressed: what can be seen and heard (expressive signals, meaning movement and vocalizations, captured by vision, body worn IMUs, and audio) and what reflects internal state (autonomic signals measured by physiology). We then discuss how prior multimodal systems have attempted to connect these channels, and highlight how leveraging pretrained representations addresses the fundamental data constraints of personalized detection. In ABA practice, BCBAs begin with what can be seen: direct observation of antecedents, behaviors, and consequences. It is therefore unsurprising that early computational systems mirrored this visual emphasis, adapting video pipelines to quantify visible topographies. A systematic review highlighted the promise and pitfalls of computer vision in autism research, underscoring the importance of robust pose estimation and ecological validity [20]. Recent work has scaled to longer, more naturalistic recordings: an open-source pipeline automatically localized stereotyped motor movements in hours of clinical video from 241 children with high segment-retrieval sensitivity [19]. Targeted detectors also show strong accuracy when tasks are narrowly defined; for example, Washington et al. achieved a macro-F1 of 90.8% for head-banging classification using head-pose keypoints with a CNN+LSTM under child-wise splits [26]. Despite these advances, real-world deployment must contend with occlusion, variable lighting, bystander privacy, and the fact that much of the day occurs outside any camera’s view [20, 19]. These constraints motivate body-worn sensing. Body-worn inertial sensors can capture posture tightening, abrupt limb actions, repetitive movements, and other motor expressions of agitation and challenging behavior. Early lab/classroom work recognized stereotypy from accelerometry with ∼89% accuracy and classroom per-class F1 of 0.74–0.94 [22]. Subsequent studies extended to naturalistic self-injurious behavior, reporting up to 99.1% individual-level accuracy [21], while design studies emphasized garmentintegrated multi-IMU arrays for comfort and adherence in schools [27]. Yet, motor changes are not the sole expressive channel during escalation as brief vocal signals often co-occur and can add complementary evidence. Microphones can capture protests, non-speech vocalizations, and prosodic shifts that accompany agitation, especially in busy classrooms where ambient mics underperform. Meta-analytic and cumulative work documents robust (though heterogeneous) acoustic differences in ASD (e.g., higher mean pitch and greater pitch variability/range) [28, 29, 30]. For behavior measurement, a neural system quantifying vocal stereotypy achieved session-wise correlations ≥0.80 in 6/8 participants (and ≥0.90 in several), showing that lapel audio can reliably track repetitive vocal topographies [31]. In practice, acoustic features (RMS, zero-crossing, spectral centroid/rolloff, MFCCs) augment IMUs when agitation is expressed vocally. Still, agitation often begins without any visible or audible change, prompting a turn to autonomic physiological signals. Sympathetic arousal (heart-rate increases, Heart-Rate Variability (HRV) reductions, skin-conductance rises) can precede outward behavior by seconds to minutes [18, 24, 32]. These signals are captured noninvasively with Photoplethysmography (PPG)/Electrocardiography (ECG) and EDA following established acquisition/quality practices [33, 34]. In inpatient/residential contexts, physiology alone has predicted imminent aggression at short horizons: a multi-site study (70 youths, 4 hospitals) reported area under the receiver operating characteristic curve (AUROC) ≈0.80 for 3-minute forecasts using logistic regression [35], while earlier work found AUROC 0.84 (person-specific) vs. 0.71 (population) 3

Detecting Agitation Before Behavioral Escalation in Autistic Youth

at 1 minute [23]. Complementary feasibility data show anticipatory EDA rises in roughly 60% of agitation episodes [36], supporting physiology as a low-salience, autonomic channel, though best used in concert with expressive cues. How these systems are evaluated shapes what their results mean. A recent systematic review of thirteen studies predicting severe behavior problems from wearables in neurodivergent people finds that methodological concerns reduce the veracity of the advance-prediction claims in this literature, and recommends cross-validation blocked by both participant and time [37]. Sliding windows are extracted back to back and conventionally overlap by half or more, so adjacent windows share many of the same measurements, and random fold assignment places near-duplicates on both sides of the split [38, 39]. Across 47 clinical wearable-sensor studies, record-wise cross-validation gave a median error of 5.60% against 13.00% for subject-wise [40]. The split is therefore a key design choice. Across these channels, most models have been trained directly on the collected data, whether from a single modality [26, 21, 31, 35] or several combined [16], rather than from encoders pretrained outside the task. Foundation models present that opportunity, and concurrent work has begun to take it [41]. Trained once on a large external corpus, such a model learns general-purpose representations that transfer to downstream tasks with little labeled data, an approach that has reshaped vision, language, and audio and is now reaching wearable and physiological sensing. This fits our setting, where each child contributes only a handful of labeled events, far too few to train an encoder from scratch. Movement encoders pretrained across many human-activity datasets recognize actions they were never trained on [42]; efficient audio networks pretrained on AudioSet transfer to a broad range of acoustic tasks [43]; self-supervised models pretrained on large unlabeled photoplethysmography corpora produce embeddings that carry cardiovascular state [44]; and a model pretrained on a large corpus of wearable electrodermal activity yields embeddings of autonomic arousal [45]. Such an encoder can be used with its weights frozen, training only a small head on top, or adapted to the task with low-rank adaptation, which inserts a few trainable parameters into the frozen backbone and tunes it without overfitting the limited data [46]. We adopt these models because they are already proven in their respective domains, using them as the encoders our own per-participant data could not train. Concurrent work applies this transfer to challenging behavior in profound autism, fine-tuning a movement encoder pretrained on a large activity corpus alongside an electrodermal autoencoder and a temperature network, fusing the three, and predicting up to ten minutes ahead in a special education classroom, reaching an AUC of 0.78 across nine participants [41]. Earlier work detected the same behaviors from the same three signals without pretraining, reaching an AUC of 0.71 [47]. That line of work predicts the challenging behavior itself, defined as categories shared across a cohort, and we detect the agitation that precedes it, defined by the signs a clinician identified for each child. We also transfer a pretrained encoder for every modality, including audio, which it does not sense. Those numbers are therefore not commensurable with ours, because the label, the negative class, the horizon, and the split all differ. We bring this transfer to the agitation-detection problem, where it directly addresses the data constraint. The difference from prior work is where the representation comes from. Most systems in this area learn their features from the study’s own recordings [16, 21, 31, 23], so the representation can be no richer than the labeled events one small cohort provides, and a child who contributes sixty events constrains it as much as they constrain the classifier. We instead take representations learned from external corpora orders of magnitude larger than any agitation dataset, and train a thin projection, fusion block, and per-participant head on top. Each child’s few labeled events then specify a decision boundary within a space that is already structured, and no longer have to build that space. The shared structure comes from pretraining, and what remains individual is a linear read-out.

3

Methods

3.1

Dataset and Data Acquisition

The data for this study were collected through two sessions of modified IISCA, a structured approach derived from PFA principles. During these sessions, participants were equipped with various wearable technologies to capture relevant data. In the following subsections, we detail the data collection protocol, the sensor modalities utilized, and the characteristics and categorization of the observed signs of agitation. Participant demographics are summarized in Table 1, with individualized agitation and challenging behavior descriptions provided in Table 4 in the Appendix. 3.1.1

Data Collection Protocol

Data collection sessions were conducted in a standardized two-room suite at affiliated therapy centers during normal clinic operations, introducing realistic ambient noise and interruptions (Figure 1). The session room, where the participant and a BCBA engaged in clinician-led sessions, was equipped with highly preferred materials identified through a caregiver interview to establish a “happy, relaxed, and engaged” (HRE) state. An adjacent observation room housed the caregiver and a session manager who monitored the session via a one-way mirror or live video feed, 4

Detecting Agitation Before Behavioral Escalation in Autistic Youth

(a) Experimental Room

(b) Observation Room [16]

Figure 1: Data collection setup, showing two connected spaces: (a) the session room for participants and BCBAs, equipped with sensors, and (b) the observation room for the research team and caregivers, with equipment for monitoring and communication.

(a) Inertial Measurement Units

(b) Physiology Sensors [48]

(c) Audio Microphone

Figure 2: Overview of the multimodal sensor array employed for data collection. (a) shows the IMU hardware unit used in the study. (b) displays the wrist- or ankle-worn device used to capture HR, EDA, and acceleration. (c) shows the wireless lavalier microphone used to record vocalizations. confirmed agitation occurrences, and relayed observations to the session-room implementer. Participants and caregivers were informed of their right to pause or terminate at any time; no participant chose to do so. Participants (or caregivers) received $40 after the first visit and $90 after the second. Sessions began with a baseline reinforcement (SR) phase providing unrestricted access to preferred items and social interaction, allowing the participant to acclimate and reach HRE (typically at least 5 minutes). An establishing operation (EO) was then introduced, consisting of individualized antecedent conditions (e.g., removing preferred items, introducing non-preferred demands) designed to evoke non-dangerous agitation. Upon the first clear sign of agitation (or challenging behavior), the EO was terminated and SR reinstated. This contingent alternation enabled repeated, controlled observation of agitation events. 3.1.2

Ethics and Consent

All procedures were approved by the Vanderbilt University Institutional Review Board (IRB #211846). Caregivers provided informed consent, and participants provided assent when able to do so. Sessions were conducted by BCBAs with real-time oversight and a predefined plan to terminate an establishing operation and return to reinforcement conditions immediately upon display of agitation. No adverse events occurred. 3.1.3

Modalities

To capture a comprehensive picture of the participants’ states and behaviors during these sessions, a multimodal sensor array was developed and used to record upper-body movement, physiological signals, and vocalizations. The sensor array wirelessly transmitted, through Bluetooth Low Energy (BLE), all data in real-time to a custom-developed iOS application for recording and monitoring. 5

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Participant 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15

Age 10 yrs 1 mo 11 yrs 11 yrs 3 yrs 5 mo 3 yrs 5 mo 13 yrs 1 mo 7 yrs 4 mo 5 yrs 10 mo 3 yrs 7 mo 17 yrs 5 mo 5 yrs 10 mo 5 yrs 9 mo 6 yrs 7 mo 10 yrs 8 mo 3 yrs 4 mo

Sex M F M F M M M F M M F M M M M

Agitation events 223 241 366 81 66 333 211 81 80 64 106 61 58 196 96

Table 1: Participant demographics and number of annotated agitation events. Individualized descriptions are provided in Table 4 in the Appendix.

A key modality, illustrated in Figure 2a, involved the use of wearable IMUs. Five IMU sensor units were used per participant. These units, with a Bosch BNO055 (9-axis accelerometer, magnetometer, and gyroscope) [49], were embedded in custom-designed wireless circuit boards. For consistent placement and participant comfort, these circuit boards were inserted into specially designed pockets sewn into a custom shirt worn by the participant. The garment material and attachment decisions were informed by a formative wearability study with autistic youth and caregivers (Section 3.1.4). Placements included one IMU on each wrist, one on each upper arm, and one on the upper torso. The torso IMU was placed either on the upper back or the upper chest depending on participant comfort and tolerance, while keeping the location fixed within a participant across sessions. This configuration was selected to capture both distal and proximal upper-body movement and overall trunk posture, as many individualized signs of agitation and challenging behaviors in our cohort (e.g., hitting, grabbing, slapping surfaces, tensing shoulders, leaning away) are expressed through the arms and upper body, while still keeping all hardware integrated into a single, tolerable garment [50, 51, 52]. Data from all IMUs were sampled at 100 Hz. The accuracy of the IMU data was validated by comparing their derived Euler angles against known positions and orientations. Physiological data, depicted in Figure 2, were captured with a wrist- or ankle-worn EmotiBit, which records electrodermal activity (15 Hz) and green-channel photoplethysmography (25–50 Hz), from which heart rate and its variability are derived [48]. Not all participants elected to wear the device, in keeping with their sensory profiles, so these signals are available for 7 of the 15 participants, and a modality a participant lacks is masked (Section 3.3). Finally, audio data, shown in Figure 2c, were captured using a wireless lavalier lapel system attached to the participant’s clothing to record vocalizations. Audio data were sampled at the typical rate of 48 kHz. 3.1.4

Formative Wearability Study to Inform Garment and Attachment Design

Wearability and sensory acceptability are practical determinants of whether autistic youth will tolerate body-worn sensors during real-world use. To inform the garment and attachment mechanisms, we conducted a formative wearability assessment with 7 autistic youth and 6 caregivers. Participants evaluated candidate shirt fabrics, wristband options, and internal vs. external sensor pocket designs, providing comfort ratings on a 7-point scale alongside qualitative feedback (Figure 3). Comfort ratings were highest for the softest fabrics and internal pocket designs, consistent with concerns about scratchiness, seams/tags, and snag risk. Wristband ratings showed greater variability, motivating alternative placements (e.g., ankle) when feasible. These preferences guided the garment material and pocket configuration adopted in our data collection. Although internal pockets were preferred, we used external pockets to enable rapid sensor access during setup and troubleshooting while keeping placement consistent. 3.1.5

Event Annotation and Labeling

The identification and precise timing of agitation events within the collected data were critical for model training and evaluation. This was achieved through a detailed post-hoc annotation procedure utilizing the synchronized video and audio recordings. 6

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 3: Wearability study components evaluated by participants, including candidate shirt fabrics, attachment bands for a watch-style wearable, and garment prototypes used to assess pocket and sensor-placement preferences.

All data collection sessions were recorded using the four webcams positioned in the session room, providing multiple viewing angles, along with the audio from the wireless lapel microphone. These data streams were synchronized with the physiological and motion sensor data. For annotation, these data were loaded into a custom-developed React-based website (Figure 4) [53]. A Registered Behavior Technician (RBT) or a BCBA used the website to review the session recordings and label the occurrences of agitation. Each identified event was timestamped directly within the website. The platform facilitated frame-by-frame video navigation and audio playback control to enhance the accuracy of temporal annotations. Inter-observer agreement (IOA) was calculated for 20% of all recorded video sessions. This involved a second independent trained annotator coding these selected sessions. Agreement for agitation event occurrence (presence/absence in a defined window) and onset/offset timing (within a ±1-second tolerance) was assessed using Cohen’s Kappa, with values consistently exceeding 0.85. Disagreements were resolved through discussion and consensus with a supervising BCBA. The frequency of annotated agitation varied widely across the cohort, reflecting the idiosyncratic nature of behavioral escalation. Table 1 reports the number of annotated agitation events per participant, highlighting the inter-individual variability in how signs of agitation manifest. 3.1.6

Pooling Individualized Signs

As introduced in Section 1, we treat agitation as a state defined by the presence of individualized, clinically identified signs of distress, and our task is to detect those signs. Agitation is expressed differently by every participant, and the individualized signs are presented in Table 4. Treating each sign as its own class would spread limited data across many labels and yield unstable performance, particularly for the rarer signs. We therefore pool every sign into a single agitation class against a non-agitation background, and the model detects that class rather than any individual sign. We next describe how these raw, continuous streams become the features and inputs the model uses. 7

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 4: The custom React-based website used to annotate agitation events. Trained annotators reviewed the synchronized four-angle video and lapel audio, navigated frame by frame, and timestamped the onset and offset of each event directly in the browser.

3.2

Signal Processing and Feature Engineering

Continuous data streams from all modalities were segmented into fixed windows, from which we extracted the engineered features described below along with the raw signals consumed by the foundation models (Section 3.3). Because our data-collection protocol’s contingent alternation repeatedly evokes agitation and ends the evoking condition at the first sign, agitation onsets recur within a session, so each window must be short enough to isolate one. A longer window would enclose more than one, breaking the correspondence between a window and a single agitation event. The same localization keeps the lead-time analysis meaningful (Section 4.3), where the label is shifted to windows before the onset and a longer window would blur advance detection together with detection at the onset. We set the length to 15 s because an agitation episode plays out over several seconds of movement and vocalization, long enough to capture that dynamic and absorb the annotation-timing jitter while keeping the onset localized without absorbing prior windows of agitation.

3.2.1

Feature Engineering

Each modality gives the model two inputs: the raw signal, which a pretrained foundation model encodes into an embedding (Section 3.3), and a set of hand-crafted features computed within the window. All four modalities use both, concatenated before projection. The engineered features are compact, interpretable descriptors computed directly on the window. For movement, photoplethysmography, and electrodermal activity, they carry a signal the foundation-model embedding does not otherwise receive, and for audio they summarize the lapel-microphone signal with standard spectral and energy descriptors. For movement, the engineered features summarize upper-body IMU motion. They include per-location speed statistics (mean, standard deviation, maximum, 90th percentile, and energy) and shape descriptors that separate discrete actions from ambient movement (jerk, kurtosis, crest factor, peak rate, and spectral entropy). We also add per-axis angular velocity and energy [54, 55]. For electrodermal activity, we compute arousal features from skin conductance: mean, variability, range, and slope. For photoplethysmography, we compute cardiac features: mean and variability of heart rate, heart-rate-variability metrics SDNN, RMSSD, and pNN50, and pulse amplitude. For audio, we compute standard descriptors of the lapel-microphone signal: root-mean-square energy, zero-crossing rate, spectral centroid, rolloff, and bandwidth, and the first thirteen mel-frequency cepstral coefficients. 8

Detecting Agitation Before Behavioral Escalation in Autistic Youth

3.2.2

Feature Set Construction

The engineered features were aggregated into a per-modality feature vector. Missing engineered values were imputed with zero prior to z-score normalization. We standardized each engineered feature per participant using that participant’s mean (µ) and standard deviation (σ): z = (x − µ)/σ. The foundation-model embeddings are used as produced, without per-participant standardization. Normalization is instead applied inside the model, by the layer normalization in each modality’s projection (Section 3.3). Note that an imputed zero in raw space becomes (0 − µ)/σ after standardization (not necessarily 0). This assigns missing entries a constant, data-driven value without fabricating temporal structure. We prefer this over forward-fill because forward-filling windowed statistics (e.g., variance, peak counts, short-term trend) can introduce artificial persistence and autocorrelation, precisely where rapid changes convey signal near agitation onsets. Zero-impute with z-score preserves marginal distributions while avoiding synthetic trajectories. To prevent leakage, µ and σ are computed from training data only (per participant) and reused for validation/test. 3.2.3

Target Label Generation for Predictive Modeling

For each participant we build a labeled set of 15 s windows. Each clinician-annotated onset contributes one agitation window, the 15 s window ending at that onset. Non-agitation windows are sampled on a 15 s grid from stretches at least 15 s from any onset. We report a single detection task over all 15 participants, signs of agitation versus non-agitation. Engineered features are z-scored per participant from training-fold statistics, so each feature reflects deviation from that child’s own baseline, and µ and σ are never computed from held-out windows. An annotation marks the moment a clinician could first identify and timestamp a sign of agitation from video. It does not mark the moment agitation began. Agitation builds over seconds and includes autonomic components that no observer can see, so the annotated onset is best understood as an upper bound on true onset: agitation is already underway when it is marked. We treat this as an explicit modeling assumption throughout, and it is the same assumption used to set label window length in prior work [16]. It has a direct consequence for evaluation. If agitation is present before it is annotated, then a model that flags a window shortly before the annotation may be detecting the state rather than making an error, and standard onset-aligned metrics will score those detections as false positives. To study how performance changes as a function of lead time, we repeated the evaluation with the input window positioned earlier relative to each onset. For a lead of ∆t, the 15 s window ends ∆t seconds before the annotated onset (∆t = 0 ends at the onset), with its label unchanged; we evaluate ∆t ∈ {5, 10, 15, 20, 25, 30} s. This measures how well the model separates pre-onset windows from non-agitation windows. They do not establish that the state present at ∆t was independently verified as agitation by a clinician, because no such annotation exists for those windows. Section 4.3 reports these results, and Section 5 discusses the interpretation this supports. 3.3

Multimodal Foundation-Model Architecture

Each participant contributes few labeled signs of agitation, so rather than learn a representation from scratch we transfer four foundation models pretrained on large external corpora, one per modality. Movement is encoded by a spatio-temporal graph network pretrained across many human-activity datasets [42], photoplethysmography by a self-supervised model [44], audio by an efficient convolutional network pretrained on AudioSet [43], and electrodermal activity by a convolutional network pretrained on a large corpus of wearable electrodermal recordings [45]. 3.3.1

Per-modality projection and fusion

Each foundation model produces a modality embedding of a different size (512, 512, 960, and 64 dimensions for movement, photoplethysmography, audio, and electrodermal activity, respectively). Each modality embedding is concatenated with that modality’s engineered features. Each modality’s representation is projected to a common 128-dimensional space through a linear layer, layer normalization, a rectified linear unit, and dropout. Projecting every modality to the same width keeps any one embedding from dominating the fusion. The four projected modalities are concatenated and passed through a fusion block and a classification head that outputs an agitation probability (Figure 5a). We use a group formulation in which a single trunk is shared across all participants and a per-participant head reads out the fused representation, so that shared structure is learned jointly while each child retains an individualized decision boundary (Figure 5b). Modalities a given participant lacks (for example physiology for a child who wore no device) are masked to zero, and the per-participant head learns to weight the modalities that participant actually has. 3.3.2

Transfer configurations

We evaluate three ways of using the pretrained backbones. In the frozen configuration the foundation models are held fixed and their embeddings are precomputed once, and only the projections, fusion block, and head are trained. In the 9

Detecting Agitation Before Behavioral Escalation in Autistic Youth

(a) Multimodal encoder E

(b) Prediction head UniMTS embedding → 512-d concat

512+47

(i) Group · one shared trunk h → a head per child

proj → 128

Hand-crafted features → 47-d head P₁ 128→2

P(agitation)

h

⋮

EfficientAT embedding → 960-d concat

960+20

softma x(2 logits)

head P₁₅

proj → 128

128→2

Hand-crafted features → 20-d

concat 4×128 = 512

fusion MLP LayerNorm(512) Linear 512→128 ReLU · Dropout 0.3

UME embedding → 64-d concat

64+6

h 128-d

(ii) Individual · a separate trunk h + head per child

proj → 128

Hand-crafted features → 6-d

h

head · W

P(agitation)

128→2

softma x(2 logits)

PaPaGei embedding → 512-d concat

512+9

proj → 128

× 15

each child trains its own encoder E and head

Hand-crafted features → 9-d

foundation model

engineered features

read-out head

Projection block = Linear → LayerNorm → ReLU → Dropout(0.2) per modality. Each head is a linear layer (128 → 2 logits); softmax is applied only at read-out (cross-entropy / log-softmax in training). No softmax layer is stored in the model.

Figure 5: Multimodal foundation-model architecture. (a) Each 15 s modality window is encoded by a pretrained foundation model (UniMTS for movement, EfficientAT for audio, PaPaGei for PPG, and UME for electrodermal activity), concatenated with that modality’s engineered features, and projected to a shared 128-dimensional space. The four projections are fused into a shared representation h. (b) The two model forms differ only at the read-out head. The group model shares one trunk and uses a separate linear head per participant, selected by participant identity. The individual model trains a separate trunk and head per participant, sharing nothing. Each head outputs two logits, and softmax gives the agitation probability at read-out and no softmax layer is stored.

fine-tuned configuration the pretrained weights are adapted with low-rank adaptation [46]: the backbone weights stay frozen and small low-rank adapters are inserted into the convolutional and linear layers, so that only a few hundred thousand parameters are trained rather than the full backbones. In the from-scratch configuration, the backbones keep the same architecture but are randomly initialized and trained end to end, isolating the contribution of pretraining. We evaluate each configuration in both the group form above and an individual form in which a separate model is trained for each participant, giving a two-by-three comparison across {group, individual} and {frozen, from-scratch, fine-tuned}. 3.4

Experimental Design and Evaluation Protocol

We evaluate every configuration the same way, so the comparisons that follow differ in the model and not in how it was scored. This section covers the cross-validation protocol, the training setup, the metrics, and the specific configurations we compare. 3.4.1

Training and Testing Strategy

Time-series physiological and kinematic data are highly autocorrelated: windows close in time share signal, so randomly assigning windows to training and testing would place near-duplicate windows on both sides of the split and leak information across it [38, 40]. We therefore evaluate each participant with five-fold chronological block cross-validation. We order each child’s session timeline in time and partition it into five contiguous blocks. Each fold holds out one block for testing and trains on the other four, an 80/20 split, and every block is held out exactly once. We use five folds rather than a single 80/20 split because a participant’s signs of agitation are heterogeneous and spread unevenly across the session. Any one held-out block may contain some kinds of signs but not others. Rotating the held-out block through all five folds tests every window. Per-participant standardization statistics are computed from the training folds only, and all reported metrics come from these held-out predictions. We adopt this protocol rather than the random, class-balanced split behind the highest accuracies reported in this literature, which a recent review of wearable prediction studies in neurodivergent populations recommends against for this reason [37]. That split reports a much higher AUC on our own data, and we give the comparison in Section 4.7. 10

Detecting Agitation Before Behavioral Escalation in Autistic Youth

3.4.2

Model Training Details

All models were trained with the Adam optimizer and a class-weighted cross-entropy loss to account for the imbalance between agitation and non-agitation windows. In the fine-tuned configuration the low-rank adapters were trained at a learning rate of 1 × 10−3 ; in the from-scratch configuration the backbones were trained at 1 × 10−4 ; the projections, fusion block, and heads were trained at 1 × 10−3 throughout. Training used mixed-precision arithmetic. Because the fine-tuned and frozen configurations start from pretrained weights, they converge in a small number of epochs rather than the long schedules required to train comparable models from random initialization. 3.4.3

Performance Metrics

We report the area under the receiver operating characteristic curve (AUC) as the primary metric, because it is thresholdindependent and directly measures how well agitation windows are ranked above non-agitation windows. For each participant we compute AUC over that child’s held-out windows, and summarize across the cohort as the macro average, weighting every participant equally. We attach a 95% bootstrap confidence interval to every reported AUC, resampling windows within a participant for per-participant intervals and resampling participants for the cohort mean. To test whether the cohort mean exceeds chance, we use a within-participant label-permutation null: within each participant, agitation and non-agitation labels are shuffled, the full pipeline is re-run, and the observed macro AUC is compared against the resulting null distribution. 3.4.4

Experimental Configurations

Our experiments answer four questions about detecting agitation in this cohort. First, how well does the model detect agitation at the clinician-annotated onset, and does detection require the external pretraining or a separate model per child? Second, how far in advance can it detect agitation, evaluated by shifting labels earlier in 5 s steps up to 30 s? Third, which sensing modalities carry the signal, evaluated by dropping each in turn? Fourth, is a watch-only configuration sufficient, evaluated on the subset whose watch recorded those signals?

4

Results

We report all results over the 15 participants as a whole. Unless stated otherwise, numbers are macro AUC across participants for the group, fine-tuned model, with 95% bootstrap confidence intervals. 4.1

Detection at the Annotated Onset

At the clinician-annotated onset, the group fine-tuned model reached an AUC of 0.724 (95% CI 0.650–0.794; pooled across all windows, 0.785). The median across participants is 0.710, robust to the two participants whose models did not clear chance. A within-participant permutation test placed this well above chance. Thirteen of fifteen participants were individually above chance. Of the remaining two, one participant’s confidence interval crosses 0.5 and the other falls entirely below it (Figure 9; per-participant values in Table 3). 4.2

Pretraining and the Group Formulation

Table 2 and Figure 6 report the two-by-three comparison. Pretraining was valuable as the from-scratch group model reached only 0.583, while both frozen (0.708) and fine-tuned (0.724) foundation-model results were far higher. Finetuning with low-rank adapters gave a gain over frozen features that fell within the confidence interval. In the individual family the ordering even reversed (frozen 0.707 versus fine-tuned 0.695). The pretrained representations are therefore already close to their ceiling on this task, and adaptation adds little. A single group model was best overall (0.724), and its confidence interval overlaps the individual model’s in every configuration (Table 2), so the group-versus-individual differences are within sampling noise. Per-child training is therefore unnecessary for detection. 4.3

Detection Before the Annotated Onset

Figure 7 shows detection AUC as the labels are shifted earlier, for the best model of each family (group fine-tuned and individual frozen). Performance is strongest at onset and decays smoothly. The group model declines from 0.724 at 0 s to 0.632 at 15 s and 0.608 at 30 s; the individual model declines from 0.707 to 0.630 to 0.586. The group model is the stronger detector at the annotated onset and stays at or above the per-child individual model across the forecasting horizon, with the two nearly coinciding through the mid leads. Both stay above the 0.5 chance level throughout. We 11

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Group Individual

From-scratch 0.583 [.53,.64] 0.650 [.59,.71]

Frozen 0.708 [.64,.77] 0.707 [.64,.77]

Fine-tuned 0.724 [.65,.79] 0.695 [.63,.76]

Table 2: Detection AUC at the annotated onset across all 15 participants (macro average), with 95% participant-bootstrap confidence intervals, for the two-by-three comparison of {group, individual} × {from-scratch, frozen, fine-tuned}. Group and individual intervals overlap within each column, and the best cell (group fine-tuned) has a within-participant permutation p = 0.0005. Per-participant confidence intervals (Figure 9) are reported for the group model only.

tested significance at the annotated onset and report the earlier leads as point estimates. As in prior work, we cannot fully separate detection of an earlier state from temporal autocorrelation with the annotated window. 4.4

Which Modalities Carry the Signal

Dropping each modality in turn (Figure 8) shows that audio dominates. Removing it cost 0.126 AUC, while removing movement, electrodermal activity, or PPG each changed it by less than 0.02. Audio, which captures vocal expression through the lapel microphone, is the primary modality through which the model recovers agitation in this cohort. This is a cohort average, and it hides wide variation across children: the cost of removing audio ranges from 0.036 to 0.351 per participant, and the near-zero aggregate cost of the other channels averages over children who rely on them and children who do not. 4.5

Watch-Only Feasibility

On the 7 participants whose watch recorded electrodermal activity, photoplethysmography, and acceleration, a watchonly configuration reached only 0.560, far below the full multimodal model’s 0.699 on the same participants (Figure 10). Two of the 7 fell below chance. A low-profile wearable therefore provides at best a weak signal for this task, and the lapel microphone remains necessary for strong detection of individualized agitation. 4.6

Comparison with Baseline Methods

The from-scratch configuration serves as a deep baseline that shares the model’s architecture but forgoes pretraining, and it reached only 0.583, below both frozen and fine-tuned transfer. Pretrained representations, and the audio modality in particular, account for most of the detectable signal. 4.7

Effect of the Evaluation Split

All results reported above come from five-fold chronological block cross-validation. Under the protocol behind the highest performance in this literature, treating each window as an independent observation with balanced classes and a random split, the same model reaches an AUC of 0.994 and 97.0% accuracy on the same data. This difference depends the split strategy as adjacent windows share signal and a random assignment places near-duplicates of the held-out windows into the training set, scoring the model on windows it has effectively already seen (Section 3.4.1). We report the chronological result because it estimates performance on a session the model has not seen, and results obtained under record-wise splits are therefore not directly comparable to ours.

5

Discussion

We adapted four pretrained foundation models into a single shared model and evaluated how well it detects individualized agitation and how far in advance, along with how much each sensing modality and the pretraining itself contribute. 5.1

Interpreting Model Behavior and Individualized Prediction

The model’s behavior raises two questions worth unpacking: why detection quality varies so widely across children, and what the pretraining and parameter sharing actually contribute. 5.1.1

Heterogeneity Across Participants

Detection quality varied widely across participants (Figure 9), from AUCs above 0.9 for two children down to one whose detection fell below chance. Some of this variability tracks how each child expresses agitation (Table 4). Children 12

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 6: Detection AUC across all 15 participants for the two-by-three transfer comparison. Pretraining (frozen, fine-tuned) is far above from-scratch. The group model is best overall and matches per-participant models closely.

Figure 7: Detection AUC as labels are shifted earlier than the annotated onset, for the best model of each family. Both decline steadily but stay above chance through 30 s. The group model leads at the annotated onset and remains at or above the individual model across the horizon.

whose signs are clear and discrete tend to separate well from non-agitation windows, while detection is weaker for children whose signs are more subtle. The pattern has exceptions, and some children with overt signs still score poorly. For this analysis we label each annotated window by whether the agitation was primarily vocal, a discrete movement, or a diffuse “other” sign that is neither (e.g., subtle postural or affective changes, withdrawal). At the window level the split is clear: discrete movement and vocal signs are detected well (AUC 0.813 and 0.786), while diffuse “other” signs are much harder (0.731). This is the audio-dominance result from the other direction, since vocal signs land squarely in the microphone, the modality that carries most of the signal, while the diffuse signs that no channel captures cleanly are where the model struggles. Between children the same tendency does not always hold. A child’s number of diffuse signs correlates only loosely with their AUC (Spearman ρ = −0.40, not significant at fifteen children). For some participants, we get unexpected results: Participant 13, the one child below chance, showed loud grunts, crying, and forceful kicking and throwing rather than quiet signs, and Participant 14, among the most vocal children in the cohort, still scored well below the participants the model detected most reliably. For them the limit is not that their signs are subtle but more likely the quality of the negative class (both had the lower-arousal proxy baseline discussed below) and how well the sensors fit their particular behaviors. The sensor ablation shows the same divide per child. Removing audio costs a group-level 0.126 in AUC, but the per-participant loss ranges from 0.036 to 0.351. Nine of the fifteen lose more than 0.10 without audio, while for the rest movement and physiology carry the signal. The multimodal stack is therefore not redundant across the cohort. 13

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 8: Drop in detection AUC when each modality is removed. Audio accounts for most of the signal.

Figure 9: Per-participant detection AUC with 95% bootstrap confidence intervals (group fine-tuned). Thirteen of fifteen participants are above chance. One participant’s interval crosses 0.5 and one falls entirely below it.

A separate factor is the quality of the negative class. The eight children with a clinician-verified calm baseline averaged 0.776, against 0.665 for the seven whose negative windows came from a lower-arousal proxy (Mann–Whitney p = 0.15); part of the difficulty is label noise, not the child. What does not explain performance is who the child is or how much data they gave: detection AUC is uncorrelated with the number of annotated windows per participant (Spearman ρ = 0.23), with age (ρ = 0.31), and with sex (0.754 AUC for the four girls versus 0.713 for the eleven boys). None of these participant-level relationships, agitation sign type and baseline included, individually clears significance at fifteen children, so we treat them as a set of trends rather than established effects. A shared model suits most of these children, but not all. The group model beats its per-child counterpart for 11 of the 15 participants, by a small margin on average (0.02 AUC). The exceptions are children whose own build-up appears more informative than cohort-level structure. 14

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 10: Watch-only detection on the 7 participants whose watch recorded electrodermal activity, photoplethysmography, and acceleration. The watch-only configuration stays near chance (0.560) but well below the full model (0.699). 5.1.2

Score Separation and Failure Modes

Figure 11 shows the model’s output score for every window, split by the true state and grouped by participant. Each participant is one split violin. Its right half (blue) is the model’s score distribution on that child’s agitation windows, its left half (gray) on the non-agitation windows. The vertical axis is the score from 0 to 1, with the dashed line at 0.5. Participants are ordered left to right by decreasing detection AUC, the same ranking as in Figure 9, with each child’s AUC printed below the label. A child is separated cleanly when the blue half piles up near the top and the gray half near the bottom. Where detection is weakest, the two halves overlap across the range. The plot also separates two failure modes that a single AUC hides. For some children, such as Participant 6, the gray non-agitation half is pushed high as well, so the model over-calls calm periods as agitation even though the ranking is only partly broken. For others, such as Participant 13, the two halves overlap and the scores barely move with the true state, so the model is closer to indifferent than miscalibrated. The two call for different fixes: over-calling is a specificity and thresholding problem that a per-participant operating point can address, while indifference means the signal for that child is not being captured, which more data or a better sensor fit would have to solve first. 5.1.3

What Pretraining and Parameter Sharing Buy

Two results reshape how individualization should be approached here. First, the gap between the from-scratch and pretrained configurations shows that most of the usable signal lives in general-purpose representations pretrained on external corpora, not in structure the model could learn from this cohort alone. Fine-tuning them added little over using them frozen, and in the individual family frozen features were strongest. Second, a single group model with per-participant heads was the best detector, so the individuality of agitation is captured well enough by a shared trunk with a light per-child readout. Together these suggest that scaling to more children is a matter of adding heads to a shared model rather than building a new model for each. 5.2

Implications for Wearable Design

Our design targets comfort, acceptability, and coverage for everyday use. The full configuration combines an IMU garment, a physiology device, and a lapel microphone. The garment was selected for autistic participants (e.g., soft seams, predictable fasteners, and fixed sensor pockets) to support tolerance and repeatable placement, guided by a formative wearability study of fabrics, attachment options, and pocket designs (Section 3.1.4). Our ablation and watch-only results speak to device choice. Audio carried most of the detectable signal, and a watch-only configuration without a microphone stayed near chance (0.560 versus 0.699 for the full model on the same 15

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Figure 11: Per-participant separation of the model’s output scores at the annotated onset, agitation (blue) versus non-agitation (gray), ordered by detection AUC. Well-detected children show cleanly split distributions, and those it detects poorly overlap. Two failure modes are visible. Participant 6 pushes many non-agitation windows to high scores (false alarms), while for Participant 13 the agitation and non-agitation scores overlap almost entirely and its detection falls below chance.

participants). A low-profile physiology device is therefore insufficient for strong detection of individualized agitation in this cohort. 5.3

Clinical and Translational Implications

The clinical value of detecting agitation is in the time it buys. Challenging behavior is usually managed reactively, once an episode is already underway, and a signal that flags rising agitation even seconds before it is visible opens a window to change the situation first. Our advance-detection results are modest at longer horizons (Figure 7), but the response to an early sign is a small change in context and not a clinical procedure, so a brief and reliable alert is actionable. We discuss three settings where such an alert would be meaningful. The first is therapy, the setting our data came from. A BCBA already watches for a child’s individualized early signs and ends the evoking condition when one appears, so a monitor here extends an existing clinical practice. The same score that drives an alert is a continuous record of when agitation arose and how it responded to what the clinician did, and a functional assessment currently has to reconstruct that from memory and hand-coded observation. A BCBA is present and calibrating to that child, so this setting also tolerates the per-participant variation in Figure 11 that a fixed threshold cannot. The second is everyday life at home and at school, where most of a child’s day happens and where trained observation is least available. Caregivers and teachers come to know a child’s signs well, but they cannot watch continuously while handling their other responsibilities. Furthermore, the autonomic component of agitation is not visible to them at all. A wearable that runs through the day extends coverage into the hours outside of clinical observation, and it does so without asking an adult to divide their attention. It would also record escalation where it actually occurs, across the settings and demands that a clinic session can only approximate. That is the context a functional assessment needs and rarely has. The third setting, and the direction we regard as the goal of this line of work, is for the person to use to alert themselves. Agitation is a state the individual is experiencing. A system that reports only to other people leaves them the subject of monitoring. A wearable that surfaces rising arousal to the person wearing it could support self-regulation directly. It could prompt a break, or a coping strategy the person chose in advance. Over time it could build awareness of a 16

Detecting Agitation Before Behavioral Escalation in Autistic Youth

build-up that is hard to notice from the inside. That is what the technology is ultimately for: widening a person’s own capacity to act before escalation. The second and third settings are not reachable with the sensing system we used here. Our full configuration combines an instrumented garment, a physiology device, and a lapel microphone. The microphone is not optional, since the watch-only configuration stayed near chance (Section 4.5). Everyday use requires hardware that is miniaturized and robust across a full day of ordinary activity. It also has to be comfortable enough that a child keeps it on without an adult managing it. This work demonstrates that the signal is present and that a shared model can recover it, not that the present system is ready to leave the clinic. Across all three settings we see this as decision support and not automation. The model outputs a graded score, and what that score should trigger depends on the child and on who is reading it. A threshold that suits one child will over-call for another (Figure 11), so the operating point belongs to the person using it. A deployed system would surface a confidence to the clinician, the caregiver, or the individual it serves, and leave the decision with them. 5.4

Ethical Considerations

Ethical principles were central in the design and implementation of this study. The data collection protocol sought to prioritize safety, dignity, and comfort by focusing on non-dangerous agitation. Caregivers and participants were informed of their right to pause or withdraw at any time, and sessions were conducted under BCBA oversight with predefined safety plans. From a modeling perspective, our reliance on pretrained representations was a deliberate choice to reduce the amount of data required per participant, avoiding extensive procedures designed solely to elicit challenging behaviors. The emphasis on personalized, agitation-focused prediction is intended to support proactive, least-restrictive interventions rather than to enable punitive or coercive responses. At the same time, any broader deployment of such systems would require ongoing attention to data privacy, secure storage and transmission of physiological and behavioral data, and transparent communication with families and clinicians about what the system does and does not infer. Continued collaboration with autistic individuals, families, and clinicians will be essential to ensure that future systems remain responsive to the communities they are intended to serve. 5.5

Limitations and Future Work

Several limitations frame where this work should go next. The cohort, while larger than prior pilots, is still small and narrow in age, developmental profile, and topography, and performance is bounded by cohort size rather than model capacity. Because heterogeneity in this cohort is individual rather than clustered, scaling the dataset would improve the shared representation without reducing the per-child variation a deployment must absorb. The labels add their own uncertainty, since subtle shifts in posture, expression, or vocal tone are hard to timestamp precisely and blur the boundary between agitation and non-agitation windows near onsets, and how much annotation granularity affects performance is itself worth studying. The advance-detection result relies on an assumption that pre-onset windows already exhibit signs of agitation. Distinguishing genuine early detection from temporal autocorrelation will require continuous intensity ratings or prospective clinician review of model-flagged windows. Detection was predominantly audio-driven, performing weakest on the facial, postural, and withdrawal signs that microphones and upper-body IMUs capture poorly. Furthermore, the watch-only configuration performed near chance level across the seven children assessed, indicating that identifying sensing modalities better suited for these signs remains an open challenge. The gain from fine-tuning over frozen features was not statistically distinguishable, so the value here is the pretrained representations rather than task adaptation. Despite these limitations, this work produces a multimodal machine-learning system for the proactive detection of agitation, using data from wearable sensors and a custom annotation tool. The goal is to give BCBAs valuable intervention time by detecting agitation before it escalates into challenging behavior. Technology of this kind could improve the safety and efficiency of treatment by making escalation visible early enough to act on, whether to an interventionist, a caregiver, or eventually the person themselves (Section 5.3).

6

Conclusion

This work addressed the detection of agitation in autistic youth, a setting constrained by data scarcity and highly individualized expression. Rather than train a task-specific model per child, we adapted four foundation models pretrained on movement, audio, photoplethysmography, and electrodermal corpora, projected each to a shared representation, and fused them into a single group model with per-participant heads. This model was then evaluated within a clinically 17

Detecting Agitation Before Behavioral Escalation in Autistic Youth

grounded modified PFA/IISCA protocol. Across all 15 participants the model detected clinician-annotated agitation with an AUC of 0.724 (95% CI 0.650–0.794), and its detectability declined smoothly as the target was shifted earlier, to 0.608 at 30 s before the annotation. Pretraining was the primary driver of these results. A from-scratch model of the same architecture achieved only 0.58, whereas both frozen and fine-tuned representations yielded better performance. Audio was the dominant modality, and a watch-only configuration stayed near chance. A single shared model was the strongest detector. These findings establish that individualized agitation is detectable from wearable and audio sensing and that foundationmodel transfer captures it with a single shared model. We view this work as a step toward assistive technologies that augment clinical judgment rather than replace it.

7

Acknowledgement

We thank the children and families who generously participated in this research. We also thank Gabi Castillo-Martinez, Gabija Zilinskaite, and Adithyan Rajaraman for their essential contributions to completing this study.

8

Funding

This work was supported by the National Science Foundation (NSF) grant 2124002.

References [1] Centers for Disease Control and Prevention. Clinical testing and diagnosis for autism spectrum disorder. https://www.cdc.gov/autism/hcp/diagnosis/index.html, 2025. Clinician page summarizing DSM5-TR diagnostic criteria. [2] World Health Organization. Autism. https://www.who.int/news-room/fact-sheets/detail/ autism-spectrum-disorders, 2023. Fact sheet, updated Nov 15, 2023. [3] K. A. Shaw et al. Prevalence and early identification of autism spectrum disorder among children aged 4 and 8 years — autism and developmental disabilities monitoring network, 16 sites, united states, 2022. MMWR Surveillance Summaries, 74(2):1–24, 2025. [4] Stephen M Kanne and Micah O Mazurek. Aggression in children and adolescents with asd: Prevalence and risk factors. Journal of autism and developmental disorders, 41:926–937, 2011. [5] O Chadwick, N Walker, S Bernard, and E Taylor. Factors affecting the risk of behaviour problems in children with severe intellectual disability. Journal of intellectual disability research, 44(2):108–123, 2000. [6] Eric Emerson. Challenging behaviour: Analysis and intervention in people with severe intellectual disabilities. Cambridge University Press, 2001. [7] Susan L Parish, Kathleen C Thomas, Roderick Rose, Mona Kilany, and Paul T Shattuck. State medicaid spending and financial burden of families raising children with autism. Intellectual and Developmental Disabilities, 50(6):441–451, 2012. [8] Sara R Jeglum, Alexandra Cicero, Jordan DeBrine, and Cynthia P Livingston. Emergency department utilization due to challenging behavior in children and adolescents diagnosed with autism spectrum disorder. Behavioral Sciences, 14(8):669, 2024. [9] Richard M Foxx. Applied behavior analysis treatment of autism: The state of the art. Child and adolescent psychiatric clinics of North America, 17(4):821–834, 2008. [10] Ashley Fee, Elizabeth Schieber, Nathan Noble, and Maria G Valdovinos. Agreement between questions about behavior function, the motivation assessment scale, functional assessment interview, and brief functional analysis of children’s challenging behaviors. Behavior Analysis: Research and Practice, 16(2):94, 2016. [11] Mark O’Reilly, Mandy Rispoli, Tonya Davis, Wendy Machalicek, Russell Lang, Jeff Sigafoos, Soyeon Kang, Giulio Lancioni, Vanessa Green, and Robert Didden. Functional analysis of challenging behavior in children with autism spectrum disorders: A summary of 10 cases. Research in Autism Spectrum Disorders, 4(1):1–10, 2010. [12] Jessica D Slaton, Gregory P Hanley, and Katherine J Raftery. Interview-informed functional analyses: A comparison of synthesized and isolated components. Journal of Applied Behavior Analysis, 50(2):252–277, 2017. 18

Detecting Agitation Before Behavioral Escalation in Autistic Youth

[13] Luigi Iovino, Floriana Canniello, Roberta Simeoli, Maria Gallucci, Rosaria Benincasa, Davide D’Elia, Gregory P Hanley, and Anthony P Cammilieri. A new adaptation of the interview-informed synthesized contingency analyses (iisca): The performance-based iisca. European Journal of Behavior Analysis, 23(2):144–155, 2022. [14] Wayne W Fisher, Brian D Greer, Patrick W Romani, Amanda N Zangrillo, and Todd M Owen. Comparisons of synthesized and individual reinforcement contingencies during functional analysis. Journal of Applied Behavior Analysis, 49(3):596–616, 2016. [15] Gregory P Hanley, Brian A Iwata, and Brandon E McCord. Functional analysis of problem behavior: A review. Journal of applied behavior analysis, 36(2):147–185, 2003. [16] Nibraas Khan, Abigale Plunk, Zhaobo Zheng, Deeksha Adiani, John Staubitz, Amy Weitlauf, and Nilanjan Sarkar. Pilot study of a real-time early agitation capture technology (react) for children with intellectual and developmental disabilities. Digital Health, 10:20552076241287884, 2024. [17] Richard McCarty. The fight-or-flight response: A cornerstone of stress research. In Stress: Concepts, cognition, emotion, and behavior, pages 33–37. Elsevier, 2016. [18] Hugo D Critchley. Electrodermal responses: what happens in the brain. The Neuroscientist, 8(2):132–142, 2002. [19] Tal Barami, Liora Manelis-Baram, Hadas Kaiser, Michal Ilan, Aviv Slobodkin, Ofri Hadashi, Dor Hadad, Danel Waissengreen, Tanya Nitzan, Idan Menashe, et al. Automated identification and quantification of stereotypical movements from video recordings of children with asd. bioRxiv, pages 2024–03, 2024. [20] Ryan Anthony J De Belen, Tomasz Bednarz, Arcot Sowmya, and Dennis Del Favero. Computer vision in autism spectrum disorder research: a systematic review of published studies from 2009 to 2019. Translational psychiatry, 10(1):333, 2020. [21] Kristine D Cantin-Garside, Zhenyu Kong, Susan W White, Ligia Antezana, Sunwook Kim, and Maury A Nussbaum. Detecting and classifying self-injurious behavior in autism spectrum disorder using machine learning techniques. Journal of autism and developmental disorders, 50(11):4039–4052, 2020. [22] Fahd Albinali, Matthew S Goodwin, and Stephen S Intille. Recognizing stereotypical motor movements in the laboratory and classroom: a case study with children on the autism spectrum. In Proceedings of the 11th international conference on Ubiquitous computing, pages 71–80, 2009. [23] Matthew S Goodwin, Carla A Mazefsky, Stratis Ioannidis, Deniz Erdogmus, and Matthew Siegel. Predicting aggression to others in youth with autism using a wearable biosensor. Autism research, 12(8):1286–1296, 2019. [24] Joachim Taelman, Steven Vandeput, Arthur Spaepen, and Sabine Van Huffel. Influence of mental stress on heart rate and heart rate variability. In 4th European Conference of the International Federation for Medical and Biological Engineering: ECIFMBE 2008 23–27 November 2008 Antwerp, Belgium, pages 1366–1369. Springer, 2009. [25] Hayden Heath Jr and Richard G Smith. Precursor behavior and functional analysis: A brief review. Journal of Applied Behavior Analysis, 52(3):804–810, 2019. [26] Peter Washington, Aaron Kline, Onur Cezmi Mutlu, Emilie Leblanc, Cathy Hou, Nate Stockham, Kelley Paskov, Brianna Chrisman, and Dennis Wall. Activity recognition with moving cameras and few training examples: applications for detection of autism-related headbanging. In Extended abstracts of the 2021 CHI conference on human factors in computing systems, pages 1–7, 2021. [27] Mindy Scheithauer, Shruthi Hiremath, Audrey Southerland, Agata Rozga, Thomas Ploetz, Chelsea Rock, and Nathan Call. Feasibility of accelerometer technology with individuals with autism spectrum disorder referred for aggression, disruption, and self injury. Research in Autism Spectrum Disorders, 98:102043, 2022. [28] Riccardo Fusaroli, Anna Lambrechts, Dan Bang, Dermot M Bowler, and Sebastian B Gaigg. Is voice a marker for autism spectrum disorder? a systematic review and meta-analysis. Autism Research, 10(3):384–407, 2017. [29] Riccardo Fusaroli, Ruth Grossman, Niels Bilenberg, Cathriona Cantio, Jens Richardt Møllegaard Jepsen, and Ethan Weed. Toward a cumulative science of vocal markers of autism: A cross-linguistic meta-analysis-based investigation of acoustic markers in american and danish autistic children. Autism Research, 15(4):653–664, 2022. [30] Wen Ma, Lele Xu, Hao Zhang, and Shurui Zhang. Can natural speech prosody distinguish autism spectrum disorders? a meta-analysis. Behavioral Sciences, 14(2):90, 2024. [31] Marie-Michèle Dufour, Marc J Lanovaz, and Patrick Cardinal. Artificial intelligence for the measurement of vocal stereotypy. Journal of the experimental analysis of behavior, 114(3):368–380, 2020. [32] Hye-Geum Kim, Eun-Jin Cheon, Dai-Seg Bai, Young Hwan Lee, and Bon-Hoon Koo. Stress and heart rate variability: a meta-analysis and review of the literature. Psychiatry investigation, 15(3):235, 2018. 19

Detecting Agitation Before Behavioral Escalation in Autistic Youth

[33] John Allen. Photoplethysmography and its application in clinical physiological measurement. Physiological measurement, 28(3):R1, 2007. [34] Junyung Park, Hyeon Seok Seok, Sang-Su Kim, and Hangsik Shin. Photoplethysmogram analysis and applications: an integrative review. Frontiers in physiology, 12:808451, 2022. [35] Tales Imbiriba, Ahmet Demirkaya, Ashutosh Singh, Deniz Erdogmus, and Matthew S Goodwin. Wearable biosensing to predict imminent aggressive behavior in psychiatric inpatient youths with autism. JAMA network open, 6(12):e2348898–e2348898, 2023. [36] Bradley J Ferguson, Theresa Hamlin, Johanna F Lantz, Tania Villavicencio, John Coles, and David Q Beversdorf. Examining the association between electrodermal activity and problem behavior in severe autism spectrum disorder: A feasibility study. Frontiers in psychiatry, 10:654, 2019. [37] Patrick W Romani, Sidney K D’Mello, Robert M Moulder, and Lily N Berkowitz. Using wearable technology to predict the occurrence of severe behavior problems among neurodiverse individuals: A systematic review. Perspectives on Behavior Science, pages 1–23, 2026. [38] Nils Y Hammerla and Thomas Plötz. Let’s (not) stick together: pairwise similarity biases cross-validation in activity recognition. In Proceedings of the 2015 ACM international joint conference on pervasive and ubiquitous computing, pages 1041–1051, 2015. [39] Akbar Dehghani, Omid Sarbishei, Tristan Glatard, and Emad Shihab. A quantitative comparison of overlapping and non-overlapping sliding windows for human activity recognition using inertial sensors. Sensors, 19(22):5026, 2019. [40] Sohrab Saeb, Luca Lonini, Arun Jayaraman, David C Mohr, and Konrad P Kording. The need to approximate the use-case in clinical machine learning. Gigascience, 6(5):gix019, 2017. [41] Yadhu Kartha, Conor Anderson, Jenny Foster, Theresa Hamlin, Johanna Lantz, Ryan Lay, Juergen Hahn, Gari D Clifford, and Hyeokhyen Kwon. Prediction of challenging behaviors associated with profound autism in a classroom setting using wearable sensors. arXiv preprint arXiv:2605.17618, 2026. [42] Xiyuan Zhang, Diyan Teng, Ranak R Chowdhury, Shuheng Li, Dezhi Hong, Rajesh K Gupta, and Jingbo Shang. Unimts: Unified pre-training for motion time series. Advances in Neural Information Processing Systems, 37:107469–107493, 2024. [43] Florian Schmid, Khaled Koutini, and Gerhard Widmer. Efficient large-scale audio tagging via transformer-to-cnn knowledge distillation. In ICASSP 2023-2023 IEEE international Conference on acoustics, Speech and signal processing (ICASSP), pages 1–5. IEEE, 2023. [44] Arvind Pillai, Dimitris Spathis, Fahim Kawsar, and Mohammad Malekzadeh. Papagei: Open foundation models for optical physiological signals. In International Conference on Learning Representations, volume 2025, pages 48230–48261, 2025. [45] Leonardo Alchieri, Matteo Garzon, Lidia Alecci, Francesco Bombassei De Bona, Martin Gjoreski, Giovanni De Felice, and Silvia Santini. A foundation model for electrodermal activity data. arXiv preprint arXiv:2603.16878, 2026. [46] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3, 2022. [47] Ali Bahrami Rad, Tania Villavicencio, Yashar Kiarashi, Conor Anderson, Jenny Foster, Hyeokhyen Kwon, Theresa Hamlin, Johanna Lantz, and Gari D Clifford. From motion to emotion: exploring challenging behaviors in autism spectrum disorder through analysis of wearable physiology and movement. Physiological Measurement, 46(1):015004, 2025. [48] Sean M Montgomery, Nitin Nair, Phoebe Chen, and Suzanne Dikker. Introducing emotibit, an open-source multi-modal sensor for measuring research-grade physiological signals. Science Talks, 6:100181, 2023. [49] Bosch Sensortec. Intelligent 9-axis absolute orientation sensor. BNO055 datasheet, November, 2014. [50] Cheol-Hong Min and Ahmed H Tewfik. Automatic characterization and detection of behavioral patterns using linear predictive coding of accelerometer sensor data. In 2010 Annual International Conference of the IEEE Engineering in Medicine and Biology, pages 220–223. IEEE, 2010. [51] Ulf Großekathöfer, Nikolay V Manyakov, Vojkan Mihajlović, Gahan Pandina, Andrew Skalkin, Seth Ness, Abigail Bangerter, and Matthew S Goodwin. Automated detection of stereotypical motor movements in autism spectrum disorder using recurrence quantification analysis. Frontiers in neuroinformatics, 11:9, 2017. [52] Melissa MB Morrow, Bethany Lowndes, Emma Fortune, Kenton R Kaufman, and M Susan Hallbeck. Validation of inertial measurement units for upper body kinematics. Journal of applied biomechanics, 33(3):227–232, 2017. 20

Detecting Agitation Before Behavioral Escalation in Autistic Youth

[53] Nibraas Khan, Ruj Haan, Ingrid Shragge, Gabija Zilinskaite, Abigale Plunk, John Staubitz, Adithyan Rajaraman, Amy Weitlauf, and Nilanjan Sarkar. A universal web-based tool for multimodal data synchronization and labeling. In Masaaki Kurosu and Ayako Hashizume, editors, Human-Computer Interaction, pages 254–264, Cham, 2025. Springer Nature Switzerland. [54] Aparajita Saraf, Seungwhan Moon, and Andrea Madotto. A survey of datasets, applications, and models for imu sensor signals. In 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), pages 1–5. IEEE, 2023. [55] Zeeshan Ahmad and Naimul Khan. A survey on physiological signal-based emotion recognition. Bioengineering, 9(11):688, 2022.

21

Detecting Agitation Before Behavioral Escalation in Autistic Youth

Participant Events Group AUC [95% CI] Individual AUC 1 223 0.790 [0.75, 0.83] 0.806 2 241 0.916 [0.89, 0.94] 0.858 3 366 0.895 [0.87, 0.92] 0.853 4 81 0.730 [0.66, 0.79] 0.637 5 66 0.710 [0.64, 0.78] 0.642 6 333 0.683 [0.62, 0.74] 0.763 7 211 0.672 [0.62, 0.72] 0.636 8 81 0.810 [0.76, 0.87] 0.747 9 80 0.859 [0.81, 0.91] 0.830 10 64 0.920 [0.87, 0.96] 0.894 11 106 0.559 [0.49, 0.63]∗ 0.478 12 61 0.632 [0.54, 0.72] 0.562 13 58 0.394 [0.32, 0.47]∗ 0.507 14 196 0.607 [0.56, 0.66] 0.587 15 96 0.681 [0.61, 0.74] 0.802

Table 3: Per-participant annotated agitation event counts and detection AUC. Group AUC is the shared fine-tuned model with a 95% bootstrap confidence interval; Individual AUC is the best per-child model (frozen features), a point estimate. ∗ Participants 11 and 13 are not above chance: Participant 11’s interval crosses 0.5 and Participant 13’s falls entirely below it; all others are above chance.

A

Participant-Level Results

Table 3 gives the per-participant breakdown behind the cohort averages in Section 4. For each child it lists the number of annotated agitation events, the group fine-tuned model’s detection AUC with a 95% bootstrap confidence interval, and the best per-child individual model as a point estimate. The two AUC columns are close for most participants, which is the per-child form of the main result that a shared trunk with per-participant heads suffices. The exceptions run in both directions: Participant 15’s own model beats the shared one (0.802 versus 0.681), as does Participant 6’s (0.763 versus 0.683), while Participant 11’s is worse (0.478 versus 0.559). Participants 11 and 13 are the two children not above chance, with Participant 11’s interval crossing 0.5 and Participant 13’s falling below it.

22

Detecting Agitation Before Behavioral Escalation in Autistic Youth

B

Participant Details

Table 4 describes how agitation and challenging behavior presented for each participant, compiled from clinician input and session review. These profiles are the context behind the per-participant variation discussed in Section 5: children with discrete, high-energy signs tend to separate well from their calm baseline, while those whose agitation is subtle, postural, or withdrawn tend to sit closer, though this is a tendency with clear exceptions rather than a rule. We include the full descriptions so the per-participant results can be read against the specific behaviors each child showed rather than a single label. ID 1

Age (yrs) 10 yrs 1 mo

Sex M

2

11 yrs

F

3

11 yrs

M

4

3 yrs 5 mo

F

5

3 yrs 5 mo

M

6

13 yrs 1 mo

M

7

7 yrs 4 mo

M

8

5 yrs 10 mo

F

Identified Signs of Agitation Facial expression–mean look, squeezing hands/balled-up fist, tugging on shirt, tensing whole body, clicking mouth/odd vocalizations, "You’re annoying me," "I’m getting mad/angry," angry vocal tone, turns back or head down, pulling item away, pulling back from space/person, suddenly stops talking Frustrated/sad face, "No!", "I don’t have to do it", louder/firmer/faster speech, hand grabbing/squeezing, repetitive speech/scripting, moves away, pulls item away Louder vocalization/screaming, "First", angry face, slapping on iPad/surfaces, pushing materials away while saying "No", pulling away iPad Crying/whining, throwing materials, trying to grab removed object, clearing items, turning away, tug-of-war with item, grabbing/squeezing to stop interference, suddenly falling to knees, throwing materials, grabbing removed items Vocal stim, throwing arms down, walking away/in circles, reaching/vocalizing after denial, sad/distressed face, quick dip to side, ball-up/squeeze body, stomping, turning/hiding item, kicking/stomping feet Facial drop, jaw clench, grabbing adult’s hand, stern "stop it/no", pushing hand away, pulling item away, scooching away, forceful toe-jumping, vocal disruptions "No, no, no" (fast, hands up), "Wait, wait, wait", balled fists, widened eyes, walking into someone’s face, turning/twisting away, head/chin tics, pushing hand away Rapid breathing, body/muscle tensing, angry protesting tone, crying, eyebrows down, lips pushed together/down, folding arms/legs, turning back

23

Identified Challenging Behaviors Eloping, physical aggression (pushing, hitting, scratching, biting), property destruction, throwing items

Throwing, crying, hitting, scratching, rare tantrums

Throwing items, hitting, kicking, property destruction (throwing iPad) Self-injury (head-banging)

Self-injury (head hitting), squeezing others’ arms, crying/meltdowns

Self-injury (head hitting)

Eloping, kicking/shoving peers, yelling, head hitting (when sick)

Elopement, hitting others, property destruction

Detecting Agitation Before Behavioral Escalation in Autistic Youth

ID 9

Age (yrs) 3 yrs 7 mo

Sex M

10

17 yrs 5 mo

M

11

5 yrs 10 mo

F

12

5 yrs 9 mo

M

13

6 yrs 7 mo

M

14

10 yrs 8 mo

M

15

3 yrs 4 mo

M

Identified Signs of Agitation Screaming, loud noises, squinting/covering eyes, covering ears, "No, no, no", growling, pushing away, pulling item away/moving away, stomping/kicking, throwing item while looking for reaction Angry talking, banging on table, knocking head on fist, head down on arm, mumbling angrily, angry eyebrows, fists balled, knocking chair over, "You better not. . . ", "You do too much. . . ", insults with expletives, side-eye, hard breathing High-pitched whine, stomping feet, crying, snatching item, walking away/pushing person out of her space, crossing arms/pouting, squinting/scrunched eyebrows, tapping with heel of palm, shaking arm of person, side-eye High-pitched whine, pushing item/hand away, grimacing with teeth shown, tensing arms, single stomp, turning/moving away from adult/out of range Crying, loud grunts, eye contact then light slap, "Oh no!", withholding/protecting items, kicking doors, laying down and kicking, throwing items Yelling, "You need to go to Siberia", stomping, accusations ("Why did you hit me?"), loud/huffy breathing, fists balled, muttering, angry face, side-eye/eye rolling, growling, tensed shoulders, lunging/charging, foot tapping, snapping, kicking items, knocking items off table, protecting iPad, "No more questions please", "You want to replace me" Head down, covering ears, cutting off eye contact, pushing hand away, walking away, isolating self, facial grimace, "I need some space", grunting

Identified Challenging Behaviors Hair pulling, punching, biting, kicking, pushing

Punching holes in wall, verbal aggression, physical aggression

Self-injury (head hitting, flopping to ground)

Throwing self onto floor/wall (low intensity), fussing

Self-injury (chin hitting, head smacking), grabbing ears, clapping intensely when upset Physical aggression, property destruction

Elopement, self-injury (throwing body), pinching, biting

Table 4: Individualized agitation and challenging behavior descriptions for each participant.

24

Record · ID 1028693 · SHA-256 c98f79bf4af50b91
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.