2026-18-05
Towards a General Intelligence and Interface for Wearable Health Data
arXiv:2605.22759v1 [cs.AI] 21 May 2026
Girish Narayanswamy◦,†,1,3 , Maxwell A. Xu◦,†,1,5 , A. Ali Heydari‡,1 , Samy Abdel-Ghaffar‡,1 , Marius Guerard‡,1 , Kara Vaillancourt‡,1 , Zhihan Zhang‡,1,3 , Jake Garrison‡,1 , Levi Albuquerque‡,1 , Dimitris Spathis‡,1 , Hong Yu‡,1 , Hamid Palangi‡,1 , Xuhai "Orson" Xu1 , David G.T. Barrett2 , Joseph Breda1 , Jed McGiffin1,3 , Yubin Kim1 , Yuwei Zhang1 , Naghmeh Rezaei1 , Samuel Solomon1 , Karan Ahuja1 , Tim Althoff1 , Jake Sunshine1,3 , Ming-Zher Poh1 , Benjamin Yetton1 , Ari Winbush4 , Nicholas B. Allen4 , James M. Rehg5 , Isaac Galatzer-Levy2 , Yun Liu1 , John Hernandez1 , Anupam Pathak1 , Conor Heneghan1 , Yuzhe Yang1 , Ahmed A. Metwally1 , Pushmeet Kohli2 , Mark Malhotra1 , Shwetak Patel1,3 , Xin Liu △,†,1,3 and Daniel McDuff △,†,1,3 ◦ Co-first, △ Co-last, ‡ Core Contributor, † Corresponding Author, 1 Google Research, 2 Google DeepMind, 3 University of Washington,
4 University of Oregon, 5 University of Illinois Urbana-Champaign
While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging. Specifically, converting low-level sensor data into representations capable of characterizing higher-level states is difficult due to high phenotypic diversity and variation in individual baseline health, physiology, and lifestyle factors. Moreover, collecting wearable data paired with health outcome annotations is laborious and expensive, and retrospective annotation remains practically unfeasible, contributing to a scarcity of data with high-quality labels. To overcome these limitations, we propose a foundation model for wearable health that is pretrained on more than one trillion minutes of unlabeled sensor signals drawn from a large cohort of five million participants. We demonstrate that the joint scaling of model capacity and pretraining data volume leads to systematic improvements in performance, as evaluated on a diverse set of 35 health prediction tasks, spanning cardiovascular, metabolic, sleep, and mental health, as well as lifestyle choices and demographic factors. We find that this population scale representation unlocks label-efficient few-shot learning and generative capabilities for robust daily metric estimation. To further leverage this learned representation, we deploy a classroom of LLM agents to autonomously search the space of downstream predictive heads built on the model embeddings, showing broad performance improvements that increase with LLM model capacity. Finally, we show how integrating these downstream predictors into a Personal Health Agent can support model responses that are more relevant, contextually aware, and safe, and we validate this via 1,860 ratings from a cohort of clinicians.
1. Introduction Health is multidimensional. It comprises everything from the functioning of our many biological systems to abstract mental states to general well-being. Historically, both gathering and processing the volume of data required to provide individualized health insights to people have presented key technological bottlenecks in the pursuit of personalized healthcare. The large-scale adoption of wearable and mobile health technologies presents an unprecedented opportunity to broaden access to personal health insights and facilitate a shift towards preventive care. The data captured by these devices, in the form of continuous and longitudinally sampled sensor streams, now enables more accessible and accurate measurements of physical activity and behavior than ever before (McDuff et al., 2025a; Munos et al., 2016; Ringeval et al., 2020). Furthermore, a growing body of evidence suggests that these high-resolution data streams may be useful in disease detection and phenotyping (Metwally et al., 2026; Yang et al., 2022), early detection and monitoring (Perez et al., 2019) and the delivery of interventions (Shah et al., 2025). However, converting low-level sensor data into representations amenable to characterizing higher-level health states is challenging. Modern solutions largely take Corresponding author(s): {girishvn, xumax, xliucs, dmcduff}@google.com
Towards a General Intelligence and Interface for Wearable Health Data
the form of bespoke supervised algorithms for individual health outcomes. Unfortunately, these algorithms are hindered by their reliance on sensor data paired with health outcome annotations, which are laborious, expensive, and infeasible to source post-hoc. Prior work has often taken a piecemeal approach, utilizing only a few sensing modalities that target a small set of health endpoints. However, this approach falls short of being able to represent the high degree of phenotypic diversity at the population level, with heterogeneity in baseline health and physiology and disparate downstream health outcomes. As such, it is yet unclear whether generalizable features can be learned from wearable sensor data that capture information useful for diverse individuals and health applications. Foundation models, a recent advancement in machine learning, promise to address these limitations and accelerate progress in the field. These models often leverage self-supervised learning (SSL) over large-scale, heterogeneous, unlabeled datasets to learn universal representations which can generalize to a broad array of tasks. In the domain of health, these methods have successfully enabled improved performance on diverse applications, including radiology (Wu et al., 2025), pathology (Xu et al., 2024), and medical reasoning (McDuff et al., 2025b; Sellergren et al., 2025). Adjacently, early time-series foundation models (Garza et al., 2023; Rasul et al., 2023) demonstrated the utility of large-scale pretraining for signal forecasting, while subsequent families of models have emphasized the utility of learning common spectral properties across disparate domains (Ansari et al., 2024; Das et al., 2023; Goswami et al., 2024). The surprising capability of large-language models (LLMs) for time-series analysis further emphasizes the utility of scaled generalized pretraining across domains (Liu et al., 2023; Merrill et al., 2024; Thukral et al., 2025). In the wearable domain, recent efforts have demonstrated the potential to learn robust representations from large corpora of multimodal sensor data employing SSL (Abbaspourazad et al., 2024; Thapa et al., 2024; Yuan et al., 2024). A pivotal advancement has been the recent establishment of scaling laws for wearable data and the introduction of architectures capable of learning directly from incomplete sensor signals (Narayanswamy et al., 2025; Xu et al., 2025). Yet, despite this progress, it remains unclear the full extent to which scaled pretraining on wearable data equates to meaningful improvements in predictive performance for diverse health outcomes and insights. Furthermore, current approaches remain constrained by two critical bottlenecks: the finite scale of pretraining data and the extensive manual engineering required to adapt a single generalist embedding to many distinct health endpoints. In this work we aim to prove that leveraging large, unlabeled streams of continuous wearable sensor data during pretraining can lead to predictable improvements in both pretraining and downstream task objectives, and that resulting representations can be efficiently adapted to many outcomes. Building on this, we aim to demonstrate that providing such a model as an inference engine to an agentic health coach is more effective than having the coach process only the wearable data directly. To that end we introduce SensorFM (Fig. 1), a Large Sensor foundation Model for wearable time-series representation learning which exhibits generalizability across health domains. By scaling pretraining to an unprecedented corpus of over one trillion minutes (1, 000, 000, 000, 000) of sensor data drawn from five million participants ( 𝑁 = 5, 000, 000) and five sensor modalities, we approach a highly adaptable, universal representation of sensed human physiology. To our knowledge, this is the largest and most diverse wearable dataset utilized to date (Erturk et al., 2025; Narayanswamy et al., 2025; Yuan et al., 2024). We evaluate SensorFM across a comprehensive suite of downstream health tasks that span cardiovascular health, metabolic risk, sleep disorders, mental health, lifestyle choices and physiologically relevant demographics. We validate our model using rigorously phenotyped datasets derived from controlled clinical and laboratory studies ( 𝑁 = 13, 985). We further explore the capabilities of SensorFM in reconstructing missing data and the resultant implications for daily health metric estimation. We comprehensively characterize the model’s capabilities through rigorous
2
Towards a General Intelligence and Interface for Wearable Health Data
a Data
Downstream Data Study 2 Mental Health
Pretraining
1 Trillion Minutes of Sensor Data
5,953 People Clinical Screeners
5 Million People 100+ Countries 50 US States 20+ Devices
b Pretraining
Study 3 Sleep Medicine
1,655 People Lab Tests & Self-Reported Diagnoses
6,377 People Patient Reported Outcomes
c Generation
Original Input
Accelerometer
Study 1 Metabolic Health
Altitude
Baseline
Our Model
Normal Wrist Temp. (<37 C)
Embedding
High Wrist Temp. (>=37 C)
EDA
High SPO2 (>90%) Minutes Anaerobic Exercise (152<=HR)
Decoder
Encoder
HR/HRV
Low SPO2 (<90%) Minutes
Aerobic Exercise (114<=HR<152)
SpO2
Light Exercise (95<=HR<114) Minutes Sleep Stage REM
Temperature Sleep Stage Deep
Sleep Stage Light
Sleep
Steps
Prediction Heads
Age
ASCVD Risk
24 Hrs
Smoking
…
0
d Finetuning
4
e Scaling
70
Relative Improvement (%)
1 2 3 Percent Error in Daily Estimate
Classi�cation
Regression
60 50 40 30 20 10
Lifestyle
Cardiovascular
Mental
Sleep Impairment
Sleep Disturbance
PSS
Sleep Disorder Treat.
PHQ-8
GAD-7
Persistent Stress
Mental Health Med.
Depress./Anxiety Dx
Mild Anxiety
Mild Depression
HbA1c
HOMA-IR
Metabolic
Triglycerides
Insulin Resistance
Pre-Diabetes
Diabetes Med.
Hyperlipidemia
Diabetes Dx
Framingham 30 Risk
ASCVD Risk
Framingham Risk
Respiratory Dx
No Medications
Hype�ension Dx
Cardiovascular Dx
Smoking
Medicaid
Dis. A�ects Work
Weight
Demographics
Disability
Currently Working
BMI
Height
Age
0
Sleep
Figure 1 | Scaling and Evaluating a Sensor Foundation Model (SensorFM) for Wearable Health. We present a versatile embedding model that scales with model and data capacity, and shows generalizability to a range of generative and discriminative tasks. (a) We pretrain this model on an unprecedented corpus of over one trillion minutes of sensor data drawn from five million participants and evaluate it on an independent set of data from 13,985 people featuring 35 clinical and behavioral discriminative tasks derived from three prospective studies. (b) The model is trained with a generative reconstruction objective with a latent “bottleneck” on features derived from five sensor modalities. We evaluate the model and present results on a set of (c) generative tasks – here the baseline represents daily estimates without generative infilling, and (d) predictive tasks – here we show relative performance improvement over a supervised model trained on engineered features. (e) Aggregated performance scores for pretraining and posttraining tasks show a linear correlation as pretrain data and model capacity are co-scaled by orders of magnitude, illustrating that a reconstruction based pretraining leads to scalable improvements in downstream tasks. 3
Towards a General Intelligence and Interface for Wearable Health Data
evaluations of its scaling, label-efficiency, and interpretability. Furthermore, to push the upper bound of predictive performance for diverse health applications, we leverage an automated agentic framework, inspired by modern self-improving code-generation systems (Aygün et al., 2025; Novikov et al., 2025), to optimally adapt the SensorFM embeddings to individual downstream tasks (Figure 2). Traditionally, adapting general representations to a variety of downstream applications has required bespoke architecture engineering to train domain-specific models on the embeddings. In contrast, we provide an agentic “classroom” where LLM agents iteratively generate, test, and refine the code to develop models on downstream tasks using these embeddings as a starting point. We demonstrate how such a system enables more scalable and systematic exploration of the downstream solution space and leveraging this method autonomously conduct over 30, 000 individual experiments. We evaluate the benefit of allowing agents to play the role of machine learning engineers and analyze the agent-derived solutions. Finally, given the significant adoption of AI language models for consumer health queries (Breda et al., 2026; McDuff et al., 2025b; Sumner et al., 2025; Tu et al., 2025), we establish the model’s endto-end utility by integrating SensorFM as tool into a Personal Health Agent (Heydari et al., 2025) and evaluate its utility to enhance the delivery of context-aware physiological insights to users (Figure 3). We conduct over 40 hours of clinical evaluations of health summaries generated with wearable data and either SensorFM predictions or gold-standard ground-truth measurements involving 1,860 individual ratings. When compared to a baseline in which the language model processes the wearable data directly, using SensorFM as a tool improves the specificity of the responses and makes them more personalized, contextually appropriate and safer. When compared to a condition in which the agent has access to ground-truth measurements we observe no statistical inferiority. In summary, our work represents the most comprehensive evaluation of pretrained wearable sensor foundation models to date, demonstrating how flexible embeddings, produced through scaled pretraining, can be efficiently fine-tuned for many downstream health applications, which can in turn be leveraged to provide valuable insights at the person level.
2. Results We first characterize the properties and evaluate the performance of SensorFM along four dimensions: resource scaling, predictive performance across discriminative health tasks sourced from multiple prospective studies, generative capabilities for infilling and forecasting sensor data, and finally interpretation of the model embeddings in a latent space. We then present results from an agentic search method used to automatically design and refine application-specific prediction “heads” and evaluate this system across classification and regression-based health tasks. Finally we evaluate SensorFM as a tool for a Personal Health Agent and recruit a cohort of clinicians to evaluate the utility of the model’s predictions when creating health summaries for a set of 31 real (i.e., non-synthetic) health profiles. 2.1. Scaling a Generalist Model for Wearable Health The establishment of scaling laws, and the resultant success of foundation models in domains such as language and vision, has shown that model performance is often driven not by architectural design, but rather by compute, model capacity, and ultimately the volume of training data (Kaplan et al., 2020; Zhai et al., 2022). Scaling laws provide empirical evidence that can be used to help anticipate the performance gains that could potentially be achieved if any (or all) of these resources are increased. Initial empirical evidence of scaling has been documented for time-series and wearable sensor signals 4
Towards a General Intelligence and Interface for Wearable Health Data
d Classroom Learning Progression
Figure 2 | Intelligent Search of Prediction Heads. We employ an LLM-driven architecture to efficiently search the space of solutions for each downstream task, mimicking the role of a machine learning expert. Specifically, (a) given a dataset of wearable embeddings, demographic features, and labels, (b) an instruction prompt, (c) a “classroom” of collaborative/competitive agents iteratively refines executable code solutions. (d) the learning progress for five “student” agents. 5
Towards a General Intelligence and Interface for Wearable Health Data
a SensorFM-augmented Health Agent
Final
Query
How can I improve my health?
Focus on increasing cardiovascular intensity and managing metabolic markers, as your predictive profile suggests a biological age higher than your actual 38 years and flags potential risks for hypertension and insulin resistance.
Gemini
Predictions:
SensorFM
Backbone & Downstream Heads
…
Predicted Age: 46 Predicted BMI: 30 p(Hypertension): 90% p(Insulin Resistance): p(Sleep Disturbance): p(Sleep Impairment): … p(Anxiety): 15%
85% 10% 12%
Wearable Data b Clinician Evaluation Experiment
93 Responses
Panel of Clinicians
Demo + Daily + Truth
...
31 Participant Profiles
Demo + Daily + SensorFM
Rubric: Context Harm Relevance Justifiability Personalization
Demo + Daily
1,860 Ratings c Clinical Evaluation Results
Context
Personalization
Harm
Justifiability Relevance
Figure 3 | Agentic Use of SensorFM as a Tool. (a) Architecture of the SensorFM-augmented workflow, which translates raw wearable sensor data into health predictions to provide context for the LLM’s response. (b) We extracted demographic information, aggregated fitbit metrics, model predictions and ground-truth data for 31 real patient profiles and used Gemini to generate health summaries. A panel of four clinicians evaluated the responses. (c) Average physician evaluation scores (Likert scale) plotted across five specific clinical rubric items (Harm, Context, Personalization, Justifiability, and Relevance) for the three evaluated conditions and stacked horizontal bar charts showing counts of Likert ratings across all rubric items. Example output is for illustrative purposes only. 6
Towards a General Intelligence and Interface for Wearable Health Data
(Narayanswamy et al., 2025; Shi et al., 2024; Zhang et al., 2025). Yet, a systematic set of scaling experiments on wearable sensor data demonstrating that progressive pretraining gains predictably translate to measurable improvements in estimating meaningful health outcomes is still absent. In this work we substantially increase the scale of pretraining and compute beyond previous efforts. Specifically, we scale over four orders of magnitude for both pretraining data volume (two million to two billion hours of multimodal sensor data) and model size (100K to 100M parameters). Our maximum data volume is a 50x increase in data-hours over and above the 40 million hours used by prior work for models trained on minute-resolution data (Narayanswamy et al., 2025). Compared to the 2.5 billion hours used to train models on hour-resolution data (Erturk et al., 2025), our higher minutely resolution data accounts for a 50x increase in total data volume. Throughout our analysis we refer to data volumes by the number of individuals sampled or the associated multimodal data hours: 5K (2 × 106 hrs.), 50K (2 × 107 hrs.), 500K (2 × 108 hrs.), 5M (2 × 109 hrs.). We refer to SensorFM model variants by their capacities: XXSmall (105 params.), XSmall (106 params.), Small (107 params.), Base (108 params.). We evaluate the effect of this scaled pretraining on classification, regression, and generative tasks. The Importance of Scaling Pretraining Data and Model Capacity. We observe that the pretraining validation loss inversely scales with increases in data volume and model capacity (Table ED.5). In so doing, we verify that scaling these resources leads to predictable improvements in model performance and that similar improvements are observed for both discriminative and generative downstream tasks. For example, when pretrained with the largest 5M subject data volume, SensorFM-B consistently outperforms the smaller SensorFM-XXS. The scaled model achieves a 31% reduced validation loss (MSE) on the reconstruction pretraining task, and a 28% reduced loss (avg. MSE) across generative tasks. On discriminative tasks, the scaled model achieves a mean improvement of Δ𝐴𝑈𝐶 = 0.09 on classification tasks and Δ𝑟 = 0.21 on regression tasks. Crucially, we observe that the most significant gains are achieved through the joint scaling of both data volume and model capacity. The proportional scaling of data and capacity by orders of magnitude, leads to near-linear improvements in both generative pretraining and discriminative post-training performance (see Figure 1e). Driven by this finding, all following results, unless explicitly stated, assume that models are trained with data volumes proportionally scaled to their capacity. The impact of this joint scaling on discriminative health tasks is further visualized in Figure 4 and presented in Tables ED.6 where across model variants, SensorFM-B boasts a task win-rate of 33/35, while XXS expectedly ranks last on 33/35 tasks. 2.2. Learning a Representation Useful Across Health Domains We evaluate the SensorFM-learned embeddings across a diverse range of 35 discriminative health tasks derived from multiple prospective studies (see Methods M.2.2). These tasks span Cardiovascular Health (6), Metabolic Health (8), Mental Health (8), Sleep (3), Demographic (4), and Lifestyle Factors (6), with the full list of tasks found in Table ED.3. In order to interrogate the quality of the pretrained embeddings, we leverage a frozen SensorFM encoder and learn a computationally efficient linear head to adapt the embeddings to individual applications. To account for the limited number of annotated examples, these heads are trained with embeddings reduced to 50 principal components. We baseline SensorFM against supervised models trained with engineered features derived from the wearable sensor streams (see Methods M.3.6). We further assess the lift of the learned sensor representation, by training models both with and without demographic features, comparing against baseline supervised models trained only with demographic features. The Learned Representation Generalizes to Diverse Health Outcomes. We find that Sen-
7
Towards a General Intelligence and Interface for Wearable Health Data
sorFM, through scaled pretraining, learns a representation capable of successfully generalizing to a broad range of health outcomes. As illustrated in Table ED.9, linear heads trained on-top of the SensorFM learned embeddings consistently outperform supervised baselines trained with engineered features. Specifically, SensorFM outperforms this supervised baseline on 34 of 35 discriminative tasks (Figure 1.d). SensorFM outperforms a baseline trained on only demographic features on 24 of 30 discriminative tasks. The Utility of Demographic Features. As highlighted in Table ED.9, we find that the predictive power of SensorFM is often, though modestly, enhanced through the addition of demographic features (22 of 30 discriminative tasks). While demographic features tend to improve the performance of SensorFM, interestingly we find that the dependence of SensorFM on demographic features decreases with scale. SensorFM-B realizes smaller gains from added demographic features as compared to both smaller model variants and supervised baselines on 33 of 35 tasks, implying that these physiologically relevant traits may be implicitly learned through pretraining at scale (Table ED.8). A similar trend is observed in the feature importance attributed to embeddings when adapted to discriminative downstream tasks alongside demographic features. Models pretrained at scale provide a more robust representation which reduces the reliance on demographic priors (Figure ED.9.b). Additionally, for some tasks (e.g., cardiovascular Dx, insulin resistance, ASCVD risk, Framingham risk, and more), we find that models trained with demographics alone provide significant predictive power, (Figure 4). For such tasks, SensorFM may only outperform these demographic baselines at extreme pretraining scales (e.g., SensorFM-B trained on 50M weeks of data). For a subset of these tasks (ASCVD Risk, Framingham Risk, Framingham 30 Risk), the strong performance of demographic-only models likely stems from the explicit use of demographic features to calculate these risk scores. Scaled Pretraining Enables Label Efficient Adaptation. We find that scaled pretraining and the resultant learned representation of SensorFM enables improved label efficiency compared to supervised baselines. As depicted in Figure ED.5 we tested this by training models with varying percentages of downstream training volumes. With very few labeled samples, demographic priors act as a strong predictor for many tasks. However, as the number of labeled samples increases, SensorFM soon outperforms demographic-only baselines and consistently outperforms the feature engineered baselines, with larger model variants (B) outperforming smaller variants (XXS). 2.3. Exploiting the Learned Generative Capabilities SensorFM leverages a reconstructive pre-text task which enables learning directly from unlabeled, passively collected sensor data. Leveraging an MAE-like architecture (He et al., 2022; Xu et al., 2025), our method natively handles the missingness inherent in wearable sensor streams, a consequence of varying sensor configurations, operating modes, and user behaviors. During pretraining, SensorFM learns a decoder capable of reconstructing ablated observations, which in-turn translates to generative capacities such as data imputation and signal forecasting. Figure ED.6 presents examples of day-long windows obtained from participants in our pretraining validation set, comparing the original model input to the reconstructed output. Figures ED.7 and ED.8 depict line graphs for individual sensor feature infilling and highlight the non-linear dynamics within the SensorFM reconstructions. SensorFM Learns to Fill and Forecast Sensor Data. SensorFM, through its generative pretraining, learns to successfully impute, interpolate, and extrapolate missing or unobserved data (see Table ED.10). Specifically, we find that SensorFM outperforms the best-performing baselines by 74.8% on random imputation, 38.8% on temporal interpolation, 39.6% on temporal extrapolation, and 83.7% on sensor signal imputation. Improved Daily Metric Estimation. Wearable sensor feeds may be intermittently interrupted 8
Towards a General Intelligence and Interface for Wearable Health Data
Classification Tasks
Regression Tasks
Currently Working
Age
Disability
BMI
Disability Affects Work
Height
Smoking
Weight
Medicaid
ASCVD Risk
No Medications
Framingham Risk
Cardiovascular Dx
Framingham 30 Risk
Hypertension
HOMA-IR
Respiratory Dx
HbA1c
Diabetes Dx
Triglycerides
Diabetes Med
PHQ-8
Hyperlipidemia
GAD-7
Pre-Diabetes
PSS
Insulin Resistance
Sleep Disturbance PRO Sleep Impairment PRO 0.0
Mild Depression Mild Anxiety
0.2
0.4 0.6 Pearson Correlation
Persistent Stress Depression/Anxiety Demo only XXS (105)
Mental Health Med. Sleep Disorder Treatment 0.5
Model (Parameters)
Demographics Lifestyle
0.6
0.7 0.8 ROC AUC
0.9
XS (106) S (107)
Task Category Cardiovascular Metabolics
0.8
1.0
B (108) B (108) & Demo Mental Health Sleep
1.0
Figure 4 | Discriminative Task Linear Probe Performance. Downstream performance across 35 discriminative tasks for SensorFM variants, pretrained with proportional data scales, and a supervised baseline trained with only demographics. In general performance improves with scale with B consistently achieving the best performance. SensorFM variants are post-trained with PCA-50 reduced embeddings. For each task, we report the average 5-fold cross validation performance. Average Receiver Operating Characteristic Area Under the Curve (ROC AUC) is calculated in the logit-transform space and back-transformed. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Error bars are standard deviations calculated in the transformed space and back-transformed to give asymmetric error values. 9
Towards a General Intelligence and Interface for Wearable Health Data
for a variety of reasons. Such interruptions can significantly skew an individual’s health summary statistics. As such, it may be advantageous to provide individuals with more accurate estimates of their summary statistics allowing them to better gauge their overall health status. Towards this end, we explore the potential of leveraging SensorFM’s generative capabilities to impute missing data in order to more realistically estimate a person’s daily metrics. Specifically, SensorFM leverages temporal interpolation to infill missing segments. As highlighted in Table ED.11, we find that in the presence of missing or ablated data, SensorFM is able to produce more reliable daily metrics. Specifically, we show that when ablating 60 contiguous minutes of data in a day, SensorFM retains 99.7% accuracy in daily step count prediction, 99.9% accuracy in deep sleep prediction, and 99.2% accuracy in light exercise tracking, mitigating the underestimation observed in baseline performance (Figure 1.c). 2.4. Understand and Quantifying the Learning Latent Space Analysis and visualization of the SensorFM embeddings in latent space provide valuable insight into the learned representation, its structure, and its application to downstream tasks. Towards this end, we project and visualize the latent space for a number of health outcomes, analyze the embedding distances and intrinsic dimensionality of the model across scales, and interpret the importance of the embeddings with respect to downstream health outcomes. Visualizing Embeddings with Task Labels. We use Uniform Manifold Approximation and Projection (UMAP) (McInnes et al., 2018) to reduce the high-dimensional latent embedding vectors from our largest model down to two dimensions (Figure ED.10). The manifold of the model provides insight into the information learned during pretraining in the absence of explicit labels. We visually annotate these points with labels from different downstream datasets. Note that although all plots use the same UMAP projection, not all participants/days have labels for all tasks. There are clear patterns of demographic shifts for BMI, Age, and Gender captured within the embedding. Ultimately, these visualizations underscore how self-supervised pretraining at this scale naturally organizes fragmented sensor streams into a physiologically meaningful topology, validating a foundation model approach as a universal representation of sensed human health. Embedding Distances. To investigate how model scale influences latent space density, we compare the dispersion of user representations across different model sizes (Figure ED.11.a), an approach to quantifying participant similarity as seen in previous works (Kiyasseh et al., 2021). We find that while all models yield unimodal, right-skewed distance distributions, their latent space dispersion varies significantly. The SensorFM-S model learns the most tightly clustered representations, whereas the SensorFM-B model produces the broadest embedding spread. The smallest model (SensorFM-XXS) also exhibits a similarly broad spread, pointing to an interplay between optimal model capacity and data volume in shaping representation density. This structural evolution of the latent space confirms that the joint scaling of data and parameters is crucial for yielding an embedding space expressive enough to capture the nuanced, inter-subject variations across diverse clinical domains. Intrinsic Dimensionality and Compressibility. We evaluate the compressibility and informational structure of the embeddings across model scales (Figure ED.11.b). An analysis reveals that the model embeddings are highly compressible; particularly for the larger models, representations can be reduced to 150-200 dimensions without significant loss of variance. Furthermore, the variance scaling behaviours differ significantly across model sizes. The smallest model (XXS) captures approximately 90% of the variance within its first 20 principal components and then flatlines, indicating dimensional collapse—an over-reliance on a restricted feature manifold. Conversely, the largest model B shows strong anisotropy. It learns a large "super-feature" in its dominant direction, with the first principal component alone explaining approximately 40% of the total variance. After this initial spike, the curve flattens out, indicating that while the large model relies heavily on a dominant primary component, it 10
Towards a General Intelligence and Interface for Wearable Health Data
importantly preserves a significant "long tail" of nuanced physiological information distributed across higher dimensions. Consequently, large-scale sensor models provide a powerful dual advantage: they distill large, continuous streams of wearable data into highly efficient, compressible interfaces for downstream modelling, while crucially preserving the subtle, long-tail signals essential for predicting heterogeneous outcomes such as mood and sleep disorders. Interpreting Embedding Importances. In order to better understand the structure of the latent embedding vector, we leverage SHapley Additive explanations (SHAP) (Lundberg and Lee, 2017) to assess the effect of individual latent dimensions. We calculate SHAP values for each linear prediction head on the model and then compute the pair-wise 𝑐𝑜𝑠𝑖𝑛𝑒 similarity between the normalized, exact SHAP attribution (weight collapse analysis) profiles of each pair of tasks. This indicates the degree to which distinct tasks leverage the same underlying embedding dimensions from the foundational model (Figure ED.9.a). Expected similarities in embeddings are observed with highly correlated labels such as ASCVD risk and Framingham risk, and weight and BMI. However, we observe links between other tasks, including sleep impairment and PSS score, and HOMA-IR and PHQ-8 score. 2.5. Agent Driven Search of Downstream Model Heads Versatile pretrained model embeddings which generalize to many predictive tasks are attractive as they enable application-specific models to be built more easily, especially when labels are sparse. As such, methods which efficiently design new prediction “heads” allow these embeddings to be more rapidly adapted to novel tasks. A bottleneck to effectively adapting embeddings for downstream predictive tasks is often the expertise and iterative work required for feature engineering, model architecture selection, and hyperparameter tuning. This is especially true in health domains, where observational data is typically subject to significant constraints (e.g., sparsity, noise, limited volume, and imbalance of labels). We implement a hybrid modeling system, adapted from Aygün et al. (2025), to leverage the reasoning and code-writing abilities of language models alongside an iterative-solution-search framework to autonomously adapt general embeddings to new domains. Specifically, we leverage a "classroom" configuration (described in M.5), a set of self-evolving algorithm generation steps in which the solution synthesis is formulated as a competitive, collaborative optimisation problem solved by parallel LLM agents (Figure 2). Leveraging this framework we rapidly iterate over 30, 000 agent-proposed solutions. Improvements Across Health Tasks. We find that this agentic approach leads to improved performance as compared to a simple linear head applied to the SensorFM embeddings across a breadth of discriminative health tasks. Specifically, classroom-derived agent solutions boast improved performance on 16 of 20 classification tasks and greater Pearson correlations on 12 of 15 regression tasks (Figure 5 and Table ED.12). Note that we report F1 for these classification results as many solutions were ensemble methods from which it is not possible to obtain a continuous output with which to compute an ROC curve. In so doing we demonstrate the potential of AI agents to act as machine learning scientists, reducing the engineering burden traditionally associated with adapting general embeddings to multiple endpoints. Solution Performance Scales with LLM Capabilities. Analyses of the agent solutions organized by LLM model variant reveals an interesting pattern, with more recent models exhibiting stronger performance (see Figure ED.12.a). When these models are characterized by the commonly used Artificial Analysis Intelligence Index1 , models with higher intelligence indexes provide better solutions on average. Furthermore, we find that collaboration events between “students” (agents), triggered 1 https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
11
Towards a General Intelligence and Interface for Wearable Health Data
when a given agent demonstrates plateauing performance and is allowed to reflect on its own solutions or the solutions of other agents, enables less intelligent models to close this performance gap. However, the best performance is typically observed from more recent versions of Gemini (Comanici et al., 2025). Analysis of the Found Solutions. Meta analysis of the best solutions found through the classroom algorithm search (see Figure ED.12.b) reveals that almost all top-scoring solutions reduced the dimensionality of the embedding feature space to between 50 − 100 dimensions, most likely to reduce the variance of the input space to match the scarcity of labeled examples. The results additionally reveal that linear models were more common than non-linear models. Ensembles were employed in just under a quarter of the best solutions. While manually searching this space of solutions is tractable, it is time consuming and often inefficient, becoming increasingly less feasible as the number of downstream prediction tasks increases. By contrast, this agent driven approach allows for efficient iteration, with the average quality of final solutions improving monotonically over time (Figure 2.d). 2.6. Agentic Use of the SensorFM as a Tool A critical open question is whether providing SensorFM, adapted to multiple endpoints, as an inference engine to a personal health agent, yields improvements in the responses to user health queries as compared to the agent reasoning limited to handcrafted features. Towards this end, We evaluate the performance of a health agent under three distinct experimental conditions: • Condition (A): Demographics + Daily Wearable Metrics + SensorFM Predictions Agent receives demographics, feature-engineered daily metrics, and SensorFM predictions (e.g., predicted hyperlipidemia state, predicted hypertension state, etc.) • Condition (B): Demographics + Daily Wearable Metrics + Ground Truth. Similar to Condition (A) but with SensorFM predictions replaced by the patient’s actual ground-truth targets (e.g., actual hyperlipidemia state, actual hypertension state, etc.). • Condition (C): Demographics + Daily Wearable Metrics Baseline comparator: agent receives only demographics and feature-engineered daily metrics. An overview of the SensorFM-augmented agent setup, alongside the evaluation results, is presented in Figure 3. A set of experienced clinicians evaluated health summary responses generated by Gemini 3 Flash from experimental conditions A, B, and C, while blinded to the conditions. Over 40 hours of expert clinical annotations yielded a total of 𝑛 = 1, 860 ratings across five rubric items (See Survey ED.1). Extra Context Improves Health Agent Performance. Pairwise comparisons using Wilcoxon signed-rank tests with a Bonferroni correction reveal that the extra context from either SensorFM or ground-truth labels leads to significant improvements over the baseline condition (C) (SensorFM A vs. C: 𝑊 = 10110, 𝑝 < 0.001; Ground truth B vs. C: 𝑊 = 9596, 𝑝 < 0.001). The pattern is consistent across all five rubric dimensions. With SensorFM predictions (A), the agent generates significantly stronger responses compared to the baseline (C) across every individual axis: Context (𝑊 = 451, 𝑝 < 0.001), Personalization (𝑊 = 378, 𝑝 < 0.001), Justifiability (𝑊 = 300, 𝑝 < 0.01), Relevance (𝑊 = 412, 𝑝 < 0.001), and Harm (𝑊 = 510, 𝑝 < 0.001). Providing the agent with ground-truth labels (Condition B) yields similar benefits across all dimensions (all 𝑝s < 0.0167 after correction). Extra Context via SensorFM Predictions Matches Extra Context via Ground Truth. We find no statistically significant differences in performance between SensorFM prediction (Condition A) and 12
Towards a General Intelligence and Interface for Wearable Health Data
Classification Tasks
Regression Tasks
Currently Working
Age
Disability
BMI
Disability Affects Work
Height
Smoking
Weight
Medicaid
ASCVD Risk
No Medications
Framingham Risk
Cardiovascular Dx
Framingham 30 Risk
Hypertension
HOMA-IR
Respiratory Dx
HbA1c
Diabetes Dx
Triglycerides
Diabetes Med
PHQ-8
Hyperlipidemia
GAD-7
Pre-Diabetes PSS
Insulin Resistance
Sleep Disturbance PRO Sleep Impairment PRO 0.0
Mild Depression Mild Anxiety
0.2
Persistent Stress
Model (Parameters)
Depression/Anxiety
Demo only XXS (105) XS (106)
Mental Health Med. Sleep Disorder Treatment 0.0
0.4 0.6 Pearson Correlation
Demographics Lifestyle
0.2
0.4 0.6 F1 Score
0.8
0.8
S (107) B (108) B (108) & Demo
Classroom Classroom & Demo
Cardiovascular Metabolics
Mental Health Sleep
Task Category
1.0
1.0
Figure 5 | Discriminative Task Agentic Classroom Solution Performance. Downstream performance across 35 discriminative tasks for SensorFM-B embeddings adapted with agentic classroom found solutions, linear probes of SensorFM variants, and a demographic baseline. In general the classroom found solutions improve upon simple linear probes. Linear probes use PCA-50 reduced embeddings, the classroom uses unreduced embeddings. For each task, we report the average 5-fold cross validation performance. Average 𝐹 1 is calculated with an arithmetic mean. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Error bars are standard deviations calculated in the transformed space and back-transformed to give asymmetric error values. 13
Towards a General Intelligence and Interface for Wearable Health Data
ground truth (Condition B) ( 𝑝 = 0.396). This demonstrates that the diagnostic inferences generated by SensorFM are comparable to the diagnostic ground truth when supplied as contextual inputs to a personal health agent.
3. Discussion 3.1. Learning a Scaled Representation of Sensed Human Physiology Our scaling experiments reveal a predictable relationship between model performance, model capacity, and pretraining data volume. The most consistent performance gains were observed when model size and data volume were increased concurrently, with Figure 1.e indicating that our pretraining method has yet to saturate. This reinforces the hypothesis that both dimensions are essential to maximize gains through self-supervised pretraining and to learn robust representations of sensed physiology. Overall, our findings demonstrate how advancements in wearable AI may be driven by leveraging expansive, diverse, high-fidelity datasets. Unlike computer vision or natural language processing, which benefit from web-scale open corpora, wearable data is inherently difficult to aggregate, standardize, and share due to privacy and technical considerations. Consequently, data fidelity, additions of new modalities and longitudinal scale will likely be a primary scaling factor for the continued evolution of sensor foundation models. 3.2. The Utility of Generalizable Embeddings for Wearable Health We demonstrate that large-scale self-supervised pretraining on a large corpora of wearable sensor data produces a robust representation of sensed human physiology and behavior that transfers effectively to diverse health domains. Across 35 distinct health and behavioral tasks, encompassing cardiovascular health, metabolic function, sleep architecture, mental health, lifestyle and demographic factors, linear probes of the SensorFM embeddings consistently outperformed supervised baselines trained with engineered features. Crucially, SensorFM achieves robust predictive accuracy even without task-specific architectures, suggesting that when trained at sufficient scale, wearable foundation models can learn representations that are broadly useful and label-efficient. This is particularly critical across healthcare domains, where high-quality, ground-truth labels are often expensive and labor-intensive to obtain. To this end, our results suggest that scaled pretraining may be particularly valuable for applications involving heterogeneous and weakly expressed phenotypes, such as mental health (e.g. Depression/Anxiety, PHQ-8 score, etc.). Mental health conditions involve a blend of subjective and objective indicators (Newson et al., 2020), present with diverse clinical manifestations (Hwang et al., 2008; Kivimäki et al., 2020; McLean et al., 2011), and are governed by complex temporal dynamics (Nelson et al., 2017). By pretraining on a large corpus with broad variation in daily routines, SensorFM may better marginalize “nuisance” variation and retain latent physiological information that generalizes across diverse populations. However, we caution that population-level models are distinct from individual-level forecasting; future research should employ longitudinal personalized modeling to better evaluate within-person changes in conditions and health over time. 3.3. Accelerating Model Development Flexible and expressive pretrained embeddings can dramatically increase how efficiently we can develop new downstream models. The few-shot performance of SensorFM speaks to the generalizability and efficient adaptability of the representation learned through scaled pretraining. However, the limited volume of data with high-quality labels continues to pose a bottleneck on downstream task 14
Towards a General Intelligence and Interface for Wearable Health Data
performance. With such sparse annotations of downstream health outcomes, significant engineering may be required to optimally adapt the general SensorFM embeddings to each endpoint. Towards this, the capabilities of AI language models to reason over problems and write code presents a promising approach to expedite model development, and partially bridge the limitations of data sparsity. In this work we establish that an agent-driven framework for iterative solution discovery is able to efficiently adapt general embeddings to diverse health domains showing sweeping benefits across discriminative health tasks as compared to a linear probe. Agent driven research and scientific discovery is a rapidly evolving field (Aygün et al., 2025; Gottweis et al., 2026) and future work should continue to explore the extent to which these methods may be exploited to provide benefit to the sparsly labeled data found in healthcare. 3.4. Implications for Digital Health and Clinical Monitoring The efficacy of SensorFM carries significant implications for the future of digital health. Our results support a transition from task-specific wearable applications toward a general-purpose interface for continuous health monitoring. SensorFM achieves robust predictive accuracy without requiring complex, task-specific architectures; applying simple linear probes to the pretrained embeddings is sufficient for a wide variety of downstream applications. By providing a common substrate for various health models, pretrained representations eliminate the need for bespoke, complex, and end-to-end machine learning pipelines for every individual outcome, streamlining the deployment of predictive analytics in digital health. Furthermore, while traditional healthcare relies on episodic, "snapshot" measurements captured during clinical visits or in laboratory settings, wearable sensors provide dense, longitudinal observations of physiology and behavior in free-living conditions. In this context, a generalist model for wearable health may be useful in identifying individuals who would benefit from confirmatory testing or early intervention (Lubitz et al., 2022; Perez et al., 2019). This is especially valuable for conditions that remain asymptomatic until advanced stages (Liu et al., 2022; Yang et al., 2022). We emphasize, however, that these predictions are intended for screening, risk stratification, and longitudinal tracking rather than as a definitive replacements for clinical diagnosis. 3.5. Towards More Grounded Question Answering for Personal Health Recent works have emphasized the scale on which people leverage AI systems for medical question answering (Costa-Gomes et al., 2026), and the efficacy with which these systems are able to provide feedback to queries (Breda et al., 2026; McDuff et al., 2025b). Wearable health data offers the opportunity to improve the quality of responses to user queries by grounding answers in measures of their own sensed physiology and behavior. Our experiments on the use of SensorFM as tool by a personal health agent demonstrate that incorporating SensorFM predictions into the agent’s context results in statistically significant improvements over baseline systems that lack this extra context. Notably, despite the inherent imperfections of model-derived predictions compared to the clinical ground truth, we observed no statistically significant difference in the quality of the resulting agentic interactions. Looking forward, systems like ours, which pair AI coaches with additional grounding, may enable the delivery of more personalized, proactive, and accessible guidance, bridging the gap between continuous health sensing and sporadic clinical consultations. 3.6. Limitations and Future Work This study has several limitations that we acknowledge. First, consumer wearable devices are heterogeneous in both hardware and signal processing, and there is limited standardization in how 15
Towards a General Intelligence and Interface for Wearable Health Data
measurements are derived across platforms. Although SensorFM was trained and evaluated on data from multiple Fitbit and Pixel Watch devices, transfer to other device ecosystems is not guaranteed and would likely require additional adaptation. Second, to support large-scale modeling on consumer devices, our input representation uses one-minute aggregated features rather than raw sensor waveforms. This allows the model to capture long-range daily and circadian context, but it necessarily discards finer temporal structure, such as beat-to-beat cardiac variability and sub-second motion patterns, that may be informative for some clinical tasks. As such, more work is needed to understand the optimal settings (sampling frequency, temporal context) for different health outcomes. We additionally note that more granular signals are also not necessarily stored due to considerations such as storage footprints, and power constraints. Third, collecting large and reliable downstream labels paired with wearable data remains difficult. While some of the outcomes studied here are lab tests, a sizeable portion rely on self-reported diagnoses, medication use or screening questionnaires. Converting continuous screening measures into binary outcomes may introduce threshold-dependent noise. Future work should therefore evaluate these models against clinically verified outcomes, including electronic health record (EHR) labels and gold-standard physiological measurements. Fourth, our evaluation of the SensorFM-augmented health agent was constrained to a static, single-turn interaction paradigm. Real-world clinical consultations and AI agent interactions are inherently multi-turn, allowing users or physicians to ask follow-up questions to clarify ambiguities or gather missing context. In our experimental design, physician evaluators were restricted solely to the static data presented to them. Furthermore, the downstream datasets utilized inherently lack comprehensive ground-truth labels across all health dimensions (e.g., missing mental health or sleep targets). While this missingness highlights a primary strength of our method, predicting unknown labels, it also means evaluators had limited holistic patient context. Although restricting physicians from asking follow-up questions was an intentional methodological choice to tightly control variables and prevent evaluation bias, it does limit the direct generalizability of these findings to dynamic, conversational real-world deployments. Finally, both the pretraining and downstream evaluation populations reflect the characteristics of Fitbit and Pixel Watch users and therefore may not be fully representative of the broader United States or global population. Although the pretraining cohort spans a broad range of age, sex and BMI, the population of wearable users do not necessarily represent the United States population distribution, such as having a greater proportion of female users than the population ratio. Reported performance may therefore not generalize directly to the overall population and future evaluation will require more targeted recruitment and subgroup-specific validation.
4. Conclusion The ubiquity of wearable health monitors offers the opportunity to improve access to personal health insights and in so doing drive preventative care. To this end, we present SensorFM, a foundation model for wearable health which generalizes across diverse health domains. While data labeled with health outcome annotations are rare, our method learns directly from over one trillion minutes of unlabeled multimodal wearable sensor data, producing a robust representation of sensed physiology with predictive benefits on tasks spanning cardiovascular, metabolic, and mental health, sleep, lifestyle factors, and physiologically-relevant demographics. Through our analysis we establish the utility of this scaled pretraining, and further investigate the label efficiency, generative capabilities, and latent structure of SensorFM. To enable more efficient adaptation of the SensorFM embeddings to diverse downstream tasks, we present an agentic code generation framework that iteratively explores each 16
Towards a General Intelligence and Interface for Wearable Health Data
solution space. Finally, to understand the end-to-end utility of SensorFM we assess the benefit of providing SensorFM predictions to a personal health agent. In summary, our work represents a shift towards more generalist models for wearable health, which enable more flexible prediction across health domains, more grounded downstream reasoning, and opportunities to provide users with tangible insights about their own health.
References S. Abbaspourazad, O. Elachqar, A. Miller, S. Emrani, U. Nallasamy, and I. Shapiro. Large-scale training of foundation models for wearable biosignals. In The Twelfth International Conference on Learning Representations, 2024. A. F. Ansari, L. Stella, C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. S. Rangapuram, S. P. Arango, S. Kapoor, et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815, 2024. E. Aygün, A. Belyaeva, G. Comanici, M. Coram, H. Cui, J. Garrison, R. J. A. Kast, C. Y. McLean, P. Norgaard, Z. Shamsi, et al. An ai system to help scientists write expert-level empirical software. arXiv preprint arXiv:2509.06503, 2025. J. Breda, F. Yousif, B. Hawkins, M. Cotoi, M. Liu, R. Luo, P.-H. C. Chen, M. Schaekermann, S. Schmidgall, X. Liu, et al. Symptomai: Towards a conversational ai agent for everyday symptom assessment. arXiv preprint arXiv:2605.04012, 2026. D. Cella, W. Riley, A. Stone, N. Rothrock, B. Reeve, S. Yount, D. Amtmann, R. Bode, D. Buysse, S. Choi, et al. The patient-reported outcomes measurement information system (promis) developed and tested its first wave of adult self-reported health outcome item banks: 2005–2008. Journal of clinical epidemiology, 63(11):1179–1194, 2010. L. Choi, Z. Liu, C. Matthews, and M. Buchowski. Validation of accelerometer wear and nonwear time classification algorithm. Medicine and Science in Sports and Exercise, 43(2):357–364, 2011. S. Cohen, T. Kamarck, and R. Mermelstein. A global measure of perceived stress. Journal of health and social behavior, 24:385–396, 1983. G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. G. Cornelissen. Cosinor-based rhythmometry. Theoretical Biology and Medical Modelling, 11:16, 2014. B. Costa-Gomes, P. Tolmachev, E. Taysom, V. Sounderajah, H. Richardson, P. Schoenegger, X. Liu, M. M. Nour, S. Spielman, S. F. Way, et al. Public use of a generalist llm chatbot for health queries. Nature Health, pages 1–8, 2026. A. Das, W. Kong, R. Sen, and Y. Zhou. A decoder-only foundation model for time-series forecasting. arXiv preprint arXiv:2310.10688, 2023. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
17
Towards a General Intelligence and Interface for Wearable Health Data
E. Erturk, F. Kamran, S. Abbaspourazad, S. Jewell, H. Sharma, Y. Li, S. Williamson, N. J. Foti, and J. Futoma. Beyond sensor data: Foundation models of behavioral data from wearables improve health predictions. In Forty-second International Conference on Machine Learning, 2025. A. Garza, C. Challu, and M. Mergenthaler-Canseco. Timegpt-1. arXiv preprint arXiv:2310.03589, 2023. M. Goswami, K. Szafer, A. Choudhry, Y. Cai, S. Li, and A. Dubrawski. Moment: A family of open time-series foundation models. arXiv preprint arXiv:2402.03885, 2024. J. Gottweis, W.-H. Weng, A. Daryin, T. Tu, P. Sirkovic, A. Myaskovsky, G. Glowaty, F. Weissenberger, A. Orlandi, D. Popovici, A. Palepu, K. Rong, R. Tanno, K. Saab, F. Zhang, J. Blum, A. Carroll, K. Kulkarni, N. Tomašev, D. Zverinski, I. Rendulic, E. Vedadi, F. Hasler, L. Rimanic, M. Boia, I. Budiselic, B. Feinstein, M. Bellaiche, T. Sheffer, J. Freyberg, J. Ratcliff, O. Bertolli, K. Chou, A. Hassidim, B. Gokturk, A. Vahdat, Y. Guan, V. Dhillon, E. D. Vaishnav, B. Lee, T. R. D. Costa, J. R. Penadés, G. Peltz, Y. Matias, J. Manyika, D. Hassabis, Y. Xu, P. Kohli, A. Pawlosky, A. Karthikesalingam, and V. Natarajan. Accelerating scientific discovery with co-scientist. Nature, May 2026. ISSN 1476-4687. doi: 10.1038/s41586-026-10644-y. URL https://doi.org/10.1038/s41586-026-10644-y. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022. A. A. Heydari, K. Gu, V. Srinivas, H. Yu, Z. Zhang, Y. Zhang, A. Paruchuri, Q. He, H. Palangi, N. Hammerquist, et al. The anatomy of a personal health agent. arXiv preprint arXiv:2508.20148, 2025. B. Hjorth. Eeg analysis based on time domain properties. Electroencephalography and Clinical Neurophysiology, 29(3):306–310, 1970. J. A. Horne and O. Ostberg. A self-assessment questionnaire to determine morningness-eveningness in human circadian rhythms. International journal of chronobiology, 4(2):97–110, 1976. W.-C. Hwang, H. F. Myers, J. Abe-Kim, and J. Y. Ting. A conceptual paradigm for understanding culture’s impact on mental health: The cultural influences on mental health (cimh) model. Clinical psychology review, 28(2):211–227, 2008. J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. M. Kivimäki, G. D. Batty, J. Pentti, M. J. Shipley, P. N. Sipilä, S. T. Nyberg, S. B. Suominen, T. Oksanen, S. Stenholm, M. Virtanen, et al. Association between socioeconomic status and the development of mental and physical health conditions in adulthood: a multi-cohort study. The Lancet Public Health, 5(3):e140–e149, 2020. D. Kiyasseh, T. Zhu, and D. A. Clifton. Clocs: Contrastive learning of cardiac signals across space, time, and patients. In International conference on machine learning, pages 5606–5615. PMLR, 2021. K. Kroenke, T. W. Strine, R. L. Spitzer, J. B. Williams, J. T. Berry, and A. H. Mokdad. The phq-8 as a measure of current depression in the general population. Journal of affective disorders, 114(1-3): 163–173, 2009. 18
Towards a General Intelligence and Interface for Wearable Health Data
M. Kwon, D.-J. Kim, H. Cho, and S. Yang. The smartphone addiction scale: development and validation of a short version for adolescents. PloS one, 8(12):e83558, 2013. X. Liu, D. McDuff, G. Kovacs, I. Galatzer-Levy, J. Sunshine, J. Zhan, M.-Z. Poh, S. Liao, P. Di Achille, and S. Patel. Large language models are few-shot health learners. arXiv preprint arXiv:2305.15525, 2023. Y. Liu, G. Zhang, C. G. Tarolli, R. Hristov, S. Jensen-Roberts, E. M. Waddell, T. L. Myers, M. E. Pawlik, J. M. Soto, R. M. Wilson, et al. Monitoring gait at home with radio waves in parkinson’s disease: A marker of severity, progression, and medication response. Science Translational Medicine, 14(663): eadc9669, 2022. D. M. Lloyd-Jones, P. W. Wilson, M. G. Larson, A. Beiser, E. P. Leip, R. B. D’Agostino, and D. Levy. Framingham risk score and prediction of lifetime risk for coronary heart disease. The American journal of cardiology, 94(1):20–24, 2004. B. Löwe, O. Decker, S. Müller, E. Brähler, D. Schellberg, W. Herzog, and P. Y. Herzberg. Validation and standardization of the generalized anxiety disorder screener (gad-7) in the general population. Medical care, 46(3):266–274, 2008. S. A. Lubitz, A. Z. Faranesh, C. Selvaggi, S. J. Atlas, D. D. McManus, D. E. Singer, S. Pagoto, M. V. McConnell, A. Pantelopoulos, and A. S. Foulkes. Detection of atrial fibrillation in a large population using wearable devices: the fitbit heart study. Circulation, 146(19):1415–1424, 2022. S. M. Lundberg and S.-I. Lee. A unified approach to interpreting model predictions. Advances in neural information processing systems, 30, 2017. D. McDuff, A. Barakat, A. Winbush, A. Jiang, F. Cordeiro, R. Crowley, L. E. Kahn, J. Hernandez, N. B. Allen, et al. The google health digital well-being study: Protocol for a digital device use and well-being study. JMIR Research Protocols, 13(1):e49189, 2024. D. McDuff, I. Galatzer-Levy, S. Thomson, A. Barakat, C. Heneghan, S. Abdel-Ghaffar, J. Sunshine, M.-Z. Poh, L. Sunden, J. B. Hernandez, et al. Evidence of differences in diurnal electrodermal, temperature and heart rate patterns by mental health status in free-living data. BMJ Mental Health, 28(1), 2025a. D. McDuff, M. Schaekermann, T. Tu, A. Palepu, A. Wang, J. Garrison, K. Singhal, Y. Sharma, S. Azizi, K. Kulkarni, et al. Towards accurate differential diagnosis with large language models. Nature, 642 (8067):451–457, 2025b. L. McInnes, J. Healy, and J. Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. C. P. McLean, A. Asnaani, B. T. Litz, and S. G. Hofmann. Gender differences in anxiety disorders: prevalence, course of illness, comorbidity and burden of illness. Journal of psychiatric research, 45 (8):1027–1035, 2011. M. A. Merrill, M. Tan, V. Gupta, T. Hartvigsen, and T. Althoff. Language models still struggle to zero-shot reason about time series. arXiv preprint arXiv:2404.11757, 2024. A. A. Metwally, A. A. Heydari, D. McDuff, A. Solot, Z. Esmaeilpour, A. Z. Faranesh, M. Zhou, G. Narayanswamy, M. A. Xu, X. Liu, et al. Insulin resistance prediction from wearables and routine blood biomarkers. Nature, pages 1–11, 2026.
19
Towards a General Intelligence and Interface for Wearable Health Data
B. Munos, P. C. Baker, B. M. Bot, M. Crouthamel, G. de Vries, I. Ferguson, J. D. Hixson, L. A. Malek, J. J. Mastrototaro, V. Misra, et al. Mobile health: the power of wearables, sensors, and apps to transform clinical trials. Annals of the New York Academy of Sciences, 1375(1):3–18, 2016. G. Narayanswamy, X. Liu, K. Ayush, Y. Yang, X. Xu, S. Liao, J. Garrison, S. A. Tailor, J. Sunshine, Y. Liu, T. Althoff, et al. Scaling wearable foundation models. In The Thirteenth International Conference on Learning Representations, 2025. B. Nelson, P. D. McGorry, M. Wichers, J. T. Wigman, and J. A. Hartmann. Moving from static to dynamic models of the onset of mental disorder: a review. JAMA psychiatry, 74(5):528–534, 2017. J. J. Newson, D. Hunter, and T. C. Thiagarajan. The heterogeneity of mental health assessment. Frontiers in psychiatry, 11:76, 2020. M. Nissen, S. Slim, K. Jäger, M. Flaucher, H. Huebner, N. Danzberger, P. A. Fasching, M. W. Beckmann, S. Gradl, B. M. Eskofier, et al. Heart rate measurement accuracy of fitbit charge 4 and samsung galaxy watch active2: device evaluation study. JMIR formative research, 6(3):e33635, 2022. A. Novikov, N. Vũ, M. Eisenberger, E. Dupont, P.-S. Huang, A. Z. Wagner, S. Shirobokov, B. Kozlovskii, F. J. Ruiz, A. Mehrabian, et al. Alphaevolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131, 2025. M. M. Ohayon, M. Paskow, A. Roach, C. Filer, D. S. Hillygus, M. C. Chen, G. Langer, M. Hirshkowitz, N. S. F. S. S. Consensus, et al. The national sleep foundation’s sleep satisfaction tool. Sleep Health, 5(1):5–11, 2019. M. V. Perez, K. W. Mahaffey, H. Hedlin, J. S. Rumsfeld, A. Garcia, T. Ferris, V. Balasubramanian, A. M. Russo, A. Rajmane, L. Cheung, et al. Large-scale assessment of a smartwatch to identify atrial fibrillation. New England Journal of Medicine, 381(20):1909–1917, 2019. K. Rasul, A. Ashok, A. R. Williams, A. Khorasani, G. Adamopoulos, R. Bhagwatkar, M. Biloš, H. Ghonia, N. V. Hassen, A. Schneider, et al. Lag-llama: Towards foundation models for time series forecasting. arXiv preprint arXiv:2310.08278, 2023. M. Ringeval, G. Wagner, J. Denford, G. Paré, and S. Kitsiou. Fitbit-based interventions for healthy lifestyle outcomes: systematic review and meta-analysis. Journal of medical Internet research, 22 (10):e23954, 2020. A. Sellergren, S. Kazemzadeh, T. Jaroensri, A. Kiraly, M. Traverse, T. Kohlberger, S. Xu, F. Jamil, C. Hughes, C. Lau, et al. Medgemma technical report. arXiv preprint arXiv:2507.05201, 2025. F. Shaffer and J. P. Ginsberg. An overview of heart rate variability metrics and norms. Frontiers in public health, 5:258, 2017. K. Shah, A. Wang, Y. Chen, J. Munjal, S. Chhabra, A. Stange, E. Wei, T. Phan, T. Giest, B. Hawkins, et al. Automated loss of pulse detection on a consumer smartwatch. Nature, 642(8066):174–181, 2025. J. Shi, Q. Ma, H. Ma, and L. Li. Scaling law for time series forecasting. Advances in Neural Information Processing Systems, 37:83314–83344, 2024. D. Spathis, I. Perez-Pozuelo, S. Brage, N. J. Wareham, and C. Mascolo. Self-supervised transfer learning of physiological representations from free-living wearable data. In Proceedings of the Conference on Health, Inference, and Learning, pages 69–78, 2021. 20
Towards a General Intelligence and Interface for Wearable Health Data
R. L. Spitzer, K. Kroenke, J. B. Williams, and B. Löwe. A brief measure for assessing generalized anxiety disorder: the gad-7. Archives of internal medicine, 166(10):1092–1097, 2006. J. Sumner, Y. Wang, S. Y. Tan, E. H. H. Chew, and A. Wenjun Yip. Perspectives and experiences with large language models in health care: Survey study. Journal of Medical Internet Research, 27:e67383, 2025. R. Thapa, B. He, M. R. Kjaer, H. Moore, G. Ganjoo, E. Mignot, and J. Zou. Sleepfm: Multi-modal representation learning for sleep across brain activity, ecg and respiratory signals. arXiv preprint arXiv:2405.17766, 2024. M. Thukral, S. G. Dhekane, S. K. Hiremath, H. Haresamudram, and T. Ploetz. Layout-agnostic human activity recognition in smart homes through textual descriptions of sensor triggers (tdost). Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 9(1):1–38, 2025. J. Torous, J. Rodriguez, and A. Powell. The new digital divide for digital biomarkers. Digital Biomarkers, 1(1):87–91, 2017. T. Tu, M. Schaekermann, A. Palepu, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, Y. Cheng, et al. Towards conversational diagnostic artificial intelligence. Nature, 642(8067):442–450, 2025. E. J. Van Someren et al. Bright light therapy: improved sensitivity to its effects on rest-activity rhythms in alzheimer patients by application of nonparametric methods. Chronobiology International, 16(4): 505–518, 1999. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. A. Winbush, D. McDuff, J. Hernandez, A. Barakat, A. Jiang, C. Heneghan, B. W. Nelson, and N. B. Allen. Smartphone use in a large us adult population: Temporal associations between objective measures of usage and mental well-being. Proceedings of the National Academy of Sciences of the United States of America, 122(43):e2427311122, 2025. C. Wu, X. Zhang, Y. Zhang, H. Hui, Y. Wang, and W. Xie. Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data. Nature Communications, 16(1):7866, 2025. H. Xu, N. Usuyama, J. Bagga, S. Zhang, R. Rao, T. Naumann, C. Wong, Z. Gero, J. González, Y. Gu, et al. A whole-slide foundation model for digital pathology from real-world data. Nature, 630 (8015):181–188, 2024. M. A. Xu, G. Narayanswamy, K. Ayush, D. Spathis, S. Liao, S. A. Tailor, A. Metwally, A. A. Heydari, Y. Zhang, J. Garrison, et al. Lsm-2: Learning from incomplete wearable sensor data. arXiv preprint arXiv:2506.05321, 2025. Y. Yang, Y. Yuan, G. Zhang, H. Wang, Y.-C. Chen, Y. Liu, C. G. Tarolli, D. Crepeau, J. Bukartyk, M. R. Junna, et al. Artificial intelligence-enabled detection and assessment of parkinson’s disease using nocturnal breathing signals. Nature Medicine, 28(10):2207–2215, 2022. L. Yu, D. J. Buysse, A. Germain, D. E. Moul, A. Stover, N. E. Dodds, K. L. Johnston, and P. A. Pilkonis. Development of short forms from the promis™ sleep disturbance and sleep-related impairment item banks. Behavioral sleep medicine, 10(1):6–24, 2012.
21
Towards a General Intelligence and Interface for Wearable Health Data
H. Yuan, S. Chan, A. P. Creagh, C. Tong, A. Acquah, D. A. Clifton, and A. Doherty. Self-supervised learning for human activity recognition using 700,000 person-days of wearable data. NPJ digital medicine, 7(1):91, 2024. X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12104–12113, 2022. Y. Zhang, K. Ayush, S. Qiao, A. A. Heydari, G. Narayanswamy, M. A. Xu, A. A. Metwally, S. Xu, J. Garrison, X. Xu, T. Althoff, Y. Liu, P. Kohli, J. Zhan, M. Malhotra, S. Patel, C. Mascolo, X. Liu, D. McDuff, and Y. Yang. Sensorlm: Learning the language of wearable sensors. In Conference on Neural Information Processing Systems (NeurIPS), 2025.
22
Towards a General Intelligence and Interface for Wearable Health Data
Methods M.1. Sensor Signals for Wearable Foundation Models The Fitbit Sense 2 and Pixel Watch 2 utilize five sensors of primary relevance to this study: Photoplethysmography (PPG), skin temperature, accelerometer, electrodermal activity (EDA), and an altimeter. From these raw inputs, we derive a set of 34 aggregate signals (features), detailed in Table ED.1. To optimize device battery life and storage, raw sensor data is not retained; instead, we rely on one-minute aggregate signals. We tested whether these signals were strongly co-linear before proceeding with model training. Figure ED.14 of Appendix G shows no pair of signals has a correlation greater than 0.8. CRD Cardiovascular. Heart rate (HR) is extracted from the PPG signal at 1 Hz using a validated algorithm (Nissen et al., 2022). Per-minute HR is computed as the mean of the instantaneous heart rate across non-overlapping one-minute windows. An on-device peak detection algorithm identifies R-wave peaks to calculate RR intervals. To mitigate noise, standard HRV metrics are computed using robust statistical methods. We derive the median RR interval, the Shannon entropy of the RR intervals (ShEnRR), and the Coherence of breathing frequency to heart rate. Time-domain variability metrics—standard deviation of RR intervals (SDNN) and root mean squared of successive differences (RMSSD)—are calculated using RR intervals between the 5th and 95th percentiles to exclude outliers. We also compute pNN20, the percentage of successive RR intervals differing by more than 20ms.In the frequency domain, we extract the power in the Very Low (VLF), Low (LF), and High (HF) frequency bands, the LF/HF ratio, and the Shannon entropy of the power spectrum (SpectralEn). Finally, the "Valid RR" metric quantifies the percentage of the 5-minute window containing valid RR intervals. PLM Cardiopulmonary. We derive two features related to blood oxygen saturation. SpO2 represents the blood oxygen saturation level, while SpO2 Confidence provides the confidence level of the reading. SpO2 is computed exclusively during stationary periods; raw sensor data is filtered to isolate segments where accelerometer variance is low. The PPG waveform is processed by a convolutional neural network to extract features, which are subsequently classified via a fully connected layer. We also track SpO2 Coverage, defined as the percentage of the minute with valid SpO2 data. SLP Sleep. Sleep metrics are inferred using a validated multi-modal algorithm that fuses accelerometer-derived actigraphy with PPG-derived heart rate and HRV data. The model classifies sleep epochs into four primary stages: Awake, Light, Deep, and Rapid Eye Movement (REM). For each one-minute window, we record the time spent in each of these stages in seconds. MTN Motion. We extract ten features from the 3-axis accelerometer to characterize motion and physical activity. These include the step count, Axis Mean (mean of the 3-axis data), and Kurtosis (of the 3-axis root mean squared magnitude).Complex signal features include Jerk Autocorrelation (ratio of lag-1 autocorrelation to energy), Log Energy, and Log Energy Ratio. We calculate the Covariance as an estimate of the condition number for the 3-axis covariance. We also examine the zero-crossings of the 1st 3-axis principal component, extracting both the Zero Crossing Average (mean time between crossings) and Zero Crossing St.Dev. (standard deviation of time between crossings).Additionally, we derive a "Sleep Coefficient" metadata feature, calculated as the sum of the 3-axis max-min range within 16 log-scaled bins.From the barometer, we compute the Altitude St.Dev., which represents the standard deviation of altimeter readings (in hP) to isolate vertical displacement from atmospheric drift. SKN Skin Surface. The device measures skin conductance and temperature to infer physiological states. The Electro-Dermal Activity (EDA) sensor measures Skin Conductance Level (SCL), which 23
Towards a General Intelligence and Interface for Wearable Health Data
correlates with sympathetic nervous system arousal. We derive the "Conductance" feature, defined as the center of the linear tonic SCL value fit (in 𝜇 Siemens). Lead Contact Counts are recorded to track the number of times sensor leads contact the wrist.Concurrently, a temperature sensor on the wrist-facing surface samples skin temperature. We report the "Temperature" feature as the mean value of skin temperature (in °C) for the minute. Table ED.1 | Sensor Feature Definitions and the Sensor they are Derived From. Feature CRD Cardiovascular Heart Rate RR Median RMSSD 05-95 SDNN 05-95 pNN20 Coherence ShEnRR VLF LF HF LF/HF SpectralEn Valid RR PLM Cardiopulmonary SpO2 SpO2 Confidence SpO2 Coverage SLP Sleep Stage Awake Stage Light Stage Deep Stage REM Sleep Coefficient MTN Motion Steps Jerk Autocorrelation Log Energy Covariance Log Energy Ratio Zero Crossing St.Dev. Zero Crossing Average Axis Mean Kurtosis Altitude St.Dev. SKN Skin Surface Temperature Conductance Lead Contact Counts
Sensor
Unit
Definition
PPG PPG PPG PPG PPG PPG PPG PPG PPG PPG PPG PPG PPG
Beats/Min Msec Msec Msec % a.u. Nats Msec2 Msec2 Msec2 a.u. Nats %
Mean of instantaneous heart rate. Median RR interval. RMSSD calculated using RR intervals between the 5𝑡ℎ and 95𝑡ℎ percentile. SDNN calculated using RR intervals between the 5𝑡ℎ and 95𝑡ℎ percentile. Percentage of successive RR interval differences greater than 20 ms. Coherence of the breathing frequency band to the heart rate. Shannon entropy of the RR intervals. Power in the Very Low Frequency band (0.003-0.04 Hz) of the RR spectrum. Power in the Low Frequency band (0.04-0.15 Hz) of the RR spectrum. Power in the High Frequency band (0.15-0.4 Hz) of the RR spectrum. Ratio of LF to HF power. Shannon entropy of the RR interval power spectrum. % of 5-minute window with valid RR intervals.
PPG PPG PPG
% a.u. %
Blood oxygen saturation level. Confidence level of the SpO2 reading. Percentage of the minute with valid SpO2 data.
PPG + ACCEL PPG + ACCEL PPG + ACCEL PPG + ACCEL ACCEL
Seconds Seconds Seconds Seconds a.u.
Time spent in the Awake sleep stage. Time spent in the Light sleep stage. Time spent in the Deep sleep stage. Time spent in the REM sleep stage. Sum of 3-axis max-min range with 16 log-scaled bins.
ACCEL ACCEL ACCEL ACCEL ACCEL
Steps a.u. a.u. a.u. a.u.
ACCEL
Seconds
ACCEL ACCEL ACCEL Barometer
Seconds a.u. a.u. hP
Number of steps. Ratio of lag=1 autocorrelation to energy in 1st 3-axis principal component. Log of sum of 3-axis root mean squared magnitude. Estimate of condition number for the 3-axis covariance. Log of ratio of sum of energy in 1st 3-axis principal component over energy of 3-axis root mean squared magnitude. Standard deviation of time between zero crossing of 1st 3-axis principal component. Mean of time between zero crossing of 1st 3-axis principal component. Mean of 3-axis Kurtosis of 3-axis root mean squared magnitude. Standard deviation of altimeter readings.
TEMP EDA EDA
°C 𝜇 Siemens Counts
Mean value of skin temperature. Center of linear tonic SCL value fit. Number of times sensor leads contacted the wrist in a minute.
M.2. Datasets M.2.1. Pretraining Dataset To build the large dataset for our experiments we sampled wearable data from 5 million participants during the period September 1𝑠𝑡 2024 to September 1𝑠𝑡 2025. As participants could opt-in using any Fitbit or Pixel Watch device, the dataset contains data from a wide range of models released between 2012 to 2024 (see Table ED.13). The most common devices were Fitbit Inspire 3, Fitbit Charge 6, Fitbit Versa 2, 3 and 4, Fitbit Sense and Pixel Watch 1, 2, 3. Participants provided voluntary consent for their de-identified data to be used for research and development of new health and wellness products and services and we obtained a secondary research exemption determination from a centralized IRB 24
Towards a General Intelligence and Interface for Wearable Health Data
Figure ED.1 | SensorFM Input Data. The model ingests 34 one-minute aggregate sensor features derived from five sensor modalities (PPG, Accelerometer, EDA, Skin Temperature, and Altimeter) organized into seven categories (Accelerometer, Altitude, EDA, HR/HRV, SpO2 , Temperature, Sleep) over a 24-hour context window, processed through outlier removal and normalization before encoding. Note that these features are derived from a set of five sensors as described in Appendix G. This figure also highlights various modes of data missingness including out of range values, sensor power cycling, and human behaviors. for this research (WCG 20253840). We sub-selected from people wearing one of these devices as older device generations included fewer sensors. The subjects were asked for self-reported sex, age and weight. Table ED.2 summarizes the characteristics of the pretraining data and Fig. ED.3 shows the distribution by age and body-mass index (BMI). All data were de-identified and not linked with any other information. To create a dataset that maximized the number of subjects we randomly sampled 10-weeks of data from each of the 5 million subjects. From this set of upto 350 million human-days, we construct a pretraining dataset encompassing over 2 billion hours of multimodal sensor data. In total our pretrain set contains over one trillion minutes of minute-resolution data observed from a suite of five wearable sensors. In preparing our data for modeling, global normalization parameters (mean and standard deviation) for each of the sensor signals were computed on this pretraining dataset. These parameters were used to normalize (z-score) the pretraining and downstream data (Figure ED.2). In addition to this pretraining set, one week of data from an independent 10,000 subjects were used for validation, and one week of data from another independent set of 10,000 subjects were used for test. M.2.2. Downstream Datasets We compiled an extensive list of downstream tasks by combining de-identified data from multiple prospective IRB-approved observational studies as described below. Definitions of these labels are provided in Table ED.3.
25
Towards a General Intelligence and Interface for Wearable Health Data
Figure ED.2 | Pretraining Pipeline. Original multimodal sensor inputs are resampled and normalized, then artificially masked before being passed to the encoder. The decoder reconstructs the masked patches, with MSE loss computed on the artificially masked regions. M.2.2.1. Metabolic, Cardiac and Respiratory Health We designed a prospective observational study and recruited adult participants from the United States Metwally et al. (2026). The study was approved by Advarra (Pro00074093). We enrolled 4,416 participants, of which 1,086 had complete data (at least 14 days of data wearable data, lab results, and completed demographics) and were included in our analysis. The study, which spanned a maximum of 70 days, was designed to gauge the feasibility of leveraging wrist-worn wearable devices to develop algorithms for assessing metabolic health, cardiovascular health, and respiratory health deriving to biological age, and regressing to blood biomarkers. Specifically, participants completed surveys, obtained a one-time blood test, and were asked to continuously wear their wearable device for the duration of the study. Table ED.3 shows the numbers of people with confirmed responses for each of the data types (including lab reports and self-reported health history. Study Population. The study was limited to Fitbit users with a heart-rate sensing capable device, living in the US, and aged 21 - 80. These users were required to have at least three months prior data, with use on at least 75% of these days. Self-Report Demographics, Biometrics, and Medical History Questionnaires. Participants completed surveys reporting their demographic data (age, sex, ethnicity, weight, height), blood pressure, waist circumference, medications, diagnosed conditions, health habits and management strategies. All collected data are listed in the Methods section. Blood Panels. Participants obtained blood testing early in the morning from Quest Diagnostics after fasting for at least 8 hours in order to minimize the effect of solar diurnal cycle. Tested blood biomarkers included insulin, HbA1c, Comprehensive Metabolic Panel (including glucose), lipids (total cholesterol, triglycerides, HDL, LDL), Complete Blood Count, hs-CRP, GGT, and total testosterone. Test results were also used to calculate a Framingham Risk Score for 10-year prediction of cardiovascular 26
Towards a General Intelligence and Interface for Wearable Health Data
event risk (Lloyd-Jones et al., 2004). Wearable Data. Data was longitudinally collected from participants’ wearable devices for the duration of the study, and up to three months of wearable data prior to study enrollment was also linked. Compensation. Participants recieved the results of their blood draw free of charge. M.2.2.2. Sleep We designed a prospective observational study and recruited over 10,000 participants from all 50 states of the United States between August 2023 and January 2024. The study was approved by a centralized IRB (Advarra) (Pro00069849). The study was designed to support improvements to the Fitbit Sleep Score and to better understand the relationship between sleep, next day sleep, and health outcomes (e.g., alertness, mood, etc.). An additional aim was to leverage the data to develop models of sleep and health (e.g. circadian rhythm), enabling users to better understand their sleep and allowing for more personalized recommendations. The study duration was 15 days (14 nights) per participant during which time participants were asked to continuously wear their Fitbit except during charging. Additionally, up to one month of prior Fitbit data was included in the study data. Study Population. The study was limited to current Fitbit users whose devices included a heart rate sensor and who owned an Android smartphone capable of installing the Google Health Studies application. Participants were further required to be age 18 - 88, and located in the United States. Baseline and Post-Study Questionnaires. Participants completed a battery of self-report questionnaires at baseline including demographics, Sleep Habits, Health Habits, Sleep Environment, the National Sleep Foundation’s Sleep Satisfaction Tool (Ohayon et al., 2019), and the MorningnessEveningness Survey (Horne and Ostberg, 1976). Morning and Evening Survey. Each morning, participants were asked to complete a 5-item survey reporting on their sleep the previous night. Each evening, participants were asked to report on that day’s activities including food intake, exercise, and other activities. Alertness and Mood Surveys. Up to three times daily, participants were asked to complete Ecological Momentary Assessment (EMA) questions on their alertness and mood. Surveys were timed for mid-morning, afternoon, and evening, and were personalized for each participants’ approximate sleep schedule. Alertness Tasks. Up to three times daily, participants were asked to complete two standardized tasks meant to gauge alertness. These included a three-minute reaction time task where participants would tap their smartphone screen in response to some stimuli, and a five-minute gaze task where participants’ gaze was tracked as stimuli were presented on screen. Compensation. Subjects were compensated with a $25 Google Merchandise Purchase Code if they attempted at least 12 of 15 ( 80%) of the days/nights of the protocol. M.2.2.3. Mental Health We designed a prospective, observational study and recruited 7,500 Fitbit users in the United States (McDuff et al., 2024; Winbush et al., 2025). The study was approved by the IRB of the University of Oregon (MOD00000379). The study was designed to investigate patterns and relationships between digital device use, sensor based measures (including both behavioral and physiological signals), and self-reported measures of mental health and well-being. The study duration was fourweeks long per participant and included a wearable for the complete four-week period. The study and 27
Towards a General Intelligence and Interface for Wearable Health Data
recruitment were designed to increase participation of under-represented groups, as defined by race and ethnicity (e.g., Caucasian, African American, Asian, Latina/Latino, Native Americans/ Indigenous Populations), biological sex at birth (Female; Male), age (18 - 40; over 40), sexual orientation/gender identity (Heterosexual; LGBTQIA+). Study Population. The study was limited to participants aged 18-80 with an Android smartphone capable of installing the Google Health Studies application. Participants were further required to be free of major health conditions which severely restricted mobility and physical activity. A subset of this group, who owned a Fitbit device, were invited to share their Fitbit data. Baseline and Post-Study Questionnaires. Participants completed a battery of self-report questionnaires at baseline and study conclusion, as follows: (Baseline only) Demographics questionnaire, Sleep and Health Habits questionnaire, Patient Health Questionnaire (PHQ-8) (Kroenke et al., 2009), Generalized Anxiety Disorder Scale (GAD-7) (Löwe et al., 2008; Spitzer et al., 2006), Patient Reported Outcomes Measurement Information System (PROMIS) Sleep Disturbance & Sleep Related Impairment short form (Cella et al., 2010; Yu et al., 2012), Shortened Smartphone Addiction Scale (SAS) (Kwon et al., 2013), Perceived Stress Scale (PSS) (Cohen et al., 1983). Ecological Momentary Assessment Surveys. During the first and last week of the study, participants were asked to fill out three EMA surveys spread throughout the day. Each EMA assessed the participant’s mood across five affects (happy, calm, anxious, sad, and stressed), each assessed on a single-select 5-item Likert scale, spanning not at all to very. Participants were also asked to report who they had spent the most time with (since the previous EMA). Options included Alone, Friends, Family, Spouse/Partner, Co-workers, Co-students. Daily Status Reports. Each morning, throughout the four-week period, participants were prompted to report how they had been feeling over the past day. Participants responded via a single-select 5-item Likert scale spanning very bad to very good. Mobile and Wearable Data. Data was longitudinally collected from participants’ devices over the course of the four-week study. Mobile phone metrics included screen on time, usage time by application category, battery status, and number of phone unlocks. Aggregated measures of geolocation were also collected, binning the participants locations at either being at home, work, or other. For a subset of participants, wearable sensor data was also logged continuously. Note that sensor data, described in this work, refers to the data from these wearable devices rather than from mobile phones. Compensation. Subjects were entered into a raffle to win a $50 gift card with an 11.5% win probability. The conditions for eligibility were to: 1) consent and enable sensor collection at study start, 2) complete the pre-study assessments, 3) complete daily status assessments for a minimum of seven study days (one week, cumulatively), and 4) to complete the post-study assessments. M.2.2.4. Downstream Data Statistics Table ED.2 shows the counts and distributions of participants who feature in the pretraining and downstream datasets. The pretraining set includes 5,020,000 unique participants and the downstream set 13,985 unique participants. We compare these to US Census and CDC statistics and also to the distribution of data in the widely used All of Us dataset. While our datasets represent a broad group from the global (pretraining) and US (downstream) populace there are some areas in which representation is still lacking, pretraining data is skewed towards women who are more frequent adopters of Fitbit devices and White/Caucasians. Figure ED.4 shows the geographic distributions of the country or US state which these participants registered as their home state. The pretraining data was sampled globally with representation from over 100 countries. The downstream data was
28
Towards a General Intelligence and Interface for Wearable Health Data
limited only to US participants with representation from all 50 states. Note that the age and weight distributions appear similar between pretraining and downstream datasets. Table ED.2 | Demographics. Counts and distributions of pretraining and downstream study populations. n.c. = Data not collected. Category
Age 18–39 40–59 60–79 80+ Unknown Gender Male Female Non-Binary Unknown BMI Underweight (<18.5) Healthy (18.5-25) Overweight (25–30) Obese (≥30) Unknown Height <170cm 170–183cm >183cm Unknown Weight <63kg 63–86kg >86kg Unknown Ethnicity White/Caucasian Black Hispanic Asian Native Mixed Race Unknown Total
Pretraining
Downstream
US Census
All of Us (%)
N
%
N
%
2023 (%)
1,944,258 1,772,338 1,083,315 104,315 115,774
38.7% 35.3% 21.6% 2.1% 2.31%
5,587 6,307 1,914 19 158
40.0% 45.1% 13.7% 0.1% 1.1%
29.7% 25.4% 17.9% 4.0%
25.0% 38.7% 32.4% 3.8%
2,009,581 2,943,605 n.c. 66,814
40.0% 58.6% n.c. 1.33%
4,835 8,601 388 387
34.6% 61.5% 2.8% 1.2%
50.8% 49.2%
34.0% 63.8%
231,866 1,776,477 1,532,795 1,400,870 77,992
4.6% 35.4% 30.5% 27.9% 1.6%
162 3,303 4,114 6,062 342
1.2% 23.6% 29.4% 43.4% 2.4%
1.6%† 27.0%† 31.1%† 40.3%†
3.7% 39.1% 22.1% 35.1%
2,647,970 1,966,545 334,550 70,935
52.8% 39.2% 6.7% 1.4%
7,112 5,561 1,091 219
50.9% 39.8% 7.8% 1.6%
56%†† 37%†† 7%††
62.8% 31.5% 5.7%
1,186,701 2,262,533 1,498,675 72,091
23.6% 45.1% 29.9% 1.4%
1,660 5,681 6,302 315
11.9% 40.7% 45.1% 2.3%
11.9% 41% 42%
25.6% 24.8% 49.6%
n.c. n.c. n.c. n.c. n.c. n.c. n.c.
11,030 563 773 582 88 60 889
78.9% 4.0% 5.5% 4.2% 0.6% 0.4% 6.4%
57.8% 18.7% 12.1% 5.9% 0.7% 4.1% 0.7%
54.7% 15.4% 14.4% 3.45%
5,020,000
13,985
7.4% 0.5%
† Source: CDC National Center for Health Statistics (NHANES 2021–2023 & 2017–2018).
†† Source: CDC Anthropometric Reference Data for Children and Adults (NHANES 2015–2018).
M.3. Modeling M.3.1. Model Architecture Prior work has shown the utility of using a masked reconstruction (He et al., 2022) pretraining objective for modeling long-context (hours or days) multimodal sensor data (Erturk et al., 2025; 29
Towards a General Intelligence and Interface for Wearable Health Data
Figure ED.3 | Global Demographic Distributions Geographic and demographic distribution of pretraining and downstream datasets. (a) Global distribution of pretraining participants across countries. (b) US state-level distribution of pretraining data. (c) US state-level distribution of downstream study participants. (d) Age and weight distributions in the pretraining cohort, (e) age and weight distributions in the downstream cohort. Note the log scale of the Y axis for (d) and (e).
30
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.3 | Downstream Tasks. Summary and statistics of downstream data labels. For task type REG is regression, CLS is binary classification. For data source SLF is self-reported, LAB is lab tested, SCR is standardized screener survey. Category
Definition
Type
Source
N (Test)
% Pos. or Mean
Chronological age in years. Body-mass index in kg/m2 . Height in centimeters. Weight in kilograms.
REG REG REG REG
SLF SLF SLF SLF
13,831 13,660 13,768 13,777
44.3 years 30.2 kg/m2 1.68 meters 85.7 kg
Working status. Disability status. Disability impacts working Smoking behavior. On Medicaid Insurance. No regular medication use in daily life.
CLS CLS CLS CLS CLS CLS
SLF SLF SLF SLF SLF SLF
5,950 5,737 956 1,500 5,940 8,034
80.8% 17.2% 70.9% 8.1% 12.3% 44.9%
Diagnosis of cardiovascular condition Diagnosis of hypertension Diagnosis of cardiovascular condition 10 year risk atherosclerotic cardiovascular disease 10 year risk of cardiovascular disease 30 year risk of cardiovascular disease
CLS CLS CLS REG
SLF SLF SLF SLF
1,524 360 261 417
2.3% 23.6% 17.1% 0.04
REG REG
SLF SLF
417 412
0.06 0.27
CLS CLS CLS CLS CLS REG
SLF SLF LAB LAB LAB LAB
1,524 1,659 1,524 778 779 779
8.6% 8.6% 21.7% 11.1% 32.2% 3.3
REG REG
LAB LAB
778 781
5.5 mmol/mol 112.2 mg/dL
Depression score (PHQ-8 ≥ 10) Anxiety disorder score (GAD-7 ≥ 10) Stress score (PSS ≥ 14) Diagnosis of depression/anxiety Depression/anxiety med. Depression score Generalized anxiety disorder score Perceived stress scale score and subscores
CLS CLS CLS CLS CLS REG REG REG
SCR SCR SCR SLF SLF SCR SCR SCR
4,241 4,267 5,955 10,615 8,034 4,241 4,267 5,955
29.2% 24.3% 64.6% 21.9% 11.4% 7.1 6.4 16.6
For individuals with a sleep disorder whether they are on a treatment Sleep disturbance screener score Sleep impairment screener score
CLS
SLF
850
65.6%
REG REG
SCR SCR
5,955 5,955
20.2 19.9
Demographics Age BMI Height Weight Lifestyle Currently Working Disability Disability Affects Work Smoking Medicaid No Medications Cardiovascular Cardiovascular Dx Hypertension Dx Respiratory Dx ASCVD Risk Framingham Risk Framingham 30 Risk Metabolics Diabetes Dx Diabetes Med Hyperlipidemia Pre-Diabetes Insulin Resistance HOMA-IR HbA1c Triglycerides Mental Health Mild Depression Mild Anxiety Persistent Stress Depression/Anxiety Dx Mental Health Med. PHQ-8 GAD-7 PSS Sleep Sleep Disorder Treatment Sleep Disturbance PRO Sleep Impairment PRO
Diagnosis of diabetes condition Use of a Diabetes Medication Hyperlipidemia Dx Classification of HbA1c ≥ 5.7 Classification of HOMA-IR ≥ 2.9 Homeostatic Model Assessment for Insulin Resistance score Hemoglobin A1c Score Triglyceride level
31
Towards a General Intelligence and Interface for Wearable Health Data
a
b
c
d
e
f
g
h
i
j
k
l
m
n
o
a
b
c
d
e
f
g
h
i
j
k
l
m
n
o
p
q
r
s
t
Figure ED.4 | Task Label Statistics. Distributions of regression (top) and classification (bottom) task labels across the 35 discriminative downstream tasks derived from multiple prospective studies.
32
Towards a General Intelligence and Interface for Wearable Health Data
Narayanswamy et al., 2025; Xu et al., 2025). An important consideration in modeling these longcontext sensor streams is the fragmentation inherent in the data, with missingness occuring for reasons such as loss of charge events, intermittent removal of the device, sensor or environmental noise, or various device operation modes (see Figure ED.1). To this end, we designed a pretraining technique (see Figure ED.2) that patches multimodal data inputs and replaces patches with missing observations with mask tokens. Following the Adaptive and Inherited Masking (AIM) strategy introduced by Xu et al. (2025), our method treats the full applied mask as the union of the missing data mask and the artificial mask generated for the reconstruction pretraining objective. Two stage token masking, leveraging token dropout and attention masking, is used to ensure flexibility under the inherent variability of real world missingness while retaining computational efficiency. Additional details regarding AIM maybe found in the original paper (Xu et al., 2025). Specifically, our model leverages a ViT-1D encoder backbone (Dosovitskiy et al., 2021), with the hyperparameters described in Table ED.4. During AIM-based pretraining the latent embeddings are passed through a ViT-1D decoder which learns to reconstruct the ablated input. Masked tokens are represented through a learnable mask token. To represent temporal and sensor feature dimensions of the input, 2D additive positional encodings are applied to the tokens. Half of the encoding dimensions, corresponding to the feature dimension, are learned. Half of the positional encoding dimensions, corresponding to the temporal dimension, are 1D sinusoidal encodings (Vaswani et al., 2017), except for eight embedding features which correspond to cyclic datetime features (Spathis et al., 2021): minute of the hour [0, 59], hour of the day [0, 23], day of the week [0, 6], and day of the year [0, 364]. All SensorFM models were pretrained with a mean-squared error (MSE) reconstruction loss. The loss was calculated exclusively on the masked tokens which were originally observed (i.e., not missing). Table ED.4 | Model Configurations for Each Size. Architecture parameters for our ViT (Dosovitskiy et al., 2021) MAE-based (He et al., 2022) SensorFM across different model size variants, along with their total parameter counts. Model Variant
Parameter XXS
XS
S
B
Encoder Hidden Size MLP Dimension Number of Heads Number of Layers
64 256 1 2
128 512 2 4
256 1024 4 8
768 3072 12 12
Decoder Hidden Size MLP Dimension Number of Heads Number of Layers
48 192 1 1
96 384 2 1
192 768 4 2
512 2048 16 8
Total Parameters
138,740
933,204
7,290,068
110,763,412
M.3.2. Data Curation and Preprocessing While large pretraining datasets are essential for building foundation models, evidence suggests that rigorous curation pipelines significantly improve model quality (Grattafiori et al., 2024). To ensure the integrity of our trillion-minute data corpus, we carefully implemented a pipeline for 33
Towards a General Intelligence and Interface for Wearable Health Data
both pretraining and downstream tasks. First, to harmonize diverse sensor streams, all inputs were resampled to a uniform one-minute resolution and corrected for timezone offsets. We then applied physiological masking to eliminate non-biological artifacts. Skin conductance level (SCL) values were restricted to the physiological range of 0–60 𝜇 Siemens, and skin temperature readings were constrained to 0 − 41◦ C. Blood oxygen saturation (𝑆 𝑝𝑂2 ) values below 70% or flagged as invalid were treated as missing to remove spurious drops, while values exceeding 100% were capped. For heart rate variability, variance metrics (RMSSD, SDNN) were capped at 125 ms to limit the impact of extreme outliers, and all HRV metrics were nullified if derived from windows with < 20% valid inter-beat intervals. Additionally, data collected during periods where the device was detected as off-wrist were discarded to prevent noise injection. Finally, to stabilize training dynamics against remaining outliers, all features were z-score normalized using global per-feature statistics derived from the pretraining corpus and clipped to the interval [−5, 5]. We applied a sliding window with consecutive windows shifted by a variable interval. This shift was randomized between 8 and 24 minutes to enhance the diversity of the generated data windows. The windows with more than 80% of inherent data missingness were removed. In total, our largest pretrain data volume contains 175, 062, 146 day-long samples. M.3.3. Pretraining We pretrain our model using a self-supervised masked autoencoder framework designed to reconstruct data from incomplete inputs. Our masking strategy AIM (Xu et al., 2025) handles both inherent ("inherited") missingness and synthetic ("artificial") masking. Crucially, the reconstruction loss is computed exclusively on the artificially masked patches where ground truth values were originally present, ignoring inherited missingness. The artificial masking strategy employs a mixed probabilistic approach to simulate diverse real-world modes of multimodal sensor data fragmentation. Specifically, for a given sample, our model randomly applies one of the following masks: 80% random patch masking, 50% temporal block masking (simulating device removal), and 50% modality block masking (simulating sensor dropout). The model utilizes a patch size of [20, 1], processing 20-minute windows per sensor feature. Training was conducted with a global batch size of 4096 and an AdamW optimizer using with a weight decay of 1 × 10−4 . We leverage a cosine annealing learning rate scheduler with a base learning rate of 5 × 10−4 and a linear warm up equal to 5% of the number of steps. We train for a maximum of 𝑁 = 1, 000, 000 steps. We train all models for 240,000 steps. The only exceptions occurr when models exhibit overfitting. In these instances, pre-training was terminated early (specifically, at 100k steps for base model with 50k subjects, 80k steps for base Model with 5k subjects, and 60k steps for small model with 5k subjects). M.3.4. Discriminative Post-Training and Evaluation We independently fine-tuned a light-weight head on top of the frozen pretrained SensorFM encoder for each downstream task to prevent conflating learning objectives. Our evaluation suite comprises a total of 35 discriminative tasks (binary and regression) which are described in more detail in Table ED.3. All the tasks correspond to person-level predictions where a single label applies to the entire participant (e.g., hypertension diagnosis, age). As such we aggregated the embeddings per person across all non-masked (no inherited missingness) tokens, computing the mean and standard deviations of each person’s embeddings across all days of their data. To better match the variance of the embeddings with the sparsity of downstream labels, we reduce the SensorFM aggregated embeddings to 50 principal components (PCA-50). For tasks where demographic features are included, age, sex, BMI, and race are concatenated with the reduced embeddings prior to the linear probe. For classification tasks, we attached a linear probe to the reduced embeddings from the frozen 34
Towards a General Intelligence and Interface for Wearable Health Data
encoder, and optionally demographic features. We trained using a logistic regression head with an AdamW optimizer (learning rate 5 × 10−3 , weight decay 1 × 10−4 ) for 500 steps. We evaluate classification performance with the Receiver Operating Characteristic Area Under the Curve ( 𝑅𝑂𝐶 𝐴𝑈𝐶 ) and the 𝐹 1 score. For regression tasks, we employed a similar setup with a linear regression head, evaluating performance via Pearson correlation (𝑟 ) and Mean Absolute Error ( 𝑀 𝐴𝐸). To ensure robust evaluation, all presented results are the aggregated out-of-fold (OOF) performance with a five-fold cross validation setup. Note as each subject accounts for a singe data sample these folds are naturally person independent. The subjects per split remain fixed across all downstreams. Additionally, while F1 and MAE are aggregated with arithmetic mean and standard deviation, Pearson (𝑟 ) and 𝑅𝑂𝐶 𝐴𝑈𝐶 were aggregated in transform spaces to account for their skewness; Pearson correlation in the z-transform space, and ROC AUC the logit-transform space, before being back-transformed. Standard deviation was also calculated in the transformed space and back-transformed to give asymmetric error values. M.3.5. Generative Evaluation Since SensorFM was trained with an MAE-like (He et al., 2022; Xu et al., 2025) objective, it naturally exhibits out-of-box generative capabilities which allow it to infill and extend unobserved multimodal sensor data. To evaluate these generative capabilities we formulate the following generative tasks. Random Imputation masks out 80% of total tokens across signals and time, emulating generic random noise (e.g., random signals missing at random times). Temporal Imputation and Temporal Extrapolation ablation of a contiguous temporal window of length [10, 30, 60 minutes], either in the middle or at the end of the sequence, emulating intermittent removal, loss of charge events, etc. (e.g., all sensor features missing for a contiguous block of time). Signal Imputation masks all time points for a random set of [2/26, 6/26, 12/26] signal channels, emulating missing sensor channels (e.g., various sensor loadouts, sensor dropout, non-random missingness). Reconstruction performance was calculated with mean squared error (MSE) calculated originally observed masked tokens, averaging only over the data points that have a ground truth. 95% confidence intervals were generated by bootstrapping the generative errors across 100 iterations. The results of this evaluation are presented in Table ED.10. To better contextualize the utility of these generative capabilities, we evaluate SensorFM’s ability to reconstruct partially observed data to provide more meaningful daily summary statistics. Specifically, we simulate missingness by masking a contiguous 1-hour duration across all sensor features. We then use SensorFM to reconstruct this missing data. We predict daily metric predictions across Steps, Sleep Stage Minutes, Exercise Minutes, SPO2 Level Minutes, and Wrist Temperature Level Minutes. We compare SensorFM to the ground truth, un-ablated data, and baseline against a method which does not infill missingness. Total recovered minutes from the hour of data loss were compared to the ground truth, with a 95% confidence intervals generated from 100 bootstrap iterations. The results of this experiment are presented in Table ED.11. As mentioned in Methods M.2.1, both the generative evaluation and the daily metrics estimation are reported on an independent test set derived from 10, 000 subjects. M.3.6. Engineered Baseline Features To contextualize the performance of SensorFM embeddings, we compare against supervised models trained on engineered features derived from the same sensor streams. These engineered features are described in detail in Table ED.14 in the Appendix, and are described in a high level below. For each participant-day, we compute a fixed-length summary vector by aggregating each of the 34 minutely sensor features (Table ED.1) over the 24-hour window using daily summary statistics, following established methodologies in wearable research and chronobiology. We extracted 20 distinct 35
Towards a General Intelligence and Interface for Wearable Health Data
features per channel (680 total) capturing distributional, volatile, and chronobiological dynamics. To preserve information about data fragmentation, the missing rate was calculated on the raw data for each channel prior to any imputation. Missing data were then resolved via linear interpolation with back/forward filling at the start and ends of sequences to ensure continuous temporal derivatives. • Distributional and behavioral: Missingness and signal sparsity were quantified as hardware failure / behavioral phenotypes (Choi et al., 2011; Torous et al., 2017). Signal dispersion was captured via the mean, median, IQR, skewness, kurtosis, CV, and RMS. To robustly capture physiological extremes against high-frequency artifacts, 5th and 95th percentiles were used in lieu of absolute minimums and maximums. • Volatility: Short-term rate-of-change was measured using mean absolute minute-to-minute differences and the root mean square of successive differences (RMSSD) (Shaffer and Ginsberg, 2017), an established metric for epoch-to-epoch variability in digital phenotyping. • Morphology: Signal fragmentation and bandwidth were quantified via the mean-centered Zero Crossing Rate and Hjorth Complexity (Hjorth, 1970). These time-domain metrics were selected to effectively capture the mean frequency and bandwidth of the signals without requiring computationally intensive Fourier transforms on the longitudinal data. • Chronobiological: Diurnal rhythms were modeled via 24-hour Cosinor rhythmometry to extract the amplitude and acrophase (Cornelissen, 2014). Circadian fragmentation was assessed through Intradaily Variability (IV) (Van Someren et al., 1999), while temporal signal persistence was measured via lag-1 autocorrelation. These features represent the conventional approach to wearable health prediction: daily summary statistics fed to a standard classifier or regressor. We denote this baseline as “FE” (Feature Engineered) throughout. The same downstream training procedure described in Section M.3.4 is applied: logistic regression for classification tasks and linear regression for regression tasks, with identical crossvalidation splits and person-independent held-out evaluation. To address the high dimensionality of the engineered feature set (680 features), Principal Component Analysis (PCA) was used to reduce the feature vector to 50 principal components, similarly to the treatment of the SensorFM embeddings. Where demographic features are included age, sex, BMI, and race are concatenated with the reduced engineered feature vector prior to model fitting. Similar to the linear probe, supervised baselines are evaluated with five-fold cross validation. M.3.7. Few Shot Experiments To evaluate label efficiency we executed few shot experiments for each SensorFM model variant and the engineered features described in Section M.3.6. For each evaluated task the downstream models were trained using 5 folds and different sample percentages (10, 20, 30, 50, 60, 70, 80, 90 and 100). Specifically, in Figure ED.5 we visualize the few-shot performance of two SensorFM variants and supervised baselines trained with only demographics or engineered features.
M.4. Analysis of the Model Embeddings in Latent Space M.4.1. SHapley Additive exPlanation Analysis To identify the latent structure and the physiological semantics encoded within the high-dimensional SensorFM embeddings, we employed SHapley Additive exPlanations (SHAP) (Lundberg and Lee, 2017) to quantify the contribution of each embedding dimension to specific downstream tasks. As previously mentioned, we utilized a Principal Component Analysis (PCA) preprocessing step before 36
Towards a General Intelligence and Interface for Wearable Health Data
the linear probing heads to reduce dimensionality to allow for more appropriate comparisons with the baseline models, as well as reducing collinearity of the raw embedding space. To analyze the latent structure of our embeddings in this setting, we leveraged the following projection mechanism to map feature importance from the reduced PCA space back to the original SensorFM latent space. For each task, we fit a linear model 𝑓𝑖 ( 𝑥 ) on the PCA-transformed dimension (𝑍 in Eq. (M.1)), (M.1)
𝑍 = 𝑋𝑉 ⊤
where 𝑋 ∈ ℝ𝑁 × 𝐷 denotes the input data that is centered (𝔼[ 𝑋 ] = 0), and 𝑉 ∈ ℝ 𝐾 × 𝐷 are the principal components; in this work, we chose 𝐾 = 50. Each linear prediction head learns a set of coefficients 𝛽 ∈ ℝ 𝐾 ×1 , predicting ˆ 𝑦 = 𝑍 𝛽 + 𝑏 for 𝑏 ∈ ℝ. From here, there are two ways of mapping the importance of the PCA-transformed dimensions back to the original SensorFM embeddings: Exact Analytical Weight Collapse. To derive exact attribution for the linear heads, we treated the PCA transformation and the linear probe as a single composite linear layer. Substituting Eq. (M.1) into the linear equation yields Eq. (M.2): ˆ𝑦 = 𝑋𝑉 ⊤ 𝛽 + 𝑏 = 𝑋 (𝑉 ⊤ 𝛽 ) + 𝑏 (M.2) We define the Effective Weight Vector 𝑊𝑒 𝑓 𝑓 ∈ ℝ 𝐷 as the projection of the probe coefficients back onto the original feature axes: 𝑊𝑒 𝑓 𝑓 = 𝑉 ⊤ 𝛽 (M.3) For a linear model with centered input features, the exact independent SHAP value for a feature is the product of the feature’s value and its corresponding weight. Thus, the exact SHAP attribution matrix Φ ∈ ℝ𝑁 × 𝐷 in the original embedding space is computed analytically as: (M.4)
Φ𝑖,𝑑 = 𝑊𝑒 𝑓 𝑓 ,𝑑 · 𝑋𝑖,𝑑
By linearity, this formulation satisfies the SHAP local accuracy (efficiency) axiom, ensuring that Í𝐷 𝑦𝑖 − 𝑏. 𝑑 =1 Φ𝑖,𝑑 = ˆ For multiclass logistic regression tasks with a set of classes 𝐶 , an effective weight vector 𝑊𝑒(𝑓𝑐 )𝑓 = (𝑐)
(𝑐)
𝑉 ⊤ 𝛽 ( 𝑐 ) is computed for each class 𝑐 ∈ 𝐶 . The class-specific SHAP value is strictly Φ𝑖,𝑑 = 𝑊𝑒 𝑓 𝑓 ,𝑑 · 𝑋𝑖,𝑑 . To
establish a singular task-level attribution, we aggregated the absolute contributions across all classes: Φ̄𝑖,𝑑 =
1 ∑︁ ( 𝑐 ) Φ | 𝐶 | 𝑐 ∈ 𝐶 𝑖,𝑑
(M.5)
To ensure robustness against data splits, we computed SHAP values for each subject based on their out-of-fold predictions across a 5-fold cross-validation scheme. The global importance 𝐼𝑑 of embedding dimension 𝑑 was defined as the mean absolute SHAP value across all 𝑁 subjects in the dataset: 1 ∑︁ 𝑁
𝐼𝑑 =
𝑁
| Φ̄𝑖,𝑑 |
(M.6)
𝑖=1
Latent Profile Similarity and Network Visualization. To visualize the shared physiological semantics encoded within the SensorFM embeddings across different clinical domains, we modeled the pairwise relationships between downstream tasks based on their exact SHAP attribution profiles. Let 𝐼𝑡 ∈ ℝ 𝐷 denote the global importance vector for task 𝑡 , computed across all embedding dimensions. To ensure scale invariance across tasks with varying absolute prediction margins, each attribution profile was first max-normalized: ˜𝐼𝑡 = 𝐼𝑡 /max( 𝐼𝑡 ). We then computed the pairwise 𝐿1 (Manhattan) distance 37
Towards a General Intelligence and Interface for Wearable Health Data
between all normalized profiles. Let Δ𝑡,𝑡′ = ∥ ˜𝐼𝑡 − ˜𝐼𝑡′ ∥ 1 represent the distance between tasks 𝑡 and 𝑡 ′ . The latent profile similarity matrix 𝑆 ∈ ℝ𝑇 ×𝑇 for the 𝑇 tasks was defined as: 𝑆𝑡,𝑡′ = 1 −
Δ𝑡,𝑡′
max𝑢,𝑣 Δ𝑢,𝑣
(M.7)
where 𝑆𝑡,𝑡′ = 1 indicates identical latent utilization of the embedding space, and 𝑆𝑡,𝑡′ = 0 indicates maximum divergence. M.4.2. Embedding Distances and Intrinsic Dimensionality To evaluate the structural density of the latent space across model scales, we computed pairwise Euclidean distances between user embeddings. To ensure computational tractability and prevent memory constraints, we randomly subsampled 5,000 users for each model size. Missing values within the embeddings were resolved using mean imputation. To denoise the representations prior to distance calculation, the embedding dimensionality was reduced using Principal Component Analysis (PCA), retaining 99% of the variance. The condensed pairwise Euclidean distances were then visualised using kernel density estimation (KDE) to assess the dispersion and clustering of the latent representations across different model capacities. The intrinsic dimensionality and compressibility of the learned representations were quantified by examining the cumulative explained variance. Following mean imputation of the embeddings, we applied PCA to calculate up to 75 principal components (or the maximum available dimensions for smaller models). The explained variance ratio for each component was extracted and cumulatively summed to generate scree plots. This allowed us to identify the presence of dimensional collapse or anisotropy across the different model sizes by observing how quickly the explained variance saturated.
M.5. Agentic Classroom Search for Predictive “Head” Development To search the space of task-specific prediction "heads" for SensorFM, we developed a framework for self-evolving algorithm generation that formulates solution synthesis as a competitive, collaborative optimisation problem solved by a “classroom” of parallel LLM agents. We note that while this framework is helpful in the context of our model which has many task “heads”, it is applicable to any task whose solution quality can be expressed as a scalar score, not just those in the domain of wearable sensing. M.5.1. Implementation Details The framework is implemented as a lightweight Python library designed to run entirely within a Google Colab notebook environment. Student agents are instantiated as parallel calls to the Gemini API. We note that this framework is model-agnostic and compatible with any LLM exposing a textgeneration endpoint. All code execution occurs within Colab’s sandboxed Python runtime with solutions executing on CPU compute. We note that such a framework may be extended to support hardware acceleration (GPU/TPU) for more computationally intensive tasks. The architecture imposes no external infrastructure requirements beyond a Colab instance and API access, and is designed to be extensible to distributed computing backends for larger-scale experiments. Hyperparameters. Specifically, for our experiments, we instantiate a classroom of 𝑁 = 5 student agents leveraging the following underlying language models2 : gemini-2.5 flash, gemini-2.5 pro, 2 Accessed via the Google Gemini API between Feb 2026 and April 2026.
38
Towards a General Intelligence and Interface for Wearable Health Data
gemini-3 flash preview, gemini-3.1 flash lite preview, and gemini-3.1 pro preview. The classroom search is set to iterate for a maximum 𝑇 = 20 learning cycles. As highlighted in Figure ED.12 we run experiments both with and without agent collaboration, events where the agents analyze their own solutions or the solutions of other students. The score on which agents hill-climb is a combination of multiple classification or regression metrics (given the task). K-Fold Cross Validation. Similar to the setting where a linear head is applied to the SensorFM learned embeddings (M.3.4), we report 5-fold cross validation performance for each of the 35 discriminative tasks, leveraging the same folds as above (M.3.4). Specifically, for a given experiment (a specific fold and a specific task) we randomly split off 20% of D𝑡𝑟𝑎𝑖𝑛 to act as D𝑣𝑎𝑙 . We refer to the out-of-fold (OOF) data as D𝑡𝑒𝑠𝑡 . Over 𝑇 learning cycles, the classroom learns from D𝑡𝑟𝑎𝑖𝑛 and makes predictions for and iterates on the performance on D𝑣𝑎𝑙 . At the end of 𝑇 iterations, a best solution 𝑠∗ is selected based on the best validation score 𝜙∗ . 𝑠∗ is then train with the entirety of the D𝑡𝑟𝑎𝑖𝑛 (including the originally split validation data), and evaluated on the OOF D𝑡𝑒𝑠𝑡 to produce the reported results. By formulating the evaluation across 𝐾 folds as 𝐾 independent experiments, we effectively prevent train-test leakage. However, it should also be noted, that this formulation results in different “found” solutions per fold for a given task. Summary and Example Artifacts. In total, leveraging this framework we efficiently conduct 30, 516 total experiments across 35 tasks x 5 folds x 5 student agents x 20 learning iterations x 2 and collaboration conditions (with and without). Note that the number of total experiments accounts for students which did not conduct a full 20 learning iterations because of the patience criteria. An example student agent prompt can be found in Code ED.1 and an example agent solution can be found in Code ED.2 of Appendix E.
M.6. Evaluating SensorFM as a Tool for a Health Agent To rigorously evaluate whether integrating SensorFM improves the clinical utility of LLM-driven health agents, we designed a blinded, comparative study centered on diverse, real-world patient profiles. The health agent was tasked with generating personalized health summaries under varied contextual conditions with a full integration of our AI inferences. These generated responses were then subjected to a blinded evaluation by a panel of board-certified physicians. By assessing the outputs across multiple dimensions of clinical utility and safety, this experimental design allows us to isolate and quantify the specific value SensorFM adds to personal health agents. M.6.1. Generating the SensorFM-Augmented Responses The health agent was tasked with synthesizing a comprehensive summary of each user’s health status based on their specific data profile. The LLM powering the health agent was kept constant as Gemini 3 Flash with a temperature of 0.2. The exact system prompt designed to govern these responses is detailed below in Table ED.3. The prompt is structured to ingest multimodal user context of demographics, aggregated wearable statistics, and SensorFM inferences, while enforcing formatting and safety guardrails (e.g. qualitative interpretation of predictive metrics). The full metabolic panel is not provided in the agent prompt. M.6.2. Testing SensorFM To evaluate the clinical utility of the inferences generated by SensorFM, we established three distinct experimental conditions. The first condition, (A) Extra Context (SensorFM Predictions), provided the Health agent with user demographics, feature-engineered daily metrics, and a diverse range of 39
Towards a General Intelligence and Interface for Wearable Health Data
SensorFM predictions (e.g. hyperlipidemia, PHQ-8, sleep disturbance). The second condition, (B) Extra Context (Available Ground Truth), mirrored condition A but replaced the model predictions with the available ground-truth targets. Because only metabolic markers were available in the downstream evaluation dataset, other targets (such as mental health and sleep disturbance metrics) were inherently absent. Finally, the baseline condition, (C) No Extra Context, supplied the health agent strictly with demographics and feature-engineered daily metrics. During evaluation, the presentation order of these conditions was randomized to prevent reviewer bias. M.6.3. User Profiles To ensure the robustness and generalizability of our evaluation, we constructed ten representative user profiles that encapsulate a diverse range of common health scenarios. These profiles were stratified into two primary categories. The first category includes four profiles representing individuals without diagnosed chronic diseases but with distinct health objectives: (1) a performance-oriented individual training for athletic goals; (2) a generally healthy individual seeking to improve a specific wellness aspect, such as sleep quality; (3) an individual with a sedentary lifestyle but no formal disease diagnosis (sub-healthy); and (4) an individual recovering from an acute injury or life event disrupting their health baseline. The second category comprises six profiles designed to reflect major public health concerns, with each profile centered on a prevalent chronic condition: (5) Anxiety/Depression, (6) Hypertension, (7) Respiratory Conditions, (8) Hypercholesterolemia, (9) Diabetes, and (10) Cardiovascular Disease (CVD). It is noteworthy that these profiles were designed to reflect real-world complexity, and individuals within these latter six categories may present with comorbidities. We pull the users used in Heydari et al. (2025) that have corresponding wearable data, resulting in 12 users from the healthy profiles and 19 from the unhealthy profiles. For each individual, we extracted their statistical aggregrate wearable data, demographic information, and available groundtruth metabolic markers. Additionally, we take their minutely wearable data and use our SensorFM for the AI model predictions. M.6.4. Physician Evaluators A cohort of four board-certified physicians with specialties in internal medicine and family medicine, with an average of 11.75 years of clinical experience was recruited to evaluate the clinical soundness of the generated summaries. The physicians reviewed the responses in a randomized order and were blind to the experimental conditions. Specifically, they were not informed how inferences were integrated into the provided text and that predictive AI models were used to generate inferences. They were tasked solely with rating the clinical quality of the responses according to the established rubric, without knowledge of the underlying system architecture. To validate the consistency of the physician ratings, we calculated the Intraclass Correlation Coefficient (ICC3k), utilizing a two-way mixed-effects model based on the average score of the raters for each rubric dimension. The analysis highlights the nuanced, highly individualized nature of expert clinical evaluation. The panel demonstrated a moderate and solid consensus on Relevance (ICC = 0.653, 95% CI: [0.52, 0.76]), indicating shared agreement on the most critical medical information. For dimensions requiring more subjective clinical interpretation, such as Context (ICC = 0.478), Harm (ICC = 0.416), and Personalization (ICC = 0.387), the scores reflect the expected, natural variance inherent to diverse medical practices. Furthermore, the broad variance in the Justifiable dimension (ICC = -0.088) underscores that physicians maintain uniquely stringent, individualized 40
Towards a General Intelligence and Interface for Wearable Health Data
thresholds for what constitutes "justifiable" clinical reasoning when operating strictly from static data profiles. M.6.5. Evaluation Rubric To evaluate the generated summaries, we developed a comprehensive rubric comprising five distinct dimensions: Context, Personalization, Justifiability, Relevance, and Harm. Each dimension assesses a unique facet of clinical utility, specifically targeting the agent’s ability to ground its responses in patient-specific data to deliver safe, tailored advice. The exact evaluation criteria for each dimension are detailed in Survey ED.1. M.6.6. Evaluation Set-up During evaluation, physician evaluators were presented with the user’s demographics, daily aggregated wearable statistics, and the full metabolic panel. Crucially, the evaluators were strictly blind to the underlying SensorFM predictions. This blinding was implemented to prevent anchoring bias, ensuring that the physicians assessed the clinical soundness of the generated responses based entirely on their own independent medical judgment of the raw patient data, rather than being influenced by the AI’s diagnostic estimates. For each patient profile, physicians were explicitly instructed to read all three standardized model responses (anonymized + randomized as Models A/B/C) side-by-side. Evaluators assessed all three Model A/B/C responses for a single rubric dimension before proceeding to the next question. This parallel presentation format enabled evaluators to score the responses using a 5-point Likert scale while inherently facilitating relative comparisons between the models, allowing for the extraction of both absolute quality metrics and comparative win-rates. Author contributions GN, MAX, XL, DM contributed to the conception and design of the work; GN, MAX, AH, SAG, BY, AW, NBA, JH, CH, AM, XL, DM contributed to the data acquisition and curation; GN, MAX, AH, SAG, MG, KV, ZZ, JG, LA, HY, AW, XL, DM contributed to the technical implementation; DS, HY, YK, YZ, SS, YY provided technical and infrastructure guidance; JS, IGL, JH, JG provided clinical inputs to the study; DS, HP, OX, DB, KA, PK contributed to the supplementary data analysis; GN, MAX, AH, SAG, MG, KV, ZZ, JG, LA, DS, HY, HP, OX, DB, JB, JM, YK, YZ, NR, SS, KA, TA, JS, MZP, BY, AW, NBA, JMR, IGL, YL, JH, AP, CH, YY, AM, PK, MM, SP, XL, DM contributed to the drafting and revising of the manuscript. Correspondence Correspondence should be addressed to {girishvn, xumax, xliucs, dmcduff}@google.com.
41
Towards a General Intelligence and Interface for Wearable Health Data
Appendix A. Scaling Results We sweep four SensorFM model variants (XXS, XS, S, B; spanning 105 to 108 parameters) across four pretraining data volumes (5K, 50K, 500K, and 5M subjects; spanning 107 to 109 data hours). Table ED.5 provides an overview of model performance across the pretraining objective (validation reconstruction loss), generative tasks (random imputation, temporal interpolation/extrapolation, signal imputation), and discriminative tasks (averaged classification ROC AUC and regression Pearson correlation). Tables ED.6 and ED.7 break the discriminative linear probe results down per-task across all 35 downstream tasks (organized by Demographics, Lifestyle, Cardiovascular, Metabolic, Mental Health, and Sleep), with Pearson 𝑟 and ROC AUC in Part I (Tables ED.6) and complementary MAE and F1 metrics in Part II (Table ED.7). Table ED.8 reports the per-task improvement from including demographic features alongside SensorFM embeddings versus a supervised feature-engineered baseline; the marginal value of demographics diminishes with model scale, and 33 of 35 tasks exhibit the lowest demographic lift at B. In these tables, the SensorFM embeddings are first reduced to 50 principal components (PCA-50) to match the lower variance of the sparse downstream labels.
B. Downstream Task Results Discriminative Tasks. Table ED.9 compares a linear probe of SensorFM-B (pretrained at 5M subjects) PCA-50 reduced embeddings against supervised baselines built on engineered features and/or demographic features across all 35 discriminative tasks; SensorFM achieves the best performance on 31 of 35 tasks. Discriminative Few-Shot Performance. Figure ED.5 shows per-task few-shot performance curves obtained by training the post-adaptation head on progressively larger fractions of the downstream training set, comparing SensorFM variants (XXS through B) against a supervised feature-engineered baseline and a demographics-only baseline across all 35 downstream tasks. In the very-low-label regime, demographic priors act as a strong predictor for many tasks, but as labeled data increases SensorFM surpasses both baselines, with the larger model variants (B) consistently outperforming smaller ones (XXS). Generative Tasks. Table ED.10 reports the generative task results, covering Random Imputation (80%), Temporal Interpolation (20/60/180 min), Temporal Extrapolation (20/60/180 min), and Signal Imputation, with SensorFM-B compared against naive baselines (Mean Fill, Nearest-Neighbor Fill, and Linear Interpolation). Table ED.11 shows the downstream impact on daily-aggregated wearable metrics (steps, sleep stage minutes, active zone minutes, SpO2, wrist temperature) under a simulated 1-hour data loss, comparing the current consumer-wearable baseline (aggregation over observed values only) to SensorFM-recovered aggregates and ground truth.
42
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.5 | The Effect of Scaling Across Pretraining, Generative and Discriminative Tasks. This table presents the performance of SensorFM across pretrain, generative, and discriminative tasks as a function of model capacity and pretrain data volume. In general larger models trained with more data achieve improved performance. In pretraining the model is tasked with reconstructing a sample ablated with either random, temporal, or signal masking. As such the validation loss is a compound generative metric consisting of Random Imp. (80%), Temporal Imp. (50%) and Signal Imp. (50%). For pretraining and generative tasks we present the average reconstruction Mean Squared Error (MSE) with 95% bootstrapped confidence intervals generated through 100 bootstrap iterations. For discriminative tasks we present the mean performance across all tasks, where each task is evaluated with 5-fold cross validation. Average Receiver Operating Characteristic Area Under the Curve (ROC AUC) is calculated in the logit-transform space and back-transformed. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Colors are normalized across each task block; best model performance is bolded and has the deepest shade. Task
Model Variant (Parameter Count)
Pretraining Data Volume (Subjects)
XXS (105 )
XS (106 )
S (107 )
B (108 )
MSE
5K 50K 500K 5M
0.428 ± 0.002 0.415 ± 0.002 0.419 ± 0.001 0.414 ± 0.002
0.364 ± 0.002 0.351 ± 0.002 0.349 ± 0.002 0.350 ± 0.002
0.402 ± 0.003 0.319 ± 0.002 0.307 ± 0.002 0.306 ± 0.001
1.082 ± 0.003 0.466 ± 0.002 0.299 ± 0.001 0.285 ± 0.002
MSE
5K 50K 500K 5M
0.389 ± 0.001 0.372 ± 0.001 0.375 ± 0.001 0.371 ± 0.001
0.303 ± 0.001 0.293 ± 0.001 0.290 ± 0.001 0.292 ± 0.001
0.321 ± 0.001 0.250 ± 0.001 0.241 ± 0.001 0.240 ± 0.001
1.077 ± 0.002 0.400 ± 0.001 0.227 ± 0.001 0.215 ± 0.001
MSE
5K 50K 500K 5M
0.478 ± 0.004 0.458 ± 0.005 0.464 ± 0.003 0.451 ± 0.003
0.562 ± 0.006 0.608 ± 0.004 0.584 ± 0.005 0.668 ± 0.006
0.499 ± 0.006 0.390 ± 0.005 0.399 ± 0.003 0.389 ± 0.004
1.077 ± 0.009 0.584 ± 0.007 0.373 ± 0.004 0.353 ± 0.002
MSE
5K 50K 500K 5M
0.608 ± 0.005 0.542 ± 0.004 0.545 ± 0.003 0.539 ± 0.004
0.639 ± 0.005 0.778 ± 0.009 0.726 ± 0.009 0.841 ± 0.006
0.597 ± 0.005 0.497 ± 0.005 0.505 ± 0.003 0.503 ± 0.003
1.100 ± 0.006 0.721 ± 0.004 0.490 ± 0.003 0.463 ± 0.004
MSE
5K 50K 500K 5M
0.321 ± 0.002 0.302 ± 0.002 0.307 ± 0.001 0.305 ± 0.003
0.250 ± 0.003 0.237 ± 0.002 0.236 ± 0.001 0.236 ± 0.001
0.270 ± 0.002 0.206 ± 0.002 0.193 ± 0.002 0.192 ± 0.001
1.093 ± 0.005 0.316 ± 0.003 0.184 ± 0.001 0.170 ± 0.001
ROC
5K 50K 500K 5M
.664 .663 .663 .663
.687 .681 .681 .682
.690 .712 .710 .716
.634 .692 .746
5K 50K 500K 5M
.386 .390 .371 .402
.426 .435 .423 .427
.453 .522 .536 .559
Metric
Pretraining Reconstruction (Val. Loss) Generative Tasks Random Imp. (80%)
Temporal Interp. (30 Mins)
Temporal Extrap. (30 Mins)
Signal Imp. (35%) Discriminative Tasks Classification (Avg. Performance)
Regression (Avg. Performance)
𝑟
.752 .314 .480 .608
.612
43
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.6 | Discriminative Task Performance Across Model Scales (Part I). The Table presents the performance of SensorFM variants, pretrained with proportional data scales, on 35 discriminative tasks. In general performance improves with scale with B consistently achieving the best performance. SensorFM variants are post-trained with PCA-50 reduced embeddings. For each task, we report the average 5-fold cross validation performance. Average Receiver Operating Characteristic Area Under the Curve (ROC AUC) is calculated in the logit-transform space and back-transformed. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Standard deviations are calculated in the transformed space and back-transformed to give asymmetric error values. Colors are normalized per row; best model performance is bolded and has the deepest shade. Prediction Task
Type
Metric
Model Variant (Parameter Count) XXS (105 )
XS (106 )
S (107 )
B (108 )
Demographics Age BMI Height Weight
REG REG REG REG
𝑟 𝑟 𝑟 𝑟
.716+−..009 009 .445+−..021 021 .485+−..013 013 .362+−..021 021
.759+−..007 008 .494+−..015 016 .518+−..013 014 .455+−..013 013
.843+−..006 006 .701+−..008 008 .634+−..015 015 .700+−..012 012
.920+−..004 005 .809+−..007 007 .675+−..012 012 .809+−..007 007
Lifestyle Currently Working Disability Disability Affects Work Smoking Medicaid No Medications
CLS CLS CLS CLS CLS CLS
ROC ROC ROC ROC ROC ROC
.763+−..021 022 .689+−..024 025 .541+−..033 034 .710+−..044 048 .699+−..036 039 .676+−..003 003
.787+−..014 015 .705+−..015 016 .616+−..051 054 .721+−..033 036 .727+−..028 030 .688+−..007 007
.829+−..014 015 .717+−..021 022 .626+−..059 063 .752+−..051 059 .762+−..022 023 .709+−..007 007
.912+−..012 014 .753+−..020 021 .699+−..040 043 .870+−..060 098 .814+−..021 023 .739+−..009 009
Cardiovascular Cardiovascular Dx Hypertension Dx Respiratory Dx ASCVD Risk Framingham Risk Framingham 30 Risk
CLS CLS CLS REG REG REG
ROC ROC ROC
.667+−..148 190 .698+−..020 021 .593+−..057 060 .542+−..080 092 .496+−..040 042 .454+−..023 024
.658+−..081 091 .737+−..026 027 .624+−..064 069 .623+−..096 119 .574+−..087 102 .580+−..078 090
.651+−..110 129 .768+−..027 030 .662+−..048 052 .675+−..093 121 .651+−..101 132 .668+−..069 082
.712+−..052 058 .786+−..023 025 .682+−..054 059 .730+−..091 127 .669+−..101 133 .714+−..052 062
Metabolic Diabetes Dx Diabetes Med. Hyperlipidemia Pre-Diabetes Insulin Resistance HOMA-IR HbA1c Triglycerides
CLS CLS CLS CLS CLS REG REG REG
ROC ROC ROC ROC ROC
.688+−..052 058 .646+−..062 067 .623+−..041 043 .636+−..070 076 .635+−..060 064 .315+−..082 087 .137+−..091 093 .088+−..084 085
.669+−..041 044 .672+−..037 039 .632+−..040 042 .658+−..015 016 .641+−..062 067 .309+−..091 097 .204+−..056 057 .091+−..056 057
.724+−..039 043 .693+−..031 033 .666+−..038 041
.706+−..020 021 .683+−..024 026 .391+−..033 034 .219+−..081 084 .207+−..036 037
.763+−..037 042 .700+−..045 049 .674+−..032 033 .704+−..068 078 .761+−..044 050 .479+−..030 031 .293+−..029 030 .269+−..047 048
Mental Health Mild Depression Mild Anxiety Persistent Stress Depress./Anxiety Dx Mental Health Med. PHQ-8 GAD-7 PSS
CLS CLS CLS CLS CLS REG REG REG
ROC ROC ROC ROC ROC
.661+−..016 016 .634+−..013 013 .659+−..021 021 .660+−..016 016 .755+−..018 019 .322+−..029 030 .273+−..025 025 .343+−..018 018
.663+−..017 017 .639+−..022 023 .660+−..028 029 .672+−..015 015 .767+−..011 011 .344+−..017 017 .293+−..013 014 .355+−..020 021
.699+−..008 008 .674+−..012 012 .689+−..030 031 .696+−..017 017 .789+−..014 014 .397+−..012 012 .346+−..011 011 .415+−..016 016
.726+−..006 006 .698+−..005 005 .712+−..026 028 .717+−..009 009 .819+−..020 022 .450+−..019 019 .400+−..025 025 .463+−..013 014
Sleep Sleep Disorder Treatment Sleep Disturbance PRO Sleep Impairment PRO
CLS REG REG
ROC
.611+−..028 028 .284+−..015 016 .346+−..019 019
.633+−..033 034 .315+−..020 020 .355+−..013 013
.651+−..050 053 .352+−..008 008 .404+−..015 015
.649+−..036 038
𝑟 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 𝑟
.390+−..016 016 .448+−..013 013
44
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.7 | Discriminative Task Performance Across Model Scales (Part II: Additional Metrics). This table extends the results presented in Part I. The Table presents the performance of SensorFM variants, pretrained with proportional data scales, on 35 discriminative tasks. In general performance improves with scale with B consistently achieving the best performance. SensorFM variants are post-trained with PCA-50 reduced embeddings. For each task, we report the average 5-fold cross validation performance. For both F1 and MAE we leverage an arithmetic mean and standard deviation across folds. Colors are normalized per row; best model performance is bolded and has the deepest shade. Prediction Task
Type
Metric
Model Variant (Parameter Count) XXS (105 )
XS (106 )
S (107 )
B (108 )
Demographics Age BMI Height Weight
REG REG REG REG
MAE MAE MAE MAE
7.035+−..135 135 5.090+−..062 062 71.226+−..499 499 15.544+−..121 121
6.539+−..101 101 4.897+−..053 053 69.460+−..348 348 14.754+−..102 102
5.405+−..082 082 3.937+−..068 068 61.985+−..784 784 11.682+−..210 210
3.865+−..075 075 3.148+−..024 024 58.773+−..829 829 9.541+−..137 137
Lifestyle Currently Working Disability Disability Affects Work Smoking Medicaid No Medications
CLS CLS CLS CLS CLS CLS
F1 F1 F1 F1 F1 F1
.805+−..002 002 .380+−..016 016 .641+−..032 032 .236+−..059 059 .312+−..021 021 .612+−..011 011
.812+−..007 007 .396+−..012 012 .690+−..031 031 .259+−..065 065 .333+−..027 027 .618+−..011 011
.838+−..005 005 .402+−..018 018 .681+−..046 046 .283+−..054 054 .367+−..029 029 .641+−..006 006
.884+−..009 009 .432+−..033 033 .727+−..023 023 .416+−..088 088 .416+−..021 021 .667+−..006 006
Cardiovascular Cardiovascular Dx Hypertension Dx Respiratory Dx ASCVD Risk Framingham Risk Framingham 30 Risk
CLS CLS CLS REG REG REG
F1 F1 F1 MAE MAE MAE
.091+−..040 040 .481+−..034 034 .313+−..058 058 .037+−..002 002 .044+−..003 003 .124+−..010 010
.090+−..043 043 .510+−..043 043 .329+−..055 055 .034+−..005 005 .040+−..006 006 .111+−..011 011
.081+−..051 051 .531+−..042 042 .353+−..060 060 .032+−..003 003 .037+−..005 005 .100+−..009 009
.108+−..061 061 .560+−..032 032 .375+−..069 069 .029+−..004 004 .035+−..005 005 .091+−..010 010
Metabolic Diabetes Dx Diabetes Med. Hyperlipidemia Pre-Diabetes Insulin Resistance HOMA-IR HbA1c Triglycerides
CLS CLS CLS CLS CLS REG REG REG
F1 F1 F1 F1 F1 MAE MAE MAE
.238+−..026 026 .225+−..044 044 .387+−..045 045 .416+−..031 031 .442+−..048 048 1.314+−..088 088 .384+−..031 031 51.480+−33..443 443
.222+−..033 033 .208+−..028 028 .397+−..045 045 .469+−..068 068 .457+−..055 055 1.311+−..101 101 .372+−..030 030 50.132+−33..860 860
.266+−..020 020 .229+−..029 029 .415+−..057 057 .469+−..053 053 .485+−..032 032 1.234+−..037 037 .372+−..025 025 48.219+−33..570 570
.300+−..038 038 .259+−..022 022 .418+−..047 047 .483+−..047 047 .557+−..053 053 1.165+−..056 056 .370+−..016 016 47.769+−33..613 613
Mental Health Mild Depression Mild Anxiety Persistent Stress Depress./Anxiety Dx Mental Health Med. PHQ-8 GAD-7 PSS
CLS CLS CLS CLS CLS REG REG REG
F1 F1 F1 F1 F1 MAE MAE MAE
.487+−..023 023 .413+−..025 025 .684+−..022 022 .414+−..018 018 .353+−..017 017 4.228+−..062 062 3.951+−..078 078 5.618+−..040 040
.482+−..021 021 .407+−..020 020 .678+−..021 021 .429+−..015 015 .358+−..019 019 4.196+−..064 064 3.925+−..079 079 5.602+−..076 076
.512+−..014 014 .442+−..024 024 .700+−..019 019 .444+−..027 027 .369+−..021 021 4.072+−..070 070 3.827+−..075 075 5.424+−..063 063
.539+−..016 016 .461+−..023 023 .717+−..022 022 .461+−..016 016 .401+−..023 023 3.948+−..078 078 3.732+−..085 085 5.272+−..048 048
Sleep Sleep Disorder Treatment Sleep Disturbance PRO Sleep Impairment PRO
CLS REG REG
F1 MAE MAE
.660+−..026 026 5.440+−..109 109 5.879+−..125 125
.684+−..037 037 5.380+−..096 096 5.844+−..122 122
.681+−..038 038 5.312+−..097 097 5.687+−..121 121
.683+−..034 034
5.205+−..084 084 5.538+−..112 112
45
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.8 | Discriminative Task Improvement Due to Demographic Features. The table presents the mean improvement in task performance caused by the inclusion of demographic features across SensorFM variants and a supervised baseline trained with engineered features. In general the effect of demographic features lessens with scale with B exhibiting the lowest change in performance on 33 of 35 tasks. SensorFM variants are post-trained with PCA-50 reduced embeddings. For each task, we report the average 5-fold cross validation performance. Average Receiver Operating Characteristic Area Under the Curve (ROC AUC) is calculated in the logit-transform space and back-transformed. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Standard deviations are calculated in the transformed space and back-transformed to give asymmetric error values. Colors are normalized per row; lowest values are bolded and have the lightest shade. Prediction Task
Type
Metric
Model Variant (Parameter Count) Feat. Eng.
XXS (105 )
XS (106 )
S (107 )
B (108 )
Lifestyle Currently Working Disability Disability Affects Work Smoking Medicaid No Medications
CLS CLS CLS CLS CLS CLS
Δ ROC Δ ROC Δ ROC Δ ROC Δ ROC Δ ROC
.014 .020 .020 .017 .003 .012
.014 .024 .027 .004 .008 .012
.009 .022 .015 .025 .004 .011
.004 .014 .014 .007 .002 .007
.000 .005 -.001 -.003 -.001 .006
Cardiovascular Cardiovascular Dx Hypertension Dx Respiratory Dx ASCVD Risk Framingham Risk Framingham 30 Risk
CLS CLS CLS REG REG REG
Δ ROC Δ ROC Δ ROC Δ𝑟 Δ𝑟 Δ𝑟
.010 .033 −0.002 .159 .182 .181
.054 .050 .011 .164 .200 .283
.015 .024 .003 .124 .154 .168
.018 .012
-0.007 .102 .098 .097
-.016 .000 -.007 .054 .055 .042
Metabolic Diabetes Dx Diabetes Med. Hyperlipidemia Pre-Diabetes Insulin Resistance HOMA-IR HbA1c Triglycerides
CLS CLS CLS CLS CLS REG REG REG
Δ ROC Δ ROC Δ ROC Δ ROC Δ ROC Δ𝑟 Δ𝑟 Δ𝑟
.006 .035 .026 .014 .007 .023 .008 .042
.025 .050 .033 .069 .040 .070 .065 .122
.011 .032 .023 .040 .025 .043 .029 .081
-.012 .016 .005 .019 .004 .006 .009 .018
−.007 .006 .002 .003 .002 .001 -.001 .018
Mental Health Mild Depression Mild Anxiety Persistent Stress Depress./Anxiety Dx Mental Health Med. PHQ-8 GAD-7 PSS
CLS CLS CLS CLS CLS REG REG REG
Δ ROC Δ ROC Δ ROC Δ ROC Δ ROC Δ𝑟 Δ𝑟 Δ𝑟
.023 .027 .033 .021 .010 .056 .050 .112
.031 .032 .034 .023 .008 .068 .062 .079
.029 .030 .036 .023
.006 .062 .058 .078
.013 .009 .017 .011 .008 .031 .024 .038
.004 .001 .006 .004 .007 .009 .004 .014
Sleep Sleep Disorder Treatment Sleep Disturbance PRO Sleep Impairment PRO
CLS REG REG
Δ ROC Δ𝑟 Δ𝑟
.015 .043 .062
.017 .022 .072
.014 .019 .077
.005 .009 .046
.004 .003 .012
46
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.9 | Discriminative Task Performance Across a Sensor Foundation Model and Baselines. The table presents the performance of SensorFM-B, pretrained with the 5M data volume, compared against supervised baselines trained with engineered features and/or demographics on 35 discriminative tasks. SensorFM (either with or without demographics) achieves the best performance on 31 of 35 tasks. SensorFM variants are post-trained with PCA-50 reduced embeddings. For each task, we report the average 5-fold cross validation performance. Average Receiver Operating Characteristic Area Under the Curve (ROC AUC) is calculated in the logit-transform space and back-transformed. Average Pearson correlation (𝑟 ) is calculated in the z-transform space and back-transformed. Standard deviations are calculated in the transformed space and back-transformed to give asymmetric error values. Colors are normalized per row; best model performance is bolded and has the deepest shade.
Prediction Task
Type
Metric
Demos. Feat. Eng. SensorFM
Demos. Feat. Eng. SensorFM
Demos. Feat. Eng. SensorFM
Demos. Feat. Eng. SensorFM
Demos. Feat. Eng. SensorFM
Demographics Age BMI Height Weight
REG REG REG REG
𝑟 𝑟 𝑟 𝑟
-
.662+−..168 279 .441+−..158 191 .409+−..174 210 .460+−..011 012
-
.920+−..004 005 .809+−..007 007 .675+−..012 012 .809+−..007 007
-
Lifestyle Currently Working Disability Disability Affects Work Smoking Medicaid No Medications
CLS CLS CLS CLS CLS CLS
ROC ROC ROC ROC ROC ROC
.637+−..028 029 .600+−..031 032 .534+−..044 044 .629+−..049 051 .589+−..020 020 .617+−..018 018
.769+−..021 022 .702+−..023 024 .605+−..050 052 .754+−..046 052 .726+−..025 027 .687+−..006 006
.782+−..022 024 .722+−..027 028 .625+−..044 047 .771+−..037 042 .729+−..025 027 .699+−..008 008
.912+−..012 014 .753+−..020 021 .699+−..040 043 .870+−..060 098 .814+−..021 023 .739+−..009 009
.912+−..012 014 .758+−..022 024 .697+−..041 044 .867+−..060 097 .814+−..022 024 .744+−..009 009
Cardiovascular Cardiovascular Dx Hypertension Dx Respiratory Dx ASCVD Risk Framingham Risk Framingham 30 Risk
CLS CLS CLS REG REG REG
ROC ROC ROC
.701+−..037 039 .762+−..038 043 .640+−..049 052 .740+−..054 066
.696+−..090 107 .747+−..019 020 .642+−..039 041 .604+−..069 079 .548+−..072 081 .592+−..045 049
.706+−..091 111 .780+−..021 022 .640+−..031 033 .764+−..058 074 .730+−..063 078 .772+−..043 052
.712+−..052 058 .786+−..023 025 .682+−..054 059 .730+−..091 127 .669+−..101 133 .714+−..052 062
.696+−..057 064 .786+−..022 023 .674+−..046 050 .784+−..059 078 .724+−..057 069 .756+−..039 046
Metabolic Diabetes Dx Diabetes Med. Hyperlipidemia Pre-Diabetes Insulin Resistance HOMA-IR HbA1c Triglycerides
CLS CLS CLS CLS CLS REG REG REG
ROC ROC ROC ROC ROC
.688+−..052 058 .704+−..042 046
.734+−..062 073
.688+−..014 015 .731+−..031 033 .717+−..044 048 .316+−..068 071 .282+−..067 069 .251+−..042 043
.727+−..059 068 .672+−..064 071 .643+−..050 053 .728+−..052 059 .710+−..050 055 .374+−..069 073 .221+−..106 112 .131+−..065 066
.706+−..069 080 .670+−..041 044 .742+−..047 054 .717+−..040 044 .397+−..042 044 .228+−..090 094 .173+−..043 044
.763+−..037 042 .700+−..045 049 .674+−..032 033 .704+−..068 078 .761+−..044 050 .479+−..030 031 .293+−..029 030 .269+−..047 048
.756+−..043 049 .705+−..040 043 .676+−..027 029 .707+−..066 075
.763+−..044 051 .480+−..029 030 .292+−..026 026 .287+−..054 056
Mental Health Mild Depression Mild Anxiety Persistent Stress Depress./Anxiety Dx Mental Health Med. PHQ-8 GAD-7 PSS
CLS CLS CLS CLS CLS REG REG REG
ROC ROC ROC ROC ROC
.653+−..009 009 .641+−..009 009 .676+−..022 023 .626+−..012 012 .594+−..015 015 .303+−..018 018 .291+−..009 009 .378+−..019 019
.682+−..014 015 .654+−..010 011 .664+−..033 034 .672+−..022 022 .773+−..006 006 .354+−..018 018 .244+−..130 140 .305+−..124 135
.705+−..006 007 .680+−..005 006 .697+−..025 026 .693+−..015 015 .783+−..007 007 .410+−..009 009 .294+−..139 152 .417+−..040 041
.726+−..006 006 .698+−..005 005 .712+−..026 028 .717+−..009 009 .819+−..020 022 .450+−..019 019 .400+−..025 025 .463+−..013 014
.730+−..011 011 .699+−..006 006 .717+−..026 027 .721+−..008 008 .826+−..017 018 .459+−..023 024 .403+−..025 026 .476+−..018 019
Sleep Sleep Disorder Treatment Sleep Disturbance PRO Sleep Impairment PRO
CLS REG REG
ROC
.653+−..037 039 .195+−..028 028 .344+−..023 023
.610+−..024 024 .283+−..105 112 .310+−..138 152
.625+−..047 049 .326+−..056 059 .372+−..143 163
.649+−..036 038 .390+−..016 016 .448+−..013 013
.653+−..040 042 .393+−..014 014 .460+−..016 016
𝑟 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 𝑟 𝑟
𝑟 𝑟
.743+−..041 047 .782+−..038 045
47
Towards a General Intelligence and Interface for Wearable Health Data
Table ED.10 | Generative Performance Across Data Imputation, Interpolation, and Extrapolation Tasks. For Random Imputation, Temporal Interpolation / Extrapolation, and Signal Imputation we report Mean Squared Error (MSE) on known non-missing values. Presented error are 95% confidence intervals generated through 100 bootstrap iterations. Method
Generative Task Mean Fill
NN Fill
Linear Interp.
SensorFM-B
Random Imp. 80%
0.915 ± 0.002
1.020 ± 0.001
0.854 ± 0.002
0.215 ± 0.001
Temporal Interp. 20 min 60 min 180 min
0.876 ± 0.008 0.904 ± 0.005 0.950 ± 0.007
0.693 ± 0.008 0.943 ± 0.007 1.163 ± 0.008
0.561 ± 0.006 0.777 ± 0.007 0.961 ± 0.008
0.353 ± 0.002 0.468 ± 0.003 0.574 ± 0.002
Temporal Extrap. 20 min 60 min 180 min
0.923 ± 0.006 0.937 ± 0.007 0.974 ± 0.006
0.846 ± 0.010 1.102 ± 0.014 1.336 ± 0.010
0.846 ± 0.014 1.102 ± 0.008 1.336 ± 0.011
0.463 ± 0.004 0.563 ± 0.004 0.646 ± 0.003
Signal Imp. 2/26 6/26 12/26 20/26
1.016 ± 0.006 1.020 ± 0.005 1.025 ± 0.005 1.022 ± 0.002
1.016 ± 0.017 1.020 ± 0.006 1.025 ± 0.003 1.022 ± 0.003
1.016 ± 0.012 1.020 ± 0.008 1.025 ± 0.003 1.022 ± 0.003
0.122 ± 0.003 0.137 ± 0.002 0.170 ± 0.001 0.236 ± 0.001
Table ED.11 | Reconstructed Daily Sum-Aggregated Metrics. Comparison of Baseline, SensorFM Recovered, and Ground Truth under a simulated 1-hour data loss across all modalities. Baseline represents current standard for consumer wearable systems where no recovery methods are used and the aggregate is calculated solely over observed values. Note that SpO2 and Wrist Temp minutes do not sum to a full day due to inherent missingness in the original ground truth data. For a controlled comparison, aggregations are only computed over known valid minutes, which include the simulated loss. Presented error are 95% confidence intervals generated through 100 bootstrap iterations. Daily Average Value
Metric Baseline
SensorFM Recovered
Ground Truth
Activity (a.u.) Steps
5958.89
6208.41 ± 35.14
6227.49
Sleep Stages (Minutes) Light Deep REM
24.57 426.24 66.71
25.43 ± 0.12 444.12 ± 1.60 69.78 ± 0.33
25.64 444.48 69.51
Exercise (Minutes) Light (95≤HR<114) Aerobic (114≤HR<152) Anaerobic (152≤HR)
125.07 21.94 1.04
129.63 ± 1.04 22.38 ± 0.32 1.05 ± 0.08
130.64 22.92 1.07
SPO2 (Minutes) High (>90%) Low (<90%)
223.68 6.48
233.40 ± 1.20 6.62 ± 0.15
233.28 6.74
Wrist Temp (Minutes) Normal (<37◦ C) High (≥37◦ C)
1065.56 1.10
1112.05 ± 4.78 1.12 ± 0.16
1112.03 1.13
48
Towards a General Intelligence and Interface for Wearable Health Data
SensorFM-XXS
SensorFM-B
Figure ED.5 | Label Efficiency. By varying the percentage of data in training set we interogate the label efficiency of the models. SensorFM demonstrates good label efficient behavior.
49
Towards a General Intelligence and Interface for Wearable Health Data
C. Reconstruction Visualization Figure ED.6 presents 24-hour multimodal reconstruction heatmaps across multiple held-out validation subjects, illustrating how SensorFM fills in fragmented multimodal sensor segments through its generative pre-text task. Figures ED.7 and ED.8 zoom into two representative example days at persignal resolution (Heart Rate, Heart Rate Variability, Electrodermal Activity, Steps, Wrist Temperature, SpO2, and Sleep Stage REM) under two qualitatively different masking regimes: high-frequency fragmented signal loss (Example I) and a single multi-hour (∼10 h) block mask (Example II). The two examples show that the model leverages both local context (Example I) and long-context internal representations to maintain physiologically plausible baselines and circadian structure (Example II).
D. Analysis of the Model Embeddings in Latent Space SHapley Additive exPlanation Analysis. Figure ED.9 reports a SHAP-based latent feature attribution analysis aggregated from out-of-fold predictions across 5-fold cross-validation. Panel (a) is a chord diagram of pairwise cosine similarity between normalized SHAP attributions across downstream tasks (post-trained without demographic features), revealing which tasks share underlying embedding dimensions; only the top 30% of similarities are plotted per task to highlight the dominant latent relationships, and the outer ring encodes each task’s average pairwise similarity. Panel (b) plots feature attribution (averaged across non-demographic downstream tasks) for linear heads adapted on PCA-50 reduced embeddings combined with demographic features; embedding attribution rises from 82.7% at SensorFM-XXS to 87.3% at SensorFM-B while the reliance on demographic features correspondingly decreases. Embedding Space Analysis and Visualization. Figure ED.10 presents UMAP projections of SensorFM-B embeddings across 15 discriminative health outcomes (with the full downstream cohort plotted in light grey and continuous outcomes colored by deviation from the population median), giving a qualitative view of how cohorts cluster in the learned representation space. Figure ED.11 then characterizes the embedding geometry quantitatively across model scales: panel (a) shows kernel density estimates of pairwise Euclidean distances between user embeddings — latent-space dispersion is non-monotonic in scale, with S yielding the tightest clusters and B yielding the broadest spread, while XXS sits surprisingly close to B — and panel (b) plots cumulative explained variance versus principal component count, where smaller models saturate variance rapidly (suggestive of dimensional collapse) while SensorFM-B exhibits a dominant first PC capturing ∼40% of variance paired with a long tail of information distributed across higher dimensions.
50
Towards a General Intelligence and Interface for Wearable Health Data
Before
After Reconstruction via SensorFM
:00
:00
14
:00
12
:00
10
:00
08
:00
06
:00
04
:00
02
:00
00
:00
22
:00
20
:00 04
:00 02
:00 00
:00 22
:00 20
:00 18
:00 16
:00
:00
14
12
:00 10
:00
:00
:00 11
:00 09
:00 07
:00 05
:00 03
:00 01
:00 23
:00
:00
21
19
:00 17
:00
:00
:00 16
:00 14
:00 12
:00 10
:00 08
:00 06
:00 04
:00
:00
02
00
:00
:00
:00
22
Time of Day
:00 04
:00 02
:00
:00
00
Time of Day
22
:00 20
:00 18
:00 16
:00
:00
14
12
:00 10
:00 06
:00 04
:00 02
:00 00
:00
:00
Time of Day
22
20
18
16
14
12
10
08
:00
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
:00
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
08
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
18
:00 16
:00 14
:00 12
:00
:00
Time of Day
10
08
06
04
02
00
22
20
:00
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
:00
Altimeter Std Tonic Level Electrode Contact %
18
Time of Day
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
20
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
13
:00 11
:00 09
:00 07
:00
:00
Time of Day
05
03
:00 01
23
21
19
17
15
13
:00
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
:00
Time of Day
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
15
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
06
:00 04
:00 02
:00 00
:00
:00
Time of Day
22
20
:00 18
16
14
12
10
08
06
:00
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
06
Time of Day
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
08
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
16
:00
:00
14
:00
12
:00
10
:00
Time of Day
08
06
04
02
00
22
20
18
16
:00
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
Awake Light Sleep Deep Sleep REM Sleep
:00
Wrist Temp.
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Heart Rate HRV Median RR Int. HRV SDNN (5th-95th) HRV RMSSD (5th-95th) HRV PNN 20 HRV RR Coherence HRV Shannon Ent. HRV LF Power HRV HF Power HRV LF/HF Power HRV VLF Power HRV Spectral Entropy HRV % Good SpO2 SpO2 Confidence SpO2 Coverage
:00
Altimeter Std Tonic Level Electrode Contact %
:00
Altimeter Std Tonic Level Electrode Contact %
:00
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
18
Steps Jerk Log Energy Covariance Log Energy Ratio Zero Crossing Avg Axis Mean Axis Kurtosis Sleep Coefficient
Figure ED.6 | Generative Reconstruction of Multimodal Sensor Data. SensorFM model reconstructions of fragmented multimodal wearable sensor data. Each row represents one 24-hour sample from the pretraining validation dataset. The plot highlights how the model, through its generative pre-text task, internalizes structures in the data and enables the filling of missing segments with plasible, non-linear reconstructions. 51
Towards a General Intelligence and Interface for Wearable Health Data