2026-5-6
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment Joseph Breda1,† , Fadi Yousif1 , Beszel Hawkins1 , Marinela Cotoi1 , Miao Liu1 , Ray Luo1 , Po-Hsuan Cameron Chen1 , Mike Schaekermann1 , Samuel Schmidgall2 , Xin Liu1 , Girish Narayanswamy1 , Samuel Solomon1 , Maxwell A. Xu1 , Xiaoran Fan1 , Longfei Shangguan1 , Anran Wang1 , Bhavna Daryani1 , Buddy Herkenham1 , Cara Tan1 , Mark Malhotra1 , Shwetak Patel1 , John B. Hernandez1 , Quang Duong1 , Yun Liu1 , Zach Wasson1 , Dimitrios Antos1 , Bob Lou1 , Matthew Thompson1 , Jonathan Richina1 , Anupam Pathak1 , Nichole Young-Lin1 , Jake Sunshine1,‡ and Daniel McDuff1,‡
arXiv:2605.04012v1 [cs.AI] 5 May 2026
‡ Equal Leadership, 1 Google Research, 2 Google DeepMind, † Work done while at Google Research
Language models excel at diagnostic assessments on currated medical case-studies and vignettes, performing on par with, or better than, clinical professionals. However, existing studies focus on complex scenarios with rich context making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. We deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants ( 𝑁 = 13, 917) to interact with five AI agents. This corpus captures diverse communication and a realistic distribution of illnesses from a real world population. A subset of 1,228 participants reported a clinician-provided diagnosis, and 517 of these were further evaluated by a panel of clinicians during over 250 hours of annotation. SymptomAI DDx were significantly more accurate (𝑂𝑅 = 2.47, 𝑝 < 0.001) than those from independent clinicians given the same dialogue in a blinded randomized comparison. Moreover, agentic strategies which conduct a dedicated symptom interview that elicit additional symptom information before providing a diagnosis, perform substantially better than baseline, user-guided conversations ( 𝑝 < 0.001). An auxiliary analysis on 1,509 conversations from a general US population panel validated that these results generalize beyond wearable device users. We used SymptomAI diagnoses as labels for all 13,917 participants to analyze over 500,000 days of wearable metrics across nearly 400 unique conditions. We identified strong associations between acute infections and physiological shifts (e.g., 𝑂𝑅 > 7 for influenza). While limited by self-reported ground truth, these results demonstrate the benefits of a dedicated and complete symptom interview compared to a user-guided symptom discussion, which is the default of most consumer LLMs.
1. Introduction Consumer health information seeking patterns have undergone a global transformation in the 21st century with the rise of the Internet (Jia et al., 2021) and, more recently, the introduction of large language models (LMs) (Gallup). With up to one in five conversational AI queries relating to medical knowledge (Sumner et al., 2025), and millions of people using it for medical advice regularly (Shahsavar et al., 2023), AI is on track to becoming a primary interface that people approach for medical information needs (Ayers et al., 2023). Infact, a recent investigation into the types of personal guidance sought through one conversational AI platform found health and wellness to be the most popular topic, representing over a quarter of guidance-seeking conversations (Shen et al., 2026). Close to 20% of health-related AI chat conversations involve symptom assessment or condition discussion (Costa-Gomes et al., 2026). This trend toward self-guided medical assessment through technology predates conversational AI. Increases in search engine queries for symptoms predict decreases in outpatient visits for the same medical concerns (Heumann and Steinhubl, 2025). Consequently, a variety of online symptom checkers have emerged that enable patients to retrieve a set of possible
Corresponding author(s): joebreda, jakesunshine, [email protected] © 2026 Google. All rights reserved
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
Diagnosis Follow-up
Symptom Summary & DDx 0-2 weeks
Diagnosis From HPC
Symptom Interview
b Mobile Interface
a Study Flow: AI Interaction & Follow-up Urgent need to pee, lower abdomen pain and bloating, cloudy pee.
They started three days ago and is constant. It has been hard to sleep.
Thank you for sharing. could you tell me when these symptoms started?
Survey
A 3. Using the bathroom helps but only for a little bit.
How is the pain (1-10)? Is there anything that makes the symptoms worse?
Increased frequency but nothing else.
Neither.
Some possible causes of these symptoms are: 1. Urinary Tract Infection 2. Kidney Stones 3. Interstitial Cystitis
Do you have a family history of similar symptoms or taking any medication?
Have you experienced any fever, chills? Any burning sensation?
Blood Ox.
ns ia in ic
Clinical Annotation
e DDx Clinical Accuracy Validation
73%
Sy m
Activity
N=1,000
Cl
Electrodermal
60%
AI
Skin Temp
pt om
SymptomAI DDx & Auto-Rater Validation
Heart Rate/HRV
DDx Top-5 Accuracy
c Example Dialogue
& Auto-Rater Design
…
DDx
DDx
…
DDx
…
DDx
DDx …
DDx
DDx
…
DDx
Auto-DDx Generation
f Population Biomarker Analysis
Bronchitis
Respiration Rate
DDx
Influenza
…
d Real-World 9 Month Deployment
N=13, 917
…
N=13, 917
…
…
…
…
…
Resting Heart Rate
30 days wearable data pre-symptom report
Odds Ratio
Figure 1 | SymptomAI Study. (a-b) Experimental deployment study procedure of SymptomAI for endto-end patient interviewing and generative AI differential diagnosis (DDx) for symptom assessment that were benchmarked against study participant-reported diagnoses recieved from a Health Care Provider (HCP). (c-d) This led to a large dataset (N=13,917) of naturalistic symptom conversations communicated by laypeople paired with recent wearable data. (e) We leveraged clinical expert annotation to validate SymptomAI DDx against and to inform the development of an LLM verifier (i.e., auto-rater) for expanding validation beyond the clinical evaluation sub-sample. (f) Leveraging SymptomAI as a phenotype labeler enables phenome-wide analysis of biosignals across the study population. candidate diagnoses from a set of self-reported symptoms; however, traditional solutions are extremely limited with diagnostic accuracies ranging between 20-40% (Gilbert et al., 2023; Semigran et al., 2015; Wallace et al., 2022). This is significant as these initial symptom assessments often serve as a primary entry point for downstream medical care. Clinical history-taking (i.e., natural language exchange) alone is estimated to provide the basis for 7580% of diagnoses (Hampton et al., 1975; Peterson et al., 1992; Roshan and Rao, 2000), representing a 2
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
significant opportunity for impact through accurate diagnostic LMs. Harnessing this opportunity could have a significant impact on public health as access to clinical expertise is both episodic and globally scarce (Crisp and Chen, 2014). More broadly, improving access to medical reasoning expertise could positively impact the quality, accessibility, consistency, and affordability of medical intervention, both through empowering self-initiated assessment and through supplementing professional face-to-face guidance when available. Towards this, LMs have demonstrated performance on par with, or better than, clinical professionals on curated medical knowledge benchmarks (Arora et al., 2025; Bedi et al., 2025; Nori et al., 2023; Saab et al., 2024) and archival medical case studies (Kanjee et al., 2023; Manrai et al., 2026; McDuff et al., 2025). However, these text corpora used for evaluation have been limited to synthetic examples (Arora et al., 2025; Liu et al., 2026), single-turn medical question-answering (Nori et al., 2023; Saab et al., 2024; Singhal et al., 2023), or highly detailed reports of atypical, challenging medical scenarios (McDuff et al., 2025), none of which are truly representative of the type of information communicated conversationally by a patient in an everyday interaction through a digital interface nor the distribution of reported symptoms and illnesses across the population. Analysis of patient interviews simulated by trained patient actors has demonstrated improved performance in more realistic multi-turn scenarios, showing the potential of conversational AI for history-taking (Tu et al., 2025). While related work has explored conversational AI for specialized care (O’Sullivan et al., 2026; Palepu et al., 2025a), disease management (Palepu et al., 2025b), and multimodal reasoning (Saab et al., 2025), as well as frameworks for physician oversight (Vedadi et al., 2025), applying these systems to general symptom assessment by laypeople introduces unique challenges. An evaluation of diagnosis through human-AI interaction revealed that the involvement of laypeople in communicating necessary context significantly degraded AI diagnostic accuracy compared to AI applied directly to clinical vignettes (from 94.5% to 34.5%) (Bean et al., 2025; Goh et al., 2024). This decrease in performance is largely due to the incomplete or misrepresented information provided by nonexperts, indicating the criticality of testing diagnostic AI, when possible, with real users in naturalistic conditions. While recent work has begun to assess conversational AI involving laypeople with real health needs on subjective measures like perceived helpfulness in modest-scale studies (Sayres et al., 2026), a large-scale evaluation of conversational AI for layperson symptom assessment using clinical measures like diagnostic accuracy has not yet been demonstrated. To comprehensively understand the performance of conversational AI for providing accessible medical information to the broader population, in real-world contexts, it is important to evaluate its performance through an integrated study involving people with real health needs, assessing the ability of a general-purpose symptom checker to: (1) conduct flexible and personalized patient interviews to elicit appropriate context, (2) produce accurate DDx given context provided by general users, and (3) maintain performance across a range of real-world conditions. We conducted a strictly experimental research study of SymptomAI, a conversational AI agent, built on top of Gemini. This study-specific system was operationalized through the Fitbit Labs research environment in the Fitbit mobile application∗ from June 2025 to April 2026. The system was designed to explore the feasibility of both guiding participants through a series of symptom-related questions and providing a set of possible associated reasons with relevant educational information to research study participants (see Figure 1). The study was undertaken under informed consent (Advarra, Maryland USA: GH-SCD-001). Our investigation resulted in 13,917 multi-turn conversations in which research study participants voluntarily described their health symptoms to SymptomAI. During this investigation, we randomized the participants across five study arms, representing different agent prompting strategies ranging from highly structured history of present illness (HPI) interviews ∗ https://play.google.com/store/apps/details?id=com.fitbit.FitbitMobile&hl=en_US
3
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
based on canonical medical history taking questions (e.g., onset, location, symptom characterization, symptom provocation/palliation, symptom quality, symptom severity and associated symptoms) to a fully dynamic conversational agent, specifically a variant of the ‘Wayfinding AI’ system from (Sayres et al., 2026). Along with these conversations, we asked research study participants to report any diagnoses they received from healthcare provider interactions at the outset of, and two weeks after, their engagement with SymptomAI through an in-app survey. To further ground our evaluation, we conducted a human-expert annotation study where a panel of three board-certified Family Medicine physicians with >35 years of post-residency experience across primary care, urgent care, and academic medicine provided an independent DDx based on AI conversation transcripts. Each clinician then assessed the top-5 accuracy of DDx provided by both real clinicians and by SymptomAI side by side, while blinded to the author of the DDx and the participant’s actual diagnosis, for a subset of conversations. Clinicians were also requested to provide quality assessments of all study materials (conversation chat logs, participant-reported diagnoses, and all DDx lists). The exploratory associations generated by the system were not shared with the participants’ actual primary care providers and had no bearing on their clinical treatment plan. Building on the clinical expert validation, we employed an auto-rater to assess the DDx accuracy of SymptomAI across the entire cohort of participants who self-reported a diagnosis (N=1,228). To then leverage our full study cohort (N=13,917), we treated the top diagnosis generated by SymptomAI as a reference label for these research study participants, enabling a phenome-wide association study (PheWAS) (Bastarache et al., 2022) of nearly 400 unique medical diagnoses with over 500,000 days of wearable biosignals. This would have been infeasible to perform at comparable scale without the use of AI-generated reference diagnoses. We show that AI generated diagnoses from SymptomAI – particularly those of acute respiratory infections – share trends with wearable biosignals, potentially enabling future research analysis of wearable biosignals for predicting symptom onset. Such predictions could be used to trigger future SymptomAI conversations amongst users suspected of respiratory (and potentially non-respiratory) infectious diseases.
2. Results 2.1. Conversations That Elicit More Information Outperform User-Guided Conversations Participants were randomly assigned to one of five study arms, each employing a different prompting strategy. All prompting strategies that explicitly elicited more information from the user through follow-up questions outperformed the base user-guided condition (Fisher’s Exact Test - 𝑝 < 0.001). Arm 1: Base was only instructed to restrict responses to medical and health topics, reflecting the base performance of Gemini 2.0 Flash without specialized prompting to guide the conversation. Arm 2: Fixed canonical questions (Fixed Canonical) & Arm 3: Flexible canonical questions (Flexible Canonical) were both based on canonical medical history taking questions, representing "low agency" condition where SymptomAI was instructed to follow a structured interview procedure with a set of prescriptive questions. Arm 2 asked a fixed set of prescriptive questions irrespective of the users responses, while Arm 3 was allowed flexibility to drop irrelevant questions during the interview. Arm 4: Dynamic with live updates (Dynamic Live) & Arm 5: Dynamic with final output only (Dynamic Final) gave SymptomAI full agency over which follow up questions to ask and only restricted the number of turns before making a final DDx. Arm 4 provided intermediate best-effort DDx at every turn, while Arm 5 provided only a final DDx at the end of the conversation. For a full description of the prompting strategies, see Appendix E. Arms 2-5 explicitly elicited more information from the user, either through prescribed follow-up questions or enforcing multi-turn conversations of a minimum number of turns. This general strategy, a model-guided interview that elicits more information, resulted in an average of 27.34% higher accuracy 4
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
than the base, user-guided, prompting strategy which did not explicitly elicit more information and relies on users to provide or not provide information as part of the symptom checking experience. Moreover, prompting strategies that did not explicitly pre-define questions (arms 4 & 5; combined accuracy 71.4%) yielded comparable performance to strategies utilizing canonical, clinician-defined history-taking questions (arms 2 & 3; combined accuracy 75.6%), with no statistically significant difference observed via Welch’s t-test ( 𝑝 = 0.155). Figure 2c shows the accuracy stratified by arm, while Figure 2d shows the distribution of total user words provided in the interview. This speaks to SymptomAI’s ability to conduct effective patient history of present illness interviews autonomously based on natural conversational trajectory, yielding accuracy results similar to those obtained through asking common medical history questions. Conversely, the results also demonstrate that models adopting canonical history questions taught in medical training perform similarly to models not specifically constrained in what they can ask but tasked with arriving at an accurate differential diagnosis. 2.2. Clinical Experts Prefer SymptomAI Over Clinician-Generated Differential Diagnoses A cohort of 517 participants were selected for clinical evaluation according to the method described in Section A.4. All clinical evaluations were conducted as retrospective reviews of the conversation transcripts by expert raters who did not engage in any direct interaction or clinical care with the study participants. For each case, we generated three DDx per case, one from SymptomAI and two from independent clinicians ("baseline clinicians") reviewing the conversation transcripts. To ensure unbiased evaluations, another independent clinician ("clinical rater") ranked these DDx while blinded to both the DDx authors and the ground-truth diagnoses, relying exclusively on their clinical judgment. The clinical raters ranked the SymptomAI DDx as the best option in over 50% of the cases, indicating a significant preference over chance (odds ratio of 2.20, one-sided binomial test against expected 𝑃 = 0.33, 𝑛 = 517, 𝑝 < 0.001). This is also supported by a Cohen’s ℎ of 0.39, indicating a clear, consistent, small-to-medium effect size, in favor of SymptomAI. Figure 2a shows the proportion of 1st , 2nd and 3rd DDx across SymptomAI and baseline clinicians rated by the clinical raters. Supplemental Figure 11 shows the clinical raters’ preference for SymptomAI across conversation quality ratings. Importantly, this preference for SymptomAI is most pronounced in the subset of conversations clinician raters deem highest quality. Because these highest-quality conversations provide the most complete clinical context, they represent the most rigorous baseline for comparison. 2.3. SymptomAI DDx is More Accurate than Clinicians To quantitatively assess the accuracy of DDx, on the cohort of 517 cases, the clinical raters reviewed the DDx alongside the ground truth diagnosis, while blinded to the DDx author. SymptomAI demonstrated higher top-5 DDx accuracy over the baseline clinician’s DDx (McNemar’s Test: Median OR = 2.47, 95% CI Cohen’s 𝑔 [0.17, 0.25], p < 0.001). Figure 2b shows the average top-5 accuracy assigned by the clinical raters for SymptomAI and the baseline clinicians. Importantly, these results hold for the subset of conversations rated as highest quality by the clinical raters. Because these conversations provide complete clinical information, they offer the most rigorous and ideal baseline for comparison where the baseline clinicians have the full context for a complete DDx. Figure 2e shows the top-5 accuracy of SymptomAI and baseline clinicians’ DDx stratified across conversation quality, defined as whether the conversation contains sufficient context to produce an accurate DDx. 2.4. Robustness of SymptomAI on Low Information Conversations Figure 2f shows the top-5 accuracy for SymptomAI and clinician DDx for conversations stratified by the clinician’s confidence in their own DDx. While clinicians and SymptomAI performed similarly well on conversations where the clinicians felt confident in their own DDx, SymptomAI significantly 5
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
52.9% 36.7%
2nd Best
***
26.7% 39.9%
3rd Best 0.0
***
0.1
0.2
b
0.3
0.4
Clinician SymptomAI
***
20.4%
0.5
0.6
0.7
Proportion of Total Selections
0.8
d
Overall Top-5 Accuracy
Total User Word Count
DDx Source
Clinician
***
SymptomAI 0.0
0.4
0.6
0.8
1.0
Clinician SymptomAI
*
0.8
***
**
0.6 0.4 0.2 0.0
1 - Not Very Confident
2
3
4 - Neutral
5
6
7 - Very Confident
Clinician's Confidence in Conversation Quality
***
0.8
Flexible Canonical Dynamic Live
Dynamic Final ***
0.6 0.4 0.2 0.0
Clinician
DDx Source
SymptomAI
User Engagement by Prompt Strategy 300 250 200 150 100 50 0
Base
Fixed Canonical
Flexible Canonical
Dynamic Live
Dynamic Final
DDx Accuracy Across Clinician's Confidence in Their Own DDx
f ***
***
1.0
Base Fixed Canonical
350
1.0
Top-5 Accuracy (Mean ± SEM)
DDx Accuracy Across Clinician's Confidence in Conversation Quality
e Top-5 Accuracy (Mean ± SEM)
0.2
Accuracy by Prompt Strategy
Top-5 Accuracy (Mean ± SEM)
23.5%
1st Best
Rank Position
c
Clinician Ranking Distribution
Top-5 Accuracy (Mean ± SEM)
a
1.0
Clinician SymptomAI ***
0.8
***
0.6 0.4 0.2 0.0
1 - Not Very Confident
2
3 - Neutral
4
Clinician's Confidence in Clinician's DDx
5 - Very Confident
Figure 2 | Clinical evaluation and user engagement of SymptomAI. (a) Proportion of SymptomAI and clinician DDx (normalized by category total) ranked by blinded clinicians as 1st, 2nd, and 3rd position amongst a randomized list of 3 possible DDx lists for each conversation (one SymptomAI and two clinician baselines per trial). (b) The average top-5 accuracy assigned by clinicians to DDx produced by clinicians and SymptomAI. (c) Top-5 accuracy of clinicians and SymptomAI stratified by conversation prompting strategy. (d) Total user words sent across all user messages across each prompting strategy. Horizontal line denotes median word count. (e) The top-5 accuracy assigned by clinicians to SymptomAI and baseline clinician-generated DDx stratified by clinician’s confidence that the conversation contained enough information to support a plausibly accurate DDx. (f) The top-5 accuracy assigned by clinicians to SymptomAI and baseline clinician-generated DDx stratified by clinician’s rating of confidence in their own DDx for that conversation.
6
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
outperformed the clinicians for the conversations where the clinicians felt neutral or not confident in their own DDx. This indicates that SymptomAI is robust to factors that degrade human expert DDx performance and confidence. 2.5. Representativeness of the Study Population The Clinical Evaluation Sub-Sample is Representative of the Full Study Population. The sample of conversations used in the clinical evaluation were randomly sampled from the subset of the study population which had self-reported a diagnosis obtained from a healthcare provider. We assessed the representativeness of this clinical evaluation subsample relative to the total study population and found low effect size and significance, and therefore no practically meaningful distributional shifts in demographic covariates: age 𝐷 = 0.0537 ( 𝑝 = 0.0026), gender 𝑉 = 0.0640 (𝜒2 (1) = 70.37, 𝑝 < 0.001) , and weight 𝐷 = 0.0250 ( 𝑝 = 0.4641). This suggests that the subsample of research study participants who self-reported a diagnosis did not introduce detectable selection bias across demographics. We find only a small shift in diagnosis category Phecode: 𝑉 = 0.1670 (𝜒2 (390) = 848.36, 𝑝 < 0.001). A model stress test predicting whether participants reported a diagnosis resulted in poor discriminative ability ( 𝐴𝑈𝐶 = 0.65), suggesting further, through multivariate control, that the propensity to self-report was not strongly associated with user demographics or the category of illness the research study participants faced. SymptomAI Top-5 Accuracy Remains Consistent with the General Population. Sampling research study participants via Fitbit Labs in the Fitbit app raises questions about whether reported symptoms and DDx accuracy is similar to those expected from a general cross-section of the US population. We collected symptom assessment surveys and self-reported diagnoses from 1,509 people via a broad general population panel provider (Toluna). We then employed an auto-rater, validated on clinical labels through the approach outlined in Section B.2, to compare DDx performance across study populations. Despite capturing a significantly different distribution of illnesses from the SymptomAI study data 𝑉 = 0.3899 (𝜒2 (410) = 2738.82, 𝑝 < 0.001), we see similar performance of 75.2% top-5 accuracy of SymptomAI’s DDx on the auxiliary study population as compared to 80.0% top-5 on the SymptomAI study population, suggesting generalizability of SymptomAI diagnostic reasoning capabilities beyond our single study environment. 2.6. Diagnoses from SymptomAI Correlate with Physiological Biosignal Onset The cost of clinical labels can often prohibit population-scale analyses. By automating clinical-quality diagnoses, systems like SymptomAI open up large-scale analyses of physiological data which may otherwise be infeasible at scale. One such example is correlating wearable biosignals with different categories of illness derived from symptoms reported by laypeople through conversational data. SymptomAI May Enable Phenome Wide Association Studies of Certain Conditions. Figure 3 illustrates the illnesses linked to significant shifts across eight wearable biosignals, comparing affected patient cohorts for each diagnosis against the remaining study population without that condition. Most significant shifts appear for respiratory and circulatory illnesses, with most notably respiratory illnesses driven by biosignals present in the recent days leading up to SymptomAI engagement. Figure 4 shows the odds ratios for every biosignal across the illness cohorts that demonstrated at least one significant biosignal association. Physiological Biosignals Correlate with SymptomAI Engagement Onset. As highlighted in Figure 3, we observe strong associations between sensed wearable biosignals and acute respiratory infections. Figure 5 shows the relative change in wearable biosignals for a cohort of 1,546 participants diagnosed by SymptomAI with respiratory infection in the days leading up to their SymptomAI engagement. We observed distinct biosignal shifts, signaling symptom onset in the days leading up to users reporting their symptoms. Importantly, the cohort was defined by grouping SymptomAI’s Top-1 candidate 7
SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment
diagnoses into a respiratory infection category, explicitly excluding non-infectious conditions such as allergic rhinitis and chronic obstructive pulmonary disease. The correlation of wearable biosignal shift peaks aligning with the date of symptom reporting for these participants supports the feasibility of using physiological signals as precursors to digital health-seeking behavior and may enable SymptomAI to initiate future conversations for proactive triage based on these indicators. Historic Driven (Chronic) Resting Heart Rate (BPM)
8