ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance: A Randomized Clinical Trial.

Shoshani A et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cognitive-psychology
cognitive psychology

Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance: A Randomized Clinical Trial - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice JAMA Netw Open . 2026 Apr 14;9(4):e266713. doi: 10.1001/jamanetworkopen.2026.6713 Search in PMC Search in PubMed View in NLM Catalog Add to search Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance A Randomized Clinical Trial Anat Shoshani Anat Shoshani , PhD 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Anat Shoshani 1, ✉ , Bar Gurfinkel Bar Gurfinkel , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Bar Gurfinkel 1 , Ariel Kor Ariel Kor , PhD 2 Faculty of Medicine, Hebrew University of Jerusalem, Jerusalem, Israel 3 School of Medicine, Yale University, New Haven, Connecticut Find articles by Ariel Kor 2, 3 , Yael Ben-Haim Yael Ben-Haim , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Yael Ben-Haim 1 , Or Kanarek Or Kanarek , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Or Kanarek 1 , Romi Segev Romi Segev , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Romi Segev 1 , Or Shafir Or Shafir , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Or Shafir 1 , Romi Arbel Romi Arbel , MA 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel Find articles by Romi Arbel 1 Author information Article notes Copyright and License information 1 Baruch Ivcher School of Psychology, Reichman University, Herzliya, Israel 2 Faculty of Medicine, Hebrew University of Jerusalem, Jerusalem, Israel 3 School of Medicine, Yale University, New Haven, Connecticut Accepted for Publication: February 19, 2026. Published: April 14, 2026. doi: 10.1001/jamanetworkopen.2026.6713 Open Access: This is an open access article distributed under the terms of the CC-BY-NC-ND License , which does not permit alteration or commercial use, including those for text and data mining, AI training, and similar technologies. © 2026 Shoshani A et al. JAMA Network Open . ✉ Corresponding Author: Anat Shoshani, PhD, Baruch Ivcher School of Psychology, Reichman University, PO Box 167, Herzliya 46150, Israel ( [email protected] ). Author Contributions: Dr Shoshani had full access to all of the data in the study and takes responsibility for the integrity of the data and the accuracy of the data analysis. Concept and design: Shoshani, Gurfinkel, Kor, Kanarek, Shafir. Acquisition, analysis, or interpretation of data: All authors. Drafting of the manuscript: Shoshani, Gurfinkel, Kor, Ben-Haim, Kanarek, Shafir. Critical review of the manuscript for important intellectual content: All authors. Statistical analysis: Shoshani, Kor, Kanarek, Shafir. Administrative, technical, or material support: Gurfinkel, Ben-Haim, Segev, Arbel. Supervision: Shoshani, Kor, Kanarek, Shafir. Conflict of Interest Disclosures: Dr Shoshani reported receiving personal fees and holding stock options in KAI.AI Ltd outside the submitted work. Dr Gurfinkel reported receiving personal fees from and being employed by KAI.AI Ltd outside the submitted work. No other disclosures were reported. Data Sharing Statement: See Supplement 3 . ✉ Corresponding author. Received 2025 Nov 3; Accepted 2026 Feb 19; Collection date 2026 Apr. Copyright 2026 Shoshani A et al. JAMA Network Open . This is an open access article distributed under the terms of the CC-BY-NC-ND License, which does not permit alteration or commercial use, including those for text and data mining, AI training, and similar technologies. PMC Copyright notice PMCID: PMC13080544  PMID: 41979879 This randomized clinical trial assesses the efficacy of a conversational artificial intelligence (AI)–based platform for anxiety, depression, posttraumatic stress disorder, well-being, and life satisfaction among university students in Israel reporting psychological distress. Key Points Question Does a conversational artificial intelligence (AI)-based platform reduce psychiatric symptoms compared with face-to-face group therapy or a waiting list control? Findings In this randomized clinical trial of 995 university students, the AI intervention was associated with greater reduction in anxiety and improvement in well-being than both comparators, and greater reductions in depression and life satisfaction than the waiting list control; no significant differences were observed for posttraumatic stress disorder symptoms. Meaning The findings of this study suggest that conversational AI interventions may represent scalable adjuncts to mental health care. Abstract Importance Accessible, scalable interventions for psychiatric symptoms are needed to address global mental health care gaps. Conversational artificial intelligence (AI) may extend access by providing personalized support, yet rigorous evidence of efficacy and therapeutic mechanisms remains scarce. Objectives To evaluate the efficacy of a conversational AI–based mental health platform for psychiatric symptoms, and to assess how perceived therapeutic alliance contributes to user engagement and psychological outcomes. Design, Setting, and Participants This 3-arm randomized clinical trial was conducted in Israel from April 1 to October 27, 2025, including the 12-week intervention and 3-month follow-up period. Participants were university students who reported psychological distress. Interventions Participants were randomly assigned 1:1:1 to a 12-week AI-based conversational platform, face-to-face group therapy, or waiting list control. Main Outcomes and Measures The primary outcomes were anxiety (Generalized Anxiety Disorder-7), depression (Patient Health Questionnaire-9), posttraumatic stress disorder (PTSD; PTSD Checklist for DSM-5 ), well-being (World Health Organization-5 Well-Being Index), and life satisfaction (Brief Multidimensional Students’ Life Satisfaction Scale) measures at the end of the 12-week intervention. Analyses followed the intention-to-treat principle. Results In total, 995 participants (mean [SD] age, 23.1 [2.4] years; 504 [50.7%] female) were included in analyses, with 336 randomized to the AI intervention, 331 to group therapy, and 328 to waiting list control. After the intervention, participants in the AI group showed greater anxiety reduction than those in group therapy (mean difference [MD], −2.17 [95% CI, −2.67 to −1.67]) or control (MD, −2.15 [95% CI, −2.65 to −1.65]) and greater depression reduction than control (MD, −1.99 [95% CI, −2.63 to −1.35]). PTSD symptoms did not differ among groups. The AI group also showed greater improvements in well-being than group therapy (MD, 5.72 [95% CI, 2.71 to 8.73]) and control (MD, 9.16 [95% CI, 6.14 to 12.18]). Structural equation modeling indicated that perceived therapeutic alliance was associated with engagement (β = 0.31 [95% CI, 0.16 to 0.43]; P < .001) and symptom improvement (β = –0.58 [95% CI, –0.69 to –0.46]; P < .001). Conclusions and Relevance In this randomized clinical trial of university students with psychological distress, the use of a conversational AI agent was associated with improvements in anxiety, depression, well-being, and life satisfaction, and its perceived therapeutic alliance was associated with engagement and psychological improvement. These findings suggest that conversational AI may serve as a scalable resource within mental health frameworks. Trial Registration ISRCTN Registry Identifier: ISRCTN61075527 Introduction The global mental health crisis is urgent and escalating. About 1 in 8 people worldwide is currently affected by a mental disorder, yet only a quarter of those who need care receive it. 1 This treatment gap reflects structural constraints in conventional 1-to-1 models of care, including a shortage of trained professionals, unequal distribution of services, and persistent stigma around help-seeking. 2 Addressing these barriers requires scalable, accessible, and high-quality interventions that can support large populations without placing additional strain on already burdened systems. Digital technologies have become a promising avenue to meet these needs. Digital mental health interventions, ranging from self-guided applications to integrated platforms that incorporate artificial intelligence (AI), can remove geographic and logistical barriers to care. 3 Despite this promise, many commercially available apps struggle with sustained engagement. Passive or impersonal formats often lead to rapid attrition, with some studies reporting fewer than 4% of users remaining active after the first month. 4 Attrition is frequently attributed to static and nonrelational designs that do not reproduce the empathic and adaptive qualities associated with effective therapy. 5 To address these shortcomings, attention has turned to interactive interventions and, in particular, to conversational agents. Unlike general purpose chatbots, newer AI companions are designed to simulate natural dialogue and to cultivate a sense of connection through empathic exchanges and personalized support. 6 Advances in large language models and natural language processing have improved contextual sensitivity and the ability to approximate key ingredients of therapeutic dialogue, including empathy, responsiveness, and personalization. 7 Consistent with the computers as social actors framework, users tend to respond to these systems as relational partners, which may support disclosure of sensitive concerns. 6 , 8 The empirical base for conversational agents is expanding. Meta-analytic evidence indicates small to moderate effects on mental health outcomes, with benefits observed most consistently for depressive symptoms and more variable effects for anxiety and well-being. 9 , 10 A recent randomized clinical trial of a generative AI system reported symptom reductions in depression and anxiety compared with a waiting list control in a clinical sample over a short-term intervention period. 11 However, much of the evidence comes from smaller or short-term studies, 9 , 10 often outside clinical practice settings. There is limited knowledge about the durability of effects, the generalizability of outcomes, and the role of relational processes such as therapeutic alliance in shaping engagement and clinical change. The present randomized clinical trial addressed these gaps by evaluating the efficacy of a conversational AI platform for emotional support, compared with face-to-face group therapy and a waiting list control among university students reporting psychological distress. Primary outcomes included symptoms of anxiety, depression, and posttraumatic stress disorder (PTSD) alongside well-being and life satisfaction. The secondary outcome was intention to seek therapy. Additionally, the study examined whether perceived therapeutic alliance with the AI platform was associated with user engagement and psychological outcomes. We hypothesized that the AI intervention would be associated with greater symptom reduction than the waiting list control and would demonstrate comparable or superior outcomes to face-to-face group therapy across selected measures. Finally, we hypothesized that a stronger perceived alliance would be associated with increased engagement and improved psychological outcomes. Methods Trial Design and Participants This 3-arm, parallel-group randomized clinical trial was conducted in Israel between April 1 and October 27, 2025. The trial followed the Consolidated Standards of Reporting Trials ( CONSORT ) reporting guideline, was registered with ISRCTN, and was approved by the Reichman University institutional ethics committee. Recruitment, screening, consent, and baseline assessments occurred remotely on a rolling basis between February 1 and March 31, 2025. Participants were recruited through university networks and online platforms via targeted advertisements seeking students experiencing current emotional distress. Interested individuals accessed a secure link to complete screening and provide electronic written informed consent. The full trial protocol is provided in Supplement 1 . Eligibility was assessed via the 4-item Patient Health Questionnaire (PHQ-4) and a clinical safety protocol. Scores on the PHQ-4 range from 0 to 12, with higher scores indicating greater psychological distress. Inclusion criteria were university students (18-35 years of age) reporting current psychological distress, defined as a score of 3 or higher on either the anxiety or depression subscales of the PHQ-4, Hebrew fluency, and internet access. Exclusion criteria included active suicidal ideation or psychiatric crisis, current psychotherapy or psychiatric medication, and a history of severe mental disorders. Individuals indicating acute risk during screening were excluded and directed to crisis resources. Participants received academic credit for completing the assessments, independent of their level of engagement with the intervention. Randomization and Masking Participants were assigned to 1 of the 3 groups (1:1:1) using computer-generated variable blocks stratified by sex. Allocation was concealed through automated assignment. Masking was not feasible for participants or therapists due to the nature of the interventions. Procedures and Interventions The 12-week intervention period commenced for participants between April 1 and April 25, 2025, depending on group scheduling and onboarding requirements. Assessments occurred at baseline, immediately following the completion of each participant’s 12-week intervention period, and at the 3-month follow-up. To maintain rigorous temporal equivalence, each participant’s assessment schedule was anchored to their specific start date, ensuring a uniform duration of exposure across the AI, face-to-face, and waiting list control conditions. AI Intervention Participants interacted with Kai, a mobile and website platform (KAI.AI Inc) delivering tailored psychological support through natural language conversations in familiar messaging environments. 12 , 13 Onboarding included assessments of emotional state, coping challenges, and preferences that informed a personalized plan drawing from cognitive behavioral therapy, acceptance and commitment therapy, dialectical behavior therapy, mindfulness, and positive psychology. Daily exchanges included reflective prompts, stress-regulation and emotion-regulation exercises, breathing practices, and motivational messages to reinforce coping and self-awareness (eTable 1 in Supplement 2 ). The platform integrated large language models with adaptive memory and user profiling to maintain continuity across sessions. Participants had unrestricted access and were encouraged to engage at least 3 times weekly. A multilayered safety framework monitored language for acute risk, providing automated crisis resources and enabling licensed clinicians to intervene in urgent cases. Representative screenshots are shown in eFigure 1 in Supplement 2 . Face-to-Face Group Therapy Participants attended 12 weekly 90-minute sessions facilitated by licensed psychologists. Groups of approximately 20 participants followed a semistructured format blending manualized elements with guided discussion. Sessions covered psychoeducation on depression, anxiety, and trauma-related stress reactions, alongside cognitive restructuring, behavioral activation, acceptance-based strategies for difficult emotions, mindfulness practices, relaxation and breathing techniques, and reflective writing (eTable 2 in Supplement 2 ). The approach matched the theoretical foundations of the AI intervention while allowing for flexibility for group process and participant needs. Waiting List Control Participants received no active intervention during the study period. However, they were offered access to the intervention after completion of the 3-month follow-up. Measures Primary outcomes were symptoms of anxiety, depression, PTSD, well-being, and life satisfaction, collected via secure online questionnaires at baseline, at completion of the 12-week intervention, and at the 3-month follow-up. Anxiety was assessed with the Generalized Anxiety Disorder-7 14 (scores range from 0 to 21, with higher scores indicating greater anxiety severity) and depression with the Patient Health Questionnaire-9 15 (scores range from 0 to 27, with higher scores indicating greater depression severity); both scales use 4-point Likert items assessing symptoms over the past 2 weeks. PTSD symptoms were measured with the brief PTSD Checklist for DSM-5 , 16 with items rated 0 to 4 for distress severity during the previous month; scores range from 0 to 16, with higher scores indicating greater PTSD symptom severity. Life satisfaction was assessed with the Brief Multidimensional Students’ Life Satisfaction Scale, 17 evaluating satisfaction across family, friends, school, self, and environment domains (item scores range from 1 [very dissatisfied] to 7 [very satisfied]; total scores range from 5 to 35, with higher scores indicating greater life satisfaction). The World Health Organization–5 Well-Being Index 18 measured positive mood and vitality (0 [at no time] to 5 [all of the time]), with scores converted to a 0 to 100 scale; higher values denote greater well-being. Internal consistency for these measures was high (α = .87-.91). The secondary outcome was intention to seek therapy, evaluated with 1 item from the General Help-Seeking Questionnaire 19 and a yes or no question regarding plans to begin therapy within 3 months. Process measures included engagement, defined as the total number of user-initiated messages sent to the AI platform during the 12-week period. To distinguish active engagement from passive responses, this metric excluded automated button clicks and reflected only unique text-based messages authored by the participant. Engagement was further summarized as the number of active days and interaction episodes per week. Perceived therapeutic alliance was assessed with 15 items adapted from the Counselor Rating Scale, 20 rating warmth, empathy, and professionalism on a 5-point scale (α = .93). Statistical Analysis Analyses followed the intention-to-treat principle using SPSS, version 28.0 (IBM). Data from all participants were analyzed in their randomized groups with no crossover. Models included age (continuous) and sex as a priori covariates to account for confounding in young adult psychological outcomes. Repeated-measures multivariate analysis of covariance (MANCOVA) was used to evaluate condition by time effects from baseline to the 12-week primary end point and 3-month follow-up. Significant omnibus interactions were followed by prespecified pairwise contrasts and Bonferroni-adjusted univariate analyses. Missing data patterns were evaluated by comparing baseline characteristics of participants with complete vs incomplete data (eTable 3 in Supplement 2 ). Missing data were addressed via multiple imputation with chained equations (20 datasets) (eMethods 2 in Supplement 2 ). Power analysis (G*Power, version 3.1.9.7) determined that a sample size of 995 provided 80% power (α = .05, 2-sided) to detect a small effect (Cohen f = 0.15) with medium correlations between measures ( r = 0.50) (eMethods 1 in Supplement 2 ), accounting for attrition (eMethods 2 in Supplement 2 ). Negligible intraclass correlation coefficients (0.01-0.02) justified nonclustered models (eMethods 3 in Supplement 2 ). Finally, path modeling in AMOS, version 28.0, was used to examine associations among therapeutic alliance, engagement, and psychological improvement. A 2-sided P < .05 was considered statistically significant. Results Study Participants Of 1050 screened individuals, 995 participants (mean [SD] age, 23.1 [2.4] years; 504 [50.7%] female and 491 [49.3%] male) were randomized to the AI intervention (n = 336), face-to-face group therapy (n = 331), or waiting list control (n = 328). The complete participant flow is presented in the diagram in Figure 1 . At baseline, mean (SD) anxiety scores ranged from 7.27 (4.66) to 7.51 (4.62), and depression scores from 7.63 (4.79) to 8.07 (4.76), consistent with mild to moderate distress relative to population norms. 14 , 15 All continuous variables fell within acceptable ranges for skewness and kurtosis, and demographic characteristics were balanced across groups. Demographic information was self-reported and is summarized in Table 1 . Figure 1. CONSORT Flow Diagram. Open in a new tab AI indicates artificial intelligence; CONSORT, Consolidated Standards of Reporting Trials. Table 1. Sample Demographic and Clinical Characteristics at Baseline. Characteristic Participants, No. (%) AI platform (n = 336) Face-to-face group therapy (n = 331) Control (n = 328) Age, mean (SD), y 22.9 (2.1) 23.2 (2.6) 23.2 (2.6) Sex Female 185 (55.1) 155 (46.8) 164 (50.0) Male 151 (44.9) 176 (53.2) 164 (50.0) Socioeconomic status High 81 (24.1) 60 (18.1) 57 (17.4) Middle 208 (61.9) 206 (62.2) 210 (64.0) Low 47 (14.0) 65 (19.6) 61 (18.6) Religion Christian 1 (0.3) 3 (0.9) 2 (0.6) Jewish 329 (98.2) 322 (97.3) 321 (97.9) Muslim 5 (1.5) 6 (1.8) 5 (1.5) Marital status Married 7 (2.1) 6 (1.8) 5 (1.5) Single 208 (62.1) 210 (63.4) 214 (65.2) Divorced 1 (0.3) 2 (0.6) 0 In relationship (not married) 119 (35.5) 113 (34.1) 109 (33.2) Life satisfaction, mean (SD) score a 27.90 (4.46) 27.85 (4.87) 27.77 (4.51) Well-being b 54.93 (16.59) 52.38 (13.07) 53.04 (12.49) Anxiety c 7.27 (4.66) 7.39 (4.90) 7.51 (5.05) Depression d 7.64 (4.79) 8.07 (5.32) 7.90 (5.38) PTSD e 4.26 (3.69) 4.36 (4.52) 4.56 (4.47) Open in a new tab Abbreviations: AI, artificial intelligence; PTSD, posttraumatic stress disorder. a Assessed with the Brief Multidimensional Students’ Life Satisfaction Scale; items scored 1 (very dissatisfied) to 7 (very satisfied), with total scores ranging from 5 to 35 and higher scores indicating greater life satisfaction. b Measured with The World Health Organization-5 Well-Being Index, scoring positive mood and vitality (0 [at no time] to 5 [all of the time]), with scores converted to a 0 to 100 scale. Higher values denote greater well-being. c Measured with the Generalized Anxiety Disorder-7; scores range from 0 to 21, with higher values indicating greater anxiety severity. d Measured with the Patient Health Questionnaire-9; scores range from 0 to 27, with higher scores indicating greater depression severity. e Measured with the brief PTSD Checklist for DSM-5 ; scores range from 0 to 16, with higher scores indicating greater PTSD symptom severity. Mental Health and Well-Being Outcomes A 3 × 2 repeated-measures MANCOVA, controlling for age and sex, revealed a significant condition by time interaction ( Table 2 ). Follow-up univariate tests used a Bonferroni-adjusted α of .01. Differences reported here are from the 12-week measurement. Table 2. Means, Standard Deviations, and Effect Sizes Outcomes by Intervention. Outcome and Group Score, mean (SD) F (Time × condition, T1-T2) Adjusted mean difference (95% CI) T1 a T2 a T3 a T2 AI vs face-to-face T2 AI vs control T3 AI vs face-to-face T3 AI vs control Anxiety b AI platform 7.27 (4.66) 6.28 (4.37) 6.16 (4.29) 73.00 c −2.17 (−2.67 to −1.67) −2.15 (−2.65 to −1.65) −1.65 (−2.42 to −0.87) −2.08 (−2.85 to −1.30) Face-to-face 7.39 (4.90) 8.57 (5.20) 8.47 (4.66) NA NA NA NA Control 7.51 (5.05) 8.67 (5.31) 8.83 (5.15) NA NA NA NA Depression d AI platform 7.64 (4.79) 6.68 (5.12) 6.52 (5.08) 27.89 c −0.68 (−1.32 to −0.04) −1.99 (−2.63 to −1.35) −0.85 (−1.69 to −0.01) −1.79 (−2.82 to −0.76) Face-to-face 8.07 (5.32) 7.80 (5.33) 7.98 (5.07) NA NA NA NA Control 7.90 (5.38) 8.94 (5.60) 8.76 (4.99) NA NA NA NA PTSD e AI platform 4.26 (3.69) 3.49 (3.59) 3.41 (3.12) 0.72 −0.29 (−0.85 to 0.27) −0.49 (−1.05 to 0.07) 0.15 (−0.40 to 0.70) −0.27 (−0.82 to 0.28) Face-to-face 4.36 (4.52) 3.79 (4.20) 3.64 (3.42) NA NA NA NA Control 4.56 (4.47) 4.11 (4.30) 4.02 (3.34) NA NA NA NA Life satisfaction f AI platform 27.90 (4.46) 29.01 (4.22) 28.78 (5.71) 21.33 c 1.19 (0.23 to 2.14) 2.58 (1.62 to 3.54) 2.17 (1.05 to 3.29) 2.79 (1.67 to 3.91) Face-to-face 27.85 (4.87) 27.78 (4.52) 26.70 (5.74) NA NA NA NA Control 27.77 (4.51) 26.31 (4.37) 26.03 (6.01) NA NA NA NA Well-being g AI platform 54.93 (16.59) 61.49 (17.43) 60.82 (18.98) 26.65 c 5.72 (2.71 to 8.73) 9.16 (6.14 to 12.18) 6.08 (3.07 to 9.10) 10.13 (7.11 to 13.15) Face-to-face 52.38 (13.07) 53.22 (12.35) 53.08 (15.32) NA NA NA NA Control 53.04 (12.49) 50.43 (12.76) 49.47 (16.00) NA NA NA NA Intention to seek therapy, No. (%) AI platform 100 (29.8) 79 (23.5) 53 (24.4) NA NA NA NA NA Face-to-face 105 (31.7) 102 (30.8) 73 (34.1) NA NA NA NA Control 111 (33.9) 121 (37.0) 80 (38.6) NA NA NA NA Open in a new tab Abbreviations: AI, artificial intelligence; NA, not applicable; PTSD, posttraumatic stress disorder; T1, baseline; T2, 12 weeks after the intervention; T3, 3-month follow-up. a Sample sizes at T1 and T2: AI platform, n = 336; face-to-face, n = 331; control, and n = 328. Sample sizes at T3: AI platform, n = 217; face-to-face, n = 214; and control, n = 208. b Measured with the Generalized Anxiety Disorder-7; scores range from 0 to 21, with higher values indicating greater anxiety severity. c P < .001. d Measured with the Patient Health Questionnaire-9; scores range from 0 to 27, with higher scores indicating greater depression severity. e Measured with the brief PTSD Checklist for DSM-5 ; scores range from 0 to 16, with higher scores indicating greater PTSD symptom severity. f Assessed with the Brief Multidimensional Students’ Life Satisfaction Scale; items scored 1 (very dissatisfied) to 7 (very satisfied), with total scores ranging from 5 to 35 and higher scores indicating greater life satisfaction. g Measured with The World Health Organization-5 Well-Being Index, scoring positive mood and vitality (0 [at no time] to 5 [all of the time]), with scores converted to a 0 to 100 scale. Higher values denote greater well-being. For anxiety, the AI group showed greater symptom reduction compared with both face-to-face group therapy (mean difference [MD], –2.17 [95% CI, –2.67 to –1.67]; P < .001) and control (MD, –2.15 [95% CI, –2.65 to –1.65]; P < .001). No significant difference was observed between group therapy and control (MD, −0.05 [95% CI, −0.76 to 0.66]; P = .89). For depression, the AI group (MD, –1.99 [95% CI, –2.63 to –1.35]; P < .001) and group therapy (MD, –1.31 [95% CI, –1.95 to –0.66]; P < .001) both showed greater reductions in symptoms than control. However, the difference between AI and group therapy (MD, –0.68 [95% CI, –1.32 to –0.04]; P = .03) was no longer statistically significant after adjustment for multiple comparisons. Regarding PTSD symptoms, no significant condition by time interaction was observed. Pairwise comparisons after the intervention showed no significant differences between the AI group and face-to-face therapy (MD, –0.29 [95% CI, –0.85 to 0.27]; P = .31), between AI and control (MD, –0.49 [95% CI, –1.05 to 0.07]; P = .09), or between face-to-face therapy and control (MD, –0.20 [95% CI, –0.76 to 0.36]; P = .48). For positive functioning, well-being scores were higher in the AI group than in group therapy (MD, 5.72 [95% CI, 2.71 to 8.73]; P < .001) and control (MD, 9.16 [95% CI, 6.14 to 12.18]; P < .001). Group therapy did not differ significantly from control (MD, 1.01 [95% CI, –0.79 to 2.80; P = .27). Life satisfaction scores were higher in the AI group than in control (MD, 2.58 [95% CI, 1.62 to 3.54]; P < .001) and group therapy (MD, 1.19 [95% CI, 0.23 to 2.14]; P = .009), and group therapy also showed higher scores than control (MD, 0.77 [95% CI, 0.21 to 1.33]; P = .007). Intention to Seek Therapy and Perceived Support While baseline therapy intentions were uniform across groups, differences emerged after the intervention: intention to seek therapy declined in the AI group (100 [29.8%] to 79 [23.5%]), remained stable in group therapy (105 [31.7%] to 102 [30.8%]), and increased in the control group (111 [33.9%] to 121 [37.0%]) ( P < .001 for all) (eTable 4 and eFigure 2 in Supplement 2 ). Perceived supportive qualities were similar across conditions; mean (SD) warmth scores were 3.86 (0.79) for AI vs 3.74 (0.99) ( P = .08) for human therapists, and professionalism scores were 3.85 (0.85) for AI vs 3.93 (0.94) for human therapists ( P = .25). User Engagement and Psychological Outcomes Participants sent a mean (SD) of 18.6 (12.4) messages per week, with engagement occurring across a mean of 3.0 active days per week; 205 of 336 participants (61.0%) remained active through week 12. This interaction level reflected substantial textual engagement, with a mean (SD) of 1645 (1348) words per week and mean (SD) of 83.1 (67.9) minutes of active platform engagement. Message frequency correlated with improvements in anxiety ( r = –0.46), depression ( r = –0.33), and well-being ( r = 0.21) (all P < .001). Structural equation modeling ( Figure 2 ) indicated good model fit (standardized root mean square residual, 0.031; comparative fit index, 0.994; and root mean square error of approximation, 0.024); therapeutic alliance was associated with engagement (β = 0.31 [95% CI, 0.16 to 0.43]; P < .001), which was associated with symptom change (β = –0.58 [95% CI, –0.69 to –0.46]; P < .001). The indirect path from alliance through engagement to symptom change was significant (β = –0.18 [95% CI, –0.27 to –0.10]; P < .001). Figure 2. Schematic Depicting Structural Equation Model Linking Therapeutic Alliance, Engagement, Symptoms, and Therapy Need in the Artificial Intelligence Group. Open in a new tab Standardized path coefficients are shown along arrows. Values above indicators represent factor loadings for latent variables. R 2 values indicate the proportion of variance explained in each endogenous construct. Mental health symptoms were modeled as a latent factor indicated by anxiety, depression, posttraumatic stress disorder (PTSD), and well-being. Life satisfaction was not retained in the latent construct because of low factor loading. Therapeutic alliance was modeled as a latent factor indicated by perceived professional competence and warmth. a P < .05. b P < .001. Follow-Up Outcomes At 3 months, attrition was similar across conditions (AI, 119 of 336 [35.4%]; group therapy, 117 of 331 [35.3%]; control, 120 of 328 [36.6%]). No significant differences emerged between participants who completed vs did not complete the intervention regarding baseline demographics, clinical measures, or treatment intention, suggesting no systematic attrition bias (eTable 3 in Supplement 2 ). A 1-way MANCOVA indicated a significant overall effect of condition. Anxiety scores remained lower in the AI group than in group therapy (MD, –1.65 [95% CI, –2.42 to –0.87]; P < .001) and control (MD, –2.08 [95% CI, –2.85 to –1.30]; P < .001). For depression, scores in the AI group remained lower than control (MD, –1.79 [95% CI, –2.82 to –0.76]; P < .001), but differences between AI and group therapy (MD, –0.85 [95% CI, –1.69 to –0.01]; P = .04) and between group therapy and control (MD, –0.94 [95% CI, –1.78 to –0.10]; P = .03) were not significant after Bonferroni adjustment. For life satisfaction, the AI group reported higher scores than group therapy (MD, 2.17 [95% CI, 1.05-3.29]; P < .001) and control (MD, 2.79 [95% CI, 1.67-3.91]; P < .001); group therapy did not differ significantly from control (MD, 0.62 [95% CI, –0.50 to 1.74]; P = .28). For well-being, the AI group reported higher scores than group therapy (MD, 6.08 [95% CI, 3.07-9.10]; P < .001) and control (MD, 10.13 [95% CI, 7.11-13.15]; P < .001); group therapy also reported higher well-being scores than control (MD, 4.05 [95% CI, 1.03-7.06]; P = .009). No significant between-group differences were observed for PTSD symptoms at 3 months: AI versus group therapy (MD, 0.15 [95% CI, –0.40 to 0.70]; P = .59), AI versus control (MD, –0.27 [95% CI, –0.82 to 0.28]; P = .33), and group therapy versus control (MD, –0.42 [95% CI, –0.97 to 0.12]; P = .13). Clinical Transitions Among participants meeting baseline clinical cutoffs (score ≥10), transition rates to the nonclinical range for anxiety after the intervention were 57.9% (55 of 95) in the AI group, 14.4% (14 of 97) in group therapy, and 9.8% (11 of 112) in control ( P < .001). For depression, transition rates were 47.0% (47 of 100) in the AI group, 26.8% (30 of 112) in group therapy, and 14.0% (18 of 129) in control ( P < .001) (eTable 5 in Supplement 2 ). At follow-up, group differences were maintained for anxiety but not for depression. No significant differences in transition rates were observed for PTSD symptoms at either time point. Safety and Adverse Events No serious adverse events were observed. Minor transient increases in emotional distress were reported by 1.2% (n = 4) of the AI group and 1.5% (n = 5) of the face-to-face group therapy arm. No participants discontinued the study due to adverse events, and no adverse events were reported in the waiting list control group. Discussion This randomized clinical trial indicated that an AI-based emotional support platform was associated with significant, albeit modest, improvements across a range of psychological outcomes. Participants using the platform reported gains in well-being and life satisfaction alongside reductions in anxiety and depression compared with participants in a waiting list control group. Notably, the AI intervention yielded favorable results in domains in which traditional face-to-face group therapy did not differ significantly from the control, such as anxiety. However, these benefits did not extend to PTSD symptoms, suggesting that the impact of the intervention may be specific to general distress and positive functioning rather than trauma-related symptomatology. The relative efficacy of the AI platform may stem from structural advantages over traditional formats. Unlike weekly sessions, the AI platform offered continuous, on-demand access, facilitating real-time reinforcement. 21 This accessibility is crucial, as traditional group therapy showed no significant benefits over control for anxiety. While symptoms slightly escalated in both the therapy and control arms, the AI intervention effectively mitigated this pattern. Furthermore, the system’s personalization and memory may have fostered an individualized therapeutic alliance, whereas group facilitators must divide attention among multiple participants. 22 Participants may also have experienced greater openness with the AI platform due to its nonjudgmental nature; the online disinhibition effect suggests individuals often disclose more sensitive information to digital systems than to human practitioners. 23 In contrast, group dynamics and fear of negative evaluation can restrict disclosure. After the intervention, participants rated the supportive qualities comparably between the AI platform and human therapists. In the AI condition, these qualities were associated with higher engagement and subsequent symptom improvement. Perceived warmth and competence were associated with engagement, which in turn was associated with reductions in anxiety and depression and gains in well-being. This pathway aligns with the concept of a digital therapeutic alliance, 24 suggesting relational processes central to psychotherapy can develop in AI interactions. The decline in therapy-seeking intention among AI participants likely reflects reduced perceived need after improvement but warrants cautious interpretation. While potentially indicating greater self-efficacy, it could signal unmet needs if individuals rely solely on digital support without escalating to human care when necessary. The absence of effects on PTSD highlights an important limitation; trauma-related symptoms often require specialized exposure-based approaches 25 not included in this platform. This underscores the need for trauma-specific modules and referral mechanisms within digital interventions. Limitations This study has limitations. Outcomes were self-reported rather than clinician-rated, limiting external validation. Reduced therapy intention reflects perceived need rather than actual service use, leaving the impact on health system demand unclear. Although effects persisted at follow-up, long-term cost-effectiveness and functional recovery remain unknown. Finally, attrition at 3 months was substantial, potentially affecting longer-term estimates. Conclusions In this randomized clinical trial of university students with psychological distress, the use of an AI-based emotional support platform was associated with modest gains in well-being and life satisfaction alongside reductions in anxiety and depression. Participants perceived the platform as possessing relational qualities typically associated with human therapists, including warmth and competence. These perceptions were associated with both participant engagement and clinical symptom change, demonstrating that fostering a digital therapeutic alliance is feasible and may be central to the clinical efficacy of automated interventions. These findings suggest that conversational AI may serve as a scalable resource within mental health frameworks, although its role is likely best suited as an adjunct or early intervention tool. Further research is required to evaluate the responsible integration of such systems into existing care models, with particular emphasis on scalability, equity, safety, and the maintenance of rigorous clinical oversight. Supplement 1. Trial Protocol jamanetwopen-e266713-s001.pdf (259.1KB, pdf) Supplement 2. eMethods 1. Sample Size Calculation and Power Analysis eMethods 2. Missing Data and Multiple Imputation eMethods 3. Statistical Modeling of Group Clustering eTable 1. Structure and Therapeutic Elements of the AI-Based Emotional Support Platform (Kai.AI) eTable 2. Structure and Core Content of the Face-to-Face Group Therapy eTable 3. Baseline Characteristics of Participants with Complete vs Missing Outcome Data eTable 4. Intention to Seek Therapy by Intervention Group at T1 and T2 eTable 5. Transitions From Clinical to Nonclinical Symptom Status at Postintervention and 3-Month Follow-up eFigure 1. Representative Screenshots From the Kai AI Digital Platform eFigure 2. Intention to Seek Therapy by Intervention Group at Baseline (T1) and Postintervention (T2) jamanetwopen-e266713-s002.pdf (494.6KB, pdf) Supplement 3. Data Sharing Statement jamanetwopen-e266713-s003.pdf (17.4KB, pdf) References 1. Nearly one billion people have a mental disorder: WHO. United Nations News. World Health Organization . June 17, 2022. Accessed September 19, 2025. https://news.un.org/en/story/2022/06/1120682 2. Fan Y, Fan A, Yang Z, Fan D. Global burden of mental disorders in 204 countries and territories, 1990-2021: results from the global burden of disease study 2021. BMC Psychiatry. 2025;25(1):486. doi: 10.1186/s12888-025-06932-y [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Adler J, Van Brunt D. It is time to realize the promise of the digital mental health revolution: application for population mental health. J Med Internet Res. 2025;27:e63791. doi: 10.2196/63791 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Baumel A. Therapeutic activities as a link between program usage and clinical outcomes in digital mental health interventions: a proposed research framework. J Technol Behav Sci. 2022;7:234-239. doi: 10.1007/s41347-022-00245-7 [ DOI ] [ Google Scholar ] 5. Torous J, Linardon J, Goldberg SB, et al. The evolving field of digital mental health: current evidence and implementation issues for smartphone apps, generative artificial intelligence, and virtual reality. World Psychiatry. 2025;24(2):156-174. doi: 10.1002/wps.21299 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Vaidyam AN, Wisniewski H, Halamka JD, Kashavan MS, Torous JB. Chatbots and conversational agents in mental health: a review of the psychiatric landscape. Can J Psychiatry. 2019;64(7):456-464. doi: 10.1177/0706743719828977 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Siddals S, Torous J, Coxon A.. “It happened to be the perfect thing”: experiences of generative AI chatbots for mental health. npj Ment Health Res. 2024;3:48. doi: 10.1038/s44184-024-00097-4 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Nass C, Moon Y. Machines and mindlessness: social responses to computers. J Soc Issues. 2000;56(1):81-103. doi: 10.1111/0022-4537.00153 [ DOI ] [ Google Scholar ] 9. Li H, Zhang R, Lee YC, Kraut RE, Mohr DC. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. NPJ Digit Med. 2023;6(1):236. doi: 10.1038/s41746-023-00979-5 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Zhang Q, Zhang R, Xiong Y, Sui Y, Tong C, Lin FH. Generative AI mental health chatbots as therapeutic tools: systematic review and meta-analysis of their role in reducing mental health issues. J Med Internet Res. 2025;27:e78238. doi: 10.2196/78238 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Heinz MV, Mackin DM, Trudeau BM, et al. Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI. Published online March 27, 2025. doi: 10.1056/AIoa2400802 [ DOI ] [ Google Scholar ] 12. Naor N, Frenkel A, Winsberg M. Improving well-being with a mobile artificial intelligence–powered acceptance commitment therapy tool: pragmatic retrospective study. JMIR Form Res. 2022;6(7):e36018. doi: 10.2196/36018 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Vertsberger D, Naor N, Winsberg M. Adolescents’ well-being while using a mobile artificial intelligence–powered acceptance commitment therapy tool: evidence from a longitudinal study. JMIR AI. 2022;1(1):e38171. doi: 10.2196/38171 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Spitzer RL, Kroenke K, Williams JB, Löwe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med. 2006;166(10):1092-1097. doi: 10.1001/archinte.166.10.1092 [ DOI ] [ PubMed ] [ Google Scholar ] 15. Kroenke K, Spitzer RL, Williams JB. The PHQ-9: validity of a brief depression severity measure. J Gen Intern Med. 2001;16(9):606-613. doi: 10.1046/j.1525-1497.2001.016009606.x [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Price M, Szafranski DD, van Stolk-Cooke K, Gros DF. Investigation of abbreviated 4 and 8 item versions of the PTSD Checklist 5. Psychiatry Res. 2016;239:124-130. doi: 10.1016/j.psychres.2016.03.014 [ DOI ] [ PubMed ] [ Google Scholar ] 17. Seligson JL, Huebner ES, Valois RF. Preliminary validation of the brief multidimensional students’ life satisfaction scale (BMSLSS). Soc Indic Res. 2003;61:121-145. doi: 10.1023/A:1021326822957 [ DOI ] [ Google Scholar ] 18. Sischka PE, Costa AP, Steffgen G, Schmidt AF. The WHO-5 well-being index – validation based on item response theory and the analysis of measurement invariance across 35 countries. J Affect Disord Rep. 2020;1:100020. doi: 10.1016/j.jadr.2020.100020 [ DOI ] [ Google Scholar ] 19. Wilson CJ, Deane FP, Ciarrochi J, Rickwood D. Measuring help-seeking intentions: properties of the general help-seeking questionnaire. Can J Counsell. 2005;39(1):15-28. [ Google Scholar ] 20. Cash TF, Kehr J, Salzbach RF. Help-seeking attitudes and perceptions of counselor behavior. J Couns Psychol. 1978;25(4):264-269. doi: 10.1037/0022-0167.25.4.264 [ DOI ] [ Google Scholar ] 21. Nahum-Shani I, Murphy SA. Just-in-time adaptive interventions: where are we now and what is next? Annu Rev Psychol. 2026;77(1):679-703. doi: 10.1146/annurev-psych-121024-044244 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Thirupathi L, Kaashipaka V, Dhanaraju M, Katakam V. AI and IoT in mental health care: from digital diagnostics to personalized, continuous support. In: Joshi H, Kumar Reddy C, Ouaissa M, Hanafiah M, Doss S, eds. Intelligent Systems and IoT Applications in Clinical Health. IGI Global Scientific Publishing; 2025:271-294. [ Google Scholar ] 23. Shim H, Cho J, Sung YH. Unveiling secrets to AI agents: exploring the interplay of conversation type, self-disclosure, and privacy insensitivity. Asian Communication Research. 2024;21(2):195-216. doi: 10.20879/acr.2024.21.019 [ DOI ] [ Google Scholar ] 24. Malouin-Lachance A, Capolupo J, Laplante C, Hudon A. Does the digital therapeutic alliance exist? integrative review. JMIR Ment Health. 2025;12:e69294. doi: 10.2196/69294 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Bisson JI, Roberts NP, Andrew M, Cooper R, Lewis C. Psychological therapies for chronic post-traumatic stress disorder (PTSD) in adults. Cochrane Database Syst Rev. 2013;(12):CD003388. doi: 10.1002/14651858.CD003388.pub4 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplement 1. Trial Protocol jamanetwopen-e266713-s001.pdf (259.1KB, pdf) Supplement 2. eMethods 1. Sample Size Calculation and Power Analysis eMethods 2. Missing Data and Multiple Imputation eMethods 3. Statistical Modeling of Group Clustering eTable 1. Structure and Therapeutic Elements of the AI-Based Emotional Support Platform (Kai.AI) eTable 2. Structure and Core Content of the Face-to-Face Group Therapy eTable 3. Baseline Characteristics of Participants with Complete vs Missing Outcome Data eTable 4. Intention to Seek Therapy by Intervention Group at T1 and T2 eTable 5. Transitions From Clinical to Nonclinical Symptom Status at Postintervention and 3-Month Follow-up eFigure 1. Representative Screenshots From the Kai AI Digital Platform eFigure 2. Intention to Seek Therapy by Intervention Group at Baseline (T1) and Postintervention (T2) jamanetwopen-e266713-s002.pdf (494.6KB, pdf) Supplement 3. Data Sharing Statement jamanetwopen-e266713-s003.pdf (17.4KB, pdf) Articles from JAMA Network Open are provided here courtesy of American Medical Association ACTIONS View on publisher site Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Related documents

Record · ID 14813 · SHA-256 9bbf6b0d9fb12c22
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.