Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice PLOS Digit Health . 2026 Apr 8;5(4):e0001354. doi: 10.1371/journal.pdig.0001354 Search in PMC Search in PubMed View in NLM Catalog Add to search A systematic review of the limitations of large language models in generating healthcare content Mohsen Khosravi Mohsen Khosravi 1 Social Determinants of Health Research Center, Birjand University of Medical Sciences, Birjand, Iran Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Writing – original draft Find articles by Mohsen Khosravi 1, * , Zahra Zamaninasab Zahra Zamaninasab 2 Department of Epidemiology and Biostatistics, School of Health, Social Determinants of Health Research Center, Birjand University of Medical Sciences, Birjand, Iran Formal analysis, Visualization, Writing – review & editing Find articles by Zahra Zamaninasab 2 , Seyyed Morteza Mojtabaeian Seyyed Morteza Mojtabaeian 3 Department of Healthcare Services Management, School of Management and Medical Informatics, Shiraz University of Medical Sciences, Shiraz, Iran Data curation, Investigation, Resources Find articles by Seyyed Morteza Mojtabaeian 3 , Emine Kübra Dindar Demiray Emine Kübra Dindar Demiray 4 Department of Infection Diseases and Clinical Microbiology, Siirt University Medical School, Siirt, Türkiye Validation, Writing – review & editing Find articles by Emine Kübra Dindar Demiray 4 , Morteza Arab-Zozani Morteza Arab-Zozani 1 Social Determinants of Health Research Center, Birjand University of Medical Sciences, Birjand, Iran Validation, Writing – review & editing Find articles by Morteza Arab-Zozani 1 Editor: Mayue Shi 5 Author information Article notes Copyright and License information 1 Social Determinants of Health Research Center, Birjand University of Medical Sciences, Birjand, Iran 2 Department of Epidemiology and Biostatistics, School of Health, Social Determinants of Health Research Center, Birjand University of Medical Sciences, Birjand, Iran 3 Department of Healthcare Services Management, School of Management and Medical Informatics, Shiraz University of Medical Sciences, Shiraz, Iran 4 Department of Infection Diseases and Clinical Microbiology, Siirt University Medical School, Siirt, Türkiye 5 University of Oxford, UNITED KINGDOM OF GREAT BRITAIN AND NORTHERN IRELAND The authors have declared that no competing interests exist. ✉ * E-mail: [email protected] Roles Mohsen Khosravi : Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Writing – original draft Zahra Zamaninasab : Formal analysis, Visualization, Writing – review & editing Seyyed Morteza Mojtabaeian : Data curation, Investigation, Resources Emine Kübra Dindar Demiray : Validation, Writing – review & editing Morteza Arab-Zozani : Validation, Writing – review & editing Mayue Shi : Editor Received 2025 Nov 5; Accepted 2026 Mar 21; Collection date 2026 Apr. © 2026 Khosravi et al This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice PMCID: PMC13061218 PMID: 41950239 Abstract Large language models (LLMs) have recently gained prominence in healthcare content provision due to their numerous advantages. Despite these benefits, LLMs exhibit notable limitations in this domain. This study aimed to systematically identify the limitations of LLMs in provision of healthcare content. This study was a systematic review conducted in September 2025, including articles published in English between 2018 and 2025. Searches were performed in PubMed, Scopus, and the Cochrane Database of Systematic Reviews. Two independent evaluators screened the references and assessed quality of the selected studies using the Authority, Accuracy, Coverage, Objectivity, Date, and Significance (AACODS) checklist. Data were analyzed using Boyatzis’s qualitative thematic approach with an inductive methodology, applying the input–process–output (IPO) model as the analytical framework. A total of 81 studies were included in the final analysis. The included studies were predominantly of high quality and demonstrated minimal risk of bias. The thematic analysis identified key themes: data limitations, dependence on input and prompt quality, accessibility issues, model design and architecture constraints, interaction challenges, response quality and comprehensiveness, and ethical, safety, and regulatory concerns. The study identified multiple limitations of LLMs in healthcare, with output issues being most common. In this regard, the most frequently cited limitation was the accuracy gap. However, these output issues were mainly resulted from flaws in input data, emphasizing the crucial role of input quality. The study also proposed strategies to address these challenges. Author summary This study systematically reviewed the existing literature regarding the limitations of large language models(LLMs) in provision of healthcare content. A group of prominent databases were searched and the screened references were assessed in terms of quality. Finally, a thematic analysis was conducted on the data derived from the included studies corresponding to the research question using the input-output model as an analytical framework. The search yielded 81 studies and the quality of the included studies was presented to be predominantly high. The thematic analysis yielded a number of themes and sub-themes. The results of the research found that while the main area of LLMs` limitations is corresponding to the outputs, and particularly the existing gaps in accuracy, these limitations are shown to be derived from the existing flaws in the input data. The study also presented some strategies to overcome these limitations based on the existing data within the literature. 1. Introduction Large language models (LLMs) are sophisticated artificial intelligence (AI) systems developed through extensive training on vast corpora of text data, enabling them to generate outputs that closely resemble human language [ 1 ]. These models have been widely utilized across diverse medical domains, including health informatics, medical imaging, clinical diagnostics, treatment planning, ophthalmology, oncology, and other specialized fields [ 2 ]. This trend signifies their broad and growing integration into medical research and clinical practice. LLMs have become pivotal in healthcare by enhancing clinical decision support, diagnostics, medical education, and patient engagement [ 2 – 4 ]. They improve diagnostic accuracy by analyzing extensive clinical data and medical literature, aiding in personalized treatment planning and patient care management [ 3 ]. LLMs are increasingly integrated into hospitals, clinical settings, academic medical centers, and virtual care platforms, supporting healthcare providers with evidence-based recommendations and facilitating patient interactions through chatbots and virtual assistants [ 2 ]. Furthermore, they assist in research by automating documentation and synthesizing biomedical information, thereby optimizing clinical workflows [ 5 ]. Notwithstanding the numerous advantages previously discussed, LLMs in healthcare exhibit several critical limitations that must be addressed to ensure their safe and effective deployment. In this regard, they are prone to generating plausible yet factually incorrect or fabricated information, a phenomenon known as hallucination, which poses significant risks in clinical settings [ 3 , 6 ]. Furthermore, LLMs often lack the depth of contextual understanding required to accurately interpret complex medical scenarios, as they may fail to integrate multifaceted clinical data or temporal information adequately [ 3 ]. The development and application of LLMs are also constrained by limited access to high-quality, diverse clinical datasets due to privacy, ethical, and legal challenges [ 7 ]. Ethical and legal issues such as bias, misinformation, data privacy, and insufficient regulatory frameworks further challenge their adoption [ 8 ]. Additionally, the opaque, “black box” nature of these models undermines transparency and interpretability, complicating trust and reliance by healthcare professionals [ 9 ]. Practical concerns include their substantial computational and energy demands, which limit feasibility in resource-constrained environments [ 3 ]. Such limitations of LLMs pose significant challenges to their adoption and utilization across various sectors of healthcare systems. In this regard, these challenges must be critically addressed to enable the successful and comprehensive integration of LLM technologies within healthcare infrastructure. While several review studies have aimed to delineate the limitations of LLMs in providing healthcare content, there remains a significant gap in the literature regarding a comprehensive and systematic presentation of these limitations for end-users [ 3 , 6 – 9 ]. In this regard, a systematic review categorized the limitations of LLMs into two primary domains: design and output. Design limitations included several items such as lack of optimization for the medical domain, data transparency issues, and accessibility challenges. Output limitations also included several items such as non-reproducibility, incompleteness, inaccuracies, safety concerns, and biases [ 6 ]. The data generated from research on the limitations of LLMs would be invaluable for technology developers to enhance the quality of their models. Additionally, healthcare policymakers and administrators could leverage this information to make informed decisions about implementing these technologies within their organizations, fully acknowledging their current constraints. Moreover, future researchers would benefit from this detailed framework by conducting focused investigations on each identified limitation, thereby contributing further insights and advancing the field for subsequent users and stakeholders. 2. Results As demonstrated in Fig 1 , the database search was conducted on September 13, 2025, yielding 409 references from the Cochrane Database of Systematic Reviews, 1,766 references from PubMed, and 3,278 references from Scopus, of which 1,798 were identified as duplicates. Following the screening of the retrieved references, a total of 81 studies were included as the final selections for the study. 20% of the included studies were published in 2023, 40% in 2024, and the remaining 40% in 2025. Fig 1. PRISMA diagram. Open in a new tab 2.1. Data quality The quality assessment of the included studies indicated that they were predominantly of high quality, with an average score of 10. Approximately 34% of the studies achieved the highest quality score of 12, while about 8% scored 7, reflecting lower quality relative to the other included studies. Moreover, as delineated by the data corresponding to the objectivity item of the AACODS checklist, the level of bias in the included studies was generally low, with approximately 59% of the studies exhibiting the minimal possible bias ( S1 Appendix ). 2.2. Data analysis The thematic analysis identified a total of eight distinct themes within the four major categories outlined by the IPO model concerning the limitations of LLMs in healthcare content generation. The themes included data limitations, dependence on input quality, dependence on prompt quality, accessibility issues, model design and architecture limitations, interaction challenges, response quality and comprehensiveness, and ethical, safety, and regulatory concerns ( Table 1 ). Table 1. Thematic analysis of findings. Category Theme Sub-theme Reference(s) Input limitations Data limitations Scarcity of Medical Dialogue Datasets [ 10 ] Dependence on Comprehensive Training Data [ 11 – 17 ] Dependence on input quality Need for Structured and Clear Inputs [ 18 , 19 ] Complexity of Language and Dialects [ 10 ] Limited Ability to Recognize non-verbal Signals [ 20 ] Session and Question Limits [ 21 ] Limited Multimodal Data Processing [ 13 , 19 , 22 – 25 ] Issues with question difficulty and length [ 26 ] Dependence on prompt quality Need for Prompt Engineering [ 13 , 24 , 25 , 27 – 38 ] Vulnerability to Adversarial prompts [ 39 ] Accessibility issues Dependency on access to model internals and updates [ 33 , 40 – 43 ] Dependency on Stable Internet [ 44 ] Process limitations Model Design and Architecture Limitations High Power Consumption [ 44 , 45 ] Algorithmic Bias [ 13 – 15 , 19 , 20 , 23 , 25 , 40 , 41 , 43 , 45 – 54 ] Hardware Constraints [ 43 , 44 ] Lack of Clinical Experience and Judgment [ 19 ] Over-reliance on imaging modalities [ 13 , 29 , 55 ] Inability for Critical Thinking [ 30 , 41 , 56 ] Reliance on Provided Data [ 56 ] Interaction Challenges Requirement for Repeated Interactions [ 16 , 24 , 57 ] Lack of interaction [ 23 , 56 , 58 ] Conversation Tracking Issues [ 21 , 59 ] Latency Issues [ 44 ] Inability to Ask Clarifying Questions [ 60 ] Output limitations Response Quality and Comprehensiveness Limited Depth of Responses in Complex Assignments [ 12 , 13 , 16 , 18 , 22 – 24 , 26 – 28 , 30 , 31 , 36 , 37 , 42 , 49 , 50 , 53 , 56 , 58 , 60 – 73 ] Restricted Response Length [ 27 , 61 , 74 ] Incomplete responses [ 40 ] Inconsistency in Responses [ 11 , 15 , 19 , 28 , 40 , 50 , 51 , 63 , 75 – 77 ] Repetitive and Vague Recommendations [ 15 , 30 , 56 ] Lack of Personalization and Clinical Nuance [ 17 , 20 , 23 , 30 , 35 , 37 , 45 , 53 , 57 , 58 , 60 , 66 , 68 , 72 , 73 , 76 , 78 – 82 ] Limited Actionability [ 11 , 14 , 17 , 25 , 27 , 31 , 50 , 55 , 63 ] Accuracy Gaps [ 12 , 13 , 15 , 16 , 19 – 21 , 23 , 25 , 26 , 28 , 29 , 31 , 34 – 37 , 40 , 42 , 43 , 46 , 48 , 49 , 51 , 53 , 55 – 58 , 64 , 65 , 68 – 70 , 72 , 73 , 75 – 77 , 79 – 86 ] Advanced Reading Level [ 15 , 17 , 25 , 27 , 31 , 32 , 34 , 50 , 67 , 70 , 81 , 87 , 88 ] Variable Translation Quality by Language [ 10 , 32 , 67 , 76 , 77 , 80 , 89 ] Outdated References [ 11 , 19 , 41 , 50 , 51 , 55 , 57 , 59 – 61 , 76 ] Ethical, Safety, and Regulatory Concerns Ethical, and Regulatory Challenges [ 11 , 13 , 14 , 19 , 20 , 23 , 35 , 39 – 41 , 43 – 48 , 50 , 52 – 54 , 75 , 76 , 84 , 86 , 90 , 91 ] Regulatory and Quality Control Issues [ 46 ] Risk of Overreliance [ 20 , 22 , 37 , 52 , 53 , 65 , 76 , 78 , 91 ] Risk of Misinterpretation [ 76 ] Lack of Transparency [ 11 , 34 , 40 , 53 , 70 ] Lack of Real-World Validation [ 16 , 23 , 25 , 35 , 43 , 66 ] Inconsistent Use of Disclaimers [ 14 , 21 , 34 , 67 ] Inability to Replace Human [ 14 , 19 , 20 , 22 , 23 , 26 , 35 , 36 , 40 , 42 , 45 – 47 , 49 , 51 , 57 , 58 , 60 , 68 , 69 , 71 , 73 , 76 , 79 – 81 , 83 , 85 , 88 ] Lack of Self-awareness [ 65 ] Risk of dangerous information [ 39 , 84 ] Risk of Misinformation in Less Common Languages [ 32 , 67 , 76 , 77 , 80 , 89 ] Open in a new tab 2.2.1. Input limitations. 2.2.1.1. Data limitations: The scarcity of medical dialogue datasets in Arabic was highlighted as a critical constraint, primarily due to privacy concerns and the sensitive nature of medical conversations, which results in limited availability of comprehensive and representative datasets [ 10 ]. Additionally, the effective functioning of these models was found to heavily depend on access to comprehensive, unbiased, and up-to-date training data. Any gaps or deficiencies in such data adversely affected the quality and reliability of the generated outputs, underscoring the essential role of robust and well-maintained datasets in ensuring model accuracy and relevance within healthcare contexts [ 11 – 17 ]. 2.2.1.2. Dependence on input quality: A key issue identified was the need for structured and clear inputs, as unstructured or ambiguous data inputs were found to reduce AI accuracy and effectiveness in clinical decision-making [ 18 , 19 ]. The complexity of some languages such as the Arabic language, characterized by its rich morphology and diverse dialects, further complicates natural language processing tool development compared to more standardized languages [ 10 ]. Additionally, the models exhibited a limited ability to recognize non-verbal signals, such as subtle crisis indicators or nuanced mental states, which restricts their utility in critical situations like suicidal or homicidal ideation [ 20 ]. Session and question limits, such as the ChatGPT 4.0 restriction of 40 questions per three hours—absent in version 3.5—also potentially hinder the model’s effectiveness as a telepharmacy or healthcare tool [ 21 ]. Moreover, current large language models possess limited multimodal data processing capabilities, primarily handling text with minimal ability to interpret images or other data types, thereby constraining their applicability in fields heavily reliant on imaging, including radiology and pathology [ 13 , 19 , 22 – 25 ]. Finally, the complexity and length of questions were negatively correlated with accuracy, as longer and more difficult questions were more likely to be answered incorrectly [ 26 ]. 2.2.1.3. Dependence on prompt quality: Optimized prompt engineering is essential to enhance clarity, actionability, and readability of model outputs while minimizing the risk of misinformation [ 13 , 24 , 25 , 27 – 38 ]. Additionally, these models demonstrated vulnerability to adversarial prompts, whereby maliciously crafted inputs could circumvent existing safeguards. For instance, even advanced models such as ChatGPT-4.0 were susceptible to manipulation that enabled generation of harmful and detailed instructions, including those that could potentially cause ocular damage through biological, chemical, or physical means [ 39 ]. 2.2.1.4. Accessibility issues: Restricted access to model internals and updates was found to limit broader implementation and comprehensive understanding of these models within healthcare settings [ 33 , 40 – 43 ]. Furthermore, most large language models require a stable internet connection for cloud-based processing, which introduces latency and significantly restricts their usability in offline environments or regions with poor connectivity [ 44 ]. 2.2.2. Process limitations. 2.2.2.1. Model design and architecture limitations: High power consumption associated with cloud-based LLM usage was identified as a constraint, particularly for portable or energy-sensitive applications [ 44 , 45 ]. Algorithmic bias present in training data, such as racial bias, resulted in models like ChatGPT producing varied recommendations based on patient race or ethnicity, thereby perpetuating healthcare disparities and obscuring such biases due to the models’ opaque “black box” nature [ 13 – 15 , 19 , 20 , 23 , 25 , 40 , 41 , 43 , 45 – 54 ]. Hardware constraints were also noted, with limited memory and processing capacity of devices like microcontrollers (e.g., ESP32 and ESP8266) posing challenges to the complexity and responsiveness of LLM-powered healthcare applications [ 43 , 44 ]. Additionally, the lack of real-world clinical experience restricted LLMs’ ability to perform complex clinical reasoning, diagnostic accuracy, and higher-order judgment essential for medical decision-making [ 19 ]. There was also an identified over-reliance on imaging modalities such as CT and MRI, without adequate customization for specific clinical contexts [ 13 , 29 , 55 ]. Furthermore, LLMs like ChatGPT demonstrated an inability for critical thinking necessary to tailor and guide patient management effectively [ 30 , 41 , 56 ]. Finally, responses generated by these models depended solely on the provided data without interpretative insight or clinical opinion, risking omission of underlying clinical nuances—such as failing to detect depression in patients presenting with non-specific symptoms like sleep disturbances [ 56 ]. 2.2.2.2. Interaction challenges: Effective use by clinicians often required repeated interactions, involving multiple queries and refined questioning to obtain accurate and relevant responses [ 16 , 24 , 57 ]. However, the models lacked the capability for dynamic interaction, as they could not gather additional information necessary for precise diagnosis and management [ 23 , 56 , 58 ]. Conversation tracking posed further limitations: ChatGPT 3.5 lacked conversation memory due to privacy and browsing restrictions, while ChatGPT 4.0, despite tracking conversations, inaccurately counted inquiries, thereby undermining feedback and dialogue continuity [ 21 , 59 ]. Latency issues arising from data transmission to cloud servers also presented challenges, potentially delaying real-time healthcare applications that demand immediate responses [ 44 ]. Moreover, these models were unable to ask clarifying questions to seek further clinical clues, which constrained their diagnostic accuracy and the precision of their advice [ 60 ]. 2.2.3. Output limitations. 2.2.3.1. Response quality and comprehensiveness: LLMs often provided limited depth in complex tasks, offering superficial or incomplete answers, particularly for clinical interventions, follow-up discussions, critical appraisals, and complex data reasoning, resulting in poorer performance compared to simpler question types [ 12 , 13 , 16 , 18 , 22 – 24 , 26 – 28 , 30 , 31 , 36 , 37 , 42 , 49 , 50 , 53 , 56 , 58 , 60 – 73 ]. Additionally, response length was constrained by a maximum word count (approximately 650 words for ChatGPT), limiting comprehensive critical analysis and extensive discussion of complex healthcare topics, though improvements are anticipated in newer model versions [ 27 , 61 , 74 ]. Incomplete responses were also noted, including occasional neglect of imaging descriptions [ 40 ]. Consistency issues arose as LLMs sometimes generated varying answers to identical questions or prompts, undermining reliability [ 11 , 15 , 19 , 28 , 40 , 50 , 51 , 63 , 75 – 77 ]. Further prompting often produced broad, repetitive, and vague recommendations lacking personalization or precision [ 15 , 30 , 56 ]. Responses frequently missed subtle clinical nuances and tailored patient-specific details, necessitating professional oversight to ensure accuracy and patient safety [ 17 , 20 , 23 , 30 , 35 , 37 , 45 , 53 , 57 , 58 , 60 , 66 , 68 , 72 , 73 , 76 , 78 – 82 ]. While chatbot outputs were generally understandable, they were often insufficiently actionable, limiting patients’ ability to take clear steps, and despite streamlining some administrative tasks, LLMs currently offer limited impact on routine clinical care without further development [ 11 , 14 , 17 , 25 , 27 , 31 , 50 , 55 , 63 ]. Accuracy gaps were evident, with occasional imprecision, indecisiveness, hallucinations, irrelevant information, and omissions of key clinical considerations, particularly for special populations such as pregnant patients. Moreover, fabricated references were also observed sporadically [ 12 , 13 , 15 , 16 , 19 – 21 , 23 , 25 , 26 , 28 , 29 , 31 , 34 – 37 , 40 , 42 , 43 , 46 , 48 , 49 , 51 , 53 , 55 – 58 , 64 , 65 , 68 – 70 , 72 , 73 , 75 – 77 , 79 – 86 ]. The model’s language was often at an advanced reading level, complicating accessibility and comprehension for many patients [ 15 , 17 , 25 , 27 , 31 , 32 , 34 , 50 , 67 , 70 , 81 , 87 , 88 ]. Moreover, translation quality varied across languages, and references cited were frequently outdated, with a tendency to neglect recent high-quality studies such as randomized controlled trials, reducing the reliability of evidence-based content, especially in critical appraisals where current literature is essential [ 10 , 11 , 19 , 32 , 41 , 50 , 51 , 55 , 57 , 59 – 61 , 67 , 76 , 77 , 80 , 89 ]. 2.2.3.2. Ethical, safety, and regulatory concerns: Ethical and regulatory challenges emphasized the need for enforceable regulations, comprehensive ethical frameworks, and proactive controls to prevent misuse, particularly given the sensitivity of AI applications in healthcare and biowarfare. Issues of authorship, accountability, and trustworthiness were raised as AI cannot assume responsibility for its outputs and may perpetuate biases embedded in training data. Patient data privacy and vulnerability to cyber threats further necessitate robust data protection measures [ 11 , 13 , 14 , 19 , 20 , 23 , 35 , 39 – 41 , 43 – 48 , 50 , 52 – 54 , 75 , 76 , 84 , 86 , 90 , 91 ]. Regulatory and quality control challenges were highlighted due to the evolving nature of AI models and the current underdevelopment of regulatory frameworks governing algorithmic medicine [ 46 ]. The risk of overreliance emerged as a critical concern, with users potentially accepting AI-generated information without critical evaluation or expert consultation, posing risks of misinformation and harm. Excessive dependence on ChatGPT was also associated with increased social isolation, potentially exacerbating depression in vulnerable populations [ 20 , 22 , 37 , 52 , 53 , 65 , 76 , 78 , 91 ]. Misinterpretation risks were identified due to variability in user comprehension, which could lead to adverse clinical outcomes [ 76 ]. The lack of transparency stemming from the model’s “black box” operation limits the ability to evaluate responses, as the rationale behind conclusions remains unclear [ 11 , 34 , 40 , 53 , 70 ]. Additionally, the absence of real-world clinical validation restricts independent clinical use [ 16 , 23 , 25 , 35 , 43 , 66 ]. The inconsistent use of disclaimers—more frequent in earlier versions like ChatGPT 3.5, but less so in newer versions—raises questions regarding role compliance [ 14 , 21 , 34 , 67 ]. AI’s inability to replace human judgment was noted, given its lack of genuine understanding, empathy, and ethical reasoning, functioning instead as a “stochastic parrot” mimicking language without true comprehension [ 14 , 19 , 20 , 22 , 23 , 26 , 35 , 36 , 40 , 42 , 45 – 47 , 49 , 51 , 57 , 58 , 60 , 68 , 69 , 71 , 73 , 76 , 79 – 81 , 83 , 85 , 88 ]. The model demonstrated limited self-awareness, seldom acknowledging when it lacked answers and frequently providing incorrect explanations without recognizing these limitations [ 65 ]. Concerns about the risk of generating dangerous information were amplified by potential exploitation by malicious actors [ 39 , 84 ]. Finally, the variability and higher error rates in translations for less common languages heighten the risk of miscommunication in clinical contexts when relying solely on machine-generated translations [ 32 , 67 , 76 , 77 , 80 , 89 ]. 3. Discussion As presented by the study findings and delineated in Fig 2 , which presents the number of citations for each sub-theme presented in the thematic analysis, the category with the greatest number of limitations identified in the literature was output limitations, underscoring the significant challenges related to the output of large language models in the provision of healthcare content. In line with our study findings, a recent review emphasized that the taxonomy of limitations associated with large language models reveals a substantially greater number of codes related to output issues compared to those connected with design or input phases. These output limitations encompass challenges such as generating accurate, contextually relevant, and comprehensive responses, as well as concerns regarding interpretability, responsiveness, and ethical considerations [ 3 ]. Fig 2. Distributions of limitations of LLMs, categorized according to the frequency of citations reported in the literature. Open in a new tab The findings of our study identified accuracy gaps (i.e., fabricated responses) as the most frequently reported limitation of LLMs. This provides the rationale for the notable increase in recent studies evaluating the accuracy of LLMs [ 92 , 93 ]. In accordance with these findings, which highlight the significance of limitations related to the output of LLMs, this section is primarily dedicated to analyzing the existing data on this subject, taking into account findings from the literature in other contexts. The findings of our study identified several output limitations of LLMs in generating healthcare content. These limitations included issues with response quality and comprehensiveness, such as limited depth of responses in complex assignments, restricted response length, incomplete answers, inconsistency, as well as repetitive and vague recommendations. Additionally, the models demonstrated a lack of personalization and clinical nuance, limited actionability, and accuracy gaps (i.e., fabricated responses). Other challenges involved advanced reading levels, variable translation quality across languages, and reliance on outdated references. Ethical, safety, and regulatory concerns were also evident, including regulatory and quality control issues, risks of overreliance and misinterpretation, as well as lack of transparency and real-world validation. The inconsistent use of disclaimers, inability to replace human expertise, lack of self-awareness, and risks of disseminating dangerous information or misinformation—particularly in less common languages—were further significant limitations noted [ 10 – 32 , 34 – 37 , 39 – 91 ]. The literature has identified multiple reasons for the issues related to response quality and comprehensiveness in LLMs. In this regard, a significant portion of these issues originates from limitations in the training data. For instance, LLMs frequently generate incomplete answers due to gaps in their training datasets and their inability to fully integrate complex or comprehensive information. This shortcoming is particularly critical in medical settings, where the omission of essential information can lead to inadequate clinical decisions or treatment recommendations [ 6 , 8 ]. Moreover, flawed, biased, or incomplete data can propagate inaccuracies and inconsistencies within the model outputs, reducing their reliability and safety, especially in healthcare applications [ 94 , 95 ]. Additionally, LLMs are limited by the cutoff date of their training data, causing reliance on outdated information, which diminishes the clinical relevance and accuracy of their responses [ 95 , 96 ]. These findings highlighted the interconnected nature of limitations in the LLMs, emphasizing that the core issues in output quality stem from deficiencies in the input data. Training data plays a crucial role, significantly impacting the quality and comprehensiveness of generated content. In this regard, our study indicated that while output limitations are numerous, their root causes lie primarily in the constraints and imperfections of the input data, underscoring the dependency of LLM outputs on input data quality. The literature presents several strategies to improve the quality and comprehensiveness of content generated by LLMs. Some strategies emphasize enhancing the input quality to ultimately improve output quality. For example, role-playing prompts significantly increase accuracy, comprehensiveness, and acceptability of responses by guiding LLMs to adopt specific personas or roles, which encourages more detailed and contextually appropriate answers, as demonstrated in models like ChatGPT-4 and ChatGPT-3.5 [ 97 ]. Additionally, prompt engineering and augmentation—such as designing carefully structured prompts and incorporating relevant context or supporting data—enhance response quality. Retrieval augmented generation (RAG) techniques, which integrate verified external data into the generation process, further improve the accuracy and comprehensiveness of responses [ 98 ]. Other strategies focus on refining the training of LLMs to elevate the quality of generated content. Multi-modal and task-specific training, which involves training models on diverse data types like text combined with images and specialized datasets, allows LLMs to better manage complex healthcare questions with greater depth and clarity [ 99 ]. Fine-tuning and adaptation using domain-specific datasets, particularly healthcare data, improve the models’ performance on specialized tasks such as medical question answering and clinical note summarization. Techniques like instruction tuning and reinforcement learning from human feedback (RLHF) enhance model alignment with clinical expectations, thereby improving precision and relevance [ 3 ]. Some strategies center on improving feedback mechanisms for LLM outputs to enhance response quality. Systematic evaluation and benchmarking with healthcare-specific metrics—focusing on factual accuracy, medical reasoning, and readability (e.g., USMLE, PubMedQA)—support the continuous assessment and improvement of model outputs [ 99 ]. Moreover, incorporating human oversight through expert review and iterative feedback during model development and deployment promotes safety, ethical integrity, and clinical appropriateness of responses [ 97 ]. The study findings also identified an insufficiency in clinical nuance as a notable limitation of LLMs [ 17 , 20 , 23 , 30 , 35 , 37 , 45 , 53 , 57 , 58 , 60 , 66 , 68 , 72 , 73 , 76 , 78 – 82 ]. In this regard, the evidence indicates that, owing to ethical constraints and the highly specialized nature of the medical domain, the acquisition, processing, and utilization of clinical data by LLMs are subject to substantial restrictions. Furthermore, LLMs are not explicitly trained on medical records curated and selected by qualified clinical professionals [ 100 ]. Collectively, these factors likely constitute the primary underlying causes of the limited clinical nuance observed in LLM outputs. The literature outlines several strategies to address ethical issues in healthcare content generation by the LLMs. These strategies can be broadly categorized into governance, technical, input-focused, and output-focused approaches. Governance strategies emphasize human oversight, including the establishment of legal and policy frameworks to ensure data governance and accountability [ 101 , 102 ]. Clinician-in-the-loop review and participatory development processes help maintain human expertise and reduce overreliance on automated outputs [ 103 ]. Continuous real-world validation is essential for ensuring reliability, especially in diverse linguistic contexts [ 104 ]. Technical strategies focus on explainable AI methods and audit trails to foster trust and accountability [ 105 ]. Ethical considerations for model inputs include safety-first prompting techniques designed to enhance ethical reasoning in AI outputs [ 106 , 107 ]. Finally, output-focused strategies advocate for standardized disclaimers that clarify AI limitations and help minimize risks of misinterpretation [ 108 ]. Our study findings indicated that LLMs face significant risks associated with the dissemination of harmful information or misinformation, particularly in less commonly used languages [ 32 , 67 , 76 , 77 , 80 , 89 ]. In this regard, evidence indicates that human evaluation methods are costly and difficult to standardize, especially for rapidly evolving LLMs [ 109 ]. On the other hand, novel approaches, enhanced with the optimized algorithm, have achieved greater accuracy, faster detection, and improved scalability than traditional methods, offering a robust and reliable defense against these threats [ 110 , 111 ]. Open-source benchmarks derived from physician certification exams—such as MedQA, PubMedQA, and the Massive Multitask Language Understanding (MMLU) dataset—are widely used to assess the knowledge and reasoning of medical LLMs, including those promoted for patient-facing use [ 112 – 115 ]. In such context, benchmark-based evaluations can serve as a scalable proxy for identifying model behaviors that could adversely affect patient care. Despite recent advancements in LLMs across various aspects of input, processing, and output capabilities, our study provides valuable insights into the limitations of LLMs in delivering healthcare content [ 116 – 118 ]. These insights can be leveraged by relevant stakeholders to develop precise, evidence-based strategies aimed at mitigating these limitations while preserving the strengths of LLMs. While newer versions of LLMs continue to be released, they still exhibit certain limitations [ 117 , 119 ]. Consequently, the data presented in our study remains highly relevant and valuable to various stakeholders, including LLM developers, healthcare policymakers, and administrators aiming to implement LLM technologies within their organizations. 4. Limitations and implications This study had a notable limitation in that it did not include data from studies published in languages other than English due to time and resource constraints. Such a limitation may have resulted in the exclusion of relevant existing data on the topic published in other languages. Consequently, future researchers are encouraged to conduct similar reviews that encompass non-English publications. The study also did not evaluate heterogeneity across different study types, nor did it distinctly delineate limitations specific to clinical, educational, and administrative contexts. On the other hand, this study offers several important implications for key stakeholders, including healthcare policymakers, administrators, researchers, and developers of large language models. Specifically, the findings highlighted output-related limitations as the most significant category of constraints affecting LLMs in the provision of healthcare content. In this regard, the most frequently cited limitation was the accuracy gap. This identification of the areas with the greatest vulnerability and weakness can guide beneficiaries in prioritizing and implementing strategic and operational plans, while also informing researchers in their efforts to mitigate these limitations in future developments. Meanwhile, our study findings indicated that although limitations in LLM outputs are numerous, their primary root causes lie in the constraints and imperfections of the input data. This underscores the critical dependency of LLM output quality on the integrity and quality of the input data. This finding also has significant implications for the beneficiaries. In this regard, our study proposed several strategies to address major issues encountered by the LLMs in the provision of healthcare content. These solutions are designed to be practical and can also be readily adopted by beneficiaries to improve healthcare outcomes and ensure safer, more reliable LLM-assisted content generation. 5. Conclusion The study identified several categories of limitations in LLMs regarding the provision of healthcare content, with output limitations being the most prevalent. In this regard, the most frequently cited limitation was the accuracy gap (i.e., fabricated responses). Our findings indicated that these numerous output limitations primarily stem from constraints and imperfections in the input data. This highlights the critical dependency of LLM output quality on input data integrity. Additionally, our study proposed several strategies to address these key challenges encountered by the LLMs. 6. Methods This study was a systematic review conducted in 2025, following the guidelines set forth by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement to ensure transparent and standardized reporting of its methodology and findings [ 120 ]. Following the completion of the systematic review, a thematic analysis was performed on the collected data. The research question was formulated using the PICOT framework, which includes Population (P), Intervention or Indicator (I), Comparison group (C), Outcome of interest (O), and Timeframe (T) [ 121 ]. In this study, the population was defined as the existing LLMs operating globally; the intervention as the limitations of LLMs in provision of healthcare content, the comparison group encompassed the various types of LLMs reported in the literature, the outcome focused on limitations occurring in the input, process, and output phases, and the timeframe was set from 2018 to 2025. The research question was formulated as follows: “What are the limitations of large language models in the generation of healthcare content in the 2018-2025 time-period?” 6.1. Information sources and search strategy A comprehensive literature search was conducted in September 13, 2025, to identify all articles published in English between 2018 and 2025 that address the limitations of large language models in generating healthcare content. The specified time span was chosen because the initial release of LLMs dates back to 2018 [ 122 ]. The databases searched included PubMed, Scopus, and Cochrane database of systematic reviews. MeSH terms were used to categorize keywords into three groups: limitations, large language models, and healthcare. Synonymous keywords within each group were combined using the logical operator “OR,” and the three groups were subsequently combined using the logical operator “AND.” The search was performed across titles, abstracts, and keywords. Reference management was performed using EndNote version 20.2.1. The detailed search strategy is presented in Table 2 . Table 2. Search strategy. Final Strategy “Limitations” AND “Large language models” AND “Healthcare” Context Keywords Limitations limitation* OR restriction* OR constraint* OR boundar* OR barrier* OR challenge* OR impediment* OR drawback* OR weakness* OR deficienc* OR defect* Large Language Models large language model* OR LLM OR LLMs OR ChatGPT OR Bard OR Claude OR Copilot OR Bing OR Gemini OR Perplexity OR ChatSonic Healthcare healthcare OR health Database-specific search strategies Database Search strategy PubMed ((limitation*[Title/Abstract] OR restriction*[Title/Abstract] OR constraint*[Title/Abstract] OR boundar*[Title/Abstract] OR barrier*[Title/Abstract] OR challenge*[Title/Abstract] OR impediment*[Title/Abstract] OR drawback*[Title/Abstract] OR weakness*[Title/Abstract] OR deficienc*[Title/Abstract] OR defect*[Title/Abstract]) AND (Large Language Models [MeSH Terms] OR large language model*[Title/Abstract] OR LLM[Title/Abstract] OR LLMs[Title/Abstract] OR ChatGPT[Title/Abstract] OR Bard[Title/Abstract] OR Claude[Title/Abstract] OR Copilot[Title/Abstract] OR Bing[Title/Abstract] OR Gemini[Title/Abstract] OR Perplexity[Title/Abstract] OR ChatSonic[Title/Abstract])) AND (Health [MeSH Terms] OR healthcare[Title/Abstract] OR health[Title/Abstract]) Scopus (TITLE-ABS-KEY (limitation* OR restriction* OR constraint* OR boundar* OR barrier* OR challenge* OR impediment* OR drawback* OR weakness* OR deficienc* OR defect*) AND TITLE-ABS-KEY (large language model* OR LLM OR LLMs OR ChatGPT OR Bard OR Claude OR Copilot OR Bing OR Gemini OR Perplexity OR ChatSonic) AND TITLE-ABS-KEY (healthcare OR health)) Cochrane database of systematic reviews limitation* OR restriction* OR constraint* OR boundar* OR barrier* OR challenge* OR impediment* OR drawback* OR weakness* OR deficienc* OR defect* in Title Abstract Keyword AND large language model* OR LLM OR LLMs OR ChatGPT OR Bard OR Claude OR Copilot OR Bing OR Gemini OR Perplexity OR ChatSonic in Title Abstract Keyword AND healthcare OR health in Title Abstract Keyword (Word variations have been searched) Open in a new tab 6.2. Study selection Following the database search, duplicate articles were removed, and the remaining records were screened based on their titles and abstracts. Articles that were not aligned with the research objectives were excluded, and the full texts of the eligible articles were thoroughly reviewed. Only those meeting the predefined inclusion criteria were incorporated into the final analysis. This entire screening and selection process was independently conducted by two authors. In cases of disagreement, a third author was consulted to resolve discrepancies and finalize the screening process. Inclusion criteria: Availability of full text. Published in English language. Published after 2018. Exclusion criteria: Published solely in conferences. This consideration arose from the perception that the quality of peer review for conference papers is generally lower compared to manuscripts published in journals indexed in prominent databases, which undergo a more rigorous and extensive peer review process prior to publication. In this context, the authors concluded that including conference papers could introduce a degree of bias and compromise the overall quality of the study data. Published as letter to editor or protocol. The exclusion was based on the insufficient or lower quality of data presented in these types of research papers. This approach was adopted to ensure the higher quality of data reported in the current study. 6.3. Data quality assessment The quality of all included studies was independently assessed by two evaluators using the Authority, Accuracy, Coverage, Objectivity, Date, and Significance (AACODS) checklist, which consists of six items [ 123 ]. The checklist was selected due to its straightforwardness and clarity, as it effectively presents the quality of the included studies in a transparent and easily understandable manner. A standardized scoring system was applied, assigning 2 points for “Yes,” 1 point for “Can’t Tell,” and 0 points for “No,” resulting in scores ranging from 0 to 12, with higher scores indicating better quality. Studies were classified into four categories based on their scores: very low quality (0–3), low quality [ 4 – 6 ], medium quality [ 7 – 9 ], and high quality [ 10 – 12 ]. In the process, only studies rated as medium or high quality were considered eligible for inclusion in the research. Any disagreements between the two evaluators were resolved through discussion and consultation with a third reviewer acting as an arbitrator. To ensure consistency, this evaluation procedure was conducted twice. 6.4. Data extraction Data extraction from the selected articles was performed independently by two authors. A third author supervised and approved the data extraction process. A data extraction form was developed using Microsoft Excel 2016, which included sections for the study citation, year of publication, and a summary of the findings. 6.5. Data analysis The data collected from the preceding steps were analyzed using the qualitative thematic approach proposed by Boyatzis, employing an inductive methodology. This approach comprises multiple steps, including familiarization with the study data, generation of initial codes, development of themes, and finally reviewing, defining, and naming these themes [ 124 ]. The thematic analysis was also performed using the input–process–output (IPO) model as the foundational framework. The IPO model, also referred to as the input-process-output pattern, is a fundamental framework extensively utilized in systems analysis and software engineering. It serves as a foundational approach for representing the structure of an information processing system or other procedural workflows. This model is commonly introduced in introductory programming and systems analysis literature as the most basic and essential method for describing and conceptualizing a process [ 125 – 128 ]. The term input refers to the resources, data, or materials introduced into a system. The process encompasses the operations or transformations applied to the input to convert it into an output. And, the output denotes the results or products generated following the processing of the input [ 129 ]. In such context, in our study, the category labeled input was defined to include a collection of themes related to the resources, data, or materials introduced into LLMs. The process category was assigned to themes corresponding to the operations or transformations applied to the input provided to the LLMs, facilitating its conversion into an output. Similarly, the output category was designated for themes associated with the results or products generated as a consequence of processing the inputs given to the LLMs. The authors systematically reviewed all the extracted data from the included studies to gain a comprehensive understanding, assigning initial codes to each significant data segment. Each code was developed to capture a unique outcome of artificial intelligence in healthcare. Subsequently, codes sharing similar concepts were grouped into sub-themes, and related sub-themes were further consolidated under overarching main themes. Prior to this categorization, all initial codes underwent thorough examination and refinement. The finalized sub-themes and main themes were then defined, described, and presented in a tabular format. Any disagreements among the authors were resolved through mutual consultation. To ensure the validity and reliability of the qualitative data analysis, the authors adhered to Lincoln and Guba’s criteria, which encompass credibility, transferability, dependability, and confirmability [ 130 ]. In this regard, credibility was established through repeated cross-referencing of the codes with their original data sources. Dependability was ensured by having two authors independently conduct the thematic analysis and compare results to identify discrepancies. Confirmability was maintained via mutual review of the themes, sub-themes, and codes by the authors. Finally, transferability was enhanced by expressing the findings in a manner that facilitates application across diverse contexts within the healthcare system. Supporting information S1 Appendix. Quality assessment. (DOCX) pdig.0001354.s001.docx (69.7KB, docx) S1 Checklist. PRISMA checklist [ 120 ]. (DOCX) pdig.0001354.s002.docx (37.7KB, docx) Acknowledgments The authors utilized ChatGPT-4 in order to rewrite the entire text of the manuscript in terms of correct grammar and wording. In this regarding, the authors validated the clarity and accuracy of the rewritten text. Data Availability The research data is presented as a supplementary file. Funding Statement The author(s) received no specific funding for this work. References 1. Omiye JA, Gui H, Rezaei SJ, Zou J, Daneshjou R. Large language models in medicine: the potentials and pitfalls : a narrative review. Ann Intern Med. 2024;177(2):210–20. doi: 10.7326/M23-2772 [ DOI ] [ PubMed ] [ Google Scholar ] 2. Gencer G, Gencer K. Large language models in healthcare: a bibliometric analysis and examination of research trends. J Multidiscip Healthc. 2025;18:223–38. doi: 10.2147/JMDH.S502351 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Maity S, Saikia MJ. Large language models in healthcare and medical applications: a review. Bioengineering (Basel). 2025;12(6):631. doi: 10.3390/bioengineering12060631 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Zhang Z, Nezhad MJM, Hosseini SMB, Zolnour A, Zonour Z, Hosseini SM, et al. A scoping review of large language model applications in healthcare. Stud Health Technol Inform. 2025;329:1966–7. doi: 10.3233/SHTI251302 [ DOI ] [ PubMed ] [ Google Scholar ] 5. Omar M, Nadkarni GN, Klang E, Glicksberg BS. Large language models in medicine: a review of current clinical trials across healthcare applications. PLOS Digit Health. 2024;3(11):e0000662. doi: 10.1371/journal.pdig.0000662 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Busch F, Hoffmann L, Rueger C, van Dijk EH, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Commun Med (Lond). 2025;5(1):26. doi: 10.1038/s43856-024-00717-2 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Yu E, Chu X, Zhang W, Meng X, Yang Y, Ji X. Large language models in medicine: applications, challenges, and future directions. Int J Med Sci. 2025;22(11):2792–801. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Kim J, Vajravelu BN. Assessing the current limitations of large language models in advancing health care education. JMIR Form Res. 2025;9:e51319. doi: 10.2196/51319 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Qin H, Tong Y. Opportunities and challenges for large language models in primary health care. J Prim Care Community Health. 2025;16:21501319241312571. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Almutairi M, AlKulaib L, Aktas MY, Alsalamah S, Lu CT. Synthetic Arabic medical dialogues using advanced multi-agent LLM techniques. In: ArabicNLP 2024 - 2nd Arabic Natural Language Processing Conference, Proceedings of the Conference. 2024. [ Google Scholar ] 11. Eggmann F, Blatz MB. ChatGPT: chances and challenges for dentistry. Compend Contin Educ Dent. 2023;44(4):220–4. [ PubMed ] [ Google Scholar ] 12. Ge J, Sun S, Owens J, Galvez V, Gologorskaya O, Lai JC, et al. Development of a liver disease-specific large language model chat interface using retrieval-augmented generation. Hepatology. 2024;80(5):1158–68. doi: 10.1097/HEP.0000000000000834 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Lavoie-Gagne OZ, Shen OY, Chen NC, Bhashyam AR. Assessing the usability of ChatGPT responses compared to other online information in hand surgery. Hand (N Y). 2025. doi: 15589447251329584 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Malik A, Ratha N, Yalavarthi B, Sharma T, Kaushik A, Jutla C. Confidential and protected disease classifier using fully homomorphic encryption. In: Proceedings - 2024 IEEE Conference on Artificial Intelligence, CAI 2024. 2024. [ Google Scholar ] 15. Riley G, Wang E, Flynn C, Lopez A, Sridhar A. Evaluating the fidelity of AI-generated information on long-acting reversible contraceptive methods. Eur J Contracept Reprod Health Care. 2025;30(2):74–7. doi: 10.1080/13625187.2025.2450011 [ DOI ] [ PubMed ] [ Google Scholar ] 16. Sovrano F, Ashley K, Bacchelli A. Toward eliminating hallucinations: GPT-based explanatory AI for intelligent textbooks and documentation. In: CEUR Workshop Proceedings. 2023. [ Google Scholar ] 17. Yetkin NA, Baran B, Rabahoğlu B, Tutar N, Gülmez İ. Evaluating the reliability and quality of sarcoidosis-related information provided by AI chatbots. Healthcare (Basel). 2025;13(11):1344. doi: 10.3390/healthcare13111344 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Bull D, Okaygoun D. Evaluating the performance of ChatGPT in the prescribing safety assessment: implications for artificial intelligence-assisted prescribing. Cureus. 2024;16(11):e73003. doi: 10.7759/cureus.73003 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Darji VN, Liao CC, Liao D. Automated interpretation of non-destructive evaluation contour maps using large language models for bridge condition assessment. In: Proceedings - 2024 IEEE International Conference on Big Data, BigData 2024. 2024. [ Google Scholar ] 20. Kalam KT, Rahman JM, Islam MR, Dewan SMR. ChatGPT and mental health: Friends or foes? Health Sci Rep. 2024;7(2):e1912. doi: 10.1002/hsr2.1912 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Bazzari FH, Bazzari AH. Utilizing ChatGPT in telepharmacy. Cureus. 2024;16(1):e52365. doi: 10.7759/cureus.52365 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Lower K, Seth I, Lim B, Seth N. ChatGPT-4: transforming medical education and addressing clinical exposure challenges in the post-pandemic era. Indian J Orthop. 2023;57(9):1527–44. doi: 10.1007/s43465-023-00967-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Singh H, Bajetha I, Pandey A, Singh S, Singh S. From text to treatment: Large language models in clinical practice and medical research. In: Proceedings - 2024 International Conference on Progressive Innovations in Intelligent Systems and Data Science, ICPIDS 2024. 2024. [ Google Scholar ] 24. Yang Z, Yao Z, Tasmin M, Vashisht P, Jang WS, Ouyang F, et al. Unveiling GPT-4V’s hidden challenges behind high accuracy on USMLE questions: observational study. J Med Internet Res. 2025;27:e65146. doi: 10.2196/65146 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Yılmaz IBE, Doğan L. Talking technology: exploring chatbots as a tool for cataract patient education. Clin Exp Optom. 2025;108(1):56–64. doi: 10.1080/08164622.2023.2298812 [ DOI ] [ PubMed ] [ Google Scholar ] 26. Maitland A, Fowkes R, Maitland S. Can ChatGPT pass the MRCP (UK) written examinations? Analysis of performance and errors using a clinical decision-reasoning framework. BMJ Open. 2024;14(3):e080558. doi: 10.1136/bmjopen-2023-080558 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Alshak MN, Cecelic J, Florissi I, Alam R, Cohen AJ. Assessing ChatGPT responses to frequently asked patient questions in reconstructive urology. Urol Pract. 2025;12(4):451–8. doi: 10.1097/UPJ.0000000000000792 [ DOI ] [ PubMed ] [ Google Scholar ] 28. Franc JM, Hertelendy AJ, Cheng L, Hata R, Verde M. Accuracy of a commercial large language model (ChatGPT) to perform disaster triage of simulated patients using the simple triage and rapid treatment (START) protocol: gage repeatability and reproducibility study. J Med Internet Res. 2024;26:e55648. doi: 10.2196/55648 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Gaebe K, van der Woerd B. Evaluation of large language models as a diagnostic tool for medical learners and clinicians using advanced prompting techniques. PLoS One. 2025;20(8):e0325803. doi: 10.1371/journal.pone.0325803 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Huang Y, Wang W, Zhou J, Zhang L, Lin J, Liu H, et al. Integrative modeling enables ChatGPT to achieve average level of human counselors performance in mental health Q&A. Inform Process Manage. 2025;62(5). [ Google Scholar ] 31. Incerti Parenti S, Bartolucci ML, Biondi E, Maglioni A, Corazza G, Gracco A, et al. Online patient education in obstructive sleep apnea: ChatGPT versus google search. Healthcare (Basel). 2024;12(17):1781. doi: 10.3390/healthcare12171781 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Joshi S, Ha E, Amaya A, Mendoza M, Rivera Y, Singh VK. Ensuring accuracy and equity in vaccination information from ChatGPT and CDC: mixed-methods cross-language evaluation. JMIR Form Res. 2024;8:e60939. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Kıyak YS, Kononowicz AA. Case-based MCQ generator: a custom ChatGPT based on published prompts in the literature for automatic item generation. Med Teach. 2024;46(8):1018–20. doi: 10.1080/0142159X.2024.2314723 [ DOI ] [ PubMed ] [ Google Scholar ] 34. Lee TJ, Campbell DJ, Rao AK, Hossain A, Elkattawy O, Radfar N, et al. Evaluating ChatGPT responses on atrial fibrillation for patient education. Cureus. 2024;16(6):e61680. doi: 10.7759/cureus.61680 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Patel A, Ajumobi A. Evaluating the reliability of OpenAI’s ChatGPT-4 in providing pre-colonoscopy patient guidance. Cureus. 2025;17(6):e86512. doi: 10.7759/cureus.86512 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Stephan D, Bertsch A, Burwinkel M, Vinayahalingam S, Al-Nawas B, Kämmerer PW, et al. AI in dental radiology-improving the efficiency of reporting with ChatGPT: comparative study. J Med Internet Res. 2024;26:e60684. doi: 10.2196/60684 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Theophilou E, Koyutürk C, Yavari M, Bursic S, Donabauer G, Telari A, et al. Learning to prompt in the classroom to understand AI limits: a pilot study. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). Springer Nature Switzerland; 2023. pp. 481–96. doi: 10.1007/978-3-031-47546-7_33 [ DOI ] [ Google Scholar ] 38. Vikan M, Aryan R, Kannelønning MS, Riegler MA, Danielsen SO. Reflecting on LLM support in reflexive thematic analysis: an exploratory study. Qual Health Res. 2026;36(2–3):191–205. doi: 10.1177/10497323251365211 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Balas M, Wong DT, Arshinoff SA. Artificial intelligence, adversarial attacks, and ocular warfare. AJO Int. 2024;1(3):100062. doi: 10.1016/j.ajoint.2024.100062 [ DOI ] [ Google Scholar ] 40. Erkan EE, Kömürcü MA, Çelikten T, Ergün AE, Onan A. Understanding large language model performance in question answering: A comparative analysis of semantic and lexical metrics. Lecture Notes in Networks and Systems. 2025. [ Google Scholar ] 41. Jeyaraman M, K SP, Jeyaraman N, Nallakumarasamy A, Yadav S, Bondili SK. ChatGPT in medical education and research: a Boon or a Bane? Cureus. 2023;15(8):e44316. doi: 10.7759/cureus.44316 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Koh MCY, Ngiam JN, Yong J, Tambyah PA, Archuleta S. The role of an artificial intelligence model in antiretroviral therapy counselling and advice for people living with HIV. HIV Med. 2024;25(4):504–8. doi: 10.1111/hiv.13604 [ DOI ] [ PubMed ] [ Google Scholar ] 43. Temsah A, Alhasan K, Altamimi I, Jamal A, Al-Eyadhy A, Malki KH. DeepSeek in healthcare: revealing opportunities and steering challenges of a new open-source artificial intelligence frontier. Cureus. 2025;17(2):e79221. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 44. Gupta E, Gupta V. A Comparative Analysis of ESP32 and ESP8266 for AI-Powered Applications. In: 2025 International Conference on Next Generation Communication & Information Processing (INCIP). 2025. 992–6. doi: 10.1109/incip64058.2025.11019649 [ DOI ] [ Google Scholar ] 45. Zaman M. ChatGPT for healthcare sector: SWOT analysis. Int J Res Ind Eng. 2023;12(3):221–33. [ Google Scholar ] 46. Bentzen SM. Artificial intelligence in health care: a rallying cry for critical clinical research and ethical thinking. Clin Oncol (R Coll Radiol). 2025;41:103798. doi: 10.1016/j.clon.2025.103798 [ DOI ] [ PubMed ] [ Google Scholar ] 47. Haider SA, Prabha S, Gomez-Cabello CA, Borna S, Genovese A, Trabilsy M. Synthetic patient-physician conversations simulated by large language models: a multi-dimensional evaluation. Sensors (Basel). 2025;25(14). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 48. Joerg L, Kabakova M, Wang JY, Austin E, Cohen M, Kurtti A, et al. AI-generated dermatologic images show deficient skin tone diversity and poor diagnostic accuracy: an experimental study. J Eur Acad Dermatol Venereol. 2025;39(12):2134–41. doi: 10.1111/jdv.20849 [ DOI ] [ PubMed ] [ Google Scholar ] 49. Kaya Kaçar H, Kaçar ÖF, Avery A. Diet quality and caloric accuracy in AI-generated diet plans: a comparative study across chatbots. Nutrients. 2025;17(2). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Kral J, Hradis M, Buzga M, Kunovsky L. Exploring the benefits and challenges of AI-driven large language models in gastroenterology: think out of the box. Biomed Pap Med Fac Univ Palacky Olomouc Czech Repub. 2024;168(4):277–83. doi: 10.5507/bp.2024.027 [ DOI ] [ PubMed ] [ Google Scholar ] 51. Morath B, Chiriac U, Jaszkowski E, Deiß C, Nürnberg H, Hörth K, et al. Performance and risks of ChatGPT used in drug information: an exploratory real-world analysis. Eur J Hosp Pharm. 2024;31(6):491–7. doi: 10.1136/ejhpharm-2023-003750 [ DOI ] [ PubMed ] [ Google Scholar ] 52. Nguyen T. ChatGPT in medical education: a precursor for automation bias? JMIR Med Educ. 2024;10:e50174. doi: 10.2196/50174 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 53. Tarris G, Martin L. Performance assessment of ChatGPT 4, ChatGPT 3.5, Gemini Advanced Pro 1.5 and Bard 2.0 to problem solving in pathology in French language. Digit Health. 2025;11. doi: 10.1177/20552076241310630 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 54. Urbina JT, Vu PD, Nguyen MV. Disability ethics and education in the age of artificial intelligence: identifying ability bias in ChatGPT and Gemini. Arch Phys Med Rehabil. 2025;106(1):14–9. doi: 10.1016/j.apmr.2024.08.014 [ DOI ] [ PubMed ] [ Google Scholar ] 55. Daza J, Bezerra LS, Santamaría L, Rueda-Esteban R, Bantel H, Girala M, et al. Evaluation of four chatbots in autoimmune liver disease: a comparative analysis. Ann Hepatol. 2025;30(1):101537. doi: 10.1016/j.aohep.2024.101537 [ DOI ] [ PubMed ] [ Google Scholar ] 56. Dergaa I, Fekih-Romdhane F, Hallit S, Loch AA, Glenn JM, Fessi MS, et al. ChatGPT is not ready yet for use in providing mental health assessment and interventions. Front Psychiatry. 2024;14:1277756. doi: 10.3389/fpsyt.2023.1277756 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 57. Şahin MF, Keleş A, Özcan R, Doğan Ç, Topkaç EC, Akgül M, et al. Evaluation of information accuracy and clarity: ChatGPT responses to the most frequently asked questions about premature ejaculation. Sex Med. 2024;12(3):qfae036. doi: 10.1093/sexmed/qfae036 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 58. Rodgers DL, Needler M, Robinson A, Barnes R, Brosche T, Hernandez J. Artificial Intelligenceand and the Simulationists. Simul Healthc. 2023;18(6):395–9. [ DOI ] [ PubMed ] [ Google Scholar ] 59. Huang W, Wei W, He X, Zhan B, Xie X, Zhang M, et al. ChatGPT-assisted deep learning models for influenza-like illness prediction in mainland China: time series analysis. J Med Internet Res. 2025;27:e74423. doi: 10.2196/74423 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 60. Jin Y, Liu H, Zhao B, Pan W. ChatGPT and mycosis- a new weapon in the knowledge battlefield. BMC Infect Dis. 2023;23(1):731. doi: 10.1186/s12879-023-08724-9 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 61. Ali K, Barhom N, Tamimi F, Duggal M. ChatGPT-A double-edged sword for healthcare education? Implications for assessments of dental students. Eur J Dent Educ. 2024;28(1):206–11. doi: 10.1111/eje.12937 [ DOI ] [ PubMed ] [ Google Scholar ] 62. Alkuraya IF. Is artificial intelligence getting too much credit in medical genetics? Am J Med Genet C Semin Med Genet. 2023;193(3):e32062. doi: 10.1002/ajmg.c.32062 [ DOI ] [ PubMed ] [ Google Scholar ] 63. Cankurtaran RE, Polat YH, Aydemir NG, Umay E, Yurekli OT. Reliability and usefulness of ChatGPT for inflammatory bowel diseases: an analysis for patients and healthcare professionals. Cureus. 2023;15(10):e46736. doi: 10.7759/cureus.46736 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 64. Chen J, Cadiente A, Kasselman LJ, Pilkington B. Assessing the performance of ChatGPT in bioethics: a large language model’s moral compass in medicine. J Med Ethics. 2024;50(2):97–101. doi: 10.1136/jme-2023-109366 [ DOI ] [ PubMed ] [ Google Scholar ] 65. Cuthbert R, Simpson AI. Artificial intelligence in orthopaedics: can Chat Generative Pre-trained Transformer (ChatGPT) pass Section 1 of the Fellowship of the Royal College of Surgeons (Trauma & Orthopaedics) examination? Postgrad Med J. 2023;99(1176):1110–4. doi: 10.1093/postmj/qgad053 [ DOI ] [ PubMed ] [ Google Scholar ] 66. Geneş M, Yaşar S, Fırtına S, Yağcı AF, Yıldırım E, Barçın C, et al. Artificial intelligence in cardiac rehabilitation: assessing ChatGPT’s knowledge and clinical scenario responses. Turk Kardiyol Dern Ars. 2025;53(3):173–7. doi: 10.5543/tkda.2025.57195 [ DOI ] [ PubMed ] [ Google Scholar ] 67. Hristidis V, Ruggiano N, Brown EL, Ganta SRR, Stewart S. ChatGPT vs Google for queries related to dementia and other cognitive decline: comparison of results. J Med Internet Res. 2023;25:e48966. doi: 10.2196/48966 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 68. Khosravi T, Al Sudani ZM, Oladnabi M. To what extent does ChatGPT understand genetics? Innov Educ Teach Int. 2024;61(6):1320–9. [ Google Scholar ] 69. Lang S, Vitale J, Galbusera F, Fekete T, Boissiere L, Charles YP, et al. Is the information provided by large language models valid in educating patients about adolescent idiopathic scoliosis? An evaluation of content, clarity, and empathy : The perspective of the European Spine Study Group. Spine Deform. 2025;13(2):361–72. doi: 10.1007/s43390-024-00955-3 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 70. Read R, Lukies M. The quality of information produced by ChatGPT about conditions managed by interventional radiologists. J Med Imaging Radiat Oncol. 2025;69(7):715–24. doi: 10.1111/1754-9485.13881 [ DOI ] [ PubMed ] [ Google Scholar ] 71. Singh S, Errampalli E, Errampalli N, Miran MS. Enhancing patient education on cardiovascular rehabilitation with large language models. Missouri Med. 2025;122(1):67–71. [ PMC free article ] [ PubMed ] [ Google Scholar ] 72. Sparks CA, Kraeutler MJ, Chester GA, Contrada EV, Zhu E, Fasulo SM. Inadequate performance of ChatGPT on orthopedic board-style written exams. Cureus. 2024;16(6):e62643. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 73. Zheng Y, Wu Y, Feng B, Wang L, Kang K, Zhao A. Enhancing diabetes self-management and education: a critical analysis of ChatGPT’s role. Ann Biomed Eng. 2024;52(4):741–4. doi: 10.1007/s10439-023-03317-8 [ DOI ] [ PubMed ] [ Google Scholar ] 74. Alkuraya IF. Is artificial intelligence getting too much credit in medical genetics? Am J Med Genet Part C: Seminars Med Genet. 2023;193(3). [ DOI ] [ PubMed ] [ Google Scholar ] 75. Grossman S, Zerilli T, Nathan JP. Appropriateness of ChatGPT as a resource for medication-related questions. Br J Clin Pharmacol. 2024;90(10):2691–5. doi: 10.1111/bcp.16212 [ DOI ] [ PubMed ] [ Google Scholar ] 76. Halaseh FF, Yang JS, Danza CN, Halaseh R, Spiegelman L. ChatGPT’s role in improving education among patients seeking emergency medical treatment. West J Emerg Med. 2024;25(5):845–55. doi: 10.5811/westjem.18650 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 77. Wang Y, Chen Y, Sheng J. Assessing ChatGPT as a medical consultation assistant for chronic hepatitis B: cross-language study of English and Chinese. JMIR Med Inform. 2024;12:e56426. doi: 10.2196/56426 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 78. Aliyeva A, Alaskarov E, Sari E. Postoperative management of tympanoplasty with ChatGPT-4.0. J Int Adv Otol. 2025;21(1):1–6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 79. Koh SJQ, Yeo KK, Yap JJ-L. Leveraging ChatGPT to aid patient education on coronary angiogram. Ann Acad Med Singap. 2023;52(7):374–7. doi: 10.47102/annals-acadmedsg.2023138 [ DOI ] [ PubMed ] [ Google Scholar ] 80. Kong M, Fernandez A, Bains J, Milisavljevic A, Brooks KC, Shanmugam A, et al. Evaluation of the accuracy and safety of machine translation of patient-specific discharge instructions: a comparative analysis. BMJ Qual Saf. 2026;35(3):150–8. doi: 10.1136/bmjqs-2024-018384 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 81. Soroudi D, Gozali A, Knox JA, Parmeshwar N, Sadjadi R, Wilson JC, et al. Comparing provider and ChatGPT responses to breast reconstruction patient questions in the electronic health record. Ann Plast Surg. 2024;93(5):541–5. doi: 10.1097/SAP.0000000000004090 [ DOI ] [ PubMed ] [ Google Scholar ] 82. Zeller NP, Shah AD, Van Heest AE, Bohn DC. Assessing accuracy of chat generative pre-trained transformer’s responses to common patient questions regarding congenital upper limb differences. J Hand Surg Glob Online. 2025;7(4):100764. doi: 10.1016/j.jhsg.2025.100764 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 83. Buhr CR, Smith H, Huppertz T, Bahr-Hamm K, Matthias C, Cuny C, et al. Assessing unknown potential-quality and limitations of different large language models in the field of otorhinolaryngology. Acta Otolaryngol. 2024;144(3):237–42. doi: 10.1080/00016489.2024.2352843 [ DOI ] [ PubMed ] [ Google Scholar ] 84. Elpasiony NMA, Sabek EM, Ibrahim SSM. Chat generative pre-trained transformers era: pros and cons between nursing researchers in Egypt. BMC Nurs. 2025;24(1):667. doi: 10.1186/s12912-025-03332-1 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 85. Gill B, Bonamer J, Kuechly H, Gupta R, Emmert S, Kurkowski S, et al. ChatGPT is a promising tool to increase readability of orthopedic research consents. J Orthopaed Trauma Rehabil. 2024;31(2):148–52. doi: 10.1177/22104917231208212 [ DOI ] [ Google Scholar ] 86. Yaş S, Yapar D, Yapar A, Özel T, Tokgöz MA, Baymurat AC, et al. Assessing the role of large language models in adolescent idiopathic scoliosis care: a comparison between ChatGPT and Google Gemini. Acta Orthop Traumatol Turc. 2025;59(4):222–9. doi: 10.5152/j.aott.2025.25279 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 87. Asfuroğlu ZM, Yağar H, Gümüşoğlu E. High accuracy but limited readability of large language model-generated responses to frequently asked questions about Kienböck’s disease. BMC Musculoskelet Disord. 2024;25(1):879. doi: 10.1186/s12891-024-07983-0 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 88. Gürbostan Soysal G, Mercanlı M, Özcan ZÖ, Yılmaz İE, Berhuni M. Evaluating the effectiveness of chatbots and traditional resources in patient education on dry eye disease. Clin Exp Optom. 2026;109(2):182–6. doi: 10.1080/08164622.2025.2517750 [ DOI ] [ PubMed ] [ Google Scholar ] 89. Brewster RCL, Gonzalez P, Khazanchi R, Butler A, Selcer R, Chu D, et al. Performance of ChatGPT and Google translate for pediatric discharge instruction translation. Pediatrics. 2024;154(1):e2023065573. doi: 10.1542/peds.2023-065573 [ DOI ] [ PubMed ] [ Google Scholar ] 90. Fraile Navarro D, Coiera E, Hambly TW, Triplett Z, Asif N, Susanto A, et al. Expert evaluation of large language models for clinical dialogue summarization. Sci Rep. 2025;15(1):1195. doi: 10.1038/s41598-024-84850-x [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 91. Tam W, Huynh T, Tang A, Luong S, Khatri Y, Zhou W. Nursing education in the age of artificial intelligence powered chatbots (AI-chatbots): Are we ready yet? Nurse Educ Today. 2023;129:105917. [ DOI ] [ PubMed ] [ Google Scholar ] 92. Wang L, Li J, Zhuang B, Huang S, Fang M, Wang C, et al. Accuracy of large language models when answering clinical research questions: systematic review and network meta-analysis. J Med Internet Res. 2025;27:e64486. doi: 10.2196/64486 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 93. Shan G, Chen X, Wang C, Liu L, Gu Y, Jiang H, et al. Comparing diagnostic accuracy of clinical professionals and large language models: systematic review and meta-analysis. JMIR Med Inform. 2025;13:e64963. doi: 10.2196/64963 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 94. Pal A, Wangmo T, Bharadia T, Ahmed-Richards M, Bhanderi MB, Kachhadiya R, et al. Generative AI/LLMs for plain language medical information for patients, caregivers and general public: opportunities, risks and ethics. Patient Prefer Adherence. 2025;19:2227–49. doi: 10.2147/PPA.S527922 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 95. Serugunda HM, Jianquan O, Kasujja Namatovu H, Ssemaluulu P, Kimbugwe N, Garimoi Orach C, et al. Using large language models for chronic disease management tasks: scoping review. JMIR Med Inform. 2025;13:e66905. doi: 10.2196/66905 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 96. Unger Z, Soffer S, Efros O, Chan L, Klang E, Nadkarni GN. Clinical applications and limitations of large language models in nephrology: a systematic review. Clin Kidney J. 2025;18(9):sfaf243. doi: 10.1093/ckj/sfaf243 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 97. Chen Y-C, Lee S-H, Sheu H, Lin S-H, Hu C-C, Fu S-C, et al. Enhancing responses from large language models with role-playing prompts: a comparative study on answering frequently asked questions about total knee arthroplasty. BMC Med Inform Decis Mak. 2025;25(1):196. doi: 10.1186/s12911-025-03024-5 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 98. Fink A, Rau A, Kotter E, Bamberg F, Russe MF. Optimized interaction with large language models : a practical guide to prompt engineering and retrieval-augmented generation. Radiologie (Heidelb). 2025;65(4):235–42. doi: 10.1007/s00117-025-01416-2 [ DOI ] [ PubMed ] [ Google Scholar ] 99. He Z, Bhasuran B, Jin Q, Tian S, Hanna K, Shavor C, et al. Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study. J Med Internet Res. 2024;26:e56655. doi: 10.2196/56655 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 100. Ullah E, Parwani A, Baig MM, Singh R. Challenges and barriers of using large language models (LLM) such as ChatGPT for diagnostic medicine with a focus on digital pathology - a recent scoping review. Diagn Pathol. 2024;19(1):43. doi: 10.1186/s13000-024-01464-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 101. Bernier A, Molnár-Gábor F, Knoppers BM. The international data governance landscape. J Law Biosci. 2022;9(1):lsac005. doi: 10.1093/jlb/lsac005 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 102. Khosravi M, Izadi R, Aghamaleki Sarvestani M, Bouzarjomehri H, Ahmadi Marzaleh M, Ravangard R. Performance of artificial intelligence large language models (Copilot and Gemini) compared to human experts in healthcare policy making: A mixed-methods cross-sectional study. Health Informatics J. 2025;31(3). doi: 10.1177/14604582251381269 40977570 [ DOI ] [ Google Scholar ] 103. Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA. 2025;333(4):319–28. doi: 10.1001/jama.2024.21700 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 104. Maiorana NV, Marceglia S, Treddenti M, Tosi M, Guidetti M, Creta MF. Large language models in neurological practice: real-world study. J Med Internet Res. 2025;27:e73212. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 105. Gondara L, Simkin J, Devji S. Clinical trial design approach to auditing language models in health care setting. JCO Clin Cancer Inform. 2025;9:e2400331. doi: 10.1200/CCI-24-00331 [ DOI ] [ PubMed ] [ Google Scholar ] 106. Esmaeilzadeh P. Ethical implications of using general-purpose LLMs in clinical settings: a comparative analysis of prompt engineering strategies and their impact on patient safety. BMC Med Inform Decis Mak. 2025;25(1):342. doi: 10.1186/s12911-025-03182-6 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 107. Farnós J, Sans Pinillos A, Costa V. Ethical prompting: toward strategies for rapid and inclusive assistance in dual-use AI systems. Front Artif Intell. 2025;8:1646444. doi: 10.3389/frai.2025.1646444 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 108. Pham T. Ethical and legal considerations in healthcare AI: innovation and policy for safe and fair use. R Soc Open Sci. 2025;12(5):241873. doi: 10.1098/rsos.241873 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 109. Alber DA, Yang Z, Alyakin A, Yang E, Rai S, Valliani AA, et al. Medical large language models are vulnerable to data-poisoning attacks. Nat Med. 2025;31(2):618–26. doi: 10.1038/s41591-024-03445-1 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 110. Addula SR, Meesala MK, Ravipati P, Sajja GS. A hybrid autoencoder and gated recurrent unit model optimized by honey badger algorithm for enhanced cyber threat detection in IoT networks. Security Privacy. 2025;8(6). doi: 10.1002/spy2.70086 [ DOI ] [ Google Scholar ] 111. Gokcimen T, Das B. A novel system for strengthening security in large language models against hallucination and injection attacks with effective strategies. Alexandria Eng J. 2025;123:71–90. doi: 10.1016/j.aej.2025.03.030 [ DOI ] [ Google Scholar ] 112. Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172–80. doi: 10.1038/s41586-023-06291-2 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 113. Luo R, Sun L, Xia Y, Qin T, Zhang S, Poon H, et al. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Brief Bioinform. 2022;23(6):bbac409. doi: 10.1093/bib/bbac409 [ DOI ] [ PubMed ] [ Google Scholar ] 114. Jin Q, Dhingra B, Liu Z, Cohen W, Lu X. PubMedQA: A Dataset for Biomedical Research Question Answering. Hong Kong, China: Association for Computational Linguistics; 2019. [ Google Scholar ] 115. Jin D, Pan E, Oufattole N, Weng WH, Fang H, Szolovits P. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Appl Sci. 2021;11(14):6421. [ Google Scholar ] 116. Abbas ASH, Ong AY, Antaki F, Akhtar HN, Shehab M, Keane PA. An updated analysis of large language model performance on ophthalmology speciality examinations. Eye. 2026;40(5):572–4. doi: 10.1038/s41433-026-04262-1 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 117. Genç A, Aydın GR, Kasap M, Özkan A, Rakhymzhanov A, Gürcan HH, et al. Comparative evaluation of ChatGPT versions in training program design: scientific approach, accuracy, and practical applicability. BMC Sports Sci Med Rehabil. 2025;18(1):19. doi: 10.1186/s13102-025-01409-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 118. Sheikhalishahi S, Haddadi A, Sadeghipour S, Rafiei F, Soltani H. Comparative performance of ChatGPT-4o, ChatGPT-5, and gemini 2.5 flash on Persian internal medicine subspecialty board exams. Sci Rep. 2025;16(1):1371. doi: 10.1038/s41598-025-31251-3 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 119. Liu M, Okuhara T, Chang X, Shirabe R, Nishiie Y, Okada H, et al. Performance of ChatGPT across different versions in medical licensing examinations worldwide: systematic review and meta-analysis. J Med Internet Res. 2024;26:e60807. doi: 10.2196/60807 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 120. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi: 10.1136/bmj.n71 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 121. Barroga E, Matanguihan GJ. A practical guide to writing quantitative and qualitative research questions and hypotheses in scholarly articles. J Korean Med Sci. 2022;37(16):e121. doi: 10.3346/jkms.2022.37.e121 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 122. Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt J-N, Laleh NG, et al. The future landscape of large language models in medicine. Commun Med (Lond). 2023;3(1):141. doi: 10.1038/s43856-023-00370-1 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 123. Tyndall J. AACODS (Authority, Accuracy, Coverage, Objectivity, Date, Significance) Checklist. 2010. Available from: http://dspaceflinderseduau/dspace [ Google Scholar ] 124. Boyatzis RE. Transforming qualitative information: thematic analysis and code development. Sage; 1998. [ Google Scholar ] 125. Grady JO. System engineering planning and enterprise identity. Crc Press; 1995. [ Google Scholar ] 126. Goel A. Computer fundamentals. Pearson Education India; 2010. [ Google Scholar ] 127. Zelle JM. Python programming: an introduction to computer science. Franklin, Beedle & Associates, Inc.; 2004. [ Google Scholar ] 128. Curry A, Flett P, Hollingsworth I. Managing information & systems: The business perspective: Routledge; 2006. [ Google Scholar ] 129. Braziller G. General system theory. Foundations, development, applications. Ludwig von Bertalanffy Prieiga internetu; 1968. [ Google Scholar ] 130. Forero R, Nahidi S, De Costa J, Mohsin M, Fitzgerald G, Gibson N, et al. Application of four-dimension criteria to assess rigour of qualitative research in emergency medicine. BMC Health Serv Res. 2018;18(1):120. doi: 10.1186/s12913-018-2915-2 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] PLOS Digit Health. doi: 10.1371/journal.pdig.0001354.r001 Decision Letter 0 Mayue Shi Mayue Shi Academic Editor Find articles by Mayue Shi Author information Copyright and License information Roles Mayue Shi : Academic Editor © 2026 Mayue ShiMayue ShiMayue ShiMayue Shi This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 28 Jan 2026 Response to Reviewers '. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.'. This file does not need to include responses to any formatting updates and technical items listed in the 'Journal Requirements' section below.-->-->* A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled ' Revised Manuscript with Track Changes '.'.-->-->* An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled ' Manuscript '.'.-->--> -->-->If you would like to make changes to your financial disclosure, competing interests statement, or data availability statement, please make these updates within the submission form at the time of resubmission. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.-->--> -->-->We look forward to receiving your revised manuscript.-->--> -->-->Kind regards,-->--> -->-->Mayue Shi, Ph.D.-->-->Academic Editor-->-->PLOS Digital Health-->--> -->-->Leo Anthony Celi-->-->Editor-in-Chief-->-->PLOS Digital Health-->-->orcid.org/0000-0001-6712-6626-->--> -->--> Journal Requirements: -->--> -->-->If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. -->--> -->--> Additional Editor Comments (if provided): -->--> -->-->This is a timely systematic review that addresses an important and rapidly evolving topic in healthcare AI. The review is comprehensive and well motivated; however, several key issues should be addressed to strengthen its rigor and clarity. I suggest authors addressing reviewers' concerns accrodingly. Methodological transparency could be improved by providing reproducible, database-specific search strategies, clarifying quality and bias assessment methods, and importantly, justifying exclusion of conference literature. The discussion would benefit from deeper analytical integration.-->--> -->-->[Note: HTML markup is below. Please do not edit.]-->--> -->--> Reviewers' Comments: -->--> -->-->Reviewer's Responses to Questions Comments to the Author 1. Does this manuscript meet PLOS Digital Health’s publication criteria ? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.-->?> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** 2. Has the statistical analysis been performed appropriately and rigorously?-->?> Reviewer #1: Yes Reviewer #2: N/A Reviewer #3: Yes ********** 3. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** Reviewer #1: 1. This study only includes articles published in English. This is a significant limitation for a systematic review because valuable research in other languages, such as Chinese, German, and Arabic, is important in artificial intelligence and healthcare. Excluding these sources may skew our understanding of global issues, especially regarding translation quality. It is crucial to reassess this limitation or at least discuss its effects clearly. 2. The Discussion section primarily repeats the results without fully exploring the connections between the input, process, and output limitations. A more in-depth analysis is needed to understand how input limitations, like a lack of Arabic datasets, lead to specific output problems, such as variable translation quality. 3. In the Conclusion and Limitations section (Section 4), it is important to state clearly that reliance on English publications is a major limitation. This focus could distort the findings and underreport language-specific biases. Future research should address this issue by including studies in other languages to improve the generalizability of the findings. 4. The authors should better emphasize the connection between input flaws and output problems. The text states that output issues "mainly resulted from flaws in input data," but it should elaborate on this point. For example, one could state, "Limited multimodal data processing (input flaw) leads to incomplete responses and neglected imaging descriptions (output limitation) in radiology contexts. "A Hybrid Autoencoder and Gated Recurrent Unit Model Optimized by Honey Badger Algorithm for Enhanced Cyber Threat Detection in IoT Networks”. This reference discusses a hybrid deep learning model in the context of IoT security and relates directly to the limitations in the review of large language models (LLMs). Specifically: 5. It shows the need for specialized hybrid architectures and optimization algorithms, in contrast to the general-purpose LLMs discussed. This reference highlights a successful application in a specific field (cyber threat detection), presenting a counterexample to the weaknesses of general LLMs in complex fields such as healthcare. 6. The need to optimize parameters reflects the LLM issue regarding reliance on general training, which leads to gaps in accuracy and biases.This reference should be included in the Discussion section while analyzing the relationship between output limitations and flaws in model design. 7. The text notes that large language models (LLMs) are prone to generating plausible but incorrect or fabricated information, a phenomenon known as It is important to emphasize this term in bold for clarity. 8. In Section 6.5, it is helpful to reference the original work that defines the IPO model to maintain academic rigor. Although the model is introduced, clarifying its ties to system analysis will improve understanding. It is suggested that the authors check the alignments in the reference section and front issues and correct anything if needed. Reviewer #2: Summary This manuscript presents a systematic review (2018–2025) that synthesizes the limitations of large language models (LLMs) in generating healthcare content, using the input–process–output (IPO) framework. The authors searched PubMed, Scopus, and Cochrane, identified 81 studies, assessed them using the AACODS checklist, and performed an inductive thematic analysis following Boyatzis’s method. The study identifies eight themes relating to data limitations, prompt/input dependence, architectural constraints, interaction challenges, output quality limitations, and ethical/regulatory risks. The authors conclude that although output errors are the most commonly cited limitations, many originate from input-quality constraints, and they propose strategies to mitigate these issues. Major Comments The search strategy is overly minimal (“Limitations” AND “Large language models” AND “Healthcare”), and Table 1 lists broad synonym groups without showing the actual Boolean-structured queries used in each database. For reproducibility, PRISMA requires database-specific queries (e.g., PubMed MeSH lines). There is no mention of PROSPERO registration. Given the systematic review methodology, the absence of protocol registration raises concerns about potential bias and post-hoc modifications. The manuscript excludes conference papers and letters but does not justify these exclusions. Many influential AI/LLM limitations papers appear in conference proceedings (ACL, NeurIPS, AAAI). While the IPO model is a general systems-analysis tool, its adoption here seems partially post-hoc; some limitations overlap categories (e.g., prompt issues can be both input and process). Consider clarifying the rationale for mapping each limitation to its category. The results extend to many pages with repetitive descriptions of similar limitations. While comprehensive, it compromises readability. A more synthesized narrative plus thematic tables would improve clarity. The figure 2 displays frequency counts but lacks: definition of how “number of references” was computed (e.g., per theme or per sub-theme), visual clarity (font size, grouping). Much of the discussion repeats results rather than integrating them with broader literature or offering conceptual insight. The section could more deeply analyze why certain limitations dominate and how emerging LLM architectures (e.g., multimodal models, RAG, domain-specific pretraining) may mitigate them. Recent advances such as GPT-5, Gemini 2.0, Qwen 2.5, smaller <13B models like Gemma and domain-specific models (Med-PaLM, MedGemma, LLaMA-family medical finetunes) are not discussed, though their capabilities meaningfully impact the identified limitations. The study acknowledges language restriction but omits several important limitations: no assessment of publication bias,heterogeneity across study types (chatbot assessments, narrative essays, technical papers), and limited granularity in separating limitations specific to clinical, educational, and administrative contexts Numerous grammatical inconsistencies remain (e.g., “LLMS” instead of LLMs; “supposedly to be derived…”).\ The abstract has several duplicated or redundant phrases. Improve flow by reducing overly long sentences throughout the manuscript. Ensure consistent referencing formatting (several citations appear as ranges with missing formatting). Citation [75] is listed as “!!! INVALID CITATION !!!” which must be corrected. 75. . !!! INVALID CITATION !!! [4, 5, 12, 24, 29, 30, 41, 43, 51, 57, 59, 62, 67-70]. Table 2 is overly long and may be better placed in an appendix. Some themes overlap significantly (e.g., “dependence on input quality” vs “dependence on prompt quality”). Consider merging or refining boundaries. Terminology such as “accuracy gap” should be defined more precisely. The authors mention using ChatGPT-4 to “rewrite the text,” which is appropriate, but PLOS typically requires specifying: which sections were rewritten, whether human authors verified accuracy. Reviewer #3: This is a great work. It is timely and important, and it has a lot of present and future implications on both healthcare practice and research. The coverage and classification of the LLM limitations and the awareness about them are essential for different backgrounds and academic levels of people involved in designing, implementing and end-using them. Going forward, this subject can not be avoided and no matter what the limitations and challenges are, we have to face them, deal with them and tame them. Thanks for sharing this important work and thanks for giving me the opportunity to review it. I learned a lot from reading it. The researchers have summarized many ideas from the literature that they have reviewed, that I would not have time to read it myself and add to my understanding. So I have to be extremely thankful. One of the injustices that afflict this study is that it is tackling a huge subject. While reading it, at some points or sub-topics, I found myself thinking that every one limitation or challenge mentioned here is worth a separate study because of the depth of the ideas discussed. I would like to put my comments in the following bullets, following quoting the manuscript text portions hoping that this will help in shaping the final publication in the best shape. • “Two independent evaluators screened the references and assessed quality of the selected studies using the Authority, Accuracy, Coverage, Objectivity, Date, and Significance (AACODS) checklist, which examines Authority, Accuracy, Coverage, Objectivity, Date, and Significance.” In this sentence, the evaluators assessed the Authority, Accuracy etc of what exactly? For readers with different backgrounds, the second subject of this sentence needs to be clearer. • “Boyatzis's qualitative thematic approach” again, for any reader, they would expect some elaboration on this approach anywhere in the manuscript. The approach has been mentioned four times, including the one in the references without at least a single line simple explanation about it. • “A total of 81 studies were included in the final analysis. The included studies were predominantly of high quality and demonstrated minimal risk of bias.” Looking for the methodology used to assess risk of bias. • “The thematic analysis identified key themes: data limitations, dependence on input and prompt quality, accessibility issues, model design and architecture constraints, interaction challenges, response quality and comprehensiveness, and ethical, safety, and regulatory concerns.” Tables? I would suggest that multiple tables should replace the huge Table 2, to ensure that the text is shown in correspondence with the relevant table contents. • “The study identified multiple limitations of LLMs in healthcare, with output issues being most common.” Reading through the text and the table, I found it necessary to repeat the point of dividing Table 2 into more tables, and this might bring even more subtopics that deserve to be included in the tables.? • “However, these output problems mainly resulted from flaws in input data, emphasizing the crucial 60 role of input quality. The study also proposed strategies to address these challenges.” Table? • “The search yielded 81 studies, and the quality of the included studies was presented to be predominantly high.” Quality assessment criteria? • “The thematic analysis yielded a number of themes and sub-themes.” Tables? • “The results of the research found that while the main area of LLMs` limitations is corresponding to the outputs, and particularly the existing gaps in accuracy, these limitations are supposedly to be derived from the existing flaws in the input data.” Important conclusion to be emphasized in the discussion. • “The study also presented some strategies to overcome these limitations 80 based on the existing data within the literature.” Important future plan or recommendation. • “Large language models (LLMs) are sophisticated artificial intelligence (AI) systems developed through extensive training on vast corpora of text data, enabling them to generate outputs that closely resemble human language(1). These models have been widely utilized across diverse medical domains, including health informatics, medical imaging, clinical diagnostics, treatment planning, ophthalmology, oncology, and other specialized fields(2). This trend signifies their broad and growing integration into medical research and clinical practice. LLMs have become pivotal in healthcare by enhancing clinical decision support, diagnostics, medical education, and patient engagement(2-4). They improve diagnostic accuracy by analyzing extensive clinical data and medical literature, aiding in personalized treatment planning and patient care management(3).” Debatable, but understood as completely dependant on the reference. • “98 Notwithstanding the numerous advantages previously discussed, LLMs in healthcare exhibit 99 several critical limitations that must be addressed to ensure their safe and effective deployment. 100 In this regard, they are prone to generating plausible yet factually incorrect or fabricated 101 information, a phenomenon known as hallucination, which poses significant risks in clinical 102 settings(3, 6).” Table for limitations, hallucinations included. • “Furthermore, LLMs often lack the depth of contextual understanding required to 103 accurately interpret complex medical scenarios, as they may fail to integrate multifaceted clinical 104 data or temporal information adequately(3).” Other limitations, lack of depth, lack of contextuality, failure to integrate clinical data, failure to integrate temporal information. Table? • “The development and application of LLMs are also 105 constrained by limited access to high-quality, diverse clinical datasets due to privacy, ethical, and 106 legal challenges(7). Ethical and legal issues such as bias, misinformation, data privacy, and 107 insufficient regulatory frameworks further challenge their adoption(8)” Constraints to LLM development and constraints to LLM adoption. • “Additionally, the opaque, 108 "black box" nature of these models undermines transparency and interpretability, complicating 109 trust and reliance by healthcare professionals(9).” Limitation table: opacity, black box, trust and reliance. • “Practical concerns include their substantial 110 computational and energy demands, which limit feasibility in resource-constrained 111 environments(3).” Practical limitations, table? • “Such limitations of LLMs pose significant challenges to their adoption and 112 utilization across various sectors of healthcare systems. In this regard, these challenges must be 113 critically addressed to enable the successful and comprehensive integration of LLM technologies 114 within healthcare infrastructure.” Consequences of limitations. Table or figure. • “115 While several review studies have aimed to delineate the limitations of LLMs in providing 116 healthcare content, there remains a significant gap in the literature regarding a comprehensive 117 and systematic presentation of these limitations for end-users” Practical limitations in LLMs, research gap in LLM limitations. • “In this regard, a systematic 118 review categorized the limitations of LLMs into two primary domains: design and output.” In this sentence, “ a systematic review”, should not it be “this systematic review”? • “Design 119 limitations included several items such as lack of optimization for the medical domain, data 120 transparency issues, and accessibility challenges.” Table? • “Output limitations also included several items 121 such as non-reproducibility, incompleteness, inaccuracies, safety concerns, and biases(6).” Table? • “The 122 data generated from such research would be invaluable for technology developers to enhance 123 the quality of their models.” In the introduction, “such research”? Again, are we still referring to this research or recommending such research in our introduction? • “Additionally, healthcare policymakers and administrators could 124 leverage this information to make informed decisions about implementing these technologies 125 within their organizations, fully acknowledging their current constraints.” Definitely • “Moreover, future 126 researchers would benefit from this detailed framework by conducting focused investigations on 127 each identified limitation, thereby contributing further insights and advancing the field for 128 subsequent users and stakeholders.” I totally agree. • “” Why is the methods or material section put after the conclusion? Little mentioned about the methodology of the research under the ‘Results’. Figure-1 • Figure 1 “Records removed before screening: Duplicate records removed (n =1798) Records marked as ineligible by automation tools (n =0) Records removed for other reasons (n =0)” Could you please elaborate on the automation tools? What are the other reasons? • “Figure 1. PRISMA diagram.” Would this be explained to give the lacking ‘Methods’ section? • “In this regard, 20% of the included studies were published in 2023, 139 40% in 2024, and the remaining 40% in 2025.” “In this regard” this phrase was overused in this script. It could be replaced by different places in different places. What is the significance of the distribution of the studies onto the years? Can this be added to the research questions and given explanation for its significance? • “2.1. Data quality 181 The quality assessment of the included studies indicated that they were predominantly of high 182 quality, with an average score of 10. Approximately 34% of the studies achieved the highest 183 quality score of 12, while about 8% scored 7, reflecting lower quality relative to the other 184 included studies. Moreover, the level of bias in the included studies was generally low, with 185 approximately 59% of the studies exhibiting the minimal possible bias (Appendix 1. Quality 186 Assessment).” The assessments of both quality and bias need to be explained in the text, in addition to any table. Also, the table in Appendix one is very brief. It shows the questions for the tabulated criteria, but I believe it needs more elaboration. The evaluation is generally subjective in nature. • “There was also an identified over-reliance on imaging modalities such as 254 CT and MRI, without adequate customization for specific clinical contexts” This is very important point. I believe it deserves wider elaboration, but I understand that it might over expand the scope of this work. • “255 Furthermore, LLMs like ChatGPT demonstrated an inability for critical thinking necessary to tailor 256 and guide patient management effectively” I believe that for an LLM to have critical thinking, they would make them more challenging than any challenge or limitation mentioned here. However, guiding their thinking by human supervised training, and training and more training would make them more useful. • “Conversation tracking posed 265 further limitations: ChatGPT 3.5 lacked conversation memory due to privacy and browsing 266 restrictions, while ChatGPT 4.0, despite tracking conversations, inaccurately counted inquiries, 267 thereby undermining feedback and dialogue continuity” Conversation memory, and tracking accuracy would be one of the biggest future solutions for many of the challenges in LLM clinical use. • “. Accuracy 293 gaps were evident, with occasional imprecision, indecisiveness, hallucinations, irrelevant 294 information, and omissions of key clinical considerations, particularly for special populations 295 such as pregnant patients; fabricated references were also observed sporadically” I realize that fabricated references are a feature of inaccuracy, but I would still suggest that this might better be a separate sentence. • “2.2.3.2. Ethical, safety, and regulatory concerns” This is one of the richest sections of the manuscript. I believe it deserves a separate table or tables. Focusing on keywords and their explanations. It would help a lot in this research to coin new definitions for many keywords to standardize the future study and tackling of the LLMs limitations and challenges • “Table 2. Thematic analysis of findings.” A repeated comment on this table is that it can have its text content converted to single keywords and this can even help create figure and charts from it. • “In this regard, in line with our study findings, a recent review emphasized 347 that the taxonomy of limitations associated with large language models reveals a substantially 348 greater number of codes related to output issues compared to those connected with design or 349 input phases.” Again, “in this regard”, I would suggest that the researchers could take the opportunity to claim the taxonomy of LLM limitations and standardize their linguistic use. • “Figure 2. Distributions of limitations of LLMs, categorized according to the frequency of citations reported in the literature.” Again, I would recommend dividing this figure into three figures. It looks very crowded, and the text font is very tiny because of that. • “381 8). Moreover, flawed, biased, or incomplete data can propagate inaccuracies and inconsistencies 382 within the model outputs, reducing their reliability and safety, especially in healthcare 383 applications” To explain this again. I give the researchers the richness of their work, but the bias and quality standards need to be explained so the reader can benefit from the manuscript and build upon it in improving their future clinical interactions with the LLMs • “. This finding also has significant 445 implicationsfor the beneficiaries. In this regard, our study proposed several strategies to address 446 major issues encountered by the LLMs in the provision of healthcare content.” Thanks again, Yasser Abdullah ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. Do you want your identity to be public for this peer review? If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public. For information about this choice, including consent withdrawal, please see our Privacy Policy ..--> Reviewer #1: No Reviewer #2: Yes: Balu BhasuranBalu BhasuranBalu BhasuranBalu Bhasuran Reviewer #3: Yes: Yasser AbdullahYasser AbdullahYasser AbdullahYasser Abdullah ********** Figure resubmission: -->--> -->--> -->While revising your submission, we strongly recommend that you use PLOS’s NAAS tool ( https://ngplosjournals.pagemajik.ai/artanalysis ) to test your figure files. NAAS can convert your figure files to the TIFF file type and meet basic requirements (such as print size, resolution), or provide you with a report on issues that do not meet our requirements and that NAAS cannot fix.--> --> Reproducibility: -->--> -->-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->--> -->-->To enhance the reproducibility of your results, we recommend that authors of applicable studies deposit laboratory protocols in protocols.io, where a protocol can be assigned its own identifier (DOI) such that it can be cited independently in the future. Additionally, PLOS ONE offers an option to publish peer-reviewed clinical study protocols. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols-->?> PLOS Digit Health. 2026 Apr 8;5(4):e0001354. doi: 10.1371/journal.pdig.0001354.r002 Author response to Decision Letter 1 Article notes Copyright and License information Collection date 2026 Apr. PMC Copyright notice 8 Feb 2026 Attachment Submitted filename: Authors` Response.docx pdig.0001354.s004.docx (35.7KB, docx) PLOS Digit Health. doi: 10.1371/journal.pdig.0001354.r003 Decision Letter 1 Mayue Shi Mayue Shi Academic Editor Find articles by Mayue Shi Author information Copyright and License information Roles Mayue Shi : Academic Editor © 2026 Mayue ShiMayue ShiMayue ShiMayue Shi This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 21 Mar 2026 A Systematic Review of the Limitations of Large Language Models in Generating Healthcare Content PDIG-D-25-00961R1 Dear Dr. Khosravi, We are pleased to inform you that your manuscript 'A Systematic Review of the Limitations of Large Language Models in Generating Healthcare Content' has been provisionally accepted for publication in PLOS Digital Health. Before your manuscript can be formally accepted you will need to complete some formatting changes, which you will receive in a follow-up email from a member of our team. Please note that your manuscript will not be scheduled for publication until you have made the required changes, so a swift response is appreciated. IMPORTANT: The editorial review process is now complete. PLOS will only permit corrections to spelling, formatting or significant scientific errors from this point onwards. Requests for major changes, or any which affect the scientific understanding of your work, will cause delays to the publication date of your manuscript. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they'll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact [email protected]. Thank you again for supporting Open Access publishing; we are looking forward to publishing your work in PLOS Digital Health. Best regards, Zhenwei Shi Section Editor PLOS Digital Health *********************************************************** Additional Editor Comments (if provided): After considering the reviewers’ comments, I am pleased to inform you that your manuscript has been accepted for publication. I have no further comments at this stage. Please follow the instructions regarding the next steps in the publication process. Congratulations, and thank you for submitting your work to us. Reviewer Comments (if any, and for reference): Reviewer's Responses to Questions Comments to the Author Reviewer #1: All comments have been addressed Reviewer #2: All comments have been addressed Reviewer #3: All comments have been addressed ********** publication criteria ? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.? Is the manuscript technically sound, and do the data support the conclusions? The manuscript must describe methodologically and ethically rigorous research with conclusions that are appropriately drawn based on the data presented.-->?> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** 3. Has the statistical analysis been performed appropriately and rigorously?-->?> Reviewer #1: Yes Reviewer #2: N/A Reviewer #3: N/A ********** 4. Have the authors made all data underlying the findings in their manuscript fully available (please refer to the Data Availability Statement at the start of the manuscript PDF file)??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception. The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes Reviewer #3: Yes ********** Reviewer #1: The authors have done a thorough job revising the manuscript after the initial peer review. They have effectively addressed all the concerns and suggestions raised by the reviewers. The current version of the manuscript is much stronger and now meets the publication standards. 1. The authors thoroughly addressed reviewer comments, showing their dedication to enhancing the study’s methods and discussion based on feedback. 2. The manuscript is now clearer and more organized, with added search strategies and better definitions of analytical frameworks. 3. The technical quality is confirmed with criteria for study inclusion/exclusion and transparent bias and quality assessments. 4. This work significantly contributes to digital health by detailing the limitations of large language models in healthcare content. 5. Results are strengthened by better data organization, with tables and figures illustrating model constraint frequencies and distributions. 6. The manuscript's presentation and quality have greatly improved, with smoother text, corrected grammar, and recent references. Reviewer #2: The authors have adequately addressed all of the points raised in the previous round of review. The manuscript is significantly improved, and I have no further concerns. I recommend the paper for publication. Reviewer #3: I repeat my appreciation for the opportunity to review this work. I find it a great work and I believe it is timely and important to be added to the literature discussing the Healthcare LLM subject. Despite the fact that as a reader, I do not agree that one huge table can improve readability, I would recommend this manuscript to be published and I wish you the best of luck in your future research and I encourage you to build on this work in particular. ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. Do you want your identity to be public for this peer review? If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public.If you choose “no”, your identity will remain anonymous but your review may still be made public. For information about this choice, including consent withdrawal, please see our Privacy Policy ..--> Reviewer #1: No Reviewer #2: Yes: Balu BhasuranBalu BhasuranBalu BhasuranBalu Bhasuran Reviewer #3: Yes: Yasser AbdullahYasser AbdullahYasser AbdullahYasser Abdullah ********** Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials S1 Appendix. Quality assessment. (DOCX) pdig.0001354.s001.docx (69.7KB, docx) S1 Checklist. PRISMA checklist [ 120 ]. (DOCX) pdig.0001354.s002.docx (37.7KB, docx) Attachment Submitted filename: Authors` Response.docx pdig.0001354.s004.docx (35.7KB, docx) Data Availability Statement The research data is presented as a supplementary file. Articles from PLOS Digital Health are provided here courtesy of PLOS ACTIONS View on publisher site PDF (1.2 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top