Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Front Digit Health . 2026 Mar 26;8:1761641. doi: 10.3389/fdgth.2026.1761641 Search in PMC Search in PubMed View in NLM Catalog Add to search Large language models in healthcare quality management: a European perspective on process automation and compliance Markus Knott Markus Knott 1 Klinikum Stuttgart, Stuttgart Cancer Center – Tumorzentrum Eva Mayr-Stihl, Stuttgart, Germany Conceptualization, Methodology, Writing – original draft, Writing – review & editing Find articles by Markus Knott 1, *, † , Markus Krebs Markus Krebs 2 Comprehensive Cancer Center Augsburg, Medical Faculty, University of Augsburg, Augsburg, Germany 3 Bavarian Cancer Research Center (BZKF), Erlangen, Germany Conceptualization, Visualization, Writing – review & editing Find articles by Markus Krebs 2, 3 , Alexander Kerscher Alexander Kerscher 3 Bavarian Cancer Research Center (BZKF), Erlangen, Germany 4 Department of Gynecology and Obstetrics, Erlangen University Hospital, Comprehensive Cancer Center Erlangen-European Metropolitan Area Nuremberg (CCC ER-EMN), Friedrich Alexander University of Erlangen-Nürnberg, Erlangen, Germany Writing – review & editing Find articles by Alexander Kerscher 3, 4 Author information Article notes Copyright and License information 1 Klinikum Stuttgart, Stuttgart Cancer Center – Tumorzentrum Eva Mayr-Stihl, Stuttgart, Germany 2 Comprehensive Cancer Center Augsburg, Medical Faculty, University of Augsburg, Augsburg, Germany 3 Bavarian Cancer Research Center (BZKF), Erlangen, Germany 4 Department of Gynecology and Obstetrics, Erlangen University Hospital, Comprehensive Cancer Center Erlangen-European Metropolitan Area Nuremberg (CCC ER-EMN), Friedrich Alexander University of Erlangen-Nürnberg, Erlangen, Germany * Correspondence: Markus Knott [email protected] † ORCID Markus Knott orcid.org/0000-0002-9147-3203 Roles Markus Knott : Conceptualization, Methodology, Writing – original draft, Writing – review & editing Markus Krebs : Conceptualization, Visualization, Writing – review & editing Alexander Kerscher : Writing – review & editing Received 2025 Dec 5; Revised 2026 Feb 20; Accepted 2026 Feb 24; Collection date 2026. © 2026 Knott, Krebs and Kerscher. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. PMC Copyright notice PMCID: PMC13062252 PMID: 41970523 Abstract Large Language Models (LLMs) are transforming back-office quality management processes in European healthcare systems through automation of compliance monitoring, quality assurance, and process optimization without direct patient interaction. This narrative review synthesizes evidence from recent systematic reviews and implementation studies (2023-2025) examining LLM deployment within the European regulatory framework encompassing the Medical Device Regulation (MDR), General Data Protection Regulation (GDPR), and the EU Artificial Intelligence Act (Regulation EU 2024/1689). Current research demonstrates meaningful efficiency gains: individual studies of AI-assisted documentation tools report improvements ranging from modest increases in documentation speed to reductions in processing time approaching 50%, while broader policy analyses estimate administrative workload reductions of up to 30% through digital health and AI solutions. Clinical trial applications show particular maturity, with LLM-generated informed consent forms demonstrating improved readability (76% vs. 67%) without compromising accuracy. However, critical gaps persist between research achievements and practical deployment. Analysis of 519 evaluation studies reveals that only 5% utilized real patient care data, while 95% focused exclusively on accuracy metrics to the neglect of fairness (16%), deployment readiness (5%), and calibration (1%). No LLM-based quality management system has yet received regulatory clearance, and implementation science frameworks remain underdeveloped. We propose a risk-stratified implementation framework emphasizing process-oriented applications—standard operating procedure automation, audit documentation, deviation management, and compliance monitoring—that avoid medical device classification while capturing substantial operational benefits. Advanced methodological approaches including retrieval-augmented generation (RAG) architectures, digital twin integration, and natural language processing-based pattern recognition offer pathways toward comprehensive quality intelligence platforms. The convergence of LLMs with emerging technologies such as knowledge graphs, digital twin architectures and multimodal analysis creates opportunities for predictive quality management that anticipates rather than merely documents quality-relevant events. Evidence supports deployment in administrative quality processes, with particular potential for applications that redirect human expertise from documentation toward quality improvement activities, though current evidence derives predominantly from non-European healthcare contexts and simulated or limited-scope settings. Success requires adapted validation methodologies addressing LLM non-determinism, robust governance structures, and comprehensive change management that maintains the high standards European healthcare systems demand. Keywords: digital health, EU AI act, healthcare quality management, large language models, medical device regulation, process automation, regulatory compliance, retrieval-augmented generation Introduction European healthcare systems allocate substantial resources to quality management and regulatory compliance, with estimates suggesting that administrative processes consume 15%-20% of operational budgets ( 1 , 2 ). The implementation of the Medical Device Regulation (EU 2017/745) has further added documentation requirements ( 54 ), while the European Health Data Space Regulation ( 56 ) promises additional complexity. This administrative burden diverts resources from direct patient care and quality improvement initiatives. As an illustrative example drawn from the authors ' professional experience, a medium- to large-sized German cancer center dedicates thousands of staff hours annually to certification preparation for the Deutsche Krebsgesellschaft (DKG), yet a substantial proportion of this effort involves repetitive document review that adds no direct clinical value. While we present this as anecdotal illustration rather than empirical evidence, it reflects a documentation burden widely recognized across European healthcare institutions. Large Language Models (LLMs) represent a paradigm shift in addressing these challenges. Unlike traditional rule-based systems requiring structured inputs and rigid templates, LLMs process natural language, interpret context, and generate coherent outputs across diverse quality management scenarios. Recent systematic reviews ( 3 , 4 ) demonstrate their potential for transforming healthcare operations, though significant gaps remain between research achievements and practical implementation. The distinction between clinical and administrative AI applications proves fundamental for regulatory navigation. Quality management tools that operate without direct patient interaction may avoid medical device classification under MDR (Medical Device Regulation), enabling faster deployment while maintaining safety. The EU AI Act ( 9 ) provides additional clarity, establishing risk-based requirements that align with existing quality management principles. This narrative review examines current evidence for LLM deployment in healthcare quality management, emphasizing process-oriented applications suitable for the European regulatory environment. We synthesize findings from peer-reviewed systematic reviews and primary studies, European regulatory documents, and selected implementation literature to analyze implementation challenges and propose frameworks for responsible deployment that maximize operational efficiency while maintaining compliance. Figure 1 provides a conceptual overview of the four thematic domains addressed in this review, which were identified through author consensus informed by the structure of the European regulatory landscape and the authors ' implementation experience at three German medical centers. Figure 1. Open in a new tab Conceptual overview of thematic domains addressed in this narrative review. Methodological approach This narrative review was conducted to synthesize current evidence on LLM applications in healthcare quality management within the European regulatory context. We searched PubMed/MEDLINE, IEEE Xplore, Google Scholar, and arXiv using combinations of terms including “large language models,” “healthcare quality management,” “process automation,” “clinical documentation,” “EU AI Act,” and “medical device regulation.” The primary search window covered publications from 2023 to 2025, supplemented by targeted citation tracking from identified systematic reviews and inclusion of foundational references from outside this search window where necessary for regulatory or methodological context. Sources were organized into four categories: (a) peer-reviewed systematic reviews and primary studies examining LLM performance, validation, and implementation in healthcare and related domains (2019–2025); (b) European regulatory and guidance documents including the EU AI Act, MDR/IVDR, General Data Protection Regulation (GDPR) ( 5 ), and associated Medical Device Coordination Group (MDCG) guidance; (c) selected grey literature, technical standards, and preprints informing implementation practice, including GxP frameworks and quality management standards; and (d) foundational textbooks and professional guidance documents providing methodological or conceptual context. All sources in categories (c) and (d) are publicly available documents; no unpublished original data were used. Thematic domains were derived through author consensus informed by the EU regulatory landscape and professional experience at three German academic medical centers. The resulting structure—regulatory framework analysis, current state of evidence, advanced methodological approaches, implementation pathway, and discussion—reflects both the literature landscape and the practical needs of healthcare organizations navigating LLM adoption. As a narrative review, this approach prioritizes breadth of coverage and contextual integration over the exhaustive search and formal quality assessment characteristic of systematic reviews. Accordingly, no formal study counting, PRISMA flow diagram, or standardized quality appraisal tool was applied; this methodological choice is appropriate given the rapidly evolving evidence base, the interdisciplinary scope spanning AI research, regulatory science, and healthcare quality management, and the perspective-driven nature of the synthesis. Readers seeking a formal systematic assessment of LLM limitations in patient care are referred to Busch et al. ( 3 ), while Bedi et al. ( 4 ) provides a systematic evaluation of LLM performance metrics across healthcare tasks. A-Regulatory framework analysis EU AI Act implementation The EU AI Act (Regulation EU 2024/1689) establishes comprehensive requirements based on risk categorization. Quality management applications typically qualify as limited or minimal risk systems, avoiding requirements for high-risk medical AI. However, organizations must ensure transparency, maintain human oversight, and implement appropriate risk management strategies. Key requirements for quality management implementations include: Data governance Training datasets must demonstrate relevance, representativeness, and accuracy. For quality applications, this necessitates using historical QMS data that accurately reflect organizational processes. The Act requires documentation of data sources, preprocessing methods, and potential biases (Article 10, EU AI Act). Technical documentation Detailed documentation encompassing system design, development methodology, and validation results. Quality professionals familiar with design history files will recognize parallels with existing medical device documentation requirements (Annex IV, EU AI Act). Human oversight Systems must enable human understanding and intervention. This aligns naturally with quality management principles emphasizing defined responsibilities and authorities (Article 14, EU AI Act). Transparency Organizations must clearly communicate when users interact with AI systems and explain capabilities and limitations (Article 13, EU AI Act). MDR and medical device classification The Medical Device Coordination Group guidance (MDCG 2025-6) ( 55 ) clarifies the interplay between MDR/IVDR and the AI Act. Quality management tools qualify as medical devices only when serving medical purposes—diagnosis, treatment support, or clinical decisions. Administrative tools without direct patient care impact typically remain exempt, though borderline cases require careful assessment. ISO 13485:2016 Section 4.1.6 mandates software validation for any system affecting product quality, regardless of medical device status ( 59 ). Research on LLM validation in regulated environments ( 6 ) proposes statistical approaches for nondeterministic systems, including confidence intervals for acceptable output variation and continuous performance monitoring protocols. GDPR compliance GDPR ( 5 ) presents unique challenges for LLM implementation in quality management. Even administrative health data constitutes “special category data” under Article 9, requiring explicit consent or alternative lawful basis. Quality management typically relies on legitimate interest (Article 6(1)(f)) or legal obligation (Article 6(1)(c)), but organizations must document their rationale comprehensively. Data Protection Impact Assessments become mandatory before deploying LLMs processing health-related data (Article 35, GDPR). Comprehensive reviews of privacy-preserving techniques for generative AI identify four principal approaches applicable to healthcare quality management: differential privacy, which adds mathematical noise to prevent individual data reconstruction; federated learning, enabling model training across distributed datasets without centralized data aggregation; homomorphic encryption, permitting computations on encrypted data; and secure multi-party computation for collaborative analysis without exposing raw data ( 7 ). These techniques specifically address risks including model inversion attacks, data leakage during training, and membership inference vulnerabilities—concerns particularly relevant when processing quality management data derived from patient records. Importantly, these privacy-preserving approaches can be aligned with EU AI Act requirements while maintaining analytical utility for quality improvement applications. B-Current state of evidence Systematic review findings Comprehensive systematic reviews reveal both promise and limitations in current LLM research. Busch et al. ( 3 ) analyzed 89 studies across 29 medical specialties, establishing a taxonomy of LLM limitations. Their PRISMA-guided review found that 87.6% of studies reported issues with non-comprehensiveness and incorrectness, while 42.7% documented reproducibility problems. Most critically, design limitations including lack of medical optimization and data transparency affected most implementations. Bedi et al. ( 4 ) examined 519 studies evaluating LLM performance in healthcare tasks, revealing fundamental methodological weaknesses: only 5% used real patient care data for evaluation, 95.4% focused exclusively on accuracy metrics while neglecting fairness (15.8%), deployment readiness (4.6%), and calibration (1.2%). Administrative tasks remained significantly understudied, with billing codes and prescriptions each representing only 0.2% of evaluated applications. Quality measurement automation Research on automated quality measurement demonstrates mature capabilities for process automation. A pivotal study from a major U.S. academic medical center ( 8 ) evaluated LLM-based automation for complex quality measure abstraction. The system achieved 90% agreement with manual abstractors ( κ =0.82; 95% CI: 0.71-0.92), substantially exceeding traditional inter-rater reliability ( κ =0.39). When discrepancies occurred, expert review identified human abstractor errors in 40% of cases, suggesting LLMs may enhance rather than merely replicate human performance. The economic implications of inefficient quality reporting and administrative workflows are substantial. Across EU health systems, workforce shortages and rising demand for care make inefficient use of staff time increasingly costly [OECD/European Commission ( 64 )]. Digital health and AI solutions can reduce administrative workload for health professionals by up to 30% and automate repetitive back-office processes, freeing capacity for patient care and quality management. At the level of individual AI-assisted tools, primary studies report task-specific documentation efficiency gains, though these derive from heterogeneous settings and technologies. Xia et al. ( 10 ) developed a speech-recognition-based electronic medical record system that reduced average processing time from 46 to 26 min which accumulates to an approximately 44% reduction in a Chinese healthcare setting. Mairittha et al. ( 11 ) reported a 15% increase in documentation speed using a spoken dialogue system for nursing care, though this involved only 12 participants in a simulated scenario. A systematic review of speech recognition for clinical documentation from 1990 to 2018 found highly variable results, with five studies reporting 19%–92% decreases in documentation time but four others reporting 13%–50% increases ( 12 ) , underscoring that efficiency gains are neither uniform nor guaranteed. Process documentation and compliance Evidence for LLM applications in pharmaceutical quality management derives primarily from industry reports and limited peer-reviewed studies. An industry collaboration between SeerPharma and the University of Melbourne (2024—grey literature) analyzed LLM implementations across pharmaceutical quality management subsystems, identifying applications in audit scheduling optimization, automated report generation, and inspection readiness evaluation ( 58 ). However, the scarcity of peer-reviewed validation studies represents a critical gap in the literature. Nelson and Aguero ( 13 ) provide one of the few peer-reviewed examinations of LLMs in pharmaceutical operations, focusing on supply chain management. Their analysis highlights both opportunities—inventory optimization, procurement automation, distribution planning—and significant risks including data leakage, prompt injection vulnerabilities, and potential for hallucinated outputs. The authors strongly recommend restricting LLM deployment to clerical tasks using historical data with mandatory human verification, emphasizing that autonomous operation remains premature given current validation gaps. Research on business process modeling demonstrates substantial LLM capabilities. Comprehensive evaluation of 16 state-of-the-art LLMs across 20 diverse business processes revealed significant but variable performance in translating natural language descriptions into formal process models ( 14 ). The evaluation framework assessed LLM capabilities in modeling business processes from descriptions, generating executable code, following embedded instructions, and incorporating feedback for iterative quality improvement. Results demonstrated positive correlation between efficient error handling and output quality, with self-improvement techniques—particularly output optimization—showing promise for enhancing model quality, especially for initially lower-performing models. Translated to quality management applications, these findings suggest that LLM-assisted standard operating procedure (SOP) generation should rely on iterative refinement cycles rather than one-shot generation approaches. The ability of LLMs to translate procedural descriptions into structured formats, combined with error handling for self-correction, aligns well with the iterative nature of SOP development in regulated environments. However, the observed performance variations across LLM types underscore the importance of systematic evaluation before production deployment. Key takeaway While peer-reviewed evidence supports LLM capabilities in process modeling, successful implementation requires appropriate model selection, iterative refinement protocols, and validation against organizational quality standards. Clinical trial documentation and good clinical practice (GCP) compliance In marked contrast, LLM applications in clinical trial management benefit from extensive peer-reviewed evidence demonstrating both efficacy and implementation pathways within existing GCP frameworks. Informed consent innovation Multiple high-quality studies validate LLM capabilities in consent documentation. Decker et al. ( 15 ) demonstrated in JAMA Network Open that ChatGPT-3.5 generated informed consent forms that were, on average, less complex (lower readability grade level) and significantly more comprehensive and accurate than surgeon-generated documentation, with no instances of clinically inaccurate information in the chatbot output. Building on this, Shi et al. ( 16 ) showed the Mistral 8 × 22B model achieved 76.39% readability scores vs. 66.67% for human-generated ICFs, with no compromise in accuracy ( p > .10 for all accuracy measures). Clinical trial operations Omar and Nadkarni ( 17 ) systematically reviewed 27 trials investigating LLM applications across healthcare applications. Ongoing interventional trials include LLM-assisted discharge summary generation ( NCT06263855 ; target n = 1,015), preoperative visit documentation ( NCT05945004 ), and ChatGPT-supported informed consent for knee arthroplasty (ChiCTR2300078274). However, accuracy remains concerning—inaccuracies appear in 36% of LLM-generated patient histories despite improved detail and comprehensiveness ( 18 ). Patient engagement Gao et al. ( 19 ) demonstrated GPT-4's transformative potential for patient education, generating trial summaries from complex consent forms that 80% of participants found enhanced their understanding. The sequential summarization approach balanced accuracy with accessibility, addressing long-standing challenges in trial recruitment and retention. A comprehensive review of LLM applications across the clinical trial lifecycle identifies implementation opportunities spanning trial design, operations, and analysis phases ( 20 ). In trial design, LLMs demonstrate capability for extracting research elements from prior studies, refining eligibility criteria, and tailoring informed consent materials. Operational applications include accelerated patient screening through automated eligibility assessment, standardized data collection, and real-time safety monitoring including adverse event detection and drug-drug interaction identification. The review emphasizes that domain-specific pretraining and fine-tuning substantially enhance LLM performance for clinical trial tasks, with patient-trial matching and trial data extraction showing particular promise for reducing time and financial costs while improving recruitment efficiency. Regulatory integration LLM implementation in clinical trials must comply with the revised ICH E6(R3) Good Clinical Practice guideline (adopted January 2025), which strengthens expectations for data governance and computerised systems ( 52 , 61 ). Allen et al. ( 21 ) outline five possible models for integrating LLMs into clinical research consent—from use as a supplementary tool to more automated configurations—and argue that more automated models require stronger oversight and, in some cases, regulatory reform. Building on broader GxP guidance [ICH E6(R3), EMA's guideline on computerised systems and electronic data, 21 CFR Part 11, and GAMP 5] ( 53 ), key compliance considerations for LLM-based consent systems include: Secure, time-stamped audit trails for consent interactions Robust version control and change management for models and prompt templates Human-in-the-loop oversight and accountability consistent with GCP principles on investigator and sponsor responsibilities Risk-based validation of LLM systems following GAMP 5 and related CSV guidance Ensuring GxP data integrity for all LLM-generated records according to ALCOA + principles (Attributable, Legible, Contemporaneous, Original, Accurate, Complete, Consistent, Enduring, Available) as outlined in WHO and PIC/S data-integrity guidance ( 57 ) Key takeaway Clinical trial applications demonstrate mature evidence with clear regulatory pathways, contrasting sharply with the validation gaps in pharmaceutical quality management. C-Advanced methodological approaches Digital twins for quality management The evolution from static metrics to dynamic quality intelligence increasingly leverages digital twin technology. Recent research demonstrates how digital twin platforms create virtual representations of physical systems spanning their lifecycles, facilitating simulation and optimization without operational risk ( 22 , 23 ). Digital twin implementations enable healthcare providers to assess and enhance processes through integration of data from diverse sources including electronic health records, medical devices, and administrative systems to identify bottlenecks and inefficiencies ( 24 ). The convergence of digital twin architecture with LLM capabilities creates powerful quality intelligence platforms. Research in manufacturing quality control demonstrates how Asset Administration Shell (AAS) frameworks combined with LLMs enable interoperable information modeling for zero defect strategies ( 25 ). This approach addresses a fundamental challenge in quality management: data interoperability across heterogeneous systems. The methodology employs fine-tuned LLMs for semantic search and entity matching, automatically referencing standardized vocabularies to maintain consistency across organizational systems. A case study in injection molding demonstrated the practical application, achieving statistical validation of LLM-based semantic search algorithms for linking product quality and process data. This integration pattern—combining standardized digital twin representations with LLM-based semantic processing—offers a transferable framework for healthcare quality management, where similar challenges of data heterogeneity and terminology standardization exist across clinical, administrative, and regulatory systems. The application of digital twins specifically within quality engineering contexts emphasizes the role of statistics in connecting virtual and physical systems, with implementations demonstrating versatility across manufacturing domains ( 26 ). This statistical foundation aligns with existing quality management competencies, suggesting natural integration pathways for healthcare quality professionals familiar with statistical process control and design of experiments methodologies. Complementing these engineering-derived digital twin approaches, recent work has proposed a systematic framework for integrating mathematical modeling—including machine learning, knowledge graphs, and digital twins—directly into healthcare quality assessment ( 27 ). This framework categorizes patient-centered quality of care into three quantifiable dimensions: patient safety, procedure accuracy, and procedure efficacy, each with corresponding mathematical descriptions and modeling tasks. By mapping quality metrics to these categories and assigning relevant computational methods to each, the framework provides a structured pathway from conceptual quality definitions to operational digital twin implementations. For quality management applications, this approach offers two advantages: first, it grounds digital twin design in established quality dimensions rather than ad hoc metrics; second, it identifies specific knowledge graph architectures suited to each quality category, enabling targeted deployment of graph-based reasoning for safety surveillance, procedural compliance monitoring, and efficacy benchmarking. The integration of such structured quality ontologies with LLM-based processing could advance the transition from retrospective documentation-oriented quality management toward predictive, model-driven quality intelligence. Digital twin implementations offer several methodological advantages: Predictive Simulation: Virtual testing of process modifications reduces validation costs and implementation risks System-Level Analysis: Modeling interconnected processes identifies cascade effects and hidden dependencies Continuous Calibration: Real-time data integration improves prediction accuracy iteratively Risk-Free Experimentation: Radical process improvements can be evaluated virtually before implementation NLP-based pattern recognition Advanced natural language processing (NLP) techniques enable sophisticated pattern recognition in quality data. Diaz Ochoa et al. ( 28 ) developed clustering algorithms using edge-betweenness methods for identifying symptom networks in emergency presentations, demonstrating methodology applicable to quality management. Their approach identified sex-specific patterns and distinct stratification patterns among polysymptomatic, oligosymptomatic, and atypical cases that traditional analysis overlooked, suggesting applications for detecting systemic biases in quality metrics. Research on ensemble methods for LLMs demonstrates improved reliability through consensus mechanisms ( 17 , 62 , 63 ). By comparing outputs from diverse models, these systems identify potential discrepancies and improve confidence estimates through iterative consensus approaches. Meta-learning approaches further enhance performance by adapting to organization-specific terminology and processes. Retrieval-augmented generation Studies employing retrieval-augmented generation (RAG) architectures report significant improvements in accuracy and reliability ( 30 ). The foundational RAG framework combines pre-trained parametric and non-parametric memory for language generation, where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index accessed with a neural retriever ( 30 ). RAG models have been shown to generate more specific, diverse, and factual language than parametric-only seq2seq baselines, particularly excelling in knowledge-intensive NLP tasks. Recent implementations demonstrate substantial performance gains. Speculative RAG achieves up to 12.97% accuracy improvement while reducing latency by 51% compared to conventional RAG systems ( 31 ). Corrective RAG (CRAG) frameworks address scenarios where retrievers return inaccurate results, significantly improving robustness across both short- and long-form generation tasks (Yan et al. 2024). The key innovation involves grounding LLM outputs in verified knowledge bases through dynamic retrieval mechanisms. Technical implementations typically combine vector databases for semantic search with LLMs for response generation. This architecture enables organizations to maintain control over source information while leveraging generative capabilities for natural language interaction, with embedding models translating queries and documents into vectors for similarity-based retrieval ( 32 ). Practical applications in manufacturing quality control demonstrate RAG's effectiveness for troubleshooting and failure analysis. An advanced RAG system designed for quality control utilized specialized bibliographic knowledge bases to diagnose defects and propose solutions, incorporating tailored preprocessing and postprocessing mechanisms to optimize document retrieval and response generation ( 33 ). The system demonstrated capability for identifying nonconformities, determining root causes, and generating actionable solutions—applications directly transferable to healthcare quality management where similar structured knowledge bases exist for deviation management, CAPA processes, and regulatory guidance interpretation. This suggests RAG architectures connecting organizational quality documentation with regulatory requirements could substantially accelerate deviation investigation and corrective action development. D-Implementation framework Translating LLM capabilities into practical quality management tools demands structured implementation strategies. The phased implementation framework presented below was developed through author consensus, informed by synthesis of the reviewed literature, existing implementation science principles, and the authors' professional experience with quality management systems at German medical centers. The specific phase timelines represent pragmatic estimates based on typical institutional change management cycles in European healthcare settings and should be adapted to local organizational contexts. This framework has not been empirically validated through prospective pilot testing, which represents an important direction for future research. Risk-stratified deployment Evidence supports graduated implementation beginning with lowest-risk applications: Phase 1: administrative documentation (months 0-6) Research indicates immediate value in automating routine documentation tasks. AI tools can generate draft SOPs that are both comprehensive and accurate, significantly reducing the time and effort involved in the creation process ( 29 ). These applications involve no patient data and minimal regulatory oversight. Phase 2: quality process automation (months 6-12) Intelligent process automation (IPA) achieves flexible and intelligent automation by combining robotic process automation (RPA), artificial intelligence (AI), and other emerging technologies ( 34 ). Studies indicate RPA helps reduce labor costs and enables businesses to reduce human error, with bots operating 24/7 to essentially eliminate downtime-induced wastage ( 35 – 37 ). Implementation studies document efficiency gains through automated audit report generation and deviation categorization ( 34 , 38 ). Research on LLM integration with industrial quality management systems provides implementation insights applicable to healthcare. Studies demonstrate that success depends on systematic approaches combining standardized information models with semantic search capabilities, enabling LLMs to navigate organizational knowledge bases while maintaining terminological consistency ( 25 ). This suggests healthcare implementations should prioritize integration with existing quality management information systems and standardized healthcare vocabularies (ICD, SNOMED CT, LOINC) before attempting autonomous quality assessment Phase 3: compliance intelligence (months 12–18) Advanced applications leverage AI for regulatory interpretation and gap analysis. The integration of artificial intelligence into clinical decision support systems has significantly enhanced diagnostic precision, risk stratification, and treatment planning, though challenges remain in achieving regulatory approval for such systems ( 39 ). A systematic review examining AI and natural language processing techniques in healthcare found that these technologies can effectively improve clinical decision systems' accuracy when combined with human criteria, optimizing clinical diagnosis and treatment flows ( 40 ). However, the unique characteristics of large language models—including probabilistic outputs and potential for emergent behaviors—present novel regulatory challenges that current frameworks struggle to address, necessitating new approaches to ensure responsible deployment and patient safety ( 41 ). Phase 4: predictive analytics (months 18+) Integration with advanced analytics enables predictive quality management. Machine learning methods support defect detection, root cause analysis, and predictive maintenance, offering opportunities to reduce production risks, minimize unexpected downtimes, and optimize processes ( 42 ). Healthcare organizations leveraging predictive modeling facilitate the transition from reactive to proactive healthcare delivery models by enabling predictive and preventive interventions; by analyzing historical data and real-time information, organizations can anticipate patient needs, identify high-risk individuals, and intervene early to prevent adverse health events ( 43 ). Validation methodology Validating nondeterministic systems requires fundamentally adapted approaches that acknowledge the unique characteristics of generative language models. Unlike deterministic AI prediction algorithms where standardized validation criteria such as discrimination metrics can be consistently applied, LLMs face a more complex validation landscape due to their output variability—a single prompt may generate grammatically distinct yet equally valid text outputs ( 44 ). This extensive output space complicates comparison to ground truth and necessitates task-specific validation frameworks. Statistical validation Research proposes using confidence intervals rather than deterministic pass/fail criteria. Studies establish acceptable variation ranges through repeated testing, typically requiring 95% CI within predetermined boundaries ( 45 ). For quality management applications, validation must distinguish between grammatical variations (acceptable) and factual inconsistencies (unacceptable), requiring domain-specific evaluation rubrics. Multi-dimensional quality assessment Validation frameworks should address multiple axes including factuality, comprehension, reasoning quality, potential for harm, and bias detection ( 44 ). For quality management specifically, additional dimensions include regulatory alignment, organizational consistency, and actionability of generated outputs. Continuous monitoring Unlike traditional software, LLMs require ongoing performance assessment. Recommended metrics include accuracy trends, drift detection, and anomaly identification using statistical process control methods ( 46 ). The non-deterministic nature of LLM outputs means that validation represents a continuous process rather than a one-time certification event. Human-in-the-loop validation Systems maintaining human review require validation of the complete human-AI system. Research emphasizes validating reviewer training, escalation procedures, and override protocols ( 47 ). Critically, the validation scope must encompass how human reviewers interact with AI-generated content, as incorrect outputs may present as plausible and require domain expertise to identify. Governance structures Systematic reviews of successful implementations identify critical governance elements: Multidisciplinary Oversight: Committees combining quality management, IT, regulatory affairs, and domain expertise are essential for comprehensive AI oversight. The integration of Quality Management System (QMS) principles into AI/ML lifecycle management requires multidisciplinary teams encompassing People & Culture, Process & Data, and Validated Technology components to bridge the translation gap between research and operational deployment ( 48 ). Governance committees should ensure appropriate deliberations regarding efficacy, effectiveness, privacy, safety, quality, and ethical factors of AI applications in quality management contexts ( 49 ). The EU AI Act specifically mandates that high-risk AI systems operate under quality management systems incorporating defined governance structures with documented responsibilities [Article 17, EU AI Act]. Risk Assessment Frameworks: Standardized evaluation of implementation risks and required controls forms a regulatory requirement under multiple frameworks. ISO/IEC 42001:2023, the first international AI management system standard, provides structured approaches for risk management, impact assessment, and lifecycle governance of AI systems [ISO/IEC 42001:2023] ( 60 ). Healthcare organizations can adapt existing QMS frameworks—analogous to those in regulated pharmaceutical industries—to establish risk-based approaches for AI technologies that complement existing governance structures ( 48 ) The FAIR-AI framework offers practical guidance integrating risk assessment throughout the AI lifecycle ( 50 ). Performance Dashboards: Real-time visibility into system performance across applications enables continuous quality assurance. Research demonstrates that LLM-based systems can achieve 90% concordance with manual quality measure abstraction ( κ =0.82) when properly monitored ( 8 ). The SALIENT framework emphasizes systematic tracking of AI system performance as integral to end-to-end implementation ( 37 ). Continuous monitoring systems should track both operational metrics (throughput, error rates, processing times) and quality indicators aligned with organizational QMS requirements. Incident Management: Clear procedures for identifying, reporting, and correcting system-related events align with established CAPA (Corrective and Preventive Action) processes in quality management systems. Organizations should integrate AI-related incident reporting into existing QMS structures, ensuring that unexpected outputs, system failures, or quality deviations trigger appropriate investigation and corrective action workflows. This approach extends traditional QMS incident management principles to address the unique challenges of AI system variability [ISO 13485:2016] ( 59 ). Change Control: Formal processes for updating models, prompts, or integration points must address AI-specific challenges while aligning with established computerized system validation requirements. GAMP 5 principles for risk-based validation of computerized systems provide a foundation for AI change control, though adaptation is required for LLM non-determinism [ISPE GAMP 5, 2022] ( 53 ). Change control procedures should document model versions, prompt modifications, training data updates, and integration changes, with appropriate testing and approval workflows before production deployment. Discussion This narrative review has examined the current landscape of LLM applications in healthcare quality management, revealing a field rich in potential but constrained by significant implementation challenges. Our analysis synthesizes evidence from multiple systematic reviews and primary studies to provide a comprehensive—though not exhaustive—overview of this rapidly evolving domain. The evidence base, while growing quickly, remains dominated by proof-of-concept studies and small-scale evaluations rather than large-scale operational deployments. Moreover, the available evidence originates predominantly from US-based and Asian healthcare contexts, with European-specific implementation data for LLM-based quality management remaining notably scarce — a gap that limits direct transferability of findings to European healthcare systems with their distinct regulatory requirements, multilingual environments, and heterogeneous quality management structures. The fundamental gap between research achievements and practical deployment stems not from technical limitations but from inadequate implementation science. Only 5% of studies use real patient care data, while critical factors like fairness, deployment readiness, and calibration remain largely unexamined. This disconnect necessitates new research priorities emphasizing real-world validation and implementation frameworks. European organizations face unique opportunities and challenges. The regulatory framework, while complex, provides clearer pathways for non-clinical applications than many other jurisdictions. By focusing on process automation rather than clinical decision support, organizations can capture substantial benefits while avoiding the most stringent regulatory requirements. The integration of LLMs with emerging technologies—digital twins, knowledge graphs, advanced clustering algorithms—points toward comprehensive quality intelligence platforms. These hybrid approaches overcome individual technology limitations while creating capabilities neither could achieve independently. Emerging evidence suggests that such integrated systems reduce quality incidents by 25%-40% through predictive intervention. Privacy considerations extend beyond regulatory compliance to fundamental system architecture decisions. Recent analysis demonstrates that absolute privacy in LLM systems remains mathematically impossible, necessitating risk-based approaches that balance privacy protection with analytical utility ( 7 ). European organizations implementing LLMs for quality management should consider hybrid approaches combining differential privacy during model training with access controls and audit logging during deployment. The emergence of post-quantum cryptography as a future-proofing measure underscores the need for architectures that can adapt to evolving security requirements without requiring complete system reconstruction. Critical success factors emerging from systematic reviews include strong governance structures, adapted validation methodologies, and comprehensive change management. Organizations must develop new competencies in prompt engineering, AI result interpretation, and system limitation recognition. These “soft” factors often determine implementation success or failure. The evolution toward multimodal LLMs—systems capable of processing text, images, and other data types simultaneously—presents additional opportunities for quality management applications ( 51 ). Healthcare quality data increasingly spans multiple modalities: textual audit reports, process flow diagrams, photographic documentation of facility conditions, and time-series sensor data from medical devices. Multimodal LLMs could enable integrated analysis across these heterogeneous data sources, potentially identifying quality patterns invisible to text-only systems. However, multimodal integration introduces additional validation complexity, as errors may arise from misalignment between modalities or from modality-specific artifacts that propagate through combined analysis. Organizations should monitor developments in multimodal medical LLMs while recognizing that current implementations should focus on demonstrating value in text-based applications before expanding to multimodal architectures. Limitations Several limitations of this narrative review should be acknowledged. Literature selection was purposive rather than systematic, guided by relevance to the European regulatory context and practical implementation considerations. While we drew on major peer-reviewed systematic reviews ( 3 , 4 ) and primary studies, the absence of a formal protocol, standardized screening, or quality assessment means that relevant studies may have been omitted and that the weight of evidence for specific claims may differ from what a systematic review would yield. The empirical evidence synthesized in this review originates predominantly from US-based studies and international systematic reviews. European-specific implementation data for LLM-based quality management remains limited, reflecting the earlier adoption of comprehensive electronic health record systems and larger health informatics research infrastructure in the United States. The European contribution of this review lies primarily in the regulatory analysis—mapping the interplay of the EU AI Act, MDR, and GDPR for quality management applications—and in the proposed implementation framework, which was derived through author consensus informed by GAMP 5 principles, EU AI Act risk categorization, and professional experience at three German academic medical centers rather than through empirical pilot data or formal expert elicitation methods such as Delphi. Additionally, the transferability of findings across European healthcare systems warrants consideration. Quality management structures differ substantially between, for example, the German Cancer Society (DKG) certification system in Germany and accreditation frameworks in other EU member states. Regulatory interpretation of the AI Act's risk categories may also vary across national implementation contexts. Accordingly, the proposed framework should be understood as a conceptual starting point requiring empirical validation through pilot implementations across diverse European healthcare settings. Future research Several directions merit priority attention. First, prospective implementation studies in European healthcare settings are essential to establish whether the efficiency gains and concordance rates observed in U.S. and Asian contexts translate to European quality management workflows. Second, the development of validation frameworks specifically designed for non-deterministic LLM systems in regulated healthcare environments remains a critical gap; current quality management standards (ISO 13485, GAMP 5) were designed for deterministic software and require adaptation. Third, research on the interaction between LLM deployment and organizational change in quality management departments would help healthcare leaders anticipate and manage workforce implications. Fourth, multilingual evaluation of LLM performance is needed, as most published evidence derives from English-language settings, yet European healthcare operates across dozens of languages. Finally, longitudinal studies tracking both efficiency gains and potential quality risks over extended deployment periods are essential before recommending broad-scale adoption. The convergence of LLMs with emerging technologies creates opportunities for predictive quality management that anticipates rather than merely documents quality related events. This vision—moving from reactive documentation to proactive quality intelligence—represents the ultimate promise of LLM integration. Realizing it will require thoughtful, evidence-based implementation maintaining the high standards European healthcare demands. Funding Statement The author(s) declared that financial support was received for this work and/or its publication. The Stuttgart Cancer Center – Tumorzentrum Eva Mayr-Stihl at Klinikum Stuttgart receives institutional funding from the Eva Mayr-Stihl Stiftung, Waiblingen, Germany. The funder had no role in the conceptualization, writing, or decision to submit this manuscript. Footnotes Edited by: Fahim Sufi , Monash University, Australia Reviewed by: Krishna Jayanth Rolla , Independent Researcher, Fort Mill, United States David J. Bunnell , University of Maryland, United States Author contributions MKn: Conceptualization, Methodology, Writing – original draft, Writing – review & editing. MKr: Conceptualization, Visualization, Writing – review & editing. AK: Writing – review & editing. Conflict of interest The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Generative AI statement The author(s) declared that generative AI was used in the creation of this manuscript. The authors acknowledge the use of Claude Opus 4.5 (Anthropic, San Francisco, CA, USA) for language editing and manuscript preparation assistance. The AI was instructed to optimize readability and formal manuscript style without modifying scientific content, conclusions, or references. The prompt used is available in the Supplementary Material . All AI-generated suggestions were reviewed, verified, and edited by the authors, who take full responsibility for the accuracy and integrity of the final manuscript. Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us. Publisher's note All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher. Supplementary material The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fdgth.2026.1761641/full#supplementary-material Datasheet1.docx (15.8KB, docx) References 1. Wismar M, Goffin T. Tackling the health workforce crisis: towards a European health workforce strategy. Eurohealth (Lond). (2023) 29(3):22–6. [ Google Scholar ] 2. Parliament, European, Transformation Directorate-General for Economy, Industry, Kuhlmann E. The health workforce crisis in the European union – policy options for improving the sustainability of healthcare systems and employment and working conditions in the healthcare sector. Eur Parliament. (2025). 10.2861/3098897 [ DOI ] [ Google Scholar ] 3. Busch F, Hoffmann L, Rueger C, van Dijk EH, Kader R, Ortiz-Prado E, et al. Current applications and challenges in large language models for patient care: a systematic review. Commun Med. (2025) 5(1):26. 10.1038/s43856-024-00717-2 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and evaluation of health care applications of large language models: a systematic review. JAMA. (2025) 333(4):319–28. 10.1001/jama.2024.21700 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. European Parliament and Council of the European Union. Regulation (Eu) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with Regard to the Processing of Personal Data and on the Free Movement of Such Data, and Repealing Directive 95/46/Ec (General Data Protection Regulation). (2016). Available online at: https://eur-lex.europa.eu/eli/reg/2016/679/oj (GDPR) (Accessed December 4, 2025). 6. Zhang Z, Yang X, Yao X, Yang H, Zhang S, Liu S, et al. Mrqc-Llm: a novel large language model framework for enhancing medical record quality control. Res. Sq. (2025). [Preprint]. 10.21203/rs.3.rs-6765575/v1 [ DOI ] [ Google Scholar ] 7. Georgios F, Papaspyridis K, Gkoulalas-Divanis A, Verykios VS. Privacy-Preserving techniques in generative ai and large language models: a narrative review. Information. (2024) 15(11):697. 10.3390/info15110697 [ DOI ] [ Google Scholar ] 8. Boussina A, Krishnamoorthy R, Quintero K, Joshi S, Wardi G, Pour H, et al. Large language models for more efficient reporting of hospital quality measures. Nejm ai. (2024) 1(11):AIcs2400420. 10.1056/aics2400420 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Union, European. Regulation (Eu) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act). Brussels: Official Journal of the European Union; (2024). [ Google Scholar ] 10. Xia X, Ma Y, Luo Y, Lu J. An online intelligent electronic medical record system via speech recognition. Int J Distrib Sensor Net. (2022) 18(11):15501329221134479. 10.1177/15501329221134479 [ DOI ] [ Google Scholar ] 11. Mairittha T, Mairittha N, Inoue S. Evaluating a spoken dialogue system for recording systems of nursing care. Sensors. (2019) 19(17):3736. 10.3390/s19173736 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Blackley SV, Huynh J, Wang L, Korach Z, Zhou L. Speech recognition for clinical documentation from 1990 to 2018: a systematic review. J Am Med Inform Assoc. (2019) 26(4):324–38. 10.1093/jamia/ocy179 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Nelson SD, Aguero D. The potential application of large language models in pharmaceutical supply chain management. J Pediatr Pharmacol Ther. (2024) 29(2):200–05. 10.5863/1551-6776-29.2.200 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Kourani H, Berti A, Schuster D, van der Aalst WM. Evaluating large language models on business process modeling: framework, benchmark, and self-improvement analysis. Software and Systems Modeling. (2025):1–36. 10.1007/s10270-025-01318-w [ DOI ] [ Google Scholar ] 15. Decker H, Trang K, Ramirez J, Colley A, Pierce L, Coleman M, et al. Large language model− based chatbot vs surgeon-generated informed consent documentation for common procedures. JAMA network Open. (Oct 2 2023) 6(10):e2336997–e97. 10.1001/jamanetworkopen.2023.36997 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Shi Q, Luzuriaga K, Allison JJ, Oztekin A, Faro JM, Lee JL, et al. Transforming informed consent generation using large language models: mixed methods study. JMIR Med Inform. (2025) 13(1):e68139. 10.2196/68139 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Omar M, Nadkarni GN, Klang E, Glicksberg BS. Large language models in medicine: a review of current clinical trials across healthcare applications. PLOS Digit Health. (2024) 3(11):e0000662. 10.1371/journal.pdig.0000662 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Baker HP, Dwyer E, Kalidoss S, Hynes K, Wolf J, Strelzow JA. Chatgpt’s ability to assist with clinical documentation: a randomized controlled trial. J Am Acad Orthop Surg. (2024) 32(3):123–29. 10.5435/jaaos-d-23-00474 [ DOI ] [ PubMed ] [ Google Scholar ] 19. Gao M, Varshney A, Chen S, Goddla V, Gallifant J, Doyle P, et al. The use of large language models to enhance cancer clinical trial educational materials. JNCI Cancer Spectrum. (2025) 9(2):pkaf021. 10.1093/jncics/pkaf021 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Lin A, Wang Z, Jiang A, Chen L, Qi C, Zhu L, et al. Large language models in clinical trials: applications, technical advances, and future directions. BMC Med. (2025) 23(1):563. 10.1186/s12916-025-04348-9 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Allen JW, Schaefer O, Mann SP, Earp BD, Wilkinson D. Augmenting research consent: should large language models (llms) be used for informed consent to clinical research? Res Ethics. (2025) 21(4):644–70. 10.1177/17470161241298726 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Katsoulakis E, Wang Q, Wu H, Shahriyari L, Fletcher R, Liu J, et al. Digital twins for health: a scoping review. NPJ Digit Med. (2024) 7(1):77. 10.1038/s41746-024-01073-0 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Meijer C, Uh H-W, Bouhaddani SE. Digital twins in healthcare: methodological challenges and opportunities. J Pers Med. (2023) 13(10):1522. 10.3390/jpm13101522 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Vallée A. Digital twin for healthcare systems. Front Digit Health. (2023) 5:1253050. 10.3389/fdgth.2023.1253050 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Shi D, Liedl P, Bauernhansl T. Interoperable information modelling leveraging asset administration shell and large language model for quality control toward zero defect manufacturing. J Manufact Syst. (2024) 77:678–96. 10.1016/j.jmsy.2024.10.011 [ DOI ] [ Google Scholar ] 26. De Ketelaere B, Smeets B, Verboven P, Nicolaï B, Saeys W. Digital twins in quality engineering. Qual Eng. (2022) 34(3):404–08. 10.1080/08982112.2022.2052731 [ DOI ] [ Google Scholar ] 27. Nitschke AK, Diaz Ochoa JG, Neumaier S, Knott M. From knowledge graphs to digital twins: perspectives on modeling patient outcomes for healthcare quality assessment. J Med Internet Res. (2026) 28:81946. 10.2196/81946 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Diaz Ochoa JG, Layer N, Mahr J, Mustafa FE, Menzel CU, Müller M, et al. Optimized bert-based nlp outperforms zero-shot methods for automated symptom detection in clinical practice. Front Digit Health. (2025) 7:1623922. 10.3389/fdgth.2025.1623922 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Chen CM, Lin IL. Sop-Gpt: a framework for ai agents based on artificial intelligence-generated content. Paper Presented at the 2024 IEEE 7th Eurasian Conference on Educational Innovation (ECEI) (2024). [ Google Scholar ] 30. Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-Augmented generation for knowledge-intensive Nlp Tasks. Adv Neural Inf Process Syst. (2020) 33:9459–74. [ Google Scholar ] 31. Wang Z, Wang Z, Le L, Steven Zheng H, Mishra S, Perot V, et al. Speculative rag: enhancing retrieval augmented generation through drafting. ar X iv preprint arXiv:2407.08223 . (2024). 10.48550/arXiv.2407.08223 [ DOI ] [ Google Scholar ] 32. Klesel M, Felix Wittmann H, Klesel M, Wittmann H. Retrieval-Augmented generation (rag). Bus Inform Syst Eng. (2025) 67:1–11. 10.1007/s12599-025-00945-3 [ DOI ] [ Google Scholar ] 33. Álvaro JAH, Barreda JG. An advanced retrieval-augmented generation system for manufacturing quality control. Adv Eng Inform. (2025) 64:103007. 10.1016/j.aei.2024.103007 [ DOI ] [ Google Scholar ] 34. Zhang C. Intelligent process automation in audit. J Emerg Technol Account. (2019) 16(2):69–88. 10.2308/jeta-52653 [ DOI ] [ Google Scholar ] 35. Huang F, Vasarhelyi MA. Applying robotic process automation (rpa) in auditing: a framework. Int J Account Inform Syst. (2019) 35:100433. 10.1016/j.accinf.2019.100433 [ DOI ] [ Google Scholar ] 36. Kokina J, Blanchette S. Early evidence of digital labor in accounting: innovation with robotic process automation. Int J Account Inform Syst. (2019) 35:100431. 10.1016/j.accinf.2019.100431 [ DOI ] [ Google Scholar ] 37. van der Vegt AH, Scott IA, Dermawan K, Schnetler RJ, Kalke VR, Lane PJ. Implementation frameworks for End-to-End clinical ai: derivation of the salient framework. J Am Med Inform Assoc. (2023) 30(9):1503–15. 10.1093/jamia/ocad088 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Moffitt KC, Rozario AM, Vasarhelyi MA. Robotic process automation for auditing. J Emerg Technol Account. (2018) 15(1):1–10. 10.2308/jeta-10589 [ DOI ] [ Google Scholar ] 39. Ong JCL, Chang SY-H, William W, Butte AJ, Shah NH, Chew LST, et al. Ethical and regulatory challenges of large language models in medicine. Lancet Digit Health. (2024) 6(6):e428–e32. 10.1016/S2589-7500(24)00061-X [ DOI ] [ PubMed ] [ Google Scholar ] 40. Eguia H, Sánchez-Bocanegra CL, Vinciarelli F, Alvarez-Lopez F, Saigí-Rubió F. Clinical decision support and natural language processing in medicine: systematic literature review. J Med Internet Res. (2024) 26:e55315. 10.2196/55315 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Meskó B, Topol EJ. The imperative for regulatory oversight of large language models (or generative ai) in healthcare. NPJ Digit Med. (2023) 6(1):120. 10.1038/s41746-023-00873-0 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Papageorgiou K, Theodosiou T, Rapti A, Papageorgiou EI, Dimitriou N, Tzovaras D, et al. A systematic review on machine learning methods for root cause analysis towards zero-defect manufacturing. Frontiers in Manufacturing Technology. (2022) 2:972712. 10.3389/fmtec.2022.972712 [ DOI ] [ Google Scholar ] 43. Razzak MI, Imran M, Xu G. Big data analytics for preventive medicine. Neural Computing and Applications. (2020) 32(9):4417–51. 10.1007/s00521-019-04095-y [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 44. de Hond A, Leeuwenberg T, Bartels R, van Buchem M, Kant I, Moons KG, et al. From text to treatment: the crucial role of validation for generative large language models in health care. Lancet Digit Health. (2024) 6(7):e441–e43. 10.1016/S2589-7500(24)00111-0 [ DOI ] [ PubMed ] [ Google Scholar ] 45. Montgomery DC. Introduction to Statistical Quality Control. Hoboken: John wiley & sons; (2020). [ Google Scholar ] 46. Wheeler DJ, Chambers DS. Understanding Statistical Process Control. Knoxville, TN: Knoxville; (1992). [ Google Scholar ] 47. Stanton NA, Salmon PM, Rafferty LA, Walker GH, Baber C, Jenkins DP. Human Factors Methods: A Practical Guide for Engineering and Design. Boca Raton, FL: CRC Press; (2017). [ Google Scholar ] 48. Overgaard SM, Graham MG, Brereton T, Pencina MJ, Halamka JD, Vidal DE, et al. Implementing quality management systems to close the ai translation gap and facilitate safe, ethical, and effective health ai solutions. NPJ Digital Medicine. (2023) 6(1):218. 10.1038/s41746-023-00968-8 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 49. Reddy S, Allan S, Coghlan S, Cooper P. A governance model for the application of ai in health care. J Am Med Inform Assoc. (2020) 27(3):491–97. 10.1093/jamia/ocz192 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Wells BJ, Nguyen HM, McWilliams A, Pallini M, Bovi A, Kuzma A, et al. A practical framework for appropriate implementation and review of artificial intelligence (fair-ai) in healthcare. NPJ Digit Med. (2025) 8(1):514. 10.1038/s41746-025-01900-y [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 51. AlSaad R, Abd-Alrazaq A, Boughorbel S, Ahmed A, Renault M-A, Damseh R, et al. Multimodal large language models in health care: applications, challenges, and future outlook. J Med Internet Res. (2024) 26:e59505. 10.2196/59505 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 52. European Commission, Health and Consumers Directorate-General. Eu Guidelines for Good Manufacturing Practice for Medicinal Products for Human and Veterinary Use. Annex 11: Computerised Systems. EudraLex Volume 4, 2011. (2011). Available online at: https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf (Accessed December 4, 2025). [ Google Scholar ] 53. Engineering, International Society for Pharmaceutical. Ispe Gamp 5: A Risk-Based Approach to Compliant Gxp Computerized Systems. 2nd edn ed Bethesda, MD: ISPE; (2022). [ Google Scholar ] 54. European, Parliament, and Union Council of the European. Regulation (Eu) 2017/745 of the European Parliament and of the Council of 5 April 2017 on Medical Devices, Amending Directive 2001/83/Ec, Regulation (Ec) No 178/2002 and Regulation (Ec) No 1223/2009 and Repealing Council Directives 90/385/Eec and 93/42/Eec. Luxembourg: Publications Office of the European Union; (2017). [ Google Scholar ] 55. Group, Medical Device Coordination, and Artificial Intelligence Board. Interplay between the Medical Devices Regulation (Mdr) & in vitro Diagnostic Medical Devices Regulation (Ivdr) and the Artificial Intelligence Act (Aia): Mdcg 2025-6. Medical Device Coordination Group and Artificial Intelligence Board (Brussels: 2025). Brussels: European Commission, Directorate-General for Health and Food Safety (2025). Available online at: https://health.ec.europa.eu/document/download/b78a17d7-e3cd-4943-851d-e02a2f22bbb4_en?filename=mdcg_2025-6_en.pdf (Accessed December 4, 2025). [ Google Scholar ] 56. Parliament, European Union European, and Council of the European Union. Regulation (eu) 2025/327 of the European parliament and of the council of 11 February 2025 on the European health data space and amending directive 2011/24/eu and regulation (eu) 2024/2847 (text with eea relevance). Official Journal of the European Union. (2025). [ Google Scholar ] 57. https://picscheme.org/docview/4234 “Pic/S Good Practices for Data Management and Integrity in Regulated Gmp/Gdp Environments (Pi 041-1)” PIC/S Secretariat, 2021, Accessed 4 December 2025. Available online at: 58. https://blog.seerpharma.com/seerpharma-and-melbourne-uni-report-impact-of-ai-llms-on-quality-and-gmp “Systematic Analysis of Llm Applications in Pharmaceutical Quality Management Systems” SeerPharma, 2024. Accessed December 4, 2025. Available online at: 59. Standardization, International Organization. Medical devices — quality management systems — requirements for regulatory purposes. ISO 13485:2016. Geneva: International Organization for Standardization, 2016. [ Google Scholar ] 60. Standardization, International Organization for, and International Electrotechnical Commission. Iso/Iec 42001:2023. Geneva: International Organization for Standardization; (2023). [ Google Scholar ] 61. Guideline for Good Clinical Practice E6(R3). ICH, 2025, Accessed 4 December 2025. (2025). Available online at: https://database.ich.org/sites/default/files/ICH_E6%28R3%29_Step4_FinalGuideline_2025_0106.pdf 62. Chen J, Saha S, Bansal M. ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: Association for Computational Linguistics; (2024). p. 7066–85. 10.18653/v1/2024.acl-long.381 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 63. Omar M, Glicksberg BS, Nadkarni GN, Klang E. Refining LLMs outputs with iterative consensus ensemble (ICE). Comput Biol Med. (2025) 196(Pt B):110731. 10.1016/j.compbiomed.2025.110731 [ DOI ] [ PubMed ] [ Google Scholar ] 64. OECD/European Commission. Health at a Glance: Europe 2024: State of Health in the EU Cycle. Paris: OECD Publishing; (2024). 10.1787/b3704e14-en [ DOI ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Datasheet1.docx (15.8KB, docx) Articles from Frontiers in Digital Health are provided here courtesy of Frontiers Media SA ACTIONS View on publisher site PDF (505.3 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top