Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice NPJ Digit Med . 2026 Feb 27;9:308. doi: 10.1038/s41746-026-02433-8 Search in PMC Search in PubMed View in NLM Catalog Add to search Advancing medical AI through benchmarking and competition for specialty triage Chao Ding Chao Ding 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China 2 Putuo Hospital, Shanghai University of Traditional Chinese Medicine, Shanghai, China Find articles by Chao Ding 1, 2, # , Mouxiao Bian Mouxiao Bian 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China Find articles by Mouxiao Bian 1, # , Minjia Yuan Minjia Yuan 3 Joint Laboratory of Biomedical Artificial Intelligence, Shanghai East Hospital, Shanghai, China Find articles by Minjia Yuan 3, # , Luyi Jiang Luyi Jiang 4 Shanghai Institute of Infectious Disease and Biosecurity, Fudan University, Shanghai, China 5 Shanghai Health Development Research Center(Shanghai Medical Information Center), Shanghai, China Find articles by Luyi Jiang 4, 5 , Kaiyi Luo Kaiyi Luo 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China Find articles by Kaiyi Luo 1 , Pengcheng Chen Pengcheng Chen 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China 6 University of Washington, Seattle, WA USA Find articles by Pengcheng Chen 1, 6 , Yuanye Jiang Yuanye Jiang 2 Putuo Hospital, Shanghai University of Traditional Chinese Medicine, Shanghai, China Find articles by Yuanye Jiang 2, ✉ , Jie Xu Jie Xu 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China Find articles by Jie Xu 1, ✉ Author information Article notes Copyright and License information 1 Shanghai Artificial Intelligence Laboratory, Shanghai, China 2 Putuo Hospital, Shanghai University of Traditional Chinese Medicine, Shanghai, China 3 Joint Laboratory of Biomedical Artificial Intelligence, Shanghai East Hospital, Shanghai, China 4 Shanghai Institute of Infectious Disease and Biosecurity, Fudan University, Shanghai, China 5 Shanghai Health Development Research Center(Shanghai Medical Information Center), Shanghai, China 6 University of Washington, Seattle, WA USA ✉ Corresponding author. # Contributed equally. Received 2025 Apr 8; Accepted 2026 Feb 1; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13066094 PMID: 41760911 Abstract Artificial intelligence holds transformative potential for clinical triage, yet challenges in accuracy, generalization, and interpretability persist. To address these gaps, we introduce MedTriage, a benchmark designed to evaluate large-scale models across diverse clinical scenarios rigorously. Leveraging this framework, we launched the Large-Model-Based Medical Triage Evaluation Competition, utilizing real-world clinician-patient dialogues from general hospitals and four specialized domains. The competition engaged numerous research teams, spurring advancements in large-model-driven triage algorithms. Building on the competition insights, we developed an enhanced model (MedGPT-Guide) employing a “10 Relevant + 10 Random + Ensemble” strategy, achieving superior accuracy on the MedTriage benchmark. Our results underscore the power of “evaluation-driven training” to improve model performance and lay the groundwork for standardized, deployable intelligent triage systems. Moving forward, priorities include enhancing data security, model generalization, and addressing legal and regulatory frameworks. Subject terms: Computational biology and bioinformatics, Data acquisition Introduction Medical triage serves as the pivotal first step in patient care, profoundly influencing treatment efficacy and outcomes 1 , 2 . By enabling rational allocation of limited medical resources, it not only ensures timely interventions but also mitigates risks from triage errors 3 , 4 . However, current systems remain constrained by heavy reliance on labor-intensive human expertise 5 , 6 , particularly in resource-constrained primary care settings where inadequate triage capabilities exacerbate healthcare disparities and misdiagnosis rates 7 , 8 . Conventional symptom checkers (SCs) exhibit substantial limitations in clinical diagnostic performance. Empirical studies reveal their diagnostic accuracy ranges between merely 14–27% across medical specialties, while demonstrating critical failure in detecting life-threatening conditions in 13–14% of clinical presentations 9 . More alarmingly, suboptimal triage accuracy may inadvertently delay critical diagnoses, potentially resulting in dual-system failures: misallocation of finite healthcare resources and deprivation of time-sensitive interventions for critically ill populations. The integration of Artificial Intelligence (AI), particularly Large Language Models (LLMs), is revolutionizing patient flow management by enhancing triage efficiency and optimizing resource allocation. LLM-driven triage systems demonstrate significant potential in automating medical history-taking, improving triage/routing performance, and ensuring timely interventions while reducing the workload of healthcare professionals 10 – 12 . Unlike conventional symptom checkers with inconsistent clinical performance, LLMs leverage advanced natural language processing to provide real-time, evidence-based recommendations 13 – 15 . However, challenges such as diagnostic reliability, integration with existing infrastructures, and clinician acceptance must be addressed to fully realize the potential of AI-driven triage in improving healthcare accessibility and patient outcomes 16 , 17 . We focus on specialty triage—routing a patient to the most appropriate clinical department(s)—rather than producing diagnoses or treatment plans. In practice, such a system can be integrated at multiple touchpoints in the patient journey: (i) pre-visit digital intake (e.g., online guidance or appointment portals), where patients describe symptoms in free text or multi-turn chats; (ii) on-site registration/triage, where only brief chief complaints and demographics are available; and (iii) during-visit decision support, where clinicians document semi-structured notes that summarize the chief complaint and history. Correspondingly, MedTriage includes three input modalities—short complaint-driven records, semi-free-text clinical notes, and multi-turn dialogue logs—allowing systematic evaluation across realistic deployment settings with varying information availability. Existing large-scale medical dialogue datasets such as MedDialog 18 contain 3.4 million Chinese and 0.26 million English patient-doctor conversations covering 172 and 96 specialties, respectively, yet focus predominantly on dialogue generation and not on multi-department triage decision routing. Another example is the HealthCareMagic 19 dataset, which consists of about 100k English online consultations but remains oriented for conversational response generation rather than structured triage outcome classification. Concretely, dialog-generation datasets optimize open-ended outputs such as “Your symptoms may be related to gastritis; consider dietary adjustments and seek medical care if they worsen,” whereas our benchmark requires a constrained categorical output (e.g., {Gastroenterology}, or a small set of departments) selected from a predefined label list under an exact-match multi-label evaluation. To address these challenges, we developed MedTriage--a novel benchmark framework for assessing LLM performance in clinical consultation guidance and medical reasoning. Through our organized “Evaluation of Medical Consultation Patterns Based on Big Models” competition, we established standardized evaluation protocols across five clinical specialties: general medicine, pediatrics, obstetrics/gynecology, stomatology, and traditional Chinese medicine (TCM). This initiative provides a unified dataset and multidimensional metrics to objectively quantify hospital-specific consultation patterns, establishing foundational infrastructure for field advancement. We further conducted an in-depth analysis of all participating teams’ results, which allowed us to identify key factors for model improvement and embodied the concept of “Evaluation for Practice” by leveraging an open evaluation platform to foster higher-level solutions. Building on these insights, we refined our algorithmic framework and engineering pipelines to develop enhanced models. We build a post-competition reference model and report controlled ablations under the official protocol. Our main contributions are as follows: We construct MedTriage, a large-scale benchmark for medical triage derived from real-world doctor–patient dialogues across five specialties. We establish strong baseline results across representative large language models (LLMs) on this task, providing transparent and reproducible evaluation. We release an open benchmarking platform (MedBench) that enables fair comparison and competition among participants. We conduct the first MedBench competition, analyze participating systems, and summarize trends for future development. Results Figure 1 provides a flowchart of the structure. Fig. 1. MedTriage framework for evaluating LLMs in clinical guidance. Open in a new tab This framework benchmarks large language models across six clinical domains using data from discharge summaries, outpatient records, and online consultations. Data undergoes strict filtering and cross-checking labeling by medical professionals to ensure quality. Based on guidance accuracy, models are evaluated through a standardized pipeline—deployment, testing, scoring, and leaderboard ranking. Analysis of the benchmark dataset The benchmark dataset includes patient interaction records from two hospital categories: general hospitals and specialized institutions. The latter is further divided into pediatric, obstetrics/gynecology, stomatology, and TCM facilities. Data collection employed robotic triage dialogue interfaces, which captured anonymized doctor-patient conversations along with demographic metadata such as gender and age. Interaction logs were filtered from operational guidance systems, with comprehensive dialogue templates. The distributions of specialty departments align with China’s Guidelines on Medical Service Capabilities of Tertiary General Hospitals, and the proportions of specialty-specific cases are visualized in Fig. 2 . Fig. 2. Distribution of medical scenarios in MedTriage dataset. Open in a new tab The dataset covers five major clinical scenario types: general outpatient (25%), specialist outpatient (25%), discharge summaries (25%), online consultations (15%), and others (10%). This balanced design supports the comprehensive evaluation of LLM performance across real-world medical guidance tasks. Data distribution In general hospitals, the top five most recommended departments are internal medicine, surgery, pain management, pediatrics, and otolaryngology. Additionally, TCM practices, acupuncture and moxibustion (116), and tuina (112) rank in the top ten, highlighting a blend of Chinese and Western medical approaches in clinical services. Among specialized institutions, in TCM hospitals, the most recommended departments are Acupuncture (116) and Tuina (112), far surpassing other areas. This focus underscores TCM hospitals’ emphasis on traditional, well-established therapies, highlighting their role in providing effective disease guidance through proven practices. Dental hospitals are highly specialized, featuring key departments like pediatric dentistry(105), endodontics(104), periodontics(72), prosthodontics(66), oral and maxillofacial surgery(62), implantology(48), and orthodontics(39). These specialties allow them to provide precise, targeted care for a wide range of oral health needs. Obstetrics and gynecology hospitals, on the other hand, 98.4% specialize in gynecology. Finally, pediatric hospitals cater to children’s unique health challenges, with respiratory medicine (108 recommendations), orthopedics (67), growth and development (54), and urology (31) emerging as the most recommended departments. These areas reflect a commitment to managing both acute conditions and long-term developmental concerns in young patients (Fig. 3 ). Fig. 3. Distribution of LLM predictions across clinical departments. Open in a new tab Each row represents a clinical department, the bars (left) show the most common true labels, and the donut chart (right) shows the number of hospital recommendations. a , b General hospitals: Predictions and recommendations span a broad set of departments, reflecting the comprehensive service scope of general institutions. c , d Traditional Chinese Medicine (TCM) hospitals: Acupuncture and Tuina emerge as dominant departments in both predictions and hospital guidance, highlighting the institutions’ emphasis on traditional therapies. e , f Dental hospitals: Fine-grained specialties such as pediatric dentistry, endodontics, and oral surgery are well captured by the model and mirrored in the hospital recommendations. g , h Obstetrics and gynecology hospitals: Predictions and hospital recommendations are highly concentrated in gynecology, indicating a strong single-specialty orientation. i , j Pediatric hospitals: Frequently predicted departments include respiratory medicine, orthopedics, and growth and development, corresponding closely to pediatric clinical needs. Model performance The MedTriage benchmark revealed substantial performance variations among five LLMs in medical guidance scenarios. GPT-4o has the highest accuracy of 0.4212 and performs the best among all the models. Claude 3.5 has the lowest accuracy of 0.2356, which is a gap compared to the other models. DeepSeek-v3 and DeepSeek-R1 have an accuracy of 0.3238, which is an average performance in this experiment. This suggests that the different versions of the DeepSeek model have similar accuracies in this experimental task and may be more stable in their reasoning abilities. O1-mini has an accuracy of 0.3769, which is the second highest accuracy in this experiment after GPT-4o. The model outperforms the DeepSeek family, but still falls short of GPT-4o (Fig. 4 ). Fig. 4. Accuracy Comparison of Leading LLMs in Medical Guidance. Open in a new tab GPT-4o achieves the highest accuracy (~43%), followed by 01-mini and DeepSeek-R1. DeepSeek-V3 performs similarly to its predecessor, while Claude-3.5 shows the lowest accuracy in this evaluation. This highlights notable variability in medical reasoning performance across models. Performance deficiency analysis Cross-validation analysis revealed distinct error patterns between development (dev) and test sets (Table 1 ). Claude 3.5 demonstrated test set overgeneralization with extreme over-inclusion errors (test:1045 vs dev:408), while DeepSeek variants maintained consistent over-recommendation tendencies across both sets (test:1082; dev:1067). GPT-4o showed superior dev set performance in error control (dev:616 vs test:743), though its test set precision remained comparatively higher than competitors. Complete misclassification rates exposed critical disparities (Fig. 5 ): Claude 3.5’s dev set mapping failures ( n = 408) tripled its test set errors ( n = 136), suggesting training data contamination. Conversely, 4o1-min achieved enhanced test set precision ( n = 164) versus dev performance ( n = 211). Omission errors highlighted divergent strategies - 4o1min’s conservative test set approach (omissions:123) intensified in dev environments ( n = 151), while DeepSeek models sustained balanced coverage (test/dev:44). Table 1. Comparison of Error Types Across LLMs in Medical Department Recommendation Model Extra departments Errors Wrong department Errors Missing department Errors DeepSeek V3 (dev) 896 142 44 DeepSeek V3 (test) 1082 194 48 DeepSeek R1(dev) 896 142 44 DeepSeek R1(test) 1082 194 48 4o1min(dev) 616 212 98 4o1min(test) 710 164 123 GPT-4o(dev) 616 212 98 GPT-4o(test) 833 225 151 Claude3.5(dev) 770 408 45 Claude3.5(test) 1045 363 54 Open in a new tab Fig. 5. Comparison of error types across LLMs in medical department recommendation. Open in a new tab The figure categorizes model errors into three types: recommending extra departments, assigning the wrong department, and omitting needed departments. Claude-3.5 and DeepSeek variants show high rates of extra and wrong department errors. GPT-4o and 01-mini reduce these errors but exhibit more missing department errors, especially in the test set. The balance of error types reveals distinct model tendencies in clinical reasoning and coverage. Case-specific analysis The patient is a 74-year-old male who underwent gastric tumor surgery and reports postoperative fatigue (Fig. 6 ). The task requires selecting the most appropriate department from a predefined list of medical specialties. Despite correctly identifying Oncology as the primary department, all models erroneously included Geriatrics and Rehabilitation Medicine, leading to “Over-inclusion” errors. Fig. 6. Example of department recommendation task for medical LLMs. Open in a new tab The task of the model is to select the appropriate department from a fixed list of candidates based on the patient’s description. In this example, a 74-year-old male patient feels weak after surgery for a stomach tumor, and the model recommends a department based on clinical expectations. Overall competition performance Competition participation and solution themes The MedBench triage competition attracted 37 teams spanning clinical institutions, research organizations, and industry participants. Table 2 reports the top 13 teams on the final leaderboard under the official evaluation protocol; lower-ranked teams are omitted for brevity. Table 2. The performance of the participating organizations in the competition Organization Name Test set Validation set Accuracy Difference (%) Xinhua Hospital, Shanghai Jiao Tong University School of Medicine 76.30 73.5 2.80 The First Affiliated Hospital of Nanchang University 73.45 74.19 −0.74 Shandong Academy of Sciences 67.75 70.75 −3.00 Shanghai Chest Hospital 65.40 61.69 3.71 North Healthcare Big Data Technology Co. 64.95 66.19 −1.24 Nanjing Jiangbei New Area Biomedical Public Service Platform 63.55 69.63 −6.08 The First Affiliated Hospital of Sun Yat-sen University-1 61.20 63.88 −2.68 Affiliated Hospital of Sichuan Bei Medical College 60.35 68.06 −7.71 The First Affiliated Hospital of Sun Yat-sen University-2 60.35 61.69 −1.34 Shanghai Jianjiao Technology Service Co. 57.20 60.62 −3.42 Shenzhen Big Data Research Institute Wuxi Innovation Center 27.55 56.56 −29.01 Beijing Friendship Hospital of Capital Medical University 18.55 69.69 −51.14 Wuhan University Central South Hospital (excluded) 0 22.13 Open in a new tab Leaderboard results and generalization The evaluation framework employed balanced validation and test sets (2,000 instances each). Xinhua Hospital (Shanghai Jiao Tong University School of Medicine) demonstrated optimal robustness with superior (test 76.30% vs validation 73.50%, ∧+2.80% ). Beijing Friendship Hospital (Capital Medical University) exhibited critical performance degradation (test18.55% vs validation 69.69%, ∨ - 51.14% ), showed large performance drops from validation to test. Shenzhen Big Data Research Institute’s Wuxi Innovation Center revealed severe generalization limitations (test 27.55% vs validation 56.56%, ∨ - 29.01% ). Wuhan University Central South Hospital failed test deployment (validation baseline:22.13%). Remaining institutions maintained <10% accuracy variance across sets: Xinhua Hospital improved its performance on the test set, indicating robustness. The First Affiliated Hospital of Nanchang University maintained a stable performance on both test sets, indicating good generalizability of the model (Table 2 ). Post-competition reference model Because participant method disclosures were not mandatory, we summarize “what works” via a reproducible reference model (MedGPT-Guide) and controlled ablations evaluated under the official protocol.The baseline model achieves an accuracy of 0.5319[0.5100, 0.5538], setting a reference point for evaluating the impact of various algorithmic enhancements. Adding Chain-of-Thought (CoT) 20 reasoning improves accuracy to 0.6369[0.6158, 0.6580], indicating that structured step-by-step reasoning significantly enhances the model’s decision-making ability. This suggests that the model benefits from explicit reasoning prompts, likely improving its capacity to extract and organize relevant information before making a prediction. Incorporating few-shot 21 examples ( 10 most relevant examples and 10 randomly selected ones ) further increases accuracy to 0.6625[0.6418, 0.6832]. However, removing CoT while keeping these 20 examples results in an even higher accuracy of 0.7131[0.6933, 0.7329], suggesting example-based pattern recognition outweighs isolated reasoning frameworks. Ensemble 22 choice optimization achieved peak accuracy 0.7830[0.7649, 0.8011], surpassing CoT-only configurations by 14.61 percentage points. Beyond accuracy, our model (MedGPT-Guide) demonstrates strong performance in both precision and recall. Specifically, it achieves a precision of 0.8419 [0.8259, 0.8579] and a recall of 0.7884 [0.7705, 0.8063]. These results indicate not only more accurate predictions, but also better coverage of relevant departments. Detailed comparisons of recall and precision across models are presented in Supplementary Fig. 1 , Supplementary Fig. 2 , and Fig. 7 . Fig. 7. Accuracy in medical department recommendation. Open in a new tab The figure shows that stepwise enhancements—adding chain-of-thought reasoning, relevant and random examples, and ensemble methods—progressively improve model accuracy from 53.19%[0.5100, 0.5538] to 78.30%[0.7649, 0.8011]. This highlights the effectiveness of prompt design and model aggregation in clinical decision support tasks. Discussion The LLMs into medical triage systems presents transformative potential for healthcare delivery. Traditional symptom checkers have demonstrated limited diagnostic accuracy and critical failures in identifying life-threatening conditions 9 , 23 . The MedTriage benchmark, showing significantly improved precision through structured clinical reasoning, establishes a critical foundation for evaluating LLMs in medical triage. It reveals substantial performance variations that underscore the necessity of standardized assessment in real-world deployments. This advancement addresses systemic inefficiencies in resource-constrained settings, where patients frequently struggle with disease identification and department selection, often leading to delayed care 24 , 25 . Our competition-driven approach demonstrates how open benchmarking accelerates algorithmic innovation, with the top-performing model achieving 78.30% accuracy through iterative refinements of ensemble learning and example-based pattern recognition. These advancements align with prior evidence on AI’s potential to optimize resource allocation in emergency care, while extending current understanding by quantifying institutional-level performance disparities 26 . Notably, the competition’s “Evaluation for Practice” paradigm yielded two critical advancements: 1) Identification of optimal knowledge integration strategies (e.g., 20-example few-shot learning outperformed CoT reasoning by 7.62%), and 2) Demonstrated feasibility of Docker-based clinical deployment pipelines. These achievements align with recent calls for transparent AI evaluation in medicine, while exposing persistent gaps in real-world generalizability—best exemplified by Shenzhen Big Data Institute’s 29.01% test degradation. While some researchers consider AI systems to be high-risk for deployment in medical diagnostics 27 , 28 . It is imperative for both researchers and policymakers to develop a comprehensive understanding of AI systems’ strengths and weaknesses, so that informed decisions can be made regarding their safe application and improvement 29 , 30 . Three systemic challenges emerge from this analysis. First, algorithmic bias remains a persistent concern, as evidenced by models’ tendency to over-recommend departments for pediatric (500 instances) versus geriatric cases (5 instances), potentially exacerbating existing healthcare disparities. Second, privacy risks inherent in training data utilization—particularly for sensitive specialties like reproductive health (1 instance)—demand robust safeguards beyond current anonymization practices. Third, regulatory fragmentation across jurisdictions (e.g., HIPAA, GDPR, PIPL compliance) complicates cross-border model deployment, necessitating harmonized governance frameworks. Leaderboards are the basis for making research progress 31 . Even a set of benchmark leaderboards, an algorithm may improve the “fairness score” by improving the performance of one benchmark while making another benchmark less performant. Teams may have overfitted to the validation/test sets by refining hyperparameters extensively. Real-world performance might drop if the model encounters distinct patient populations or health systems. AI medical consultation models may inadvertently reinforce disparities in healthcare access. Not everyone can benefit equally from these technologies, leading to a “digital divide” in health care. For example, poor people, women, or black people are less likely to complete telemedicine visits, and people in already poor health tend to benefit less from new digital tools, which can widen healthcare disparities. The United Nations has warned that the lack of digital access is “a matter of life and death” for people who do not have access to basic health information, calling it “the new face of inequality” because it exacerbates existing social disadvantages 32 . Some studies have revealed such biases in healthcare algorithms, with real-world consequences. It routinely allowed healthier white patients into high-risk care programs ahead of sicker Black patients, because the algorithm used healthcare spending as a proxy for need by Obermeyer, Z et al. 33 . AI consultation models rely on large amounts of patient data. Protecting patient privacy while using this data is a major challenge. Techniques such as data masking, generalization or perturbation (adding slight noise) can help but are not foolproof 34 . Navigating the legal landscape is another major limitation for AI in healthcare. There are strict regulatory frameworks protecting patient information – such as the Health Insurance Portability and Accountability Act (HIPAA) in the U.S. 35 , the General Data Protection Regulation (GDPR) in the EU 36 , and the Personal Information Protection Law (PIPL) in China 37 – and AI models must be developed and deployed in compliance with these laws. Maintaining patient privacy in AI goes beyond technical compliance; it requires earnest respect for patient autonomy and control over personal health data. In summary, this study provides a systematic evaluation of a large model-based medical consultation model assessment competition, summarizing the current state and advancements in intelligent triage algorithms. More importantly, the competition established a benchmark for the field, driving research efforts toward enhancing the clinical utility of AI-driven consultation. As such evaluations continue and superior models emerge, intelligent triage systems are expected to evolve toward higher accuracy, improved generalization, and greater reliability. In addition, AI counseling systems to be useful in healthcare, they must be safe, fair, and responsible. This requires an interdisciplinary effort: engineers to improve model robustness and privacy, clinicians to ensure medical relevance and fairness, and ethicists and regulators to develop guidelines for responsible use. It is encouraging that frameworks focusing on the ethical and equity issues of AI are being developed 38 . By acknowledging and actively mitigating the limitations discussed – bias and inequity, privacy risks, and security threats – we can work towards AI solutions that augment healthcare without compromising patients’ rights or well-being. The path forward requires caution and care: deploying AI in medicine is not just a technical endeavor but a social one, where success will be measured not only by efficiency or accuracy but also by how well these tools serve all patients securely and ethically. Although this study showcases notable advancements in LLM-based triage systems, their relationship with established ‘old-school’ specialized systems is not one of mutual exclusion but rather of potential synergy. LLM-driven systems, as evaluated by MedTriage, demonstrate strength in broad, dialogue-based initial triage by understanding diverse, unstructured patient complaints to guide patients towards relevant medical fields. ‘Old-school’ specialized systems, like IDx-DR for detecting diabetic retinopathy from retinal images, conversely, offer high accuracy and reliability for specific, well-defined diagnostic tasks, often based on different data modalities. Therefore, an efficient synergistic workflow could involve LLMs performing initial wide-net triage, identifying patients who would benefit from the targeted diagnostics offered by these specialized AI tools. Such integration leverages the flexibility and broad coverage of LLMs with the depth and accuracy of specialized systems, potentially leading to more efficient and accurately resourced healthcare pathways. Exploring these hybrid approaches remains a valuable direction for future research. Methods Task definition We formulate specialty triage as a multi-label classification task, given patient-provided information collected at a specific point of the outpatient journey, the system predicts the most appropriate clinical department(s) from a predefined department list. This task is intended to support patient routing (e.g., online self-triage before registration, registration/triage desk decision support, or assisting clinicians when summarizing initial assessment notes), rather than to provide diagnoses or treatment recommendations. All data sources used in this benchmark are routinely generated in participating hospitals’ operational systems; the benchmark was constructed via retrospective export, de-identification, and quality control. In particular, the dialogue data come from online guidance systems (virtual-doctor/robot triage), whose text logs are routinely stored by the system; we did not collect audio/video recordings of in-person consultations. Benchmark design To construct the benchmark dataset, we retrospectively sourced data from both general hospitals and specialty institutions (pediatric, obstetrics and gynecology hospitals, stomatology, and TCM hospitals) to cover the five clinical domains evaluated in MedTriage. Department scope followed the National Health Commission guideline on tertiary general hospital service capabilities (2016). Overview of data sources Data were collected through three primary channels, each corresponding to a different stage of information availability in the patient journey: (1) Initial-visit intake (complaint-driven) records, which contain brief chief-complaint text and demographics captured at registration/triage; (2) Online guidance dialogues, consisting of multi-turn conversations between patients and a virtual doctor in hospitals equipped with online guidance systems, representing pre-visit digital triage; and (3) Outpatient clinical notes (semi-free text) documenting the initial assessment (e.g., chief complaint and history of present illness), reflecting richer in-visit information. Leakage prevention and labels For all channels, we treat diagnosis/assessment/management fields as metadata only and exclude them from the model input. Specifically, outpatient diagnosis, preliminary diagnosis/assessment, management, referral, and any department-indicative fields are removed from the input text before model inference. The ground-truth label is the department assignment recorded in the corresponding operational system: registration or encounter department metadata for Channels 1 and 3, and the final department recommended by the online guidance bot (stored as a structured field in the system) for Channel 2. Channel 1: initial-contact intake records (complaint-driven) This channel represents information available before clinician evaluation at first contact (e.g., online self-triage or front-desk registration/triage). Inputs are limited to patient demographics (age/sex) and short chief-complaint text captured at intake. Diagnosis, management, referral, and any department-related fields are excluded from the model input to avoid leakage. The label is obtained from structured department assignment metadata for the encounter. Inclusion criteria (1) The data should be selected from the category of initial diagnosis, and it is not recommended to include the data of outpatient clinics that have already had a clear diagnosis in the lower hospitals to seek further treatment. (2) The extracted data should ensure that the patient information is complete, and it is not recommended to include unknown gender and unspecified gender. (3) The data should cover common diseases and typical symptoms, and it is recommended to include some difficult cases. Exclusion criteria (1) Data on the existence of a clear diagnosis for further treatment. (2) Conversations involving the interpretation of pictures such as “see attached picture/see picture/following picture/picture/report/slice ……”. Conversations that involve interpretation of pictures, etc. (3) Non-disease and non-symptom data, such as booking an appointment or asking when the report will come out should not be included. (4) The extracted data should ensure that the patient information is complete, and it is not recommended to include unknown gender and unspecified gender. Example of data Patient information Sex: Male Age: 10 years Chief complaint: Burn injury to the right hand for several years, with hypertrophic scars on the middle and ring fingers. Open in a new tab Channel 2: online guidance dialogues This channel comprises multi-turn dialogue logs from hospitals operating online guidance systems, capturing natural-language conversations between patients and a virtual doctor (guidance bot). These dialogues correspond to pre-visit or pre-registration self-triage, and the text logs are routinely stored by the guidance platform as part of normal service operation (not transcriptions of in-person consultations). Inclusion criteria (1) The data should be selected from the category of initial diagnosis, and it is not recommended to include the data of outpatient clinics that have already had a clear diagnosis in the lower hospitals in order to seek further treatment. (2) The extracted data should ensure that the patient information is complete, and it is not recommended to include unknown gender and unspecified gender. (3) The data should cover common diseases and typical symptoms, and it is recommended to include some difficult cases. (4) Data on the dialog of dispensing and prescribing of medicines Exclusion criteria (1) Data on the existence of a clear diagnosis for further treatment (2) Conversations involving the interpretation of pictures, such as “see attached picture/see picture/following picture/picture/report/slice ……”. Conversations that involve interpretation of pictures, etc. (3) Non-disease and non-symptom data such as booking an appointment or asking when the report will come out should not be included. (4) The extracted data should ensure that the patient information is complete, and it is not recommended to include unknown gender and unspecified gender. Example of data: Dialogue content Recommended department Patient: “My left big toe feels hot and painful. Male, 27 years old.” Doctor: “Hello, did the pain start after trauma or a sprain?” Patient: “Yesterday, my big toe hit the leg of the bed.” Label (metadata; not part of input): Orthopedics. Open in a new tab Channel 3: outpatient clinical notes (semi-free text) Inputs include chief complaint and history of present illness (and other narrative fields when available), while preliminary diagnosis/assessment fields are treated as metadata and excluded from the model input to prevent trivially revealing the target specialty. This channel consists of initial-visit outpatient notes from electronic medical records, reflecting information available during the visit after clinicians have taken history and documented the encounter. It enables evaluation of department recommendation under richer, semi-structured clinical narratives. Inclusion criteria (1) It is recommended to select the data from the outpatient medical records of the initial diagnosis category, and to extract the patient information from the medical records with more standardized writing. The main complaint should include the main symptoms (or signs) and duration; the current medical history should cover the occurrence, evolution, diagnosis and treatment of the disease in detail, including: onset and urgency of the disease, antecedent symptoms, possible triggers, the location, nature, duration, degree of the main symptoms, factors that alleviate or exacerbate the disease and its evolution and development, concomitant symptoms, diagnosis and treatment and the general situation since the onset of the disease, and so on. (2) The completeness of patient information should be ensured. (3) The data should cover common diseases and typical symptoms, and it is recommended to include some difficult cases. Exclusion criteria (1) Data on the existence of a clear diagnosis for further treatment. (2) Conversations involving the interpretation of pictures such as “see attached picture/see picture/following picture/picture/report/slice ……”. Conversations that involve interpretation of pictures, etc. (3) Non-disease and non-symptom data such as booking an appointment or asking when the report will come out should not be included. (4) The extracted data should ensure that the patient information is complete, and it is not recommended to include unknown gender and unspecified gender. Example of data Patient information Recommended department Sex: Female Age: 35 years Chief complaint: Recurrent episodic headache for 3 years, worsened over the past 2 days. Current medical history: The patient first developed intermittent headaches 3 years ago, without obvious triggers. Headaches typically occur during periods of high work stress and insufficient sleep. Pain is usually located in both temporal regions, pulsatile, of moderate intensity, and accompanied by nausea, vomiting, photophobia, and phonophobia. Over the past 2 days, the headache has worsened, with increased pain intensity. Daily activities are restricted, and self-administered analgesics provide poor relief, prompting presentation to the hospital. Past medical history: The patient denies a history of hypertension, diabetes, or other chronic diseases. Auxiliary examination: None. Label (metadata; not part of input): Neurology. …… …… Open in a new tab Labeling principles (1) Recommended departments should be labeled according to the standard list of departments given. (2) All relevant departments should be labeled if the same complaint involves a combination of multiple diseases. (3) If the same complaint can be treated by more than one department, one department should be labeled in order of clinical priority. Based on the information in the sample data, the patient’s chief complaint or typical symptoms are converted into a description that includes the patient’s gender and age, and the data to be labeled are referenced as follows: Raw record example Outpatient medical record example: Patient information Sex: Male Age: 10 years Chief complaint: Burn injury to the right hand for several years, with hypertrophic scars on the middle and ring fingers. Outpatient diagnosis: Scar of the right hand. (metadata; excluded from model input) Management: Hospital admission for surgical treatment was recommended.(metadata; excluded from model input) Open in a new tab Sample labeling Patient information Recommended department Male,10 years, burn injury to the right hand for several years, with hypertrophic scars on the middle and ring fingers. Surgery Open in a new tab Example of dialog Dialogue content Recommended department Patient: “My left big toe feels hot and painful. Male, 27 years old.” Doctor: “Hello, did the pain start after trauma or a sprain?” Patient: “Yesterday my big toe Label (metadata; not part of input):Orthopedics. hit the leg of the bed.” Label (metadata; not part of input):Orthopedics. Open in a new tab Note The recommended sections for the above dialog are labeled with the results: Orthopaedic Patient information Recommended department Sex: Female Age: 35 years Chief complaint: Recurrent episodic headache for 3 years, worsened over the past 2 days. Current medical history: The patient first developed intermittent headaches 3 years ago, without obvious triggers. Headaches typically occur during periods of high work stress and insufficient sleep. Pain is usually located in both temporal regions, pulsatile, of moderate intensity, and accompanied by nausea, vomiting, photophobia, and phonophobia. Over the past 2 days, the headache has worsened, with increased pain intensity and accompanying blurred vision. Daily activities are restricted, and self-administered analgesics provide poor relief, prompting presentation to the hospital. Past medical history: The patient denies a history of hypertension, diabetes, or other chronic diseases. Auxiliary examination: None. Preliminary diagnosis: Migraine. Label (metadata; not part of input): Neurology. …… …… Note: The “Recommended Department” is to be marked. Open in a new tab Principles of organizing “patient information” (1) Anonymize personally identifiable information. (2) Ensure data quality by removing duplicates, checking for medical or common sense errors, and correcting them. (3) List the sections of gender, age, chief complaint, current medical history, past history, and auxiliary examination; diagnosis/assessment fields (including preliminary diagnosis) are retained only as metadata and excluded from model input. Principles of labeling “recommended department” (1) Recommended departments should be labeled according to the standard list of departments given. (2) All relevant departments should be labeled if the same complaint involves a combination of diseases. (3) If the same complaint can be treated by more than one department, one department should be labeled in order of clinical priority. Data quality criteria Data Completeness • Individual data: Ensure that the data is complete and there are no missing values or blank fields; ensure that each piece of data contains the necessary information, such as patient information, symptoms, and labeled recommended departments, etc. • Dataset: whether the dataset covers a sufficient range of departments; whether it covers common and typical symptoms for each department Data Accuracy • Whether the data content is correct e.g. whether there are errors in medical description, common sense errors • Whether the recommended departments are correct and comprehensive Data Uniqueness • Whether there is duplicate data or information. • Are the recommended sections correct and comprehensive? Language Security • Whether the language is discriminatory or not humane. • Whether there are sensitive words, etc. • Whether the data reveals personal privacy. • Whether the data contains unnecessary information. Data format standardization Whether the data format meets the requirements of the specification. Open in a new tab Data security Data collectors are required to comply with relevant laws and regulations to ensure compliance and legality of data collection, use and storage. The data will be used in a way that respects patients’ individual privacy rights and data protection regulations, and protects the privacy and security of the data. Restrictions on the use of data The data collected will only be used for the establishment of the medical guidance dataset, and the data shall not be used for other commercial purposes without authorization; the data shall not be made available for use by a third party without the authorized consent of the data provider. The ownership of all collected data belongs to the data collector, and the data provider reserves the right to interpret the data. Model inference and output stability To ensure reproducibility and minimize variance in model predictions, we adopted deterministic inference protocols for all large language models (LLMs). For open-source models, we fixed decoding parameters, including random seed, temperature, and sampling strategy (Extended Data Table 1 ), and disabled beam search. These settings ensured stable and repeatable outputs across runs. For proprietary models (e.g., GPT-4, Qwen-Turbo), which allow limited control over internal decoding parameters, we employed two complementary strategies. Structured prompting Prompts were standardized with fixed output formats and multiple in-context examples. Model responses were constrained to a predefined candidate set of departments, and any output falling outside this set was considered invalid and excluded from analysis. Self-consistency aggregation To mitigate position bias and residual stochasticity, we performed multiple perturbed runs per input—altering candidate order and random seed—and aggregated results. This approach draws on the Self-Consistency Decoding framework. Empirically, model predictions exhibited high output stability. In 10 independent trials (5 randomly selected cases and 5 high-frequency cases), model outputs remained structurally consistent, with stable department selections. For instance, for the input I have felt dizzy for more than twenty years, and it seems to have worsened this month [female, 66 years old]’ ”, all trials yielded overlapping predictions including “ Department of Internal Medicine ”. Data sources and portfolio The evaluation dataset was obtained from the evaluation dataset collected by our team. The dataset is divided into a validation set and a test set with 2000 and 2000 data respectively (Supplementary Table 1 ) to ensure the generalization ability of the model to different data distributions. Among them, one-fifth of the data have answers and the rest have no answers, and the teams can validate the model effect of the data in the no-answer part in the MedBench https://medbench.opencompass.org.cn/ platform. Submission and deployment requirements Before the competition, validation sets will be distributed and teams can use the validation sets to test the validity of the larger model, where one-fifth of the data has an answer and the rest of the data does not have an answer, and teams can validate the validity of the model for the portion of the data that does not have an answer on the Submit Results page, and the results of the validation sets will not be used as a measure of the final score. The test set will not be sent to the teams during the competition phase, and we will evaluate the teams’ models based on the test set. Teams are required to deploy their model to a Docker image, provide a repository address that can be accessed via Docker pull, or package and compress the image and provide a downloadable link. Once running, the image should provide offline calls to the model interface. Teams may manage their own uploaded mirrors. Teams are allowed to submit up to 2 valid Docker images to the organizers during the entire finals period. The algorithm score uses Accuracy as the evaluation metric. Accuracy is defined as Accuracy = TP + TN TP + TN + FP + FN 1 Accuracy is defined as the proportion of cases for which the predicted department set exactly matches the gold-standard set. In this triage task, each question (case) may correspond to one or multiple correct departments. A prediction is considered correct only if the model outputs the complete set of all correct departments. Thus, for questions with multiple valid departments, the model must correctly identify every department in the gold standard list for that question to be counted as accurate. For example: Gold = {Internal Medicine, Ophthalmology}; Prediction = {Internal Medicine} → Incorrect Gold = {Internal Medicine, Ophthalmology}; Prediction = {Internal Medicine, Ophthalmology} → Correct This exact-match criterion ensures that the evaluation faithfully reflects real-world clinical triage requirements, where incomplete department recommendations could mislead patient routing. In addition to Accuracy, we also report Precision and Recall, together with their 95% confidence intervals (CIs) computed using the normal approximation method. Precision and Recall are defined as Precision = TP TP + FP 2 Recall = TP TP + FP 3 The 95% confidence interval for a proportion p is calculated via the normal approximation method as CI = p ± Z p ( 1 − p ) n denominator 4 Post-competition modeling In our post-competition modeling, we employed a specific few-shot sampling strategy to improve performance. First, we constructed an example library of ~1000 manually annotated reference questions. For each original question (or test query), we used the BGE-m3 model to calculate sentence vector similarities between it and all questions in the example library. Based on these similarities, we carefully selected few-shot examples for each original question:(1) Positive Samples—the 10 examples most similar to the query and their corresponding standard answers; (2) Random Samples—10 additional examples randomly selected (with a fixed seed of 42 for reproducibility) from the remainder of the library. These 20 examples were integrated into the prompt to provide contextual cues for the model(see Supplementary Table 2 ). To prevent potential data leakage, we performed strict filtering and inspection. Vector similarity checks ensured no selected example had a similarity score greater than 0.99 with the test query, which typically indicates near-duplicate content. In addition, manual review of ~45% of test cases confirmed no instance of complete overlap between the query and the selected few-shot examples. To further improve robustness, we applied a self-consistency strategy to the candidate department list. Specifically, we randomly shuffled the department order nine times, resulting in ten versions of the input for each test query. The model generated predictions for each version independently. Final predictions were aggregated through frequency voting, retaining only those departments that appeared in at least six out of ten outputs. This threshold was optimized through validation experiments to ensure maximal accuracy under a strict correctness criterion. MedGPT-Guide also followed a structured chain-of-thought (CoT) reasoning process: it first interpreted user-specific context, then analyzed symptom and history information step by step, and finally mapped this reasoning process to a department prediction. This multi-stage reasoning improved the interpretability and faithfulness of model outputs, making the decision process more aligned with clinical thinking in real-world triage scenarios. Given the inherently multi-label nature of department prediction, we further benchmarked encoder-only baselines such as BGE-m3 and ERNIE-B2. These models demonstrated strong performance, suggesting that non-generative approaches remain highly effective for fine-grained medical triage, especially when interpretability is less critical. Competition overview The Large-Model-Based Medical Triage Evaluation Competition was held from May 2024 to October 2024, attracting 37 participating teams from medical institutions, research laboratories, and industry partners across China. Each team was required to deploy its model as a Docker image and submit up to two valid versions during the final phase. The evaluation framework employed balanced validation and test sets (2,000 instances each). Supplementary information Supplementary information (166KB, pdf) Acknowledgements Supported by Shanghai Artificial Intelligence Laboratory. Author contributions C.D., M.X.B., M.J.Y, K.Y.L., C.P.C, Y.Y.J., and J.X.conceived the study. C.D., M.X.B., and J.X. designed the study, collected data, and conducted data analyses. M.X.B., M.J.Y., and K.Y.L. drafted the manuscript. K.Y.L., C.P.C, Y.Y.J., and J.X. supervised the study. All authors–C.D., M.X.B., M.J.Y, K.Y.L., C.P.C, Y.Y.J., and J.X.- have read and approved the manuscript. Data availability The evaluation scripts and baseline configurations used in this study are integrated within the MedBench platform. To comply with hospital data security policies, detailed source code and Docker deployment scripts will be released upon final acceptance at https://github.com/OpenMedZoo . All prompts, metrics, and evaluation procedures are documented in the Methods section to ensure reproducibility. Code availability The evaluation scripts and baseline configurations used in this study are integrated within the MedBench platform. To comply with hospital data security policies, detailed source code and Docker deployment scripts will be released upon final acceptance at https://github.com/OpenMedZoo . All prompts, metrics, and evaluation procedures are documented in the Methods section to ensure reproducibility. Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. These authors contributed equally: Chao Ding, Mouxiao Bian, Minjia Yuan. Contributor Information Yuanye Jiang, Email: [email protected]. Jie Xu, Email: [email protected]. Supplementary information The online version contains supplementary material available at 10.1038/s41746-026-02433-8. References 1. Hsia, R. Y., Thind, A., Zakariah, A., Hicks, E. R. & Mock, C. Prehospital and emergency care: updates from the disease control priorities, version 3. World J. Surg. 39 , 2161–2167 (2015). [ DOI ] [ PubMed ] [ Google Scholar ] 2. Kollberg, B., Dahlgaard, J. & Brehmer, P.-O. Measuring lean initiatives in health care services: Issues and findings. J. Prod. Perform. Manag 56 , 7–24 (2006). [ Google Scholar ] 3. Hamilton, W. Diagnosis: cancer diagnosis in UK primary care. Nat. Rev. Clin. Oncol. 9 , 251–252 (2012). [ DOI ] [ PubMed ] [ Google Scholar ] 4. Iserson, K. V. & Moskop, J. C. Triage in medicine, part I: concept, history, and types. Ann. Emerg. Med. 49 , 275–281 (2007). [ DOI ] [ PubMed ] [ Google Scholar ] 5. Jia, H. et al. Analysis of factors affecting medical personnel seeking employment at primary health care institutions: developing human resources for primary health care. Int. J. Equity Health 21 , 37 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Zachariasse, J. M. et al. Validity of the Manchester Triage System in emergency care: a prospective observational study. PloS one 12 , e0170811 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Friedman, A. B., Delgado, M. K. & Weissman, G. E. Artificial intelligence for emergency care triage-much promise, but still much to learn. JAMA Netw. open 7 , e248857 (2024). [ DOI ] [ PubMed ] [ Google Scholar ] 8. Bierer, B. E. & Truog, R. D. The unresolved challenge of triage. JAMA Netw. open 6 , e2329676 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 9. Knitza, J. et al. Comparison of two symptom checkers (ada and symptoma) in the emergency department: randomized, crossover, head-to-head, double-blinded study. J. Med. Internet Res. 26 , e56514 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Dawoodbhoy, F. M. et al. AI in patient flow: applications of artificial intelligence to improve patient flow in NHS acute mental health inpatient units. Heliyon 7 , e06993 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Siira, E., Tyskbo, D. & Nygren, J. Healthcare leaders’ experiences of implementing artificial intelligence for medical history-taking and triage in Swedish primary care: an interview study. BMC Prim. Care 25 , 268 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Fraser, H. et al. Comparison of diagnostic and triage accuracy of Ada Health and WebMD symptom checkers, ChatGPT, and physicians for patients in an emergency department: clinical data analysis study. JMIR mHealth uHealth 11 , e49995 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Mika, A. P., Martin, J. R., Engstrom, S. M., Polkowski, G. G. & Wilson, J. M. Assessing ChatGPT responses to common patient questions regarding total hip arthroplasty. J. bone Jt. Surg. Am. Vol. 105 , 1519–1526 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 14. Chen, M. & Decary, M. Artificial intelligence in healthcare: an essential guide for health leaders. Healthc. Manag. forum 33 , 10–18 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 15. Masanneck, L. et al. Triage performance across large language models, ChatGPT, and untrained doctors in emergency medicine: comparative study. J. Med. Internet Res. 26 , e53297 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Levine, D. M. et al. The diagnostic and triage accuracy of the GPT-3 artificial intelligence model: an observational study. Lancet Digital Health 6 , e555–e561 (2024). [ DOI ] [ PubMed ] [ Google Scholar ] 17. Griot, M., Hemptinne, C., Vanderdonckt, J. & Yuksel, D. Large Language Models lack essential metacognition for reliable medical reasoning. Nat. Commun. 16 , 642 (2025). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Chen, S. et al. Meddialog: a large-scale medical dialogue dataset. 3 (2020). 19. Chen, Y. et al. Bianque: Balancing the questioning and suggestion ability of health llms with multi-turn health conversations polished by ChatGPT. (2023). 20. Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. 35 , 24824-24837 (2022). 21. Tang, L. et al. Evaluating large language models on medical evidence summarization. NPJ Digital Med. 6 , 158 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Cao, R. et al. Development and interpretation of a pathomics-based model for the prediction of microsatellite instability in colorectal cancer. Theranostics 10 , 11080–11091 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Schmieding, M. L., Mörgeli, R., Schmieding, M. A. L., Feufel, M. A. & Balzer, F. Correction: benchmarking triage capability of symptom checkers against that of medical laypersons: survey study. J. Med. Internet Res. 23 , e30215 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Fernandes, M. et al. Clinical decision support systems for triage in the emergency department using intelligent systems: a review. Artif. Intell. Med. 102 , 101762 (2020). [ DOI ] [ PubMed ] [ Google Scholar ] 25. Nguyen, H., Meczner, A., Burslam-Dawe, K. & Hayhoe, B. Triage errors in primary and pre-primary care. J. Med. Internet Res. 24 , e37209 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Sharma, M. et al. Artificial intelligence applications in health care practice: scoping review. J. Med. Internet Res . 24 , e40238 (2022). [ DOI ] [ PMC free article ] [ PubMed ] 27. Townsend, B. A., Plant, K. L., Hodge, V. J., Ashaolu, O. & Calinescu, R. Medical practitioner perspectives on AI in emergency triage. Front. Digital health 5 , 1297073 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Ahun, E. et al. Perceptions and concerns of emergency medicine practitioners about artificial intelligence in emergency triage management during the pandemic: a national survey-based study. Front. Public Health 11 , 1285390 (2023). [ DOI ] [ PMC free article ] [ PubMed ] 29. Burnell, R. et al. Rethink reporting of evaluation results in AI. Science 380 , 136–138 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 30. Wang, A., Hertzmann, A. & Russakovsky, O. Benchmark suites instead of leaderboards for evaluating AI fairness. Patterns 5 , 101080 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Algorithms of Oppression: How Search Engines Reinforce RacismAlgorithms of Oppression: How Search Engines Reinforce Racism Safiya Umoja Noble NYU Press, 2018. 256 Science . 374 , 542 (2021). [ DOI ] [ PubMed ] 32. Saeed, S. A. & Masters, R. M. Disparities in Health Care and the Digital Divide. Curr. Psychiatry Rep. 23 , 61 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366 , 447-453 (2019). [ DOI ] [ PubMed ] 34. Yadav, N. et al. Data privacy in healthcare: in the era of artificial intelligence. Indian Dermatol. online J. 14 , 788–792 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Miles, G. & Quinlan, A. HIPAA and video recordings in the clinical setting. Nursing 53 , 15–19 (2023). [ DOI ] [ PubMed ] [ Google Scholar ] 36. Orel, A. & Bernik, I. GDPR and health personal data; tricks and traps of compliance. Stud. Health Technol. Inform. 255 , 155–159 (2018). [ PubMed ] [ Google Scholar ] 37. Yao, Y. & Yang, F. Overcoming personal information protection challenges involving real-world data to support public health efforts in China. Front. public health 11 , 1265050 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Al Kuwaiti, A. et al. A review of the role of artificial intelligence in healthcare. J. Personal. Med. 13 , 951 (2023). [ DOI ] [ PMC free article ] [ PubMed ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary information (166KB, pdf) Data Availability Statement The evaluation scripts and baseline configurations used in this study are integrated within the MedBench platform. To comply with hospital data security policies, detailed source code and Docker deployment scripts will be released upon final acceptance at https://github.com/OpenMedZoo . All prompts, metrics, and evaluation procedures are documented in the Methods section to ensure reproducibility. The evaluation scripts and baseline configurations used in this study are integrated within the MedBench platform. To comply with hospital data security policies, detailed source code and Docker deployment scripts will be released upon final acceptance at https://github.com/OpenMedZoo . All prompts, metrics, and evaluation procedures are documented in the Methods section to ensure reproducibility. Articles from NPJ Digital Medicine are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (2.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top