Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

St. Gallen International Breast Cancer Consensus-Based Clinical Decision Validation: Concordance Assessment Between Deep Large Language Model Outputs and Global Expert Panel Recommendations.

Pan Y et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
computer-science-education
computer science education

St. Gallen International Breast Cancer Consensus-Based Clinical Decision Validation: Concordance Assessment Between Deep Large Language Model Outputs and Global Expert Panel Recommendations - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Ann Surg Oncol . 2026 Feb 10;33(5):4518–4529. doi: 10.1245/s10434-026-19176-1 Search in PMC Search in PubMed View in NLM Catalog Add to search St. Gallen International Breast Cancer Consensus-Based Clinical Decision Validation: Concordance Assessment Between Deep Large Language Model Outputs and Global Expert Panel Recommendations Yi Pan Yi Pan , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Yi Pan 1, # , Chenglong Duan Chenglong Duan , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Chenglong Duan 1, # , Jinsui Du Jinsui Du , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Jinsui Du 1 , Jianing Zhang Jianing Zhang , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Jianing Zhang 1 , Keyuan Du Keyuan Du , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Keyuan Du 1 , Chenrong Zhang Chenrong Zhang , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Chenrong Zhang 1 , Zhihao Liu Zhihao Liu , MS 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Zhihao Liu 1 , Wei Zhang Wei Zhang , MD, PhD 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Wei Zhang 1 , Bin Wang Bin Wang , MD , PhD 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Bin Wang 1 , Yu Ren Yu Ren , MD, PhD 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Yu Ren 1 , Zhao Sun Zhao Sun , MD , PhD 2 School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou, Henan Province China 3 Systems Engineering Institute, Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Zhao Sun 2, 3, ✉ , Lizhe Zhu Lizhe Zhu , MD , PhD 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China Find articles by Lizhe Zhu 1, ✉ Author information Article notes Copyright and License information 1 Department of Breast Surgery, The First Affiliated Hospital of Xi’an Jiaotong University, Xi’an, Shaanxi Province China 2 School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou, Henan Province China 3 Systems Engineering Institute, Xi’an Jiaotong University, Xi’an, Shaanxi Province China ✉ Corresponding author. # Contributed equally. Received 2025 Sep 24; Accepted 2026 Jan 17; Issue date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ . PMC Copyright notice PMCID: PMC13083474  PMID: 41667891 Abstract Background The newly developed large language model (LLM) DeepSeek has shown potential for application in other medical fields. However, few systematic studies have assessed its concordance with international expert consensus or compared its performance with leading models such as Gemini 2.0 Pro and ChatGPT-4o in breast cancer. Materials and Methods A total of 139 consensus questions from the 19th St. Gallen International Breast Cancer Conference (SG-BCC) were included into analysis. Each model was trained to answer each consensus question five times. The DeepSeek model was compared with the expert panel consensus in terms of concordance rate, robustness of the answers, Pearson correlation coefficient r for non-binary questions, and absolute proportion difference for binary questions. At the same time, a horizontal comparison was made with the previous LLMs Gemini 2.0 Pro and ChatGPT-4o. Results The overall concordance rate between DeepSeek-V3 and the expert panel consensus was 63.31%, and the average answer robustness (i.e., its self-consistency across repeated queries) of DeepSeek-V3 was 86.69%. In addition, DeepSeek-V3 performed similarly to Gemini 2.0 Pro and ChatGPT-4o in terms of concordance rate of the most frequent answers ( p = 0.849). In terms of model robustness, there were significant statistical differences among the models ( p < 0.001), with DeepSeek-V3 significantly outperforming Gemini 2.0 Pro ( p = 0.005) and ChatGPT-4o ( p < 0.001). Conclusions DeepSeek models showed moderate concordance in following the consensus of breast cancer expert panel and showed significant advantages in answer robustness, suggesting that DeepSeek has great application potential in the field of clinical decision-making for breast cancer. Supplementary Information The online version contains supplementary material available at 10.1245/s10434-026-19176-1. Keywords: DeepSeek, St. Gallen international breast cancer conference, Breast cancer, Large language models, Clinical decision-making Breast cancer is one of the most common malignancies in women worldwide. 1 – 3 This heterogeneous disease contains multiple subtypes and has complex and diverse treatment methods, posing a major threat to women’s health. 4 – 6 The accelerated development of many effective new therapies has significantly improved the survival rates of all breast cancer subtypes. 7 – 12 To optimize individualized treatment strategies, it is essential to continuously monitor and integrate the latest research results. Several international organizations, including the American Society of Clinical Oncology (ASCO), the European Society for Medical Oncology (ESMO), the National Comprehensive Cancer Network (NCCN), and the St. Gallen International Breast Cancer Conference (SG-BCC) expert panel, regularly publish evidence-based clinical practice guidelines. 13 – 16 These organizations rigorously evaluate and discuss important international clinical studies to develop consensus recommendations for breast cancer diagnosis and treatment, thereby providing expert guidance for personalized treatment methods. The biennial St. Gallen International Breast Cancer Conference convenes a global panel of breast cancer experts who conduct anonymous real-time voting on controversial or evidence-limited aspects of early breast cancer management. The final consensus reflects the collective opinions of top international experts, simulates the complex clinical decision-making processes, and provides valuable guidance for clinical practice. 16 However, it remains a challenge for surgical oncologists to apply evolving and detailed consensus recommendations to individual patients. The increasing volume and complexity of clinical data require considerable time for review. This is especially evident when clinicians prepare for multidisciplinary tumor board discussions or manage uncommon clinical presentations. At the same time, artificial intelligence technologies represented by large language models (LLMs) have made breakthrough progress in recent years, and have shown great application potential in the medical field. 17 – 22 They can rapidly integrate guideline-based treatment options for complex cases and support the preparation of multidisciplinary tumor board discussions. 23 In regions with limited medical resources, LLMs may provide standardized and accessible clinical reference information. They can also serve as educational tools for trainees and assist in developing patient-oriented instructional materials. 24 However, their clinical value depends on a fundamental requirement: they must be able to reproduce expert consensus in a consistent and accurate way. Therefore, careful evaluation of their performance against established clinical standards is essential before they can be considered for routine use. Among them, the newly developed and cost-effective large language model DeepSeek has also performed well in other medical fields. 25 – 28 However, few systematic studies have assessed its concordance with international expert consensus in breast cancer clinical decision-making. According to the developer’s documentation, the model is available in two versions: DeepSeek-V3 and DeepSeek-R1. Each version has a different performance focus, and their usefulness in breast cancer clinical decision-making remains unclear. Therefore, this study evaluates both versions and compares the better-performing one with leading models such as Gemini 2.0 Pro and ChatGPT-4o. Materials and Methods Data Source The question set for this study was derived from the consensus voting results presented at the 19th SG-BCC, which were obtained by one of the authors in this study who attended the conference. After excluding questions about previous clinical practice choices that could not be answered by LLMs, 139 consensus questions with clear answer options were finally included in the study. At the same time, an expert panel cohort was established on the basis of the expert panel voting results. Large Language Model Testing We selected representative LLMs released before 15 March 2024 to evaluate their clinical decision-making performance. The models included DeepSeek, ChatGPT-4o, and Gemini 2.0 Pro. DeepSeek provides two versions: DeepSeek-V3, a general-purpose model, and DeepSeek-R1, which is optimized for complex reasoning. These four models have distinct features within the fast-evolving field of medical artificial intelligence (AI). Detailed information on their technical specifications and previous medical evaluations is provided in Supplementary Text 1. All LLMs were tested through their official application programming interfaces (APIs) or publicly available online platforms. Before testing, we standardized the training process by requiring each model to use only knowledge available before 15 March 2025 (i.e., before the 19th SG-BCC). We provided each model with the NCCN Clinical Practice Guidelines: Breast Cancer (3rd edition, 2025) 15 and the Chinese Society of Clinical Oncology (CSCO) Breast Cancer Diagnosis and Treatment Guidelines 2024 29 for reference when answering questions. To minimize bias from question phrasing, we reformatted each consensus question into standardized instructions before entering it to the LLMs. The prompts were as follows: “Please answer the following questions based on the above two guidelines, using data before March 15, 2025, and clearly provide the letter corresponding to the option you selected.” During testing, each model was independently run five times per question to establish a cohort of LLMs. 30 The answer that appeared most frequently across these runs was defined as the “most frequent answer.” If the most frequent answer did not emerge in the first five tests, we conducted a sixth test. Each test started a new round of dialogue to eliminate memory interference caused by the previous question. Metrics Concordance rate: The proportion of questions where the model’s most frequent answer agrees with the option with the highest proportion of expert panel votes. Robustness: The proportion of the model’s most frequent answer that was selected in five (or six) conversations for a specific question. This metric can provide important supplementary information. It reflects the reliability and stability of the model’s outputs, which is essential for any tool intended for clinical use. A high robustness score indicates that the model produces consistent answers and that its outputs are not driven by randomness. It also suggests stronger internal confidence in the conclusions generated. In addition, examining robustness across different clinical topics helps identify areas where the model is stable or uncertain. Such findings are useful for guiding the development of safer and more effective AI systems for clinical decision support. Absolute proportional difference: Specifically used for binary problems. This metric uses the option with the highest proportion of expert votes as the reference and calculates the absolute difference between the proportion of the option selected by the model and the proportion of experts voting for that option. The smaller the value, the closer the proportion predicted by the model is to the expert consensus. Pearson correlation coefficient r : Specifically used for nonbinary problems. It measures the linear correlation between the distribution vector of the model’s answers and the distribution vector of the proportion of expert votes, quantifying the strength of the association between them. Statistical Analysis In this study, since most measurement data (including robustness, absolute proportional difference, and Pearson correlation coefficient r ) were non-normally distributed, they were expressed as medians and interquartile ranges (IQRs), and nonparametric tests were employed. Only in Table 2 , to present detailed differences, were some indicators reported as means without accompanying statistical inference. Categorical data were expressed as frequencies ( n ) and percentages (%). The Pearson chi-squared test was used to compare concordance rates between models. For nonparametric analyses, the Kruskal–Wallis test was applied. When Kruskal–Wallis results were statistically significant ( p < 0.05), post hoc comparisons were performed using the Dwass–Steel–Critchlow–Fligner (DSCF) method. All statistical tests were two-sided, with a significance level of α = 0.05. Analyses were conducted using SPSS PASW Statistics 18, R software (version 4.2.2), and MSTATA software ( https://www.mstata.com/ ). Table 2. Detailed comparison of DeepSeek-V3 with the 19th St. Gallen International Breast Cancer Conference Expert Panel DeepSeek-V3 Expert panel Overall performance (N = 139) Common responses 88 (63.31%) Average robustness/majority in all questions ( N = 139) 86.69% 62.01% Average robustness/majority in common answers ( N = 88) 90.23% 65.62% Questions answered with average robustness/majority of ≥ 80% 109 (78.42%) 25 (17.99%) Performance by question type Binary questions (N = 37) Common responses 25 (67.57%) Absolute proportional difference 1 0.30 (0.13, 0.50) Nonbinary questions (N = 102) Common responses 63 (61.76%) Pearson correlation coefficient r 1 0.82 (0.20, 0.98) Comparison of robustness among LLMs by topic classification Topic “Systemic therapy: ER positive, HER2 negative breast cancers” ( N = 44) Common responses 22 (50.00%) Average robustness/majority in this topic 84.09% 58.34% Topic “Genetic testing” (N = 17) Common responses 13 (76.47%) Average robustness/majority in this topic 87.06% 67.71% Topic “Radiation therapy” (N = 28) Common responses 15 (53.57%) Average robustness/majority in this topic 85.00% 64.04% Open in a new tab 1 Median (M) and interquartile range (P25, P75) of the above indicators In this table, most of the measurement data are not normally distributed, but robustness is reported in the form of mean to reflect the specific differences without statistical inference. Absolute proportional difference and Pearson correlation coefficient r are still expressed in median and interquartile range LLMs large language models Ethical Approval This study does not involve any ethnic human participant data. According to national laws and regulations and institutional requirements, this study does not need to be reviewed and approved by the ethics committee. Results Characteristics of the Research Question Set A total of 139 questions were included in this study, covering various areas of breast cancer diagnosis and treatment. According to the number of options for each question, the questions were divided into binary questions ( n = 37, 26.62%) and nonbinary questions ( n = 102, 73.38%). In addition, from a clinical perspective, the questions were divided into nine main topics (Table 1 ). Table 1. Distribution of research question characteristics Category Subcategory Number of cases ( N ) Percentage (%) Total 139 100.00 Problem Type Binary questions 37 26.62 Nonbinary questions 102 73.38 Topic Ductal carcinoma in situ 10 7.19 Genetic testing 17 12.23 Breast and axillary surgery 16 11.51 Radiation therapy 28 20.14 Systemic therapy: ER positive, HER2 negative breast cancers 44 31.65 Systemic therapy: HER2 positive breast cancer 10 7.19 Oligometastatic breast cancer 7 5.04 Treatment of local-regional recurrence of breast cancer 2 1.44 Survivorship 5 3.60 Open in a new tab Binary questions are questions with two options and nonbinary questions are those with three or more options Comparison of DeepSeek Model with Expert Panel The first stage of the analysis compared DeepSeek-V3 and DeepSeek-R1 to determine which model was more suitable for further evaluation. Across the 139 SG-BCC questions, DeepSeek-V3 showed a higher overall concordance rate with the expert consensus and greater robustness than DeepSeek-R1. It also performed better in key stratified analyses. On the basis of these findings, DeepSeek-V3 was selected as the primary model for all subsequent analyses in the main text. Detailed results for DeepSeek-R1 are provided in Supplementary Table 1. Among the 139 questions, DeepSeek-V3’s answers were consistent with the options with the highest proportion of expert panel votes in 88 questions (63.31%). For these 88 questions, the model showed an average robustness of 90.23%, which means that in five (or six) rounds of answering, DeepSeek-V3 selected the most frequent expert vote in 90.23% of cases. Across all 139 questions, the answers provided by DeepSeek-V3 had an average robustness of 86.69%. Out of a total of 139 questions, DeepSeek-V3 answered 109 questions (78.42%) with a robustness of at least 80%. Among the 88 consistent questions, expert panel voted with an average majority of 65.62%. Among all the questions, 25 questions had a voting rate of more than 80% (Table 2 ). In addition, this study conducted a stratified analysis of different question types and specific clinical hot topics, revealing subtle differences in the performance of DeepSeek-V3. Among the 37 binary questions, 25 answers of DeepSeek-V3 were consistent with the options with the highest proportion of expert panel votes, accounting for about 67.57%. The median absolute proportion difference of the answers was 0.30. Among the 102 nonbinary questions, 63 answers of DeepSeek-V3 (accounting for 61.76%) were consistent with the options with the highest proportion of expert panel votes, with a median Pearson correlation coefficient r of 0.82. In addition, in the three clinical hot topics with the largest number of questions, “Systemic therapy: ER positive, HER2 negative breast cancers,” “Genetic testing,” and “Radiation therapy,” the concordance rates of DeepSeek-V3 answers were 50.00%, 76.47%, and 53.57%, respectively (Table 2 ). The results suggest that the DeepSeek-V3 model exhibits moderate concordance in following the opinions of the SG-BCC expert panel and has high robustness in its responses. Horizontal Comparison of Concordance and Robustness of LLMs To further evaluate the potential performance differences between DeepSeek model and leading models Gemini 2.0 Pro and ChatGPT-4o in following expert consensus, this study conducted a comprehensive comparative analysis from multiple dimensions including consistency and robustness. The concordance analysis showed that there was no significant difference in the overall concordance rate among the three tested models ( p = 0.849), and the concordance rate of all models remained between 60% and 64%. Subgroup analysis of binary and nonbinary questions showed no significant difference in concordance rates, absolute proportional differences, or Pearson correlation coefficient r among the models (all p > 0.05). These results indicate that DeepSeek has comparable concordance with other evaluated models (Table 3 ). Table 3. Horizontal comparison of consistency among large language models DeepSeek-V3 ChatGPT-4o Gemini 2.0 Pro p -Value Overall concordance rate (N = 139) 88 (63.31%) 84 (60.43%) 88 (63.31%) 0.849 1 Binary questions (N = 37) Common responses 25 (67.56%) 29 (78.38%) 23 (62.16%) 0.305 1 Absolute proportional difference 3 0.30 (0.13, 0.50) 0.24 (0.12, 0.34) 0.22 (0.13, 0.37) 0.397 2 Nonbinary questions (N = 102) Common responses 63 (61.76%) 55 (53.92%) 65 (63.73%) 0.319 1 Pearson correlation coefficient r 3 0.82 (0.20, 0.98) 0.65 (−0.05, 0.94) 0.72 (0.25, 0.96) 0.117 2 Open in a new tab 1 Pearson’s chi-squared test 2 Kruskal–Wallis test 3 Median (M), interquartile range (P25, P75) The robustness analysis showed that all tested LLMs exhibited median robustness values ranging from 0.80 to 1.00, while the median voting rate of the expert panel on the top-voted answers was 0.60. The DeepSeek-V3 achieved a median robustness of 1.00 (IQR 0.80, 1.00), which is significantly better than Gemini 2.0 Pro ( p = 0.005) and ChatGPT-4o ( p < 0.001). There is no significant difference between Gemini 2.0 Pro and ChatGPT-4o ( p = 0.557). In addition, subgroup analysis of multiclassification problems also revealed a similar trend, and DeepSeek models maintained their robustness advantage (Table 4 , Fig. 1 ). Table 4. Horizontal comparison of robustness among large language models DeepSeek-V3 1 ChatGPT-4o 1 Gemini 2.0 Pro 1 p -Value 2 Overall performance (N = 139) 3 1.00 (0.80, 1.00) 0.80 (0.60, 1.00) 0.80 (0.60, 1.00) < 0.001* Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – < 0.001* 0.005* Versus ChatGPT-4o – – 0.557 Versus Gemini 2.0 Pro – – – Binary questions (N = 37) 3 1.00 (0.80, 1.00) 0.80 (0.60, 1.00) 0.80 (0.60, 1.00) 0.083 Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.133 0.114 Versus ChatGPT-4o – – 0.993 Versus Gemini 2.0 Pro – – – Nonbinary questions (N = 102) 3 1.00 (0.60, 1.00) 0.80 (0.60, 1.00) 0.80 (0.60, 1.00) < 0.001* Post hoc pairwise comparisons p -values 4 Versus DeepSeek-V3 – < 0.001* 0.036* Versus ChatGPT-4o – – 0.431 Versus Gemini 2.0 Pro – – – Open in a new tab 1 Represents the robustness of the large language models, defined as the percentage of times the model selected its most frequent answer in five (or six) runs 2 Kruskal–Wallis test 3 Median (M) and interquartile range (P25, P75) of robustness 4 Dwass–Steel–Critchlow–Fligner test * Statistical significance was indicated as p < 0.05 Fig. 1. Open in a new tab Comparison of robustness among large language models; a overall robustness across all consensus questions, b robustness for binary questions, c robustness for non-binary questions; x -axis represents the large language models tested, including DeepSeek-V3, ChatGPT-4o, and Gemini 2.0 Pro, and y -axis represents the robustness of the large language models, defined as the percentage of times the model selected its most frequent answer in 5 (or 6) runs; statistical comparisons were performed by the Kruskal–Wallis test, and post hoc pairwise comparisons were assessed by the Dwass–Steel–Critchlow–Fligner method; * p < 0.05; ** p < 0.01; *** p < 0.001; **** p < 0.0001; ns not significant Horizontal Comparison by Topic Classification Furthermore, we performed a stratified analysis of the concordance rate and robustness of each model in nine clinical topics. The concordance analysis showed that there was no significant difference in the consistency rate between models in different topics (all p > 0.05) (Fig. 2 ). Fig. 2. Open in a new tab Comparison of large language models performance by clinical topic classification; a comparison of concordance among Large Language Models by clinical topic classification, b comparison of robustness among Large Language Models by clinical topic classification; x -axis represents the p -value of the comparison between models after the Pearson’s chi-squared test (in panel A) and after the Kruskal–Wallis test (in panel B) and y -axis represents the different clinical topics; each cell in the heatmap reflects either the concordance rate (panel A) or robustness score (panel B) for different clinical topics; color intensity corresponds to the magnitude of the value; statistical comparisons were performed by Pearson’s chi-squared test (in panel A) and Kruskal–Wallis test (in panel B); the number of questions on the subject “Treatment of locally recurrent breast cancer” is too small, thus no statistical analysis is required; x -axis has been expanded in the 0–0.2 range to emphasize differences in p -values in panel B; an axis break is indicated However, we observed heterogeneity in model robustness. In the topic of “Ductal carcinoma in situ,” DeepSeek-V3 showed superior robustness with a median of 1.00 (IQR 0.95, 1.00), significantly better than ChatGPT-4o ( p = 0.048). In the topics of “Systemic therapy: ER positive, HER2 negative breast cancers” and “Systemic therapy: HER2 positive breast cancer,” DeepSeek-V3 showed better robustness than ChatGPT-4o ( p = 0.016 and p = 0.003, respectively). The number of questions on the topic of “Treatment of local-regional recurrence of breast cancer” was too small, so only the median was shown and no statistical analysis was required. It is noteworthy that no single model performs best on all topics, suggesting that there may be differences in the performance of LLMs on different medical topics, which may be due to differences in the distribution of training data and the complexity of the problems 31 (Table 5 ). Table 5. Comparison of robustness among large language models by subject classification DeepSeek-V3 1 ChatGPT-4o 1 Gemini 2.0 Pro 1 p -Value 2 Topic “Ductal carcinoma in situ” (N = 10) 3 1.00 (0.95, 1.00) 0.80 (0.60, 1.00) 1.00 (1.00, 1.00) 0.012* Post hoc pairwise comparisons p -values 4 Versus DeepSeek-V3 – 0.048* 0.878 Versus ChatGPT-4o – – 0.042* Versus Gemini 2.0 Pro – – – Topic “Genetic testing” (N = 17) 3 1.00 (0.80, 1.00) 0.80 (0.80, 1.00) 1.00 (0.60, 1.00) 0.882 Post hoc pairwise comparisons p -values 4 Versus DeepSeek-V3 – 0.993 0.892 Versus ChatGPT-4o – – 0.918 Versus Gemini 2.0 Pro – – – Topic “Breast and axillary surgery” (N = 16) 3 0.90 (0.65, 1.00) 0.80 (0.60, 1.00) 0.80 (0.65, 0.95) 0.511 Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.566 0.603 Versus ChatGPT-4o – – 0.953 Versus Gemini 2.0 Pro – – – Topic “Radiation therapy” (N = 28) 3 0.90 (0.80, 1.00) 0.80 (0.60, 1.00) 0.80 (0.60, 1.00) 0.194 Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.526 0.183 Versus ChatGPT-4o – – 0.706 Versus Gemini 2.0 Pro – – – Topic “Systemic therapy: ER positive, HER2 negative breast cancers” (N = 44) 3 1.00 (0.60, 1.00) 0.60 (0.60, 1.00) 0.60 (0.60, 1.00) 0.013* Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.016* 0.074 Versus ChatGPT-4o – – 0.730 Versus Gemini 2.0 Pro – – – Topic “Systemic therapy: HER2 positive breast cancer” (N = 10) 3 1.00 (0.80, 1.00) 0.60 (0.60, 0.80) 0.90 (0.60, 1.00) 0.006* Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.003* 0.474 Versus ChatGPT-4o – – 0.142 Versus Gemini 2.0 Pro – – – Topic “Oligometastatic breast cancer” (N = 7) 3 1.00 (0.80,1.00) 0.80 (0.80,1.00) 1.00 (0.80,1.00) 0.319 Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – 0.417 > 0.999 Versus ChatGPT-4o – – 0.417 Versus Gemini 2.0 Pro – – – Topic “Treatment of local-regional recurrence of breast cancer” (N = 2) 5 0.60 0.60 0.90 – Topic “Survivorship” (N = 5) 3 1.00 (0.70,1.00) 1.00 (0.70,1.00) 0.50 (0.50,1.00) 0.857 Post hoc pairwise comparisons p-values 4 Versus DeepSeek-V3 – > 0.999 0.884 Versus ChatGPT-4o – – 0.884 Versus Gemini 2.0 Pro – – – Open in a new tab 1 Represents the robustness of the LLMs, defined as the percentage of times the model selected its most frequent answer in five (or six) runs 2 Kruskal–Wallis test 3 Median (M) and interquartile range (P25, P75) of robustness 4 Dwass–Steel–Critchlow–Fligner test 5 The number of questions related to the topic “Treatment of local-regional recurrence of breast cancer” was too small. Only the median is reported, and statistical analysis was not performed * Statistical significance was indicated as p < 0.05 Comparison of Concordance between LLMs and Expert Consensus across Evidence Levels To further assess how evidence strength affects concordance between LLMs and expert consensus, we performed a sub-analysis on the basis of the concentration of expert voting. We divided the 139 questions into two groups: evidence-based questions and experience-based questions. Evidence-based questions were those with ≥ 80% expert voting rate and were generally addressed clearly in the guidelines ( n = 25). Experience-based questions had < 80% expert voting rate and were not directly answered in the guidelines, requiring judgment on the basis of expert clinical experience ( n = 114). As shown in Fig. 3 , all three models achieved higher concordance rate on evidence-based questions than on experience-based questions. DeepSeek-V3 reached 84.00% concordance rate for evidence-based questions and 58.77% for experience-based questions ( p = 0.018). ChatGPT-4o showed a similar pattern (88.00% versus 54.39%, p = 0.002), as did Gemini 2.0 Pro (84.00% versus 58.77%, p = 0.018). These results suggest that LLMs perform well in areas supported by strong evidence but perform less concordance when clinical decisions depend mainly on expert experience. Fig. 3. Open in a new tab Comparison of concordance between LLMs and expert consensus across evidence levels; x -axis represents the large language models tested, including DeepSeek-V3, ChatGPT-4o, and Gemini 2.0 Pro and y -axis represents the concordance rate (%) of each LLM with the expert consensus; bars: each bar indicates the concordance rate for a specific model within a specific evidence category; dark blue bars represent the “Evidence-based” questions (i.e., questions with ≥ 80% expert voting rate), and light blue bars represent the “Experience-based” questions (i.e., questions with < 80% expert voting rate); numerical labels: the percentage value for each concordance rate is displayed directly within the corresponding bar for clarity; statistical analysis: statistical comparisons between the “Evidence-based” and “Experience-based” groups within each model were performed using the Pearson’s chi-squared test; the results of these comparisons are indicated by the significance markers (asterisks) above the bars; * p < 0.05, ** p < 0.01 Discussion This study systematically evaluated the emerging LLM DeepSeek model in terms of concordance with the expert panel consensus and the robustness of its response in clinical decision recommendations, using the discussion questions of the 19th SG-BCC meeting as a benchmark. We compared its performance with internationally representative models Gemini 2.0 Pro and ChatGPT-4o. The results showed that DeepSeek-V3 had a concordance rate of 63.31% with the experts panel consensus. This performance was not significantly different from Gemini 2.0 Pro and ChatGPT-4o, indicating that DeepSeek exhibited a comparable capacity to interpret and agree with complex medical consensus. In addition, the overall concordance rate of all tested models (60–64%) was slightly improved compared with previous similar studies, 30 which may be attributed to the developers’ continuous optimization work. However, the current concordance level is still far from the needs of clinical applications, and the models still face challenges in simulating expert decision-making on complex clinical problems. In terms of robustness, the median response robustness of all tested models reached or exceeded 80% (0.8). Among them, DeepSeek-V3 performed particularly well, with a significantly higher robustness than other models. This suggests that the model can maintain consistent output when facing repeated queries. Although high robustness is of great benefit to tasks that require repetition, such as batch generation of standardized documents, in the field of clinical decision support, it does not mean high accuracy, 32 and the model may stably give answers that do not conform to best clinical practice. 32 – 34 Therefore, the high robustness of DeepSeek only reflects the stability of its output and does not directly demonstrate its clinical decision-making abilities. 35 This study showed that the median voting rate of the expert panel on the top-voted answers was 60% (0.60), highlighting the inherent diversity of clinical opinions. This highlights the uncertainty and diversity of perspectives in real-world practice. It is worth noting that the internal robustness exhibited by the LLMs exceeds the level of expert panel consensus, indicating that there is a potential risk in directly equating its output results with established expert opinions. To better interpret this finding, it is necessary to distinguish consensus-based decision-making from algorithm-driven guidelines. SG-BCC reflects the heterogeneity of expert opinions. In areas where there are not currently many clinical trials, multiple clinical management approaches may be acceptable. NCCN guidelines are presented in flowchart form. It is not a simple vote, but a decision-making process based on the results of numerous clinical trials. According to the results in Fig. 3 , all models can answer evidence-based questions on the basis of algorithm-driven guidelines with a higher concordance rate than experience-based questions on the basis of consensus-based decision-making. The design mechanism of LLM may be the reason. LLMs are designed to utilize their own material library and user-uploaded data to find answers with the optimal probability of matching keywords in the user's questions. 36 , 37 When addressing evidence-based questions, clear guidelines allow LLMs to identify and reproduce a high-probability answer. In contrast, experience-based questions contain diverse and sometimes conflicting expert views, making it difficult for LLMs to determine a consistent response. This difference reflects a significant limitation of LLMs. While they can effectively summarize and integrate structured guideline information, they currently fail to reflect the nuances of clinical judgment. These differences often arise in areas with limited or conflicting evidence, and they help surgeons develop personalized treatment plans for specific patient groups. Therefore, at this stage, LLMs should be strictly limited to auxiliary roles, 38 – 40 such as retrieving disease knowledge, displaying guideline information, or generating preliminary plans. They should not completely replace clinicians’ professional judgment tailored to individual patient needs. 41 , 42 In addition, this study found topic differences in the application performance of the models. The performance of the models may be affected by the coverage of specific domain knowledge in their training corpus. This highlights the need to subdivide the domain and design scenario-specific tests when evaluating and deploying such models in the medical field. 31 This study has several limitations. First, the research questions were derived from a single international conference consensus, which inevitably led to limited coverage and restricted the testing scenarios of the models. Second, the evaluation method was confined to text-based question answering, while real-world clinical practice requires integration of multimodal data (such as imaging results, ultrasound results, and pathology results). This may affect the effectiveness of the model in real-world applications. Third, our research methodology provides the model with a limited set of reference guidelines, which aims to ensure the standardization of the testing environment and reduce errors caused by the model accessing different guidelines with varying priorities. 43 At the same time, surgeons and patients often lack the time and resources to download a large number of guidelines and input them into LLM, so our study simulates real clinical scenarios to some extent. However, future research should undoubtedly explore the impact of providing a wider range of clinical guidelines on model performance. Furthermore, although the study evaluated concordance and robustness, it did not directly evaluate the clinical plausibility or safety of the model output. 44 Future research in this field should monitor the progress of LLMs and establish a comprehensive and stable evaluation framework through continuous assessment. This system should rigorously review the clinical efficacy, security, and privacy safeguards of the model. 41 , 45 At the same time, prospective research should be conducted under a strict ethical and regulatory framework to objectively evaluate the application value and potential risks of such models in actual medical scenarios. 46 – 48 Conclusions The results of this study show that DeepSeek exhibited moderate concordance in following the 19th SG-BCC expert panel consensus, and its performance is comparable to leading international models such as Gemini 2.0 Pro and ChatGPT-4o. Furthermore, DeepSeek significantly outperformed other models in terms of answer robustness, highlighting its potential for application in clinical decision-making for breast cancer. In addition, given that all LLMs output results have moderate concordance but higher robustness than the expert panel, it is necessary to rigorously validate the clinical decision-making results output by LLMs. Supplementary Information Below is the link to the electronic supplementary material. Supplementary file1 (DOCX 19 KB) (19KB, docx) Supplementary file2 (DOCX 14 KB) (13.6KB, docx) Acknowledgement We sincerely thank all the experts who participated in the 19th St. Gallen International Breast Cancer Conference (SG-BCC) for their valuable contributions to reaching a global consensus on breast cancer treatment. We also sincerely thank the SG-BCC organizing committee for providing a high-level academic platform that promotes international cooperation and evidence-based clinical decision-making. In addition, we thank the publicly accessible large language model platforms used in this study for their support, including DeepSeek ( https://www.deepseek.com ), ChatGPT by OpenAI ( https://chat.openai.com ), and Gemini by Google ( https://gemini.google.com ). These platforms provide the basis for data generation and evaluation of the models in this study. Author Contributions Concept and design: L.Z. and Z.S. Drafting of the manuscript: Y.P. and C.D. Analysis and interpretation of data: Y.P., C.D., and J.Z. Statistical analysis: Y.P. and J.D. Acquisition of large language model data: K.D., C.Z., and Z.L. Critical revision of the manuscript for important intellectual content: W.Z. and B.W. Attendance at the 19th St. Gallen International Breast Cancer Conference (SG-BCC) and acquisition of expert voting data: Y.R. Supervision: L.Z. Funding This work was supported by National Natural Science Foundation of China (no. 82172798), Key Research and Development Program of Shaanxi (grant no.: 2023-YBSF-622), Key Research and Development Program of Shaanxi (grant no.: 2024SF-ZDCYL-02-08), and the National Natural Cultivation Youth Project of the First Affiliated Hospital of Xi’an Jiaotong University (no. 2024-QN-30). Disclosure The authors have no conflicts of interest to report. Ethical Approval This study did not involve any human participant data, so informed consent was not required. Footnotes Publisher's Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Yi Pan and Chenglong Duan have contributed equally to this work. Contributor Information Zhao Sun, Email: [email protected]. Lizhe Zhu, Email: [email protected]. References 1. Giaquinto AN, Sung H, Newman LA, et al. Breast cancer statistics 2024. CA Cancer J Clin . 2024;74(6):477–95. 10.3322/caac.21863. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Coles CE, Anderson BO, Cameron D, et al. The Lancet Breast Cancer Commission: tackling a global health, gender, and equity challenge. Lancet . 2022;399(10330):1101–3. 10.1016/S0140-6736(22)00184-2. [ DOI ] [ PubMed ] [ Google Scholar ] 3. Britt KL, Cuzick J, Phillips KA. Key steps for effective breast cancer prevention. Nat Rev Cancer . 2020;20(8):417–36. 10.1038/s41568-020-0266-x. [ DOI ] [ PubMed ] [ Google Scholar ] 4. Breast cancer: pathogenesis and treatments - PubMed. Accessed July 16, 2025. https://pubmed.ncbi.nlm.nih.gov/39966355/ 5. Derakhshan F, Reis-Filho JS. Pathogenesis of triple-negative breast cancer. Annu Rev Pathol . 2022;17:181–204. 10.1146/annurev-pathol-042420-093238. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Li YW, Dai LJ, Wu XR, et al. Molecular characterization and classification of HER2-positive breast cancer inform tailored therapeutic strategies. Cancer Res . 2024;84(21):3669–83. 10.1158/0008-5472.CAN-23-4066. [ DOI ] [ PubMed ] [ Google Scholar ] 7. Heater NK, Warrior S, Lu J. Current and future immunotherapy for breast cancer. J Hematol Oncol . 2024;17(1):131. 10.1186/s13045-024-01649-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Slamon DJ, Neven P, Chia S, et al. Overall survival with ribociclib plus fulvestrant in advanced breast cancer. N Engl J Med . 2020;382(6):514–24. 10.1056/NEJMoa1911149. [ DOI ] [ PubMed ] [ Google Scholar ] 9. Sledge GW, Toi M, Neven P, et al. The effect of abemaciclib plus fulvestrant on overall survival in hormone receptor-positive, ERBB2-negative breast cancer that progressed on endocrine therapy-MONARCH 2: a randomized clinical trial. JAMA Oncol . 2020;6(1):116–24. 10.1001/jamaoncol.2019.4782. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Im SA, Lu YS, Bardia A, et al. Overall survival with ribociclib plus endocrine therapy in breast cancer. N Engl J Med . 2019;381(4):307–16. 10.1056/NEJMoa1903765. [ DOI ] [ PubMed ] [ Google Scholar ] 11. Hortobagyi GN, Stemmer SM, Burris HA, et al. Overall survival with ribociclib plus letrozole in advanced breast cancer. N Engl J Med . 2022;386(10):942–50. 10.1056/NEJMoa2114663. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Geyer CE, Garber JE, Gelber RD, et al. Overall survival in the OlympiA phase III trial of adjuvant olaparib in patients with germline pathogenic variants in BRCA1/2 and high-risk, early breast cancer. Ann Oncol . 2022;33(12):1250–68. 10.1016/j.annonc.2022.09.159. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Bedrosian I, Somerfield MR, Achatz MI, et al. Germline testing in patients with breast cancer: ASCO-Society of Surgical Oncology Guideline. J Clin Oncol . 2024;42(5):584–604. 10.1200/JCO.23.02225. [ DOI ] [ PubMed ] [ Google Scholar ] 14. Loibl S, André F, Bachelot T, et al. Early breast cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol . 2024;35(2):159–82. 10.1016/j.annonc.2023.11.016. [ DOI ] [ PubMed ] [ Google Scholar ] 15. Gradishar WJ, Moran MS, Abraham J, et al. Breast cancer, 2024, NCCN clinical practice guidelines in oncology. J Natl Compr Canc Netw . 2024;22(5):331–57. 10.6004/jnccn.2024.0035. [ DOI ] [ PubMed ] [ Google Scholar ] 16. Curigliano G, Burstein HJ, Gnant M, et al. Understanding breast cancer complexity to improve patient outcomes: the St Gallen International Consensus Conference for the Primary Therapy of Individuals with Early Breast Cancer 2023. Ann Oncol . 2023;34(11):970–86. 10.1016/j.annonc.2023.08.017. [ DOI ] [ PubMed ] [ Google Scholar ] 17. Stanley J, Rabot E, Reddy S, Belilovsky E, Mottron L, Bzdok D. Large language models deconstruct the clinical intuition behind diagnosing autism. Cell . 2025;188(8):2235-2248.e10. 10.1016/j.cell.2025.02.025. [ DOI ] [ PubMed ] [ Google Scholar ] 18. Luo X, Rechardt A, Sun G, et al. Large language models surpass human experts in predicting neuroscience results. Nat Hum Behav . 2025;9(2):305–15. 10.1038/s41562-024-02046-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Bluethgen C, Van Veen D, Zakka C, et al. Best practices for large language models in radiology. Radiology . 2025;315(1):e240528. 10.1148/radiol.240528. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Lane TR, Vignaux PA, Harris JS, Snyder SH, Urbina F, Ekins S. Machine learning and large language models for modeling complex toxicity pathways and predicting steroidogenesis. Environ Sci Technol . 2025;59(27):13844–56. 10.1021/acs.est.5c04054. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Bazoge A, Wargny M, Constant Dit Beaufils P, et al. Assessing large language models for acute heart failure classification and information extraction from French clinical notes. Comput Biol Med . 2025;195:110609. 10.1016/j.compbiomed.2025.110609. [ DOI ] [ PubMed ] [ Google Scholar ] 22. Holmes G, Tang B, Gupta S, Venkatesh S, Christensen H, Whitton A. Applications of large language models in the field of suicide prevention: scoping review. J Med Internet Res . 2025;27:e63126. 10.2196/63126. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med . 2023;29(8):1930–40. 10.1038/s41591-023-02448-8. [ DOI ] [ PubMed ] [ Google Scholar ] 24. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health . 2023;2(2):e0000198. 10.1371/journal.pdig.0000198. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Tordjman M, Liu Z, Yuce M, et al. Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning. Nat Med . 2025. 10.1038/s41591-025-03726-3. [ DOI ] [ PubMed ] [ Google Scholar ] 26. Zhao Y, Wang S, Feng G. Insights into ChatGPT and DeepSeek application in osteoporotic vertebral compression fractures. Int J Surg . 2025. 10.1097/JS9.0000000000002512. [ DOI ] [ PubMed ] [ Google Scholar ] 27. Zhang J, Liu J, Guo M, Zhang X, Xiao W, Chen F. DeepSeek-assisted LI-RADS classification: AI-driven precision in hepatocellular carcinoma diagnosis. Int J Surg . 2025. 10.1097/JS9.0000000000002763. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Shang L, Sha S, Hou Y. Evaluating cardiovascular-kidney-metabolic syndrome knowledge in large language models: a comparative study of ChatGPT, Gemini, and DeepSeek. Diabetes Technol Ther . 2025. 10.1089/dia.2025.0216. [ DOI ] [ PubMed ] [ Google Scholar ] 29. Li J, Jiang Z. Chinese society of clinical oncology breast cancer (CSCO BC) guidelines in 2024: international contributions from china. Cancer Biol Med . 2024;21(10):838–43. 10.20892/j.issn.2095-3941.2024.0374. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Nabieva N, Brucker SY, Gmeiner B. ChatGPT’s agreement with the recommendations from the 18th St. Gallen International Consensus Conference on the Treatment of Early Breast Cancer. Cancers ( Basel ). 2024;16(24):4163. 10.3390/cancers16244163 [ DOI ] [ PMC free article ] [ PubMed ] 31. Alber DA, Yang Z, Alyakin A, et al. Medical large language models are vulnerable to data-poisoning attacks. Nat Med . 2025;31(2):618–26. 10.1038/s41591-024-03445-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Davenport T, Kalakota R. The potential for artificial intelligence in healthcare. Future Healthc J . 2019;6(2):94–8. 10.7861/futurehosp.6-2-94. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Griot M, Hemptinne C, Vanderdonckt J, Yuksel D. Large language models lack essential metacognition for reliable medical reasoning. Nat Commun . 2025;16(1):642. 10.1038/s41467-024-55628-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Chen Q, Hu Y, Peng X, et al. Benchmarking large language models for biomedical natural language processing applications and recommendations. Nat Commun . 2025;16(1):3280. 10.1038/s41467-025-56989-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Johnson KB, Wei WQ, Weeraratne D, et al. Precision medicine, AI, and the future of personalized health care. Clin Transl Sci . 2021;14(1):86–93. 10.1111/cts.12884. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature . 2023;620(7972):172–80. 10.1038/s41586-023-06291-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Bender EM, Gebru T, McMillan-Major A, Shmitchell S. On the dangers of stochastic parrots: can language models be too big? . In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . FAccT ’21. Association for Computing Machinery; 2021:610-623. 10.1145/3442188.3445922 38. Moy L. Guidelines for use of large language models by authors, reviewers, and editors: considerations for imaging journals. Radiology . 2023;309(1):e239024. 10.1148/radiol.239024. [ DOI ] [ PubMed ] [ Google Scholar ] 39. Thorp HH. ChatGPT is fun, but not an author. Science . 2023;379(6630):313. 10.1126/science.adg7879. [ DOI ] [ PubMed ] [ Google Scholar ] 40. Stokel-Walker C. ChatGPT listed as author on research papers: many scientists disapprove. Nature . 2023;613(7945):620–1. 10.1038/d41586-023-00107-z. [ DOI ] [ PubMed ] [ Google Scholar ] 41. Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA . 2023;330(9):866–9. 10.1001/jama.2023.14217. [ DOI ] [ PubMed ] [ Google Scholar ] 42. Bhattacharya S, Pradhan KB, Bashar MA, et al. Artificial intelligence enabled healthcare: a hype, hope or harm. J Family Med Prim Care . 2019;8(11):3461–4. 10.4103/jfmpc.jfmpc_155_19. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Chiang CH, Lee H yi. 2023 Can large language models be an alternative to human evaluations? arXiv . Preprint posted online. 10.48550/arXiv.2305.01937 44. He J, Baxter SL, Xu J, Xu J, Zhou X, Zhang K. The practical implementation of artificial intelligence technologies in medicine. Nat Med . 2019;25(1):30–6. 10.1038/s41591-018-0307-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Vayena E, Blasimme A, Cohen IG. Machine learning in medicine: addressing ethical challenges. PLoS Med . 2018;15(11):e1002689. 10.1371/journal.pmed.1002689. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 46. Liang S, Zhang J, Liu X, et al. The potential of large language models to advance precision oncology. EBioMedicine . 2025;115:105695. 10.1016/j.ebiom.2025.105695. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 47. Rengers TA, Thiels CA, Salehinejad H. Academic surgery in the era of large language models: a review. JAMA Surg . 2024;159(4):445–50. 10.1001/jamasurg.2023.6496. [ DOI ] [ PubMed ] [ Google Scholar ] 48. Ong JCL, Chang SYH, William W, et al. Ethical and regulatory challenges of large language models in medicine. Lancet Digit Health . 2024;6(6):e428–32. 10.1016/S2589-7500(24)00061-X. [ DOI ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary file1 (DOCX 19 KB) (19KB, docx) Supplementary file2 (DOCX 14 KB) (13.6KB, docx) Articles from Annals of Surgical Oncology are provided here courtesy of Springer ACTIONS View on publisher site PDF (845.8 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 25838 · SHA-256 d6ccfacd1380b97f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.