ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability.

Liu Y et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice JCO Clin Cancer Inform . 2026 Apr 8;10(2):e2500284. doi: 10.1200/CCI-25-00284 Search in PMC Search in PubMed View in NLM Catalog Add to search Toward Automating the Summarization of Cancer Pathology Reports Using Large Language Models to Improve Clinical Usability Yirong Liu Yirong Liu , MD, PhD 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Yirong Liu 1 , Jacob John Jacob John , MS 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Jacob John 1 , Sagnik Sarkar Sagnik Sarkar , MS 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Sagnik Sarkar 1 , Abdul Zakkar Abdul Zakkar , MD 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Abdul Zakkar 1 , Paul Kinkopf Paul Kinkopf , BS 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Paul Kinkopf 1 , P Troy Teo P Troy Teo , PhD 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by P Troy Teo 1, ✉ , Mohamed E Abazeed Mohamed E Abazeed , MD, PhD 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL Find articles by Mohamed E Abazeed 1 Author information Article notes Copyright and License information 1 Department of Radiation Oncology, Northwestern University Feinberg School of Medicine, Chicago, IL ✉ Corresponding author: P. Troy Teo, PhD; e-mail: [email protected] . Received 2025 Sep 14; Revised 2026 Jan 21; Accepted 2026 Feb 27; Collection date 2026. © 2026 by American Society of Clinical Oncology Creative Commons Attribution Non-Commercial No Derivatives 4.0 License: https://creativecommons.org/licenses/by-nc-nd/4.0/ PMC Copyright notice PMCID: PMC13068448  PMID: 41950436 Abstract PURPOSE Reviewing pathology reports requires physicians to integrate complex histopathologic, immunohistochemical, and molecular findings from multiple reports and institutions, often under time constraints that increase the risk of error and fatigue. Large language models (LLMs) offer a potential solution by generating concise, coherent summaries from complex pathology data. METHODS Patients who underwent initial consultation in a thoracic clinic between January 2019 and July 2023 were included. Original pathology reports and corresponding physician pathology summaries from consultation notes were extracted and anonymized. Six open-source LLMs (Llama 3.0, Llama 3.1, Llama 3.2, Mistral, Gemma, and DeepSeek-R1) generated pathology summaries directly from the original reports. Objective and subjective evaluations were performed using the original reports as the ground truth. LLM-generated summaries were compared with physician summaries for correctness, completeness, and conciseness. Additional subjective assessments with multiple evaluators were conducted for Llama 3.1. RESULTS Ninety-four cases met the eligibility criteria. Using the original pathology reports as the ground truth, the LLM-generated summaries achieved higher scores across all objective evaluation metrics compared with physician pathology summaries ( P < .0001). In the subjective evaluation, DeepSeek, Mistral, Llama 3.1, and Llama 3.2 achieved higher ratings for completeness ( P = .017, P < .0001, P < .0001, and P < .0001, respectively) while maintaining comparable correctness relative to physician pathology summaries ( P = 1.000). The results remained consistent in additional subjective analyses involving multiple evaluators for Llama 3.1. CONCLUSION LLM-generated summaries demonstrated better performance in objective metrics and greater completeness in subjective evaluations compared with physician summaries. These results highlight the potential of LLMs as valuable tools for enhancing clinical documentation and workflow efficiency in oncology practice. INTRODUCTION Reviewing a multitude of pathology reports is a vital yet demanding task for physicians, requiring the synthesis of fragmented and complex information. These reports often span multiple dates, clinical sites, and laboratories, making it difficult to consolidate findings into a cohesive clinical picture. 1 Physicians must interpret detailed data, including histopathologic findings, immunohistochemical results, and more extensive genetic and molecular test results. 2 , 3 The variability in report formats and terminologies across institutions adds further complexity, requiring additional effort to reconcile differences and accurately adjudicate findings. 4 Time constraints amplify the challenge as physicians must make timely decisions that directly affect patient care. The need to review and cross-reference multiple reports leaves little room for error, yet the risk of oversight increases with the volume of information. 5 Subtle but significant details, such as specific genetic markers or test variations, can easily be missed under pressure, potentially leading to suboptimal treatment decisions. In addition, the repetitive and detail-oriented nature of pathology review often results in cognitive fatigue, especially when managing a high patient volume. Fatigue can impair focus, increase the likelihood of errors, and further complicate decision making in an already challenging process. 6 , 7 CONTEXT Key Objective To determine whether open-source large language models (LLMs) can accurately and comprehensively summarize complex oncologic pathology reports, providing a more complete and efficient alternative to physician summaries. Knowledge Generated Across 94 cases and six LLMs, model-generated summaries showed higher lexical and semantic similarity to original pathology reports. Most LLMs also produced more complete summaries, particularly capturing genetic findings, while maintaining comparable correctness and low potential harm. Relevance (M. Behera) This work highlights LLM-generated summaries can reduce documentation burden and improve information accessibility for clinicians, potentially allowing more time for direct patient care while maintaining or enhancing the quality of clinical communication.* *Relevance section written by JCO CCI Associate Editor Madhusmita Behera, PhD. Large language models (LLMs), based on transformer architectures like generative pre-trained transformer, represent a leap in natural language processing. 8 , 9 Trained on vast data sets, they enable applications in health care, education, finance, and more. 8 , 10 , 11 Using LLMs to summarize multiple pathology reports might offer significant advantages in streamlining clinical workflows and enhancing medical decision making. LLMs can efficiently synthesize these complex, multidimensional pathology results into concise, coherent summaries, reducing the cognitive load on health care professionals. This capability allows clinicians to quickly grasp the key findings and trends without manually reviewing each report in detail. By also integrating genetic data, LLMs create the platform to elucidate potential associations between molecular alterations, histopathologic findings, and therapeutic approaches. Moreover, their scalability makes them particularly advantageous in high-volume clinical settings, such as large hospitals, where automating the summarization process can alleviate physicians' workload, allowing more time for complex analyses and direct patient care. In this study, we explored the feasibility of using LLMs to summarize oncologic pathology reports, aiming to assess their ability to generate complete, accurate, and concise summaries without missing critical information. METHODS This study was designed and conducted in strict compliance with the Health Insurance Portability and Accountability Act (HIPAA) regulations to safeguard the privacy and security of protected health information (PHI). All collected data were deidentified to ensure participant confidentiality, and access to PHI was restricted exclusively to authorized personnel directly engaged in the research. The study protocol received approval from the Institutional Review Board (IRB) at Northwestern University (STU00212113). A schematic of the study workflow is presented in Figure 1 . FIG 1. Open in a new tab LLM pathology report summarization workflow. Oncology pathology reports are processed by various LLMs, including Llama 3.0, Llama 3.1, Llama 3.2, Mistral, Gemma, and DeepSeek, to produce output summaries. These outputs are evaluated through objective metrics and subjective assessments conducted by clinical experts. The combined evaluations provide a comprehensive analysis of the LLMs' summarization performance. LLM, large language model. Patients Patients seen for initial consultation in the thoracic radiation oncology clinic at Northwestern Memorial Hospital between January 2019 and July 2023 were included. Eligibility required a pathologically confirmed diagnosis, pathology reports available in the Northwestern Memorial Hospital electronic medical record (EMR), and a physician-authored pathology summary (Notes) within the consultation note. Eligible reports included histopathologic findings, immunohistochemical results, and molecular or genetic testing (eg, next-generation sequencing [NGS]) within 6 months from the date of consultation. All qualifying reports per patient were included. To account for the temporal lag in molecular pathology testing such as NGS, reports were included only if results were finalized before the date of consultation. This ensured that the clinical summaries and LLM-generated summaries were compared against the same set of finalized pathology data. External image-only pathology reports were excluded. All data were extracted from the EMR and anonymized. Anonymization and LLM Deployment All pathology reports and consultation notes were deidentified using obi/deid_roberta_i2b, 12 a RoBERTa-based PHI detection model trained on the I2B2 data set, 13 integrated with a SpaCy preprocessing pipeline, Presidio's Analyzer, and Anonymizer Engines for automated removal or masking of PHI. 14 - 16 Summaries were then generated using six open-source LLMs (DeepSeek-R1, Gemma 2, Llama 3.0/3.1/3.2, and Mistral 7.2B, Data Supplement, Table S1) deployed on-premise within a HIPAA-compliant environment using Ollama. 17 Standardized inference parameters ensured reproducibility, and inputs exceeding model context limits were automatically truncated without additional preprocessing. See Supplementary Methods for full technical details. Objective Evaluation We systematically evaluated the feasibility of LLM-generated oncopathology summaries using a multimodal framework that incorporated both standard natural language processing (NLP) metrics and domain-specific semantic assessments. Given the heterogeneity and technical density of pathology reports, including histopathologic descriptions, immunohistochemical profiles, and molecular findings, we assessed summary quality across dimensions of lexical fidelity, syntactic structure, semantic similarity, and medical relevance. Lexical fidelity was quantified using ROUGE (ROUGE-1/2/L), which measures n-gram overlap and longest common subsequence alignment to evaluate surface-level completeness and structural correspondence. 18 Syntactic precision was further examined using BLEU (BLEU-1 to BLEU-4) with default brevity-penalty settings to capture word- and phrase-level similarity. 19 Because n-gram metrics do not reflect conceptual accuracy, we incorporated METEOR to account for synonymy, stemming, and word reordering 20 using the Bio_ClinicalBERT tokenizer to ensure appropriate handling of biomedical terminology. 21 To evaluate semantic fidelity, we used BERTScore, which computes cosine similarity between contextual embeddings of reference and generated summaries. 22 Given the length of pathology reports, embeddings were derived using the Llama-3-8B-ProLong-64k model, enabling full-context evaluation without truncation. All metrics were implemented using PyTorch and NLTK toolkits. 23 Statistical significance was assessed using paired Wilcoxon signed-rank tests with Bonferroni correction. Expanded methodological details are provided in the Supplementary Methods. Subjective Evaluation Subjective evaluation consisted of two complementary analyses: (1) factual relation with the original pathology reports and (2) alignment of information with physician summaries. Using the pathology report as ground truth, evaluators assessed factual correctness and completeness on a standardized 5-point Likert scale, along with a separate 5-point metric for potential harm. Potential harm was categorized as either omission of clinically relevant information or introduction of incorrect, contradictory, or hallucinated content. Because omissions in pathology summaries may be mitigated by additional clinical data, harm ratings accounted for the degree to which missing information could directly influence treatment decisions. By contrast, hallucinated or incorrect details, particularly staging- or genomics-related content, were considered higher risk. Expanded definitions and examples are provided in the Data Supplement (Table S2). To further explore the clinical utility of the LLM summaries, additional comparisons were made between LLM summaries and physician notes across correctness, completeness, and conciseness. Group differences were tested using the Kruskal-Wallis method followed by Dunn's post hoc analysis with Bonferroni correction. Subjective evaluation proceeded in two phases: an initial screening by a senior radiation oncology resident, followed by independent assessments by two additional residents for the selected model. Llama 3.1 was chosen for detailed evaluation based on consistent performance across metrics. Inter-rater agreement was quantified using Fleiss' Kappa, and consensus ratings were assigned when at least two evaluators agreed. Full methodological details are provided in the Supplementary Methods. RESULTS A total of 94 cases were included in this study. On average, the preprocessed pathology reports from each case produced 6,152 tokens as input for LLMs after anonymization, with a maximum token count of 17,312. Six open-source LLMs, Llama 3.0, Llama 3.1, Llama 3.2, Mistral, Gemma, and DeepSeek, were used to generate summaries of the pathology reports. Figure 2 provides illustrative and representative examples of pathology summaries derived from Notes and LLM outputs. The Notes included relevant histopathologic and biomarker information but omitted detailed genetic test results. The majority of LLMs demonstrated satisfactory performance, both objectively and subjectively, effectively summarizing key clinicopathologic information and presenting genetic findings in concise, albeit slightly varied, formats. However, Llama 3.0 exhibited an error by misinterpreting biomarker quality control data as positive test results, incorrectly reporting positive NTRK1, NTRK1, NTRK3, and RET fusions. After discussion within the research group, this type of error was penalized under correctness rather than hallucination. Another common error observed in the cohort was the generation of unusable outputs by certain LLMs, characterized by either meaningless text or empty files. Specifically, Llama 3.0 produced 21 unusable cases (22%, 21 of 94), whereas Gemma generated 14 (15%, 14 of 94). By contrast, Mistral produced only three such cases (3%, 3 of 94), and no unusable outputs were observed for DeepSeek, Llama 3.1, or Llama 3.2. These unusable outputs were assigned a score of 1 on the Likert scales for correctness and completeness in the subjective evaluation. FIG 2. Open in a new tab Samples of physician summaries and LLM-generated summaries. AP, anatomical pathology; CNV, copy number variation; FDA, US Food and Drug Administration; FFPE, formalin-fixed paraffin-embedded; FNA, fine-needle aspiration; LLM, large language model; MSI, microsatellite instability; NGS, next generation sequencing; NM, Northwestern Memorial; RUL, right upper lobe; SNV, single nucleotide variant; TMB, tumor mutation burden; TTF-1, thyroid transcription factor 1; VAF, variant allele frequency. Objective Evaluation We compared various LLMs using standard summarization metrics: BLEU, ROUGE (ROUGE-1, ROUGE-2, ROUGE-L), modified BERTScore, and METEOR (Fig 3 , Data Supplement, Table S3). The Notes demonstrated relatively lower values compared with LLMs across all metrics ( P < .0001), indicating lower alignment of physician pathology summaries with original pathology reports as ground truth. Mistral showed moderate performance but did not consistently outperform other LLMs. In BLEU, LLMs, including Mistral, Llama 3.0, Llama 3.1, Llama 3.2, DeepSeek, and Gemma, achieved higher median scores than Notes. A similar trend was observed for ROUGE-1, ROUGE-2, and ROUGE-L, where LLMs outperformed Notes but showed no significant differences among themselves. METEOR, which integrates synonym matching and stemming, demonstrated a similar pattern, with LLMs outperforming Notes while showing minimal variance between models. Comparable findings were observed with the modified BERTScore. Overall, modern LLMs generated summaries well-aligned with pathology reports but lacked consistent performance differentiation. FIG 3. Open in a new tab Objective evaluation metrics for LLM-generated summaries. Boxplots display the performance of different LLMs across six standard natural language generation evaluation metrics: BLEU, ROUGE-1, ROUGE-2, ROUGE-L, METEOR, and Llama_BERT. Each box represents the IQR of performance scores, with the horizontal line within each box indicating the median value. Whiskers extend to 1.5 times the IQR, and points beyond this range are considered outliers. The evaluated models include Notes (representing physician pathology summaries as a baseline), Mistral, Llama 3.0, Llama 3.1, Llama 3.2, Gemma, and Deepseak. “ns” denotes nonsignificant differences ( P > .05) between models; statistically significant comparisons are not displayed in the figure. LLM, large language model. Subjective Evaluation LLM-generated summaries were evaluated through two distinct analyses: (1) factual relation with the original pathology reports and (2) alignment of information with physician summaries. Figures 4 A- 4 C shows the results of the factual relation metrics using original pathology reports as the ground truth. DeepSeek and Mistral exhibited the highest correctness, with “Same” ratings of 97% (91 of 94) and 91% (86 of 94), respectively, whereas Llama 3.0 and Gemma received the highest proportions of “Significantly worse” ratings (21% (20 of 94) and 14% (13 of 94), respectively). In terms of completeness, DeepSeek, Mistral, Llama 3.1, and Llama 3.2 were rated as “Same” in the majority of cases (69%-83%), whereas Llama 3.0 and Gemma demonstrated higher proportions of “Significantly worse” and “Slightly worse” ratings (21%-22% and 26%-30%, respectively). Although Notes exhibited high correctness (99% [93 of 94] rated as “Same”), it underperformed in completeness, with 56% of cases rated as “Slightly worse” or “Significantly worse” primarily because of the omission of genetic information. FIG 4. Open in a new tab Subjective evaluation metrics for LLM-generated summaries. Stacked bar charts display subjective evaluations of LLM-generated summaries using two ground truth references: original pathology reports ((A-C): path-correctness, path-completeness, potential harm) and physician pathology summaries ((D-F): notes-correctness, notes-completeness, notes-conciseness). Results are presented as percentages of responses categorized into five levels: “Significantly worse,” “Slightly worse,” “Same,” “Slightly better,” and “Significantly better” for performance metrics, and “Not at all,” “Slightly,” “Moderately,” “Very,” and “Extremely” for harm assessment. For metrics using original pathology reports as ground truth, pairwise comparisons were conducted using physician pathology summaries (Notes) as the baseline. Percentages over left bars indicate the proportion rated ≥slightly worse; percentages over right bars indicate the proportion rated ≥slightly better. Total cases: 94. Statistically significant differences are indicated by asterisks: * P < .05, ** P < .01, *** P < .001, **** P < .0001. Comparisons yielding nonsignificant results are not displayed in the figure. LLM, large language model. To assess the alignment of information with physician summaries, we used the Notes as the ground truth and evaluated the LLM-generated summaries for correctness, completeness, and conciseness (Figs 4 D- 4 F). In terms of correctness, DeepSeek, Mistral, Llama 3.2, and Llama 3.1 demonstrated high fidelity to the Notes, with “Same” ratings ranging from 89% to 96%. By contrast, Llama 3.0 and Gemma exhibited a greater proportion of “Significantly worse” ratings (15%-21%). Regarding completeness, DeepSeek, Mistral, Llama 3.2, and Llama 3.1 significantly outperformed Notes, with a combined “Slightly better” and “Significantly better” rating of 52%-56%. Conversely, Llama 3.0 and Gemma received a higher percentage of “Slightly worse” and “Significantly worse” ratings (15%-24% and 16%-21%, respectively). The increased completeness seen in most LLMs is primarily due to their ability to extract and present more comprehensive genetic information. Since many LLM-generated summaries demonstrated greater completeness than Notes, we further evaluated conciseness to determine whether LLMs could deliver succinct summaries. DeepSeek summaries maintained conciseness relative to Notes, with the majority rated as “Same” (92%, 87 of 94). By contrast, Llama 3.0 and Gemma demonstrated higher redundancy, with “Slightly worse” and “Significantly worse” ratings ranging from 33% to 39% and 13%-24%, respectively. Recognizing that not all pathology information carries equal weight in clinical decision making, we evaluated the potential harm associated with both Notes and LLM-generated summaries. Notes contained pertinent information, with only 3% of cases exhibiting slight potential harm. DeepSeek, Mistral, Llama 3.2, and Llama 3.1 showed slightly higher rates of potential harm but without statistical significance (9%-16%, P = .673, 1.000, 1.000, 0.673, respectively). By contrast, Llama 3.0 and Gemma generated summaries with a higher incidence and severity of potential harm ( P < .0001). We further examined the impact of evaluator variation on subjective assessments of Llama 3.1–generated summaries (Fig 5 A). The analysis revealed varied degree of agreements among evaluators, with the summaries receiving comparable ratings across all dimensions (Fig 5 A, Data Supplement, Table S4, Fleiss' Kappa 0.182-0.464). Notably, LLM-generated summaries demonstrated improved completeness while maintaining correctness and exhibiting no increase in perceived potential harm. These findings remained robust after consolidating evaluator assessments (Fig 5 B), indicating that individual evaluator differences did not significantly alter the overall evaluation outcomes. FIG 5. Open in a new tab Subjective evaluation metrics for Llama 3.1–generated summaries across different evaluators. (A) Results from individual evaluator assessments. (B) Consolidated results combining assessments from all evaluators. Percentages over left bars indicate the proportion rated ≥slightly worse; percentages over right bars indicate the proportion rated ≥slightly better. Total cases: 94. DISCUSSION The present study represents, to our knowledge, one of the first comprehensive evaluations of LLMs for summarization of oncopathologic reports in a clinical application, 24 directly benchmarking against physician-authored summaries. While recent work has explored the use of LLMs for synoptic or structured reporting of cancer pathology data, this study is distinct in assessing free-text summarization for clinical practice, incorporating both objective NLP metrics and subjective clinical evaluations. We benchmarked the performance of six open-source LLMs against physician-generated summaries. Our findings indicate that most LLMs effectively summarized pathology reports, accurately extracting key histopathologic and genetic data. LLMs demonstrated higher scores than physician summaries on objective metrics reflecting language similarity and completeness. Recent studies highlight both the promise and challenges of applying LLMs to clinical text summarization. Van Veen et al 25 demonstrated that adapted LLMs can match or surpass expert-generated summaries across multiple clinical tasks. Oliveira et al 26 developed discharge summary systems achieving near–human-level diagnostic interpretation, underscoring practical potential. However, a scoping review by Bednarczyk et al 27 found that most studies remain exploratory, with limited validation and safety evaluation. Tang et al cautioned that LLMs remain prone to factual inconsistencies and misinformation, emphasizing the need for robust evaluation before clinical integration. 28 On closer examination of these studies, it becomes evident that even within the domain of clinical text summarization, task-specific tools and evaluation frameworks are necessary. For instance, the objectives and contextual emphases of clinical note summarization may differ substantially from those of pathology report summarization. Furthermore, the performance of various LLMs demonstrates considerable variability across tasks. This is corroborated by findings from our study. Although all six tested LLMs demonstrated statistically superior performance compared with Notes in objective evaluations, Gemma and Llama 3.0 exhibited relatively lower performance than the other models. These two models also generated a higher number of unusable outputs, where no meaningful text was generated or only empty files were returned. Llama 3.0 generated 21 unusable cases, whereas Gemma produced 14. By contrast, Mistral had only 3 cases and no unusable cases were observed in DeepSeek, Llama 3.1, or Llama 3.2. A potential explanation for these differences is the context length limitation imposed by the models' processing capacity. Cases with unusable summaries in Llama 3.0 and Gemma exhibited significantly longer token lengths compared with those with normal summaries (Data Supplement, Table S5). Gemma, designed as a lightweight LLM, supports a context length of 8,192 tokens. 29 Llama 3.0, despite having a larger model size (8B and 70B parameters available), shares the same 8,192-token context length. 30 By contrast, Mistral supports a 32K token context, 31 whereas Llama 3.1, Llama 3.2, and DeepSeek each accommodate up to 128k tokens. 32 - 34 When input text exceeds a model's token limit, LLMs may truncate information, return errors or failures, or produce unpredictable outputs. Prior work has demonstrated a significant decline in performance as input length increases, even before reaching the model's maximum context capacity. 35 Several strategies have been proposed to mitigate issues arising from excessive input length, aiming to reduce computational load and graphics processing unit memory demands, including input truncation, pruning, splitting, and other computational optimizations. 36 , 37 Although context length limitations explain unusable outputs in LLaMA 3.0 and Gemma, the three unusable summaries observed with Mistral could not be attributed to input length as all reports were well within the supported context window. These failures likely reflect model-specific generation instabilities or differences in instruction adherence rather than truncation effects and warrant further investigation. By contrast, we observed that DeepSeek outperformed other LLMs. This superior performance may reflect its later-generation post-training paradigm, which leverages large-scale reinforcement learning (RL) for reasoning rather than relying solely on supervised fine-tuning. The integration of cold-start data and RL-driven structured chain-of-thought reasoning likely enhances the model's ability to prioritize key clinical information, minimize omissions, and reduce the risk of harmful interpretations. However, because detailed architectural and training specifications are not fully disclosed, definitive attribution of these performance differences remains limited by model transparency. In potential harm evaluation, we classified risks into two primary categories: omissions of relevant information and incorrect, contradictory, or hallucinated content. Concerning omission, our analysis revealed varying degrees of omission across LLMs and cases. As to hallucination, it was uncommon overall although this finding should be interpreted cautiously given the limited sample size and corresponding constraints in detecting rare errors. Only Gemma and Llama 3.0 produced clinically relevant errors, whereas the other models performed well without major inaccuracies that could affect clinical decisions. Altogether, the more frequent omission of information findings compared with hallucinations highlights the importance of prioritizing efforts to minimize the omission of pertinent information to further reduce potential harm, while continuing to monitor for rare but potentially consequential hallucinations that may emerge with larger data sets. Task-specific fine-tuning has been suggested as a potential way to reduce data omissions and improve model performance. 38 While subjective assessments revealed some heterogeneity in LLM performance, our data support that the majority of LLMs demonstrate enhanced completeness, largely driven by their ability to present more comprehensive genetic information. The greater genetic information completeness observed for LLMs may partly reflect the use of an explicit, structured prompt, whereas physician summaries were generated without standardized instructions and reflect routine clinical documentation. With the increasing adoption of NGS in clinical practice, physicians are increasingly inundated with extensive pathologic data. Our observations indicate substantial variability in how physicians incorporate relevant genetic findings, influenced by their individual backgrounds and expertise. Information flow between pathology and oncology is often suboptimal, and strategies to streamline this communication are needed. 39 LLMs have the potential to provide significant value by delivering standardized, accurate, and clinically pertinent genetic data, supporting physicians in navigating the growing complexity of genetic information in modern oncology. The current study has several limitations. Because of financial constraints and institutional requirements, our analysis was limited to open-source LLMs, which may not perform as well on some tasks as proprietary models. To mitigate this, we incorporated a diverse selection of models, including all available versions from the Llama 3 series during the study period and DeepSeek, a cutting-edge LLM in the field. By evaluating six distinct models, we aimed to capture a broad representation of current LLM capabilities, and the resulting data demonstrate robustness in supporting LLM performance for oncopathologic report summarization. Second, because of resource limitations, we were unable to fine-tune the LLMs specifically for this task. Nevertheless, several models, particularly newer ones like DeepSeek, Llama 3.1, and Llama 3.2, demonstrated strong performance. We also observed a performance plateau among these models. Given the diminishing gains observed among newer models, future work should investigate whether task-specific fine-tuning can yield clinically meaningful improvements in performance. Third, this study was conducted at a single center and external validation is necessary to establish the generalizability of the findings across diverse institutions and data sources. Furthermore, part of the evaluation relied on language-based metrics such as BLEU and ROUGE, which may introduce bias by emphasizing surface-level text similarity rather than true clinical accuracy or interpretability. To mitigate this limitation, physician-generated summaries were incorporated as the ground truth in the subjective evaluation. In addition, the absence of pathologist involvement might have limited detection of certain histopathologic or molecular interpretation errors and reproducibility was not fully assessed. Future research should incorporate multicenter data sets, validated clinical end points, and more comprehensive assessments of inter-rater reliability to further strengthen the evidence base for LLM-assisted pathology summarization. In conclusion, this study demonstrates that open-source LLMs can accurately condense oncopathologic reports while preserving key clinical information and improving completeness relative to physician summaries. This work establishes a foundation for prospective evaluation of LLM-assisted summarization to reduce documentation burden and improve clinical workflow. Paul Kinkopf Employment: Memorial Sloan-Kettering Cancer Center P. Troy Teo Employment: Northwestern Medicine Research Funding: Canadian Institutes of Health Research (CIHR) Patents, Royalties, Other Intellectual Property: Pending patent for 63/531?036 (provisional patent) with Northwestern University IP disclosure Disc-ID-23-05-25-001 (Inst) Mohamed E. Abazeed Stock and Other Ownership Interests: Atomo, Inc Honoraria: IDEOlogy Health Research Funding: Varian Medical Systems, Inc (Inst) Patents, Royalties, Other Intellectual Property: iSeg, Autosegementation Technology for Radiotherapy Open Payments Link: https://openpaymentsdata.cms.gov/physician/1194464 No other potential conflicts of interest were reported. SUPPORT Supported in part by a fellowship from the Canadian Institutes of Health Research (CIHR-472392) and Social Impact funding from Amazon Web Services for Dr P.T.T. * Y.L. and J.J. contributed equally to this work. P.T.T. and M.E.A. contributed equally to this work. AUTHOR CONTRIBUTIONS Conception and design: Yirong Liu, Jacob John, P. Troy Teo, Mohamed E. Abazeed Administrative support: Mohamed E. Abazeed Provision of study materials or patients: Mohamed E. Abazeed Collection and assembly of data: Yirong Liu, Jacob John, Paul Kinkopf Data analysis and interpretation: All authors Manuscript writing: All authors Final approval of manuscript: All authors Accountable for all aspects of the work: All authors AUTHORS' DISCLOSURES OF POTENTIAL CONFLICTS OF INTEREST The following represents disclosure information provided by authors of this manuscript. All relationships are considered compensated unless otherwise noted. Relationships are self-held unless noted. I = Immediate Family Member, Inst = My Institution. Relationships may not relate to the subject matter of this manuscript. For more information about ASCO's conflict of interest policy, please refer to www.asco.org/rwc or ascopubs.org/cci/author-center . Open Payments is a public database containing information reported by companies about payments made to US-licensed physicians ( Open Payments ). Paul Kinkopf Employment: Memorial Sloan-Kettering Cancer Center P. Troy Teo Employment: Northwestern Medicine Research Funding: Canadian Institutes of Health Research (CIHR) Patents, Royalties, Other Intellectual Property: Pending patent for 63/531?036 (provisional patent) with Northwestern University IP disclosure Disc-ID-23-05-25-001 (Inst) Mohamed E. Abazeed Stock and Other Ownership Interests: Atomo, Inc Honoraria: IDEOlogy Health Research Funding: Varian Medical Systems, Inc (Inst) Patents, Royalties, Other Intellectual Property: iSeg, Autosegementation Technology for Radiotherapy Open Payments Link: https://openpaymentsdata.cms.gov/physician/1194464 No other potential conflicts of interest were reported. REFERENCES 1. Ektefaie Y, Yuan W, Dillon DA, et al. : Integrative multiomics-histopathology analysis for breast cancer classification. NPJ Breast Cancer 7:147, 2021 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Burke Y, Parens W, Chung WK, et al. : The challenge of genetic variants of uncertain clinical significance. Ann Intern Med 175:994-1000, 2022 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Harris TJR, McCormick F: The molecular pathology of cancer. Nat Rev Clin Oncol 7:251-265, 2010 [ DOI ] [ PubMed ] [ Google Scholar ] 4. Morgan S, Hanna J, Yousef GM: Knowledge translation in oncology: The bumpy ride from bench to bedside. Am J Clin Pathol 153:5-13, 2020 [ DOI ] [ PubMed ] [ Google Scholar ] 5. Yip S, Christofides A, Banerji S, et al. : A Canadian guideline on the use of next-generation sequencing in oncology. Curr Oncol 26:e241-e254, 2019 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. West CP, Tan AD, Habermann TM, et al. : Association of resident fatigue and distress with perceived medical errors. JAMA 302:1294-1300, 2009 [ DOI ] [ PubMed ] [ Google Scholar ] 7. Ancker JS, Edwards A, Nosal S, et al. : Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med Inform Decis Mak 17:36-39, 2017 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. : Large language models in medicine. Nat Med 29:1930-1940, 2023 [ DOI ] [ PubMed ] [ Google Scholar ] 9. Naveed H, Khan AU, Qiu S, et al. : A comprehensive overview of large language models. arXiv 10.1145/3744746 [ DOI ] [ Google Scholar ] 10. Wu S, Irsoy O, Lu S, et al. : Bloomberggpt: A large language model for finance. arXiv 10.48550/arXiv.2303.17564 [ DOI ] [ Google Scholar ] 11. Yalamanchili A, Sengupta B, Song J, et al. : Quality of large language model responses to radiation oncology patient care questions. JAMA Netw Open 7:e244630, 2024 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Liu Y, Ott M, Goyal N, et al. : Roberta: A robustly optimized BERT pretraining approach. arXiv 10.48550/arXiv.1907.11692 [ DOI ] [ Google Scholar ] 13. Stubbs A, Uzuner Ö: Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. J Biomed Inform 58:S20-S29, 2015. (suppl) [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Microsoft: Presidio—Data Protection and De-identification SDK. https://microsoft.github.io/presidio/ 15. Patchipala S: Data anonymization in AI and ML engineering: Balancing privacy and model performance using presidio. IRE J 7:1-10, 2023 [ Google Scholar ] 16. Böhlin F: Detection & Anonymization of Sensitive Information in Text: Ai-Driven Solution for Anonymization, 2024 [ Google Scholar ] 17. Marcondes FS, Gala A, Magalhães R, et al. Using Ollama, Natural Language Analytics with Generative Large-Language Models: A Practical Approach with Ollama and Open-Source LLMs. Heidelberg, Germany: Springer; 2025:23–35. [ Google Scholar ] 18. Lin C-Y: Rouge: A package for automatic evaluation of summaries, text summarization branches out. 2004, pp 74-81 [ Google Scholar ] 19. Papineni K, Roukos S, Ward T, et al. : Bleu: A method for automatic evaluation of machine translation, in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp 311-318 [ Google Scholar ] 20. Banerjee S, Lavie A: METEOR: An automatic metric for MT evaluation with improved correlation with human judgments, in Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005, pp 65-72 [ Google Scholar ] 21. Alsentzer E, Murphy JR, Boag W, et al. : Publicly available clinical BERT embeddings. arXiv 10.48550/arXiv.1904.03323 [ DOI ] [ Google Scholar ] 22. Zhang T, Kishore V, Wu F, et al. : Bertscore: Evaluating text generation with BERT. arXiv 10.48550/arXiv.1904.09675 [ DOI ] [ Google Scholar ] 23. Ding N, Chen Y, Xu B, et al. : Enhancing chat language models by scaling high-quality instructional conversations. arXiv 10.48550/arXiv.2305.14233 [ DOI ] [ Google Scholar ] 24. Rajaganapathy S, Chowdhury S, Buchner V, et al. Synoptic reporting by summarizing cancer pathology reports using large language models. medRxiv. 2024. 10.1101/2024.04.26.24306452. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Van Veen D, Van Uden C, Blankemeier L, et al. : Adapted large language models can outperform medical experts in clinical text summarization. Nat Med 30:1134-1142, 2024 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Oliveira JD, Santos HD, Ulbrich AHD, et al. : Development and evaluation of a clinical note summarization system using large language models. Commun Med 5:376, 2025 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Bednarczyk L, Reichenpfader D, Gaudet-Blavignac C, et al. : Scientific evidence for clinical text summarization using large language models: Scoping review. J Med Internet Res 27:e68998, 2025 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Tang L, Sun Z, Idnay B, et al. : Evaluating large language models on medical evidence summarization. NPJ Digit Med 6:158, 2023 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Team G : Gemma 3 Technical Report. 2025 [ Google Scholar ] 30. Grattafiori A, Dubey A, Jauhri A, et al. : The llama 3 herd of models. arXiv 10.48550/arXiv.2407.21783 [ DOI ] [ Google Scholar ] 31. Jiang AQ, Sablayrolles A, Roux A, et al. : Mixtral of experts. arXiv 10.48550/arXiv.2401.04088 [ DOI ] [ Google Scholar ] 32. Team ML: Introducing Llama 3.1: Our most capable models to date. 2024. https://ai.meta.com/blog/meta-llama-3-1/ [ Google Scholar ] 33. Team ML: Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. 2024. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ [ Google Scholar ] 34. Guo D, Yang D, Zhang H, et al. : Deepseek-r1: Incentivizing reasoning capability in LLMS via reinforcement learning. arXiv 10.48550/arXiv.2501.12948 [ DOI ] [ Google Scholar ] 35. Levy M, Jacoby A, Goldberg Y: Same task, more tokens: The impact of input length on the reasoning performance of large language models. arXiv 10.48550/arXiv.2402.14848 [ DOI ] [ Google Scholar ] 36. Xiao G, Tang J, Zuo J, et al. : Duoattention: Efficient long-context LLM inference with retrieval and streaming heads. arXiv 10.4855/arXiv.2410.10819 [ DOI ] [ Google Scholar ] 37. Liu D, Chen M, Lu B, et al. : Retrievalattention: Accelerating long-context LLM inference via vector retrieval. arXiv 10.48550/arXiv.2409.10516 [ DOI ] [ Google Scholar ] 38. Anisuzzaman DM, Malins JG, Friedman PA, et al. : Fine-tuning large language models for specialized use cases. Mayo Clinic Proc Digit Health 3:100184, 2025 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Mirham L, Hanna J, Yousef GM: Addressing the diagnostic miscommunication in pathology: Old challenges and innovative solutions. Am J Clin Pathol 156:521-528, 2021 [ DOI ] [ PubMed ] [ Google Scholar ] Articles from JCO Clinical Cancer Informatics are provided here courtesy of Wolters Kluwer Health ACTIONS View on publisher site PDF (2.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 713 · SHA-256 42bfa08ca47964d7
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.