Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

Data, Models, and Visuals: How Data Science Methods Can Augment (Geriatric) Oncology Research.

Ramsdale E et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Data, Models, and Visuals: How Data Science Methods Can Augment (Geriatric) Oncology Research - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice J Clin Oncol . Author manuscript; available in PMC: 2026 Apr 20. Published in final edited form as: J Clin Oncol. 2025 Mar 6;43(12):1404–1407. doi: 10.1200/JCO-25-00053 Search in PMC Search in PubMed View in NLM Catalog Add to search Data, Models, and Visuals: How Data Science Methods Can Augment (Geriatric) Oncology Research Erika Ramsdale Erika Ramsdale , MD 1 James P. Wilmot Cancer Center, University of Rochester Medical Center, NY, USA Find articles by Erika Ramsdale 1 , Supriya Mohile Supriya Mohile , MD 1 James P. Wilmot Cancer Center, University of Rochester Medical Center, NY, USA Find articles by Supriya Mohile 1 Author information Article notes Copyright and License information 1 James P. Wilmot Cancer Center, University of Rochester Medical Center, NY, USA ✉ Address correspondence to: Supriya Mohile, MD, James P. Wilmot Cancer Center, 601 Elmwood Avenue, Box 704, Rochester, NY 14642, Fax: 585-273-1051, [email protected] Issue date 2025 Apr 20. PMC Copyright notice PMCID: PMC12003067  NIHMSID: NIHMS2056259  PMID: 40048687 The publisher's version of this article is available at J Clin Oncol Abstract In the article that accompanies this editorial, Etienne Audureau and co-authors highlight the importance of GA variables in predicting prognosis in two large observational French cohorts of older adults with cancer. Beyond the specific problem (predicting prognosis in older adults with cancer) and results, they provide a second type of useful model: an illustration of how data science, machine learning, and many-model thinking can augment clinical research amidst a shifting data paradigm. Most patients diagnosed with cancer are older than 65. 1 Older adults with cancer are heterogeneous, with a complex mix of attributes that contribute to health, functioning, and cancer outcomes. They are under-represented in clinical trials, particularly those with common age-related conditions such as frailty, cognitive impairment, and multimorbidity. 2 Geriatric assessment (GA), which measures multiple dimensions of older adults’ health, can improve prediction of prognosis and cancer-related outcomes. 3 Utilized within an integrated intervention strategy, GA can improve communication and cancer treatment outcomes. 4 , 5 GA is recommended for all patients ≥65 years starting cancer treatment. 6 Yet, its uptake remains limited in oncologic practice. 7 In the article that accompanies this editorial, Etienne Audureau and co-authors highlight the importance of GA variables in predicting prognosis in two large observational French cohorts of older adults with cancer. Furthermore, they recognize that capturing the complex potential interactions of clinical variables is challenging for common statistical modeling methods such as regression and Cox proportional hazards analysis. For their analysis, they utilize decision-tree based machine learning methods alongside a Cox proportional hazards model to understand which factors contribute to prognosis within their cohorts, and to compare model predictive performance. This analysis not only contributes an externally validated clinical decision tool which could improve decision-making in older adults with cancer, but also provides an early roadmap for how to re-think analytic strategies in the era of rapidly emerging technologies such as machine learning, artificial intelligence, and multimodal data (i.e., datasets comprised of multiple data types such as text, numbers, and images). Although structured clinical trial data for older adults are scarce, other data sources are proliferating exponentially. According to some estimates, around a third of all data produced in the world is healthcare data. 8 Leveraging real-world data has been proposed as a strategy to close the knowledge gap created by the low inclusion of older adults in clinical trials. 2 However, in the United States these data are often trapped in the “walled gardens” of non-interoperable systems such as electronic health record systems, lack a common data model, and are “noisy” and error-prone. Large prospective observational cohorts, such as those analyzed by Audureau et al, are a valuable intermediary between clinical trials and potentially problematic real-world data, providing structured datasets reviewed for data validity and integrity. These datasets are often considered less valuable in the hierarchy of evidence compared to the gold standard of hypothesis-driven, randomized controlled trials. However, they offer a bridge to a newer paradigm of data-driven scientific analysis based on machine learning and artificial intelligence (AI) approaches. Although hypothesis-driven science appropriately remains the foundation of medical research, it does not fully prepare us to handle and make use of the massive influx of scientific data that characterizes our current era. It also does not allow us to understand and deploy some of the newest analytic methods, those centered on large AI models such as OpenAI’s Generative Pre-trained Transformer (GPT) models. These models can ingest massive amounts of multimodal data, feed them to inscrutable deep-learning algorithms, and generate outputs that are potentially biased, flawed, and opaque. 9 The emergence of data science has occurred alongside these developments, with its focus on how data are managed, interpreted, and shared, independent from their use as evidence for a specific hypothesis. Data science prioritizes a data-centric approach whereby data are not simply generated by the researcher to answer a specific question, but are independent objects which are shareable across researchers and contexts. 10 The recent emphasis on open data sharing and findable, accessible, interoperable, and reproducible (FAIR) data highlights this paradigm shift. 11 Audureau et al illustrate some of the key principles that guide a data-centric approach to answering scientific questions such as how to predict outcomes in older adults with cancer. First and foremost, data-centric approaches complement rather than supplant hypothesis-driven approaches. In this analysis, the authors still incorporate a priori judgments about which variables to include, and they use traditional statistical methods which are well-understood and highly interpretable (unlike “black box” AI methods such as GPT). However, the examination and comparison of multiple modeling approaches and the conceptualization of variable selection as a separate analytic step for comparison across models should be replicated in more studies, particularly for populations with limited prospective trial data. Most published clinical science relies on single-model analytic methods such as regression or proportional hazards. Particularly when data are noisy, limited, or error-prone (such as in real-world data), it may be better to employ what complexity theorist Scott Page calls “many-model thinking,” using different models with different relative strengths and weaknesses to better understand a complex problem. 12 A single model’s assumptions are likely to be flawed in some way, but comparison of diverse models offers more robust conclusions. In fact, the authors use “many-model” thinking in two ways: they incorporate an ensemble model, a random forest, which simulates and aggregates multiple decision tree models; and they compare the outputs of two types of models with different assumptions and relative strengths (Cox proportional hazards and decision trees/forests). In the comparison between the single decision tree and the ensemble random forest, it is unsurprising that the random forest had better performance, given that many-model thinking can mitigate the bias in a single model. 13 The use of multiple models allows for not just performance comparisons (i.e., their accuracy at predicting prognosis) but also the relative importance of the variables contributing to these predictions. Although variable selection (or, in machine learning parlance, “feature selection”) is an inherent part of model-building in hypothesis-driven science, it is not always seen as a separate step worthy of many-model comparison. In a data science paradigm, feature selection is a separate step, as is exploratory data analysis, a data-driven analysis without specific reference to a governing hypothesis or assumptions ( Figure 1 ). In a traditional statistical analysis, variable selection methods (such as stepwise methods or dimension reduction) are often incorporated into the analytic workflow, but typically further work is not done to ensure that the variables selected or created by the model are robust and truly the most meaningful variables to include. In Audureau et al’s analysis, we see the rudiments of a true “feature selection” step, comparing the variables selected as important by each of the comparator models. For example, the G8 score was retained by each of the models, substantiating its likely importance as a prognostic indicator. Figure 1. Open in a new tab Data science lifecycle However, the authors’ analysis seems to be limited to variables selected based on a priori judgments of significance. This is inherently hypothesis-driven, assuming that we know what variables are important to answer a particular question (such as what drives prognosis in older adults with cancer). In data-driven and data-centric variable selection approaches, we have the opportunity to explore beyond what is “known” to be true. Some machine learning and AI models can handle high-dimensional data (i.e., a high number of variables) with complex, non-linear interactions, and methods exist to understand variable importance even in “black box” models. 14 Combining many-model thinking with high-dimensional data permits another way to understand what variables may be important for answering a scientific question, and may open new avenues for hypothesis-driven research. Applying machine learning feature selection models followed by hypothesis testing and causal inference methods has been proposed as a powerful means of utilizing and understanding real-world observational data. 15 Again, machine learning methods are complementary, not antagonistic, to other scientific analytic methods within a many-model thinking framework. Clinically, the results have further solidified the value of GA variables for identifying older adults at highest risk of mortality. As demonstrated in other models, variables from GA consistently are related to mortality in older patients with cancer (i.e., impaired function, poor mobility, low G8 score, weight loss, polypharmacy, and others). This study adds value to the literature given the robustness of the datasets, which includes thousands of older adults with varying levels of functional status, mobility, and nutritional status; the sample is older and more frail than older adults included in clinical trials. These data can help guide decision making for older adults considering curative intent therapies by identifying those patients who may not live long enough to benefit. Further, individuals with a high risk of mortality may benefit more from supportive care approaches than palliative cancer treatment. To facilitate use in the clinical setting, Audureau et al went beyond typical results reporting and generated interactive visualizations via a web application (app). Data visualization is prioritized within the data science lifecycle, and the authors demonstrate one way research impact can be maximized through visualization. Clinicians and patients can interact with their model directly via an intuitive graphical user interface (GUI) that democratizes access to research products and insight. Apps such as these lower barriers to knowledge access, including cognitive barriers created by dense academic jargon and time/resource limitations in locating relevant research. More research teams should consider data visualization as a critical research product to enable broad access to insight. As the authors discuss, more data on how integrating the tool in the cancer care delivery influences other outcomes such as decision making, hospitalizations, and prevention of functional decline would be helpful for clinicians and health care systems. Further research could evaluate how this app could be used at the bedside to assist in prognostic communication and decision-making, and how it may encourage more uptake of guidelines related to GA in older adults with cancer. These tools likely need to be refined for different health care systems or settings. Outcomes such as mortality are influenced by availability of treatments and health care system factors; machine learning methods may be used to help refine predictive models for local settings, using data from that specific health care system. Moreover, these “local” models could be updated over time with new data, creating a learning system to account for evolving patient characteristics, expansion of new treatments, and clinician and system related factors. This manuscript is a first step in this direction. A famous aphorism attributed to statistician George Box states, “All models are wrong, but some are useful.” 16 The authors of the accompanying article do not purport to fully understand a situation as complex as outcomes in heterogeneous older adults. However, their models are useful. They used large cohort datasets for both model training and validation. They utilized and compared multiple models to understand the robustness of the conclusions. They generated interactive and accessible visualizations to help their readers understand research insights. Beyond the specific problem (predicting prognosis in older adults with cancer) and results, they provide a second type of useful model: an illustration of how data science, machine learning, and many-model thinking can augment clinical research amidst a shifting data paradigm. We may never have enough older adults enrolling in gold-standard randomized clinical trials, even if we relax inclusion criteria or address other barriers. However, if we can learn to better leverage new approaches and the proliferating real-world observational data about these patients, we can create new synergies to address knowledge gaps. Funding: The work was funded through NCI K08CA248721 (E.R.), NIA R03AG067977 (E.R.), U01CA233167 (PI Supriya Mohile). References: 1. Smith BD, Smith GL, Hurria A, et al. : Future of cancer incidence in the United States: burdens upon an aging, changing nation. J Clin Oncol 27:2758–65, 2009 [ DOI ] [ PubMed ] [ Google Scholar ] 2. Sedrak MS, Freedman RA, Cohen HJ, et al. : Older adult participation in cancer clinical trials: A systematic review of barriers and interventions. CA Cancer J Clin 71:78–92, 2021 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Magnuson A, Loh KP, Stauffer F, et al. : Geriatric assessment for the practicing clinician: The why, what, and how. CA: A Cancer Journal for Clinicians 74:496–518, 2024 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Mohile SG, Epstein RM, Hurria A, et al. : Communication With Older Patients With Cancer Using Geriatric Assessment: A Cluster-Randomized Clinical Trial From the National Cancer Institute Community Oncology Research Program. JAMA Oncol:1–9, 2019 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Mohile SG, Mohamed MR, Xu H, et al. : Evaluation of geriatric assessment and management on the toxic effects of cancer treatment (GAP70+): a cluster-randomised study. Lancet 398:1894–1904, 2021 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Mohile SG, Dale W, Somerfield MR, et al. : Practical Assessment and Management of Vulnerabilities in Older Patients Receiving Chemotherapy: ASCO Guideline for Geriatric Oncology. J Clin Oncol 36:2326–2347, 2018 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Dale W, Williams GR, A RM, et al. : How Is Geriatric Assessment Used in Clinical Practice for Older Adults With Cancer? A Survey of Cancer Providers by the American Society of Clinical Oncology. JCO Oncol Pract 17:336–344, 2021 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Reinsel D, Gantz J, Rydning J: The Digitization of the World From Edge to Core. IDC White Paper – #US44413318, Sponsored by Seagate, 2018 [ Google Scholar ] 9. Omiye JA, Gui H, Rezaei SJ, et al. : Large language models in medicine: the potentials and pitfalls: a narrative review. Annals of Internal Medicine 177:210–220, 2024 [ DOI ] [ PubMed ] [ Google Scholar ] 10. Leonelli S: Data Governance is Key to Interpretation: Reconceptualizing Data in Data Science. Harvard Data Science Review 1, 2019 [ Google Scholar ] 11. Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. : The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 3:160018, 2016 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Page SE: The Model Thinker: What You Need to Know to Make Data Work for You, Basic Books, Inc., 2018 [ Google Scholar ] 13. Dietterich TG: Ensemble Methods in Machine Learning. Multiple Classifier Systems. Berlin, Heidelberg, Springer Berlin Heidelberg, 2000, pp 1–15 [ Google Scholar ] 14. Ramsdale E, Kunduru M, Smith L, et al. : Supervised learning applied to classifying fallers versus non-fallers among older adults with cancer. J Geriatr Oncol 14:101498, 2023 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Crown WH: Real-World Evidence, Causal Inference, and Machine Learning. Value in Health 22:587–592, 2019 [ DOI ] [ PubMed ] [ Google Scholar ] 16. Box GEP: Science and Statistics. Journal of the American Statistical Association 71:791–799, 1976 [ Google Scholar ] ACTIONS View on publisher site PDF (311.9 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 67432 · SHA-256 2e78f18a0d00ea09
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.