Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice letter Brief Bioinform . 2026 Apr 8;27(2):bbag168. doi: 10.1093/bib/bbag168 Search in PMC Search in PubMed View in NLM Catalog Add to search Addressing biases and limitations in feature attribution for circRNA modification profiling Souichi Oka Souichi Oka 1 Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan Find articles by Souichi Oka 1, ✉ , Kota Takemura Kota Takemura 2 Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan Find articles by Kota Takemura 2 , Yoshiyasu Takefuji Yoshiyasu Takefuji 3 Department of Data Science, Faculty of Data Science, Musashino University, 3-3-3 Ariake Koto-ku, Tokyo 135-8181, Japan Find articles by Yoshiyasu Takefuji 3 Author information Article notes Copyright and License information 1 Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan 2 Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan 3 Department of Data Science, Faculty of Data Science, Musashino University, 3-3-3 Ariake Koto-ku, Tokyo 135-8181, Japan ✉ Corresponding author. Science Park Corporation, 3-24-9 Iriya-Nishi Zama-shi, Kanagawa 252-0029, Japan. E-mail: [email protected] Received 2026 Feb 27; Accepted 2026 Mar 17; Collection date 2026 Mar. © The Author(s) 2026. Published by Oxford University Press. This is an Open Access article distributed under the terms of the Creative Commons Attribution License ( https://creativecommons.org/licenses/by/4.0/ ), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited. PMC Copyright notice PMCID: PMC13069895 PMID: 41955026 Abstract Li et al . (CircRM: Profiling circular RNA modifications from nanopore direct RNA sequencing. Brief Bioinform 2026; 27 :bbaf726.) introduced Circular RNA Modifications (CircRM), a computational framework employing eXtreme Gradient Boosting and SHapley Additive exPlanations (SHAP) to profile RNA modifications in circular RNAs, achieving high predictive accuracy. However, we argue that strong predictive performance does not validate the biological reliability of the resulting feature-importance rankings. In heterogeneous feature spaces, tree-based models exhibit inherent biases, favoring continuous, high-cardinality variables—such as genomic position—over sparse sequence patterns, potentially obscuring true biological determinants. Furthermore, reliance on SHAP introduces theoretical vulnerabilities; recent findings on attribution limitations indicate that baseline sensitivity can decouple explanations from local mechanistic behavior. To address these analytical pitfalls, we advocate for a robust framework incorporating Highly Variable Gene Selection and Feature Agglomeration to mitigate multicollinearity, complemented by model-agnostic non-parametric methods such as Spearman’s rho and Kendall’s tau. Adopting these strategies ensures that computational profiling yields biologically actionable insights rather than reflecting statistical artifacts. Keywords: circularRNA, epitranscriptomics, machine learning, feature importance, model bias Dear Editors, Li et al . introduced CircRM, a computational framework that combines eXtreme Gradient Boosting (XGBoost) with SHapley Additive exPlanations (SHAP) to profile RNA modifications in circular RNAs (circRNAs) from raw nanopore direct RNA sequencing data [ 1 ]. Using this framework, they reported 427 high-confidence circRNAs—substantially exceeding commonly used baseline approaches—and developed modification-detection models with Receiver Operating Characteristic (ROC) Area Under the Curve (AUC) values of 0.855 (m5C), 0.817 (m6A), and 0.769 (m1A). A SHAP-based analysis further highlighted several highly ranked predictors, such as relative position within the annotated 5' Untranslated Region (UTR) and distance to nearby DRACH motifs, which were used to motivate reported differences between linear-RNA and circRNA modification patterns. Collectively, these results suggest that CircRM is a promising tool for characterizing the circRNA epitranscriptome at single-molecule resolution; however, the validity and biological interpretability of the resulting feature-importance rankings warrant further discussion. This concern follows well-recognized distortions in end-to-end analytical pipelines that can decouple predictive performance from mechanistic biological interpretation. Reliance on predictive accuracy as a proxy for biological feature relevance has been repeatedly cautioned against, particularly in high-dimensional transcriptomic settings where correlated predictors enable models to exploit statistical regularities that need not reflect causal determinants [ 2 , 3 ]. The risk is amplified when training labels are defined through probabilistic score thresholds, in which case the model may preferentially learn threshold-induced proxies rather than bona fide mechanistic signals. As summarized in our Supplementary Material , a survey of more than 300 peer-reviewed studies reports systematic skew in feature-importance estimates under closely related conditions. Accordingly, strong AUC values alone provide limited assurance that the ranking of predictors reflects distinct biological mechanisms. Tree-based learning algorithms, including XGBoost, also exhibit method-specific biases in feature-importance estimation [ 4–6 ]. In particular, split-based learning can favor variables that facilitate early partitioning, an effect that can be exacerbated by multicollinearity and complex dependency structure in transcriptomic feature spaces [ 7 ]. In Li et al .’s specific modeling setting, several highly ranked predictors—such as relative positions within annotated UTRs/Coding Sequence (CDS) and distances to nearby motifs—may be susceptible to this issue, thereby complicating their mechanistic interpretation. Moreover, CircRM employs a heterogeneous feature space that mixes dense continuous nanopore-signal statistics (e.g. mean, dwell time) with sparse one-hot encoded sequence features. In such mixed feature spaces, importance can be preferentially attributed to continuous, high-cardinality variables, potentially obscuring subtle but biologically decisive sequence patterns. Given the absence of definitive ground truth for circRNA-specific feature relevance in RNA modifications, and the difficulty of disentangling correlated sequence effects, the reported rankings should be interpreted with appropriate caution. In addition to correlation-driven concerns, SHAP-based attribution can be vulnerable to theoretical limitations that constrain the reliability of post hoc explanations in certain settings [ 8 , 9 ]. Recent results sometimes referred to as ‘impossibility’ theorems indicate that, under particular desiderata for additive attributions, local explanatory faithfulness can be fundamentally limited, depending on modeling assumptions and the choice of baseline or reference distribution [ 10 ]. In practice, the completeness constraint can induce attributions that reflect global properties of the fitted model and background distribution rather than strictly local behavior around a given observation. Consequently, SHAP values may be sensitive to baseline specification, and their magnitude or sign need not map straightforwardly onto mechanistic influence in the underlying biological system. In the CircRM application context, features identified as key predictors (e.g. distance to DRACH motifs) could therefore partially reflect baseline- or distribution-sensitive effects rather than direct modification determinants. This possibility is particularly salient where explanations are used to motivate read-level mechanistic narratives, and it motivates explicit validation beyond attribution scores alone. To address these analytical pitfalls, a more robust inferential framework is required—one that does not hinge on model-dependent attributions and that explicitly addresses the structural complexity and redundancy of sequencing-derived features. We therefore propose combining Highly Variable Gene Selection with Feature Agglomeration to mitigate multicollinearity and feature redundancy [ 11 , 12 ]. In particular, for high-dimensional sequence-derived inputs such as the one-hot encoded 5-mer representations used in CircRM, these procedures can consolidate correlated patterns and reduce the arbitrariness with which importance is distributed across redundant nucleotide indicators. In addition, complementing these steps with non-parametric association measures—specifically Spearman’s rho and Kendall’s tau—enables the assessment of monotonic and rank-based relationships without assuming linearity, thereby providing a more interpretable, model-agnostic perspective on feature–signal associations [ 13 , 14 ]. In summary, while CircRM exhibits strong predictive performance, the biological interpretation of its feature-importance rankings remains provisional. The combination of algorithm-specific biases in tree-based models and the context-dependent limitations of SHAP indicates that highly ranked predictors may, in part, reflect statistical structure rather than causal mechanisms. Predictive accuracy should therefore not be conflated with explanatory adequacy. Progress in circRNA epitranscriptomic profiling will benefit from validation strategies that emphasize robustness, stability, and model-agnostic corroboration to ensure that computational inferences translate into biologically actionable insight. Key Points Although CircRM demonstrates certain levels of predictive performance, such metrics may not necessarily ensure the reliability of the resulting feature-importance rankings for RNA modifications. Tree-based models like XGBoost are prone to specific biases in high-dimensional settings, which can prioritize high-cardinality features over biologically decisive sequence patterns. The application of additive attribution methods, such as SHAP, warrants careful consideration as rankings can be sensitive to baseline specifications rather than directly reflecting mechanistic biological influence. To mitigate these issues, we advocate for a robust analytical framework primarily based on Highly Variable Gene Selection and Feature Agglomeration, complemented by non-parametric statistical methods to provide a model-agnostic perspective. Supplementary Material Supplementary_material_bbag168 supplementary_material_bbag168.docx (74.3KB, docx) Contributor Information Souichi Oka, Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan. Kota Takemura, Research and Development Planning Department, Science Park Corporation, 3-24-9 Iriya-Nishi, Zama-shi, Kanagawa 252-0029, Japan. Yoshiyasu Takefuji, Department of Data Science, Faculty of Data Science, Musashino University, 3-3-3 Ariake Koto-ku, Tokyo 135-8181, Japan. Conflicts of interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this article. Author contributions Souichi Oka (Conceptualization, Writing—original draft), Kota Takemura (Investigation), and Yoshiyasu Takefuji (Project administration, Supervision, Writing—review & editing) Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Data availability No new data were generated or analyzed in support of this research. References 1. Li J, Chen S, Wu Z et al. CircRM: profiling circular RNA modifications from nanopore direct RNA sequencing. Brief Bioinform 2026;27:bbaf726. 10.1093/bib/bbaf726 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Fisher A, Rudin C, Dominici F. All models are wrong, but many are useful: learning a variable’s importance by studying an entire class of prediction models simultaneously. J Mach Learn Res 2019;20:177. 10.48550/arXiv.1801.01489 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Lipton ZC. The mythos of model interpretability: in machine learning, the concept of interpretability is both important and slippery. Queue 2018;16:31–57. 10.1145/3236386.3241340 [ DOI ] [ Google Scholar ] 4. Adler AI, Painsky A. Feature importance in gradient boosting trees with cross-validation feature selection. Entropy 2022;24:687. 10.3390/e24050687 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Lenhof K, Eckhart L, Rolli LM et al. Trust me if you can: a survey on reliability and interpretability of machine learning approaches for drug sensitivity prediction in cancer. Brief Bioinform 2024;25:bbae379. 10.1093/bib/bbae379 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Ugirumurera J, Bensen EA, Severino J et al. Addressing bias in bagging and boosting regression models. Sci Rep 2024;14:18452. 10.1038/s41598-024-68907-5 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Alaimo Di Loro P, Scacciatelli D, Tagliaferri G. 2-step gradient boosting approach to selectivity bias correction in tax audit: an application to the VAT gap in Italy. Stat Methods Appl 2023;32:237–70. 10.1007/s10260-022-00643-4 [ DOI ] [ Google Scholar ] 8. Hooshyar D, Yang Y. Problems with SHAP and LIME in interpretable AI for education: a comparative study of post-hoc explanations and neural-symbolic rule extraction. IEEE Access 2024;12:137472–90. 10.1109/ACCESS.2024.3463948 [ DOI ] [ Google Scholar ] 9. Huang X, Marques-Silva J. On the failings of Shapley values for explainability. Int J Approx Reason 2024;171:109112. 10.1016/j.ijar.2023.109112 [ DOI ] [ Google Scholar ] 10. Bilodeau B, Jaques N, Koh PW et al. Impossibility theorems for feature attribution. Proc Natl Acad Sci U S A 2024;121:e2304406120. 10.1073/pnas.2304406120 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Xie Y, Jing Z, Pan H et al. Redefining the high variable genes by optimized LOESS regression with positive ratio. BMC Bioinformatics 2025;26:104. 10.1186/s12859-025-06112-5 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Zhang J, Wu X, Hoi SCH et al. Feature agglomeration networks for single stage face detection. Neurocomputing 2020;380:180–9. 10.1016/j.neucom.2019.10.087 [ DOI ] [ Google Scholar ] 13. Okoye K, Hosseini S (eds). Correlation tests in R: Pearson Cor, Kendall’s tau, and Spearman’s rho. In: Okoye K, Hosseini S (eds) R Programming: Statistical Data Analysis in Research. Cham: Springer Nature, 2024, 247–77. 10.1007/978-3-031-45753-8_10 [ DOI ] [ Google Scholar ] 14. Yu H, Hutson AD. A robust spearman correlation coefficient permutation test. Commun Stat Theory Methods 2024;53:2141–53. 10.1080/03610926.2022.2121144 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary_material_bbag168 supplementary_material_bbag168.docx (74.3KB, docx) Data Availability Statement No new data were generated or analyzed in support of this research. Articles from Briefings in Bioinformatics are provided here courtesy of Oxford University Press ACTIONS View on publisher site PDF (287.8 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top