ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Advanced Bayesian kernel machine regression for large-scale exposome studies: Making the impossible possible.

Guo Y et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
computerscienceeducation
computer science education

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Innovation (Camb) . 2026 Jan 3;7(4):101248. doi: 10.1016/j.xinn.2025.101248 Search in PMC Search in PubMed View in NLM Catalog Add to search Advanced Bayesian kernel machine regression for large-scale exposome studies: Making the impossible possible Yi Guo Yi Guo 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China Find articles by Yi Guo 1, 6 , Huixun Jia Huixun Jia 2 Department of Ophthalmology, Shanghai General Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai 200080, China Find articles by Huixun Jia 2, 6 , Ziwei Peng Ziwei Peng 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China Find articles by Ziwei Peng 1 , Xinming Xu Xinming Xu 3 Department of Nutrition and Food Hygiene, Ministry of Education Key Laboratory of Public Health Safety, School of Public Health, Institute of Nutrition, Fudan University, Shanghai 200030, China Find articles by Xinming Xu 3 , Zhicheng Zhang Zhicheng Zhang 3 Department of Nutrition and Food Hygiene, Ministry of Education Key Laboratory of Public Health Safety, School of Public Health, Institute of Nutrition, Fudan University, Shanghai 200030, China Find articles by Zhicheng Zhang 3 , Keyu Pan Keyu Pan 4 Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University, Jinan 250012, China Find articles by Keyu Pan 4 , Yuqin Zhou Yuqin Zhou 5 Huangpu District Center for Disease Prevention and Control (Huangpu District Health Supervision Institute), Shanghai 200001, China Find articles by Yuqin Zhou 5 , Haidong Kan Haidong Kan 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China Find articles by Haidong Kan 1 , Zhenyu Wu Zhenyu Wu 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China Find articles by Zhenyu Wu 1, ∗∗ , Cong Liu Cong Liu 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China Find articles by Cong Liu 1, ∗ Author information Article notes Copyright and License information 1 School of Public Health, Key Lab of Public Health Safety of the Ministry of Education and NHC Key Laboratory of Health Technology Assessment, Fudan University, Shanghai 200032, China 2 Department of Ophthalmology, Shanghai General Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai 200080, China 3 Department of Nutrition and Food Hygiene, Ministry of Education Key Laboratory of Public Health Safety, School of Public Health, Institute of Nutrition, Fudan University, Shanghai 200030, China 4 Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University, Jinan 250012, China 5 Huangpu District Center for Disease Prevention and Control (Huangpu District Health Supervision Institute), Shanghai 200001, China ∗ Corresponding author [email protected] ∗∗ Corresponding author [email protected] 6 These authors contributed equally Received 2025 May 21; Accepted 2025 Dec 31; Collection date 2026 Apr 6. © 2026 The Author(s) This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). PMC Copyright notice PMCID: PMC13069415  PMID: 41970720 Abstract Exposome studies involve analyzing numerous exposures with complex interactions and potential collinearity, presenting challenges for conventional statistical methods. While Bayesian kernel machine regression (BKMR) has emerged as a promising solution, its widespread adoption has been hindered by high computational costs and restricted interpretability. To address these critical limitations in large-scale exposome studies, we developed an advanced BKMR (A-BKMR) model. The Gaussian predictive process and matrix decomposition were used to reduce both processing time and memory requirements. Additionally, we employed the parametric g-formula to generate interpretable statistics, including joint and univariate effects as well as bivariate and multivariate interactions. Across various scenarios with different sample sizes and numbers of exposures, A-BKMR demonstrated both high computational efficiency and model performance. Previously, analyzing datasets with sample sizes of 100,000 was unfeasible for traditional BKMR. The current A-BKMR can complete such analyses in 1 h on a personal computer, making it over 700,000 times faster than conventional BKMR implementations. Additionally, A-BKMR can accurately identify important exposure while preserving an area under the curve (AUC) > 0.99 and an R 2 > 0.97 across scenarios with varying sample sizes and numbers of exposures. Furthermore, A-BKMR introduces novel quantitative metrics for effect estimates and interaction analyses, substantially enhancing interpretability. These advancements establish A-BKMR as an excellent statistical framework for future large-scale exposome studies. Keywords: exposome, Bayesian kernel machine regression, computational efficiency, quantitative estimate, interaction Graphical abstract Open in a new tab Public summary • An advanced Bayesian kernel machine regression (A-BKMR) was developed for large-scale exposome studies. • A-BKMR can process datasets of 100,000 samples in an hour, over 700,000-fold faster than traditional BKMR. • A-BKMR maintains high accuracy for variable selection and effect estimation. • A-BKMR provides quantitative metrics for effects and interactions, boosting interpretability. Introduction The exposome encompasses all non-genetic exposures affecting human health, 1 including lifestyles, metabolic profiles, and environmental factors. 2 , 3 Despite its pivotal role, the effective analytical framework for exposome data remains elusive. 4 , 5 Conventional analytical paradigms predominantly employ either single-exposure models or aggregated exposure effect approaches. The former isolates individual exposures, neglecting their complex interactions, 6 , 7 while the latter risks overestimating due to overlapping correlations. Given the high collinearity among exposure categories and their intricate interactions, traditional statistical methods are not recommended for exposome data analyses. 8 , 9 , 10 To address these challenges, two broad statistical categories have been proposed for multiple exposure analysis, i.e., interpretable statistical methods and machine learning techniques. Interpretable statistical methods, such as shrinkage regression (e.g., lasso, ridge, and elastic net), 11 , 12 weighted quantile sum regression (WQS), 13 and quantile-based g-computation (QGC), 14 , 15 generate quantitative estimates while preserving interpretability. However, these methods rely on explicit regression formulas, which may be unverifiable and vulnerable to violations of fundamental model assumptions. In contrast, machine learning methodologies offer enhanced flexibility by circumventing the need for predefined model specifications. 16 , 17 , 18 , 19 However, they are often criticized for their “black-box” nature, limiting interpretability and transparency. 20 Beyond the above models, Bayesian kernel machine regression (BKMR) has emerged as a promising methodological innovation. 21 , 22 By integrating the flexibility of machine learning with enhanced interpretability, BKMR represents a balanced analytical approach for exposome studies. As a Gaussian process regression method, BKMR circumvents the need to explicitly specify the exposure-outcome relationship. Consequently, it can identify complex associations and model flexible dose-response relationships, including nonlinear, non-monotonic, and threshold-like patterns. Additionally, it enhances interpretability through posterior inclusion probabilities ( PIP s) and visualized association plots. Recognized and advocated by the Health Effects Institute (HEI) 23 , 24 and the Exposome Project for Health and Occupational Research (EPHOR), 25 BKMR has been extensively implemented in exposome studies for risk factor identification and health impact assessment. Despite its advantages, BKMR has inherent limitations that restrict its broader application. The method’s substantial computational requirements, characterized by slow processing speeds and large memory requirements, 26 , 27 render it impractical for large-scale studies. This limitation has become increasingly salient as sample sizes of exposome studies progressively increase. Consequently, the majority of BKMR-based studies are restricted to sample sizes below 2,000 participants ( Figure S1 ). 28 , 29 , 30 , 31 , 32 In studies exceeding 10,000 participants, BKMR remains highly relevant but computationally prohibitive. 33 Researchers have consequently adopted alternative strategies, such as estimating site-specific effects followed by meta-analysis 27 or applying BKMR to a randomly selected subset (e.g., 10% 34 , 35 or even 1% 36 , 37 , 38 of the original dataset). Furthermore, BKMR’s reliance on graphical representations rather than quantitative statistical metrics limits its capacity for precise interpretation. 39 , 40 , 41 To address these gaps, we propose an advanced version of BKMR (advanced BKMR [A-BKMR]; Figure 1 ), specifically designed to (1) enhance computational efficiency to facilitate large-scale exposome studies with common personal computers and (2) incorporate quantitative statistical metrics to enable precise effect estimation. This methodological advancement represents a significant progression in exposome studies, transforming previously impossible analytical approaches into possible ones. Figure 1. Open in a new tab Workflow of A-BKMR and its Python-based acceleration process (A) The flowchart of the A-BKMR framework. (B) The acceleration process implemented in Python for A-BKMR. n and s represent the number of observations in the original and sampled knot datasets, respectively. Materials and methods This section provides a comprehensive overview of the methodological framework. We first briefly review the original BKMR method and its partially accelerated extension, partially accelerated BKMR (PA-BKMR). Then, we introduce our newly proposed method, A-BKMR, followed by a detailed description of the simulation studies. BKMR and PA-BKMR BKMR can be expressed as y = h X + C β + ε , where y denotes the outcome, h ( X ) the flexible kernel function, X the exposures, C the covariates, β the regression coefficients of covariates, and ϵ the residuals. For the dataset comprising n observations, BKMR requires iterative computations involving an n × n matrix V −1 ( Figure 2 A). Consequently, for large-scale datasets (e.g., n > 10,000), each iteration requires hours to days of computational time. Given that BKMR typically necessitates over 10,000 iterations, the cumulative computational burden becomes prohibitive. To mitigate this issue, Jennifer et al. developed PA-BKMR, which employs a Gaussian predictive process model to reduce computational complexity ( Figure 2 B). 22 Briefly, PA-BKMR uses a subset of s observations ( s ≪ n ) sampled from the full dataset, referred to as the “knot dataset,” to approximate V −1 as V − 1 ˜ 42 : V − 1 ˜ = I n − I n K T ( X , X ∗ ) [ K ( X ∗ ) + K ( X , X ∗ ) I n K T ( X , X ∗ ) ] − 1 K ( X , X ∗ ) I n , (Equation 1) where I n is the unit matrix; X and X ∗ represent the full and knot datasets, respectively; and K ( X , X ∗ ) denotes their pairwise distance matrix. Although obtaining of V − 1 ˜ is more efficient, the matrix operations involving V − 1 ˜ remain highly time-consuming for large n , as V − 1 ˜ retains an n × n dimensionality. PA-BKMR needs to repeatedly perform matrix operations involving V − 1 ˜ in each iteration, resulting in limited acceleration for large datasets. Further technical details regarding BKMR and PA-BKMR are provided in Texts S1–S6 . Figure 2. Open in a new tab The estimation process of BKMR, PA-BKMR, and A-BKMR (A) BKMR. (B) PA-BKMR. (C) A-BKMR. PIPs refers to the posterior inclusion probabilities of all exposure variables. n , m , and s represent the number of observations and exposures in the original dataset and the number of observations in the sampled knot dataset, respectively. Plot 1 , Plot 2 , etc., denote the output plots for exposure-response curves, interaction effects, and so on. Sta 1 , Sta 2 , etc., denote the quantitative statistics. A-BKMR We propose the A-BKMR to address the limitations of the previous BKMR methods. As illustrated in Figure 1 A, A-BKMR introduces two major advancements: (1) computational acceleration and (2) the provision of quantitative statistical metrics. For computational acceleration, a weighted knot sampling approach is employed to generate a representative subset (knot dataset) from the full dataset. This subset comprises s observations selected from the full dataset of n observations ( s ≪ n ). The large n × n matrix is then decomposed into three small matrices ( n × s , s × s , and s × n ), which are utilized for estimation in subsequent iterations. Upon completion of all iterations, PIPs of all exposure variables are estimated, and exposure-response curves are generated. Finally, a parametric g-formula method is applied to derive quantitative statistical estimates from the exposure-response curves. Computational acceleration As aforementioned, PA-BKMR employs the Gaussian predictive process for acceleration, starting with knot sampling. However, PA-BKMR’s knot sampling method cannot handle datasets containing duplicate observations, 22 , 43 such as when multiple participants share identical exposure values, which are common in exposome studies, such as those examining air pollution, where people living nearby may be exposed to the same pollutant concentrations. To address this limitation, we developed a weighted knot sampling method. This method assigns weights to observations based on their frequency in the dataset: a weight of 1 for unique observations, 2 for duplicates, and so on. 44 , 45 During the subsequent sampling process, observations with higher weights are more likely to be selected, ensuring better representation of repeated exposures. For detailed information and implementation, please refer to Text S8.1 . Following the generation of the knot dataset via weighted knot sampling, A-BKMR utilizes a Gaussian predictive process to simplify computations, similar to PA-BKMR. However, unlike PA-BKMR, A-BKMR avoids direct storage and manipulation of the large matrix V − 1 ˜ to enhance computational efficiency. Specifically, according to Equation 1 , let A denote K ( X , X ∗ ) and B denote [ K ( X ∗ ) + K ( X , X ∗ ) I n K T ( X , X ∗ ) ] − 1 . Then, V − 1 ˜ can be expressed as V − 1 ˜ = I n − A T B A . (Equation 2) Thus, V − 1 ˜ can be decomposed into smaller matrices A T , B , and A , with dimensions n × s , s × s , and s × n , respectively. Since V − 1 ˜ serves as an intermediate statistic ( Figure 2 C), its explicit computation is unnecessary. Instead, computations involving V ˜ y − 1 can be simplified using the associative law of matrix multiplication. For example, C T V − 1 ˜ C can be expressed as C T ( I n − A T B A ) C = C T C − C T A T B A C = C T C − ( C T A T ) × B × ( A C ) , where C T C , C T A T , B , and AC are matrices with dimensions of c × c , c × s , s × s , and s × c , respectively. These small matrices enable rapid computation of C T V − 1 ˜ C . Similar simplifications are applied to other statistical computations (see Text S8.2 ). Despite avoiding direct storage and manipulation of large matrices, the iterations remained computationally intensive due to the numerous iterations. To address this, we leveraged Python for accelerated matrix inversion and multiplication tasks ( Figure 1 B; Text S8.3 ). 46 Provision of quantitative statistics In addition to computational acceleration, we enhanced BKMR to generate quantitative statistics, including estimates of effects (joint and univariate effects) and interactions (bivariate and multivariate interactions). The joint effect quantifies the average change in outcome when all exposures increase by the same scale (e.g., a 10% increase). The univariate effect measures the average change in the outcome when a single exposure increases, while all other exposures are fixed at a special value (e.g., their 50 th percentile). The bivariate interaction effect examines the interaction between two exposures X 1 and X 2 , reflecting the variability in the X 1 - y association as X 2 changes. The multivariate interaction effect examines the interaction between a given exposure X i and all other exposures, reflecting the variability in the X i - y association when all other exposures change (e.g., from their 25 th to 50 th percentiles). Further details about these quantitative statistics can be found in Text S8.4 . Simulation We conducted two simulations to evaluate the effectiveness of A-BKMR and compare it with four other BKMR-related methods: original BKMR, PA-BKMR, BKMRhat (a parallel computation extension of BKMR; Text S7 ), 47 and generalized BKMR (GBKMR; a newly developed BKMR method based on generalized linear mixed models). 48 The two simulations were designed to examine different aspects of data variability: sample size and number of exposure variables. Both continuous and binary outcome scenarios were considered, with 10,000 iterations for all analyses. We also tested the maximum number of exposures A-BKMR can handle ( Text S10 ; Figures S2 and S3 ). Simulation with varying sample sizes The first simulation focused on varying sample sizes, generating data from a multivariable normal distribution with a known covariance-variance matrix. We examined sample sizes of 50, 100, 500, 1,000, 5,000, 10,000, 50,000, and 100,000 to represent the common scales in current exposome studies. 49 The outcome distribution was set to both Gaussian and binomial, resulting in 16 scenarios (8 sample sizes × 2 outcome types). Each scenario included 5 exposure variables, X 1 – X 5 , and 3 covariates, C 1 – C 3 , where C 1 was generated based on X 1 and C 2 and C 3 were generated randomly. Details on the exposures, covariates, and outcomes can be found in Text S9.1 . Simulation with varying numbers of exposures In the second simulation, we varied the number of exposure variables, using data from 14,073 participants in the 2018 China Health and Nutrition Survey (CHNS). 50 Two sets of exposures were considered, with either a continuous outcome following a normal distribution or a binary outcome following a Bernoulli distribution in set 1 and set 2, respectively ( Text S9.2 ). For set 1, 10 trace elements (K, Na, Ca, Mg, Fe, Mn, Zn, Cu, P, and Se) were used as exposures, with 9 covariates including demographic data (age, sex, marriage, and BMI), lifestyle (smoke, alcohol, and physical activity), and socioeconomic status (education and city index). For set 2, 21 food intake variables (vegetation, exp (vegetation), legume, 1/legume, fruit, fruit 2 , log(fruit 2 ), nut, nut 3 , fish, fish 2 , meat, exp (meat), salt, grain, milk, rice, cake, cereal, tuber, and sugar) were used as exposures, using the same 9 covariates from set 1. Evaluation and comparison We compared the 4 methods (A-BKMR, BKMR, PA-BKMR, and BKMRhat) by assessing their computational efficiency and model performance. Computational efficiency was measured by processing time, including the total time for sampling (0 for BKMR), estimation, and output generation. Model performance was assessed through estimation and prediction accuracy. Estimation accuracy was determined by whether the PIPs reflected the importance of the exposures, while predictive performance was measured using R 2 for Gaussian outcomes and area under the curve (AUC) for binomial outcomes. 51 All analyses were performed on a laptop computer with 32 GB of random-access memory (RAM) and an Intel Core i9-185H CPU (2.30 GHz). BKMR, PA-BKMR, BKMRhat, and GBKMR analyses were performed using the R packages “bkmr,” “bkmr,” “bkmrhat,” and “GBKMR,” respectively. Results Computational efficiency We compared the performance of four methods: A-BKMR, PA-BKMR, original BKMR, and BKMRhat. As demonstrated in Table 1 , A-BKMR consistently outperformed the other methods in terms of computational efficiency, with the performance advantage becoming more pronounced as the sample size increased. For example, for a dataset of 5,000 observations, the original BKMR required approximately 92 days to complete the calculation, while PA-BKMR took over 10 h. In contrast, A-BKMR completed the task in just 2 min. Notably, for larger sample sizes, e.g., exceeding 5,000 for BKMR and 10,000 for PA-BKMR, the exact processing time could not be directly measured due to their excessive computational demands. Instead, these processing times were estimated by extrapolating the time per iteration over 10,000 iterations. Additionally, due to multi-chain design and system variability, the processing time of BKMRhat for larger sample sizes is difficult to determine accurately. Table 1. Processing time of different BKMR methods for Gaussian and binomial outcomes in simulations with various sample sizes n A-BKMR PA-BKMR a BKMR BKMRhat Gaussian outcome 50 0 h, 0 min, 11.86 s 0 h, 1 min, 2.73 s 0 h, 1 min, 18.8 s 0 h, 0 min, 17.48 s 100 0 h, 0 min, 22.36 s 0 h, 2 min, 32.68 s 0 h, 4 min, 30.72 s 0 h, 0 min, 15.46 s 500 0 h, 0 min, 13.37 s 0 h, 34 min, 42.98 s 1 h, 14 min, 16.04 s 0 h, 15 min, 26.95 s 1,000 0 h, 1 min, 31.48 s 1 h, 2 min, 14.7 s 8 h, 8 min, 30.46 s 1 h, 55 min, 4.97 s 5,000 0 h, 2 min, 6.44 s 10 h, 45 min, 21.42 s ∼92 days b – 10,000 0 h, 9 min, 43.43 s 36 h, 13 min, 18.61 s ∼727 days b – 50,000 0 h, 14 min, 58.18 s ∼57 days b >90 years b – 100,000 1 h, 4 min, 23.67 s ∼225 days b >90 years b – Binomial outcome 50 0 h, 0 min, 13.36 s 0 h, 2 min, 9.21 s 0 h, 3 min, 22.4 s 0 h, 0 min, 4.37 s 100 0 h, 0 min, 15.2 s 0 h, 7 min, 13.37 s 0 h, 4 min, 54.99 s 0 h, 0 min, 7.82 s 500 0 h, 1 min, 21.31 s 0 h, 42 min, 9.46 s 1 h, 13 min, 20.79 s 0 h, 5 min, 27.07 s 1,000 0 h, 2 min, 54.6 s 1 h, 24 min, 6.05 s 11 h, 16 min, 21.55 s 1 h, 19 min, 59.42 s 5,000 0 h, 9 min, 30.67 s 10 h, 29 min, 29.6 s ∼92 days b – 10,000 0 h, 15 min, 3.99 s 26 h, 44 min, 14.74 s ∼727 days b – 50,000 1 h, 3 min, 32.11 s ∼57 days b >90 years b – 100,000 2 h, 51 min, 34.59 s ∼225 days b >90 years b – Open in a new tab All analyses were performed on a laptop computer with 32 GB of random-access memory (RAM) and an Intel Core i9-185H CPU (2.30 GHz), using 10,000 iterations. BKMR, Bayesian kernel machine regression; A-BKMR, accelerated BKMR; PA-BKMR, partially accelerated BKMR; BKMRhat, BKMR estimated using parallel chains. a Some processing time was absent for PA-BKMR, BKMR, and BKMRhat. This occurred because, as the sample size increased, these methods exceeded the allowable computation time. b These processing times are approximately estimated by multiplying the time per iteration by 10,000. The computational burden became increasingly prohibitive with larger sample sizes. Specifically, with 50,000 observations, running BKMR was almost impossible, and with 100,000 observations, the estimated computation time for BKMR and PA-BKMR was over 90 years and 225 days, respectively. In contrast, A-BKMR completed the task in just 1 h. The processing time of BKMRhat was comparable to that of PA-BKMR. A visual comparison of computational efficiency is presented in Figure 3 , illustrating that the processing times of BKMR, PA-BKMR, and BKMRhat increased sharply with larger sample sizes, while the processing time of A-BKMR remained stable, consistently staying below 2 h. Detailed breakdowns of time consumption for different BKMR methods, including time of sampling, fit, and output, can be found in Text S9.3 . It is noteworthy that GBKMR requires an extensive number of iterations not only for model fitting but also for generating plots. Additionally, we found that GBKMR failed to generate plots due to its inability to handle collinearity. As a result, it cannot be directly compared with the other four methods, and all the results for GBKMR are provided separately in Tables S1–S3 . Figure 3. Open in a new tab Processing time of different BKMR methods with various sample sizes (A) Gaussian outcome with samples from 50 to 100,000. (B) Binomial outcome with samples from 50 to 100,000. (C) Gaussian outcome with samples from 50 to 1,000. (D) Binomial outcome with samples from 50 to 1,000. For the simulation based on CHNS data ( n = 14,073), only A-BKMR was used for estimation due to time constraints. In scenarios involving 10 exposures (set 1) and 21 exposures (set 2), the total processing times of A-BKMR were 46 min and 25.2 s and 2 h, 3 min, and 51.5 s, respectively. Model performance In the first simulation with varying sample sizes, outcomes were generated using only X 1 and X 2 . A-BKMR uses PIP s to measure the importance of exposures. The PIP of an exposure represents the probability that it has a meaningful effect on the outcome. For example, a PIP of 1 indicates that the exposure definitely has an effect on the outcome, while a PIP of 0 suggests that the exposure has no influence on the outcome. Ideally, the PIPs of X 1 and X 2 ( PIP 1 and PIP 2 ) should be 1, while the PIPs of the other exposures ( PIP 3 – PIP 5 ) should be 0. For Gaussian outcomes, all four methods successfully identified the important exposures, with PIP 1 and PIP 2 close to 1 and PIP 3 – PIP 5 close to 0 ( Table 2 ). Specifically, while PIP 1 and PIP 2 remained consistently close to 1 across nearly all scenarios, PIP 3 – PIP 5 were initially higher when the sample sizes were ≤100 but gradually decreased toward 0 as the sample size exceeded 100. The PIPs from A-BKMR were highly similar to those from PA-BKMR. The minor difference (0.0001) between A-BKMR and BKMR was attributed to differences between Python and R implementations and is negligible in practice. GBKMR was able to estimate PIP s for the exposures but failed to generate visual plots when applied to this dataset due to collinearity issues. Additionally, the output process of GBKMR requires numerous iterations (over 10,000), similar to its estimation process. As a result, only the model fitting time is reported in Tables S1 and S2 , while the PIP s are provided in Table S3 . Table 2. PIP s estimated using different BKMR methods for Gaussian and binomial outcomes in simulations with various sample sizes N Gaussian outcome Binomial outcome PIP 1 PIP 2 PIP 3 PIP 4 PIP 5 PIP 1 PIP 2 PIP 3 PIP 4 PIP 5 A-BKMR a 50 0.953 0.9816 0.5156 0.4398 0.2106 0.5588 0.5496 0.5772 0.5824 0.5498 100 0.989 0.9924 0.3852 0.2686 0.3692 0.4676 0.4492 0.491 0.5108 0.5326 500 1 1 0.111 0.0874 0.063 0.7624 0.6594 0.4598 0.4576 0.4444 1,000 1 1 0.034 0.027 0.0374 0.8892 0.8278 0.4648 0.3552 0.954 5,000 1 1 0.0068 0.0116 0 1 0.645 0.341 0.2378 0.2792 10,000 1 1 0.0006 0 0 1 1 0.038 0.2964 0 50,000 1 1 0 0 0 1 1 0.0004 0.096 0 100,000 1 1 0 0 0 1 1 0 0 0 PA-BKMR 50 0.953 0.9816 0.5156 0.4398 0.2106 0.5588 0.5496 0.5772 0.5824 0.5498 100 0.989 0.9924 0.3852 0.2686 0.3692 0.4676 0.4492 0.491 0.5108 0.5326 500 1 1 0.111 0.0874 0.063 0.7326 0.8558 0.4558 0.572 0.477 1,000 1 1 0.0298 0.0264 0.0374 0.8154 0.9614 0.4652 0.3906 0.8986 5,000 1 1 0.0004 0 0.0004 1 0.8706 0.3372 0.2152 0.2196 10,000 1 1 0.0044 0 0 1 1 0.13 0.175 0.0588 BKMR 50 0.9392 0.9902 0.4636 0.3992 0.1792 0.5068 0.5502 0.5062 0.5022 0.5094 100 0.9902 0.965 0.354 0.2254 0.2942 0.51 0.5452 0.5736 0.575 0.545 500 1 1 0.1026 0.0556 0.0752 0.7696 0.8246 0.5326 0.558 0.5518 1,000 1 1 0.025 0.0146 0.0224 0.9356 0.9172 0.636 0.6606 0.9572 BKMRhat 50 0.91 0.9806 0.4766 0.4478 0.182 0.6082 0.6822 0.6206 0.6546 0.6156 100 0.9722 0.9708 0.38 0.2228 0.3518 0.6796 0.754 0.7902 0.7868 0.7964 500 1 1 0.1332 0.0818 0.0816 0.964 0.9102 0.905 0.9552 0.992 1,000 1 1 0.0656 0.0382 0.0058 0.9546 0.9692 0.8914 0.8576 0.964 Open in a new tab PIP 1–5 , Posterior inclusion probabilities for the first to fifth exposures. a Some estimated PIP s were absent for PA-BKMR, BKMR, BKMRhat, and GBKMR. This occurred because, as the sample size increased, these methods exceeded the allowable computation time. For binomial outcomes, A-BKMR performed similarly to both BKMR and PA-BKMR. However, for sample sizes below 5,000, all four methods failed to reliably identify the important exposures ( PIP 3 – PIP 5 were large, even exceeding PIP 1 and PIP 2 ). This issue was alleviated as the sample size increased, highlighting the necessity of sample sizes greater than 5,000 for robust identification of important exposures in binomial outcome scenarios. As shown in Table 3 , A-BKMR demonstrated excellent predictive performance, achieving high R 2 and AUC for Gaussian and binomial outcomes, respectively. Although the other three methods also performed well for smaller sample sizes (e.g., ≤5,000), A-BKMR was the only method capable of handling large datasets (e.g., ≥50,000). Table 3. Prediction performance of different BKMR methods for Gaussian and binomial outcomes in simulations with various sample sizes N A-BKMR PA-BKMR a BKMR BKMRhat Gaussian outcome b 50 0.9848 0.9848 0.9848 0.9852 100 0.9794 0.9794 0.9793 0.9794 500 0.9806 0.9806 0.9806 0.9806 1,000 0.9772 0.9772 0.9772 0.9772 5,000 0.9782 0.9782 – – 10,000 0.9786 0.9786 – – 50,000 0.9785 – – – 100,000 0.9782 – – – Binomial outcome c 50 1.0000 1.0000 1.0000 1.0000 100 1.0000 1.0000 1.0000 1.0000 500 0.9958 0.9959 0.9960 0.9992 1,000 0.9978 0.9976 0.9976 0.9994 5,000 0.9940 0.9940 – – 10,000 0.9937 0.9937 – – 50,000 0.9939 – – – 100,000 0.9940 – – – Open in a new tab a Some prediction performance metrics were absent for PA-BKMR, BKMR, and BKMRhat. This occurred because, as the sample size increased, these methods exceeded the allowable computation time. b The prediction performance was evaluated by R 2 for the Gaussian outcome. c The prediction performance was evaluated by the area under the curve (AUC) for the binomial outcome. In the CHNS-based simulations with 10 and 21 exposures, A-BKMR successfully identified all important exposures. In the 10-exposure scenario (set 1), the PIPs for the important exposures (K, Na, Ca, Mg, and Fe) were 1, and the PIPs for the unimportant exposures (Mn, Zn, Cu, P, and Se) were 0. In the 21-exposure scenario (set 2), the PIPs for the important exposures (vegetation, legume, fruit, nut, fish, meat, and salt) were 1, while the PIP s for the unimportant exposures (grain, milk, rice, cake, cereal, tuber, and sugar) were 0. The prediction performance in both scenarios was excellent, with an R 2 of 0.889 for Gaussian outcomes and an AUC of 0.894 for binomial outcomes. Case study In the case study, we applied A-BKMR to the National Health and Nutrition Examination Survey (NHANES), a publicly available dataset comprising 5,900 participants from 2007 to 2016. The exposure variables were 24 polychlorinated biphenyls (PCBs). The binary outcome was diabetes diagnosis, defined by any of the following criteria: (1) self-reported physician-diagnosed diabetes condition, (2) use of oral hypoglycemic agents, (3) fasting plasma glucose levels exceeding 125 mg/dL, or (4) hemoglobin A1c (HbA1c) levels ≥6.5%. 52 A previous study 52 has investigated the association between these 24 PCBs and diabetes using the original BKMR. Our analysis identified PCB209, PCB66, PCB156, PCB138, and PCB167 as the most important 5 PCBs ( Figure 4 A), yielding results consistent with those obtained from the original BKMR. The joint exposure-response relationship between the 24 PCBs and diabetes risk is shown in Figure 4 B. The standard error of the predicted diabetes risk does not overlap with the x axis, suggesting that the simultaneous increase in all PCBs is associated with an increased diabetes risk. The estimated joint effect of a 1% simultaneous increase in all PCB concentrations was 0.015 (95% confidence interval [CI]: 0.013–0.017). The univariate effects of the 24 PCBs are shown in Figure 5 A. Each exposure-response curve represents the association between a selected PCB and diabetes risk, with the other PCBs held at their median values. The exposure-response curves for the 24 PCBs and diabetes generated by A-BKMR are nearly identical to those obtained by traditional BKMR. 52 We also conducted interaction analyses. For instance, the bivariate interaction between PCB209 and PCB156 was significant ( p = 0.005, Figure 5 B). However, when considering multivariate interactions with all other exposures, the association became nonsignificant ( p = 0.760, Figure 5 C). Figure 4. Open in a new tab Importance profile of PCBs and their joint effect on diabetes in the NHANES study (A) Estimated PIP s of 24 PCBs using A-BKMR. (B) Summary plot representing the joint effect of the 24 PCBs on diabetes based on A-BKMR. Dots and error bars represent the relative risks and 95% confidence intervals. Figure 5. Open in a new tab Univariate and interaction effects of PCBs in the NHANES study (A) Univariate exposure-response functions for the 24 PCBs estimated using A-BKMR. Shaded area represents the 95% confidence interval. (B) Bivariate interaction effect between PCB209 and PCB156 estimated using A-BKMR. P interaction denotes the significance of bivariate interaction. (C) Multivariate interaction effect between PCB209 and other PCBs estimated using A-BKMR. P trend denotes the significance of multivariate interaction. Discussion Principal findings In this study, we developed A-BKMR, an advanced version of BKMR that dramatically improves computational efficiency—reducing processing time from months or years to mere hours—while maintaining high performance for large-scale exposome studies. This methodological advancement enables the analysis of extensive datasets on standard personal computers, a task previously deemed unfeasible. Furthermore, A-BKMR provides quantitative statistics, enhancing result interpretability and offering a standardized framework for future exposome studies. BKMR is recognized as one of the state-of-the-art techniques recommended by leading exposome research authorities. 23 , 24 , 25 Up to now, it remains widely used in exposome studies to explore complex relationships between multiple exposures and health outcomes. In 2025 alone, notable applications include examining the link between volatile organic compound pollutants and chronic kidney disease, 53 assessing associations between metal exposures and thyroid cancer, 54 and investigating the impacts of body fat distribution on blood pressure. 55 Despite its analytical strengths, BKMR’s computational intensity has often restricted its application to studies with relatively small sample sizes. For instance, in a cohort of 69,210 adults, researchers encountered significant challenges in implementing BKMR due to its prohibitively slow processing speed. 33 This limitation underscores the urgent need for an optimized and computationally efficient version of BKMR that can accommodate large datasets. Compared to traditional BKMR, A-BKMR improves the slow computational speed and avoids huge computational resource requirements, therefore enabling its application for large-scale studies. Moreover, A-BKMR provides more interpretable estimates, such as various joint effects and interactions, making the method more accessible to researchers and facilitating wider adoption. Several extensions of BKMR, including PA-BKMR and BKMRhat, have been developed to mitigate its computational intensity. PA-BKMR, integrated into the original “bkmr” package, offers some improvements but remains computationally demanding. 22 For example, processing a dataset with 50,000 participants can take over 2 months. Additionally, PA-BKMR assumes that no observations in the dataset have identical exposure profiles. For instance, if participant A is exposed to PM 2.5 , NO 2 , and O 3 levels of 20, 10, and 30 μg/m 3 , respectively, then no other participant in the dataset can have the same exposure values for these three air pollutants (PM 2.5 , NO 2 , and O 3 ) of 20, 10, and 30 μg/m 3 . This assumption restricted its flexibility in handling repeated exposures commonly encountered in exposome studies. In contrast, A-BKMR incorporates multiple acceleration techniques, including Python-based optimizations and the Gaussian predictive process, to dramatically enhance efficiency. Furthermore, the weighted knot sampling enables A-BKMR to be applicable for datasets with repeated exposures. These advancements not only reduce processing time but also lower memory requirements, enabling the analysis of datasets with up to 100,000 observations on standard personal computers. Another critical limitation of BKMR is its reliance on graphical visualizations for interpreting exposure-outcome relationships, which often results in qualitative rather than quantitative conclusions. Traditional BKMR analyses lack numerical statistics, complicating the interpretation of results. 41 , 56 Some studies have attempted to quantify associations using arbitrary quantiles to estimate relative risks, 41 , 57 but this approach can introduce inconsistencies and reduce precision. Additionally, BKMR does not inherently provide statistical tests for interaction effects, leading many studies to rely on qualitative descriptions of interactions. 58 , 59 Some studies even adopted interaction test statistics from generalized additive models (GAMs), 39 , 57 but GAMs operate under different statistical assumptions than BKMR, potentially introducing bias and affecting the reliability of interaction estimates. To address these limitations, our A-BKMR introduces a comprehensive framework for quantitative analysis. It provides direct statistical estimates for both joint and univariate effects while incorporating rigorous tests for bivariate and multivariate interactions. These enhancements significantly improve the accuracy and interpretability of results. Other interpretable methods, such as WQS, QGC, and shrinkage regression, have been widely used for mixture exposure analysis. 60 , 61 While these methods offer faster computation and provide quantitative estimates, they rely on predefined regression formulas to model the relationship between multiple exposures and outcomes. This reliance limits their flexibility and may overlook essential statistical features, such as nonlinear relationships and interactions between exposures. In contrast, A-BKMR does not require predefined regression formulas, allowing for a more data-driven approach that can flexibly capture complex nonlinearities and interactions while still delivering the same quantitative insights comparable to those from WQS and QGC. Furthermore, A-BKMR is designed to be accessible to researchers without requiring extensive prior expertise in defining regression formulas, facilitating its broader application in exposome studies. Although A-BKMR can perform robust variable selection, its effectiveness may diminish when sample sizes are too small or when the number of exposures is too large. In such cases, some studies apply shrinkage-based regressions or random forests as a pre-screening step before conducting BKMR analysis. These techniques can be useful when the number of exposures is large relative to the sample size. However, shrinkage-based regressions assume a predefined regression formula, which can introduce bias if the assumptions are violated. Similarly, random forests can be sensitive to variance, the number of value categories, 62 and correlations among variables. 63 Therefore, researchers should apply BKMR directly when sample size permits and, if pre-screening is necessary, carefully evaluate each method’s assumptions to avoid excluding important exposures. In practice, WQS can yield results comparable to those from BKMR, particularly when every exposure is positively related to the outcome. This similarity likely arises from the “positive assumption” of WQS. 13 However, WQS may sometimes underestimate the effect, primarily because it does not fully account for interactions and nonlinearity. Exploring the mechanisms behind the similarity and clarifying the magnitude of the underestimation still remain important tasks for future studies. Implications A-BKMR has significant implications for large-scale exposome studies. First of all, it enables the completion of tasks on personal computers that traditional BKMR cannot handle, even with expensive computational resources such as high-memory servers. Second, A-BKMR provides a unified framework that integrates key statistical components, including the PIP for exposure importance, joint and univariate effects, and interaction effects. This comprehensive approach enhances the ability to dissect complex exposure-outcome relationships with greater precision and interpretability. Specifically, A-BKMR can be extended to various other methods, including multiple-imputed datasets, 64 mediation analysis, 65 distributed lag nonlinear models (DLNMs), 66 and longitudinal datasets with repeated measurements. 22 While these extensions based on the traditional BKMR framework have been developed with some progress, they have been limited by computational slowness and the lack of interpretability inherent to BKMR. A-BKMR can address these limitations and enhance the performance of these extensions, and we are actively working on these developments. Additionally, beyond environmental research, A-BKMR is equally well suited to other omics domains, including metabolomics and proteomics. In fact, recent work has begun to use BKMR for metabolomic analyses: for example, Mario et al. applied BKMR to investigate the impacts of lipoprotein profiles on diabetes risk, 67 and Yui et al. used BKMR to identify steroid metabolites associated with HbA1c and homeostatic model assessment 2-beta cell function (HOMA2-B). 68 Last but not least, A-BKMR’s model-independent quantitative framework extends beyond its own application, offering a means to enhance the interpretability of other machine learning methods with traditionally opaque black-box models. Its flexible kernel function enables it to capture complex relationships across various types of regression, such as single-variable, multiple linear, polynomial, spline, or other complex formulas, without requiring prior specification of an explicit regression formula. This adaptability makes A-BKMR a powerful and versatile tool for exposome studies and beyond. We also developed the R package “aBKMR” and provided a user guide ( Text S11 ; Figures S4–S10 ). Limitations of this study Despite these notable strengths, we acknowledge that the current version of A-BKMR has some limitations. First, the weighted knot sampling step introduces a computational burden, particularly for large datasets with numerous exposures, and future work should focus on optimizing this step. Secondly, A-BKMR is limited to Gaussian and binomial outcomes with single measurements per participant. Expanding its applicability to other outcome types, such as Poisson and survival outcomes, as well as longitudinal datasets, could further broaden its utility. Conclusion In conclusion, A-BKMR represents a significant advancement in the BKMR method, enabling large-scale exposome studies with substantially improved computational efficiency and interpretability. By facilitating complex exposure-outcome analyses on standard personal computers, even for extensive datasets, A-BKMR removes the computational barriers that previously limited BKMR’s application. Moreover, it provides rigorous quantitative insights into exposure effects and interactions, enhancing the depth and precision of exposome studies. A-BKMR paves the way for comprehensive assessments of exposures and holds great potential for large-scale modeling in exposome studies. By enabling the analysis of massive datasets and delivering quantitative statistics, A-BKMR transforms what was once impossible into something possible. Resource availability Materials availability This study did not generate new unique materials or reagents. Data and code availability Data from NHANES and CNHS were used in this study, which can be downloaded at https://wwwn.cdc.gov/nchs/nhanes and www.cpc.unc.edu/projects/china , respectively. The R package “aBKMR” and user guide are available on GitHub ( https://github.com/Guo-yi-y/A-BKMR ). Funding and acknowledgments This work was supported by the National Natural Science Foundation of China (grant numbers 82422065, 82173613, 82373681, and 82471130), the Shanghai 3-year Public Health Action Plan (grant number GWVI-11.2-YQ32), and the National Key R&D Program of China (grant number 2022YFC2704604). Author contributions Conceptualization, C.L. and Z.W.; methodology, Y.G., H.J., Z.P., K.P., X.X., Y.Z., and Z.Z.; visualization, Y.G., H.J., Z.P., and K.P.; supervision, C.L., Z.W., and H.K.; writing – original draft, C.L. and H.J.; writing – review & editing, Y.G., H.J., Z.P., X.X., Z.Z., K.P., Y.Z., H.K., Z.W., and C.L. All authors contributed to the manuscript and approved the final version. Declaration of interests The authors declare no conflicts of interest. Published Online: January 3, 2026 Footnotes It can be found online at https://doi.org/10.1016/j.xinn.2025.101248 . Contributor Information Zhenyu Wu, Email: [email protected]. Cong Liu, Email: [email protected]. Supplemental information Document S1. Figures S1–S10, Tables S1–S3, and Texts S1–S11 mmc1.pdf (962KB, pdf) Document S2. Article plus supplemental information mmc2.pdf (4.8MB, pdf) References 1. Wild C.P. The exposome: from concept to utility. Int. J. Epidemiol. 2012;41:24–32. doi: 10.1093/ije/dyr236. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Brauer M., Roth G.A. Global burden and strength of evidence for 88 risk factors in 204 countries and 811 subnational locations, 1990–2021: a systematic analysis for the Global Burden of Disease Study 2021. Lancet. 2024;403:2162–2203. doi: 10.1016/s0140-6736(24)00933-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Shaffer R.M., Sellers S.P., Baker M.G., et al. Improving and Expanding Estimates of the Global Burden of Disease Due to Environmental Health Risk Factors. Environ. Health Perspect. 2019;127 doi: 10.1289/ehp5496. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Vermeulen R., Schymanski E.L., Barabási A.-L., et al. The exposome and health: Where chemistry meets biology. Science. 2020;367:392–396. doi: 10.1126/science.aay3164. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Siroux V., Agier L., Slama R. The exposome concept: a challenge and a potential driver for environmental health research. Eur. Respir. Rev. 2016;25:124–129. doi: 10.1183/16000617.0034-2016. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Santos S., Maitre L., Warembourg C., et al. Applying the exposome concept in birth cohort research: a review of statistical approaches. Eur. J. Epidemiol. 2020;35:193–204. doi: 10.1007/s10654-020-00625-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Warembourg C., Anguita-Ruiz A., Siroux V., et al. Statistical Approaches to Study Exposome-Health Associations in the Context of Repeated Exposure Data: A Simulation Study. Environ. Sci. Technol. 2023;57:16232–16243. doi: 10.1021/acs.est.3c04805. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Carlin D.J., Rider C.V. Combined Exposures and Mixtures Research: An Enduring NIEHS Priority. Environ. Health Perspect. 2024;132 doi: 10.1289/ehp14340. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Bopp S.K., Barouki R., Brack W., et al. Current EU research activities on combined exposure to multiple chemicals. Environ. Int. 2018;120:544–562. doi: 10.1016/j.envint.2018.07.037. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Yu L., Liu W., Wang X., et al. A review of practical statistical methods used in epidemiological studies to estimate the health effects of multi-pollutant mixture. Environ. Pollut. 2022;306 doi: 10.1016/j.envpol.2022.119356. [ DOI ] [ PubMed ] [ Google Scholar ] 11. McEligot A.J., Poynor V., Sharma R., et al. Logistic LASSO Regression for Dietary Intakes and Breast Cancer. Nutrients. 2020;12 doi: 10.3390/nu12092652. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Copas J.B. Regression, Prediction and Shrinkage. J. Roy. Stat. Soc. B Stat. Methodol. 1983;45:311–335. doi: 10.1111/j.2517-6161.1983.tb01258.x. [ DOI ] [ Google Scholar ] 13. Carrico C., Gennings C., Wheeler D.C., et al. Characterization of Weighted Quantile Sum Regression for Highly Correlated Data in a Risk Analysis Setting. J. Agric. Biol. Environ. Stat. 2015;20:100–120. doi: 10.1007/s13253-014-0180-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Wang L., Yu C., Zhang Y., et al. Associations of the intake of individual and multiple fatty acids with depressive symptoms among adults in NHANES 2007–2018. J. Affect. Disord. 2024;365:364–374. doi: 10.1016/j.jad.2024.08.089. [ DOI ] [ PubMed ] [ Google Scholar ] 15. Keil A.P., Buckley J.P., O’Brien K.M., et al. A Quantile-Based g-Computation Approach to Addressing the Effects of Exposure Mixtures. Environ. Health Perspect. 2020;128 doi: 10.1289/ehp5838. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Zhang T., Zhu C., Zhao Y., et al. Deep Learning Model to Classify and Monitor Idiopathic Scoliosis in Adolescents Using a Single Smartphone Photograph. JAMA Netw. Open. 2023;6 doi: 10.1001/jamanetworkopen.2023.30617. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Chen Z., Dazard J.-E., Khalifa Y., et al. Deep Learning–Based Assessment of Built Environment From Satellite Images and Cardiometabolic Disease Prevalence. JAMA Cardiol. 2024;9:556–564. doi: 10.1001/jamacardio.2024.0749. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Zhang Z., Wang S., Zhu Z., et al. Identification of potential feature genes in non-alcoholic fatty liver disease using bioinformatics analysis and machine learning strategies. Comput. Biol. Med. 2023;157 doi: 10.1016/j.compbiomed.2023.106724. [ DOI ] [ PubMed ] [ Google Scholar ] 19. Chen X., Wang C.-C., Yin J., et al. Novel Human miRNA-Disease Association Inference Based on Random Forest. Mol. Ther. Nucleic Acids. 2018;13:568–579. doi: 10.1016/j.omtn.2018.10.005. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 2019;1:206–215. doi: 10.1038/s42256-019-0048-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Bobb J.F., Valeri L., Claus Henn B., et al. Bayesian kernel machine regression for estimating the health effects of multi-pollutant mixtures. Biostatistics. 2015;16:493–508. doi: 10.1093/biostatistics/kxu058. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Bobb J.F., Claus Henn B., Valeri L., et al. Statistical software for analyzing the health effects of multiple concurrent exposures via Bayesian kernel machine regression. Environ. Health. 2018;17 doi: 10.1186/s12940-018-0413-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Park E.S., Symanski E., Han D., et al. Vol. 183. Health Effects Institute; Boston, MA: 2015. Part 2. Development of Enhanced Statistical Methods for Assessing Health Effects Associated with an Unknown Number of Major Sources of Multiple Air Pollutants; pp. 51–113. (Development of Statistical Methods for Multipollutant Research. Research Report). [ PubMed ] [ Google Scholar ] 24. Coull B.A., Bobb J.F., Wellenius G.A., et al. Vol. 183. Health Effects Institute; Boston, MA: 2015. Part 1. Statistical Learning Methods for the Effects of Multiple Air Pollution Constituents; pp. 5–50. (Development of Statistical Methods for Multipollutant Research Research Report). [ PubMed ] [ Google Scholar ] 25. Peters S., Wan W., Portengen L., et al. Report on tutorial for the application of a suite of “multiple exposure methods”. The Exposome Project for Health and Occupational Research (EPHOR), Horizon 2020 (Grant Agreement No. 874703) 2022. https://www.we-expose.eu/multipleexposure_0102022.pdf 35 pp. 26. Taylor K.W., Joubert B.R., Braun J.M., et al. Statistical Approaches for Assessing Health Effects of Environmental Chemical Mixtures in Epidemiology: Lessons from an Innovative Workshop. Environ. Health Perspect. 2016;124:A227–A229. doi: 10.1289/ehp547. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Li H., Deng W., Small R., et al. Health effects of air pollutant mixtures on overall mortality among the elderly population using Bayesian kernel machine regression (BKMR) Chemosphere. 2022;286 doi: 10.1016/j.chemosphere.2021.131566. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Shi Y., Wang H., Zhu Z., et al. Association between exposure to phenols and parabens and cognitive function in older adults in the United States: A cross-sectional study. Sci. Total Environ. 2023;858 doi: 10.1016/j.scitotenv.2022.160129. [ DOI ] [ PubMed ] [ Google Scholar ] 29. Yao W., Liu C., Qin D.-Y., et al. Associations between Phthalate Metabolite Concentrations in Follicular Fluid and Reproductive Outcomes among Women Undergoing in Vitro Fertilization/Intracytoplasmic Sperm Injection Treatment. Environ. Health Perspect. 2023;131 doi: 10.1289/ehp11998. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Cheng X., Wei Y., Wang R., et al. Associations of essential trace elements with epigenetic aging indicators and the potential mediating role of inflammation. Redox Biol. 2023;67 doi: 10.1016/j.redox.2023.102910. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Li K., Yang Y., Zhao J., et al. Associations of metals and metal mixtures with glucose homeostasis: A combined bibliometric and epidemiological study. J. Hazard. Mater. 2024;470 doi: 10.1016/j.jhazmat.2024.134224. [ DOI ] [ PubMed ] [ Google Scholar ] 32. Goodrich A.J., Kleeman M.J., Tancredi D.J., et al. Ultrafine particulate matter exposure during second year of life, but not before, associated with increased risk of autism spectrum disorder in BKMR mixtures model of multiple air pollutants. Environ. Res. 2024;242 doi: 10.1016/j.envres.2023.117624. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Li S., Guo B., Jiang Y., et al. Long-term Exposure to Ambient PM2.5 and Its Components Associated With Diabetes: Evidence From a Large Population-Based Cohort From China. Diabetes Care. 2023;46:111–119. doi: 10.2337/dc22-1585. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Tang C., Zhang Y., Yi J., et al. The association between ozone exposure and blood pressure in a general Chinese middle-aged and older population: a large-scale repeated-measurement study. BMC Med. 2024;22 doi: 10.1186/s12916-024-03783-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Fan J., Liu S., Wei L., et al. Relationships between minerals’ intake and blood homocysteine levels based on three machine learning methods: a large cross-sectional study. Nutr. Diabetes. 2024;14 doi: 10.1038/s41387-024-00293-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Takatani T., Eguchi A., Yamamoto M., et al. Individual and mixed metal maternal blood concentrations in relation to birth size: An analysis of the Japan Environment and Children’s Study (JECS) Environ. Int. 2022;165 doi: 10.1016/j.envint.2022.107318. [ DOI ] [ PubMed ] [ Google Scholar ] 37. Cai C., Zhu S., Qin M., et al. Long-term exposure to PM2.5 chemical constituents and diabesity: evidence from a multi-center cohort study in China. Lancet Reg. Health West. Pac. 2024;47 doi: 10.1016/j.lanwpc.2024.101100. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Iwata H., Kobayashi S., Itoh M., et al. The association between prenatal per-and polyfluoroalkyl substance levels and Kawasaki disease among children of up to 4 years of age: A prospective birth cohort of the Japan Environment and Children’s study. Environ. Int. 2024;183 doi: 10.1016/j.envint.2023.108321. [ DOI ] [ PubMed ] [ Google Scholar ] 39. Zhai S., Zeng J., Zhang Y., et al. Combined health effects of PM2.5 components on respiratory mortality in short-term exposure using BKMR: A case study in Sichuan, China. Sci. Total Environ. 2023;897 doi: 10.1016/j.scitotenv.2023.165365. [ DOI ] [ PubMed ] [ Google Scholar ] 40. Tan L., Liu Y., Liu J., et al. Associations of individual and mixture exposure to volatile organic compounds with metabolic syndrome and its components among US adults. Chemosphere. 2024;347 doi: 10.1016/j.chemosphere.2023.140683. [ DOI ] [ PubMed ] [ Google Scholar ] 41. Che Z., Jia H., Chen R., et al. Associations between exposure to brominated flame retardants and metabolic syndrome and its components in U.S. adults. Sci. Total Environ. 2023;858 doi: 10.1016/j.scitotenv.2022.159935. [ DOI ] [ PubMed ] [ Google Scholar ] 42. Deng C.Y. A generalization of the Sherman–Morrison–Woodbury formula. Appl. Math. Lett. 2011;24:1561–1564. doi: 10.1016/j.aml.2011.03.046. [ DOI ] [ Google Scholar ] 43. Johnson M.E., Moore L.M., Ylvisaker D. Minimax and maximin distance designs. J. Stat. Plann. Inference. 1990;26:131–148. doi: 10.1016/0378-3758(90)90122-b. [ DOI ] [ Google Scholar ] 44. Kalal Z., Matas J.G., Mikolajczyk K. Procedings of the British Machine Vision Conference 2008. 2008. Weighted Sampling for Large-Scale Boosting. [ Google Scholar ] 45. Pfeffermann D. The Role of Sampling Weights When Modeling Survey Data. Int. Stat. Rev./Rev. Int. Stat. 1993;61:317. doi: 10.2307/1403631. [ DOI ] [ Google Scholar ] 46. Harris C.R., Millman K.J., van der Walt S.J., et al. Array programming with NumPy. Nature. 2020;585:357–362. doi: 10.1038/s41586-020-2649-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 47. Keil A. CRAN; 2020. Bkmrhat: Parallel Chain Tools for Bayesian Kernel Machine Regression. [ Google Scholar ] 48. Mou X., Zhang H., Arshad S.H. Generalized Bayesian kernel machine regression. Stat. Methods Med. Res. 2025;34:243–257. doi: 10.1177/09622802241280784. [ DOI ] [ PubMed ] [ Google Scholar ] 49. Hoofs H., van de Schoot R., Jansen N.W.H., et al. Evaluating Model Fit in Bayesian Confirmatory Factor Analysis With Large Samples: Simulation Study Introducing the BRMSEA. Educ. Psychol. Meas. 2018;78:537–568. doi: 10.1177/0013164417709314. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Zhang B., Zhai F.Y., Du S.F., et al. The China Health and Nutrition Survey, 1989–2011. Obes. Rev. 2014;15:2–7. doi: 10.1111/obr.12119. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 51. Steyerberg E.W., Vickers A.J., Cook N.R., et al. Assessing the Performance of Prediction Models. Epidemiology. 2010;21:128–138. doi: 10.1097/EDE.0b013e3181c30fb2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 52. Pan K., Jia H., Chen R., et al. Sex-specific, non-linear and congener-specific association between mixed exposure to polychlorinated biphenyls (PCBs) and diabetes in U.S. adults. Ecotoxicol. Environ. Saf. 2024;272 doi: 10.1016/j.ecoenv.2024.116091. [ DOI ] [ PubMed ] [ Google Scholar ] 53. Zhang Y., Qiu X., Wu Z., et al. Comprehensive analysis of the association between volatile organic compound pollutants and chronic kidney disease in hypertensive populations: insights from multi-omics approaches and identification of potential therapeutic targets. Environ. Res. 2025;282 doi: 10.1016/j.envres.2025.121966. [ DOI ] [ PubMed ] [ Google Scholar ] 54. Yu M., Xun J., Ge Y., et al. Relationship between internal metal exposure and thyroid cancer incidence: a case-control study simultaneously validated by BKMR and WQS models. Food Chem. Toxicol. 2025;201 doi: 10.1016/j.fct.2025.115443. [ DOI ] [ PubMed ] [ Google Scholar ] 55. Chen M., Wang X., Li Y., et al. Identifying joint association between body fat distribution with high blood pressure among 7 ∼ 17 years using the BKMR model: findings from a cross-sectional study in China. BMC Public Health. 2025;25 doi: 10.1186/s12889-024-20702-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 56. Tian Y., Luan M., Zhang J., et al. Associations of single and multiple perfluoroalkyl substances exposure with folate among adolescents in NHANES 2007–2010. Chemosphere. 2022;307 doi: 10.1016/j.chemosphere.2022.135995. [ DOI ] [ PubMed ] [ Google Scholar ] 57. Wen Y., Wang Y., Chen R., et al. Association between exposure to a mixture of organochlorine pesticides and hyperuricemia in U.S. adults: A comparison of four statistical models. Eco-Environ. Health. 2024;3:192–201. doi: 10.1016/j.eehl.2024.02.005. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 58. Zhang Y., Dong T., Hu W., et al. Association between exposure to a mixture of phenols, pesticides, and phthalates and obesity: Comparison of three statistical models. Environ. Int. 2019;123:325–336. doi: 10.1016/j.envint.2018.11.076. [ DOI ] [ PubMed ] [ Google Scholar ] 59. Tan Y., Fu Y., Yao H., et al. Relationship between phthalates exposures and hyperuricemia in U.S. general population, a multi-cycle study of NHANES 2007–2016. Sci. Total Environ. 2023;859 doi: 10.1016/j.scitotenv.2022.160208. [ DOI ] [ PubMed ] [ Google Scholar ] 60. Chen Y., Pan Z., Shen J., et al. Associations of exposure to blood and urinary heavy metal mixtures with psoriasis risk among U.S. adults: A cross-sectional study. Sci. Total Environ. 2023;887 doi: 10.1016/j.scitotenv.2023.164133. [ DOI ] [ PubMed ] [ Google Scholar ] 61. Bigambo F.M., Zhang M., Zhang J., et al. Exposure to a mixture of personal care product and plasticizing chemicals in relation to reproductive hormones and menarche timing among 12–19 years old girls in NHANES 2013–2016. Food Chem. Toxicol. 2022;170 doi: 10.1016/j.fct.2022.113463. [ DOI ] [ PubMed ] [ Google Scholar ] 62. Strobl C., Boulesteix A.-L., Zeileis A., et al. Bias in random forest variable importance measures: Illustrations, sources and a solution. BMC Bioinf. 2007;8:25. doi: 10.1186/1471-2105-8-25. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 63. Gregorutti B., Michel B., Saint-Pierre P. Correlation and variable importance in random forests. Stat. Comput. 2016;27:659–678. doi: 10.1007/s11222-016-9646-1. [ DOI ] [ Google Scholar ] 64. Bauer J.A., Devick K.L., Bobb J.F., et al. Associations of a Metal Mixture Measured in Multiple Biomarkers with IQ: Evidence from Italian Adolescents Living near Ferroalloy Industry. Environ. Health Perspect. 2020;128 doi: 10.1289/ehp6803. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 65. Devick K.L., Bobb J.F., Mazumdar M., et al. Bayesian kernel machine regression-causal mediation analysis. Stat. Med. 2022;41:860–876. doi: 10.1002/sim.9255. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 66. Wilson A., Hsu H.-H.L., Chiu Y.-H.M., et al. Kernel machine and distributed lag models for assessing windows of susceptibility to environmental mixtures in children’s health studies. Ann. Appl. Stat. 2022;16:1090–1110. doi: 10.1214/21-aoas1533. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 67. Delgado-Velandia M., Gonzalez-Marrachelli V., Domingo-Relloso A., et al. Healthy lifestyle, metabolomics and incident type 2 diabetes in a population-based cohort from Spain. Int. J. Behav. Nutr. Phys. Activ. 2022;19 doi: 10.1186/s12966-021-01219-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 68. Nakano Y., Yokomoto-Umakoshi M., Nakatani K., et al. Plasma Steroid Profiling Between Patients With and Without Diabetes Mellitus in Nonfunctioning Adrenal Incidentalomas. J. Endocr. Soc. 2024;8 doi: 10.1210/jendso/bvae140. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Document S1. Figures S1–S10, Tables S1–S3, and Texts S1–S11 mmc1.pdf (962KB, pdf) Document S2. Article plus supplemental information mmc2.pdf (4.8MB, pdf) Data Availability Statement Data from NHANES and CNHS were used in this study, which can be downloaded at https://wwwn.cdc.gov/nchs/nhanes and www.cpc.unc.edu/projects/china , respectively. The R package “aBKMR” and user guide are available on GitHub ( https://github.com/Guo-yi-y/A-BKMR ). Articles from The Innovation are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (3.7 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 4124 · SHA-256 8cd31ebc70bbe86b
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.