Scale reliant mixed effects models enhance microbiome data analysis - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Microbiome . 2026 Mar 25;14:117. doi: 10.1186/s40168-026-02377-x Search in PMC Search in PubMed View in NLM Catalog Add to search Scale reliant mixed effects models enhance microbiome data analysis Kyle C McGovern Kyle C McGovern 1 Program in Bioinformatics and Genomics, Pennsylvania State University, University Park, PA USA Find articles by Kyle C McGovern 1 , Justin D Silverman Justin D Silverman 1 Program in Bioinformatics and Genomics, Pennsylvania State University, University Park, PA USA 2 College of Information Sciences and Technology, Pennsylvania State University, University Park, PA USA 3 Department of Statistics, Pennsylvania State University, University Park, PA USA 4 Department of Medicine, Pennsylvania State University, Hershey, PA USA 5 Institute for Computational and Data Sciences, Pennsylvania State University, Hershey, PA USA Find articles by Justin D Silverman 1, 2, 3, 4, 5, ✉ Author information Article notes Copyright and License information 1 Program in Bioinformatics and Genomics, Pennsylvania State University, University Park, PA USA 2 College of Information Sciences and Technology, Pennsylvania State University, University Park, PA USA 3 Department of Statistics, Pennsylvania State University, University Park, PA USA 4 Department of Medicine, Pennsylvania State University, Hershey, PA USA 5 Institute for Computational and Data Sciences, Pennsylvania State University, Hershey, PA USA ✉ Corresponding author. Received 2025 Aug 18; Accepted 2026 Feb 9; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13085695 PMID: 41882807 Abstract Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional – they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package. Supplementary Information The online version contains supplementary material available at 10.1186/s40168-026-02377-x. Background A key goal in microbiome data analysis (e.g., 16S rRNA-seq or shotgun metagenomics) is to identify how experimental and biological factors (e.g., health vs. disease) affect the abundance of different taxa. Linear models are widely used for this task. Yet current linear modeling methods often inflate false-positive and false-negative rates because they cannot accommodate complex study designs or account for limitations of the sequencing measurement process [ 1 – 3 ]. Linear mixed-effects models (MEMs) extend linear models by incorporating random effects to model structured sources of variation, improving inference in hierarchical or longitudinal studies. For example, random intercepts can account for cage effects in co-housed mice [ 4 , 5 ], batch effects and contamination [ 6 – 8 ], or temporal autocorrelation in repeated measurements [ 9 , 10 ]. Several MEM-based tools have been developed for sequence count data, including NBZIMM [ 11 ], MaAsLin3 [ 12 ], lmerSeq [ 13 ], ANCOM-BC2 [ 14 ], and LinDA [ 15 ]. However, these methods remain fundamentally limited by compositionality: sequencing data capture only relative–not absolute–abundances and provide no information about the total biological scale (e.g., microbial load), which is essential for accurate inference [ 3 , 16 , 17 ]. Most methods attempt to circumvent this with normalizations such as Centered Log Ratio (ALDEx2), mean-variance weighting (lmerSeq, limma-voom), or sequencing-depth offsets (NBZIMM). Yet normalizations impose strong, hidden assumptions about scale, biasing estimates and inflating error rates [ 2 , 18 , 19 ]. More recent bias-correction approaches, such as ANCOM-BC2 and LinDA, still depend on unverifiable assumptions (e.g., that a stable set of taxa remains unchanged in absolute abundance across conditions) and fail when those assumptions are violated [ 15 , 20 ]. Here, we introduce Scale-Reliant Mixed-Effects Models (SR-MEM), a general framework for mixed-effects modeling of microbiome sequence count data. SR-MEM extends our recently developed Scale-Reliant Inference (SRI) framework, which uses specialized Bayesian partially identified models (PIMs) to avoid normalization and instead explicitly model uncertainty in the unmeasured biological scale through a user-defined prior, termed a scale model. By treating scale as an uncertain latent variable rather than assuming it can be recovered through normalization, SR-MEM provides a flexible and robust framework for mixed-effects modeling that is robust to errors in scale assumptions. Across simulations and real data analyses, SR-MEM consistently reduces false-positive and false-negative rates compared to existing methods. For example, in reanalyses of a longitudinal study on antibiotic effects in the gut microbiome, SR-MEM corrected previous failures in which normalization-based methods erroneously concluded that carbapenems had no effect other than increasing a single taxon—despite this class of antibiotics having broad activity against gut bacteria. Moreover, SR-MEM improves reproducibility: in replication studies, SR-MEM produces results consistent with prior findings, whereas normalization-based methods yield contradictory conclusions. We provide an open-source implementation of SR-MEM in the ALDEx3 R package, making the method broadly accessible for rigorous and reproducible mixed effects modeling of microbiome data. Results Scale-reliant mixed-effects models (SR-MEM) SR-MEM builds on the theory of Scale-Reliant Inference (SRI) , which provides a principled framework for estimating quantities that depend on absolute abundances when only noisy, compositional count data are observed [ 3 ]. The SRI framework Let W denote a D × N matrix of absolute abundances, where w dn is the abundance of taxon d in biological system n (e.g., a gut microbiome). W can be uniquely described by its composition (relative abundances) and scale (total abundances) via: w dn = w dn ‖ w n ⊥ , 1 where W ‖ is the composition: a D × N matrix of proportions satisfying ∑ d = 1 D w dn ‖ = 1 for each n ; and w ⊥ is the scale: a positive-valued N -vector of total abundances satisfying w n ⊥ = ∑ d = 1 D w dn . In practice, the absolute abundances W are not directly observed. Instead, we observe a D × N count table Y , where y dn is the number of sequencing reads assigned to taxon d in sample n . These counts provide noisy, sparse information about the composition W ‖ but little to no information about the scale w ⊥ . Specifically, SRI defines “little to no information” in terms of identification restrictions (see Supplementary file 1 for a formal definition) [ 3 ]. The goal of an SRI analysis is formalized as a target estimand : a quantity θ : = θ ( W ) that depends on the unobserved absolute abundances W and may therefore depend on both the composition W ‖ and the scale w ⊥ . Prior work has studied a range of estimands under the SRI framework, including the gene set enrichment estimand [ 19 ], correlations between taxa [ 3 ], and log-fold changes [ 2 , 3 , 19 ]. Notably, log-fold changes are themselves a special case of linear model estimands, which in turn are special cases of the mixed-effects model estimands we introduce next. The target estimand in SR-MEM SR-MEM extends SRI to mixed-effects modeling, enabling inference under complex experimental designs. Suppose we had access to the absolute abundances W . For each taxon d , our goal would be to estimate the fixed-effect parameters θ d in the standard linear mixed-effects model: log 2 w dn = x n θ d + z n u d + ε dn , 2 where: x n is a P -vector of fixed-effect covariates (observed), θ d is the P -vector of fixed effects (the target estimand; estimated), z n is a Q -vector of random-effect covariates (observed; e.g., subject, cage, collection site), u d ∼ N ( 0 , G ) is a Q -vector capturing structured random variation (estimated), and ε d · ∼ N ( 0 , R ρ ) is a residual error term (estimated). More formally, the target estimand is actually an estimator of θ d within this model (see Methods section). We model W on the log 2 scale because absolute abundances are non-negative and microbial communities are more naturally described on a log scale, where changes are often multiplicative rather than additive [ 21 ]. A full review of linear mixed-effects models is beyond the scope of this work (see, [ 22 ] for a general overview and [ 23 ] for a narrower overview with application to microbiome data analysis). Here, we briefly illustrate how the interpretation of the fixed effects θ d depends on the choice of fixed-effect covariates x n , random-effect covariates z n , and the covariance structures G and R ρ , all of which are determined by the experimental design: Binary covariates. If x n = ( 1 , x n ) , where x n = 0 for control and x n = 1 for treatment, then θ d = ( α d , β d ) , with α d the baseline log 2 -abundance in the control group and β d the log 2 -fold change in absolute abundance between treatment and control. That is, β d is the standard estimand used in differential abundance analyses. Continuous covariates. If x n = ( 1 , age n ) , then θ d = ( α d , γ d , age ) , where γ d , age represents the change in log 2 -abundance per unit increase in age, holding other covariates constant. Longitudinal studies. If x n = ( 1 , treatment n ) , where treatment n = 0 for control and 1 for treatment, then θ d , treatment represents the average log 2 -fold change in absolute abundance between groups. Suppose each subject is measured repeatedly over time. Then R ρ is often specified with an AR1 structure, where for two samples n and n ′ , cor ( ε dn , ε d n ′ ) = ρ | t - t ′ | , to account for temporal correlations within subjects. Hierarchical or nested designs. In clustered designs (e.g., mice co-housed in cages), z n could be a vector of cluster-level indicators (e.g., cage membership), and u d could represent the corresponding random effects. Together, z n u d accounts for variation introduced by the hierarchy, such as cage-to-cage differences. The fixed effects θ d are then interpreted as population-level effects after adjusting for this structured variability. Inference in SR-MEM Our goal is to estimate the fixed effects θ d from Eq. ( 2 ) as if the absolute abundances W were observed. Two sources of uncertainty complicate this task. First, the observed sequence counts Y are noisy and sparse, introducing uncertainty in the compositional component W ‖ . Second, and more fundamentally, the scale component w ⊥ is unidentifiable from Y alone because sequence counts provide no direct information about total microbial load. To address the first source of uncertainty, we follow the strategy used in ALDEx2: N independent Bayesian multinomial–Dirichlet models are fit to the observed counts Y , and posterior samples from these models are treated as plausible realizations of the underlying composition W ‖ (see Methods section). Prior work has shown that, in sequence count data, zeros arising from biological absence (“structural zeros”) and zeros arising from limited sequencing depth are statistically indistinguishable, as both correspond to taxa with very low absolute abundance [ 24 ]. Under this framework, posterior realizations of W ‖ are strictly positive, ensuring that log-scale modeling is well-defined. Zeros in the observed count matrix are therefore interpreted as reflecting uncertainty about very low absolute abundances rather than as exact absences, allowing inference on θ d to proceed without requiring an explicit distinction between structural and sampling zeros. The key innovation of SRI, and by extension SR-MEM, is to also account for uncertainty in the unobserved scale w ⊥ —a source of variation that normalization-based methods typically ignore. Standard normalizations implicitly assume that w ⊥ can be recovered without error (e.g., through sequencing depth offsets or total-sum scaling) [ 2 , 3 , 18 ]. Because Y contains no scale information, any single imputation of w ⊥ introduces bias. Instead, SRI introduces a scale model P : a user-defined probability distribution over w ⊥ that explicitly encodes prior knowledge and uncertainty about scale. This formulation places SR-MEM within the broader class of Bayesian partially identified models (PIMs), where inference proceeds by integrating over distributions of unidentifiable components rather than fixing them [ 3 ]. SR-MEM inference proceeds by sampling separate realizations of W ‖ (from multinomial-Dirichlet models) and w ⊥ (from the scale model P ). These are combined via Eq. ( 1 ) to generate realizations of the absolute abundances W . For each realization of W , a mixed-effects model (Eq. 2 ) is fit, yielding a corresponding realization of the target estimand θ d . Repeating this procedure across Monte Carlo samples produces a posterior distribution for θ d and for the corresponding p -values (adjusted for multiple testing as described in Methods section). Overall the SR-MEM inference procedure is intuitive, combining uncertainty about W ‖ and w ⊥ during inference of θ d . This inference procedure is also theoretically rigorous. Together this procedure defines a Scale Simulation Random Variable (SSRV) which is itself an example of a Bayesian Partially Identified Model [ 3 ]. Designing scale models for SR-MEM At first glance, defining a scale model for w ⊥ = w 1 ⊥ , ⋯ , w N ⊥ may seem daunting because it appears to require specifying a joint distribution over N unobserved scale parameters. However, the field of Scale-Reliant Inference (SRI) has developed practical strategies to make this task tractable. In Supplementary file 1, we review three general approaches: (1) constructing scale models directly from external measurements (e.g., flow cytometry or qPCR); (2) reformulating standard normalization procedures as scale models; and (3) avoiding explicit modeling of the full vector w ⊥ by instead specifying how covariates affect scale. To illustrate the third approach, consider a longitudinal study with repeated, biweekly measurements of the oral microbiome before and after tooth brushing. The goal is to identify taxa that change in abundance after brushing (differential abundance analysis). Rather than specifying a full joint distribution over w ⊥ , we can define a simpler model for the change in scale between conditions. Let θ ⊥ = mean x n = After log 2 w n ⊥ - mean x n = Before log 2 w n ⊥ 3 denote the difference in average log-scale between the two conditions. Because oral microbial load is expected to decrease after brushing, we anticipate θ ⊥ < 0 . A simple choice is an asymmetric prior, θ ⊥ ∼ Uniform ( - α , 0 ) , with α chosen conservatively based on prior literature. For instance, brushing has been reported to reduce oral microbial load by up to 99%, corresponding to α = 6.6 [ 25 , 26 ]. As discussed in the Methods section specifying a model for θ ⊥ implicitly defines a model for w ⊥ through its relationship in Eq. ( 3 ). Thus we have simplified the problem: we implicitly specify a scale model over w ⊥ by specifying a simpler, one-dimensional, model for θ ⊥ . SR-MEM substantially reduces false positives while maintaining power We benchmarked SR-MEM against five widely used mixed-effects modeling tools: two employing distinct normalization strategies (NBZIMM [ 11 ] and lmerSeq [ 13 ]) and three using bias-correction approaches (ANCOM-BC2 [ 14 ], MaAsLin3 [ 12 ], and LinDA [ 15 ]). Simulated 16S rRNA count data were generated with SparseDOSSA2, which reproduces realistic sparsity, variance, and taxon–taxon correlations based on real microbiome data [ 27 ]. We selected SparseDOSSA2 because its generative model is independent of all methods tested, including SR-MEM. Two simulation models were used, trained on oral [ 28 ] and skin [ 29 ] microbiome datasets. To make the scenarios biologically interpretable, we modeled a hypothetical prebiotic intervention in which participants in the treatment arm received a broad-spectrum prebiotic expected to increase total microbial load in the gut. We simulated studies with two equally sized groups (treatment and control), 6–300 participants, and 20 repeated measurements per participant (see Methods section). Three scenarios were evaluated: two based on the oral microbiome model with 50% and 80% of taxa differentially abundant (DA) and one based on the skin microbiome model with 50% DA. The oral microbiome simulations included on average 135 taxa, whereas the skin microbiome simulation included 35. The number of taxa varied slightly between simulations due to sparsity-based filtering, which removed different taxa in each simulation (see Methods section). Across all scenarios, the average log-scale difference in total microbial load between conditions was within θ ⊥ ∈ [ 1 , 2 ] . For SR-MEM, we used a moderately informative scale model θ ⊥ ∼ N ( 1.3 , 0 . 25 2 ) , which reflected an assumption that with 95% probability θ ⊥ ∈ [ 0.32 , 2.27 ] . SR-MEM was the only method that consistently controlled FDR across all study sizes and scenarios (Fig. 1 ). Supplementary Fig. S1 compares SR-MEM to ALDEx3 using the same scale model but with linear modeling (as implemented in ALDEx2). Unlike SR-MEM, ALDEx3 with linear modeling failed to control FDR, demonstrating that SR-MEM’s performance stems from integrating the scale model with mixed-effects modeling. Fig. 1. Open in a new tab SR-MEM uniquely controls false discoveries across diverse study designs while maintaining power. We benchmarked SR-MEM (solid blue line) against two normalization-based methods (lmerSeq, NBZIMM; black/gray dotted lines) and three bias-correction methods (ANCOM-BC2, MaAsLin3, LinDA; black/gray dashed lines). Simulated data were generated with SparseDOSSA2 trained on oral and skin microbiome datasets [ 28 , 29 ]. False discovery rate (FDR) and power were averaged over 20 simulations. Only SR-MEM consistently maintained FDR at or below the nominal 0.05 level (gray horizontal line). All methods used the Benjamini–Hochberg procedure for multiple hypothesis correction. A Oral microbiome model with 50% DA taxa. B Oral microbiome model with 80% DA taxa. C Skin microbiome model with 50% DA taxa Bias-correction methods (ANCOM-BC2, MaAsLin3, LinDA) failed to consistently control FDR. ANCOM-BC2 pools information across taxa to estimate scale differences between conditions, MaAsLin3 estimates bias in abundance estimates from the median regression coefficients after total sum scaling (TSS), and LinDA estimates bias from centered log-ratio (CLR) transformed data using the mode of estimated fixed effects. All three approaches failed to reliably recover bias terms from observed counts. ANCOM-BC2 controlled FDR only in the oral microbiome scenario with 50% DA taxa (Fig. 1 A), likely because a larger proportion of non-DA taxa stabilized its bias estimation. Its performance deteriorated when most taxa were DA (Fig. 1 B) or when fewer taxa were present (Fig. 1 C). SR-MEM generally had higher power than ANCOM-BC2, and although LinDA often achieved slightly higher power, it consistently failed to control FDR. In smaller studies ( < 80 participants) with 80% DA taxa, SR-MEM outperformed LinDA in both FDR and power (Fig. 1 B). MaAsLin3 failed to control FDR in all simulations, and had lower power. Normalization-based methods typically exhibited inflated FDR and generally lower power. NBZIMM’s failure aligns with prior findings that zero-inflation assumptions can increase false positives and negatives [ 24 ]. The lmerSeq method normalizes to a pseudo-reference derived from observed counts before applying a variance-stabilizing transformation [ 11 , 30 ]. These assumptions introduce bias, which lead to inflated false discoveries. We further evaluated these methods using additional simulation scenarios (Supplementary Figs. 2 and 3). Supplementary Fig. 2 presents gut microbiome simulations with increased taxonomic diversity ( 198 taxa), scale differences θ ⊥ ∈ [ 1 , 2 ] , and either 20 % or 80 % of taxa differentially abundant (DA). When 20 % of taxa are DA, SR-MEM, ANCOM-BC2, and MaAsLin3 approximately control FDR; when 80 % of taxa are DA, only SR-MEM consistently maintains FDR control. In both scenarios, SR-MEM achieved power comparable to or higher than competing methods. Supplementary Fig. 3 considers simulations with more modest signal, including gut and oral microbiome settings with 40 % DA taxa and minimal scale differences. Here we use a scale model θ ⊥ ∼ N ( 0 , 0 . 1 2 ) for SR-MEM, to reflect weak prior knowledge of the minimal scale differences between conditions. In this regime, where θ ⊥ ≈ 0 , SR-MEM, ANCOM-BC2, LinDA, and MaAsLin3 all control FDR. However, SR-MEM exhibits lower power ( ≈ 57 % ) compared to ANCOM-BC2, LinDA, and MaAsLin3 ( ≈ 65 % , 74 % , and 69 % , respectively). This behavior reflects the conservative nature of SR-MEM when scale differences are negligible and most taxa are not DA. In such settings, methods that impose stronger assumptions–such as normalization procedures that implicitly assume θ ⊥ = 0 –can achieve higher power. Together, these results highlight a trade-off between robustness and power. SR-MEM prioritizes valid FDR control under weak or uncertain scale assumptions, whereas normalization- and bias-correction–based methods can appear more powerful when their assumptions are satisfied. In these simulations, this robustness leads to conservative behavior, with near-zero FDR across scenarios (Fig. 1 ). This is expected rather than a numerical artifact: in the absence of identifiable scale information, exact FDR control (e.g., 0.05) is not achievable for differential abundance analysis [ 3 ]. Consequently, methods that achieve FDR control must be conservative in such settings. Importantly, this conservatism does not imply low power. As shown below, SR-MEM achieves comparable or higher power than competing methods in real data analyses while maintaining error control. SR-MEM enhances statistical power in an oral microbiome study To assess SR-MEM’s performance on real data, we reanalyzed an oral microbiome dataset where unstimulated saliva samples were collected from 28 participants across four perturbations: water (control), antiseptic mouthwash, alcohol-free mouthwash, and soda [ 31 ]. Microbial composition and total microbial load (scale) were measured via 16S rRNA sequencing and flow cytometry at three time points: baseline, 15 min post-perturbation, and 2 h post-perturbation. Two technical replicate flow cytometry measurements of total salivary microbial load were taken for each sample. We used these flow cytometry measurements to define a scale model that accounted for the technical variability in the measurement process (see Methods section). Using SR-MEM, we specified a mixed effects model with time, perturbation, and their interaction as fixed effects and with a participant-specific random intercept. Figure 2 shows results for the effect of alcohol-free mouthwash 15 min post-perturbation, the only condition with significant genus-level absolute abundance changes. See Supplementary file 2 for full results. Fig. 2. Open in a new tab SR-MEM improves detection of genera that differ in abundance after treatment with alcohol-free mouthwash. SR-MEM, QMP, and MaAsLin3 leveraged external microbial load measurements, while NBZIMM, LinDA, lmerSeq, and ANCOM-BC2 relied on normalization and bias correction. All methods used a mixed-effects model with perturbation, time point, and their interaction as fixed effects and a participant-specific random intercept. p -values were adjusted for multiple hypothesis testing using the Benjamini-Hochberg procedure. Sparsity indicates the percentage of samples where a given taxon was unobserved (zero counts), the last row represents the log 2 -transformed mean read count per taxon across all samples SR-MEM identified 17 genera which decreased in abundance 15 min after alcohol-free mouthwash (Fig. 2 ). Decreasing microbial abundances after mouthwash aligns with flow cytometry data which suggested a 3.8 -log decrease in microbial load in the post-mouthwash condition. We compared the SR-MEM results to six alternative methods: Quantitative Microbial Profiling (QMP) and MaAsLin3, which also used flow cytometry data, and four normalization- and bias correction-based approaches: ANCOM-BC2, LinDA, lmerSeq, and NBZIMM [ 16 ]. QMP identified only 5 of the 17 genera detected by SR-MEM. The reduced power of QMP compared to SRI-based methods is well established [ 3 ] and likely stems from differences in how the two approaches model the compositional component W ‖ . Both QMP and SR-MEM incorporate flow cytometry data to inform scale modeling; however, QMP first rarefies the data and then applies total-sum scaling to estimate W ‖ . Rarefaction discards a substantial fraction of reads, decreasing statistical power [ 3 , 32 ], whereas SR-MEM uses a Bayesian multinomial–Dirichlet model that retains all available counts and has been shown to improve power [ 3 ]. Consistent with this expectation, rarefaction in QMP increased the sparsity of the count matrix Y from 37 to 69%, indicating that a large proportion of data was effectively discarded. Many of the genera identified as decreasing by SR-MEM were also identified by MaAsLin3. However, only SR-MEM identified several high-count, low-sparsity genera (e.g., Campylobacter and Leptotrichia ), whereas only MaAsLin3 identified several low-count, high-sparsity genera (e.g., Peptostreptococcaceae and Bacteroides ). These differences arise from fundamental model assumptions. MaAsLin3 first models only the observed non-zero counts (augmented by total microbial load) and then separately models prevalence (i.e., zero occurrences); the resulting p -values are subsequently combined [ 12 ]. In contrast, SR-MEM models the full dataset jointly. Its multinomial-Dirichlet model treats zeros as low-count observations while explicitly modeling sequencing count noise. Consequently, SR-MEM penalizes low-count, high-sparsity genera and interprets zeros as evidence of reduced abundance. In addition, SR-MEM explicitly models per-sample scale uncertainty, whereas MaAsLin3 uses the average flow cytometry measurement for each sample (see Methods section). The normalization- and bias-correction methods produced results consistent with their simulated performance (Fig. 1 ). ANCOM-BC2 showed lower power than SR-MEM and detected no significant genus-level changes. LinDA and lmerSeq each identified only Alloprevotella as decreasing. NBZIMM’s zero-inflated model, which included an offset for total sequencing depth, frequently flagged high-sparsity, low-count genera as differentially abundant, likely inflating false positives (Fig. 2 ). For instance, it reported unrealistic increases of 4- to 30-fold in the absolute abundances of taxa with 64–69% sparsity ( Mogibacterium , Ottowia , and Peptidiphaga ) only 15 min after perturbation. This result also contradicts the expected broad-spectrum antimicrobial effects of alcohol-free mouthwash [ 33 ]. These findings mirror our simulation results, where NBZIMM consistently exhibited the highest FDR. Overall, these results suggest that SR-MEM has improved power compared with QMP and other normalization- and bias-correction–based methods. Relative to MaAsLin3, SR-MEM does not identify highly sparse, low-count taxa as significantly decreasing. Although the true fixed-effect coefficients are unknown, we expect inferences for low-sparsity, high-count taxa to be more reliable due to the greater amount of observed data. SR-MEM yields more reproducible results than normalization-based methods To evaluate the reproducibility of SR-MEM, we compared its results on a large metagenomic cohort of Crohn’s disease patients (the Inflammatory Bowel Disease Multi-omics Database; IBDMDB) to results from a smaller, independent 16S rRNA study (Vandeputte et al. [ 16 ]). The IBDMDB dataset comprised 944 samples from 27 healthy and 50 Crohn’s disease (CD) subjects, whereas the Vandeputte study included 95 samples from 29 healthy and 66 CD subjects. The Vandeputte study, which included paired flow cytometry measurements, was previously analyzed using scale reliant linear models [ 2 ]—appropriate for its cross-sectional design with a single measurement per subject. In contrast, the IBDMDB data include repeated biweekly measurements collected over 1 year across five body sites, making those earlier linear-model-based SRI methods inappropriate. Instead, mixed-effects modeling is required to account for participant- and site-specific variation. We therefore treated the Vandeputte results as a reference and evaluated the reproducibility of SR-MEM in this more complex longitudinal metagenomic setting. Because the IBDMDB dataset lacked paired flow cytometry or qPCR measurements, we followed Nixon et al. [ 2 ] and defined a scale model based on prior literature which specifically studied how total microbial load varied between healthy subjects and subjects with Crohn’s disease [ 34 ] (see Methods section). This scale model was centered at θ CD ⊥ = - 1.7 , suggesting total microbial load is decreased in Crohn’s disease and has been independently validated [ 2 ]. The IBDMDB data included biweekly samples collected for one year across five body sites, with each sample classified as dysbiotic or non-dysbiotic to reflect the episodic nature of CD [ 35 ]. We applied SR-MEM with fixed effects for antibiotic use and a disease–dysbiosis interaction, including participant- and site-specific random intercepts. Comparisons focused on dysbiotic samples from CD patients versus non-dysbiotic samples from healthy participants. For comparison, we analyzed two alternative methods labeled: (1) Lloyd-Price et al., which applied a mixed effects model (without SR-MEM’s multinomial-Dirichlet resampling) using arcsine square-root (ASR) normalization as performed in the original IBDMDB study [ 35 ]; (2) CLR-MEM , which used the same SR-MEM mixed effects modeling approach, but replaced the literature-based scale model with Centered Log-Ratio (CLR) normalization as implemented [ 30 ]. The results are shown in Fig. 3 . Fig. 3. Open in a new tab SR-MEM analysis of Crohn’s disease metagenomic data is more reproducible than normalization-based methods. A Differential abundance results for the 25 genera with the highest mean log 2 abundance (after adding a pseudo-count of 0.5 and applying a log 2 transformation). SR-MEM, using a scale model based on microbial load estimates from Sarrabayrouse et al. [ 34 ], produced results consistent with a prior ALDEx2 analysis of the full Vandeputte dataset when applied to the IBDMDB cohort. In contrast, two alternative analyses of the same IBDMDB data—(1) CLR-MEM , which replaced the scale model with a Centered Log-Ratio (CLR) normalization, and (2) Lloyd-Price et al., a mixed-effects model applied after arcsine square-root normalization as in the original IBDMDB study [ 35 ]—yielded results that differed both from SR-MEM and from the Vandeputte analysis. Notably, CLR normalization implicitly assumed a microbial load increase in CD patients of approximately θ CD ⊥ ≈ 3 , contradicting qPCR and flow cytometry measurements showing a substantial decrease [ 16 , 34 ]. B Sensitivity analysis of SR-MEM results. To assess the impact of potential misspecification of the scale model, a bias term b was introduced to shift the assumed scale difference to θ CD ⊥ = - 1.7 + b , and differential abundance was re-estimated across b ∈ [ - 5 , 5 ] . The case b = 0 corresponds to the SR-MEM analysis in part A The SR-MEM analysis of the IBDMDB dataset produced results largely consistent with the prior Vandeputte analysis, albeit with higher power, likely due to the substantially larger IBDMDB sample size (Fig. 3 A). Across all genera present in both studies, SR-MEM and the Vandeputte analysis agreed in both significance and direction of effect for all but five genera, which were significant in the IBDMDB analysis but not in Vandeputte. To assess whether these discrepancies reflected true positives (detected by SR-MEM due to increased power) or spurious false positives, we conducted a sensitivity analysis, previously introduced by Nixon et al. in ALDEx2 [ 2 ]. Specifically, we introduced a bias parameter b , shifting the assumed scale difference to θ CD ⊥ = - 1.7 + b , and re-estimated differential abundance across b ∈ [ - 5 , 5 ] (Fig. 3 B). Two of the five genera, Bacteroides and Ruthenibacterium , were highly robust, remaining significant unless the scale model was misspecified by more than b > 2 . Practically, b = 2 would imply an eight-fold higher microbial load in CD patients than supported by existing qPCR and flow cytometry data—an implausible scenario. These two genera therefore likely represent true positives that were undetected in the Vandeputte analysis due to either lower power or cohort-specific differences. In contrast, Odoribacter , Collinsella , and Parasutterella were more sensitive to perturbations in the scale model, suggesting that these results, while plausible, warrant further investigation and independent validation. Whereas the results of SR-MEM were closely aligned to the prior analysis of Vandeputte, the results of Lloyd-Price et al. and CLR-MEM were most similar to each other and markedly different than the Vandeputte analysis. Further investigation revealed that this similarity was primarily driven by normalization: in this study, CLR normalization implied the implausible assumption that gut microbial load was substantially increased in Crohn’s disease patients (i.e., θ CD ⊥ = 3 ), whereas qPCR and flow cytometry quantification of microbial load revealed a substantial decrease [ 16 , 34 ]. This suggests that the normalizations used in CLR-MEM and Lloyd-Price et al. likely introduced substantial bias. This conclusion is reinforced by noting that many of the differences between SR-MEM, Lloyd-Price et al., and CLR-MEM were highly sensitive to perturbations of the scale model ( b ). For example, Lloyd-Price et al. and CLR-MEM suggest that Blautia is increased in Crohn’s but this result only holds in the unlikely scenario b = 5 [ 16 , 34 ]. Our findings suggest that normalization-based methods can introduce substantial bias into analyses, degrading reproducibility. In contrast, SR-MEM combined with scale models built simply from prior literature can mitigate this bias and enhance reproducibility. SR-MEM provides more accurate antibiograms from longitudinal data We reanalyzed a longitudinal study of 128 cancer patients hospitalized to receive allogeneic hematopoietic cell transplantation [ 36 ]. The dataset included 2589 fecal samples with paired 16S rRNA sequencing and qPCR measurements, providing information on both gut microbiome composition and total microbial load. During hospitalization, patients received various oral and intravenous antibiotics. The original authors used these data to estimate antibiograms (tables summarizing the effect of each antibiotic on each gut microbe) but many of their findings were inconsistent with the known spectrum of activity of these drugs. We hypothesized that these inconsistencies arose from limitations in the statistical models used in the original analysis, and that applying SR-MEM would yield results more consistent with the known spectrum of antibiotic activity. We used SR-MEM with subject-specific random intercepts to estimate the fixed effect of each antibiotic. In longitudinal studies, fixed-effects models that include time, treatment, and their interaction are commonly used to assess differential temporal trends (as in the oral perturbation study). In the antibiotic dataset, however, sample collection times were highly irregular, with substantial variation in spacing and duration across subjects. Moreover, our primary interest was in estimating overall antibiotic perturbation effects rather than modeling linear time trends. As a result, we modeled temporal dependence using autocorrelated residuals, without linear time-treatment fixed effects, to account for within-subject correlation among repeated measurements (see Methods section). We used the paired qPCR measurements to define a scale model (see Methods section). In Fig. 4 we compare the results of SR-MEM with the results from the original authors methods for the antibiotics vancomycin and carbapenems [ 36 ]. A full comparison of all antibiotics are presented in Supplementary Figs. S4 and S5. Fig. 4. Open in a new tab SR-MEM improves inference in a longitudinal study of antibiotic effects. Comparison of SR-MEM and the original analysis by Liao et al. [ 36 ] on the impact of orally- and intravenously-administered vancomycin and intravenously-administered carbapenems on the 20 most abundant gut microbial taxa at the ASV level. SR-MEM used a mixed-effects model with antibiotics as fixed effects, subject-level random intercepts, and autocorrelated residuals to capture temporal relationships between repeated measurements. The scale model was based on paired qPCR measurements. In contrast, Liao et al. estimated absolute abundances by multiplying ASV relative abundances by sample-specific qPCR totals, followed by ridge-penalized regression on time-differenced abundances. The reported antibiotic effect sizes for the Liao et al. analysis are the effects of the antibiotics over 21days. Zero-count samples were excluded. Effect sizes are shown as heatmaps: red indicates increased abundance associated with antibiotic use. Cells are unannotated and shown in white if either the Benjamini–Hochberg adjusted p -value exceeds 0.05 or the absolute effect size is below 1. The latter filter is used because Liao et al. did not report p -values Liao et al. [ 36 ] reported that intravenous vancomycin had a stronger effect on the gut microbiota than oral vancomycin—a conclusion that contradicts well-established pharmacokinetic data. Vancomycin is a water-soluble glycopeptide with poor gastrointestinal absorption; when administered intravenously, only trace amounts reach the gut [ 37 , 38 ]. In contrast, oral vancomycin exerts a potent effect on the gut microbiota due to its broad-spectrum activity against gram-positive bacteria [ 39 ]. Its use as part of standard preoperative gastrointestinal decontamination further underscores its localized gut activity [ 40 ]. In line with these principles, SR-MEM yielded more biologically plausible results, indicating that oral vancomycin has a strong depleting effect on gram-positive taxa in the gut, while intravenous vancomycin has a comparatively modest impact. Another notable inconsistency in the Liao et al. analysis concerns the carbapenem class of antibiotics. Carbapenems are potent, broad-spectrum antibiotics known to significantly deplete gut microbial diversity [ 38 , 41 ]. Yet, their model identified only a single association: a modest increase in the abundance of Lachnospiraceae , an obligate anaerobic family well within the known spectrum of carbapenem activity [ 38 , 42 ]. In contrast, SR-MEM found that carbapenem exposure was associated with broad depletion of gut microbes, including Lachnospiraceae , consistent with prior literature and pharmacological expectations. Together, these results suggest SR-MEM improves inference in longitudinal studies. Unlike the original analysis, SR-MEM simultaneously addresses sparsity, counting-uncertainty, time-dependent correlation, and scale measurement uncertainty. Discussion In this article, we introduced Scale Reliant Mixed-Effects Models (SR-MEM) . Unlike bias-correction– or normalization-based approaches, SR-MEM enables estimation of fixed effects while explicitly modeling uncertainty in the unmeasured scale and accommodating repeated measures in nested, hierarchical, or longitudinal study designs. SR-MEM extends the scale-simulation random variable (SSRV) framework [ 3 ], previously limited to linear models, to mixed-effects models required for longitudinal and hierarchical microbiome studies. We show that SSRV-based inference remains stable and computationally tractable despite the non-linear optimization and random-effect structure of mixed-effects estimators, enabling scale-reliant inference for complex study designs. Moreover, we show that the same simple and interpretable scale models developed for linear SSRVs remain valid in SR-MEM, without requiring additional modeling of the random-effect structure. Across simulated and real data analyses, SR-MEM consistently reduced false discovery rates relative to normalization- and bias-correction–based methods while maintaining or improving statistical power, and yielded more plausible and reproducible results on previously published datasets. SR-MEM’s multinomial-Dirichlet modeling approach requires fitting mixed effects models across all D taxa for a total of S Monte Carlo samples (which are typically around 2000 , see Methods section). While computationally more expensive than many existing methods that fit only a few models per taxon, computational complexity scales only linearly in N , D , and S . To address this, we make an implementation of SR-MEM that uses parallelization of the S Monte Carlo samples available in the ALDEx3 R software package [ 43 ]. The SR-MEM approach with scale models contributes more broadly to advancing rigor and reproducibility in statistical modeling for scientific research [ 44 ]. Similar to the principles of veridical data science, scale models and sensitivity analyses in SR-MEM address data perturbations and model stability, respectively [ 45 , 46 ]. Conceptually, SR-MEM is a type of Bayesian partially identified model (PIM) [ 3 ]. By explicitly accounting for uncertainty in key parameters, Bayesian PIMs facilitate inference under weaker assumptions than fully identified models and have found use across diverse fields, including econometrics, political science, and public health [ 47 – 50 ]. While SR-MEM provides a robust and practical framework for estimating fixed effects in the presence of repeated measures and scale uncertainty, future refinements could further enhance its utility and scope. We highlight two promising directions for future work. First, although SR-MEM is designed to account for complex experimental structures using random effects, our current formulation focuses exclusively on inference for fixed effects. Inference on random effects remains an open challenge. This is due to a subtle but important feature of the scale models used in this work: by design, these models become more conservative as the variance of the scale distribution increases. That is, as uncertainty in the scale (e.g., due to imprecision in flow cytometry measurements) increases, uncertainty in the estimated fixed effects increases monotonically. This desirable property stems from the linear structure of fixed effects within mixed-effects models. However, this monotonicity may not hold for random effects. In such settings, increasing uncertainty in scale could paradoxically reduce uncertainty in the random effects or even induce bias. These complexities suggest that our current scale models may not be appropriate for inference on random effects. Future work should investigate the conditions under which reliable inference is possible and develop specialized scale models tailored to random effects estimation. A second promising, though less explored, direction involves leveraging recent work on predicting microbial load directly from sequence count data. For instance, Nishijima et al. [ 51 ] developed a machine learning model that predicts microbial load using features derived from sequence counts. While promising, their model achieved only modest performance, with out-of-sample R 2 values in the range of 0.3–0.4, even within similar study populations. Nonetheless, such predictive models could serve as the basis for novel scale models. Rather than treating predicted microbial load values as ground truth, they could be incorporated probabilistically, wrapped in a scale model to acknowledge their uncertainty. This approach would provide a principled way to integrate machine learning predictions into statistical analyses, without requiring the unrealistic assumption that these predictions are error-free. Methods SR-MEM The target estimand and its estimator In SR-MEM the goal is to infer the fixed effects term θ d as in Eq. ( 2 ). Core to the SRI approach is the target estimand : the quantity dependent on the absolute abundances W we wish to estimate. For SR-MEM the target estimand is explicitly defined as the generalized least squares estimator of the fixed effects θ d : θ d = ( X ⊤ V - 1 X ) - 1 X ⊤ V - 1 log 2 W d · , 4 where V = Z G Z ⊤ + R ρ is estimated using the restricted maximum likelihood (REML) estimator to reduce bias [ 52 ]. The SR-MEM algorithm The SR-MEM algorithm produces S independent Monte Carlo samples (referred to as “realizations” in the main text). For each Monte Carlo sample, the following steps are performed: Estimate composition: for each sample n ∈ { 1 , ⋯ , N } , SR-MEM draws a posterior sample of the composition W · n ‖ from a multinomial–Dirichlet distribution fit to the observed sequence counts in sample n , y · n : W · n ‖ ∼ Dirichlet y · n + 0.5 . Sample scale: the total scale w ⊥ = ( w 1 ⊥ , ⋯ , w N ⊥ ) is drawn from a user-defined scale model P : w ⊥ ∼ P . Construct absolute abundances: the sampled compositions and scales are combined using Eq. ( 1 ) to generate an estimate of the absolute abundances W . Fit mixed-effects model: a mixed-effects model (Eq. ( 2 )) is fit to W using either the nlme or lme4 R package interfaces. Across all Monte Carlo samples, SR-MEM summarizes the fixed effects for each taxon or gene d by reporting the arithmetic mean of the S posterior estimates of θ d . Following Nixon et al. [ 2 ], SR-MEM also summarizes p -values for each covariate. For Monte Carlo sample s and taxon or gene d , let p d , l ( s ) and p d , u ( s ) be the one-sided, lower-tail and upper-tail Student’s t -test p -values, respectively, calculated from the mixed effects model t -statistic using degrees of freedom calculated using the Welch–Satterthwaite method. The average lower-tail and upper-tail p -values across Monte Carlo samples are p d , l = 1 S ∑ s = 1 S p d , l ( s ) and p d , u = 1 S ∑ s = 1 S p d , u ( s ) , respectively. The final p -value for taxon d is p d = 2 × min ( p d , l , p d , u ) . Multiple hypothesis test correction is applied across all taxa for each individual Monte Carlo sample prior to averaging and final p -value calculation. Scale models in SR-MEM SR-MEM requires estimates of the N -length vector of scales, w ⊥ . Any scale model P specified over the scale parameters θ ⊥ (i.e., θ ⊥ ∼ P ) can be expressed equivalently as a model over w ⊥ via the following relationship for each sample n : log 2 w n ⊥ = ∑ p = 1 P x np θ p ⊥ , 5 where x np is the value of the p -th covariate for sample n . The resultant vector log 2 w ⊥ is then transformed to w ⊥ and used to compute the absolute abundances W (Eq. ( 1 )), enabling inference within the SR-MEM framework. Further details are provided in Supplementary file 1. Simulation benchmarking Simulation details All simulations were performed using SparseDOSSA2 (v0.99.2), which generates ground-truth absolute abundances and sequence count data [ 27 ]. Data were simulated with equal-sized control and treatment groups. Each participant contributed 20 replicate samples and was assigned a random intercept drawn from N ( 0 , 1 ) . Random intercepts were incorporated into the simulation using SparseDOSSA2’s spike-in abundance parameter. For all simulation scenarios, SparseDOSSA2 models were trained using a correlation penalization parameter λ = 0.3 , allowing for inter-microbe correlation. For the oral microbiome scenarios, a model was trained on 100 healthy human buccal mucosa 16S rRNA-seq samples and 197 genera, pre-processed as previously described [ 28 ]. For the skin microbiome scenario, the model was trained on 33 healthy human armpit samples and 40 families, also pre-processed as described [ 29 ]. For the gut microbiome scenarios, a model was trained on 100 16S rRNA-seq samples from individuals with type 2 diabetes and 200 genera. In all cases, only taxa with no more than 70% sparsity were included for model training, and only samples with sequencing depth ≥ 1000 reads were included. For each scenario and each study size (ranging from 6 to 300 participants), we ran 20 independent simulations. Within each simulation, treatment effects and differentially abundant (DA) taxa were randomly assigned. Fixed treatment effects were applied using SparseDOSSA2’s spike-in abundance parameter, while the prevalence spike-in parameters were set to one-quarter of the fixed effect size to reflect changes in sparsity associated with the fixed effects. For the oral microbiome scenario with 50% DA taxa, fixed effects were drawn from Uniform ( 0 , 3 ) for 40% of taxa and from Uniform ( - 2 , 0 ) for 10%. In the oral and gut scenarios with 80% DA taxa, 70% were simulated from Uniform ( 0 , 3 ) and 10% from Uniform ( - 2 , 0 ) . For the skin microbiome scenario, fixed effects were drawn from Uniform ( 0 , 3 ) for 50% of taxa. For the oral and gut microbiome scenarios with 40% taxa DA, 20% were simulated from Uniform ( - 2.5 , 0 ) and 20% from Uniform ( 0 , 2.5 ) . In the gut scenario with 20% DA taxa, fixed effects were drawn from Uniform ( 0 , 3 ) . To compute the true θ ⊥ for each simulation, ground-truth absolute abundances W were first transformed to w ⊥ using Eq. ( 1 ). A mixed-effects model was then fit to log 2 w ⊥ , including a fixed effect for treatment and a random intercept for participants. Simulations were constrained such that θ ⊥ ∈ [ 1 , 2 ] , except when 40% of taxa were DA the constraint was θ ⊥ ∈ [ - 0.25 , 0.25 ] . All samples were generated with a fixed sequencing depth of 80,000 reads. Analysis details Before model fitting, taxa with greater than 70% sparsity were amalgamated into a single “other” group. All methods, except NBZIMM and ALDEx3 with linear modeling, used the same mixed-effects model structure, which included a fixed effect for treatment and a random intercept for each participant. NBZIMM additionally included a fixed-effect offset for log 2 sequencing depth, as recommended [ 11 ]. ALDEx3 with linear modeling did not include a random intercept. Both SR-MEM and ALDEx3 assumed the scale model θ ⊥ ∼ N ( 1.3 , 0 . 25 2 ) , except in the 40% taxa DA scenarios where the scale model was θ ⊥ ∼ N ( 0 , 0 . 1 2 ) , which were transformed to corresponding scale models over w ⊥ using Eq. ( 5 ). SR-MEM and ALDEx3 with linear modeling were implemented using ALDEx3 (v0.4.0) with 300 Monte Carlo samples. SR-MEM used the lme4 package for model fitting. ANCOM-BC2 (v2.6.0) was run with default parameters; taxa that failed the pseudo-count sensitivity test were assigned p -values of 1 after multiple hypothesis test correction [ 14 ]. MaAsLin3 (v1.2.0) was run with total sum scaling (TSS) normalization, median bias correction for the abundance test, and Benjamini-Hochberg hypothesis test correction on the joint abundance/prevalence p -values. NBZIMM (v1.0) used the mms function with a zero-inflated negative binomial model. LinDA (v1.2) was run with zero imputation enabled. lmerSeq (v0.1.7) applied DESeq2 size-factor normalization followed by variance-stabilizing transformation (VST), as recommended [ 13 , 53 ]. For all methods, taxa grouped as “other” were assigned p -values of 1 following multiple testing correction. A taxon was considered a true positive if the null hypothesis of no treatment effect was rejected at FDR ≤ 0.05 and the sign of the estimated treatment effect matched the ground truth. If the sign was incorrect, the taxon was included in the denominator for power calculations but excluded from the numerator. All methods used the Benjamini-Hochberg procedure for multiple testing correction. Reanalysis of oral microbiome perturbation study Sequence count data were pre-processed following Marotz et al. [ 31 ], and analyses were conducted at the genus level. Taxa with more than 70% sparsity were grouped into a single “other” category, resulting in a total of 59 taxa. The dataset comprised 81 samples from 21 participants. For each sample, oral microbial load was measured in duplicate using flow cytometry. Denoting μ n ⊥ as the mean and σ n ⊥ as the standard deviation of the two log-transformed replicate cell counts for sample n , we modeled the absolute abundance as: w n ⊥ ∼ N μ n ⊥ , σ n ⊥ 2 . All methods except NBZIMM used the same mixed-effects model structure, including fixed effects for time point, perturbation, and their interaction, as well as a random intercept for each participant. The fixed effects intercept corresponded to the baseline (pre-treatment) time point under the reference perturbation (water). NBZIMM additionally included a fixed-effect offset for log 2 sequencing depth. For SR-MEM, 2000 Monte Carlo samples were drawn using the lme4-based implementation in ALDEx3 (v0.4.0). A prior of 0.5 was used (see Supplementary Fig. S6 for fit of posterior predictive). For the MaAsLin3 and QMP analyses the average of the two microbial load replicates per sample were used for the scale. For QMP, absolute abundance estimates W were estimated using the QMP method of Vandeputte et al. [ 16 ], A pseudo-count of 0.5 was added prior to log-transformation. The same mixed-effects model described above was applied to each taxon. For MaAsLin3, non-log–scale average measurements were provided as input to augment the observed data to absolute abundances, as described in [ 12 ]. Median bias correction was not used, as it is not recommended for absolute abundances. Benjamini–Hochberg correction was applied to the joint abundance/prevalence p -values. NBZIMM (v1.0) used the mms function with a zero-inflated negative binomial model. ANCOM-BC2 (v2.2.1) was run with default settings; taxa failing the pseudo-count sensitivity test were assigned a p -value of 1 following multiple testing correction. LinDA (v1.2) applied zero imputation. lmerSeq (v0.1.7) used DESeq2 size-factor normalization followed by variance-stabilizing transformation (VST), as recommended [ 13 , 53 ]. All methods applied multiple testing correction using the Benjamini–Hochberg procedure. Reanalysis of IBDMDB data Metagenomic data pre-processing Raw metagenomic sequence data were obtained from the IBDMDB database [ 35 ]. Host DNA was removed using Kneaddata (v0.12.0), and taxonomic profiling with read count estimation was performed using MetaPhlAn (v4.0.6) and the CHOCOPhlAn database (v30) [ 54 , 55 ]. Genera with zero counts in more than 75% of samples were grouped into a single “other” category, resulting in 43 taxa analyzed across 1,297 samples from 106 individuals. Sample metadata, including gut dysbiosis classification (dysbiotic vs. non-dysbiotic), were also obtained from the IBDMDB database. Analyses in the main text compared 582 Crohn’s disease (CD) samples to 362 non-IBD samples; the remaining 353 samples from ulcerative colitis (UC) patients are reported separately (Supplementary file 3). Analyses of IBDMDB data Disease diagnosis (non-IBD, UC, or CD) and dysbiosis status were combined into a single covariate labeled disease-dysbiosis . All analyses used a mixed-effects model including fixed effects for antibiotic use and disease-dysbiosis, and random intercepts for participant and site. The fixed-effects intercept corresponded to non-IBD samples collected during non-dysbiosis without antibiotic use. SR-MEM used 1000 Monte Carlo samples and a prior scale model informed by visual inspection of Fig. 2 in Sarrabayrouse et al. [ 34 ]: θ CD ⊥ θ UC ⊥ ∼ N - 1.7 - 3.1 , 0.5 0 0 0.5 , where θ CD ⊥ and θ UC ⊥ represent the difference in average log 2 scale for CD and UC samples during dysbiosis, relative to non-IBD samples during non-dysbiosis. CD and UC samples taken during non-dysbiosis and non-IBD samples taken during dysbiosis were assumed to have differences in average log 2 scale of zero. This scale model was transformed to one over w ⊥ using Eq. ( 5 ). In the Lloyd-Price et al. analysis, observed counts were converted to relative abundances w dn ‖ = y dn / ∑ y · n , then transformed using the arcsine square-root (ASR) transformation: w dn = arcsin w dn ‖ . These transformed values were used in the same mixed-effects model as SR-MEM. The CLR-MEM analysis was identical to SR-MEM, except that absolute abundances were estimated using centered log-ratio normalization: w dn = w dn ‖ gm ( w · n ‖ ) , where gm denotes the geometric mean across taxa within each sample. Multiple hypothesis testing correction for all analyses was performed using the Benjamini-Hochberg procedure. Reanalysis of vandeputte data Sequence count data were pre-processed as described by Vandeputte et al. [ 16 ]. Genera with zero counts in more than 75% of samples were aggregated into a single “other” category, resulting in 30 taxa analyzed across 95 samples. The scale model, previously defined by Nixon et al. [ 2 ], was: w n ⊥ ∼ N log 2 q n , 0.7 , where q n denotes the flow cytometry measurement of cells per gram of frozen feces for sample n . ALDEx2 (v1.40.0) was used to compare disease status (CD vs. control) with 1000 Monte Carlo samples. Multiple testing correction was performed using the Benjamini-Hochberg procedure. A prior of 0.5 was used (see Supplementary Fig. S7 for fit of posterior predictive). Reanalysis of longitudinal antibiotic dataset Data pre-processing 16S rRNA sequencing count data were pre-processed by Liao et al. and obtained from Figshare (version 6) [ 36 ]. Taxa were analyzed at the amplicon sequence variant (ASV) level. Antibiotics were categorized by route of administration: oral (glycopeptides, penicillins, quinolones, sulfonamides, and macrolide derivatives) or intravenous (aztreonam, carbapenems, cephalosporins, glycopeptides, metronidazole, oxazolidinones, penicillins, and quinolones). All other antibiotics were grouped into an “other” category. Antiviral and antifungal treatments were excluded from filtering and modeling. SR-MEM analysis Samples lacking paired qPCR measurements were excluded. A sample was labeled as unaffected by antibiotics if it was collected either before antibiotic administration or at least 21 days afterward. Samples collected within 21 days post-antibiotic-treatment were excluded. Patients with fewer than 10 samples were also removed. After filtering, the final dataset included 2589 samples from 144 patients. The SR-MEM model included 14 fixed effects: one for each antibiotic-route combination and one for the “other” antibiotic category. A random intercept was included for each patient. SR-MEM analysis was performed using the nlme-based implementation of ALDEx3 (v0.4.0). A first-order autoregressive correlation structure with the formula ∼ TimePoint | PatientID was used. This defines the residual covariance matrix as R ρ = σ 2 L ρ , where each element l xy of L ρ is l xy = ρ | t x - t y | , if P x = P y 0 , if P x ≠ P y , with t x and t y representing the collection time points (in days), and P x and P y denoting the patients corresponding to samples x and y , respectively. The autocorrelation parameter ρ was estimated via restricted maximum likelihood (REML). We used 3,000 Monte Carlo samples in the SR-MEM analysis. For each sample n with an observed qPCR value q n , the scale model was: w n ⊥ ∼ N ( log 2 q n , 0.25 ) , where q n is the qPCR-measured microbial load (cells per gram of stool). Multiple testing correction was performed using the Benjamini-Hochberg procedure. A prior of 0.5 was used (see Supplementary Fig. S8 for fit of posterior predictive). Liao analysis The original Liao et al. analysis was reproduced as described in their publication, using the provided MATLAB code [ 36 ]. Multiple testing correction was performed using the Benjamini-Hochberg procedure. Supplementary information Supplementary Material 1. (2.9MB, zip) Acknowledgements The authors would like to thank Dr. Rachel Silverman for her manuscript comments. Authors’ contributions KCM wrote all code and performed all analyses in the manuscript. JDS obtained funding. KCM and JDS contributed the idea for the scale-reliant mixed effects modeling statistical framework. All authors read and approved the manuscript. Funding KCM and JDS were supported by NIGMS R01GM148972-01. Data availability All datasets used in this study are publicly available and have been previously published. The oral microbiome data used to train SparseDOSSA2 are available from Qiita (study ID 10370), and the skin microbiome training data are accessible via the HMP2Data R package. Data from the oral microbiome perturbation study are available from Qiita (study ID 11899). The IBDMDB dataset is available at https://ibdmdb.org . Data from the longitudinal antibiotic study are available on Figshare at (DOI 10.6084/m9.figshare.c.5271128.v6) Code availability All code used to generate figures, supplementary figures, and supplementary files is available through Zenodo at https://zenodo.org/records/18480042 (DOI 10.5281/zenodo.18480042). The implementation of SR-MEM is included in the ALDEx3 R package, available on the Comprehensive R Archive Network (CRAN) and on GitHub at https://github.com/jsilve24/ALDEx3 . Declarations Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Footnotes Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. References 1. Roche KE, Mukherjee S. The accuracy of absolute differential abundance analysis from relative count data. PLoS Comput Biol. 2022;18(7):e1010284. 10.1371/journal.pcbi.1010284. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Nixon MP, Gloor GB, Silverman JD. Incorporating scale uncertainty in microbiome and gene expression analysis as an extension of normalization. Genome Biol. 2025. 10.1186/s13059-025-03609-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Nixon MP, Letourneau J, David LA, Lazar NA, Mukherjee S, Silverman JD. Scale Reliant Inference. 2023. Preprint at arXiv:2201.03616 . 4. Barlow JT, Bogatyrev SR, Ismagilov RF. A quantitative sequencing framework for absolute abundance measurements of mucosal and lumenal microbial communities. Nat Commun. 2020;11(1):2590. 10.1038/s41467-020-16224-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Singh G, Brass A, Cruickshank SM, Knight CG. Cage and maternal effects on the bacterial communities of the murine gut. Sci Rep. 2021. 10.1038/s41598-021-89185-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Salter SJ, Cox MJ, Turek EM, Calus ST, Cookson WO, Moffatt MF. Reagent and laboratory contamination can critically impact sequence-based microbiome analyses. BMC Biol. 2014. 10.1186/s12915-014-0087-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Brooks JP, Edwards DJ, Harwich MD, Rivera MC, Fettweis JM, Serrano MG, et al. The truth about metagenomics: quantifying and counteracting bias in 16S rRNA studies. BMC Microbiol. 2015. 10.1186/s12866-015-0351-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Randall DW, Kieswich J, Swann J, McCafferty K, Thiemermann C, Curtis M. Batch effect exerts a bigger influence on the rat urinary metabolome and gut microbiota than uraemia: a cautionary tale. Microbiome. 2019. 10.1186/s40168-019-0738-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Mu A, Carter GP, Li L, Isles NS, Vrbanac AF, Morton JT, et al. Microbe-metabolite associations linked to the rebounding murine gut microbiome postcolonization with vancomycin-resistant Enterococcus faecium . mSystems. 2020. 10.1128/msystems.00452-20. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Romero R, Hassan SS, Gajer P, Tarca AL, Fadrosh DW, Nikita L, et al. The composition and stability of the vaginal microbiota of normal pregnant women is different from that of non-pregnant women. Microbiome. 2014. 10.1186/2049-2618-2-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Zhang X, Yi N. NBZIMM: negative binomial and zero-inflated mixed models, with application to microbiome/metagenomics data analysis. BMC Bioinformatics. 2020. 10.1186/s12859-020-03803-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Nickols WA, Kuntz T, Shen J, Maharjan S, Mallick H, Franzosa EA, et al. MaAsLin 3: Refining and extending generalized multivariable linear models for meta-omic association discovery. bioRxiv. 2024. Preprint at 10.1101/2024.12.13.628459. [ DOI ] [ PMC free article ] [ PubMed ] 13. Vestal BE, Wynn E, Moore CM. lmerSeq: an R package for analyzing transformed RNA-Seq data with linear mixed effects models. BMC Bioinformatics. 2022. 10.1186/s12859-022-05019-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Lin H, Peddada SD. Multigroup analysis of compositions of microbiomes with covariate adjustments and repeated measures. Nat Methods. 2023;21(1):83–91. 10.1038/s41592-023-02092-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Zhou H, He K, Chen J, Zhang X. LinDA: linear models for differential abundance analysis of microbiome compositional data. Genome Biol. 2022;23(1):95. 10.1186/s13059-022-02655-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Vandeputte D, Kathagen G, D’hoe K, Vieira-Silva S, Valles-Colomer M, Sabino J, et al. Quantitative microbiome profiling links gut community variation to microbial load. Nature. 2017;551(7681):507–11. [ DOI ] [ PubMed ] [ Google Scholar ] 17. Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ. Microbiome datasets are compositional: and this is not optional. ISME J. 2017;8:2224. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. McGovern KC, Silverman JD. Replacing normalizations with interval assumptions enhances differential expression and differential abundance analyses. BMC Bioinformatics. 2025. 10.1186/s12859-025-06177-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. McGovern KC, Nixon MP, Silverman JD. Addressing erroneous scale assumptions in microbe and gene set enrichment analysis. PLoS Comput Biol. 2023;19(11):1–16. 10.1371/journal.pcbi.1011659. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Lin H, Peddada SD. Analysis of compositions of microbiomes with bias correction. Nat Commun. 2020;11(1):3514. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Silverman JD, Roche K, Holmes ZC, David LA, Mukherjee S. Bayesian multinomial logistic normal models through marginally latent matrix-T processes. J Mach Learn Res. 2022;23(7):1–42. [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Gelman A, Hill J. Data Analysis Using Regression and Multilevel/Hierarchical Models. Cambridge University Press; 2007. [ Google Scholar ] 23. Xia Y, Sun J. Linear Mixed-Effects Models for Longitudinal Microbiome Data. In: Bioinformatic and Statistical Analysis of Microbiome Data. Springer; 2023. 24. Silverman JD, Roche K, Mukherjee S, David LA. Naught all zeros in sequence count data are the same. Comput Struct Biotechnol J. 2020;18:2789–98. 10.1016/j.csbj.2020.09.014. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Slot D, Wiggelinkhuizen L, Rosema N, der Van Weijden G. The efficacy of manual toothbrushes following a brushing exercise: a systematic review. Int J Dent Hyg. 2012;10(3):187–97. 10.1111/j.1601-5037.2012.00557.x. [ DOI ] [ PubMed ] [ Google Scholar ] 26. Hope CK, Petrie A, Wilson M. Efficacy of removal of sucrose-supplemented interproximal plaque by electric toothbrushes in an in vitro model. Appl Environ Microbiol. 2005;71(2):1114–6. 10.1128/aem.71.2.1114-1116.2005. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Ma S, Ren B, Mallick H, Moon YS, Schwager E, Maharjan S, et al. A statistical model for describing and simulating microbial community profiles. PLoS Comput Biol. 2021;17(9):e1008913. 10.1371/journal.pcbi.1008913. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Fettweis JM, Serrano MG, Brooks JP, Edwards DJ, Girerd PH, Parikh HI, et al. The vaginal microbiome and preterm birth. Nat Med. 2019;25(6):1012–21. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Bouslimani A, da Silva R, Kosciolek T, Janssen S, Callewaert C, Amir A, et al. The impact of skin care products on skin chemistry and microbiome dynamics. BMC Biol. 2019. 10.1186/s12915-019-0660-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Fernandes AD, Reid JN, Macklaim JM, McMurrough TA, Edgell DR, Gloor GB. Unifying the analysis of high-throughput sequencing datasets: characterizing RNA-seq, 16S rRNA gene sequencing and selective growth experiments by compositional data analysis. Microbiome. 2014;2(1):1–13. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Marotz C, Morton JT, Navarro P, Coker J, Belda-Ferre P, Knight R, et al. Quantifying live microbial load in human saliva samples over time reveals stable composition and dynamic load. mSystems. 2021. 10.1128/msystems.01182-20. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. McMurdie PJ, Holmes S. Waste not, want not: why rarefying microbiome data is inadmissible. PLoS Comput Biol. 2014;10(4):e1003531. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Brookes Z, McGrath C, McCullough M. Antimicrobial mouthwashes: an overview of mechanisms-what do we still need to know? Int Dent J. 2023;73:S64-8. 10.1016/j.identj.2023.08.009. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Sarrabayrouse G, Elias A, Yáñez F, Mayorga L, Varela E, Bartoli C, et al. Fungal and bacterial loads: noninvasive inflammatory bowel disease biomarkers for the clinical setting. mSystems. 2021;6(2):10–1128. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Lloyd-Price J, Arze C, Ananthakrishnan AN, Schirmer M, Avila-Pacheco J, Poon TW, et al. Multi-omics of the gut microbial ecosystem in inflammatory bowel diseases. Nature. 2019;569(7758):655–62. 10.1038/s41586-019-1237-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Liao C, Taylor BP, Ceccarani C, Fontana E, Amoretti LA, Wright RJ, et al. Compilation of longitudinal microbiota data and hospitalome from hematopoietic cell transplantation patients. Sci Data. 2021;8(1). 10.1038/s41597-021-00860-8. [ DOI ] [ PMC free article ] [ PubMed ] 37. Eubank TA, Hu C, Gonzales-Luna AJ, Garey KW. Detectable vancomycin stool concentrations in hospitalized patients with diarrhea given intravenous vancomycin. Pharmacoepidemiol Drug Saf. 2023;2(4):283–8. 10.3390/pharma2040024. [ Google Scholar ] 38. Gallagher JC, MacDougall C. Antibiotics simplified. Jones & Bartlett Learning; 2022. [ Google Scholar ] 39. Kim AH, Lee Y, Kim E, Ji SC, Chung JY, Cho JY. Assessment of oral vancomycin-induced alterations in gut bacterial microbiota and metabolome of healthy men. Front Cell Infect Microbiol. 2021;11:629438. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Ambe PC, Zarras K, Stodolski M, Wirjawan I, Zirngibl H. Routine preoperative mechanical bowel preparation with additive oral antibiotics is associated with a reduced risk of anastomotic leakage in patients undergoing elective oncologic resection for colorectal cancer. World J Surg Oncol. 2019;17:1–6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Kuang H, Yang Y, Luo H, Lv X. The impact of three carbapenems at a single-day dose on intestinal colonization resistance against carbapenem-resistant Klebsiella pneumoniae . mSphere. 2023;8(6):e00479–23. 10.1128/msphere.00479-23. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Hayase E, Hayase T, Jamal MA, Miyama T, Chang CC, Ortega MR, et al. Mucus-degrading Bacteroides link carbapenems to aggravated graft-versus-host disease. Cell. 2022;185(20):3705-3719.e14. 10.1016/j.cell.2022.09.007. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Silverman J, Gloor G, McGovern K. ALDEx3: Linear Models for Sequence Count Data. The R Foundation; 2026. 10.32614/cran.package.aldex3 44. Ioannidis JPA. Why most published research findings are false. PLoS Med. 2005;2(8):e124. 10.1371/journal.pmed.0020124. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Yu B, Kumbier K. Veridical data science. Proc Natl Acad Sci U S A. 2020;117(8):3920–9. 10.1073/pnas.1901326117. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 46. Yu B. Stability. Bernoulli. 2013;19(4). 10.3150/13-BEJSP14. 47. Manski CF, Molinari F. Estimating the COVID-19 infection rate: anatomy of an inference problem. J Econ. 2021;220(1):181–92. 10.1016/j.jeconom.2020.04.041. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 48. Manski CF, Pepper JV. How do right-to-carry laws affect crime rates? Coping with ambiguity using bounded-variation assumptions. Rev Econ Stat. 2018;100(2):232–44. 10.1162/rest_a_00689. [ Google Scholar ] 49. Silverman JD, Hupert N, Washburne AD. Using influenza surveillance networks to estimate state-specific prevalence of SARS-CoV-2 in the United States. Sci Transl Med. 2020;12(554):eabc1126. 10.1126/scitranslmed.abc1126. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Kline B, Tamer E. Bayesian inference in a class of partially identified models: Bayesian inference in partially identified models. Quant Econ. 2016;7(2):329–66. 10.3982/qe399. [ Google Scholar ] 51. Nishijima S, Stankevic E, Aasmets O, Schmidt TSB, Nagata N, Keller MI, et al. Fecal microbial load is a major determinant of gut microbiome variation and a confounder for disease associations. Cell. 2025;188(1):222-236.e15. 10.1016/j.cell.2024.10.022. [ DOI ] [ PubMed ] [ Google Scholar ] 52. Bates D, Mächler M, Bolker B, Walker S. Fitting Linear Mixed-Effects Models Using lme4. J Stat Softw. 2015;67(1). 10.18637/jss.v067.i01 53. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15(12):1–21. 10.1186/s13059-014-0550-8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 54. Blanco-Míguez A, Beghini F, Cumbo F, McIver LJ, Thompson KN, Zolfo M, et al. Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nat Biotechnol. 2023;41(11):1633–44. 10.1038/s41587-023-01688-w. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 55. Beghini F, McIver LJ, Blanco-Míguez A, Dubois L, Asnicar F, Maharjan S, et al. Integrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3. eLife. 2021. 10.7554/elife.65088. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Material 1. (2.9MB, zip) Data Availability Statement All datasets used in this study are publicly available and have been previously published. The oral microbiome data used to train SparseDOSSA2 are available from Qiita (study ID 10370), and the skin microbiome training data are accessible via the HMP2Data R package. Data from the oral microbiome perturbation study are available from Qiita (study ID 11899). The IBDMDB dataset is available at https://ibdmdb.org . Data from the longitudinal antibiotic study are available on Figshare at (DOI 10.6084/m9.figshare.c.5271128.v6) All code used to generate figures, supplementary figures, and supplementary files is available through Zenodo at https://zenodo.org/records/18480042 (DOI 10.5281/zenodo.18480042). The implementation of SR-MEM is included in the ALDEx3 R package, available on the Comprehensive R Archive Network (CRAN) and on GitHub at https://github.com/jsilve24/ALDEx3 . Articles from Microbiome are provided here courtesy of BMC ACTIONS View on publisher site PDF (3.6 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top