Unusual outcome variances as a method to identify potentially problematic clinical trials - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice PLoS One . 2026 Apr 15;21(4):e0346238. doi: 10.1371/journal.pone.0346238 Search in PMC Search in PubMed View in NLM Catalog Add to search Unusual outcome variances as a method to identify potentially problematic clinical trials Philippe P Hujoel Philippe P Hujoel 1 Department of Epidemiology, School of Public Health, University of Washington, Seattle, Washington, United States of America 2 Department of Oral Health Sciences, School of Dentistry, University of Washington, Seattle, Washington, United States of America Conceptualization, Data curation, Formal analysis, Methodology, Writing – original draft, Writing – review & editing Find articles by Philippe P Hujoel 1, 2, * , Margaux LA Hujoel Margaux LA Hujoel 3 Department of Human Genetics, University of California, Los Angeles, California, United States of America 4 Department of Computational Medicine, University of California, Los Angeles, California, United States of America Formal analysis, Methodology, Supervision, Writing – review & editing Find articles by Margaux LA Hujoel 3, 4 Editor: Robin Haunschild 5 Author information Article notes Copyright and License information 1 Department of Epidemiology, School of Public Health, University of Washington, Seattle, Washington, United States of America 2 Department of Oral Health Sciences, School of Dentistry, University of Washington, Seattle, Washington, United States of America 3 Department of Human Genetics, University of California, Los Angeles, California, United States of America 4 Department of Computational Medicine, University of California, Los Angeles, California, United States of America 5 Max Planck Institute for Solid State Research, GERMANY Competing Interests: The authors have declared that no competing interests exist. ✉ * E-mail: [email protected] Roles Philippe P Hujoel : Conceptualization, Data curation, Formal analysis, Methodology, Writing – original draft, Writing – review & editing Margaux LA Hujoel : Formal analysis, Methodology, Supervision, Writing – review & editing Robin Haunschild : Editor Received 2025 May 15; Accepted 2026 Mar 13; Collection date 2026. © 2026 Hujoel, Hujoel This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice PMCID: PMC13082665 PMID: 41984903 Abstract An unusual outcome variance contributed to uncovering major cases of research misconduct, leading to over 200 retractions. Detecting such problematic randomized trials early – before they unduly influence clinical guidelines – remains challenging. Empirical evidence indicates that differences in variances between trial arms (DiVBTAs) are usually small and non-significant in properly conducted trials. This study investigated whether the converse – unusually large and statistically significant DiVBTAs - can serve as a red flag for potentially problematic trials. We conducted simulations to assess the sensitivity and specificity of a DiVBTA-based decision rule under realistic scenarios, including proper randomization, heterogeneous treatment effects, and missing-not-at-random data. In parallel, we applied the rule in a real-world analysis of 226 systematically sampled randomized trials in diabetes research to assess whether unusually large and statistically significant DiVBTAs occur with sufficient frequency to warrant screening. Unusually large DiVBTA values were defined as those falling outside the 3-sigma prediction limits. Simulations demonstrated high specificity, with legitimate trials rarely flagged (low false-positive rate), and adequate sensitivity for detecting a specific form of severe fabrication. In the empirical analysis, 19 out of 226 trials (8%) were flagged as potentially problematic demonstrating utility to screening trials for unusually large and statistically significant DiVBTAs. Subsequent screening of the identified trials revealed additional concerns in 18 (out of 19) flagged trials. These findings suggest that screening for unusually large, statistically significant DiVBTAs offers a simple, low-effort tool to identify trials warranting further scrutiny, potentially strengthening the reliability of evidence used in clinical guidelines. Introduction Problematic clinical trials are widespread and erode the credibility of health information. It is estimated that 25% of published clinical trials may be flawed or fraudulent [ 1 ]. In absolute terms, hundreds of thousands of trials are believed to lack credibility [ 2 ]. These unreliable trials have permeated meta-analyses and clinical guidelines [ 3 – 6 ], which has led to concerted efforts at curbing their influence. These efforts have included global regulatory guidelines which have imposed requirements on data collection that are designed to minimize fraud, misrepresentation, and data integrity issues [ 7 ]. The conduct of systematic reviews presents itself with its own set of challenges to prevent potentially problematic trials from infiltrating clinical guidelines. A 2021 Cochrane editorial warned of the threat posed by untrustworthy or “problematic” studies, highlighting that retracted studies are only the tip of the iceberg. The Cochrane editorial coincided with the release of new Cochrane guidelines on how to handle concerns about the trustworthiness of a publication when no formal post-publication correction exists [ 8 ]. Checklists to identify problematic trials [ 9 – 12 ] and recommendations on how these tools should become integrated during research synthesis followed [ 13 ]. The 2021 Cochrane editorial also underscored the urgent need for validated statistical methods to reliably and fairly detect trials with statistical irregularities [ 8 ]. One particularly promising method involves identifying improbable distributions of baseline data across trial arms [ 14 ]. This method exploits the cornerstone assumption that participants are allocated randomly to interventions leading to predictable distributions of baseline variables across groups. Trials flagged as having highly unusual distributions when compared to those expected by chance have shown a higher likelihood of retraction [ 15 – 17 ]. Additional statistical methods have become available and at least two statistical packages integrate methods to re-appraise the publication integrity in groups of randomized controlled trials [ 18 – 21 ]. An unusual outcome variance – described as a “very small standard deviation” by the first whistleblower– helped in the discovery of the largest fraudulent research body identified to date [ 22 ]. Building on the informativeness of unusual standard deviations to inform on fraud, we propose that Differences in Variance Between Trial Arms (DiVBTAs) of continuous outcome measures can offer the basis to develop an objective method to identify potentially problematic trials. Empirical evidence in support of this proposal is that a preponderance of meta-analyses of DiVBTAs within the setting of clinical trials demonstrated that DiVBTAs are typically not statistically significant, or, when statistically significant, are small in size [ 23 – 27 ]. The power of statistical tests furthermore to detect significant DiVBTA is low in both meta-analyses of clinical trials, and, especially so within the setting of single clinical trials [ 21 ]. The discovery of statistically significant DiVBTAs within the setting of a single clinical trial can thus be regarded as a somewhat unexpected finding. Explanations for such unexpected findings currently focus on genuine design and analysis issues such as randomization, heterogeneous treatments effects, informative trial participant dropout, compliance issues, or floor and ceiling effects of the outcome variable. We suggest here that fraud or unintentional error needs to be included in the list of plausible explanations. As such, we (1) describe methods to identify statistically significant DiVBTA outliers, (2) perform simulations to assess how likely genuine design and analysis issues can cause statistically significant DiVBTA outliers and (3) provide a case-study on clinical trials included in systematic reviews on diabetes. Methods The methods section is presented in two parts (i) simulations to assess the sensitivity and specificity of the proposed DiVBTA decision rule, and (ii) a diabetes case study to assess whether the proposed decision rule has real-world clinical utility. The following background presents DiVBTA terms discussed in the two proceeding subsections. Background: The proposed decision rule to flag a potentially problematic trial has two elements: (1) the DiVBTA has to be statistically significant, and (2) the DiVBTA needs to be an outlier, i.e., fall outside a tolerance band. A first step is selecting a DiVBTA statistic among those available (for a review of DiVBTA statistics see [ 21 ]). For the detection of potentially problematic trials, it is advantageous to focus on those DiVBTA estimators (a) which can be derived from published summary statistics, (b) which are standardized by the mean, and (c) which are normally distributed and robust. First, selecting a DiVBTA estimator which can be derived from published summary statistics is crucial given that it remains uncommon for authors of clinical trials to share individual participant data. A 2019 survey of authors from 619 randomized controlled trials published in high-impact anesthesiology journals (2014–2016) found that only about 4% provided individual participant data upon request [ 28 ]. A 2019 randomized controlled trial assessing the impact of financial incentives to encourage data sharing reported that none of the investigators provided individual participant data [ 29 ]. As a result, DiVBTA estimators requiring individual participant data for calculations remain currently of little value in identifying potentially problematic trials. Potential summary statistics of interest to calculate DiVBTA measures are the standard deviation (s), the mean ( x ¯ ), and the derived measure of coefficient of variation (CV). For a two-arm trial, C V T = s T x ― T ― and C V C = s C x ― C ― where subscripts T and C denote treatment and control, respectively. Second, DiVBTA measures which minimize assumption about mean-variance relationships have been recommended over measures which are built on the assumption that no mean-variance relationships exists [ 30 ]. The log coefficient of variation ratio ( l n ( C V T C V C ) or lnCVR) is from this perspective a conservative choice as it explicitly normalizes variability by the mean. The lnCVR has the other advantage of being a “master” statistic, a statistic which simplifies to other DiVBTA statistics when no mean-variance relationship exists. The log of the variability ratio ( l n ( s T s C ) or lnVR) is a special DiVBTA case of lnCVR when group means are equal. The F-ratio ( s T 2 s C 2 ), another DiVBTA ratio measure, is a log transformation of VR (ln(F)=2 lnVR). And third, normally distributed ratio DiVBTA measures are preferable when it comes identifying outliers. A log transformation can achieve this goal by reducing skewness which is, for instance, inherent to the F-statistic. Ratio DiVBTA measures are preferable to difference DiVBTA measures because they are scale-invariant, robust across heterogeneous studies, and largely unaffected by errors in publications that mislabel standard errors as standard deviations or fail to label the reported measures of variability as either standard errors or standard deviations. (i) The first criterion needed for a DiVBTA-based decision rule is to establish statistical significance of the DiVBTA. Methods to test the statistical significance of lnCVR are presented in the section on simulation methods and in the Supplementary Materials for lnVR and the F statistic ( S1 Text ). (ii) The second criterion needed for a DiVBTA-based decision rule is to define an outlier, i.e., to construct a DiVBTA tolerance band. Standard trial dynamics (which can lead to statistically significant DiVBTAs) should not be flagged as potentially problematic. Randomization, for instance, will in and of itself lead to 5% of the DiVBTAs to be statistically significant when the type I error rate is set at 5%. Thus, 5% of the trials would be falsely flagged as potentially problematic without setting a DiVBTA tolerance band. Other legitimate trial dynamics such as heterogeneous treatment effects may further increase the proportion of falsely flagged trials as potentially problematic. Construction of a tolerance band reduces such false alarms. The wider the tolerance band, the fewer false alarms, but at the cost of fewer truly problematic trials being captured. A DiVBTA tolerance band can be determined using parametric or non-parametric methods based on the standard trial dynamics of a given clinical response variable (e.g., blood pressure or quality of life) and its statistical characteristics (e.g., absence of mean-variance relationships or ceiling effects). A parametric approach to define tolerance bands which account for standard trial dynamics is to construct the 1-α DiVBTA prediction intervals based on a meta-regression [ 31 ]. Let DiVBTA ijk and sd ijk be the estimator and standard deviation for the i th variance difference or variance ratio between the control arm and treatment arm i, at the j th post-intervention time, and for the k th trial. A meta-analysis of these DiVBTA ijk leads to (1) the mean DiVBTA for the group of randomized trials (M*), (2) the sample estimate of the variance of the true effect sizes (T 2 ), and (3) the variance (V M* ) of the mean effect sizes, M* [ 32 ]. The prediction interval of the differences in variance between trial arms can be calculated as: DiVBTA U , L = M * ± t d f α T 2 + V M * where t α is the t-value corresponding to the desired prediction interval (e.g., α ≈ 0.0027 for a 3-sigma probability). For independent DiVBTAs (i.e., one DiVBTA per clinical trial), the degrees of freedom (df) is 2 less than the number of clinical trials, and V M* can be derived from a model-based estimate [ 32 ]. For correlated DiVBTAs (e.g., trials with more than 2 arms), robust meta-regression methods can be used where df is typically recommended to be calculated using a Satterthwaite approximation, and V M* is estimated empirically using a cluster-robust “sandwich” estimator [ 33 ]. A non-parametric approach to define tolerance bands which account for standard trial dynamics is to derive the median DiVBTA for each included trial (median across all DiVBTA ijk for a given trial k ). The median and the interquartile range of these median DiVBTAs can then be used as basis to construct DiVBTA Tukey inner and outer fences [ 34 ]. Potentially problematic trials can then be defined as statistically significant DiVBTAs which fall outside of these inner and outer Tukey fences. The parametric or non-parametric cutoff values (e.g., α ≈ 0.0027 or inner Tukey fences) determine the width of the DiVBTA tolerance bands which in turn determine the sensitivity and specificity of the decision rule to identify potentially problematic trials. On one hand, defining a narrow tolerance band will increase sensitivity and decrease specificity; standard trial dynamics will frequently be falsely flagged as potentially problematic. On the other hand, setting a wide tolerance band will decrease sensitivity but increase specificity; potentially problematic trials will frequently fail to be flagged for further evaluation. Simulation methods The simulation methods are presented using the ADEMP structure [ 35 ]. Aims. To evaluate the specificity and sensitivity of the proposed DiVBTA-based decision rule for detecting anomalous variance patterns under realistic and manipulated trial scenarios. Specifically, the simulations assess: Specificity (true negative rate): the probability that the test correctly classifies non-problematic trials as non-problematic (i.e., avoids false positives/ false alarms) under standard/legitimate trial dynamics and non-standard/questionable trial dynamics, including (i) proper randomization, (ii) heterogeneous treatment effects (HTE), and (iii) missing-not-at-random (MNAR) dropout. Sensitivity (true positive rate): the probability that the test correctly identifies potentially problematic trials. For the simulations, a specific form of data manipulation (i.e., fraud) was modeled, namely deletion of a fraction of the worst responders in the treatment arm followed by replacement (duplication) of those values with copies of the best-responder observations. How trial data is manipulated is largely unknown, making the relevance of this specific form of simulated fraud data questionable. As the mechanisms of fraud are largely unknown, the sensitivity of any method for identifying potentially problematic trials is difficult to quantify. Proposed decision rules will have utility if it has high specificity (i.e., few false positives), thus allowing it to be a screening tool to rule in potential fraud. A secondary aim is to examine the performance of the proposed decision rule under both large-sample and small-sample scenarios. These aims focus on assessing robustness against false positive conclusions (specificity under realistic trial features that should not trigger alarms) and detection power (sensitivity to detect a targeted fraudulent mechanism), treating a statistically significant DiVBTA outlier as a diagnostic flag for potential issues (unintentional errors or fraud). Data-generating mechanisms. The data-generating mechanisms for the simulations to assess these specific aims were based on five parameter estimates; two parameter estimates characterizing the probability distribution of the clinical response variable (e.g., the mean and the standard deviation of normal distribution) and three parameter estimates derived from a meta-analysis of clinical trials on the clinical response variable (mean (M*), variance of the mean (V M* ), and variance of the true effect size(τ 2 )). Almost all clinical outcomes will be able to be modeled based on 5 parameter estimates, making the provided R-program versatile. These 5 parameter estimates are specific to the (continuous) clinical response under investigation. Clinical outcomes such as patient-reported outcome measures can have ceiling or floor effects (which can create statistically significant DiVBTAs) whereas other clinical outcomes such as blood pressure do not suffer from this effect. Other standard trial dynamics such as the presence of heterogeneous treatment effects (which can create statistically significant DiVBTAs) can also differ depending on the selected clinical outcome. In other words, the 5 parameter estimates are specific to a specific response variable or domain. The 5 parameters in the simulations presented here are diabetes-specific. Data were generated for simulated two-arm randomized trials with a continuous primary outcome (post-intervention HbA1c), modeled parametrically as a gamma distribution. Shape and rate parameters for the gamma distribution were estimated from individual participant data in a pivotal NIH-funded diabetes trial [ 36 ]. The standard trial dynamics in diabetes clinical trials were derived from a meta-analysis of post-intervention HbA1c standard deviations in 175 trials. The between-trial lnCVR heterogeneity was modeled as a baseline layer that existed in every simulated trial by sampling a trial-specific true lnCVR from N(M*, τ 2 ). Three standard (legitimate) trial dynamics were modeled to evaluate specificity: Randomization — Participants randomly assigned to treatment or control arms (assessed under the assumption of no heterogeneous treatment effects and no missing data to isolate randomization effects on lnCVR variability). Heterogeneous treatment effects (HTE) — Systematic variation in treatment response in the intervention arm, modeled via two parameters: (1) proportion of participants responding to treatment (varied 0%–50%), and (2) magnitude of treatment effect among responders (HbA1c treatment effect from −0.2% to −2.0%). Standard and non-standard treatment effect sizes were defined as between -0.2 to −1.4% and −1.4% to −2.0%, respectively. (−1.4% is the 3-sigma bound for the treatment effect sizes in a systematic review of diabetes trials; next section). A systematic review of diabetes trials failed to provide strong evidence in support of heterogeneous treatment effects. Standard and non-standard trial dynamics for the proportion of participants responding to treatment were defined as 0% to 20% and 30% to 50%, respectively. Missing-not-at-random (MNAR) dropout — Dropout probability dependent on unobserved (missing) outcomes modeled as deletion of the worst responders (highest HbA1c values) from the control group. This is an extreme mechanism unlikely in real settings and specified as such to stress-test specificity. Less than 3% of the trials included in Cochrane reviews have a dropout rate of 30% of more [ 37 ] and the fraction of these trials having the extreme mechanism of informative dropout modeled in this study is likely to very small. We classified the simulated dropout rate of 0% to 20% as a standard trial dynamic, and a simulated dropout of 30% to 50% or larger (both with the extreme mechanism described above) as a non-standard trial dynamic. Given the extreme form of dropout modeled this is likely to a conservative definition of standard/non-standard. One fraudulent mechanism was modeled to evaluate sensitivity: Deletion of a fraction of the worst responders (highest HbA1c) in the treatment arm, replaced by duplication of the single best responder (lowest post-treatment HbA1c) observation. The fraction replaced was varied between 50% and 90%. Simulations were run in both large-sample (n = 250 per arm) and small-sample (n = 20 per arm) scenarios, reflecting approximate 95th and 25th quantiles of sample sizes in the motivating diabetes trial meta-analysis. For each scenario/combination, 1,000 independent trials (iterations) were simulated which leads to a Monte Carlo standard error for a sensitivity or specificity of 0.95 of ~0.007. The 3-sigma tolerance bounds for outlier classification were derived from a robust meta-analysis of 175 diabetes trials, specifically using the variance of true effects (τ²) and variance of the mean effect (V M* ) estimated from lnCVR measures in those trials. All simulation code (R), random seeds, and modifiable key parameters are available at https://www.github.com/mhujoel/DIVBTA enabling verification and adaptation to other clinical settings. Estimand. The estimand is the trial-specific true lnCVR. The estimator is the log coefficient of variation ratio (lnCVR), which quantifies relative variability between treatment and control arms. LnCVR is defined as: as l n ( C V T C V C ) + 1 2 ( 1 n T − 1 − 1 n C − 1 ) + 1 2 ( s C 2 n C x ― C 2 − s T 2 n T x ― T 2 ) . The approximate sampling variance of lnCVR is s C 2 n C x ― C 2 + s C 4 2 n C 2 x ― C 4 + n C 2 ( n C − 1 ) 2 + s T 2 n T x ― T 2 + s T 4 2 n T 2 x ― T 4 + n T 2 ( n T − 1 ) 2 . Methods. For each simulated trial (iteration), lnCVR and its sampling variance were computed using the formulas above. The decision rule to classify a trial as potentially problematic consisted of two criteria: lnCVR is statistically significant (p ≤ 0.05) using a two-sample t-test with Welch–Satterthwaite degrees of freedom for parallel-arm trials. lnCVR exceeds the 3-sigma tolerance bound, derived from the meta-analysis of 175 diabetes trials defined as reliable (see next section). No comparator methods were evaluated because comparators such as lnVR or F are biased since mean-variance relationships are present for the clinical outcome selected (Hemoglobin A1c). Performance measures. The primary performance measures were specificity and sensitivity (framed in diagnostic testing terminology, where the “disease” is a problematic trial due to unintentional errors or fraud, and a positive test for “disease” is a statistically significant 3-sigma lnCVR outlier). These measures directly address the aims: high specificity indicates low risk of false alarms under legitimate trial features; high sensitivity indicates good detection of the targeted fraud mechanism. Methods for case-study assessing real-world utility A systematic search was performed of the Cochrane library for systematic reviews with the key words of diabetes and glycaemic ((“The Cochrane database of systematic reviews”[Journal]) AND ((“2010/1/1”[Date – Publication]: “2023/05/23”[Date – Publication]))) AND (diabetes[Title] OR diabetic[Title]) AND (glycaemic OR “glucose-lowering drugs”). Key trial characteristics of the identified trials were abstracted and screened for the availability of post-intervention standard deviations or standard errors. Variance measures derived from statistical methods which assume a homogeneity of variance were excluded. The Carlisle-Stouffer-Fisher p-value, a measure of the plausibility of randomization given the baseline data, was calculated for trials reporting baseline data [ 15 , 38 ]. Risk of bias scores for randomization, blinding, and ascertainment were abstracted from the Cochrane reviews and assigned the values of 0 for high risk, 1 for unclear risk, and 2 for low risk, respectively. Data source (“published data only” or relying on “published and unpublished data”), funding source, sample sizes, number of trials arms, trial duration, effect size, PubMed identification numbers (PMID) were abstracted. Trials without PMID included Ph.D. or Master’s theses, grey literature, clinical trial registries which posted results on the clinical trial registration website, and other data sources not indexed in PubMed. The origin of the trial data was reclassified from “published data only” to “published and unpublished data” for 19 trials included in two Cochrane reviews because it was an organization other than Cochrane (the Institut für Qualität und Wirtschaftlichkeit im Gesundheitswesen or (IQWiG)) which obtained unpublished data for respectively 6 and 13 trials in two Cochrane reviews [ 39 , 40 ]. Trials with no or improbable baseline data were identified (operationalized as a one-sided Carlisle-Stouffer-Fisher p-value which was < 0.025, > 0.975, or missing) and were excluded from the set of trials for defining the tolerance band. Bootstrapping sampling assessed the robustness of excluding trials with no or improbable baseline data on the width of the tolerance band. Parametric and non-parametric tolerance bands were estimated as described in the previous section. Statistical tests reported in the trials were described as questionable when the trial failed to provide a description of the primary statistical test or reported any of the following tests without accommodation for the extreme variance heterogeneity (1) the use of standard Student’s t-test (or equivalent pooled-variance method), (2) reliance on standard ANOVA (including repeated measures ANOVA), (3) use of standard repeated measures ANCOVA (or ANCOVA), (4) reporting of parametric or non-parametric tests “as appropriate” as this assumes readers can retroactively infer the decision rule, or (5) reporting of p-values without reporting the statistical tests used. Countries with a retraction rate above 0.10% in the field of medicine were labeled as having a high retraction ranking [ 41 ]. Results Simulation results The specificity and sensitivity of the proposed decision rule to flag potentially problematic trials are presented for standard and non-standard trial dynamics ( Tables 1–3 for a 3-sigma decision rule, and S1 – S3 Tables for a 4-sigma decision rule). Table 1. Specificity of 3-sigma statistically-significant lnCVR when randomization and heterogeneous treatment effects occur in legitimate trials, by sample size per trial arm. HbA1c Effect size Subgroup with HTE 10% Subgroup with HTE 20% Subgroup with HTE 30%* Subgroup with HTE 40% * Subgroup with HTE 50% * n = 20 n = 250 n = 20 n = 250 n = 20 n = 250 n = 20 n = 250 n = 20 n = 250 −0.0% 91.8% 99.4% 92.9% 99.6% 92.2% 99.5% 93.5% 99.4% 92.8% 99.7% −0.2% 93.7% 99.6% 92.3% 99.5% 93.0% 100.0% 92.8% 99.5% 92.4% 99.3% −0.4% 92.8% 99.1% 92.2% 99.2% 91.0% 99.4% 94.1% 99.4% 90.7% 99.4% −0.6% 91.9% 99.1% 92.0% 99.5% 92.7% 99.9% 92.1% 99.3% 93.9% 99.6% −0.8% 92.1% 99.4% 92.1% 99.8% 93.3% 99.2% 90.8% 98.7% 91.4% 99.5% −1.0% 92.6% 99.5% 92.6% 99.0% 90.8% 99.0% 89.1% 99.3% 91.3% 97.7% −1.2% 93.7% 99.6% 90.5% 99.2% 89.7% 98.4% 87.9% 97.3% 86.5% 97.2% −1.4% * 91.2% 99.1% 90.7% 98.5% 88.3% 98.2% 84.9% 96.6% 82.3% 94.3% −1.6% * 90.0% 99.3% 86.4% 98.3% 85.5% 96.2% 81.5% 94.0% 75.0% 89.6% −1.8% * 90.6% 99.4% 85.8% 96.9% 80.6% 94.1% 73.7% 89.0% 69.4% 81.5% −2.0% * 90.6% 98.2% 84.1% 95.4% 73.5% 89.4% 66.1% 79.0% 61.8% 65.0% Open in a new tab * The systematic review of diabetes trials included in this report suggests that treatment effects larger than a −1.2% and subgroup effects larger than 30% may not reflect standard trial dynamics. Specifically, the 3-sigma bounds of treatment effect sizes ranged from −1.4% to +0.8% and the meta-analysis of lnCVRs did not find convincing evidence of substantial heterogeneity of variances, suggesting that trial where large subgroups (e.g., 30%+) respond very differently in the intervention arm are unlikely. Table 2. Specificity of 3-sigma statistically significant lnCVR when randomization and an extreme form of MNAR dropout (0–50%) occur in trials, by sample size per trial arm. Prevalence of missing-not-at-random dropout level Sample size per clinical trial arm n = 20 n = 250 0 93.3% 99.6% 10% 89.9% 96.6% 20% 83.0% 89.1% 30% * 77.3% 79.0% 40% * 73.5% 63.3% 50% * 69.0% 51.5% Open in a new tab * 30% and larger highly informative dropout rates are not a standard trial dynamic. Table 3. Sensitivity of 3-sigma statistically significant lnCVR for detecting simulated fraud (50–90% worst HbA1c scores in intervention arm replaced by best-responder value), by sample size per trial arm. Proportion ‘worst’ responders deleted and replaced by “best” response. Sample size per clinical trial arm n = 20 n = 250 50% * 35.0% 10.5% 60% * 49.0% 8.7% 70%* 70.4% 10.5% 80% * 89.8% 34.6% 90% * 99.2% 85.4% Open in a new tab * Deleting 50% or more of the individual participant data cannot be considered subtle or minimal data fabrication. Stated differently, these simulation statistics evaluate the sensitivity to detect massive data fraud. Specificity of 3- and 4-sigma significant lnCVRs in the presence of randomization only: Randomization alone leads to few false alarms. In large samples, the specificity of 3- and 4-sigma significant lnCVRs exceeds 99.4% and 99.9%, respectively (very few false alarms). In small samples, the specificity of 3- and 4-sigma lnCVRs exceeds 91.8% and 97.2%, respectively. Specificity of 3- and 4-sigma significant lnCVRs in the presence of randomization and heterogeneous treatment effects: (i) Under standard trial dynamics (heterogeneous treatment effects combined with randomization), 3-sigma significant lnCVRs yielded 99.0% to 99.8% specificity in large-sample settings and 90.5% to 93.7% specificity in small-sample settings. By contrast, 4-sigma significant lnCVRs yielded 99.9% to 100% specificity in large-sample settings and 96.7% to 98.5% specificity in small-sample settings. (ii) Under non-standard trial dynamics (heterogeneous treatment effects combined with randomization), 3-sigma significant lnCVRs yielded 65.0% to 100% specificity in large-sample settings and 61.8% to 94.1% specificity in small-sample settings. By contrast, 4-sigma significant lnCVRs yielded 96.6% to 100% specificity in large-sample settings and 80.1% to 98.6% specificity in small-sample settings. Specificity of 3- and 4-sigma significant lnCVRs in the presence of randomization and missing-not-at-random dropout: (i) Under standard-trial dynamics, when the prevalence of the extreme form of dropout modeled was 20% or less, the specificity exceeded 89.1% in large-sample settings and 83.0% in small-sample settings for 3-sigma significant lnCVRs. For 4-sigma significant lnCVRs, the specificity exceeded 98.2% and 90.9% in large- and small-sample settings, respectively. (ii) Under non-standard trial dynamics, for 3-sigma bounds, the specificity ranged from 69.0% to 77.3% for small-sample settings, and from 51.5% to 79.0% for large-sample settings. Under non-standard trial dynamics, for 4-sigma bounds, the specificity ranged from 69.0% to 84.8% for small-sample settings, and from 82.5% to 95.4% for large-sample settings. Sensitivity of 3- and 4-sigma significant lnCVRs to duplication of responses: The proposed decision rule to flag potentially problematic trials is not sensitive to detecting a 50% duplication of responses in a trial arm. It is only when the rate of data duplication in the intervention arm reaches 80% or higher that the sensitivity becomes larger than 89.8% for small-sample trials. In large-sample settings, a data duplication rate of 90% leads to a sensitivity of 85.4%. The sensitivity decreased when a 4-sigma significant lnCVR tolerance was selected for the decision rule. A case study assessing utility of the proposed decision rule The PRISMA flow diagram shows the systematic selection process which led to a sample of 305 trials, 226 of which with an ability to calculate lnCVR ( S1 Fig ). Trials reporting calculable lnCVRs (n = 226) (versus those where no lnCVRs can be calculated (n = 79)) were more likely not to report funding, to be of shorter duration, to have fewer authors, not to be indexed in PubMed, to have a smaller sample size, to have fewer trial arms, and to originate from a country with a high retraction ranking ( S4 Table ). The 3-sigma prediction interval for lnCVR (i.e., the selected DiVBTA) was −0.54 to 0.47 for 175 trials for trials with plausible baseline data (175 trials;229 lnCVR estimates). Bootstrap sampling showed a modest impact of restricting the estimation of the tolerance bands to trials reporting plausible baseline data ( S5 Table ). Bootstrapping from the available 226 trials led to 3-sigma prediction interval where the lower bound of the 95% confidence interval ranged from −0.67 to −0.46 and the upper bound ranged from 0.39 to 0.65. Nineteen trials reported statistically significant lnCVRs falling outside the parametric 3-sigma prediction interval of −0.54 to 0.47 ( Fig 1 ). These trials, when compared to the 207 trials with either non-significant LnCVRs or LnCVRs not falling outside of the 3-sigma prediction interval, were more likely to report baseline data distributions which are inconsistent with randomization and smaller sample size ( S6 Table ). 18 of 19 trials reported at least one additional potentially problematic feature ( Table 4 ): (1) improbable or no baseline data (n = 7), (2) calculation or data errors in glycemic responses (n = 3), (3) 0% dropout (n = 4), (4) high risk of attrition bias (n = 2) of allocation concealment bias (n = 2), (5) a larger than 3-sigma effect size for HbA1c improvement (n = 1), and (6) reporting of a questionable statistical test (e.g., standard Student t test), or no statistical test (n = 13). Eleven of the 19 trials reported statistically significant lnCVRs falling outside the parametric 4-sigma bounds. Fig 1. Nineteen trials flagged as potentially problematic because their LnCVRs are (i) statistically significant and (ii) outside of the 3-sigma lnCVR prediction intervals estimated based on 175 trials. Open in a new tab The left side of the graph shows the parametric approach to screen for potentially problematic trials; construct the 3-sigma lnCVR prediction interval. The right side of the graph shows the non-parametric approach with the construction of Tukey inner and outer fences to screen for potentially problematic trials. The parametric 3-sigma prediction interval and the Tukey inner fences are remarkably similar. Table 4. Nineteen trials with DIVBTAs outside of the 3-sigma bounds – 7 checks on trustworthiness. First Author Ref. Problematic or no baseline data 1 Calculation/data error in glycemic respons 2 Zero patients lost to follow-up 3 High risk for allocation concealment bias 4 High risk for attrition bias 5 Statistically significant HbA1c effect size> than 3-sigma 6 Questionable statistical test 7 Meschi [ 46 ] ✔ ✔ Homko [ 47 ] ✔ Vincent [ 48 ] ✔ Kiran [ 49 ] ✔ ✔ Macedo [ 50 ] ✔ Agurs-Collins* [ 51 ] Zieve* [ 52 ] ✔ Durán* [ 43 ] ✔ Tsalikis* [ 53 ] ✔ ✔ Schiel* [ 54 ] ✔ ✔ Huang* [ 55 ] ✔ ✔ ✔ ✔ Hirsch* [ 56 ] ✔ Schade* [ 57 ] ✔ Ma [ 58 ] ✔ ✔ ✔ Petrovski [ 59 ] ✔ ✔ Al-Zahrani [ 60 ] ✔ Fang [ 61 ] ✔ ✔ Cohen [ 62 ] ✔ ✔ Dans [ 63 ] ✔ ✔ ✔ Open in a new tab * 8 trials with LnCVR sizes between the 3 and 4-sigma bounds. The checks for other problematic trial features in this table were in part derived from three currently suggested checklists to assess trustworthiness in randomized controlled trials [ 10 – 12 ] 1. Problematic baseline data: no baseline characteristics provided in the publication (leading to Carlisle-Stouffer-Fisher statistics which is missing), or a close to perfect balance for multiple baseline characteristics, or significant/large differences between baseline characteristics (Carlisle-Stouffer-Fisher p-value < 0.025, or > 0.975). 2. A (un)intentional data error or calculation error in the primary outcome (HbA1c): (a) standard deviations (before, after, and of the change) resulted in a correlation less than −1 or greater than +1, or (b) discrepancy in reported before, after, and HbA1c change. 3. Trial duration of 3 months or longer with Cochrane reported 0% dropout. 4. Cochrane reported a “High risk of bias for allocational concealment”. 5. Cochrane reported a “High risk of attrition bias”. 6 HbA1c effect size fell outside the 3-sigma bounds (−1.4% to 0.8%) and was statistically significant. 7. The reported statistical tests was unspecified or failed to specify that the heterogeneity of variances was addressed. The high prevalence of potentially problematic trials reported here (~8%) reflects on calendar years when awareness and scrutiny of problematic trials was low. Included Cochrane reviews were published starting in 2010; all but one of the Cochrane reviews included were published before Cochrane’s 2021 editorial policy on managing potentially problematic studies. Authors of the most recent Cochrane review included in this report may have been unaware of the 2021 editorial policy, in part because it exists separately from the Cochrane Handbook. Discussion Three lines of evidence support the view that trials with unusually large and statistically significant Differences in Variances Between Trial Arms (DiVBTAs) can be flagged as potentially problematic. First, empirical evidence has shown that DiVBTAs are typically small and statistically insignificant [ 21 , 23 – 27 , 42 ], making trials with large and statistically significant DiVBTAs unexpected. Second, simulations demonstrated that the decision rule to flag studies with statistically significant DiVBTA outliers as potentially problematic has high specificity, especially when evaluated under standard trial dynamics. A high specificity implies that a flagged trial can reliably be ruled in as potentially problematic. Third, the real-world case study showed that the effort to calculate and analyze DiVBTAs is worthwhile. About 8% of the trials were flagged even when the definition of a large DiVBTA was defined as a 3-sigma event. These 3 lines of evidence suggest that identifying studies with statistically significant DiVBTA outliers is a worthwhile screening tool and offers a fair and objective flag for potentially problematic studies. An illustrative example demonstrates how the proposed DiVBTA method provides an objective and fair alternative to the subjective impression of unusual variances that flagged problematic trials in the past. The 2009 Boldt et al. trial—pivotal in exposing one of the largest research misconduct scandals to date (22)—was originally questioned in part based on perceived anomalies in reported variances. Our method retrospectively confirms this concern objectively. The proposed methods here flagged the trial as potentially problematic. The DiVBTA was statistically significant (p = 3.1 × 10 ⁻ 13 ), a first criterion to flag a trial. The DiVBTA was also an outlier, the second criterion to flag a trial. This example illustrates that what was once flagged subjectively as potentially problematic can now be detected reliably and transparently using statistical criteria, thereby enhancing fairness and reproducibility in identifying potentially fraudulent or problematic studies. Further real-world examples of confirmed fraudulent trials (e.g., from retracted studies or audits) may reveal how often fraud manifests itself as outlying significant variance differences across trial arms. It is essential to clarify that a statistically significant and unusually large DiVBTA identifies trials as potentially problematic but does not confirm fraud, fabrication, unintentional error, or any specific form of misconduct. Such outliers may occasionally arise from genuine, albeit rare, trial dynamics. As illustrated with the simulations with 3- and 4-sigma outliers as flags for identifying potentially problematic trials, the wider the tolerance band, the more unlikely a flagged trial is to have arisen from rare trial dynamics. DiVBTA outliers may also arise from unintentional errors beyond the control of the trial investigators (e.g., transcription mistakes during the publication process or in intermediary steps between trial publication by the clinical trial team and meta-analysis publication by a separate research team). In one such case, the published trial report presented only non-parametric statistics [ 43 ]. The trial investigators later provided unpublished parametric statistics directly to the meta-analysis team [ 44 ]—statistics now flagged as outliers in this report. When unpublished data are used in this way, it becomes impossible to determine whether an error occurred or, if it did, which team (trial investigators or meta-analyst team) was responsible. These considerations underscore the role of DiVBTA as a screening tool rather than a definitive diagnostic, complementing rather than replacing other data integrity tools. Spotting DIVBTA outliers could become integrated into checklists to assess the trustworthiness of clinical trial reports. Checks of baseline data such as the Carlisle-Stouffer-Fisher method focuses on detecting randomization failures [ 15 ]. DiVBTA methods extends scrutiny to post-randomization outcome data, capturing problematic data issues in post-baseline data, that may not be manifest at baseline. Its compatibility with summary statistics commonly reported in clinical trials makes it practical for implementation. The method’s applicability to non-normal outcomes via lnCVR further enhances its versatility. By integrating parametric (robust meta-regression) and non-parametric (Tukey fences) approaches, and assessing different widths of tolerance bands, a fair and objective framework can be constructed to assess the robustness of labeling trials as potentially problematic. The case study presented suggests that trials with DiVBTA outliers frequently fail to report statistical tests which considers the extreme variance heterogeneity. This absence of an appropriate analysis can lead to biased p-values, risk invalid inferences and consequently expose patients to ineffective/harmful treatments or deny access to beneficial ones. Despite its strengths, the DiVBTA method has limitations. First, detecting DiVBTA becomes challenging when reported standard deviations are imprecise. This can occur when small standard deviations (SD ≈ 1–2) are excessively rounded, an issue that becomes more pronounced when values are reported as standard errors, where rounding can result in a greater loss of information. Second, no meaningful DiVBTA can be computed for trials reporting standard deviations or standard errors calculated under the assumption of a homogeneity of variances. Third, simulation studies showed that small sample sizes increase the risk of false alarms (ruling in a trial as being potentially problematic, when in fact standard trial dynamics may have been responsible). Fourth, in the unlikely scenario in which there are no systematic reviews of clinical trials available, the proposed methods needs to start with a search for trials reporting variability estimates. By offering a scalable, objective method, DiVBTAs could be integrated into the checklists of journal editors, peer reviewers, and meta-analysts to flag trials for further scrutiny, potentially reducing or preventing the impact of problematic studies on the meta-analyses which impact clinical guidelines. Several such workflows are already in place [ 13 , 45 ]. Prospective studies tracking flagged trials for retraction rates could quantify the method’s predictive accuracy, while integration into automated tools or AI-assisted review systems could streamline its application. Collaborative efforts to combine DiVBTA with multimodal integrity checks (e.g., baseline anomalies, effect size outliers) could further improve specificity. Ultimately, the DiVBTA method may offer a robust, transparent approach to bolstering research integrity, addressing the urgent need for validated statistical tools to safeguard the credibility of clinical evidence. Supporting information S1 Text. DiVBTA measures other than lnCVR. (DOCX) pone.0346238.s001.docx (16.4KB, docx) S1 Table. Specificity of 4-sigma statistically significant lnCVR when randomization and heterogeneous treatment effects occur in legitimate trials, by sample size per trial arm. (DOCX) pone.0346238.s002.docx (17.9KB, docx) S2 Table. Specificity of 4-sigma statistically significant lnCVR when randomization and MNAR dropout (0–50%) occur in legitimate trials, by sample size per trial arm. (DOCX) pone.0346238.s003.docx (14.8KB, docx) S3 Table. Sensitivity of 4-sigma statistically significant lnCVR for detecting simulated fraud (50–90% worst HbA1c scores in intervention arm replaced by best-responder value), by sample size per trial arm. (DOCX) pone.0346238.s004.docx (15KB, docx) S1 Fig. Flow diagram of study selection for the meta-analysis of lnCVR estimates. From 58 Cochrane reviews identified via search terms (n = 57) and follow-up (n = 1), 23 were excluded due to absence of HbA1c outcome or being reviews of reviews. 338 trials were identified in the remaining 35 Cochrane reviews, yielding 305 unique trials. 79 trials lacked informative SD estimates and were excluded, leaving 226 trials for the lnCVR meta-analysis. (PNG) pone.0346238.s005.png (151.3KB, png) S4 Table. Characteristics of diabetes intervention trials stratified by reporting of outcome standard deviations. (DOCX) pone.0346238.s006.docx (18KB, docx) S5 Table. Characteristics of diabetes intervention trials stratified by reporting of statistically significant lnCVR outliers. (DOCX) pone.0346238.s007.docx (17.9KB, docx) S6 Table. Assessment of the robustness towards the selection of RCTs for inclusion into the lnCVR meta-analysis. Bootstrap summary of lnCVR effect estimates and 3σ prediction interval bounds. Results from 1000 bootstrap replications (with replacement at the study level). The original values are based on the full dataset (n = 226 studies). The 95% confidence intervals are percentile-based. (DOCX) pone.0346238.s008.docx (13.6KB, docx) Data Availability All relevant data are within the paper and its Supporting Information files. Data, analysis code, and documentation for both the simulations, the case-study, and the Boldt trial presented in the 2nd paragraph of the discussion are available at: https://github.com/mhujoel/DIVBTA . Funding Statement The author(s) received no specific funding for this work. References 1. Van Noorden R, Thompson B. Audio long read: Medicine is plagued by untrustworthy clinical trials. How many studies are faked or flawed?. Nature. 2023;:10.1038/d41586-023-02627–0. doi: 10.1038/d41586-023-02627-0 [ DOI ] [ PubMed ] [ Google Scholar ] 2. Ioannidis JPA. Hundreds of thousands of zombie randomised trials circulate among us. Anaesthesia. 2021;76(4):444–7. doi: 10.1111/anae.15297 [ DOI ] [ PubMed ] [ Google Scholar ] 3. Anonymous. There is a worrying amount of fraud in medical research and a worrying unwillingness to do anything about it. The Economist. 2023. [ Google Scholar ] 4. Blake H, Watt H, Winnett R. Millions at risk in drug fraud scandal: Investigation prosecutors investigate doctor. The Daily Telegraph. 2011. [ Google Scholar ] 5. Subbaraman N. The band of debunkers busting bas scientists; Stanford’s president and a high-profile physicist are among those taken down by a growing wave of volunteers who expose faulty or fraudulent research. Wall Street Journal. 2023. [ Google Scholar ] 6. McKie R. Peer review and scientific publishing “The situation has become appalling’: fake scientific papers push research credibility to crisis point. The Guardian. 2024. [ Google Scholar ] 7. International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH). Guideline for Good Clinical Practice E6(R2). 2016. https://www.ema.europa.eu/en/ich-e6-good-clinical-practice-scientific-guideline [ Google Scholar ] 8. Boughton SL, Wilkinson J, Bero L. When beauty is but skin deep: dealing with problematic studies in systematic reviews. Cochrane Database Syst Rev. 2021;6(6):ED000152. doi: 10.1002/14651858.ED000152 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Hunter KE, Webster AC, Clarke M, Page MJ, Libesman S, Godolphin PJ, et al. Development of a checklist of standard items for processing individual participant data from randomised trials for meta-analyses: Protocol for a modified e-Delphi study. PLoS One. 2022;17(10):e0275893. doi: 10.1371/journal.pone.0275893 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Mol BW, Lai S, Rahim A, Bordewijk EM, Wang R, van Eekelen R, et al. Checklist to assess Trustworthiness in RAndomised Controlled Trials (TRACT checklist): concept proposal and pilot. Res Integr Peer Rev. 2023;8(1):6. doi: 10.1186/s41073-023-00130-8 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Weibel S, Popp M, Reis S, Skoetz N, Garner P, Sydenham E. Identifying and managing problematic trials: A research integrity assessment tool for randomized controlled trials in evidence synthesis. Res Synth Methods. 2023;14(3):357–69. doi: 10.1002/jrsm.1599 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Wilkinson J, Heal C, Antoniou GA, Flemyng E, Ahnström L, Alteri A, et al. Assessing the feasibility and impact of clinical trial trustworthiness checks via an application to Cochrane Reviews: Stage 2 of the INSPECT-SR project. J Clin Epidemiol. 2025;184:111824. doi: 10.1016/j.jclinepi.2025.111824 [ DOI ] [ PubMed ] [ Google Scholar ] 13. Mousa A, Flanagan M, Tay CT, Norman RJ, Costello M, Li W, et al. Research Integrity in Guidelines and evIDence synthesis (RIGID): a framework for assessing research integrity in guideline development and evidence synthesis. EClinicalMedicine. 2024;74:102717. doi: 10.1016/j.eclinm.2024.102717 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Carlisle JB, Dexter F, Pandit JJ, Shafer SL, Yentis SM. Calculating the probability of random sampling for continuous variables in submitted or published randomised controlled trials. Anaesthesia. 2015;70(7):848–58. doi: 10.1111/anae.13126 [ DOI ] [ PubMed ] [ Google Scholar ] 15. Carlisle JB. Data fabrication and other reasons for non-random sampling in 5087 randomised, controlled trials in anaesthetic and general medical journals. Anaesthesia. 2017;72(8):944–52. doi: 10.1111/anae.13938 [ DOI ] [ PubMed ] [ Google Scholar ] 16. Estruch R, Ros E, Salas-Salvadó J, Covas M-I, Corella D, Arós F, et al. Primary prevention of cardiovascular disease with a Mediterranean diet. N Engl J Med. 2013;368(14):1279–90. doi: 10.1056/NEJMoa1200303 [ DOI ] [ PubMed ] [ Google Scholar ] 17. Editor’s note. N Engl J Med. 2018;378(25):2442. [ DOI ] [ PubMed ] [ Google Scholar ] 18. Bolland MJ, Avenell A, Grey A. Statistical techniques to assess publication integrity in groups of randomized trials: a narrative review. J Clin Epidemiol. 2024;170:111365. doi: 10.1016/j.jclinepi.2024.111365 [ DOI ] [ PubMed ] [ Google Scholar ] 19. Hunter KE, Aberoumand M, Libesman S, Sotiropoulos JX, Williams JG, Aagerup J, et al. The Individual Participant Data Integrity Tool for assessing the integrity of randomised trials. Res Synth Methods. 2024;15(6):917–39. doi: 10.1002/jrsm.1738 [ DOI ] [ PubMed ] [ Google Scholar ] 20. Hunter KE, Aberoumand M, Libesman S, Sotiropoulos JX, Williams JG, Li W, et al. Development of the individual participant data integrity tool for assessing the integrity of randomised trials using individual participant data. Res Synth Methods. 2024;15(6):940–9. doi: 10.1002/jrsm.1739 [ DOI ] [ PubMed ] [ Google Scholar ] 21. Mills HL, Higgins JPT, Morris RW, Kessler D, Heron J, Wiles N, et al. Detecting Heterogeneity of Intervention Effects Using Analysis and Meta-analysis of Differences in Variance Between Trial Arms. Epidemiology. 2021;32(6):846–54. doi: 10.1097/EDE.0000000000001401 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Wise J. Boldt: the great pretender. BMJ (Clinical research ed). 2013;346:f1738. [ DOI ] [ PubMed ] [ Google Scholar ] 23. Munkholm K, Winkelbeiner S, Homan P. Individual response to antidepressants for depression in adults-a meta-analysis and simulation study. PLoS One. 2020;15(8):e0237950. doi: 10.1371/journal.pone.0237950 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Plöderl M, Hengartner MP. What are the chances for personalised treatment with antidepressants? Detection of patient-by-treatment interaction with a variance ratio meta-analysis. BMJ Open. 2019;9(12):e034816. doi: 10.1136/bmjopen-2019-034816 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Guo X, McCutcheon RA, Pillinger T, Mizuno Y, Natesan S, Brown K, et al. The magnitude and heterogeneity of antidepressant response in depression: A meta-analysis of over 45,000 patients. J Affect Disord. 2020;276:991–1000. doi: 10.1016/j.jad.2020.07.102 [ DOI ] [ PubMed ] [ Google Scholar ] 26. Volkmann C, Volkmann A, Müller CA. On the treatment effect heterogeneity of antidepressants in major depression: A Bayesian meta-analysis and simulation study. PLoS One. 2020;15(11):e0241497. doi: 10.1371/journal.pone.0241497 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Alsaeid M, Sung S, Bai W, Tam M, Wong YJ, Cortes J, et al. Heterogeneity of treatment response to beta-blockers in the treatment of portal hypertension: A systematic review. Hepatol Commun. 2024;8(2):e0321. doi: 10.1097/HC9.0000000000000321 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Gabelica M, Cavar J, Puljak L. Authors of trials from high-ranking anesthesiology journals were not willing to share raw data. J Clin Epidemiol. 2019;109:111–6. doi: 10.1016/j.jclinepi.2019.01.012 [ DOI ] [ PubMed ] [ Google Scholar ] 29. Veroniki AA, Ashoor HM, Le SPC, Rios P, Stewart LA, Clarke M, et al. Retrieval of individual patient data depended on study characteristics: a randomized controlled trial. J Clin Epidemiol. 2019;113:176–88. doi: 10.1016/j.jclinepi.2019.05.031 [ DOI ] [ PubMed ] [ Google Scholar ] 30. Senior AM, Viechtbauer W, Nakagawa S. Revisiting and expanding the meta-analysis of variation: The log coefficient of variation ratio. Res Synth Methods. 2020;11(4):553–67. doi: 10.1002/jrsm.1423 [ DOI ] [ PubMed ] [ Google Scholar ] 31. Tanner-Smith EE, Tipton E. Robust variance estimation with dependent effect sizes: practical considerations including a software tutorial in Stata and spss. Res Synth Methods. 2014;5(1):13–30. doi: 10.1002/jrsm.1091 [ DOI ] [ PubMed ] [ Google Scholar ] 32. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley. 2021. [ Google Scholar ] 33. Tipton E. Small sample adjustments for robust variance estimation with meta-regression. Psychol Methods. 2015;20(3):375–93. doi: 10.1037/met0000011 [ DOI ] [ PubMed ] [ Google Scholar ] 34. Tukey JW. Exploratory data analysis. Reading, Mass.: Addison-Wesley Pub. Co. 1977. [ Google Scholar ] 35. Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38(11):2074–102. doi: 10.1002/sim.8086 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Engebretson SP, Hyman LG, Michalowicz BS, Schoenfeld ER, Gelato MC, Hou W, et al. The effect of nonsurgical periodontal therapy on hemoglobin A1c levels in persons with type 2 diabetes and chronic periodontitis: a randomized clinical trial. JAMA. 2013;310(23):2523–32. doi: 10.1001/jama.2013.282431 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Babic A, Tokalic R, Amílcar Silva Cunha J, Novak I, Suto J, Vidak M, et al. Assessments of attrition bias in Cochrane systematic reviews are highly inconsistent and thus hindering trial comparability. BMC Med Res Methodol. 2019;19(1):76. doi: 10.1186/s12874-019-0717-9 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Carlisle JB. R code for calculating Carlisle-Stouffer-Fisher statistic. 2023. [ Google Scholar ] 39. Fullerton B, Siebenhofer A, Jeitler K, Horvath K, Semlitsch T, Berghold A, et al. Short-acting insulin analogues versus regular human insulin for adults with type 1 diabetes mellitus. Cochrane Database Syst Rev. 2016;2016(6):CD012161. doi: 10.1002/14651858.CD012161 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Semlitsch T, Engler J, Siebenhofer A, Jeitler K, Berghold A, Horvath K. (Ultra-)long-acting insulin analogues versus NPH insulin (human isophane insulin) for adults with type 2 diabetes mellitus. Cochrane Database Syst Rev. 2020;11(11):CD005613. doi: 10.1002/14651858.CD005613.pub4 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Sebo P, Sebo M. Geographical Disparities in Research Misconduct: Analyzing Retraction Patterns by Country. J Med Internet Res. 2025;27:e65775. doi: 10.2196/65775 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Senior AM, Gosby AK, Lu J, Simpson SJ, Raubenheimer D. Meta-analysis of variance: an illustration comparing the effects of two dietary interventions on variability in weight. Evol Med Public Health. 2016;2016(1):244–55. doi: 10.1093/emph/eow020 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Durán A, Martín P, Runkle I, Pérez N, Abad R, Fernández M, et al. Benefits of self-monitoring blood glucose in the management of new-onset Type 2 diabetes mellitus: the St Carlos Study, a prospective randomized clinic-based interventional study with parallel groups. J Diabetes. 2010;2(3):203–11. doi: 10.1111/j.1753-0407.2010.00081.x [ DOI ] [ PubMed ] [ Google Scholar ] 44. Malanda UL, Welschen LM, Riphagen II, Dekker JM, Nijpels G, Bot SD. Self-monitoring of blood glucose in patients with type 2 diabetes mellitus who are not using insulin. Cochrane Database of Systematic Reviews. 2012;2012(1):Cd005060. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Cochrane. Policy for managing potentially problematic studies: implementation guidance. Chichester, UK: John Wiley & Sons. 2024. https://www.cochranelibrary.com/cdsr/editorial-policies#problematic-studies [ Google Scholar ] 46. Meschi F, Beccaria L, Vanini R, Szulc M, Chiumello G. Short-term subcutaneous insulin infusion in diabetic children. Comparison with three daily insulin injections. Acta Diabetol Lat. 1982;19(4):371–5. doi: 10.1007/BF02629260 [ DOI ] [ PubMed ] [ Google Scholar ] 47. Homko CJ, Santamore WP, Whiteman V, Bower M, Berger P, Geifman-Holtzman O, et al. Use of an internet-based telemedicine system to manage underserved women with gestational diabetes mellitus. Diabetes Technol Ther. 2007;9(3):297–306. doi: 10.1089/dia.2006.0034 [ DOI ] [ PubMed ] [ Google Scholar ] 48. Vincent D, Pasvogel A, Barrera L. A feasibility study of a culturally tailored diabetes intervention for Mexican Americans. Biol Res Nurs. 2007;9(2):130–41. doi: 10.1177/1099800407304980 [ DOI ] [ PubMed ] [ Google Scholar ] 49. Kiran M, Arpak N, Unsal E, Erdoğan MF. The effect of improved periodontal health on metabolic control in type 2 diabetes mellitus. J Clin Periodontol. 2005;32(3):266–72. doi: 10.1111/j.1600-051X.2005.00658.x [ DOI ] [ PubMed ] [ Google Scholar ] 50. Macedo GO, Novaes AB Jr, Souza SLS, Taba M Jr, Palioto DB, Grisi MFM. Additional effects of aPDT on nonsurgical periodontal treatment with doxycycline in type II diabetes: a randomized, controlled clinical trial. Lasers Med Sci. 2014;29(3):881–6. doi: 10.1007/s10103-013-1285-6 [ DOI ] [ PubMed ] [ Google Scholar ] 51. Agurs-Collins TD, Kumanyika SK, Ten Have TR, Adams-Campbell LL. A randomized controlled trial of weight reduction and exercise for diabetes management in older African-American subjects. Diabetes Care. 1997;20(10):1503–11. doi: 10.2337/diacare.20.10.1503 [ DOI ] [ PubMed ] [ Google Scholar ] 52. Zieve FJ, Kalin MF, Schwartz SL, Jones MR, Bailey WL. Results of the glucose-lowering effect of WelChol study (GLOWS): a randomized, double-blind, placebo-controlled pilot study evaluating the effect of colesevelam hydrochloride on glycemic control in subjects with type 2 diabetes. Clin Ther. 2007;29(1):74–83. doi: 10.1016/j.clinthera.2007.01.003 [ DOI ] [ PubMed ] [ Google Scholar ] 53. Tsalikis L, Sakellari D, Dagalis P, Boura P, Konstantinidis A. Effects of doxycycline on clinical, microbiological and immunological parameters in well-controlled diabetes type-2 patients with periodontal disease: a randomized, controlled clinical trial. J Clin Periodontol. 2014;41(10):972–80. doi: 10.1111/jcpe.12287 [ DOI ] [ PubMed ] [ Google Scholar ] 54. Schiel R, Müller UA. Efficacy and treatment satisfaction of once-daily insulin glargine plus one or two oral antidiabetic agents versus continuing premixed human insulin in patients with type 2 diabetes previously on long-term conventional insulin therapy: the Switch pilot study. Exp Clin Endocrinol Diabetes. 2007;115(10):627–33. doi: 10.1055/s-2007-984445 [ DOI ] [ PubMed ] [ Google Scholar ] 55. Huang X, Song L, Li T. Effect of health education and psychosocial intervention on depression in patients with type 2 diabetes. Zhongguo-xinli-weisheng-zazhi. 2002;16(3):149–51. [ Google Scholar ] 56. Hirsch IB, Abelseth J, Bode BW, Fischer JS, Kaufman FR, Mastrototaro J, et al. Sensor-augmented insulin pump therapy: results of the first randomized treat-to-target study. Diabetes Technol Ther. 2008;10(5):377–83. doi: 10.1089/dia.2008.0068 [ DOI ] [ PubMed ] [ Google Scholar ] 57. Schade DS, Mitchell WJ, Griego G. Addition of sulfonylurea to insulin treatment in poorly controlled type II diabetes. A double-blind, randomized clinical trial. JAMA. 1987;257(18):2441–5. [ PubMed ] [ Google Scholar ] 58. Ma W-J, Huang Z-H, Huang B-X, Qi B-H, Zhang Y-J, Xiao B-X, et al. Intensive low-glycaemic-load dietary intervention for the management of glycaemia and serum lipids among women with gestational diabetes: a randomized control trial. Public Health Nutr. 2015;18(8):1506–13. doi: 10.1017/S1368980014001992 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 59. Petrovski G, Dimitrovski C, Bogoev M, Milenkovic T, Ahmeti I, Bitovska I. Is there a difference in pregnancy and glycemic outcome in patients with type 1 diabetes on insulin pump with constant or intermittent glucose monitoring? A pilot study. Diabetes Technol Ther. 2011;13(11):1109–13. doi: 10.1089/dia.2011.0081 [ DOI ] [ PubMed ] [ Google Scholar ] 60. Al-Zahrani MS, Bamshmous SO, Alhassani AA, Al-Sherbini MM. Short-term effects of photodynamic therapy on periodontal status and glycemic control of patients with diabetes. J Periodontol. 2009;80(10):1568–73. doi: 10.1902/jop.2009.090206 [ DOI ] [ PubMed ] [ Google Scholar ] 61. Fang Q, Fang M, Yao Y, Feng S, Yang Y, Xue L, et al. Efficacy and safety of pioglitazone for intervention therapy of impaired glucose regulation [吡格列酮⼲预糖调节受损的疗效和安全性]. Journal of Clinical Research. 2013;30(2):239–42. [ Google Scholar ] 62. Cohen D, Weintrob N, Benzaquen H, Galatzer A, Fayman G, Phillip M. Continuous subcutaneous insulin infusion versus multiple daily injections in adolescents with type I diabetes mellitus: a randomized open crossover trial. J Pediatr Endocrinol Metab. 2003;16(7):1047–50. doi: 10.1515/jpem.2003.16.7.1047 [ DOI ] [ PubMed ] [ Google Scholar ] 63. Dans AML, Villarruz MVC, Jimeno CA, Javelosa MAU, Chua J, Bautista R, et al. The effect of Momordica charantia capsule preparation on glycemic control in type 2 diabetes mellitus needs further studies. J Clin Epidemiol. 2007;60(6):554–9. doi: 10.1016/j.jclinepi.2006.07.009 [ DOI ] [ PubMed ] [ Google Scholar ] PLoS One. doi: 10.1371/journal.pone.0346238.r001 Decision Letter 0 Robin Haunschild Robin Haunschild Academic Editor Find articles by Robin Haunschild Author information Copyright and License information Roles Robin Haunschild : Academic Editor © 2026 Robin Haunschild This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 29 Jul 2025 Dear Dr. Hujoel, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Sep 12 2025 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at [email protected] . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .... We look forward to receiving your revised manuscript. Kind regards, Robin Haunschild Academic Editor PLOS ONE Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1.Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf 2. Please note that your Data Availability Statement is currently missing the repository name and/or the DOI/accession number of each dataset OR a direct link to access each database. If your manuscript is accepted for publication, you will be asked to provide these details on a very short timeline. We therefore suggest that you provide this information now, though we will not hold up the peer review process if you are unable. 3. We note you have included a table to which you do not refer in the text of your manuscript. Please ensure that you refer to Table 3 in your text; if accepted, production will need this reference to link the reader to the Table. 4. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? Reviewer #1: Partly Reviewer #2: Yes ********** 2. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #1: No Reviewer #2: Yes ********** 3. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: No Reviewer #2: No ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes ********** Reviewer #1: The authors propose an estimator for implausible study outcomes which is a commendable initiative. The approach is beyond comparing means and less subjective. Nevertheless, the manuscript is of a very general nature. The applied setting is intuitive, however, the findings might be supported by a simulation study with all underlying assumptions being varied (treatment effects, dropout rates, the set of trustworthy trials, etc.). Major - page 4, 1st paragraph: respective DIVBTA estimators should be mentioned in more detail. The paper could also benefit from benchmarking results of alternative estimates in the applied setting - page 4, 2nd paragraph: “The set of randomized controlled trials at the basis of clinical guidelines” and, in particular, the definition of a “set of trustworthy trials” should be extended. So far authors remain vague on this definition albeit it might have crucial relevance for application of DIVBTAs. Which criteria qualify to be part of the set, placebo-controlled trials, trials with active control (head-to-head comparisons), efficacy/safety trials, sample size of trials? When is the size of the set sufficient? How to flag a trial as untrustworthy in absence of such a set? - The first assumption “First, there is no treatment effect heterogeneity – the treatment effect is the same for every individual” I'm not sure if this is mentioned here as a more general or specific assumption for the estimator? This cannot apply to placebo-controlled studies. If otherwise, please comment. - Please specify in more detail the second assumption “Second, the error of the outcome does not depend on the treatment or the outcome”. How is error defined? Is that a plausible assumption for all kinds of outcomes? For example, single- or double-bounded outcomes should be considered, please see: https://doi.org/10.3102/1076998610396895 . Variances should also change, in particular, when a considerable proportion of patients achieve remission. - Please also be more detailed about randomization. The phrasing “A particularly unlucky randomization may lead to …” appears not well aligned with common terminology. - The phrasing “the variability of DIVTBA can be impacted by the noise created by participant dropout” is unclear. In which form do dropouts add noise? To this reviewers experience, they introduce imbalanced patient characteristics and missingness. In case of selective dropout this leads to upwards/downwards biased results. - The whole paragraph at the end of page 5/1st page 6 is unclear. - The section “Interpretation of the DIVBTA variance observed in a set of trustworthy trials” might benefit from a table summarizing the assumptions for DIVBTA such as expectations regarding the mean/variability in trustworthy vs. untrustworthy trials, impact of dropout, interaction, by chance, trial size, randomization, and further biases. Minor Abstract: Please correct typos “. to illustrate this approach, we assessed the DIVBTAs for Hemoglobin A1cin a systematic sample” Please revise the use and setting of references with/without leading spaces. Reviewer #2: I very much enjoyed reading your manuscript and look forward to seeing it published soon. With kind regards, Emma Sydenham (Senior Editor, Cochrane Database of Systematic Reviews) Major comments: 1) Concerning point 3 above 'Have the authors made all data underlying the findings of their manuscript fully available' I have answered no as the list of reviews identified in 'II. Problematic differences in variance between trial arms (DIVBTA): Application' is not given. It would be helpful if the thirty-five Cochrane reviews that were identified and reported HbA1c could be referenced in this section, or listed in a table. Minor comments: 1) The third sentence of the introduction does not follow. I suggest re-phrasing based on the following comment, and those below. While it is true that the unreliable trials have permeated clinical guidelines, I disagree that this has contributed to a growing crisis in research integrity. Rather, unreliable trials have permeated guidelines because research integrity and regulatory compliance criteria were not previously included in meta-analysis and guideline development processes. This is partly due to the information not being requested or made available by publishers in previous decades, and the issue not having been taken up by a working group in evidence based medicine years ago (i.e. in the early 2000s). I don't think there is a growing crisis in research integrity - I think research integrity was never raised as an issue until recently. Research integrity issues were highlighted during the COVID pandemic in particular, in 2020, which was some 30 years after evidence based medicine started taking hold as an academic discipline. I think the possibility of tackling research integrity issues in evidence based medicine has never been better than at present, because now there are a handful of research integrity checklists, policies, and methods (such as this paper) which are available for researchers to use which didn't previously exist. Also, looking back ten years, the critics of research integrity were not morally wrong or factually incorrect in their views. (Here I refer to an old debate ( https://www.bmj.com/content/350/bmj.h2463 ) and a more recent commentary ( https://www.jclinepi.com/article/S0895-4356(25)00003-4/pdf ).) Clinical trials are highly regulated and so it should indeed be safe to assume that they were conducted according to the relevant domestic and international regulations. It should be the case that study reports accurately reflect lawfully collected clinical trial data. One of a number of problems is that new computing technologies have been developed which make it easy to generate fake data, and these new technologies have become more accessible over the last ten years. Multiple issues relating to computer generated data, changes in medical journal publishing, trial registration, international harmonisation of clinical trial conduct and electronic data collection exist in parallel and are constantly evolving. The more recent acknowledgement that research integrity can be a problem, and the use of solutions, are making evidence based clinical guidelines safer. Here is one such example: https://www.cochrane.org/about-us/news/cochrane-launches-new-feature-identify-retracted-publications So I would encourage you to reconsider using the phrase 'a growing crisis in research integrity'. 2) The following sentence 'A 2021 Cochrane editorial...' could be expanded. This editorial by Boughton, Wilkinson and Bero was the announcement of Cochrane's Editorial Policy on Managing Potentially Problematic Studies, which was developed over 5 years with broad consultation. I assume I have been invited to comment on your work as I am listed as a member of the Policy advisory committee. Unfortunately the Policy document does not have its own DOI, because it is included in an online policy manual and website, so the editorial is often referenced instead of the policy implementation guidance web link. The editorial you reference was only one means of disseminating the policy, there are also training materials produced by Cochrane, there was a popular blog post by Richard Smith 'Time to assume that health research is fraudulent until proven otherwise?' https://blogs.bmj.com/bmj/2021/07/05/time-to-assume-that-health-research-is-fraudulent-until-proved-otherwise/ , among other resources and conference round table discussions. However, rather than referring to prior calls for action, you could present the work as a contribution to the other new developments which have followed. For example, a research integrity tool with an associated R package was developed by Hunter et al: https://doi.org/10.1002%2Fjrsm.1738 and there is also an R package for the statistical checks of the REAPPRAISED checklist: https://reappraised.wordpress.com/2023/03/28/the-reappraised-r-package/ Ideally this piece of work will inform new R packages, which are currently being used to automate research integrity checks in the field of Data Science. Are you aware that an R package has already been developed for this analysis? It might be worth referencing, too: https://github.com/harrietlmills/DetectingDifferencesInVariance 3) You have chosen to examine a cohort of Cochrane reviews for your example. The analysis that you have done for these trials is good, I just think there is a slight problem in the way you have explained the rationale for selecting these trials which should be reconsidered. The search for reviews starts in 2010 which is prior to the development of Cochrane's Editorial Policy for Managing Potentially Problematic Studies. So the trials that are included in your analysis are not ones which went through a Trustworthiness Screening Tool (such as: Identifying and handling potentially untrustworthy trials – Trustworthiness Screening Tool (TST). Developed by the Cochrane Pregnancy and Childbirth Group. Alfirevic Z, Kellie FJ, Weeks J, Stewart F, Jones L, Hampson L, on behalf of the Pregnancy and Childbirth Editorial Board and 10.1002/cesm.12037). The content of your work is fine and should be published, I just think the way you have framed the issue of the reviews being trustworthy because they are Cochrane reviews is not quite right if those reviews were published prior to Cochrane's policy and the reviews didn't incorporate a trustworthiness screening tool or another research integrity assessment tool or strategy. 4) Further to point 3) with regards to the paragraph 'Trustworthy trials as a source of DIVBTA estimates' (p.4), 'By focusing on a set of trials viewed as trustworthy by an authoritative organization' (p.6), and 'A non-parametric approach to define unusual DIVBTAs...' (p.7) it's worth noting that not all Cochrane authors are aware of the Policy on Managing Potentially Problematic Trials. This is partly due to the fact that the Cochrane Handbook is long, and the Policy is described in a separate online manual covering Cochrane's editorial policies. Even at present (July 2025), not every trial included in a Cochrane review is assessed using a research integrity tool. I have no doubt that will come in the future, but we aren't there yet and very few trials were statistically checked in the past. Furthermore, teaching about meta-analysis varies in scope and research integrity may not be included in the curriculum. Research integrity was not commonly taught prior to 2020, until the COVID pandemic brought issues concerning research integrity into public discourse. 5) I encourage you to re-phrase the sentences 'A non-parametric approach to define unusual DIVBTAs is to derive the median DIVBTA for each trustworthy trial' (p.7) and 'The two criteria for defining the set of trustworthy clinical trials were...' (p.10). I understand what you mean; however, I would like to point out that the trials you have selected for analysis have not been through a formal trustworthiness assessment prior to publication of the Cochrane reviews. The only Cochrane editorial group that routinely assessed all studies for trustworthiness prior to inclusion in the review was the Pregnancy and Childbirth Group, which developed and applied their Cochrane PCG-TST tool in the reviews for which they held editorial responsibility. (Note, this has all changed now as there is a 'new' Cochrane Central Editorial Service.) So while your work presented in this paper is true and valid in terms of the actual analysis, it is not the case that the studies included in the 35 Cochrane reviews were assessed as being trustworthy to start with. No trustworthiness assessment was done, apart from possibly filtering out the retracted studies as part of the searching procedures (as per the Mandatory standard C48 Examining Errata, Cochrane Handbook section 4.4.6 and the technical supplement section 3.9). The work you have done and presented in this paper is good and should be published, I'm just not sure you should say that the trials are trustworthy if they have not been through a trustworthiness screening tool or a research integrity checklist (such as PCG-TST, the Reappraised checklist (Grey et al), RIA (Weibel et al), TRACT (Mol et al), or a procedure as described in the RIGID Framework (Mousa et al)). These research integrity tools hadn't been developed in 2010 which is the start date of your search. 6) In the first paragraph of the Discussion, you could make reference to some other work in this area. For example, in order to understand the errors identified, the RIGID Framework and the Cochrane Policy recommend contacting trial authors to request clarification of the reasons for possible errors. RIGID Framework: https://www.thelancet.com/pdfs/journals/eclinm/PIIS2589-5370(24)00296-7.pdf Cochrane Policy: https://www.cochranelibrary.com/cdsr/editorial-policies/problematic-studies-implementation-guidance Very sadly there have been a few cases internationally of researchers taking their own life after their work was found to have problems, and one of the reasons for liaising with them is to make them aware their work is under review and to give them an opportunity to explain any problems that might be identified. (You can look up the case of Yoshiki Sasai, for example, which was highly publicised. However, there are other cases which have received no publicity so the actual number of cases is slightly higher than one can find in a web search.) 7) Your example is in diabetes research, but you could also reference where a similar analysis has been used in other areas of medicine, such as: https://doi.org/10.1097/EDE.0000000000001401 and https://doi.org/10.1002/bimj.202200116 among others. ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our Privacy Policy ..--> Reviewer #1: Yes: Dr. Adrian RichterDr. Adrian RichterDr. Adrian RichterDr. Adrian Richter Reviewer #2: Yes: Emma Sydenham (Senior Editor, Cochrane Database of Systematic Reviews)Emma Sydenham (Senior Editor, Cochrane Database of Systematic Reviews)Emma Sydenham (Senior Editor, Cochrane Database of Systematic Reviews)Emma Sydenham (Senior Editor, Cochrane Database of Systematic Reviews) ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/ . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at . PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at [email protected] . Please note that Supporting Information files do not need this step.. Please note that Supporting Information files do not need this step.. Please note that Supporting Information files do not need this step.. Please note that Supporting Information files do not need this step. PLoS One. 2026 Apr 15;21(4):e0346238. doi: 10.1371/journal.pone.0346238.r002 Author response to Decision Letter 1 Article notes Copyright and License information Collection date 2026. PMC Copyright notice 30 Nov 2025 Response to reviewers document was uploaded. Attachment Submitted filename: Response to reviewers.docx pone.0346238.s010.docx (281.7KB, docx) PLoS One. doi: 10.1371/journal.pone.0346238.r003 Decision Letter 1 Robin Haunschild Robin Haunschild Academic Editor Find articles by Robin Haunschild Author information Copyright and License information Roles Robin Haunschild : Academic Editor © 2026 Robin Haunschild This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 13 Jan 2026 Dear Dr. Hujoel, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Feb 27 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at [email protected] . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .... We look forward to receiving your revised manuscript. Kind regards, Robin Haunschild Academic Editor PLOS One Journal Requirements: If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice. [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author Reviewer #1: (No Response) Reviewer #2: All comments have been addressed ********** 2. Is the manuscript technically sound, and do the data support the conclusions??> Reviewer #1: Yes Reviewer #2: Yes ********** 3. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #1: I Don't Know Reviewer #2: Yes ********** 4. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes ********** Reviewer #1: The authors have provided a very thorough and carefully considered revision, which has substantially improved the manuscript. The simulation studies convincingly underscore the relevance of an estimator that addresses implausible variance differences between trial arms and even provide novel insights in relation to previous methodological work, while also pointing to potentially problematic randomized controlled trials. The discussion of the proposed method is well balanced and appropriately acknowledges its limitations. It would therefore be highly desirable for this method to be considered in applied settings, for example in meta-analyses, to support the assessment of the underlying evidence base and its credibility, ideally via an openly available R package rather than a Git repository alone. However, as the manuscript has undergone substantial revisions, there are remaining issues related to structure and naming conventions that at times make it difficult to follow. In particular, the manuscript appears to pursue four distinct and only loosely connected objectives, each addressed in a separate section: (1) identification of statistically significant DiVBTA outliers, (2) simulation studies, (3) a case study on a landmark fraud case, and (4) a case study on clinical trials included in systematic reviews on diabetes management. As a consequence, the manuscript deviates from the conventional IMRaD structure. For example, methods are introduced in both Sections 1 and 2, and the latter simultaneously presents results. Moreover, the two sections appear somewhat disconnected, as estimates or methods introduced in Section 1 do not clearly reappear in Section 2. In this context, closer adherence to established reporting guidelines would substantially strengthen the manuscript. In particular, the recommendations for simulation studies proposed by Boulesteix et al. (1) and Morris et al. (2) (ADEMP framework, endorsed by the STRATOS initiative) would provide a helpful structure. At present, key elements required for transparent reporting of simulation studies are missing or insufficiently described, including: - Why are results presented only for LnCVRs? - How was the simulation setup defined with respect to the number of iterations, software packages used, distributional assumptions, and used seeds? - How were noise levels and signal-to-noise ratios handled? - Why was n = 20 chosen for small-sample trials, and is n = 250 a realistic or appropriate choice for large trials? - Is the chosen range of treatment effect variability (0.2% to 1.4%) plausible for small randomized controlled trials? In addition, clear performance metrics such as sensitivity and specificity are currently missing. Potential users of the proposed method need to understand both the risk of false-positive findings and the probability that an untrustworthy trial remains undetected. Minor comments - Inconsistent naming conventions are used, for example: “Large samples (250/trial arm) and HTE” versus “Small samples (n = 20 per group) and HTE”. - The correct reference to Tukey’s original work should be included and how fences or spread are defined should be added. (1) Boulesteix A-L, Groenwold RH, Abrahamowicz M, Binder H, Briel M, Hornung R, Morris TP, Rahnenführer J, Sauerbrei W. Introduction to statistical simulations in health research. BMJ Open. 2020;10(12):e039921. https://doi.org/10.1136/bmjopen-2020-039921 . (2) Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Statistics in Medicine. 2019;38(11):2074-102. https://doi.org/10.1002/sim.8086 . Reviewer #2: (No Response) ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our Privacy Policy ..--> Reviewer #1: Yes: Dr. Adrian RichterDr. Adrian RichterDr. Adrian RichterDr. Adrian Richter Reviewer #2: Yes: Emma SydenhamEmma SydenhamEmma SydenhamEmma Sydenham ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation . NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. PLoS One. 2026 Apr 15;21(4):e0346238. doi: 10.1371/journal.pone.0346238.r004 Author response to Decision Letter 2 Article notes Copyright and License information Collection date 2026. PMC Copyright notice 10 Mar 2026 Response to reviewer was uploaded as a word document Attachment Submitted filename: response to reviewers_PH.docx pone.0346238.s011.docx (33KB, docx) PLoS One. doi: 10.1371/journal.pone.0346238.r005 Decision Letter 2 Robin Haunschild Robin Haunschild Academic Editor Find articles by Robin Haunschild Author information Copyright and License information Roles Robin Haunschild : Academic Editor © 2026 Robin Haunschild This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 17 Mar 2026 Unusual Outcome Variances as a Method to Identify Potentially Problematic Clinical Trials PONE-D-25-25920R2 Dear Dr. Hujoel, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support .... If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact [email protected]. Kind regards, Robin Haunschild Academic Editor PLOS One Additional Editor Comments (optional): Reviewers' comments: PLoS One. doi: 10.1371/journal.pone.0346238.r006 Acceptance letter Robin Haunschild Robin Haunschild Academic Editor Find articles by Robin Haunschild Author information Copyright and License information Roles Robin Haunschild : Academic Editor © 2026 Robin Haunschild This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice PONE-D-25-25920R2 PLOS One Dear Dr. Hujoel, I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team. At this stage, our production department will prepare your paper for publication. This includes ensuring the following: * All references, tables, and figures are properly cited * All relevant supporting information is included in the manuscript submission, * There are no issues that prevent the paper from being properly typeset You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps. Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact [email protected]. You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing . If we can help with anything else, please email us at [email protected]. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Robin Haunschild Academic Editor PLOS One Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials S1 Text. DiVBTA measures other than lnCVR. (DOCX) pone.0346238.s001.docx (16.4KB, docx) S1 Table. Specificity of 4-sigma statistically significant lnCVR when randomization and heterogeneous treatment effects occur in legitimate trials, by sample size per trial arm. (DOCX) pone.0346238.s002.docx (17.9KB, docx) S2 Table. Specificity of 4-sigma statistically significant lnCVR when randomization and MNAR dropout (0–50%) occur in legitimate trials, by sample size per trial arm. (DOCX) pone.0346238.s003.docx (14.8KB, docx) S3 Table. Sensitivity of 4-sigma statistically significant lnCVR for detecting simulated fraud (50–90% worst HbA1c scores in intervention arm replaced by best-responder value), by sample size per trial arm. (DOCX) pone.0346238.s004.docx (15KB, docx) S1 Fig. Flow diagram of study selection for the meta-analysis of lnCVR estimates. From 58 Cochrane reviews identified via search terms (n = 57) and follow-up (n = 1), 23 were excluded due to absence of HbA1c outcome or being reviews of reviews. 338 trials were identified in the remaining 35 Cochrane reviews, yielding 305 unique trials. 79 trials lacked informative SD estimates and were excluded, leaving 226 trials for the lnCVR meta-analysis. (PNG) pone.0346238.s005.png (151.3KB, png) S4 Table. Characteristics of diabetes intervention trials stratified by reporting of outcome standard deviations. (DOCX) pone.0346238.s006.docx (18KB, docx) S5 Table. Characteristics of diabetes intervention trials stratified by reporting of statistically significant lnCVR outliers. (DOCX) pone.0346238.s007.docx (17.9KB, docx) S6 Table. Assessment of the robustness towards the selection of RCTs for inclusion into the lnCVR meta-analysis. Bootstrap summary of lnCVR effect estimates and 3σ prediction interval bounds. Results from 1000 bootstrap replications (with replacement at the study level). The original values are based on the full dataset (n = 226 studies). The 95% confidence intervals are percentile-based. (DOCX) pone.0346238.s008.docx (13.6KB, docx) Attachment Submitted filename: Response to reviewers.docx pone.0346238.s010.docx (281.7KB, docx) Attachment Submitted filename: response to reviewers_PH.docx pone.0346238.s011.docx (33KB, docx) Data Availability Statement All relevant data are within the paper and its Supporting Information files. Data, analysis code, and documentation for both the simulations, the case-study, and the Boldt trial presented in the 2nd paragraph of the discussion are available at: https://github.com/mhujoel/DIVBTA . Articles from PLOS One are provided here courtesy of PLOS ACTIONS View on publisher site PDF (719.5 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top