Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice PLoS One . 2026 Apr 10;21(4):e0346038. doi: 10.1371/journal.pone.0346038 Search in PMC Search in PubMed View in NLM Catalog Add to search Integrative bioinformatics and machine learning identify shared molecular mechanisms and diagnostic biomarkers between Helicobacter pylori infection and atrial fibrillation Aojian Deng Aojian Deng 1 Department of Gastroenterology, The Third Xiangya Hospital, Central South University, Changsha, Hunan, Chinas Funding acquisition, Investigation, Resources, Software, Supervision, Validation, Visualization, Writing – original draft Find articles by Aojian Deng 1 , Wei Wang Wei Wang 2 Department of Fourth Internal Medicine, The Third People's Hospital Health Care Group of Cixi, Ningbo, China Conceptualization, Data curation, Formal analysis, Investigation, Project administration, Resources, Visualization, Writing – original draft, Writing – review & editing Find articles by Wei Wang 2, * Editor: Tomasz W Kaminski 3 Author information Article notes Copyright and License information 1 Department of Gastroenterology, The Third Xiangya Hospital, Central South University, Changsha, Hunan, Chinas 2 Department of Fourth Internal Medicine, The Third People's Hospital Health Care Group of Cixi, Ningbo, China 3 Versiti Blood Research Institute, UNITED STATES OF AMERICA ✉ * E-mail: [email protected] Competing Interests: The authors have declared that no competing interests exists. Roles Aojian Deng : Funding acquisition, Investigation, Resources, Software, Supervision, Validation, Visualization, Writing – original draft Wei Wang : Conceptualization, Data curation, Formal analysis, Investigation, Project administration, Resources, Visualization, Writing – original draft, Writing – review & editing Tomasz W Kaminski : Editor Received 2025 Nov 19; Accepted 2026 Mar 14; Collection date 2026. © 2026 Deng, Wang This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice PMCID: PMC13068215 PMID: 41961800 Abstract Background Helicobacter pylori ( H. pylori ) infection and atrial fibrillation(AF) are major global health concerns. Emerging evidence has suggested a potentially chronic inflammation-mediated link between them, but the shared genetic mechanisms remain unclear. Methods We analyzed multidataset gene expression profiles from the Gene Expression Omnibus (GEO) database. Differential expression analysis, weighted gene co-expression network analysis (WGCNA), functional enrichment, and machine learning were employed to identify common genes, pathways, and diagnostic biomarkers. Protein-protein interaction (PPI) networks, drug-gene analysis, and molecular docking were used to identify hub genes and potential therapeutics. Results We identified 73 common differentially expressed genes (DEGs) between H. pylori infection and AF, which were predominantly enriched in immune-related processes including leukocyte activation, neutrophil migration, and myeloid cell-mediated immunity. Machine learning identified 15 and 23 key feature genes for H. pylori and AF, respectively, with S100A8 emerging as a shared diagnostic biomarker. Ten hub genes including TYROBP, ITGB2, and SPI1, were identified from the PPI network. Drug repositioning analysis suggested retinoic acid, indirubin, and ropivacaine as candidate therapeutics targeting these key hub genes. Conclusion Our integrative analysis highlights the central role of immune-inflammatory pathways in linking H. pylori infection to AF. We propose S100A8 and other identified hub genes as potential biomarkers and therapeutic targets. The predicted candidate therapeutics, particularly retinoic acid, may offer novel avenues for intervention, warranting further experimental validation. Introduction Helicobacter pylori ( H. pylori ) is a gram-negative, microaerophilic, spiral-shaped bacterium that colonizes the human gastric mucosa and represents one of the most prevalent chronic bacterial infections worldwide [ 1 , 2 ] . It is estimated that more than half of the global population is infected, with acquisition typically occurring during childhood. Without treatment, the infection often persists through life [ 2 , 3 ] . H. pylori is a primary etiological agent of chronic active gastritis and is strongly associated with peptic ulcer disease—which causes approximately 90% of duodenal ulcers and 70–80% of gastric ulcers—as well as gastric mucosa-associated lymphoid tissue (MALT) lymphoma [ 4 ] . Beyond its direct gastrointestinal manifestations, accumulating evidence hsa indicated potential links between H. pylori infection and various extragastric disorders, including unexplained iron-deficiency anemia, idiopathic thrombocytopenic purpura, and cardiometabolic diseases. This broad influence positions H. pylori as a systemic pathogen capable of affecting multiple organ systems [ 5 ] . Atrial fibrillation (AF) is the most prevalent sustained cardiac arrhythmia and id characterized by rapid and disorganized atrial electrical activity resulting in ineffective atrial contraction. Clinical presentation varies widely; patients may experience palpitations, fatigue, and dyspnea, or remain entirely asymptomatic [ 6 , 7 ] . The global prevalence of AF is steadily increasing, affecting tens of millions of adults and rising significantly with age, imposing a substantial and growing public health burden [ 8 ] . The clinical significance of AF extends beyond its symptomatic burden. Its most critical complication is intracardiac thrombus formation secondary to atrial blood stasis; subsequent embolization can lead to systemic embolism, particularly ischemic stroke, increasing the risk of stroke five-fold [ 9 , 10 ] . Additionally, AF is independently associated with heart failure, cognitive impairment, and elevated all-cause mortality [ 11 – 13 ] . The pathophysiology of AF is multifactorial. The established risk factors include advanced age, hypertension, heart failure, valvular heart disease, diabetes, obesity, and sleep apnea. The initiation and perpetuation of AF are driven primarily by electrophysiological and structural remodeling processes, including atrial fibrosis, ion channel dysfunction, inflammation, and oxidative stress [ 13 ] . Notably, systemic inflammation has been recognized as an independent contributor to AF pathogenesis. Inflammatory mediators can promote atrial fibrosis and alter electrophysiological properties, thereby creating a substrate conducive to AF initiation and maintenance [ 14 ] . Emerging clinical evidence has suggested a potential association between H. pylori infection and an increased risk of AF, implicating that it is involved in the development of arrhythmia. However, the specific underlying mechanisms remain insufficiently elucidated [ 15 – 17 ] . In this study, we addressed this gap by leveraging public sequencing data for H. pylori and AF from the GEO database. Through integrated bioinformatic analyses, machine learning approaches, and molecular prediction methods, we aimed to identify shared hub genes, elucidate common pathways, and explore potential therapeutic molecules associating these two prevalent conditions. Materials and methods Data acquisition and preprocessing All gene expression data were sourced from the GEO database. Four datasets were included for AF: GSE41177 and GSE176166 constituted the discovery cohort, while GSE108660 and GSE128188 formed the independent validation cohort [ 18 , 19 ] . Five datasets were selected for H. pylori infection: GSE27411 , GSE60662 , and GSE233973 comprised the discovery cohort, and GSE5081 and GSE60427 serving as the validation cohort [ 20 – 24 ] . Detailed information for each dataset, including platform, sample sizes (disease vs. control), and source, is summarized in Table 1 . Raw data (CEL files) were processed using the affy or oligo R packages for background correction and normalization. For the discovery cohorts within each condition (H. pylori infection and AF), the sva R package was used to perform ComBat batch effect correction, integrating multiple datasets into a single expression matrix for subsequent differential expression and WGCNA analyses 25. Validation cohorts were processed individually without batch correction to ensure independent assessment. Table 1. Basic information of the GEO datasets used in the study. GSE series Disease Samples Platform Group GSE41177 AF 32 AF samples and 6 healthy controls GPL570 Discovery cohort GSE176166 AF 3 AF samples and 4 health controls GPL23126 Discovery cohort GSE108660 AF 5 AF samples and 5 healthy controls GPL19612 Validation cohort GSE128188 AF 5 AF samples and 5 healthy controls GPL18573 Validation cohort GSE27411 Hp 6 Hp samples and 6 healthy controls GPL6255 Discovery cohort GSE60662 Hp 12 Hp samples and 4 healthy controls GPL13497 Discovery cohort GSE233973 Hp 13 Hp samples and 9 healthy controls GPL21185 Discovery cohort GSE5081 Hp 16 Hp samples and 16 healthy controls GPL570 Validation cohort GSE60427 Hp 24 Hp samples and 8 healthy controls GPL17077 Validation cohort Open in a new tab Screening and functional enrichment of DEGs DEGs between disease and control samples in the discovery cohorts were identified separately for H. pylori infection and AF using the limma R package. Datasets related to AF and H. pylori infection were analyzed separately to identify both up- and down-regulated DEGs for each condition [ 24 ]. Genes with an adjusted p-value (adj.P.Val) < 0.05 and |log2 fold change (FC)| > 0.5 were considered significant. Commonly co-expressed DEGs between AF and H. pylori infection were extracted on the basis of their overlapping expression patterns (shared up- or down-regulation). Functional enrichment analysis of these common DEGs was performed using the org.Hs.e.g.,db and clusterProfiler R package for Gene Ontology (GO) terms and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. WGCNA WGCNA was employed on the batch-corrected discovery cohort expression matrices to construct scale-free co-expression networks. The soft-thresholding power was chosen based of the criterion of approximate scale-free topology. A topological overlap matrix (TOM) was constructed, and genes were clustered into modules using dynamic tree cutting. Module eigengenes (MEs) were computed, and their correlations with the clinical trait (disease status) were assessed. Modules with the highest absolute correlation (|r| > 0.4, p < 0.05) were selected as trait-relevant. Genes with high module membership (MM > 0.8) and gene significance (GS > 0.2) within these key modules were considered core genes. Application of multiple machine-learning (ML) methods Data Preparation and Modeling Strategy: The union set of genes from common Differentially expressed genes DEGs and shared WGCNA module genes served as the initial feature space. For each condition ( H. pylori and AF), the batch-corrected discovery cohort was used for model training and feature selection. We employed more than 20 ML algorithms, including Random Forest (RF), least absolute shrinkage and selection operator (LASSO) regression, and eXtreme Gradient Boosting (XGBoost), which were implemented using R packages such as glmnet, randomForest, and XGBoost, respectively. A 10-fold cross-validation (CV) repeated 5 times within the discovery cohort was used for hyperparameter tuning and to prevent overfitting. The optimal hyperparameters (e.g., mtry for RF, lambda for LASSO) were selected on the basis of the highest mean AUC from the CV process. Each algorithm generated a ranked list of important features. Gene Selection and Integrated Model Building: The top-performing features from each single algorithm were intersected to derive a consensus set of pivotal diagnostic genes. An integrated model was then constructed using a two-step approach: first, a meta-classifier (either naïve Bayes for H. pylori or linear discriminant analysis (LDA) for AF) was trained on the predictions of the base RF model as new input features. The diagnostic performance of the final model was evaluated in the independent validation cohorts using the area under the receiver operating characteristic curve (AUROC). The contribution of each selected gene to the RF model's prediction was quantified using SHapley Additive exPlanations (SHAP) analysis. PPI network construction and hub gene identification A PPI network for the co-expressed DEGs between AF and H. pylori as input was constructed using the STRING database (v12.0) with a confidence score threshold > 0.4. Hub genes were identified using the CytoHubba plugin, applying the maximal clique centrality (MCC) and degree algorithms. The top 10 overlapping candidates from both algorithms were selected as hub genes. The functional interactions among these hub genes were further explored using the GeneMANIA web tool [ 25 ]. Screening and structural characterization of small molecules The identified hub genes were submitted to the DSigDB database of the ENRICHR platform to predict potential therapeutic compounds. Compounds with an adjusted p-value < 0.05 were considered significant [ 26 – 28 ]. The chemical structures of the top candidates were retrieved from PubChem. Molecular docking was performed to evaluate the binding affinity between the candidate drugs and their predicted target proteins (hub genes). Protein structures were obtained from the RCSB Protein Data Bank (PDB). If an experimental structure was unavailable for a human target, homology modeling was performed using the SWISS-MODEL server. Docking simulations were conducted using the CB-Dock2 web server, which predicts binding sites and calculates Vina scores. The conformation with the lowest (most negative) Vina score for each pair was selected for analysis and visualization using PyMOL. Results A schematic workflow of the present study was depicted in Fig 1 , which systematically summarized the entire analytical pipeline from raw data acquisition to the final model performance evaluation. Briefly, the study commenced with the collection and preprocessing of original multi-batch omics data, where batch effects were effectively eliminated using the ComBat batch correction method ( S2 Fig ) to ensure the reliability and comparability of subsequent analyses. After data preprocessing, key feature screening and dimensionality reduction were conducted to extract the most informative variables for model construction. Subsequently, the binary classification model was established based on the optimized feature set, and the Matthews Correlation Coefficient (MCC Score) was employed as the core metric to comprehensively evaluate the predictive performance of the model. Finally, the robustness and generalization ability of the model were further verified through cross-validation and independent external validation, thus completing the whole research process of data analysis, model building and performance assessment. Fig 1. Schematic Workflow of the Research. Open in a new tab Identification of common DEGs in H. pylori infection and AF To establish the foundational transcriptional overlap between H. pylori infection and AF, we first performed differential expression analysis on their respective discovery cohorts. Regarding H. pylori infection, we identified 2,160 DEGs (1,334 upregulated and 826 downregulated). In terms of AF, 414 DEGs were identified (210 up- and 204 down-regulated) ( Fig 2A – D ). Functional enrichment revealed that DEGs related to both conditions were significantly associated with immune cell regulation and proliferation ( S1 Fig ). Critically, an intersection of these gene sets revealed 73 common DEGs shared between the two diseases, comprising 64 consistently upregulated genes and 9 downregulated DEGs ( Fig 2E , F ). Functional enrichment analysis indicated that the DEGs for H. pylori were primarily associated with immune cell proliferation, differentiation, and regulation. Similarly, DEGs for AF were enriched in immune cell-related pathways, as well as those implicated in hematologic disorders ( S1 Fig ). This core gene set provided the initial evidence for a shared transcriptional response. Fig 2. Identification of common DEGs in H. pylori infection and AF. Open in a new tab (A–B) Volcano plots of all DEGs in the H. pylori and AF discovery cohorts. Red dots represent upregulated genes, green dots represent downregulated genes, and gray dots represent genes with no differential expression (adjusted p value < 0.05, |log2FC| > 0.5). (C–D) Heatmaps of the top 50 DEGs in the H. pylori and AF discovery cohort. (E–F) Venn diagram showing 64 coupregulated and 9 codownregulated DEGs. WGCNA identifies trait-relevant gene modules To move beyond individual genes and identify functionally coherent gene networks associated with each condition, we employed WGCNA. In the H. pylori cohort (soft threshold β = 20), the black (r = 0.77, p = 1 × 10−10), brown (r = 0.68, p = 5 × 10−8), and gray (r = 0.52, p = 1 × 10−4) modules strongly positively correlated with infection status, whereas the yellow module (r = –0.68, p = 5 × 10−8) was negatively correlated ( Fig 3A ‒C). In the AF cohort (soft threshold β = 5), the brown (r = 0.45, p = 0.002) module was positively correlated, whereas the black (r = –0.57, p = 5 × 10−5) and magenta (r = –0.41, p = 0.005) modules were negatively correlated with the disease trait ( Fig 3D ‒F). The intersection of genes from these key, trait-relevant modules across both diseases yielded an additional 5 shared genes ( Fig 3G ). This network-based approach independently reinforced the existence of a common genetic substrate. Fig 3. WGCNA identifies trait-relevant gene modules. Open in a new tab Analysis of network topology for various soft-thresholding powers (β) to select the optimal value for constructing a scale-free network in the H. pylori infection (A) and AF (D) cohorts. Hierarchical clustering dendrograms of genes for H. pylori infection (B) and AF (E) by the dynamic tree cut algorithm. Heatmaps of the module–trait relationships showing the association (Pearson’s r) between module eigengenes and clinical traits in H. pylori infection (C) and AF (F). Each cell contains the correlation coefficient and corresponding p-value. The darker the color is, the stronger the correlation, red indicates a positive correlation, and blue indicates a negative correlation. (G) Venn diagram demonstrating the intersection of common genes obtained by WGCNA from the key disease-correlated modules identified in both H. pylori infection and AF analyses, yielding 5 shared genes. Functional enrichment of the integrated gene set implicates immune-inflammatory pathways We subsequently merged the 73 common DEGs with the 5 shared WGCNA genes into a unified set of 77 unique genes (including 1 intersecting gene) for comprehensive functional annotation. GO enrichment analysis revealed that these genes are overwhelmingly involved in immune effector processes, including leukocyte-mediated immunity, leukocyte proliferation, myeloid leukocyte activation, and neutrophil chemotaxis/migration ( Fig 4B , C ). KEGG pathway analysis further highlighted the roles of these genes in innate immune pathways, such as the complement and coagulation cascades, hematopoietic cell lineage, cell adhesion molecules, natural killer cell-mediated cytotoxicity and neutrophil extracellular trap formation ( Fig 4D , E ). These results strongly suggest that dysregulated immune-inflammatory signaling forms a central biological link between these two conditions. Fig 4. Functional enrichment of the integrated gene set implicates immune–inflammatory pathways. Open in a new tab (A) Venn diagram visualizing the composition of the 77-gene set derived from common DEGs and shared WGCNA genes. (B–C) Results of the enrichment analysis of the enriched genes in the GO pathway. Dot plot (B) and circular plot (C) of the top significantly enriched GO terms. BP, biological process; CC, cellular component; MF, molecular function. (D–E) Bar plot (D) and dot plot (E) displaying the KEGG pathway enrichment analysis results. The color scale represents the adjusted p value, and the dot size corresponds to the gene count. ML identifies robust diagnostic biomarkers To distill the shared gene set into a minimal, high-fidelity diagnostic signature, we conducted a comprehensive comparison of more than 20 ML algorithms. The feature selection process for H. pylori infection and AF is detailed in Fig 5A and 5F . For H. pylori infection, an integrated random forest + naïve model identified 15 pivotal genes, with SHAP analysis ranking S100A9 (SHAP = 0.359), S100A8 (0.308), C1QA (0.261), HLA-DPA1 (0.211), and CD74 (0.181) as the top contributors ( Fig 5B ). This model demonstrated excellent diagnostic performance (AUROC = 0.93) in the independent validation cohort ( Fig 5E ). For AF, an optimal random forest + LDA model selected 23 pivotal genes, with PGAM1 (SHAP = 0.0237), S100A8 (0.0202), SLA (0.0195), C5AR1 (0.018), and CD28 (0.017) being the most informative features according to the SHAP values ( Fig 5G ). This finding was also robust (AUROC = 0.88) ( Fig 5J ). The differential expression patterns of these key genes are summarized in Fig 5D and 5I . The consistent appearance of S100A8 as a key feature in both disease-specific models underscores its potential as a shared diagnostic biomarker. Fig 5. ML identifies robust diagnostic biomarkers. Open in a new tab (A, F) Model performance comparison: Heatmap showing AUC values for various models across cohorts in H. pylori infection (A) and AF (F). Right column: models; middle column: AUC. Colors indicate cohort sources. (B, G) Bar plot of gene SHAP values incorporated in the optimal ML methods for H. pylori infection (B) and AF (G). Features are ordered by their mean absolute SHAP value. The larger the bar is, the greater the contribution to the model. (C, H) Violin plots showing gene expression distributions across conditions in H. pylori infection (C) and AF (H). Width represents the density of the data, and color represents the level of expression. (D, I) Volcano plot of the top 5 genes incorporated in the optimal ML methods for H. pylori infection (D) and AF (I). (E, J) ROC curve of the top 5 genes incorporated in the optimal ML methods for H. pylori infection (E) and AF (J). The AUC values are displayed. Hub gene extraction and enrichment analysis To elucidate the core regulatory machinery within the shared gene set, we constructed a PPI network. CytoHubba analysis revealed the top 10 hub genes, including TYROBP, ITGB2, ITGAM, and SPI1 ( Fig 6B and Table 2 ). Functional analysis of these hub genes confirmed and refined our earlier findings, showing concentrated enrichment in myeloid leukocyte activation and neutrophil degranulation ( Fig 6D ‒F). KEGG pathways such as neutrophil extracellular trap formation and natural killer cell-mediated cytotoxicity were again prominent ( Fig 6G ), indicating that these hub genes occupy central positions in the inflammatory networks linking H. pylori infection to AF. Fig 6. Hub gene extraction and enrichment analysis. Open in a new tab (A) PPI network analysis of DEGs (STRING database). (B) The top 10 hub genes identified by the degree and the maximal clique centrality (MCC) algorithms via the CytoHubba plugin using the MCC method in Cytoscape. (C) GeneMANIA diagram showing the coexpression interactions between the hub genes and their neighboring genes. The color codes indicate the functions shared by the genes. (D–F) GO enrichment analysis of the hub genes, showing the top terms for BP (D), MF (E), and CC (F). (G) KEGG enrichment analyses of the hub genes. Table 2. Top 10 hub genes of the shared gene set. Rank Name MCC Score 1 TYROBP 2.05E + 07 2 ITGB2 2.05E + 07 3 ITGAM 1.99E + 07 4 FCGR3A 1.98E + 07 5 CSF3R 1.42E + 07 6 NCF2 1.39E + 07 7 SPI1 1.38E + 07 8 HCK 1.35E + 07 9 FCER1G 1.10E + 07 10 C5AR1 9999366 Open in a new tab MCC Score: Matthews Correlation Coefficient, a comprehensive performance metric for binary classification models, with a value range of [−1, 1]. A higher value indicates better model prediction performance. Prediction and validation of potential therapeutic compounds Finally, to explore translational implications, we leveraged the identified top 10 hub genes for drug repositioning via Enrichr. Querying the DSigDB database predicted several candidate compounds, including phorbol 12-myristate 13-acetate, indirubin (CHEMBL35349), tamibarotene, retinoic acid, pergolide, ropivacaine, lidocaine, aspirin, fenbuconazole, and methotrexate ( Table 3 ). Search Tool for Interactions of Chemicals (STITCH) database analysis revealed interactions between retinoic acid and SPI1/ITGAM, between indirubin and SPI1, and between ropivacaine and ITGAM [ 29 ] ( Fig 7A ). The chemical structures of retinoic acid, indirubin, and ropivacaine are shown in Fig 7B ‒D. Molecular docking simulations confirmed stable binding conformations for these candidate drug‒target pairs, with favorable binding energies ( Fig 7E ‒H). These in silico results suggest testable hypotheses for modulating the shared pathogenic network. Table 3. AF and Hp gene-targeted drugs. Term P value Adjusted P value Combined Score Genes Phorbol 12-myristate 13-acetate (phorbol 12. 13.) 3.72E-08 1.19E-05 1049.705997 HCK; TYROBP; ITGAM; SPI1; NCF2; ITGB2 CHEMBL35349 (indirubin) 9.50E-07 1.52E-04 3120.361178 CSF3R; SPI1; ITGB2 Tamibarotene 3.76E-06 4.02E-04 438.2370113 HCK; TYROBP; ITGAM; FCER1G; ITGB2 Retinoic acid 7.22E-06 5.79E-04 394.727841 HCK; FCGR3A; TYROBP; ITGAM; CSF3R; SPI1; FCER1G; NCF2; ITGB2 pergolide 1.13E-05 7.27E-04 485.1570488 HCK; TYROBP; FCER1G; NCF2 Ropivacaina (ropivacaine) 2.04E-05 0.001091906 4494.847104 ITGAM; ITGB2 Lidocaine 9.07E-05 0.004159024 1720.515946 ITGAM; ITGB2 aspirin 1.12E-04 0.004259775 211.5027674 HCK; CSF3R; FCER1G; ITGB2 Fenbuconazole 1.33E-04 0.004259775 1349.727435 FCER1G; C5AR1 methotrexate 1.64E-04 0.004798672 182.995811 TYROBP; CSF3R; FCER1G; C5AR1 Open in a new tab Fig 7. Prediction and validation of potential therapeutic compounds. Open in a new tab (A) PPI network diagram of the 10 compounds and their predicted target hub genes, visualized using the STITCH database. (B–D) Chemical structures and 3D structures of indirubin (B), retinoic acid (C), and ropivacaine (D). (E–H) Molecular docking model showing the predicted binding conformations of indirubin and SPI1 (E), retinoic acid and ITGAM (F), retinoic acid and SPI1 (G), and ropivacaine and ITGAM (H). Discussion Our integrative bioinformatics and ML study elucidates the shared molecular landscape between H. pylori infection and AF, with a focus on robust immune-inflammatory activation. This computational approach moves beyond reported epidemiological links [ 15 , 30 – 32 ] to define a precise, genetically based framework that could explain the clinical association. The identification of 77 unique genes at the intersection of both conditions, significantly enriched in pathways governing neutrophil chemotaxis, myeloid leukocyte activation, and leukocyte-mediated immunity, provides strong molecular evidence supporting chronic inflammation as a key mechanistic bridge. These findings align with and extend the current understanding that systemic inflammation induced by H. pylori [ 33 – 35 ] creates a proarrhythmic substrate, facilitating the electrophysiological and structural remodeling central to AF pathogenesis [ 36 , 37 ] . A pivotal discovery from our multialgorithm ML pipeline is the prominence of S100A8 as a shared diagnostic biomarker. S100A8, which typically functions as a heterodimer with S100A9, is a potent damage-associated molecular pattern (DAMP) protein crucial to innate immunity [ 38 , 39 ] . Its role appears to be context-dependent, linking local infection to systemic cardiac vulnerability. During H. pylori infection, S100A8/A9 contributes to gastric mucosal inflammation and host defense via nutritional immunity [ 40 – 43 ] , potentially increasing its systemic expression. In the context of AF, S100A8/A9 promotes oxidative stress, cardiomyocyte dysfunction, and electrical remodeling [ 44 – 46 ] . Thus, we propose that S100A8 serves as a measurable pathogenic link: its systemic expression increase due to chronic gastric infection may directly exacerbate atrial inflammation and oxidative injury, thereby lowering the threshold for arrhythmogenesis. Our models suggest that quantifying S100A8 expression, potentially in conjunction with that of other key biomarkers such as FCER1G or ITGB2, could enhance risk stratification for AF in patients with H. pylori infection. Beyond S100A8, the PPI network revealed a core set of hub genes—including TYROBP, ITGB2, SPI1, ITGAM, and C5AR1—that further underscore the centrality of leukocyte adhesion, signaling, and activation in this shared pathology. The enrichment of these hubs in pathways such as “neutrophil extracellular trap formation” and “myeloid leukocyte-mediated immunity” indicates that a sustained, coordinated innate immune response is a plausible unifying mechanism. This network analysis shifts the focus from a single biomarker to a dysfunctional immune module, offering a broader set of potential therapeutic targets for intervention. The translational promise of our findings is highlighted by the drug repositioning analysis. Candidate compounds such as indirubin and ropivacaine, which are predicted to target hub genes such as SPI1 and ITGAM, merit investigation given their known anti-inflammatory properties in other contexts [ 47 , 48 ] . Particularly intriguing is the connection with retinoic acid, which has experimentally been shown to downregulate S100A8 expression [ 49 ] . This suggests a testable hypothesis: retinoic acid could attenuate the H. pylori -AF link by mitigating the S100A8-driven inflammatory cascade. The stable binding conformations of these candidates with their target proteins, validated by molecular docking, provide a structural rationale for further preclinical studies. Limitations and future perspectives We acknowledge several limitations inherent to this in silico study. First, while we employed rigorous ComBat batch correction and independent validation cohorts, the analysis is based on heterogeneous public datasets, and the findings require confirmation in prospective, uniformly processed clinical cohorts. Second, our study demonstrated an association but could not establish causality between H. pylori infection, the identified gene signatures, and AF. Third, the diagnostic models, although performant, need validation in larger, multicenter studies to assess their real-world clinical utility. Finally, the predicted therapeutic candidates need thorough experimental validation in vitro and in vivo to confirm their efficacy and mechanism. Future research should employ Mendelian randomization to infer causality, utilize single-cell sequencing to pinpoint the specific immune cell populations driving these signatures, and conduct functional experiments to delineate the precise roles of hub genes such as S100A8 in atrial pathobiology. Conclusion In conclusion, by integrating bioinformatics and machine learning, we constructed a shared immune‒inflammatory pathway network linking H. pylori infection to AF, with S100A8 emerging as a central biomarker. The identified hub genes and the drug repositioning candidates, especially retinoic acid, provide a foundational framework for understanding the pathophysiology and for developing novel diagnostic and therapeutic strategies. This work translates epidemiological observations into a molecular hypothesis, offering concrete targets for future mechanistic and clinical investigations. Supporting information S1 Fig. Functional enrichment and pathway enrichment analysis of DEGs in H. pylori infection and AF. Dot plot (A, D) and circular plot (B, E) displaying the results of the GO enrichment analysis of DEGs specific to H. pylori infection and AF. KEGG enrichment analysis results of DEGs specific to H. pylori infection (C) and AF (F). (TIF) pone.0346038.s001.tif (9.8MB, tif) S2 Fig. Batch effect correction of AF and H. pylori datasets. (TIF) pone.0346038.s002.tif (2.9MB, tif) S1 Table. 77 union genes for comprehensive functional annotation and machine leaming. (DOCX) pone.0346038.s003.docx (14.4KB, docx) S2 Table. AUC values for various models across cohorts in H. pylori infection. (DOCX) pone.0346038.s004.docx (15.3KB, docx) S3 Table. AUC values for various models across cohorts in AF. (DOCX) pone.0346038.s005.docx (16.9KB, docx) S4 Table. The AUC values of the top 5 genes incorporated in the optimal ML methods for H. pylori infection. (DOCX) pone.0346038.s006.docx (10.9KB, docx) S5 Table. The AUC values of the top 5 genes incorporated in the optimal ML methods for AF. (DOCX) pone.0346038.s007.docx (10.9KB, docx) Data Availability All sequencing data were sourced from the Gene Expression Omnibus (GEO) database. A total of four datasets related to AF were included: GSE41177 and GSE176166 constituted the discovery cohort, while GSE108660 and GSE128188 formed the validation cohort. For H. pylori infection, five datasets were selected: GSE27411, GSE60662, and GSE233973 comprised the discovery cohort, and GSE5081 and GSE60427 served as the validation cohort. Funding Statement The author(s) received no specific funding for this work. References 1. Zeng R, Gou H, Lau HCH, Yu J. Stomach microbiota in gastric cancer development and clinical implications. Gut. 2024;73(12):2062–73. doi: 10.1136/gutjnl-2024-332815 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Hooi JKY, Lai WY, Ng WK, Suen MMY, Underwood FE, Tanyingoh D, et al. Global Prevalence of Helicobacter pylori Infection: Systematic Review and Meta-Analysis. Gastroenterology. 2017;153(2):420–9. doi: 10.1053/j.gastro.2017.04.022 [ DOI ] [ PubMed ] [ Google Scholar ] 3. Yuan C, Adeloye D, Luk TT, Huang L, He Y, Xu Y, et al. The global prevalence of and factors associated with Helicobacter pylori infection in children: a systematic review and meta-analysis. Lancet Child Adolesc Health. 2022;6(3):185–94. doi: 10.1016/S2352-4642(21)00400-4 [ DOI ] [ PubMed ] [ Google Scholar ] 4. Malfertheiner P, Megraud F, O’Morain CA, Gisbert JP, Kuipers EJ, Axon AT, et al. Management of Helicobacter pylori infection-the Maastricht V/Florence Consensus Report. Gut. 2017;66(1):6–30. doi: 10.1136/gutjnl-2016-312288 [ DOI ] [ PubMed ] [ Google Scholar ] 5. Franceschi F, Annalisa T, Teresa DR, Giovanna D, Ianiro G, Franco S, et al. Role of Helicobacter pylori infection on nutrition and metabolism. World J Gastroenterol. 2014;20(36):12809–17. doi: 10.3748/wjg.v20.i36.12809 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Staerk L, Sherer JA, Ko D, Benjamin EJ, Helm RH. Atrial Fibrillation: Epidemiology, Pathophysiology, and Clinical Outcomes. Circ Res. 2017;120(9):1501–17. doi: 10.1161/circresaha.117.309732 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 7. Kirchhof P, Benussi S, Kotecha D, Ahlsson A, Atar D, Casadei B, et al. 2016 ESC Guidelines for the management of atrial fibrillation developed in collaboration with EACTS. Eur Heart J. 2016;37(38):2893–962. doi: 10.1093/eurheartj/ehw210 [ DOI ] [ PubMed ] [ Google Scholar ] 8. Chugh SS, Havmoeller R, Narayanan K, Singh D, Rienstra M, Benjamin EJ. Worldwide epidemiology of atrial fibrillation: a Global Burden of Disease 2010 Study. Circulation. 2014;129(8):837–47. doi: 10.1161/circulationaha.113.005119 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Lane DA, Skjøth F, Lip GYH, Larsen TB, Kotecha D. Temporal Trends in Incidence, Prevalence, and Mortality of Atrial Fibrillation in Primary Care. J Am Heart Assoc. 2017;6(5):e005155. doi: 10.1161/JAHA.116.005155 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Kornej J, Börschel CS, Benjamin EJ, Schnabel RB. Epidemiology of Atrial Fibrillation in the 21st Century: Novel Methods and New Insights. Circ Res. 2020;127(1):4–20. doi: 10.1161/CIRCRESAHA.120.316340 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Reddy YNV, Obokata M, Verbrugge FH, Lin G, Borlaug BA. Atrial Dysfunction in Patients With Heart Failure With Preserved Ejection Fraction and Atrial Fibrillation. J Am Coll Cardiol. 2020;76(9):1051–64. doi: 10.1016/j.jacc.2020.07.009 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Zhang MJ, Norby FL, Lutsey PL, Mosley TH, Cogswell RJ, Konety SH, et al. Association of Left Atrial Enlargement and Atrial Fibrillation With Cognitive Function and Decline: The ARIC-NCS. J Am Heart Assoc. 2019;8(23):e013197. doi: 10.1161/JAHA.119.013197 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Ko D, Chung MK, Evans PT, Benjamin EJ, Helm RH. Atrial Fibrillation: A Review. JAMA. 2025;333(4):329–42. doi: 10.1001/jama.2024.22451 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Andersen JH, Andreasen L, Olesen MS. Atrial fibrillation-a complex polygenetic disease. Eur J Hum Genet. 2021;29(7):1051–60. doi: 10.1038/s41431-020-00784-8 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Montenero AS, Mollichelli N, Zumbo F, Antonelli A, Dolci A, Barberis M, et al. Helicobacter pylori and atrial fibrillation: a possible pathogenic link. Heart. 2005;91(7):960–1. doi: 10.1136/hrt.2004.036681 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Bunch TJ, Day JD, Anderson JL, Horne BD, Muhlestein JB, Crandall BG, et al. Frequency of helicobacter pylori seropositivity and C-reactive protein increase in atrial fibrillation in patients undergoing coronary angiography. Am J Cardiol. 2008;101(6):848–51. doi: 10.1016/j.amjcard.2007.09.118 [ DOI ] [ PubMed ] [ Google Scholar ] 17. Andrew P, Montenero AS. Is Helicobacter pylori a cause of atrial fibrillation? Future Cardiol. 2006;2(4):429–39. doi: 10.2217/14796678.2.4.429 [ DOI ] [ PubMed ] [ Google Scholar ] 18. Yeh Y-H, Kuo C-T, Lee Y-S, Lin Y-M, Nattel S, Tsai F-C, et al. Region-specific gene expression profiles in the left atria of patients with valvular atrial fibrillation. Heart Rhythm. 2013;10(3):383–91. doi: 10.1016/j.hrthm.2012.11.013 [ DOI ] [ PubMed ] [ Google Scholar ] 19. Thomas AM, Cabrera CP, Finlay M, Lall K, Nobles M, Schilling RJ, et al. Differentially expressed genes for atrial fibrillation identified by RNA sequencing from paired human left and right atrial appendages. Physiol Genomics. 2019;51(8):323–32. doi: 10.1152/physiolgenomics.00012.2019 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Nookaew I, Thorell K, Worah K, Wang S, Hibberd ML, Sjövall H, et al. Transcriptome signatures in Helicobacter pylori-infected mucosa identifies acidic mammalian chitinase loss as a corpus atrophy marker. BMC Med Genomics. 2013;6:41. doi: 10.1186/1755-8794-6-41 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Hanada K, Uchida T, Tsukamoto Y, Watada M, Yamaguchi N, Yamamoto K, et al. Helicobacter pylori infection introduces DNA double-strand breaks in host cells. Infect Immun. 2014;82(10):4182–9. doi: 10.1128/IAI.02368-14 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 22. Takeuchi C, Sato J, Yamamichi N, Kageyama-Yahara N, Sasaki A, Akahane T, et al. Marked intestinal trans-differentiation by autoimmune gastritis along with ectopic pancreatic and pulmonary trans-differentiation. J Gastroenterol. 2024;59(2):95–108. doi: 10.1007/s00535-023-02055-x [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Galamb O, Gyõrffy B, Sipos F, Dinya E, Krenács T, Berczi L, et al. Helicobacter pylori and antrum erosion-specific gene expression patterns: the discriminative role of CXCL13 and VCAM1 transcripts. Helicobacter. 2008;13(2):112–26. doi: 10.1111/j.1523-5378.2008.00584.x [ DOI ] [ PubMed ] [ Google Scholar ] 24. Nagashima H, Iwatani S, Cruz M, Jiménez Abreu JA, Uchida T, Mahachai V, et al. Toll-like Receptor 10 in Helicobacter pylori Infection. J Infect Dis. 2015;212(10):1666–76. doi: 10.1093/infdis/jiv270 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Warde-Farley D, Donaldson SL, Comes O, Zuberi K, Badrawi R, Chao P, et al. The GeneMANIA prediction server: biological network integration for gene prioritization and predicting gene function. Nucleic Acids Res. 2010;38(Web Server issue):W214–20. doi: 10.1093/nar/gkq537 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Chen EY, Tan CM, Kou Y, Duan Q, Wang Z, Meirelles GV, et al. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinformatics. 2013;14:128. doi: 10.1186/1471-2105-14-128 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Kuleshov MV, Jones MR, Rouillard AD, Fernandez NF, Duan Q, Wang Z, et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 2016;44(W1):W90-7. doi: 10.1093/nar/gkw377 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Xie Z, Bailey A, Kuleshov MV, Clarke DJB, Evangelista JE, Jenkins SL, et al. Gene Set Knowledge Discovery with Enrichr. Curr Protoc. 2021;1(3):e90. doi: 10.1002/cpz1.90 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Szklarczyk D, Santos A, von Mering C, Jensen LJ, Bork P, Kuhn M. STITCH 5: augmenting protein-chemical interaction networks with tissue and affinity data. Nucleic Acids Res. 2016;44(D1):D380-4. doi: 10.1093/nar/gkv1277 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Farah R, Hanna T, Levin G. Is there a link between atrial fibrillation and Helicobacter pylori infections? Minerva Gastroenterol. 2024;70(2):177–80. doi: 10.23736/s2724-5985.23.03323-5 [ DOI ] [ PubMed ] [ Google Scholar ] 31. Wang D-Z, Chen W, Yang S, Wang J, Li Q, Fu Q, et al. Helicobacter pylori infection in Chinese patients with atrial fibrillation. Clin Interv Aging. 2015;10:813–9. doi: 10.2147/CIA.S72724 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Rivington J, Twohig P. Quantifying Risk Factors for Atrial Fibrillation: Retrospective Review of a Large Electronic Patient Database. J Atr Fibrillation. 2020;13(3):2365. doi: 10.4022/jafib.2365 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Fei X, Chen S, Li L, Xu X, Wang H, Ke H, et al. Helicobacter pylori infection promotes M1 macrophage polarization and gastric inflammation by activation of NLRP3 inflammasome via TNF/TNFR1 axis. Cell Commun Signal. 2025;23(1):6. doi: 10.1186/s12964-024-02017-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Choi YH, Lai J, Kim M-A, Kim A, Kim J, Su H, et al. CagL polymorphisms between East Asian and Western Helicobacter pylori are associated with different abilities to induce IL-8 secretion. J Microbiol. 2021;59(8):763–70. doi: 10.1007/s12275-021-1136-2 [ DOI ] [ PubMed ] [ Google Scholar ] 35. Guo X, Tang P, Zhang X, Li R. Causal associations of circulating Helicobacter pylori antibodies with stroke and the mediating role of inflammation. Inflamm Res. 2023;72(6):1193–202. doi: 10.1007/s00011-023-01740-0 [ DOI ] [ PubMed ] [ Google Scholar ] 36. Shitole SG, Heckbert SR, Marcus GM, Shah SJ, Sotoodehnia N, Walston JD, et al. Assessment of Inflammatory Biomarkers and Incident Atrial Fibrillation in Older Adults. J Am Heart Assoc. 2024;13(24):e035710. doi: 10.1161/JAHA.124.035710 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Li X, Peng S, Wu X, Guan B, Tse G, Chen S, et al. C-reactive protein and atrial fibrillation: Insights from epidemiological and Mendelian randomization studies. Nutr Metab Cardiovasc Dis. 2022;32(6):1519–27. doi: 10.1016/j.numecd.2022.03.008 [ DOI ] [ PubMed ] [ Google Scholar ] 38. Pruenster M, Vogl T, Roth J, Sperandio M. S100A8/A9: From basic science to clinical application. Pharmacol Ther. 2016;167:120–31. doi: 10.1016/j.pharmthera.2016.07.015 [ DOI ] [ PubMed ] [ Google Scholar ] 39. Wang S, Song R, Wang Z, Jing Z, Wang S, Ma J. S100A8/A9 in Inflammation. Front Immunol. 2018;9:1298. doi: 10.3389/fimmu.2018.01298 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Gaddy JA, Radin JN, Loh JT, Piazuelo MB, Kehl-Fie TE, Delgado AG, et al. The host protein calprotectin modulates the Helicobacter pylori cag type IV secretion system via zinc sequestration. PLoS Pathog. 2014;10(10):e1004450. doi: 10.1371/journal.ppat.1004450 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Diaz-Ochoa VE, Jellbauer S, Klaus S, Raffatellu M. Transition metal ions at the crossroads of mucosal immunity and microbial pathogenesis. Front Cell Infect Microbiol. 2014;4:2. doi: 10.3389/fcimb.2014.00002 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Inciarte-Mundo J, Frade-Sosa B, Sanmartí R. From bench to bedside: Calprotectin (S100A8/S100A9) as a biomarker in rheumatoid arthritis. Front Immunol. 2022;13:1001025. doi: 10.3389/fimmu.2022.1001025 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Kao KD, Grasberger H, El-Zaatari M. The Cxcr2+ subset of the S100a8+ gastric granylocytic myeloid-derived suppressor cell population (G-MDSC) regulates gastric pathology. Front Immunol. 2023;14:1147695. doi: 10.3389/fimmu.2023.1147695 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 44. Meng S, Huang T, Zhou Z, Yu L, Wang H. Chronic mild stress exacerbates atrial fibrillation and neutrophil extracellular traps formation through S100A8/A9 signaling. Signal Transduct Target Ther. 2025;10(1):108. doi: 10.1038/s41392-025-02199-7 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Liang S, Zhang X, Chen J, He Y, Lai J. Co-expression and interaction network analysis identifies neutrophil-related genes as the core mediator of atrial fibrillation. Gen Physiol Biophys. 2024;43(3):209–19. doi: 10.4149/gpb_2024004 [ DOI ] [ PubMed ] [ Google Scholar ] 46. Yarovinsky TO, Su M, Chen C, Xiang Y, Tang WH, Hwa J. Pyroptosis in cardiovascular diseases: Pumping gasdermin on the fire. Semin Immunol. 2023;69:101809. doi: 10.1016/j.smim.2023.101809 [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 47. Rabinow B, Werling J, Bendele A, Gass J, Bogseth R, Balla K, et al. Intra-articular (IA) ropivacaine microparticle suspensions reduce pain, inflammation, cytokine, and substance p levels significantly more than oral or IA celecoxib in a rat model of arthritis. Inflammation. 2015;38(1):40–60. doi: 10.1007/s10753-014-0006-z [ DOI ] [ PubMed ] [ Google Scholar ] 48. Miyoshi K, Takaishi M, Digiovanni J, Sano S. Attenuation of psoriasis-like skin lesion in a mouse model by topical treatment with indirubin and its derivative E804. J Dermatol Sci. 2012;65(1):70–2. doi: 10.1016/j.jdermsci.2011.10.001 [ DOI ] [ PubMed ] [ Google Scholar ] 49. Li D, Li H, Cheng C, Li G, Yuan F, Mi R, et al. All-trans retinoic acid enhanced the antileukemic efficacy of ABT-199 in acute myeloid leukemia by downregulating the expression of S100A8. Int Immunopharmacol. 2022;112:109182. doi: 10.1016/j.intimp.2022.109182 [ DOI ] [ PubMed ] [ Google Scholar ] PLoS One. doi: 10.1371/journal.pone.0346038.r001 Decision Letter 0 Tomasz Kaminski Tomasz Kaminski Academic Editor Find articles by Tomasz Kaminski Author information Copyright and License information Roles Tomasz Kaminski : Academic Editor © 2026 Tomasz KaminskiTomasz Kaminski This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 17 Dec 2025 Dear Dr. wang, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Jan 31 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at [email protected] . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .. We look forward to receiving your revised manuscript. Kind regards, Tomasz W. Kaminski Academic Editor PLOS One Journal Requirements: When submitting your revision, we need you to address these additional requirements. 1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf 2. Please note that PLOS One has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse. 3. Please note that funding information should not appear in any section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form. Please remove any funding-related text from the manuscript. 4. Thank you for stating the following in your Competing Interests section: “None” Please complete your Competing Interests on the online submission form to state any Competing Interests. If you have no competing interests, please state "The authors have declared that no competing interests exist.", as detailed online in our guide for authors at http://journals.plos.org/plosone/s/submit-now This information should be included in your cover letter; we will change the online submission form on your behalf. 5. Please note that your Data Availability Statement is currently missing a direct link to access each database. If your manuscript is accepted for publication, you will be asked to provide these details on a very short timeline. We therefore suggest that you provide this information now, though we will not hold up the peer review process if you are unable. 6. Please include your full ethics statement in the ‘Methods’ section of your manuscript file. In your statement, please include the full name of the IRB or ethics committee who approved or waived your study, as well as whether or not you obtained informed written or verbal consent. If consent was waived for your study, please include this information in your statement as well. 7. We note you have included a table to which you do not refer in the text of your manuscript. Please ensure that you refer to Table 1 in your text; if accepted, production will need this reference to link the reader to the Table. 8. Please include captions for your Supporting Information files at the end of your manuscript, and update any in-text citations to match accordingly. Please see our Supporting Information guidelines for more information: http://journals.plos.org/plosone/s/supporting-information .. 9. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. Additional Editor Comments: Dear Authors, Thank you for your submission. The study addresses an interesting question, and the overall analytical approach is promising. However, substantial revisions are needed to improve methodological clarity, manuscript structure, figures, formatting and language quality before the work can be reconsidered. We look forward to receiving your re-submission. Best Regards, Tomasz W Kaminski [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author 1. Is the manuscript technically sound, and do the data support the conclusions? Reviewer #1: Partly Reviewer #2: Yes ********** 2. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #1: No Reviewer #2: Yes ********** 3. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #2: Yes ********** 4. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: No Reviewer #2: Yes ********** Reviewer #1: This manuscript, in its current form, does not meet the minimum presentation, structure, and formatting standards expected of a PLOS ONE research article, independently of the scientific content. Even before considering the validity of the analyses, the paper does not look or read like a publication-ready journal article. This is a major concern and would normally justify rejection without peer review or a request for complete reformatting and resubmission. 1. Overall Structure Does Not Follow Journal Standards While the manuscript nominally contains the standard sections (Abstract, Introduction, Methods, Results, Discussion), the structure is not consistently applied in a journal-ready way: The Methods section reads like a rough technical checklist, not a properly written scientific methods narrative. The Results section lacks clear subheadings and logical flow, making it difficult to follow the analytical progression. The Discussion repeatedly re-states the Results instead of critically interpreting them, and frequently shifts into speculation without clear boundaries. The manuscript currently reads more like a bioinformatics project report or preprint draft, not a polished research article. 2. Figures and Tables Are Not Presented at Journal Standard The figures and tables exhibit multiple serious presentation problems: Repeated typographical errors such as "Fiugre" instead of "Figure" throughout the document are unprofessional and unacceptable at submission stage. Figure references appear as “Click here to access/download;Figure;Figure X.tif”, which is not appropriate for a manuscript text and suggests that the document was exported incorrectly. The figures themselves are not described with sufficient interpretive captions — many captions merely restate what the plot type is, rather than what the figure shows scientifically. Table formatting is inconsistent, and some entries appear cluttered, poorly aligned, and difficult to read. At minimum, the authors must: Correct all figure and table labeling errors, Embed figures and tables properly in publication format, Rewrite all figure legends to be self-contained and explanatory, not technical placeholders. 3. Language, Grammar, and Editorial Quality Are Substandard While the general meaning of the text is understandable, the manuscript contains: Numerous grammatical errors, awkward phrasing, and unnatural sentence construction, Inconsistent spacing, punctuation, and formatting, Repetitive wording and poorly structured paragraphs, Informal or imprecise scientific phrasing. PLOS ONE explicitly states that it will not copyedit accepted manuscripts, and the current language level is not acceptable for publication. The manuscript requires professional, full-scale English language editing, not light proofreading. 4. Visual and Logical Flow Is Disrupted The overall visual and logical presentation is weak: Figures, tables, and text are not well integrated. The narrative jumps between bioinformatics steps without smooth transitions. The manuscript lacks a clear storyline that would guide a reader from biological question → computational strategy → biological interpretation. As it stands, the paper is difficult to read continuously as a coherent article. 5. Professional Presentation and Journal Readiness Taken together: Formatting problems, Figure embedding errors, Language quality, Structural weaknesses, mean that this manuscript does not resemble a finished journal article. It appears closer to a working draft or early internal report rather than a manuscript ready for peer-reviewed publication. - Final Recommendation on Presentation Grounds Alone Regardless of the scientific analyses, I believe this manuscript requires complete reformatting and professional language editing before it can be properly evaluated as a journal article. At this stage, my recommendation based purely on presentation, structure, and formatting is: Major Revision (borderline Reject / Resubmit as New Submission after full reformatting) Reviewer #2: Overall, this manuscript is clearly organized and presents an interesting integrative analysis combining bioinformatics and machine learning to explore shared molecular features between H. pylori infection and atrial fibrillation. To strengthen transparency and reproducibility, it would be helpful to add a concise dataset and preprocessing summary, including for each GEO dataset the platform, source, sample sizes per group, and the key preprocessing steps, as well as a clear description of how batch effects were addressed when combining discovery datasets. The machine-learning component is promising, and a few additional reporting details would make the work easier to reproduce and interpret. Specifically, (i) how the discovery datasets were harmonized/combined prior to modeling, (ii) which hyperparameters were evaluated and the criterion used to choose the final settings, (iii) how the final gene panels were derived, and (iv) how the integrated models were implemented (stacking, voting, or sequential). If cross-validation was used within the discovery cohort during tuning or feature selection, please specify the CV design; if not, stating that explicitly would also help readers interpret the reported AUCs. Providing the final gene lists, key model settings, and any relevant code and parameters in supplementary material would further increase confidence in the robustness of the proposed markers. The manuscript is generally intelligible and readable, and with a careful proofread it could be even stronger. I suggest revising minor typographical and formatting issues (like "Fiugre")and ensuring consistent use of abbreviations, spacing, and figure labels throughout the text. ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our Privacy Policy .--> Reviewer #1: No Reviewer #2: Yes: Yuchen ZhangYuchen Zhang ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation . NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. PLoS One. 2026 Apr 10;21(4):e0346038. doi: 10.1371/journal.pone.0346038.r002 Author response to Decision Letter 1 Article notes Copyright and License information Collection date 2026. PMC Copyright notice 12 Jan 2026 Manuscript Title: Integrative Bioinformatics and Machine Learning Identify Shared Molecular Mechanisms and Diagnostic Biomarkers Between Helicobacter pylori Infection and Atrial Fibrillation We sincerely thank the editor and both reviewers for their valuable time and insightful comments, which have greatly helped us improve the quality and clarity of our manuscript. We have carefully considered each point and have substantially revised the manuscript accordingly. In direct response to your critique that the manuscript did not meet minimum journal standards, we undertook a comprehensive overhaul. We have systematically addressed every point raised by carefully re-examining, re-writing, and professionally refining the entire manuscript. Below, we provide a detailed point-by-point account of the actions taken. Response to Reviewer #1: Comment 1: On overall structure, language, grammar, and editorial quality. Response: We agree completely with the reviewer's assessment that the original manuscript's language and structure were substandard and read more like a project report. To rectify this, we have not merely edited but have substantially re-written the manuscript to conform to the standards of a research article. Global Language and Formatting Revision: The entire text has undergone rigorous professional language editing and proofreading. We have corrected grammatical errors, improved sentence fluency and clarity, eliminated awkward phrasing and repetition, and ensured consistent and formal scientific terminology throughout. Restructured Methods Section: The "Materials and Methods" section has been completely re-organized and re-written. It now provides a coherent, logical narrative of our analytical workflow in dedicated subsections, moving decisively away from the initial checklist format to offer clear, reproducible descriptions. Re-organized and Re-written Results Section: We have introduced clear, descriptive subheadings that reflect the analytical progression The text has been carefully re-crafted to ensure a logical flow, explicitly connecting each analytical step to the next and guiding the reader clearly through the sequence of discoveries. Expanded and Deepened Discussion Section: We have significantly expanded the "Discussion" to move decisively beyond re-stating results. It now provides a critical interpretation of our findings within the broader context of existing literature, elaborates in depth on the biological and clinical implications (e.g., the dual contextual role of S100A8), and clearly differentiates between evidence-based conclusions and speculative points for future inquiry. A dedicated "Limitations and Future Perspectives" subsection has been added to frame the study appropriately. Correction of Editorial Artifacts: Regarding the specific typographical error ("Fiugre") and improper figure references noted, we have meticulously proofread our source documents. We believe these may have been introduced during the initial file export or submission process. We will exercise utmost care during re-submission to ensure all text, labels, and citations are error-free. Comment 2: On the presentation of figures/tables and the overall logical/visual flow. Response: We thank the reviewer for this critical feedback on presentation, which we have addressed comprehensively. Enhanced Figures and Tables: All figures have been re-created and formatted strictly according to typical journal requirements concerning font consistency, size, resolution, and visual clarity. The legends have been entirely re-written to be self-contained and explanatory, detailing what each panel shows and summarizing the key scientific finding it conveys. Improved Logical Narrative and Cohesion: To address the disjointed flow, we have strengthened the connective tissue of the manuscript. This includes refining the narrative arc from the Introduction through to the Discussion and incorporating more transitional phrasing between sections and paragraphs. This ensures a smoother, more coherent reading experience that logically connects the biological question, the computational strategy, and the final interpretation. We are profoundly grateful for the reviewer's rigorous assessment. The critical feedback provided was essential for us to elevate the quality of our work. We believe the extensive revisions detailed above have thoroughly addressed all the concerns raised and have transformed the manuscript into a polished, coherent, and publication-ready research article. We hope the revised version now meets the expected standards of your journal. Response to Reviewer #2: We sincerely thank the reviewer for the positive assessment and exceptionally constructive suggestions, which have been instrumental in guiding our comprehensive revision. As noted in our response to Reviewer #1, we have re-written the manuscript in its entirety to address overarching presentation issues. Your specific comments were central to shaping the new, detailed methodological reporting within this re-written framework, greatly enhancing the work's transparency and reproducibility. Comment 1: Strengthening transparency in dataset description and preprocessing. Response: We have added a new subsection, and a Supplementary Table 1 summarizing each dataset (accession, platform, sample sizes). We now clearly state that ComBat batch correction was applied to the discovery cohorts only, while validation cohorts were processed independently. Comment 2: Clarifying the machine learning workflow. Response: We have expanded the methods subsection “Machine Learning-Based Diagnostic Model Construction and Validation” to detail the process: Hyperparameter tuning was performed using 10-fold cross-validation repeated 5 times within the discovery cohort. Final gene panels were derived by intersecting top features from multiple algorithms. The integrated models were implemented via a two-step approach (base RF + meta-classifier). All reported AUROC values are from the independent validation cohorts, ensuring an unbiased performance estimate. Comment 3: Improving language and formatting consistency. Response: The manuscript has undergone thorough proofreading. All typographical errors have been corrected, and formatting (abbreviations, spacing, figure labels) has been standardized throughout. We are grateful for your insightful comments, which have strengthened the manuscript. Attachment Submitted filename: To the Reviewers.docx pone.0346038.s009.docx (14.5KB, docx) PLoS One. doi: 10.1371/journal.pone.0346038.r003 Decision Letter 1 Tomasz Kaminski Tomasz Kaminski Academic Editor Find articles by Tomasz Kaminski Author information Copyright and License information Roles Tomasz Kaminski : Academic Editor © 2026 Tomasz KaminskiTomasz Kaminski This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 27 Jan 2026 Dear Dr. wang, Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process. Please submit your revised manuscript by Mar 13 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at [email protected] . When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file. A letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'. A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'. An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'. If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at . Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols .. We look forward to receiving your revised manuscript. Kind regards, Tomasz W. Kaminski Academic Editor PLOS One Journal Requirements: 1. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise. 2. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice. Additional Editor Comments: Dear Authors, Thank you for your revised submission and for the careful responses provided in the previous round. One reviewer finds that the manuscript has adequately addressed all earlier comments and considers the study technically sound and suitable for publication. To further strengthen the manuscript, I encourage you to focus the revision on improving the clarity and structure with which the key results are presented and interpreted, as suggested Reviewer. In particular, summarizing the main findings in a more structured manner and expanding the analytical depth of the Discussion - especially the limitations - would help readers more clearly assess the scope, robustness, and implications of the work. These revisions are intended to enhance transparency and balance rather than to request additional analyses. Please feel free to reach out if any points in the reviews would benefit from clarification. Kind regards, Tomasz W Kaminski [Note: HTML markup is below. Please do not edit.] Reviewers' comments: Reviewer's Responses to Questions Comments to the Author Reviewer #1: All comments have been addressed Reviewer #2: All comments have been addressed ********** 2. Is the manuscript technically sound, and do the data support the conclusions??> Reviewer #1: Partly Reviewer #2: Yes ********** 3. Has the statistical analysis been performed appropriately and rigorously? -->?> Reviewer #1: Yes Reviewer #2: Yes ********** 4. Have the authors made all data underlying the findings in their manuscript fully available??> The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.--> Reviewer #1: Yes Reviewer #2: Yes ********** 5. Is the manuscript presented in an intelligible fashion and written in standard English??> Reviewer #1: Yes Reviewer #2: Yes ********** Reviewer #1: The manuscript lacks clear, publication-quality figures and tables that adequately support the results. At present, the results are presented largely in narrative form, with minimal quantitative visualization. Key findings (e.g., diagnostic model performance, hub gene importance, validation outcomes) are not summarized in structured tables or interpretable figures. The absence of well-designed figures and summary tables significantly limits the evidential value of the manuscript and makes it difficult for readers to independently assess the results. The authors should include properly labelled, publication-standard figures and tables that directly correspond to each major result and hypothesis. While the Discussion is extensive, it remains largely narrative and would benefit from deeper analytical expansion rather than additional speculative content. In particular, the limitations section should be substantially expanded to address dataset heterogeneity, model stability, lack of causal inference, and the non-predictive nature of molecular docking. Strengthening these sections would significantly improve the scientific rigor and interpretive balance of the manuscript. Reviewer #2: The authors have adequately addressed all comments raised in the previous review. The revised manuscript is clearly written, technically sound, and presents a coherent and well-documented analytical workflow. The data support the conclusions, and the study meets the standards for publication. ********** what does this mean? ). If published, this will include your full peer review and any attached files.). If published, this will include your full peer review and any attached files. If you choose “no”, your identity will remain anonymous but your review may still be made public. Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our For information about this choice, including consent withdrawal, please see our Privacy Policy .--> Reviewer #1: No Reviewer #2: Yes: yuchen zhangyuchen zhang ********** [NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.] To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation . NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications. PLoS One. 2026 Apr 10;21(4):e0346038. doi: 10.1371/journal.pone.0346038.r004 Author response to Decision Letter 2 Article notes Copyright and License information Collection date 2026. PMC Copyright notice 12 Mar 2026 Response to Reviewer #1: Dear Reviewer, We sincerely thank you for re-reviewing our manuscript and for providing such valuable feedback. Your professional insights have been instrumental in further improving the quality of our work. In this second round of revision, we have carefully considered all the issues raised and have made corresponding improvements to the manuscript. Below is our point-by-point response to your specific comments. Response to Reviewer #1: Comment 1: The manuscript lacks clear, publication-ready data and tables to adequately support the results. Key findings are not summarized in structured tables or interpretable figures, which limits the evidentiary value of the manuscript. Response: We fully agree with your assessment that clear data visualization and structured tables are fundamental to supporting scientific conclusions. To address this, we have comprehensively supplemented and optimized the figures and tables in our manuscript: Addition of a study flowchart (now Figure 1): We have included a comprehensive flowchart that systematically illustrates the entire research workflow, from data acquisition and preprocessing, through differential expression analysis, WGCNA, machine learning, and hub gene screening, to drug prediction. This visual summary helps readers quickly grasp the overall study design. Tabular presentation of machine learning features and performance (new S2 Table and S3 Table): The pivotal diagnostic genes for H. pylori and AF (previously shown only in Figure 4A and 4F) are now summarized, which also lists each gene's SHAP value and ranking, enabling quantitative assessment of feature importance. The diagnostic performance metrics (AUROC) of the final models are now organized in S4 Table and S5 Table, with results clearly labeled for both the discovery cohort (internal cross-validation) and independent validation cohorts. This makes the model validation process more transparent and credible. Visualization of batch effect correction (new Fig S2): To address concerns regarding dataset heterogeneity, we have added PCA plots and boxplots for both AF and H. pylori datasets before and after ComBat correction. This figure intuitively demonstrates that batch effects were effectively eliminated and samples from different datasets were well-integrated, providing visual justification for subsequent combined analyses. Complete list of core genes (new S1 Table): We have provided the full list of 77 union genes (including gene symbols, full names, and expression patterns) as a supplementary table, facilitating direct access and analysis by readers. All figures and tables have been prepared according to journal formatting guidelines (appropriate resolution, font size, layout), and the legends have been rewritten to be self-explanatory. We believe these additions substantially enhance the readability and evidentiary value of our results. Comment 2: The Discussion remains largely narrative. It would benefit from deeper analysis rather than speculative expansion. In particular, the Limitations section should be substantially expanded to address dataset heterogeneity, model stability, lack of causal inference, and the non-predictive nature of molecular docking. Response: We thank you for your insightful comments on the Discussion. In our revision, we have placed particular emphasis on enhancing the depth and critical analysis of this section: Strengthened interpretation grounded in results: The newly added figures and tables (Figure 1 flowchart, Tables 2 and 3, Supplementary Figure 2, etc.) provide a more robust data foundation for the core arguments in the Discussion. We have closely integrated these new results into the Discussion, offering in-depth interpretations of key findings (such as the dual role of S100A8 and the immune network status of hub genes) while avoiding unsubstantiated speculation. Clear distinction between conclusions and hypotheses: We have carefully reviewed all discussion statements to ensure that every inference is grounded in our results or published literature. Hypothetical points (such as the therapeutic potential of candidate drugs) are expressed with appropriate caution, explicitly noting the need for future validation. Expanded limitations discussion: Following your suggestion, we have further emphasized the following aspects in the "Limitations and Future Perspectives" subsection: Dataset heterogeneity: Although ComBat correction was applied, the original data originated from different platforms; future validation using prospective, uniformly processed cohorts is still warranted. Model stability: While the models performed well in independent validation, broader external validation in multi-center studies would further strengthen their generalizability. Lack of causal inference: This study is associational and cannot establish causality; future studies employing Mendelian randomization or longitudinal designs are needed to explore causal relationships. Predictive nature of molecular docking: The docking results are computational predictions; their biological activity, specificity, and safety require experimental validation. These additions have made the Limitations section more comprehensive and balanced, reflecting our clear awareness of the scientific boundaries of our conclusions. We once again thank you for taking the time to review our manuscript. Your rigorous and detailed feedback has pushed us to pursue higher standards. We would also like to express our sincere gratitude to Reviewer #2 for their positive evaluation and recommendation for publication. The professional insights from both reviewers have jointly contributed to the improvement of this study, and we are deeply appreciative. We hope that the revised manuscript now meets your expectations. Sincerely. Attachment Submitted filename: To_the_Reviewers_auresp_2.docx pone.0346038.s010.docx (13.9KB, docx) PLoS One. doi: 10.1371/journal.pone.0346038.r005 Decision Letter 2 Tomasz Kaminski Tomasz Kaminski Academic Editor Find articles by Tomasz Kaminski Author information Copyright and License information Roles Tomasz Kaminski : Academic Editor © 2026 Tomasz KaminskiTomasz Kaminski This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice 15 Mar 2026 Integrative Bioinformatics and Machine Learning Identify Shared Molecular Mechanisms and Diagnostic Biomarkers Between Helicobacter pylori Infection and Atrial Fibrillation PONE-D-25-59195R2 Dear Dr. wang, We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements. Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication. An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support .. If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact [email protected]. Kind regards, Tomasz W. Kaminski Academic Editor PLOS One PLoS One. doi: 10.1371/journal.pone.0346038.r006 Acceptance letter Tomasz Kaminski Tomasz Kaminski Academic Editor Find articles by Tomasz Kaminski Author information Copyright and License information Roles Tomasz Kaminski : Academic Editor © 2026 Tomasz KaminskiTomasz Kaminski This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited., which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. PMC Copyright notice PONE-D-25-59195R2 PLOS One Dear Dr. Wang, I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team. At this stage, our production department will prepare your paper for publication. This includes ensuring the following: * All references, tables, and figures are properly cited * All relevant supporting information is included in the manuscript submission, * There are no issues that prevent the paper from being properly typeset You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps. Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact [email protected]. You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing . If we can help with anything else, please email us at [email protected]. Thank you for submitting your work to PLOS ONE and supporting open access. Kind regards, PLOS ONE Editorial Office Staff on behalf of Dr. Tomasz W. Kaminski Academic Editor PLOS One Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials S1 Fig. Functional enrichment and pathway enrichment analysis of DEGs in H. pylori infection and AF. Dot plot (A, D) and circular plot (B, E) displaying the results of the GO enrichment analysis of DEGs specific to H. pylori infection and AF. KEGG enrichment analysis results of DEGs specific to H. pylori infection (C) and AF (F). (TIF) pone.0346038.s001.tif (9.8MB, tif) S2 Fig. Batch effect correction of AF and H. pylori datasets. (TIF) pone.0346038.s002.tif (2.9MB, tif) S1 Table. 77 union genes for comprehensive functional annotation and machine leaming. (DOCX) pone.0346038.s003.docx (14.4KB, docx) S2 Table. AUC values for various models across cohorts in H. pylori infection. (DOCX) pone.0346038.s004.docx (15.3KB, docx) S3 Table. AUC values for various models across cohorts in AF. (DOCX) pone.0346038.s005.docx (16.9KB, docx) S4 Table. The AUC values of the top 5 genes incorporated in the optimal ML methods for H. pylori infection. (DOCX) pone.0346038.s006.docx (10.9KB, docx) S5 Table. The AUC values of the top 5 genes incorporated in the optimal ML methods for AF. (DOCX) pone.0346038.s007.docx (10.9KB, docx) Attachment Submitted filename: To the Reviewers.docx pone.0346038.s009.docx (14.5KB, docx) Attachment Submitted filename: To_the_Reviewers_auresp_2.docx pone.0346038.s010.docx (13.9KB, docx) Data Availability Statement All sequencing data were sourced from the Gene Expression Omnibus (GEO) database. A total of four datasets related to AF were included: GSE41177 and GSE176166 constituted the discovery cohort, while GSE108660 and GSE128188 formed the validation cohort. For H. pylori infection, five datasets were selected: GSE27411, GSE60662, and GSE233973 comprised the discovery cohort, and GSE5081 and GSE60427 served as the validation cohort. Articles from PLOS One are provided here courtesy of PLOS ACTIONS View on publisher site PDF (3.9 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top