ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Mapping disease loci to biological processes via joint pleiotropic and epigenomic partitioning.

Kerner G et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed systems architecture

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Cell Genom . 2026 Jan 26;6(4):101138. doi: 10.1016/j.xgen.2025.101138 Search in PMC Search in PubMed View in NLM Catalog Add to search Mapping disease loci to biological processes via joint pleiotropic and epigenomic partitioning Gaspard Kerner Gaspard Kerner 1 Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA 2 Broad Institute of MIT and Harvard, Boston, MA, USA Find articles by Gaspard Kerner 1, 2, 6, ∗ , Nolan Kamitaki Nolan Kamitaki 3 Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA Find articles by Nolan Kamitaki 3 , Benjamin Strober Benjamin Strober 1 Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA 4 Computational Health Informatics Program, Boston Children’s Hospital, Boston, MA, USA Find articles by Benjamin Strober 1, 4 , Alkes L Price Alkes L Price 1 Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA 2 Broad Institute of MIT and Harvard, Boston, MA, USA 5 Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA Find articles by Alkes L Price 1, 2, 5, ∗∗ Author information Article notes Copyright and License information 1 Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA 2 Broad Institute of MIT and Harvard, Boston, MA, USA 3 Department of Biomedical Informatics, Harvard Medical School, Boston, MA, USA 4 Computational Health Informatics Program, Boston Children’s Hospital, Boston, MA, USA 5 Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA ∗ Corresponding author [email protected] ∗∗ Corresponding author [email protected] 6 Lead contact Received 2025 Jun 4; Revised 2025 Oct 9; Accepted 2025 Dec 30; Collection date 2026 Apr 8. © 2026 The Authors This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). PMC Copyright notice PMCID: PMC13069866  PMID: 41592567 Previous version available: This article is based on a previously available preprint posted on medRxiv on May 22, 2025: " Mapping disease loci to biological processes via joint pleiotropic and epigenomic partitioning ". Summary Genome-wide association studies have identified thousands of disease-associated loci, yet their biological interpretation remains limited. We propose joint pleiotropic and epigenomic partitioning (J-PEP), a clustering framework that integrates pleiotropic SNP effects on auxiliary traits and tissue-specific epigenomic data to partition disease-associated loci into biologically distinct clusters. We introduce a metric—pleiotropic and epigenomic prediction accuracy (PEPA)—that evaluates how well the clusters predict SNP-to-trait and SNP-to-tissue associations in off-chromosome data. Analyzing summary statistics for 165 diseases/traits (average N = 290,000), J-PEP attained 16%–30% higher PEPA than pleiotropic or epigenomic partitioning approaches, with larger improvements for well-powered traits, consistent with simulations; these gains arise from J-PEP’s tendency to upweight signals present in both auxiliary trait and tissue data, emphasizing shared components. Notably, integrating single-cell chromatin accessibility data refined bulk-based clusters, enhancing cell-type resolution and specificity. For type 2 diabetes, hypertension, and other diseases/traits, J-PEP clusters recapitulated known pathways while revealing underexplored biological processes. Keywords: GWAS, pleiotropic partitioning, epigenomics, polygenic risk scores, partitioned polygenic risk scores, biobank, complex diseases, type 2 diabetes, hypertension, prediction Graphical abstract Open in a new tab Highlights • J-PEP integrates pleiotropic and epigenomic data to cluster GWAS loci • Cross-modal PEPA metric quantifies trait-tissue prediction accuracy • J-PEP improves biological interpretability across 165 diseases/traits • Single-cell epigenomics refines clusters to cell-type resolution Kerner et al. introduce J-PEP, a framework that integrates pleiotropic and epigenomic data to partition GWAS loci into biologically meaningful clusters. Using a new cross-modal prediction metric and single-cell epigenomics, they reveal tissue- and cell-type-specific processes underlying diverse complex diseases. Introduction A central challenge in complex disease genomics is to decipher the biological processes underlying the thousands of loci identified by genome-wide association studies (GWASs). 1 , 2 , 3 , 4 , 5 , 6 While interpretation of individual loci has yielded important insights into disease biology and functionally validated putative causal genes, 7 , 8 , 9 , 10 , 11 findings at individual loci remain limited, 1 , 12 motivating genome-wide approaches to map disease loci to biological processes. Extensive pleiotropy between diseases and complex traits 13 , 14 , 15 , 16 , 17 , 18 , 19 , 20 , 21 , 22 motivates pleiotropic partitioning 23 , 24 , 25 (see also Zhang et al., 26 , 27 , 28 , 29 , 30 , 31 Kim et al., 26 , 27 , 28 , 29 , 30 , 31 Reeve et al., 26 , 27 , 28 , 29 , 30 , 31 Qi et al., 26 , 27 , 28 , 29 , 30 , 31 Ghatan et al., 26 , 27 , 28 , 29 , 30 , 31 and Li et al. 26 , 27 , 28 , 29 , 30 , 31 )—leveraging pleiotropic associations with auxiliary traits to group disease loci into clusters that often align with disease-critical cell types and biological pathways. However, interpreting clusters based solely on pleiotropy can be challenging when biological processes are partially overlapping or produce similar multi-trait signatures. Epigenomic annotations provide valuable information about disease biology 32 , 33 , 34 , 35 , 36 , 37 and have been used to retrospectively interpret clustering results 23 , 24 , 25 but have not been directly integrated into clustering models. Furthermore, no quantitative metric has been proposed to evaluate clustering performance. Here, we introduce joint pleiotropic and epigenomic partitioning (J-PEP), a method that jointly partitions disease loci into clusters using both pleiotropic associations with auxiliary traits and tissue-specific epigenomic profiles under a tissue sparsity constraint that improves algorithmic stability and interpretability; J-PEP clusters capture both SNP-to-trait and SNP-to-tissue associations, enhancing cluster interpretability. We also introduce a new evaluation metric, pleiotropic and epigenomic prediction accuracy (PEPA), which evaluates how well the clusters predict SNP-to-trait and SNP-to-tissue associations. We compare J-PEP to pleiotropic and epigenomic partitioning approaches via simulations and analyses of 165 GWAS diseases/traits using the PEPA metric. We find that J-PEP outperforms other approaches using the PEPA metric, and we highlight novel biological insights for type 2 diabetes (T2D), hypertension (HTN), and neutrophil count (NC). Finally, we show that by integrating single-cell chromatin accessibility data, J-PEP enables the association of clusters—i.e., subsets of loci—with fine-grained cell types, refining broader bulk-derived associations. Results Overview of methods J-PEP partitions disease loci into biologically interpretable clusters by jointly leveraging pleiotropic data (across auxiliary traits) and epigenomic data (across tissues) ( Figure 1 ). For a given focal disease/trait, J-PEP inputs a SNP-to-trait matrix ( V trait ) derived from fine-mapping results 38 , 39 across auxiliary traits, with each entry defined as the product of posterior inclusion probabilities (PIPs) for a given fine-mapped SNP (focal trait PIP > 0.01) across the focal trait and a given auxiliary trait ( STAR Methods ). To accommodate the sign of pleiotropic effects, each auxiliary trait is split into two components: one representing concordant (or positive) pleiotropic effects and the other representing discordant (or negative) pleiotropic effects ( STAR Methods ). J-PEP also inputs a SNP-to-tissue matrix ( V tissue ) derived from epigenomic data across tissues, where each entry reflects the normalized strength of association between a fine-mapped SNP and a tissue, computed using an expectation maximization (EM) algorithm ( STAR Methods ). J-PEP then jointly factorizes V trait and V tissue to produce estimates of a SNP-to-cluster membership matrix ( W ), an auxiliary trait-to-cluster profile matrix ( H trait ), and a tissue-to-cluster profile matrix ( H tissue ), using an extended version of joint Bayesian non-negative matrix factorization (bNMF). 40 , 41 Standard bNMF decomposes a single non-negative data matrix into lower-dimensional non-negative factors under Bayesian priors; J-PEP extends this framework to two matrices, encouraging cross-data alignment across clusters (see below). To enhance cluster interpretability, J-PEP restricts each tissue to at most a single cluster, whereas auxiliary traits may span multiple clusters. This iterative optimization process continues until convergence, yielding clusters that integrate pleiotropic and epigenomic information. To enhance stability, we perform ten optimization runs with different random seeds and select the run with the highest average projection correlation between modalities (i.e., traits and tissue) among those yielding the most consistent number of identified clusters ( STAR Methods ). We note that of all input/output matrices, only V tissue has normalized rows that specifically represent probabilities. Figure 1. Open in a new tab Schematic representation of the J-PEP model The J-PEP model performs an extended version of joint Bayesian non-negative matrix factorization (bNMF) to fine-mapped SNPs from a given focal disease/trait. It jointly factorizes two input matrices: a SNP-to-trait matrix ( V trait ) and a SNP-to-tissue matrix ( V tissue ). J-PEP infers a shared SNP-to-cluster membership matrix ( W ), an auxiliary trait-to-cluster profile matrix ( H trait ) and a tissue-to-cluster profile matrix ( H tissue ). In detail, J-PEP solves the following optimization problem: { W , H t r a i t , H t i s s u e } = argmin W ≥ 0 , H t r a i t ≥ 0 , H t i s s u e ≥ 0 { 1 2 ‖ V t r a i t − W H t r a i t ‖ 2 + 1 2 ‖ V t i s s u e − W H t i s s u e ‖ 2 + ∑ k λ ( ( V t r a i t H t r a i t T ) k , ( V t i s s u e H t i s s u e T ) k ) } , (Equation 1) where W , H trait , H tissue , V trait , and V tissue are as defined above, and ( V t r a i t H t r a i t T ) k and ( V t i s s u e H t i s s u e T ) k are vectors of length M (the number of focal disease/trait fine-mapped SNPs with PIP > 0.01) representing the projections of fine-mapped SNPs onto the k th cluster as defined by auxiliary traits and tissues, respectively. To guide the factorization process, the first two terms of the optimization function minimize the reconstruction error of the input matrices, while the third term is a penalization term that promotes alignment between pleiotropic and epigenomic data and suppresses irrelevant clusters. Specifically, λ ( A , B ) = − log 10 ( p c o r r ( A , B ) ) − 1 , where p corr indicates the one-sided Pearson correlation p value for positive correlation. Additionally, J-PEP guarantees sparsity in tissue profiles by setting all non-maximal tissue-to-cluster memberships (columns of H tissue ) to zero, leaving each tissue assigned to at most a single cluster. Conceptually, J-PEP optimizes the objective function of Equation 1 under the constraint that H tissue ∈ C , where C is the set of matrices whose columns each have at most one non-zero entry. The initial number of clusters (set to 15 following prior work 23 ) serves as an upper bound and may be reduced during optimization as uninformative clusters are pruned. We systematically compared J-PEP to alternative methods for partitioning disease-associated loci. First, we considered two simpler single-modality approaches: pleiotropic partitioning, 23 , 24 , 25 which factorizes a SNP-to-trait matrix (e.g., V trait ) without incorporating tissue information; and epigenomic partitioning, which factorizes a SNP-to-tissue matrix (e.g., V tissue ) without incorporating pleiotropic information ( STAR Methods ). In both approaches, partitioning is performed using standard bNMF. For pleiotropic partitioning, we learn H tissue from V tissue and W using gradient descent by solving V tissue = W × H tissue . For epigenomic partitioning, we learn H trait from V trait and W using gradient descent by solving V trait = W × H trait . To support comparisons between methods, the sparsity constraint on tissue profiles applied to J-PEP was also imposed on the two simpler approaches ( STAR Methods ). Second, we benchmarked J-PEP against two independently developed software packages that perform pleiotropic partitioning: FactorGO 26 and Flashier. 42 FactorGO is a variational Bayesian factor analysis model that decomposes GWAS Z -score matrices into latent pleiotropic factors by modeling true genetic effects with uncertainty, using priors to prune uninformative factors. Flashier is an empirical Bayes matrix factorization framework that assumes the observed data can be approximated by the product of two low-rank latent matrices (with Gaussian noise) and places priors on factors; we benchmarked three versions of Flashier that differ in their prior specifications, regularization strategies, and input data. We applied J-PEP, pleiotropic partitioning, and epigenomic partitioning to GWAS summary statistics for 165 focal diseases/traits (average of 290,000 samples and 77 fine-mapped loci; a fine-mapped locus is defined as a locus containing at least one SNP with PIP > 0.5), including 57 UK Biobank diseases/traits 43 and 108 non-UK Biobank diseases/traits ( Table S1 ; see data and code availability ). We integrated EpiMap data 35 comprising 1,595 chromatin tracks, each defined by a unique combination of one of six histone or accessibility marks (DNase HS, H3K4me1, H3K4me2, H3K4me3, H3K9ac, and H3K27ac) and a cell-type label. These tracks are grouped into 32 broad tissue categories defined by EpiMap 35 ( STAR Methods ; see data and code availability ); we refined our interpretation of J-PEP clusters using single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq), 44 profiling 615,998 nuclei across 30 adult human tissues, ultimately resolving 111 fine-grained cell types. We restricted most analyses to 55 approximately independent focal diseases/traits (pairwise genetic correlation [ r g ] < 0.5), prioritizing traits with greater statistical power (defined by the number of fine-mapped loci); focal traits yielding fewer than two clusters for any of the three partitioning methods were excluded from most comparisons, yielding 38 focal diseases/traits ( STAR Methods ). Auxiliary traits for each focal trait were selected from the full set of 165 diseases/traits based on two criteria: (1) a genetic correlation of <0.5 with the focal trait (other thresholds were also tested); and (2) the existence of at least one shared causal SNP (PIP > 0.5) with the focal trait. On average, a focal disease/trait had 58 auxiliary traits. To improve cluster interpretability in selected real trait examples, we also conducted Gene Ontology (GO) enrichment 45 informed by SNP-to-gene linking scores from Gazal et al. 46 ( STAR Methods ). To assess the performance of J-PEP and compare it to other clustering approaches, we developed PEPA, a new validation metric that evaluates how well a method predicts pleiotropic information (SNP-to-trait associations; V trait ) from epigenomic information (SNP-to-tissue associations; V tissue ) and vice versa. PEPA employs an off-chromosome prediction framework, 47 , 48 , 49 , 50 where entries in V trait and V tissue , respectively, are predicted from the corresponding entries in V tissue and V trait , respectively, using cluster-specific profiles ( H trait and H tissue ) estimated using off-chromosome data. We note that the off-chromosome prediction framework ensures that this metric is robust to overfitting. PEPA is defined as the geometric mean of two constituent metrics: pleiotropic prediction accuracy (PPA), which measures how well SNP-to-trait associations ( V trait ) can be predicted from SNP-to-tissue associations ( V tissue ); and epigenomic prediction accuracy (EPA), which measures how well SNP-to-tissue associations ( V tissue ) can be predicted from SNP-to-trait associations ( V trait ). We use the geometric mean of PPA and EPA, rather than the arithmetic mean, because it prioritizes clusters with balanced support from both auxiliary traits and tissues while penalizing imbalance between modalities, reflecting the overarching goal of J-PEP. We also note that even though different inference approaches can differ in the number of discovered clusters, the PEPA metric ensures comparability across methods because it evaluates prediction accuracy in the observed data space, independent of the number of inferred clusters. Standard errors on each metric are computed using a genomic block jackknife, similarly to Finucane et al. 34 Further details are provided in STAR Methods . Simulations To evaluate the performance of different clustering methods, we performed simulations by simulating GWAS summary statistics at real SNPs, incorporating genetic architectures informed by predefined cluster profiles ( STAR Methods ). Our primary simulations included 32 tissues and 20 auxiliary traits, all of which were genetically uncorrelated with one another. In brief, we simulated data using K causal = 5 predefined causal clusters and sampled 100–400 causal SNPs from chromosomes 1–2 of the 1000 Genomes Project 51 (1KG) according to their aggregated cluster memberships, generating correlated SNP-to-trait ( V trait ) and SNP-to-tissue ( V tissue ) profiles. Using 1KG linkage disequilibrium (LD), we then simulated GWAS summary statistics and applied single-causal-variant fine mapping to match our real-data analyses ( STAR Methods ; see Figure S1 for an illustrative example). Our primary simulations also included uncorrelated structure, i.e., structure within V tissue and V trait , respectively, that is not correlated to any structure in V trait and V tissue , respectively, by injecting random subsets of s non-causal tissues and s non-causal traits with permuted profiles from causal ones ( STAR Methods ); we also performed simulations with other parameter configurations ( STAR Methods ). We note that uncorrelated structure is likely to arise in real data, e.g., if the relevant tissues or cell types driving pleiotropic associations to auxiliary traits are missing or under-represented in epigenomic annotations, or analogously if a cell-type-specific disease process is not captured by the available auxiliary traits. We first evaluated the accuracy of pleiotropic partitioning, epigenomic partitioning, and J-PEP in reconstructing the auxiliary trait-to-cluster profile matrix ( H trait ) and tissue-to-cluster profile matrix ( H tissue ) by computing the Frobenius norm error (defined as the root square of the sum of squared element-wise differences). These errors were averaged across 100 simulation replicates under four simulation settings, corresponding to m = 100, 200, 300, or 400 causal SNPs for the focal disease/trait. We determined that J-PEP attained the lowest average Frobenius norm error for both H tissue and H trait , with improvements becoming more pronounced as the number of causal SNPs increased ( Figure 2 A and Table S2 ). This indicates that J-PEP’s enhanced predictive performance is not solely driven by the constraint of identifying shared structures between traits and tissues, which is enforced by the third term of the optimization function in Equation 1 , but also benefits from the first two terms of Equation 1 , which ensure that the inferred cluster memberships ( W ) remain closely aligned with the auxiliary trait and tissue data. Figure 2. Open in a new tab J-PEP outperforms pleiotropic partitioning and epigenomic partitioning in simulations (A) Frobenius norm errors between the true vs. inferred auxiliary trait-to-cluster profile matrix ( H trait ) and tissue-to-cluster profile matrix ( H tissue ). (B) PEPA, PPA, and EPA prediction accuracy metrics. (C) PEPA metric across downsampling proportions of non-causal SNPs. (D) PEPA metric across various PIP thresholds. In (A)–(D), results represent averages across 100 simulation replicates under four simulation settings with 100, 200, 300, or 400 causal SNPs for the focal disease/trait (denoted along the x axis as m = 100, 200, 300, or 400, respectively). (E) Comparison between pleiotropic partitioning, epigenomic partitioning, J-PEP, FactorGO, and three versions of Flashier across three metrics: PEPA (left), Frobenius norm error for H trait (center), and Frobenius norm error for H tissue (right). Error bars denote standard errors across replicates. p values were computed using paired two-sided Wilcoxon signed-rank tests. ∗ p < 0.05, ∗∗ p < 0.01, ∗∗∗ p < 0.001. Numerical results are reported in Table S2 . Next, we evaluated the accuracy of pleiotropic partitioning, epigenomic partitioning, and J-PEP using our new metric (PEPA) and its two constituent metrics (PPA and EPA). The average PEPA achieved by J-PEP increased with larger values of m , showing progressively greater and more significant improvements over both pleiotropic and epigenomic partitioning as m increased ( Figure 2 B and Table S2 ). At m = 200, J-PEP outperformed pleiotropic and epigenomic partitioning by 18% (paired t test p = 1.2 × 10 −4 ) and 8% ( p = 0.04), respectively; at m = 400, these gains rose to 19% ( p = 2.0 × 10 −10 ) and 38% ( p = 7.8 × 10 −20 ). Among the two constituent methods, pleiotropic partitioning performed better at predicting SNP-to-tissue associations ( V tissue ) from SNP-to-trait associations ( V trait ) (EPA), while epigenomic partitioning performed better at predicting SNP-to-trait associations ( V trait ) from SNP-to-tissue associations ( V tissue ) (PPA); this is consistent with pleiotropic partitioning and epigenomic partitioning yielding lower error when reconstructing the SNP-to-cluster membership matrix ( W ) from SNP-to-trait and SNP-to-tissue association data, respectively, as evaluated in the respective EPA and PPA metrics—compared to reconstructing W from SNP-to-tissue and SNP-to-trait data as evaluated in PPA and EPA, respectively. To investigate the impact of LD and PIP thresholds on model performance, we conducted two additional analyses. First, we downsampled non-causal SNPs, thereby reducing the number of tagging SNPs contributing LD structure ( Figure 2 C and Table S2 ). As expected, performance of all three methods improved as downsampling increased, and the data more closely reflected the true causal signals; importantly, J-PEP consistently maintained its relative advantage over pleiotropic and epigenomic partitioning across all settings. Second, we varied the PIP threshold used to define the fine-mapped SNPs included as input to each method, testing both stricter (≥0.05) and more lenient (down to 0.001) thresholds ( Figure 2 D and Table S2 ). This analysis revealed that applying stricter thresholds reduced prediction accuracy while including additional low-PIP SNPs improved prediction accuracy across all three methods (except that accuracy plateaued for epigenomic partitioning). This suggests that retaining a larger set of SNPs, weighted by their posterior probabilities, allows the models to capture additional true information and attain improved accuracy (consistent with expectations under well-calibrated fine mapping 52 ). J-PEP continued to outperform pleiotropic and epigenomic partitioning at more lenient PIP thresholds, but not at stricter PIP thresholds, perhaps because J-PEP benefits from a sufficiently large set of causal SNPs to reveal concordant structure between pleiotropic effects and epigenomic annotations; when too many informative SNPs are removed, the shared signal across the two data sources becomes too sparse to provide additional gains over single-modality approaches. Based on these results, we used a PIP threshold of 0.01 for all real trait analyses in this study (limiting computational cost). Finally, we benchmarked FactorGO 26 and three versions of Flashier 42 against J-PEP ( Figure 2 E and Table S2 ). With respect to reconstruction error (Frobenius norm), FactorGo performed worse than J-PEP, but Flashier (particularly the semi-non-negative version, Flashier_snn) performed almost as well as J-PEP in some cases, consistent with its design to optimize reconstruction of the single modality being partitioned. However, FactorGO and Flashier both performed far worse than J-PEP with respect to our cross-modal PEPA metric, consistent with the fact that FactorGO and Flashier do not consider tissue data. This suggests that unlike FactorGO and Flashier, J-PEP recovers clusters that align traits with tissues, making them more biologically meaningful by capturing disease processes in a tissue context. We performed several secondary analyses. First, we formally evaluated receiver operating characteristic area under the curve (ROC AUC), sensitivity, and specificity using a co-clustering framework: for each pair of SNPs, we assessed the extent to which they co-cluster, taking into account their probabilistic cluster assignments ( Figure S2 ). J-PEP attained a mean ROC AUC of 0.5237, closely followed by pleiotropic partitioning (0.5236, two-sided paired t test p = 0.38 for the difference with J-PEP across replicates) and well above epigenomic partitioning (0.5017, p < 10 −12 ). In terms of sensitivity and specificity, J-PEP consistently attained higher sensitivity than pleiotropic partitioning and higher specificity than epigenomic partitioning. Pleiotropic partitioning attained high specificity but low sensitivity, consistent with the fact that shared genetic effects across traits are sparse, hence stringent; epigenomic partitioning attained high sensitivity (at low PIP thresholds) but low specificity, as SNPs tend to overlap many regulatory annotations, capturing true co-clustering while also introducing false co-clustering. Across six additional simulation analyses, we found that J-PEP’s advantages were robust across generative settings ( Figures S3 and S4 ). Performance gains increased when uncorrelated structure was present and persisted across variations in the number of causal clusters, maximum allowed K , and the proportion of missing causal tissues; reducing K improved PEPA, and when half of the causal tissues were absent but auxiliary traits were available, J-PEP achieved even larger improvements over pleiotropic partitioning. We then considered a version of J-PEP in which the tissue sparsity constraint is directly included in the objective function, encouraging but not enforcing sparsity in H tissue (J-PEP_L1; STAR Methods ), instead of the default “hard max sparsity constraint” of J-PEP. J-PEP_L1 produced very similar results to J-PEP when convergence was achieved ( Figure S5 A). However, J-PEP_L1 failed to converge or did not identify two distinct clusters in 37% of simulation replicates, whereas J-PEP converged in all replicates ( Figure S5 B). Finally, across simulations varying dimensionality (numbers of auxiliary traits and tissues), SNP-inclusion thresholds, and partitioning methods ( Figures S6 , S7 A, and S7B), J-PEP showed stable PEPA behavior and remained computationally tractable, except at very low PIP thresholds (<0.01) where computational cost became prohibitive. Application of J-PEP to 165 diseases and complex traits We applied J-PEP, pleiotropic partitioning, and epigenomic partitioning to single causal variant fine-mapping results of GWAS summary statistics for 165 focal diseases/traits (average of 290,000 samples and 77 fine-mapped loci; Table S1 ), focusing our comparisons on 38 approximately independent diseases/traits for which all three methods identified at least two clusters. We performed single causal variant fine mapping instead of multiple causal variant fine mapping, as this is the recommended approach when in-sample LD is not available. 52 Across these 38 traits, J-PEP attained significantly higher and lower) PEPA than pleiotropic partitioning (genomic block jackknife p < 0.05) for nine traits and two traits, respectively, and significantly higher and lower PEPA than epigenomic partitioning for six traits and no traits, respectively ( Figure 3 A and Table S3 ). J-PEP attained the most significant improvements for immune-related traits (e.g., blood cell counts; Figure 3 A) and metabolic traits (e.g., total cholesterol; Table S3 ). In contrast, pleiotropic partitioning outperformed J-PEP for one brain-related trait (fractional anisotropy; Table S3 ) and one lung-related trait (FEV 1 ) ( Table S3 ); in each case, J-PEP’s inferred cluster tissue profiles ( H tissue ) lacked associations with brain and lung annotations, respectively, despite significant enrichment of these and other tissues in the corresponding SNP-to-tissue association matrix ( V tissue ) matrices. Figure 3. Open in a new tab J-PEP outperforms pleiotropic partitioning and epigenomic partitioning in analyses of real diseases/traits (A) PEPA prediction accuracy metric for 20 focal diseases/traits with the largest number of fine-mapped loci (denoted in parentheses). Error bars denote standard errors (via genomic block jackknife). (B) PEPA, PPA, and EPA prediction accuracy metrics averaged across 38, 33, 24, or 16 focal traits (of a total of 38) with at least 0, 10, 50, or 100 fine-mapped loci, respectively. (C and D) Analogous to (B), with the genetic correlation ( r g ) threshold for auxiliary traits changed from r g < 0.5 to r g < 0.2 and r g < 0.9, respectively. Error bars denote standard errors (via genomic block jackknife). p values were computed using two-sided genomic block jackknife tests on the difference between J-PEP and each other method. ∗ p < 0.05, ∗∗ p < 0.01, ∗∗∗ p < 0.001 (vs. each other method). Numerical results for all 165 focal diseases/traits are reported in Table S3 . MSCV, mean sphered corpuscular volume; RDWCV, red cell distribution width—coefficient of variation; PDW, platelet distribution width; HLRP, high light reticulocyte proportion; FEV1FVC, forced expiratory volume in 1 s/forced vital capacity. Averaging across the 38 focal diseases/traits, J-PEP attained 16% higher PEPA than pleiotropic partitioning (genomic jackknife on the difference p = 2.4 × 10 −2 ; p from inverse-variance weighting [ p w ] = 1.2 × 10 −3 ; genomic block jackknife on the average across 38 traits; STAR Methods ) and 22% higher PEPA than epigenomic partitioning ( p = 2.6 × 10 −4 ; p w = 3.2 × 10 −5 ) ( Figure 3 B and Table S3 ). These improvements are consistent with our primary simulations, suggesting that real disease/trait data harbor uncorrelated structure, i.e., structure within V tissue and V trait that is not correlated to any structure in V trait and V tissue , respectively. Notably, J-PEP’s performance advantage grew more pronounced in analyses of well-powered traits with more fine-mapped loci, also consistent with simulation results. For example, J-PEP outperformed pleiotropic partitioning by 27% ( p = 1.6 × 10 −3 ; p w = 2.0 × 10 −3 ) and epigenomic partitioning by 30% ( p = 1.4 × 10 −8 ; p w = 3.1 × 10 −8 ) across 16 focal traits with at least 100 fine-mapped loci ( Figure 3 B and Table S3 ). We generally observed stronger concordance between J-PEP and epigenomic partitioning than between J-PEP and pleiotropic partitioning (average maximum SNP-level correlation of 0.86 between J-PEP and epigenomic partitioning clusters and 0.25 between J-PEP and pleiotropic partitioning clusters across all 38 focal diseases/traits; Table S4 ; see below). Pleiotropic partitioning performed better at predicting SNP-to-tissue associations ( V tissue ) from SNP-to-trait associations ( V trait ) (EPA), while epigenomic partitioning performed better at predicting SNP-to-trait associations ( V trait ) from SNP-to-tissue associations ( V tissue ) (PPA), consistent with our simulations. As expected, a given constituent method tends to perform well when J-PEP significantly outperforms the other constituent method ( Figure S8 ). These results demonstrate that J-PEP consistently improves tissue-trait prediction accuracy across a broad set of diseases/traits, particularly in analyses of well-powered traits. The above analyses restricted auxiliary traits to those with r g < 0.5 with the focal trait (average of 58 auxiliary traits per focal trait). To assess the impact of this choice, we repeated our analyses using a more stringent threshold ( r g < 0.2; average of 55.3 auxiliary traits) or a more lenient threshold ( r g < 0.9; average of 62.5 auxiliary traits) relative to the default threshold ( r g < 0.5; average of 61.3 auxiliary traits), restricting all comparisons to the 32 focal traits for which all tested thresholds yielded at least two clusters ( Figures 3 C and 3D; Table S3 ). We determined that the more stringent threshold ( r g < 0.2) led to a reduction in accuracy as quantified by the PEPA metric (e.g., J-PEP PEPA = 0.040 vs. 0.050), with J-PEP continuing to outperform pleiotropic partitioning and epigenomic partitioning by a similar margin. The more lenient threshold ( r g < 0.9) slightly increased prediction accuracy (e.g., J-PEP PEPA = 0.056 vs. 0.050), with J-PEP continuing to outperform pleiotropic partitioning and epigenomic partitioning by a similar margin. However, six of the 38 focal traits that produced at least two clusters at r g < 0.5 failed to do so at r g < 0.9, all of which had gained new auxiliary traits. Conversely, no focal trait yielded at least two clusters at r g < 0.9 and not at r g < 0.5. This suggests that inclusion of highly correlated traits may destabilize the model. For this reason, we decided not to increase the default threshold to r g < 0.9. We performed several secondary analyses. First, we investigated whether the stronger concordance between J-PEP and epigenomic partitioning than between J-PEP and pleiotropic partitioning ( Table S4 ) could arise from differences in the scaling of the two input matrices. In the default version of J-PEP, entries of the SNP-to-trait matrix ( V trait ) are constructed from products of (focal trait PIP) × (auxiliary trait PIP), whereas entries of the SNP-to-tissue matrix ( V tissue ) are individual probabilities that are not weighted by (focal trait PIP). We performed a new analysis of real traits in which we weighted V tissue by (focal trait PIP), so that V trait and V tissue are on the same scale with respect to focal trait PIPs. In this analysis, we observed stronger concordance between J-PEP and pleiotropic partitioning than between J-PEP and epigenomic partitioning ( Figure S5 C). We elected to continue to not weight V tissue by (focal trait PIP) in the default version of J-PEP because J-PEP is explicitly designed to integrate tissue and pleiotropic signals indirectly through joint factorization, and weighting V tissue by (focal trait PIP) may suppress much of the cross-tissue structure captured by the epigenomic data, likely contributing to the frequent failures to converge or recover clusters. We also investigated whether the stronger concordance between J-PEP and epigenomic partitioning than between J-PEP and pleiotropic partitioning could arise from the sparsity constraint on tissue-cluster membership but concluded that this was not the case ( Figures S5 C and S5D). Across an additional comprehensive set of secondary analyses ( Figures S7–S14 ; Tables S5 , S6 , and S7 ), J-PEP substantially outperformed FactorGO 26 and Flashier 42 with respect to the cross-modal PEPA metric, consistent with simulations and with the fact that these methods do not use tissue data. J-PEP also behaved robustly across traits and parameter settings: cluster resolution increased with GWAS power; SNPs and loci were well differentiated; aggregated tissue profiles recapitulated expected biology; reducing the maximum number of clusters modestly improved PEPA without changing method rankings; tissue-to-cluster scores reflected annotation density; running times remained low; and all runs converged. Single-cell accessibility profiles enhance cell-type resolution and specificity We sought to assess whether single-cell chromatin accessibility profiles could help interpret the J-PEP clusters for 38 focal disease/trait clusters identified by J-PEP using bulk epigenomic data. We analyzed a single-cell atlas of chromatin accessibility comprising 615,998 cells across 111 fine-grained adult human cell types. 44 To derive cell-type-to-cluster profile matrices ( H cell-type ), we used a two-step procedure (J-PEP-CT): (1) apply J-PEP to bulk data in order to infer the shared SNP-to-cluster membership matrix W ; and (2) solve V cell-type = W × H cell-type using a non-negative least-squares approach with multiplicative update rules analogous to those used in bNMF, where V cell-type denotes the SNP-to-cell-type association matrix obtained using the EM algorithm on scATAC-seq data across the 111 cell types (analogous to V tissue as defined above; Table S8 and STAR Methods ). We compared H cell-type to H tissue by collapsing both to 24 unified tissue categories ( Table S9 ); we observed high concordance (mean r = 0.55; Figure S15 ), confirming high concordance between the two modalities. We did not include direct application of J-PEP to single-cell data (i.e., applying J-PEP to V trait and V cell-type instead of the two-step J-PEP-CT procedure) in our primary analyses; due to increased sparsity, cell-type complexity, and technical variability of single-cell assays, this led to lower prediction accuracy than application of J-PEP to bulk data (although J-PEP still substantially outperformed pleiotropic partitioning and epigenomic partitioning in direct application to single-cell data; Figure S13 ), such that we do not recommend direct application of J-PEP to single-cell data. We evaluated whether scATAC-seq data could improve the resolution and specificity of cell-type associations. Restricting to the 11 of 24 unified tissue categories with at least four annotated cell types, we determined that focal trait-cluster pairs positively associated with these tissue categories were associated with only 21% of the corresponding cell types on average ( Figures 4 and S15 ; Table S8 ), indicating a high degree of specificity. For example, in the immune category, cluster 1 of lymphocyte count was predominantly associated with T cell subtypes, whereas cluster 4 of platelet count was exclusively linked to plasma B cells. In the brain and nervous system category, cluster 2 of neuroticism was specifically associated with a glutamatergic neuron subtype—excitatory cells critical for brain signaling— whereas cluster 1 of each white blood cell trait (NC, monocyte count, and basophil count) showed distinct associations with microglia, the brain’s resident immune cells. In the skin category, which comprises four annotated cell types, cluster 1 of hair pigmentation was uniquely associated with melanocytes, the pigment-producing cells of the epidermis ( Figure S15 ). These findings demonstrate that single-cell epigenomic data can resolve the cellular basis of trait-tissue associations, enhancing the interpretability of J-PEP-derived clusters. Figure 4. Open in a new tab Single-cell data improve resolution and specificity of cell-type associations We report normalized cell-type scores for each of 20 selected focal trait-cluster pairs (selected as described below; rows) across 24 single-cell-derived immune cell types (left) and brain/nervous system cell types (right). Normalized scores are defined as the proportion of each cluster assigned to a given cell type, normalized across all cell types for that cluster. We selected focal trait-cluster pairs with at least one normalized score >1% and sorted them by identifying the tissue category (immune or brain) with the highest aggregate score for each pair, then sorting first by tissue category and second by the magnitude of the maximum aggregate score within that category. Black borders denote focal trait/cluster/cell type triplets discussed in the main text, with corresponding focal trait-cluster pairs in bold font. Numerical results for 38 diseases/traits are reported in Table S6 . RDWCV, red cell distribution width—coefficient of variation. J-PEP identifies non-canonical axes of T2D pathogenesis T2D is a multifactorial disorder involving heterogeneous pathological mechanisms including pancreatic β cell dysfunction, impaired insulin production by excess hepatic fat accumulation, and obesity-related insulin resistance. 23 , 24 , 25 , 53 , 54 Prior efforts to partition T2D loci—often via pleiotropic partitioning—have consistently identified clusters reflecting these distinct processes. 23 , 24 , 25 , 27 , 30 , 31 , 55 However, these approaches have often lacked resolution in the specific cellular contexts in which they occur, a limitation that J-PEP is designed to address. To enable direct comparison with previous work, all results in this section are based on the 109 auxiliary traits from the most comprehensive pleiotropic partitioning study to date 25 ( Table S10 ), which employed a method closely aligned with pleiotropic partitioning as defined in the current study. We applied J-PEP and J-PEP-CT to T2D summary statistics from Suzuki et al. 24 (428,452 cases and 2,107,149 controls). J-PEP identified seven clusters, capturing both established and underexplored biological processes. These clusters are reported in Figure 5 , sorted in decreasing order of their joint proportion of variance explained across both the trait and tissue matrices, computed as the squared magnitude of each cluster’s reconstruction relative to the total squared magnitude of the observed data (see STAR Methods ). With respect to the PEPA metric, J-PEP (PEPA = 0.03) did not significantly differ from epigenomic partitioning (PEPA = 0.04, genomic jackknife on the difference p = 0.28; five clusters) and was outperformed by pleiotropic partitioning (PEPA = 0.06, p = 5.1 × 10 −4 ; five clusters) ( Table S3 ), consistent with the trade-off between the biological information in additional clusters vs. optimizing PEPA ( Figure S13 ). We detail J-PEP’s results below, as its clusters integrate elements from both pleiotropic and epigenomic partitions, offering a broad perspective on disease processes. Figure 5. Open in a new tab Comparison of T2D clusters identified by J-PEP and J-PEP-CT with other clustering methods For each T2D cluster identified by J-PEP and J-PEP-CT, the pie chart (left) shows the pleiotropic auxiliary trait profile (top four traits), and the bulk and single-cell columns show the associated epigenomic tissue and cell-type profiles, respectively. The “+” or “−” symbol appended to auxiliary trait names denotes pleiotropic associations that are concordant or discordant, respectively, with the focal trait ( STAR Methods ). Single-cell refinement results show the top five contributing cell types within the most enriched bulk tissue category, shaded by relative cell proportion. NA indicates that the single-cell atlas that we employed lacks cell types from the implicated bulk tissue. Cell types that do not correspond to a refinement of the implicated bulk tissue are denoted in parentheses. Cluster labels indicate the projection correlation ( r proj ), quantifying internal concordance between trait and tissue profiles and the proportion of total variance ( ptv ) they explain jointly across traits and tissues ( STAR Methods ). Clusters are displayed in decreasing order of ptv . The heatmap shows the maximum correlation between each J-PEP cluster and clusters obtained from pleiotropic partitioning, epigenomic partitioning, or two previous T2D clustering studies (Smith et al. 25 and Suzuki et al. 24 ). See data and code availability for numerical results. MCV, mean corpuscular volume; MRV, mean reticulocyte volume; ProInsulinadjBMI, proinsulin adjusted for body mass index; HOMAB, homeostasis model assessment of β cell function; WCadjBMI, waist circumference adjusted for body mass index; HLRC, high light reticulocyte count. For each J-PEP cluster, we computed SNP-level correlations between each cluster produced by a different method (defined as the correlation between the corresponding columns of the respective SNP-to-cluster membership matrices W ) and assessed the maximum correlation. For clusters 2, 4, and 5, SNP-level correlations were strongest between J-PEP and pleiotropy-based methods (including Smith et al. 25 and Suzuki et al. 24 ), yet overall correlations remained modest (mean max r = 0.34; Figure 5 ); for clusters 1, 3, 6, and 7, correlations were strongest between J-PEP and epigenomic partitioning, with high correlations (mean max r = 0.96; Figure 5 ), underscoring the informativeness of tissue data. Clusters 2–4 align with canonical T2D pathways 23 , 24 , 25 , 27 , 30 , 31 , 55 ( Figure 5 ). Cluster 2 is associated with insulin secretion traits and endocrine/pancreatic tissues and is enriched for GO terms related to insulin signaling and glucose regulation ( Table S11 and STAR Methods ). Notably, single-cell data refined this association to pancreatic islet β cells, supporting a β cell dysfunction mechanism contributing to insulin deficiency. 23 , 24 , 56 , 57 Cluster 3 exhibits a consistent liver signature across bulk and single-cell profiles with the single-cell association refined to hepatocytes—the main liver cell type and the only one represented in the single-cell data—and GO enrichment for cholesterol efflux and triglyceride metabolism, indicative of hepatic lipid dysregulation contributing to insulin resistance. 23 , 24 , 25 , 27 Cluster 4 is associated with liver enzyme and adiposity traits and maps to digestive tissues. Single-cell data refined this digestive association to multiple gastrointestinal cell types, and GO term enrichment for hydrogen peroxide catabolism points to hepatic oxidative stress linked to lipid overload and steatosis—a known mechanism of insulin resistance. 7 Clusters 1 and 5–7 reflect non-canonical components of T2D biology ( Figure 5 ). Cluster 1 is associated with red blood cell traits and epithelial tissues. Single-cell data refined this association to thyroid follicular cells, and GO enrichment revealed epithelial-to-mesenchymal transition and oxidative stress, suggesting a stress-related process potentially affecting insulin signaling. 58 Cluster 5 lacks a clear pathological interpretation but is enriched for chromatin organization and associated with embryonic stem cells. Refinement using single-cell data was not feasible, as the single-cell atlas that we employed lacks embryonic cell types. This cluster may reflect early developmental or transcriptional regulatory influences on T2D risk. Cluster 6 is associated with obesity-related traits and immune tissues. Single-cell data refined this association to macrophages, with GO enrichment for inflammatory signaling and thermogenesis, suggesting a role for immune-metabolic regulation of energy balance in T2D. 59 , 60 , 61 Finally, cluster 7 is characterized by associations with liver enzymes and triglycerides, GO enrichment for lipid homeostasis pathways, and a stromal bulk tissue profile refined by single-cell data to a mix of adipocyte and neuronal cell types, suggesting a metabolic mechanism. J-PEP identifies stromal and endocrine axes of HTN pathogenesis HTN is a highly polygenic disease with contributions from multiple biological pathways. 62 , 63 Prior clustering efforts have often focused on shared metabolic comorbidities such as obesity, lipid levels, or T2D, providing limited insight into tissue-specific regulatory contributions to blood pressure control. 64 , 65 We applied J-PEP and J-PEP-CT to HTN summary statistics from Bycroft et al. 43 (3,174 cases and 455,380 controls) using 68 auxiliary traits ( Table S10 ). J-PEP identified two clusters described in detail in Figure 6 A, sorted in decreasing order of their joint proportion of variance explained across both the trait and tissue matrices (see above). With respect to the PEPA metric, J-PEP (PEPA = 0.07) outperformed both pleiotropic partitioning (PEPA = 0.04, genomic jackknife on the difference p = 0.07; three clusters) and epigenomic partitioning (PEPA = 0.03, p = 2.1 × 10 −3 ; five clusters) ( Figure 3 A and Table S3 ). Figure 6. Open in a new tab Comparison of HTN and NC clusters identified by J-PEP and J-PEP-CT with other clustering methods For each hypertension (HTN) cluster (A) and neutrophil count (NC) cluster (B) identified by J-PEP and J-PEP-CT, the pie chart (left) shows the pleiotropic auxiliary trait profile (top four traits), and the bulk and single-cell columns show the associated epigenomic tissue and cell-type profiles, respectively. Single-cell refinement results show the top five contributing cell types within the most enriched bulk tissue category, shaded by relative cell proportion. The “+” or “−” symbol appended to auxiliary trait names denotes pleiotropic associations that are concordant or discordant, respectively, with the focal trait ( STAR Methods ). Cluster labels indicate the projection correlation ( r proj ), quantifying internal consistency between tissue and trait profiles and the proportion of total variance ( ptv ) they explain jointly across traits and tissues ( STAR Methods ). Clusters are displayed in decreasing order of ptv . The heatmap shows the maximum correlation between each J-PEP cluster and clusters obtained from pleiotropic partitioning or epigenomic partitioning (or Vaura et al. 64 for HTN). See data and code availability for numerical results. HLR, high light reticulocyte count; PDW, platelet distribution width; BP, blood pressure; MCV, mean corpuscular volume; PEF, peak expiratory flow. Cluster 1 is associated with bulk stromal tissue, refined by single-cell data to mesothelial cells, with GO enrichment for Wnt signaling and cytoskeletal remodeling ( Table S11 ), suggesting mesothelial-driven extracellular matrix (ECM) remodeling and vascular stiffening, a known contributor to elevated vascular resistance in HTN. 66 While ECM remodeling is a recognized mechanism in HTN pathophysiology, 67 , 68 the specific role of mesothelial cells in this process remains speculative. 69 Cluster 2 is associated with endocrine bulk tissue, refined by single-cell data to adrenal cortex cells, with GO terms highlighting neutrophil-mediated immunity, indicating a potential feedback between aldosterone signaling and local immune activity. 70 Although aldosterone is known to modulate various immune pathways involved in inflammation, and immune-targeted therapies have shown benefit in managing HTN, the precise mechanisms underlying its crosstalk with the immune system remain unknown. 71 For both cluster 1 and cluster 2, SNP-level correlations were strongest between J-PEP and epigenomic partitioning ( Figure 6 A), highlighting the benefit of incorporating epigenomic data. Compared to the four metabolically themed clusters identified by Vaura et al., 64 who applied a method similar to pleiotropic partitioning to FinnGen data 72 —comprising obesity, two lipid-related, and one stature-associated clusters—J-PEP produced a distinct set of clusters with minimal overlap, likely reflecting its integrative design that incorporates epigenomic information. J-PEP identifies immune, hepatic, and neuroinflammatory axes of NC NC is a key hematological trait central to immune and inflammatory processes, 58 , 59 yet to our knowledge it has not previously been analyzed using clustering methods that integrate either pleiotropic or epigenomic data. We applied J-PEP and J-PEP-CT to NC summary statistics from Vuckovic et al. 73 (563,085 samples) using 83 auxiliary traits ( Table S10 ). J-PEP identified three clusters reflecting immune, hepatic, and neuroinflammatory axes ( Figure 6 B, sorted in decreasing order of their joint proportion of variance explained across both the trait and tissue matrices; see above). With respect to the PEPA metric, J-PEP (PEPA = 0.14) significantly outperformed pleiotropic partitioning (PEPA = 0.05, genomic jackknife on the difference p = 3.8 × 10 −9 ; five clusters) and epigenomic partitioning (PEPA = 0.06, p = 6.4 × 10 −6 ; three clusters) ( Figure 3 A and Table S3 ). Cluster 1 is associated with bulk hematopoietic stem and progenitor cells, refined by single-cell data to alveolar and general macrophages, with GO enrichment for chemotaxis, immune activation, and antimicrobial responses ( Table S11 ), highlighting canonical immune regulation of neutrophils. 74 Cluster 2 exhibits a consistent liver signature across bulk and single-cell profiles (refined to hepatocytes in single-cell data) and is linked to platelet-related traits, suggesting a speculative hepatic-inflammatory connection to neutrophil regulation, consistent with neutrophil activation in metabolic disorders. 75 Cluster 3 exhibits a bulk brain profile refined by single-cell data to glutamatergic neurons and oligodendrocytes, along with associations to hematocrit and platelet count, pointing to emerging but unconfirmed links between neural regulation and immune tone. 76 , 77 SNP-level correlations indicate that cluster 1 aligns closely with epigenomic partitioning, whereas clusters 2 and 3 are only weakly recovered by pleiotropic partitioning—despite their most distinctive features (compared to cluster 1 and to each other) being reflected in their tissue and cell-type profiles, which notably do not implicate immune-related components. Discussion Summary of findings We have developed J-PEP, a method that integrates pleiotropic and epigenomic information to partition disease loci into biologically interpretable clusters. By leveraging shared structure across auxiliary traits and tissues, J-PEP improves the accuracy of predicting SNP-to-trait and SNP-to-tissue associations, which we quantify using our new PEPA metric. These gains are most pronounced for focal diseases/traits whose underlying biology is well captured by available epigenomic datasets, such as immune- and liver-related traits. In our applications to GWAS summary statistics for a broad set of diseases and traits, J-PEP recapitulates canonical disease pathways, e.g., endocrine/pancreatic dysfunction and hepatic-driven insulin resistance for T2D, 23 , 24 , 25 but also reveals underexplored biological processes, e.g., immune-metabolic and developmental components for T2D. Notably, incorporating single-cell epigenomic profiles refines tissue associations to cell-type resolution, e.g., resolving an endocrine/pancreatic cluster for T2D to pancreatic β cells, and resolving an endocrine/renal cluster for HTN to specific adrenal cortex cell types. Alternative approaches J-PEP shares conceptual goals with several recent approaches aimed at disentangling shared genetic architecture across complex traits. DeGAs 17 applies truncated singular value decomposition (tSVD) to GWAS summary statistics to extract latent components of genetic variation. While scalable and effective for broad surveys, it lacks structured priors to promote sparsity, usually requiring post hoc enrichment analyses for interpretation. FactorGo 26 addresses this by modeling uncertainty in GWAS effect estimates within a variational Bayesian framework and promoting sparsity through automatic relevance determination. However, it operates on LD-pruned data, which limits compatibility with regulatory annotations that do not depend on LD structure, such as epigenomic marks. The V2G2P framework 37 takes a more mechanistic route, linking GWAS variants to genes through enhancer maps and to regulatory programs derived from single-cell perturbation screens. This strategy enables interpretable, cell-type-specific disease insights, but it requires substantial experimental resources and is currently limited to well-characterized cellular systems, limiting its scalability. Downstream implications Our results have several downstream implications. From a translational perspective, J-PEP provides a scalable means to distill complex polygenic signals into a small number of interpretable axes of biological activity. This may aid in prioritizing genes, tissues, and pathways for functional follow-up and therapeutic development. 4 , 78 , 79 Also, by uncovering shared axes across diseases/traits, J-PEP may help identify opportunities for drug repurposing—for example, targeting pathways implicated in both T2D and cardiovascular disease. 80 , 81 In the context of precision medicine, clusters inferred by J-PEP could aid the construction of partitioned polygenic risk scores, previously shown to support disease subtyping. 24 , 25 Finally, our PEPA metric provides a principled framework for benchmarking future methods that integrate pleiotropic and epigenomic data. Limitations of the study We note six limitations of J-PEP and this study. First, although J-PEP outperforms single-modality methods, overall PEPA scores remain modest, reflecting inherent biological complexity, measurement noise, and incomplete regulatory and phenotypic coverage—gaps that are likely to narrow as new genetic and epigenomic resources continue to emerge. 4 , 82 , 83 , 84 , 85 Second, while J-PEP captures trait-tissue correlations, it does not explicitly infer causality between them; nonetheless, clusters derived from joint partitioning are less prone to spurious associations compared to single-modality approaches, as they better reflect disease-relevant biological context. Third, in bulk epigenomic data, the results of J-PEP (and other methods) are biased toward tissues with denser annotations; in single-cell epigenomic data, direct application of J-PEP (instead of J-PEP-CT) is limited by sparsity and technical variability, and we do not recommend this approach. Ongoing expansion of tissue-specific and single-cell epigenomic datasets will help mitigate both issues. 4 , 82 , 83 , 84 , 85 Fourth, the algorithm can be computationally demanding for large-scale datasets, particularly when incorporating high-dimensional modalities like single-cell multi-omics; however, the method is computationally tractable ( Figure S14 ), and efficiency can be further enhanced by restricting inputs to critical tissues or a curated set of auxiliary traits. Fifth, J-PEP employs a non-standard procedure to enforce sparsity in tissue-to-cluster scores at each iteration. However, our comparative analyses demonstrate that this sparsity constraint is essential for ensuring convergence, computational efficiency, and biologically meaningful clusters, and it outperforms alternative strategies ( Figures S5 A and S5B), such that we recommend it as our default approach. Finally, we did not evaluate J-PEP under multiple-causal variant fine mapping with in-sample LD; because such methods can be miscalibrated with out-of-sample LD, 52 and our empirical analyses relied primarily on out-of-sample LD, we used single-causal variant fine mapping for both simulations and real data to ensure calibration and consistency. Despite these limitations, J-PEP provides a generalizable and interpretable approach for partitioning GWAS loci into biologically distinct clusters. Resource availability Lead contact Requests for further information and resources should be directed to and will be fulfilled by the lead contact, Gaspard Kerner ( [email protected] ). Materials availability This study did not generate new unique reagents. Data and code availability All data used in this study and the corresponding access information are listed in the key resources table . This includes the jpepR package, developed as part of this work, which provides an environment to run J-PEP, as well as GWAS summary statistics for 165 traits and the corresponding J-PEP results. Links to the public software and algorithms used in the present study are listed in the key resources table . Acknowledgments We thank Kushal Dey and all members of Alkes Price’s lab for helpful discussions and data sharing. This research was funded by NIH grants R01 HG006399, U01 HG012009, R01 MH101244, R01 HG013083, R37 MH107649, and T32 HG002295. Author contributions G.K. and A.L.P. conceived and designed the study. G.K. was the lead analyst, with important contributions from A.L.P., B.S., and N.K. N.K. and B.S. conceived and designed the SNP-to-tissue scoring algorithm using the EM algorithm. G.K. wrote the manuscript with substantial contributions from A.L.P. Declaration of interests The authors declare no competing interests. STAR★Methods Key resources table REAGENT or RESOURCE SOURCE IDENTIFIER Deposited data EpiMap bulk epigenomic tracks Boix et al. 35 https://compbio.mit.edu/epimap/ Single-cell ATAC-seq Zhang et al. 44 GEO: GSE184462 GWAS summary statistics for 165 diseases/traits This paper https://alkesgroup.broadinstitute.org/J-PEP Auxiliary traits (109 traits from ref. 25) This paper https://alkesgroup.broadinstitute.org/J-PEP cS2G (SNP-to-gene linking annotations) Gazal et al. 46 J-PEP input/output matrices and PEPA/PPA/EPA This paper https://alkesgroup.broadinstitute.org/J-PEP Software and algorithms jpepR (J-PEP, J-PEP-CT, PEPA and bNMF code) This paper https://github.com/kernerg/jpepR ( https://doi.org/10.5281/zenodo.17704042 ) Flashier Willwerscheid et al. 42 https://doi.org/10.48550/arXiv.2110.00152 FactorGO Zhang et al. 26 https://github.com/mancusolab/FactorGo MACS2 (peak calling) Github https://macs3-project.github.io/MACS/ UCSC bigWigToBedGraph UCSC Genome Browser https://genome.ucsc.edu/goldenpath/help/bigWig.html R (statistical computing environment) R Core Team Version 4.1.2 Open in a new tab Method details J-PEP method: Input and output J-PEP inputs V trait and V tissue and outputs W , H trait and H tissue . V t r a i t ∈ R M × A is a SNP-to-auxiliary trait matrix defined for a given focal trait. M is the number of fine-mapped SNPs (PIP >0.01) for the focal trait and A is the number of auxiliary traits. The entries of this matrix are defined as: ( V t r a i t ) m , a = P I P m f × P I P m a (Equation 2) where P I P m f and P I P m a denote the posterior inclusion probabilities (PIPs) for SNP m in the focal trait f and auxiliary trait a , respectively. Because J-PEP relies on non-negative matrix factorization, each auxiliary trait is split into two components: a concordant pleiotropic component (denoted by a “+” sign in figures) and a discordant pleiotropic component (denoted by a “−” sign). V t i s s u e ∈ R M × T is a SNP-to-tissue matrix defined for a given focal trait. M is the number of fine-mapped SNPs (PIP >0.01) for the focal trait and T is the number of tissues. The entries of this matrix are defined as: ( V t i s s u e ) m , t = a ˆ m , t (Equation 3) where a ˆ m , t is a normalized score quantifying the strength of association between SNP m and tissue t derived from tissue-specific epigenomic profiles. We generate this using the E-M algorithm. Details are below (see EpiMap dataset ). W ∈ R M × K is a SNP-to-cluster membership matrix defined for a given focal trait. M is the number of fine-mapped SNPs (PIP >0.01) for the focal trait and K is the number of clusters identified by the J-PEP method. H t r a i t ∈ R K × A is a cluster-specific profile matrix for auxiliary traits. K is the number of clusters identified by the J-PEP method and A is the number of auxiliary traits, as in V trait . H t i s s u e ∈ R K × T is a cluster-specific profile matrix for tissues. K is the number of clusters identified by the J-PEP method and T is the number of tissues, as in V tissue . We note that neither the columns of W nor the rows of H matrices are constrained to sum to 1, as is the case in other factorization approaches, 26 , 42 since the partitioning approach preserves relative contributions of clusters to the overall partition. Reconstructed V matrices are not used for benchmarking but only inferred W and H matrices, since their reconstruction would only reproduce observed data. J-PEP method: Objective function J-PEP performs joint non-negative matrix factorization on V trait and V tissue , enforcing a shared SNP-to-cluster membership matrix W . The objective function is: { W , H t r a i t , H t i s s u e } = argmin W ≥ 0 , H t r a i t ≥ 0 , H t i s s u e ≥ 0 { 1 2 ‖ V t r a i t − W H t r a i t ‖ 2 + 1 2 ‖ V t i s s u e − W H t i s s u e ‖ 2 + ∑ k λ ( ( V t r a i t H t r a i t T ) k , ( V t i s s u e H t i s s u e T ) k ) } (Equation 4) where k is taken across the K clusters (columns of V t r a i t H t r a i t T and V t i s s u e H t i s s u e T ) from the current iteration of the algorithm and the penalty function λ promotes alignment between pleiotropic and epigenomic data: λ ( A , B ) = − log 10 ( p c o r r ( A , B ) ) − 1 (Equation 5) where p corr indicates the one-sided Pearson’s correlation p value for positive correlation. We note that since V t r a i t H t r a i t ′ and V t i s s u e H t i s s u e ′ approximate the shared cluster membership matrix W (and are equal to it when H trait and H tissue are orthogonal matrices, which is rare in practice, and otherwise approximate W by projecting the observed data into the cluster membership space spanned by H , as is standard in matrix factorization approaches 86 , 87 ), the penalty function λ encourages clusters where SNP memberships are consistent across modalities (i.e., trait and tissue data). We refer to λ as the penalty function and the entries of the vector λ as the penalty weights. We also note that λ depends on the number of fine-mapped SNPs, so the method uses a scaling parameter to adjust the strength of the penalty and can be tuned by the user. J-PEP method: Optimization We optimized the J-PEP objective function using an iterative algorithm alternating updates of the shared SNP-to-cluster membership matrix W , the trait-specific cluster profile matrix H trait , and the tissue-specific profile matrix H tissue . All matrices were initialized with small positive random values scaled to the input range. Updates were performed using multiplicative update rules derived from the gradient of the objective function 41 (also referred to as gradient descent updates), as done in bNMF. 40 Specifically, these updates can be derived from the gradient of the objective function with respect to the H and W matrices, and the multiplicative update rules correspond to a projected gradient descent step under non-negativity constraints. Updates also incorporate the penalization function λ , whose length is equal to the inferred number of clusters K . Entries of the vector λ (penalty weights) enforce consistency between trait-based and tissue-based cluster assignments by penalizing clusters whose SNP memberships are poorly correlated across the two modalities. To maintain numerical stability, a small constant ε = 10 −50 was added to all matrix approximations. Convergence was assessed by monitoring relative changes in the penalization weights, i.e., the entries of the vector λ, applied to each cluster. The algorithm terminated when the maximum relative change of the penalty weights across iterations fell below a predefined threshold (default: 10 −7 ) or when the maximum number of iterations (default: 5,000) was reached (very rare in practice, see Figure S7 C). The iteration cap served as a fail-safe rather than a convergence condition; if the cap is hit without meeting the tolerance, the routine stops with an explicit “maximum iterations without convergence” error and the run is not used. We discarded solutions in which fewer than two non-redundant clusters remained active (defined by row sums of H trait and H tissue exceeding a minimal threshold, defined by default to 10 −10 ). Additionally, we discarded solutions in which at least one method failed to identify at least two clusters in a minimum of 18 out of the 22 model runs required for computing prediction accuracy metrics (one per chromosome removed; see below). Applying these two criteria retained 38 of the original 55 focal traits from the full set of 165 defined as approximately genetically independent (pairwise r g < 0.5). Sparsity constraint To improve stability and interpretability, J-PEP imposes tissue-to-cluster memberships (columns of H tissue ) to be either zero or positive for at most one cluster, as previously proposed, 88 whereas auxiliary traits remain free to associate with multiple clusters. Importantly, a tissue remains unassigned if its highest tissue-to-cluster value is zero (or for practical purposes, very close to 0). Conceptually, J-PEP optimizes the objective function of Equation 1 under the constraint that H tissue ∈ C , where C is the set of matrices whose columns each have at most one non-zero entry. For robustness, we perform ten optimization runs, as previously proposed, 23 with different random seeds and retain the solution that yielded the highest average projection correlation between modalities, i.e., the highest average correlation between V t r a i t H t r a i t ′ and V t i s s u e H t i s s u e ′ across clusters, among those yielding the most consistent number of clusters; in case of ties, we prioritize the run with fewer clusters. Cluster ordering In Figures 5 and 6 , we ordered clusters based on the proportion of total variance they jointly explained across both trait and tissue association matrices. Specifically, for each cluster k , we computed its contribution to the SNP-to-trait matrix V trait and the SNP-to-tissue matrix V tissue using rank-1 reconstructions: V trait , k = W . k H trait [ k , .] and V tissue , k = W . k H tissue [ k ,]. The total variance explained by cluster k was defined as the sum of squared entries in V trait , k and V tissue , k , normalized by the total variance of the original matrices. Clusters were then ranked in decreasing order of this joint variance contribution. Pleiotropic partitioning method The pleiotropic partitioning method factorizes V trait using bNMF, that is, by solving the optimization problem: { W , H t r a i t } = argmin W ≥ 0 , H t r a i t ≥ 0 { ‖ V t r a i t − W H t r a i t ‖ 2 + ∑ k λ s i n g l e ( W k , H t r a i t , k ) } (Equation 6) where W and H trait are defined as above, and λ single is a penalty term that reduces overfitting by pruning irrelevant clusters. 40 The maximum number of clusters K is predefined (typically K = 15), but the actual number of identified clusters by the partitioning method is often reduced as irrelevant components are penalized to zero. M includes all fine-mapped SNPs with PIP values ≥ 0.01 for the focal trait. The code to run this step was directly taken from ref. 23 without modification. On the other hand, the following steps to obtain H tissue (which is not produced by pleiotropic partitioning methods but is required here to enable comparison to other methods using our new PEPA metric) are original to this work. We generate cluster-specific tissue profiles ( H tissue ) while ensuring that each tissue associates with a single cluster, as follows. 1) We estimate H tissue from V tissue and W using gradient descent by solving V tissue = W × H tissue , without applying any constraint on H tissue . 2) Clusters with highly similar tissue profiles are merged by averaging their respective cluster memberships. Similarity is assessed using the simil function from the R v.4.2 proxy package, with “high similarity” defined as a similarity score > 0.5. 3) We repeat step (1) but now enforce the sparsity condition, as implemented in the J-PEP method. Epigenomic partitioning method The epigenomic partitioning method factorizes V tissue using bNMF, that is, by solving the analogous optimization problem to the pleiotropic partitioning method, using as input V tissue instead of V trait : { W , H t i s s u e } = argmin W ≥ 0 , H t i s s u e ≥ 0 { ‖ V t i s s u e − W H t i s s u e ‖ 2 + ∑ k λ s i n g l e ( W k , H t i s s u e , k ) } (Equation 7) The code to run this step was directly taken from ref. 23 without modification. We learn cluster-specific auxiliary trait profiles ( H trait ) from V trait and W using gradient descent by solving V trait = W × H trait . Prediction accuracy metrics PPA. Let V t r a i t ∈ R M × A and V t i s s u e ∈ R M × T represent the SNP-to-trait and SNP-to-tissue matrices, respectively, where M is the number of fine-mapped SNPs (PIP >0.01) for the focal trait, A is the number of auxiliary traits, and T is the number of tissues. The prediction process aims to estimate V trait using the SNP-to-cluster membership matrix W and the tissue cluster profiles H tissue and H trait . Essentially, PPA follows a two-step approach. 1) For a given cluster k and a held-out chromosome c , we predict the cluster membership W k ( c ) based on V t i s s u e ( c ) and H t i s s u e k − ( c ) , the SNP-to-tissue associations for chromosome c and the tissue profiles inferred from the remaining 21 chromosomes, using gradient descent. 2) Then, the predicted SNP-to-trait associations for cluster k and chromosome c , i.e., V ˆ t r a i t k ( c ) , are calculated as: V ˆ t r a i t k ( c ) = W k ( c ) H t r a i t k − ( c ) (Equation 8) The Pearson correlation between the predicted values ( V ˆ t r a i t k ( c ) ) and the true values ( V t r a i t ( c ) ) for each cluster is computed as: r k ( c ) = c o r ( V ˆ t r a i t k ( c ) , V t r a i t ( c ) ) (Equation 9) The weighted mean correlation across all clusters in chromosome c is computed as: r t r a i t ( c ) = ∑ k r k ( c ) w k ( c ) ∑ k w k ( c ) (Equation 10) where w k ( c ) is defined as: w k ( c ) = V a r ( W k ( c ) ) V a r ( ∑ j W j ( c ) ) (Equation 11) EPA Computation for EPA is analogous to PPA, inverting the roles of trait and tissue data (e.g., V tissue for V trait and V trait for V tissue ). The Pearson correlation between the predicted values ( V ˆ t i s s u e k ( c ) ) and the true values ( V t i s s u e ( c ) ) for each cluster is computed as: r k ( c ) = c o r ( V ˆ t i s s u e k ( c ) , V t i s s u e ( c ) ) (Equation 12) The weighted mean correlation across all clusters in chromosome c is computed as: r t i s s u e ( c ) = ∑ k r k ( c ) w k ( c ) ∑ k w k ( c ) (Equation 13) PEPA The overall PEPA score is the sum of the weighted average correlations for both prediction tasks across all chromosomes: P E P A = ∑ c ϵ C ( ( r t i s s u e ( c ) × r t r a i t ( c ) ) × p p i p ( c ) ) (Equation 14) where ppip ( c ) denotes the proportion of fine-mapping PIPs on chromosome c relative to the total across all held-out chromosomes in the set C, where C indicates the group of chromosomes containing at least five GWAS loci. PEPA was only computed in cases in which |C| > 18. Although inference approaches can differ in the number of discovered clusters, the PEPA metric ensures comparability across methods because it evaluates prediction accuracy in the observed data space, independent of the number of inferred clusters. Projection correlation To validate trait–tissue coherence across identified clusters, we computed a projection correlation metric beyond the PEPA metric, defined as r p r o j = c o r ( V t r a i t × H t r a i t t , V t i s s u e × H t i s s u e t ) . Methodological considerations and limitations of J-PEP We highlight three methodological considerations. First, J-PEP assumes a linear shared structure between modalities, which may overlook non-linear interactions between pleiotropy and epigenomic data; however, the linearity assumption is likely to be a reasonable approximation given the largely additive nature of genetic effects. 22 , 89 , 90 Second, the requirement for concordant clustering across traits and tissues may suppress modality-specific associations (e.g., clusters detectable only through pleiotropy or epigenomics alone); this constraint also underlies the definition of the PEPA metric. However, this restriction likely improves robustness by focusing on clusters supported by multiple lines of evidence. Similarly, the geometric mean used in the PEPA metric penalizes imbalance between modalities (traits and tissues), thus failing to prioritize solutions where one modality performs well while the other performs poorly. Nevertheless, we chose the geometric mean specifically because J-PEP aims to identify clusters with concordant support across modalities. Third, the different statistical properties of V trait and V tissue arising from their different scaling with respect to focal trait PIPs is a primary factor impacting the relative concordance of J-PEP with pleiotropic partitioning vs. epigenomic partitioning; however, the alternative that rescales V tissue by focal trait PIPs performed worse than the default approach ( Figure S5 C; Table S3 ). Simulation framework Generating simulated data consists of 8 steps: Step 1: simulating cluster-specific profiles, Step 2: estimating the SNP-to-cluster membership matrix, Step 3: computing SNP-to-trait associations, Step 4: sampling causal SNPs, Step 5: sampling causal effect sizes, Step 6: sampling GWAS summary statistics, Step 7: fine-mapping, Step 8: generating uncorrelated structure. Step 1. Simulating cluster-specific profiles We initialized cluster-specific trait and tissue profiles as matrices H t r a i t ∈ R A × K and H t i s s u e ∈ R T × K , where A is the number of auxiliary traits, T the number of tissues, and K the number of causal clusters. Each matrix was initialized with standard uniform values, and a subset of traits and tissues was randomly assigned as causal for each cluster (one trait and one tissue per cluster) by retaining non-zero values only in the corresponding columns: H t r a i t [ i , j ] = { U [ 0.1 , 1 ] , i f i ∈ c a u s a l t r a i t f o r c l u s t e r j 0 , o t h e r w i s e ) (Equation 15) H t i s s u e [ i , j ] = { U [ 0.1 , 1 ] , i f i ∈ c a u s a l t i s s u e f o r c l u s t e r j 0 , o t h e r w i s e ) (Equation 16) Each row was normalized to sum to one. We used K = 5 by default and tested scenarios with K = 3 and K = 8. Step 2. Estimating SNP-to-cluster membership matrix The SNP-to-cluster membership W , which probabilistically maps SNPs to clusters, was inferred using gradient descent, a SNP-to-tissue matrix from real data analysis and simulated H tissue , by solving the factorization problem: { W ˆ } = argmin W ≥ 0 { 1 2 ‖ V ˜ t i s s u e − W H t i s s u e ‖ 2 } where V ˜ t i s s u e is a SNP-to-tissue matrix constructed by intersecting SNPs from chromosomes 1 and 2 of 1000 Genomes Phase 3 43 with EpiMap 35 annotations (see Epimap dataset ). Only SNPs with MAF >0.05 were retained. Step 3. Computing SNP-to-trait associations The SNP-to-trait matrix V trait was obtained by matrix multiplication: W × H trait . Step 4. Sampling causal SNPs Causal SNPs for the focal and auxiliary traits were sampled from distinct causal loci, ensuring independence between all causal SNPs. These were sampled with probability proportional to: P ( c a u s a l S N P i f o r t r a i t j ) ∝ V t r a i t [ i , j ] × M A F i (Equation 17) where MAF i is the MAF of SNP i . We assumed the first column of V trait represents the focal trait. This sampling scheme has the consequence that common variants explain more trait variance (analogous to real traits, due to negative selection 91 , 92 , 93 ) and also the consequence that common variant architectures are more polygenic (analogous to real traits, due to negative selection 94 ). For auxiliary traits contributing to clusters, we enforced at least 25% overlap in causal SNPs with the focal trait. We varied the number of total causal SNPs m ∈ {100; 200; 300; 400}, sampling each from different LD blocks on chromosomes 1 and 2. Step 5. Sampling causal effect sizes Effect sizes were sampled from an independent standard normal distribution. Each trait’s total SNP-heritability ( h j 2 = 0.5) was distributed equally across its causal SNPs by rescaling effect sizes such that their squared sum matched h j 2 . Step 6. Simulating GWAS summary statistics We then simulated summary statistics using the RSS likelihood, 95 β ˆ ∼ N ( S R S − 1 β , S R S ) , where β is the vector of true causal effect sizes, R is the LD matrix from chromosomes 1 and 2 of the 1000 Genomes Project Phase 3, S is a diagonal matrix with elements S i i = 1 N 2 p i ( 1 − p i ) , N is the sample size and p i the minor allele frequency of SNP i . We used a regularized LD matrix, R ×(1– λ )+ λI , where λ = 0.001. Step 7. Fine-mapping We performed single causal variant fine-mapping 38 , 39 on all simulated GWAS summary statistics as in ref. 20 . We note that multiple causal variant fine-mapping is not recommended when in-sample LD is not available. 52 Step 8. Generating uncorrelated structure We define uncorrelated structure as a reproducible, modality-specific component of variation, i.e., a pattern present in one modality that has no counterpart in the other. To simulate trait- or tissue-specific variation unrelated to the predefined cluster structure, we added uncorrelated, yet “important”, features (i.e., auxiliary traits for V trait and tissues for V tissue ) to both V trait and V tissue . We first ranked auxiliary traits and tissues by their total contribution to the simulated clusters, then randomly selected a subset of low-contributing (“non-important”) features, defined as features that were not associated to the predefined clusters and thus had little or no contribution to the simulated clusters, and an equally sized subset of high-contributing (“important”) features. For each selected non-important feature, we replaced its profile (i.e., the SNP-to-feature scores) by resampling values from one of the sampled top-contributing feature of the same modality—maintaining realistic signal strength while disrupting alignment with the cluster structure. The number of such substituted features (traits and tissues) defines the level of uncorrelated structure simulated ( s ). This is an approach original to this paper. Representative simulation example We selected a representative simulation scenario with K causal = 5 and m = 200 for visualization in Figure S1 . To improve clarity, only 4 of the 5 clusters are shown, excluding the cluster with the weakest trait–tissue correlation based on r proj . To align colors between true and inferred clusters across methods, we computed correlations between the true and inferred trait and tissue profiles, assigning each inferred cluster the color of its best-matching true cluster. Additional simulation analyses Downsampling To evaluate the impact of LD among non-causal SNPs on model performance, we performed a downsampling procedure using the same simulation parameters as in the main simulations presented in the main text (e.g., K causal = 5 and s = 2). In each replicate, all true causal SNPs were retained, while the set of non-causal SNPs was progressively reduced. Specifically, we randomly removed 0% (default), 20%, 40%, or 60% of the remaining non-causal SNPs. We then applied J-PEP, pleiotropic partitioning, and epigenomic partitioning to the downsampled data and evaluated performance using the PEPA metric. We present results for scenarios with m = 200 or m = 400 causal SNPs. Varying PIP threshold To assess how the choice of posterior inclusion probability (PIP) threshold influences model performance, we conducted additional simulations under the same parameter settings as in the main simulations presented in the main text (e.g., K causal = 5 and s = 2). For each replicate, we varied the minimum PIP threshold used to retain SNPs in the input matrices. In addition to the default threshold of 0.01, we tested stricter thresholds (0.05, 0.1, 0.2, 0.5) as well as more lenient thresholds (0.005 and 0.001). For each threshold, we constructed V trait and V tissue accordingly, ran J-PEP, pleiotropic partitioning, and epigenomic partitioning, and evaluated prediction accuracy using the PEPA metric. We present results for scenarios with m = 200 or m = 400 causal SNPs. FactorGO and Flashier implementations To benchmark J-PEP against independently developed approaches, we conducted additional analyses with FactorGO 26 and three variants of Flashier 42 (Flashierz, Flashier, and Flashier_snn), applied both in simulations (under the default primary simulations procedure) and in real data. FactorGO FactorGO was run on a z-score–based V trait matrix, where entries are z-scores of auxiliary traits and the focal trait for fine-mapped SNPs of the focal trait. We used default settings, with a maximum of K = 15 initial factors, consistent with the maximum cluster number allowed for J-PEP, pleiotropic, and epigenomic partitioning. Flashierz Flashier applied to the same z-score–based V trait matrix with K = 15. Flashier Flashier directly applied to the original PIP-based V trait . Flashier_snn A semi-non-negative version of Flashier, enforcing non-negativity on one of the inferred factor matrices, applied to the PIP-based V trait . For this configuration we additionally applied non-negative gradient descent to obtain H tissue . This is the only Flashier variant used in real data analysis as it was the best performing in simulations. For all Flashier runs, we set backfit = TRUE. Since these methods only factorize the trait matrix, they do not provide H tissue . To enable comparison with J-PEP, we projected tissues onto the inferred SNP-to-cluster membership matrix W using gradient descent, applying a non-negativity constraint only for the Flashier_snn configuration (as also done in pleiotropic partitioning). Performance was evaluated using two criteria: (i) the PEPA metric (in both simulations and real data) and (ii) matrix reconstruction error, assessed via the Frobenius norm error of the reconstructed H trait and H tissue matrices, and only assessed in simulations. GWAS summary statistics We applied partitioning methods to GWAS summary statistics for 165 focal diseases/traits (average of 290K samples and 77 fine-mapped loci), including 57 UK Biobank diseases/traits 43 and 108 non-UK Biobank diseases/traits ( Table S1 ; see Data Availability). Fine mapping We conducted single causal variant fine-mapping on all tested GWAS traits as in ref. 20 Combining results across GWAS diseases/traits To evaluate significant differences between J-PEP’s PEPA average across GWAS diseases/traits with respect to other partitioning strategies, we used two approaches. First, we applied a genomic block jackknife to the unweighted average PEPA difference across traits (denoted p in the main text). Second, we performed a jackknife using trait-specific weights based on the inverse variance of each trait’s PEPA difference (denoted p w in the main text). EpiMap dataset Track calling We obtained EpiMap 35 p value signal tracks from the Washington University Epigenome for downstream analysis. We restricted our selection to tracks that met two criteria: (1) the tracks were directly observed rather than imputed, and (2) the tracks corresponded to one of six epigenomic annotations —DNase hypersensitive sites (DNase HS), H3K4me1, H3K4me2, H3K4me3, H3K9ac, or H3K27ac. The selected tracks were initially downloaded in bigWig format and converted to bedGraph format using the UCSC command-line utility bigWigToBedGraph. Following this conversion, we performed peak calling using MACS2 to identify significant regions of epigenomic activity. Peaks were called using the macs2 bdgpeakcall command with a significance threshold of –log 10 ( p ) > 2, ensuring that only high-confidence peaks were included in the subsequent analyses. SNP-to-tissue scoring using the EM algorithm To compute SNP-to-tissue scores from epigenomic annotations, we applied an expectation maximization (EM) algorithm to fine-mapped variants overlapping epigenomic peaks from the EpiMap project. This approach is original to this paper. Briefly, the EM algorithm is used to infer a latent, unobserved variable a pt that indicates if SNP p is active (i.e., overlaps a causal regulatory element) in annotation t . Let V ϵ {0, 1} M × t denote the binary overlap matrix of M SNPs and t epigenomic annotations. Then, formally, the algorithm maximizes the following likelihood: L ( θ ) = ∏ p = 1 M ∑ t = 1 T θ t V p t w p b t (Equation 18) where V pt indicates the observed overlap of SNP p and annotation t , w p is the PIP weight for SNP p , b t indicates the background expectation for annotation t , and θ t is the annotation-level enrichment weight. In practice, we initialize posterior weights a pt across annotations for each SNP p by normalizing overlap rows: a p t ( 0 ) = V p t ∑ t ′ V p t ′ (Equation 19) At each iteration z, we re-estimate annotation weights θ t as a weighted sum over SNPs (M-step): θ t ( z ) = 1 Z ∑ p = 1 M a p t ( z − 1 ) w p b t (Equation 20) where w p is the SNP-specific weight based on its PIP, b t is the background expectation for annotation t , and Z is a normalization constant ensuring ∑ t θ t = 1 . Posterior weights a pt are then updated as (E-step): a p t ( z ) = V p t θ t ( z ) ∑ t ′ V p t ′ θ t ′ ( z ) (Equation 21) Iterations proceed until convergence is achieved (mean squared difference between a p t ( z ) and a p t ( z − 1 ) below 10 −8 , or a maximum of 200 iterations). SNPs were retained if their posterior inclusion probability (PIP) exceeded a default PIP threshold (default: 0.01). Functional annotations were filtered to exclude cancer samples and tracks with extreme background enrichment values (computed on chromosome 1 of the 1000 Genomes Project 51 ). Track-level scores were aggregated into tissue-level scores by summing the normalized posterior probabilities from the EM algorithm across all tracks assigned to predefined tissue groups (as defined in EpiMap metadata 35 ), yielding a final SNP-to-tissue matrix V t i s s u e ϵ R M × T , where T is the number of grouped tissues. Single-cell dataset We analyzed a single-cell atlas of chromatin accessibility comprising 615,998 cells across 111 fine-grained adult human cell types. 44 Cell-type-specific peak annotations were downloaded using the following command: wget -r --level=1 -np -R index . html http://catlas.org/catlas_downloads/humantissues/Peaks/ . We retained the original cell-type labels as provided by the authors. The SNP-to-cell-type association matrix V cell-type was constructed using the same procedure applied to the bulk tissue annotations (see EpiMap dataset ) but substituting in the single-cell–derived peak tracks. J-PEP-CT: Inference of cell-type-derived cluster profiles We refer to the extension of J-PEP to single-cell resolution as J-PEP-CT. J-PEP-CT consists of two steps. (i) First, we applied J-PEP to bulk data in order to infer the shared SNP-to-cluster membership matrix W ; (ii) then, we derived cell-type–specific cluster profiles, denoted H cell-type , by solving V c e l l − t y p e ≈ W × H c e l l − t y p e , where V cell-type denotes the SNP-to-cell-type association matrix obtained from the EM algorithm on scATAC-seq data across 111 cell types (analogous to V tissue ). Estimation of H cell-type was done using a non-negative least squares procedure with multiplicative update rules analogous to those used in bNMF. Importantly, we maintained the same tissue-sparsity constraint as in the bulk analysis, restricting each cell type to be assigned to at most one cluster, while allowing the possibility of remaining unassigned. This approach enables projection of cell-type data onto bulk-derived clusters, ensuring consistency between bulk- and cell-type–level analyses. Gene ontology (GO) enrichment analysis To investigate the biological processes associated with each cluster, we conducted Gene Ontology (GO) term enrichment analysis using the “Enrichr” R package. SNP-to-cluster scores were mapped to genes using pre-annotated SNP-to-gene mappings, 46 and gene-level scores were calculated by summing the scores of all SNPs assigned to each gene. To prioritize the most informative signals, genes were ranked by their cluster scores, and only the top genes cumulatively contributing to 80% of the total cluster weight were retained. GO terms were considered significant if they included at least four overlapping genes and had a Benjamini-Hochberg adjusted p -value below 0.05. Quantification and statistical analysis The quantitative and statistical analyses are described thoroughly in the relevant sections of the method details . Published: January 26, 2026 Footnotes Supplemental information can be found online at https://doi.org/10.1016/j.xgen.2025.101138 . Contributor Information Gaspard Kerner, Email: [email protected]. Alkes L. Price, Email: [email protected]. Supplemental information Document S1. Figures S1–S15 mmc1.pdf (2.6MB, pdf) Table S1. Summary of GWAS diseases/traits, related to STAR Methods The column “Approx. genet. indep.” denotes whether a trait was part of the set of 55 approximately genetically independent traits selected for benchmarking analyses. The column “Criteria met” identifies the subset of 38 traits for which at least two robust clusters (see STAR Methods ) were identified consistently across all three partitioning methods: pleiotropic partitioning, epigenomic partitioning, and J-PEP. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. mmc2.xlsx (18.5KB, xlsx) Table S2. Simulation benchmarking: Frobenius norm error, EPA, PPA, and PEPA values, related to Figure 2 mmc3.xlsx (658KB, xlsx) Table S3. PEPA metric values across 165 diseases/traits for J-PEP, pleiotropic partitioning, and epigenomic partitioning, related to Figure 3 Missing values indicate cases where prediction was not performed because fewer than two clusters were identified in at least 5 out of the 22 chromosome-specific runs. Number of clusters is indicated for each trait-model pair (K_JPEP, K_EPI, and K_PLEIO). The column “Criteria met” identifies the subset of 38 traits for which at least two robust clusters (see STAR Methods ) were identified consistently across all three partitioning methods: pleiotropic partitioning, epigenomic partitioning, and J-PEP. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. Results for r g < 0.2 and r g < 0.9 are also shown below the default criterion of r g < 0.5. Results for the alternative J-PEP variant that inputs V tissue × (focal trait PIP) instead of the default V tissue are also reported. mmc4.xlsx (44.5KB, xlsx) Table S4. SNP-level correlation for pairs of J-PEP and pleiotropic or epigenomic partitioning clusters, related to Figures 5 and 6 SNP loadings ( W ) from J-PEP, pleiotropic partitioning, and epigenomic partitioning methods were merged by SNP. For each J-PEP cluster, we computed Pearson correlations between its W column and every cluster column from each comparison method. mmc5.xlsx (59.6KB, xlsx) Table S5. Projection correlation scores across 38 diseases/traits, related to STAR Methods Average across clusters for each focal trait. mmc6.xlsx (11.6KB, xlsx) Table S6. Epigenomic tracks used in bulk-level analyses, related to STAR Methods Number of tissue-specific epigenomic tracks and tissue-to-cluster scores for focal trait-model pairs are shown. mmc7.xlsx (349KB, xlsx) Table S7. Dominant tissues across focal trait-cluster pairs, related to STAR Methods Tissues with their proportion of times they appear to be the dominant tissue (tissue with maximum score) across focal trait-cluster pairs across 38 diseases/traits. mmc8.xlsx (11.3KB, xlsx) Table S8. Single-cell-derived cluster profiles across 38 diseases/traits, related to Figure 4 Single-cell-derived cluster profiles are shown sequentially and were obtained using the SNP-to-cluster membership matrix W derived from the respective partitioning method using bulk-level data. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. mmc9.xlsx (137.2KB, xlsx) Table S9. Cell-type-to-tissue mapping for single-cell chromatin accessibility data, related to Figure 4 H tissue matrices obtained from single-cell data are sequentially shown for different focal traits. mmc10.xlsx (10.9KB, xlsx) Table S10. SNP-to-trait matrices for T2D, HTN, and NC analyses, related to Figures 5 and 6 V trait matrices are shown one after the other for T2D (as presented in the main text), HTN, and NC. mmc11.xlsx (5.6MB, xlsx) Table S11. GO enrichment results for clusters in T2D, HTN, and NC, related to Figures 5 and 6 Significant GO terms for T2D (as presented in the main text), HTN, and NC clusters. GO terms were considered significant if they included at least four overlapping genes and had a Benjamini-Hochberg adjusted p value below 0.05. mmc12.xlsx (11.4KB, xlsx) Document S2. Transparent peer review records for Kerner et al. mmc13.pdf (4.4MB, pdf) Document S3. Article plus supplemental information mmc14.pdf (11MB, pdf) References 1. Tam V., Patel N., Turcotte M., Bossé Y., Paré G., Meyre D. Benefits and limitations of genome-wide association studies. Nat. Rev. Genet. 2019;20:467–484. doi: 10.1038/s41576-019-0127-1. [ DOI ] [ PubMed ] [ Google Scholar ] 2. Shendure J., Findlay G.M., Snyder M.W. Genomic Medicine-Progress, Pitfalls, and Promise. Cell. 2019;177:45–57. doi: 10.1016/j.cell.2019.02.003. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Claussnitzer M., Cho J.H., Collins R., Cox N.J., Dermitzakis E.T., Hurles M.E., Kathiresan S., Kenny E.E., Lindgren C.M., MacArthur D.G., et al. A brief history of human disease genetics. Nature. 2020;577:179–189. doi: 10.1038/s41586-019-1879-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Abdellaoui A., Yengo L., Verweij K.J.H., Visscher P.M. 15 years of GWAS discovery: Realizing the promise. Am. J. Hum. Genet. 2023;110:179–194. doi: 10.1016/j.ajhg.2022.12.011. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Sollis E., Mosaku A., Abid A., Buniello A., Cerezo M., Gil L., Groza T., Güneş O., Hall P., Hayhurst J., et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51:D977–D985. doi: 10.1093/nar/gkac1010. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Southam L., Zeggini E. Twenty years of genome-wide association studies. Nature. 2025;641:47–49. doi: 10.1038/d41586-025-01128-6. [ DOI ] [ PubMed ] [ Google Scholar ] 7. Musunuru K., Strong A., Frank-Kamenetsky M., Lee N.E., Ahfeldt T., Sachs K.V., Li X., Li H., Kuperwasser N., Ruda V.M., et al. From noncoding variant to phenotype via SORT1 at the 1p13 cholesterol locus. Nature. 2010;466:714–719. doi: 10.1038/nature09266. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Bauer D.E., Kamran S.C., Lessard S., Xu J., Fujiwara Y., Lin C., Shao Z., Canver M.C., Smith E.C., Pinello L., et al. An erythroid enhancer of BCL11A subject to genetic variation determines fetal hemoglobin level. Science. 2013;342:253–257. doi: 10.1126/science.1242088. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Claussnitzer M., Dankel S.N., Kim K.-H., Quon G., Meuleman W., Haugen C., Glunk V., Sousa I.S., Beaudry J.L., Puviindran V., et al. FTO Obesity Variant Circuitry and Adipocyte Browning in Humans. N. Engl. J. Med. 2015;373:895–907. doi: 10.1056/NEJMoa1502214. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Downes D.J., Cross A.R., Hua P., Roberts N., Schwessinger R., Cutler A.J., Munis A.M., Brown J., Mielczarek O., de Andrea C.E., et al. Identification of LZTFL1 as a candidate effector gene at a COVID-19 risk locus. Nat. Genet. 2021;53:1606–1615. doi: 10.1038/s41588-021-00955-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Morris J.A., Caragine C., Daniloski Z., Domingo J., Barry T., Lu L., Davis K., Ziosi M., Glinos D.A., Hao S., et al. Discovery of target genes and pathways at GWAS loci by pooled single-cell CRISPR screens. Science. 2023;380:eadh7699. doi: 10.1126/science.adh7699. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Derks E.M., Thorp J.G., Gerring Z.F. Ten challenges for clinical translation in psychiatric genetics. Nat. Genet. 2022;54:1457–1465. doi: 10.1038/s41588-022-01174-0. [ DOI ] [ PubMed ] [ Google Scholar ] 13. Cotsapas C., Voight B.F., Rossin E., Lage K., Neale B.M., Wallace C., Abecasis G.R., Barrett J.C., Behrens T., Cho J., et al. Pervasive sharing of genetic effects in autoimmune disease. PLoS Genet. 2011;7 doi: 10.1371/journal.pgen.1002254. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Solovieff N., Cotsapas C., Lee P.H., Purcell S.M., Smoller J.W. Pleiotropy in complex traits: challenges and strategies. Nat. Rev. Genet. 2013;14:483–495. doi: 10.1038/nrg3461. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Bulik-Sullivan B., Finucane H.K., Anttila V., Gusev A., Day F.R., Loh P.-R., ReproGen Consortium. Psychiatric Genomics Consortium. Genetic Consortium for Anorexia Nervosa of the Wellcome Trust Case Control Consortium 3. Duncan L. An atlas of genetic correlations across human diseases and traits. Nat. Genet. 2015;47:1236–1241. doi: 10.1038/ng.3406. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Pickrell J.K., Berisa T., Liu J.Z., Ségurel L., Tung J.Y., Hinds D.A. Detection and interpretation of shared genetic influences on 42 human traits. Nat. Genet. 2016;48:709–717. doi: 10.1038/ng.3570. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Tanigawa Y., Li J., Justesen J.M., Horn H., Aguirre M., DeBoever C., Chang C., Narasimhan B., Lage K., Hastie T., et al. Components of genetic associations across 2,138 phenotypes in the UK Biobank highlight adipocyte biology. Nat. Commun. 2019;10:4064. doi: 10.1038/s41467-019-11953-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Grotzinger A.D., Rhemtulla M., de Vlaming R., Ritchie S.J., Mallard T.T., Hill W.D., Ip H.F., Marioni R.E., McIntosh A.M., Deary I.J., et al. Genomic structural equation modelling provides insights into the multivariate genetic architecture of complex traits. Nat. Hum. Behav. 2019;3:513–525. doi: 10.1038/s41562-019-0566-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Ballard J.L., O’Connor L.J. Shared components of heritability across genetically correlated traits. Am. J. Hum. Genet. 2022;109:989–1006. doi: 10.1016/j.ajhg.2022.04.003. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. Grotzinger A.D., Mallard T.T., Akingbuwa W.A., Ip H.F., Adams M.J., Lewis C.M., McIntosh A.M., Grove J., Dalsgaard S., Lesch K.-P., et al. Genetic architecture of 11 major psychiatric disorders at biobehavioral, functional genomic and molecular genetic levels of analysis. Nat. Genet. 2022;54:548–559. doi: 10.1038/s41588-022-01057-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Romero C., Werme J., Jansen P.R., Gelernter J., Stein M.B., Levey D., Polimanti R., de Leeuw C., Posthuma D., Nagel M., van der Sluis S. Exploring the genetic overlap between twelve psychiatric disorders. Nat. Genet. 2022;54:1795–1802. doi: 10.1038/s41588-022-01245-2. [ DOI ] [ PubMed ] [ Google Scholar ] 22. Mackay T.F.C., Anholt R.R.H. Pleiotropy, epistasis and the genetic architecture of quantitative traits. Nat. Rev. Genet. 2024;25:639–657. doi: 10.1038/s41576-024-00711-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 23. Udler M.S., Kim J., von Grotthuss M., Bonàs-Guarch S., Cole J.B., Chiou J., Anderson C.D., on behalf of METASTROKE and the ISGC. Boehnke M., Laakso M., Atzmon G. Type 2 diabetes genetic loci informed by multi-trait associations point to disease mechanisms and subtypes: A soft clustering analysis. PLoS Med. 2018;15 doi: 10.1371/journal.pmed.1002654. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. Suzuki K., Hatzikotoulas K., Southam L., Taylor H.J., Yin X., Lorenz K.M., Mandla R., Huerta-Chagoya A., Melloni G.E.M., Kanoni S., et al. Genetic drivers of heterogeneity in type 2 diabetes pathophysiology. Nature. 2024;627:347–357. doi: 10.1038/s41586-024-07019-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 25. Smith K., Deutsch A.J., McGrail C., Kim H., Hsu S., Huerta-Chagoya A., Mandla R., Schroeder P.H., Westerman K.E., Szczerbinski L., et al. Multi-ancestry polygenic mechanisms of type 2 diabetes. Nat. Med. 2024;30:1065–1074. doi: 10.1038/s41591-024-02865-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Zhang Z., Jung J., Kim A., Suboc N., Gazal S., Mancuso N. A scalable approach to characterize pleiotropy across thousands of human diseases and complex traits using GWAS summary statistics. Am. J. Hum. Genet. 2023;110:1863–1874. doi: 10.1016/j.ajhg.2023.09.015. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Kim H., Westerman K.E., Smith K., Chiou J., Cole J.B., Majarian T., von Grotthuss M., Kwak S.H., Kim J., Mercader J.M., et al. High-throughput genetic clustering of type 2 diabetes loci reveals heterogeneous mechanistic pathways of metabolic disease. Diabetologia. 2023;66:495–507. doi: 10.1007/s00125-022-05848-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Reeve M., Kanai M., Graham D., Karjalainen J., Luo S., Kolosov N., Adams C., Ritari J., Karczewski K., Kiiskinen T., et al. Autoimmune hypothyroidism GWAS reveals independent autoimmune and thyroid-specific contributions and an inverse relation with cancer risk. Res. Sq. 2024 doi: 10.21203/rs.3.rs-4626646/v1. Preprint at. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Qi G., Chhetri S.B., Ray D., Dutta D., Battle A., Bhattacharjee S., Chatterjee N. Genome-wide large-scale multi-trait analysis characterizes global patterns of pleiotropy and unique trait-specific variants. Nat. Commun. 2024;15:6985. doi: 10.1038/s41467-024-51075-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Ghatan S., van Rooij J., van Hoek M., Boer C.G., Felix J.F., Kavousi M., Jaddoe V.W., Sijbrands E.J.G., Medina-Gomez C., Rivadeneira F., Oei L. Defining type 2 diabetes polygenic risk scores through colocalization and network-based clustering of metabolic trait genetic associations. Genome Med. 2024;16:10. doi: 10.1186/s13073-023-01255-7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Li Y., Chen G.-C., Moon J.-Y., Arthur R., Sotres-Alvarez D., Daviglus M.L., Pirzada A., Mattei J., Perreira K.M., Rotter J.I., et al. Genetic Subtypes of Prediabetes, Healthy Lifestyle, and Risk of Type 2 Diabetes. Diabetes. 2024;73:1178–1187. doi: 10.2337/db23-0699. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Roadmap Epigenomics Consortium. Kundaje A., Meuleman W., Ernst J., Bilenky M., Yen A., Heravi-Moussavi A., Kheradpour P., Zhang Z., Wang J. Integrative analysis of 111 reference human epigenomes. Nature. 2015;518:317–330. doi: 10.1038/nature14248. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Farh K.K.-H., Marson A., Zhu J., Kleinewietfeld M., Housley W.J., Beik S., Shoresh N., Whitton H., Ryan R.J.H., Shishkin A.A., et al. Genetic and epigenetic fine mapping of causal autoimmune disease variants. Nature. 2015;518:337–343. doi: 10.1038/nature13835. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Finucane H.K., Bulik-Sullivan B., Gusev A., Trynka G., Reshef Y., Loh P.-R., Anttila V., Xu H., Zang C., Farh K., et al. Partitioning heritability by functional annotation using genome-wide association summary statistics. Nat. Genet. 2015;47:1228–1235. doi: 10.1038/ng.3404. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Boix C.A., James B.T., Park Y.P., Meuleman W., Kellis M. Regulatory genomic circuitry of human disease loci by integrative epigenomics. Nature. 2021;590:300–307. doi: 10.1038/s41586-020-03145-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Gaulton K.J., Preissl S., Ren B. Interpreting non-coding disease-associated human variants using single-cell epigenomics. Nat. Rev. Genet. 2023;24:516–534. doi: 10.1038/s41576-023-00598-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. Schnitzler G.R., Kang H., Fang S., Angom R.S., Lee-Kim V.S., Ma X.R., Zhou R., Zeng T., Guo K., Taylor M.S., et al. Convergence of coronary artery disease genes onto endothelial cell programs. Nature. 2024;626:799–807. doi: 10.1038/s41586-024-07022-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Wellcome Trust Case Control Consortium. Maller J.B., McVean G., Byrnes J., Vukcevic D., Palin K., Su Z., Howson J.M.M., Auton A., Myers S. Bayesian refinement of association signals for 14 loci in 3 common diseases. Nat. Genet. 2012;44:1294–1301. doi: 10.1038/ng.2435. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Schaid D.J., Chen W., Larson N.B. From genome-wide associations to candidate causal variants by statistical fine-mapping. Nat. Rev. Genet. 2018;19:491–504. doi: 10.1038/s41576-018-0016-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Schmidt M.N., Winther O., Hansen L.K. In: Independent Component Analysis and Signal Separation. Adali T., Jutten C., Romano J.M.T., Barros A.K., editors. Springer; 2009. Bayesian Non-negative Matrix Factorization; pp. 540–547. [ DOI ] [ Google Scholar ] 41. Tan V.Y.F., Févotte C. Automatic Relevance Determination in Nonnegative Matrix Factorization with the/spl beta/-Divergence. IEEE Trans. Pattern Anal. Mach. Intell. 2013;35:1592–1605. doi: 10.1109/TPAMI.2012.240. [ DOI ] [ PubMed ] [ Google Scholar ] 42. Willwerscheid J., Carbonetto P., Stephens M. ebnm: An R Package for Solving the Empirical Bayes Normal Means Problem Using a Variety of Prior Families. arXiv. 2024 doi: 10.48550/arXiv.2110.00152. Preprint at. [ DOI ] [ Google Scholar ] 43. Bycroft C., Freeman C., Petkova D., Band G., Elliott L.T., Sharp K., Motyer A., Vukcevic D., Delaneau O., O’Connell J., et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562:203–209. doi: 10.1038/s41586-018-0579-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 44. Zhang K., Hocker J.D., Miller M., Hou X., Chiou J., Poirion O.B., Qiu Y., Li Y.E., Gaulton K.J., Wang A., et al. A single-cell atlas of chromatin accessibility in the human genome. Cell. 2021;184:5985–6001.e19. doi: 10.1016/j.cell.2021.10.024. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Young M.D., Wakefield M.J., Smyth G.K., Oshlack A. Gene ontology analysis for RNA-seq: accounting for selection bias. Genome Biol. 2010;11:R14. doi: 10.1186/gb-2010-11-2-r14. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 46. Gazal S., Weissbrod O., Hormozdiari F., Dey K.K., Nasser J., Jagadeesh K.A., Weiner D.J., Shi H., Fulco C.P., O’Connor L., et al. Combining SNP-to-gene linking strategies to identify disease genes and assess disease omnigenicity. Nat. Genet. 2022;54:827–836. doi: 10.1038/s41588-022-01087-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 47. Chen K.M., Wong A.K., Troyanskaya O.G., Zhou J. A sequence-based global map of regulatory activity for deciphering human genetics. Nat. Genet. 2022;54:940–949. doi: 10.1038/s41588-022-01102-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 48. Gao V.R., Yang R., Das A., Luo R., Luo H., McNally D.R., Karagiannidis I., Rivas M.A., Wang Z.-M., Barisic D., et al. ChromaFold predicts the 3D contact map from single-cell chromatin accessibility. Nat. Commun. 2024;15:9432. doi: 10.1038/s41467-024-53628-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 49. Fu X., Mo S., Buendia A., Laurent A.P., Shao A., Alvarez-Torres M.D.M., Yu T., Tan J., Su J., Sagatelian R., et al. A foundation model of transcription across human cell types. Nature. 2025;637:965–973. doi: 10.1038/s41586-024-08391-z. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 50. Dorans E., Jagadeesh K., Dey K., Price A.L. Linking regulatory variants to target genes by integrating single-cell multiome methods and genomic distance. Nat. Genet. 2025;57:1649–1658. doi: 10.1038/s41588-025-02220-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 51. 1000 Genomes Project Consortium. Auton A., Brooks L.D., Durbin R.M., Garrison E.P., Kang H.M., Korbel J.O., Marchini J.L., McCarthy S., McVean G.A., Abecasis G.R. A global reference for human genetic variation. Nature. 2015;526:68–74. doi: 10.1038/nature15393. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 52. Weissbrod O., Hormozdiari F., Benner C., Cui R., Ulirsch J., Gazal S., Schoech A.P., van de Geijn B., Reshef Y., Márquez-Luna C., et al. Functionally informed fine-mapping and polygenic localization of complex trait heritability. Nat. Genet. 2020;52:1355–1363. doi: 10.1038/s41588-020-00735-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 53. Tobias D.K., Merino J., Ahmad A., Aiken C., Benham J.L., Bodhini D., Clark A.L., Colclough K., Corcoy R., Cromer S.J., et al. Second international consensus report on gaps and opportunities for the clinical translation of precision diabetes medicine. Nat. Med. 2023;29:2438–2457. doi: 10.1038/s41591-023-02502-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 54. Misra S., Wagner R., Ozkan B., Schön M., Sevilla-Gonzalez M., Prystupa K., Wang C.C., Kreienkamp R.J., Cromer S.J., Rooney M.R., et al. Precision subclassification of type 2 diabetes: a systematic review. Commun. Med. 2023;3:138. doi: 10.1038/s43856-023-00360-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 55. DiCorpo D., LeClair J., Cole J.B., Sarnowski C., Ahmadizar F., Bielak L.F., Blokstra A., Bottinger E.P., Chaker L., Chen Y.-D.I., et al. Type 2 Diabetes Partitioned Polygenic Scores Associate With Disease Outcomes in 454,193 Individuals Across 13 Cohorts. Diabetes Care. 2022;45:674–683. doi: 10.2337/dc21-1395. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 56. Leibiger I.B., Leibiger B., Berggren P.-O. Insulin Signaling in the Pancreatic β-Cell. Annu. Rev. Nutr. 2008;28:233–251. doi: 10.1146/annurev.nutr.28.061807.155530. [ DOI ] [ PubMed ] [ Google Scholar ] 57. Dooley J., Tian L., Schonefeldt S., Delghingaro-Augusto V., Garcia-Perez J.E., Pasciuto E., Di Marino D., Carr E.J., Oskolkov N., Lyssenko V., et al. Genetic predisposition for beta cell fragility underlies type 1 and type 2 diabetes. Nat. Genet. 2016;48:519–527. doi: 10.1038/ng.3531. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 58. Ghionescu A.-V., Sorop A., Linioudaki E., Coman C., Savu L., Fogarasi M., Lixandru D., Dima S.O. A predicted epithelial-to-mesenchymal transition-associated mRNA/miRNA axis contributes to the progression of diabetic liver disease. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-77416-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 59. Donath M.Y., Shoelson S.E. Type 2 diabetes as an inflammatory disease. Nat. Rev. Immunol. 2011;11:98–107. doi: 10.1038/nri2925. [ DOI ] [ PubMed ] [ Google Scholar ] 60. Luchsinger J.A. Type 2 diabetes and cognitive impairment: linking mechanisms. J. Alzheimers Dis. 2012;3:S185–S198. doi: 10.3233/JAD-2012-111433. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 61. Roh E., Song D.K., Kim M.-S. Emerging role of the brain in the homeostatic regulation of energy and glucose metabolism. Exp. Mol. Med. 2016;48:e216. doi: 10.1038/emm.2016.4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 62. Evangelou E., Warren H.R., Mosen-Ansorena D., Mifsud B., Pazoki R., Gao H., Ntritsos G., Dimou N., Cabrera C.P., Karaman I., et al. Genetic analysis of over 1 million people identifies 535 new loci associated with blood pressure traits. Nat. Genet. 2018;50:1412–1425. doi: 10.1038/s41588-018-0205-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 63. Keaton J.M., Kamali Z., Xie T., Vaez A., Williams A., Goleva S.B., Ani A., Evangelou E., Hellwege J.N., Yengo L., et al. Genome-wide analysis in over 1 million individuals of European ancestry yields improved polygenic risk scores for blood pressure traits. Nat. Genet. 2024;56:778–791. doi: 10.1038/s41588-024-01714-w. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 64. Vaura F., Kim H., Udler M.S., Niiranen T., FinnGen. Salomaa V., Lahti L. Multi-Trait Genetic Analysis Reveals Clinically Interpretable Hypertension Subtypes. Circ. Genom. Precis. Med. 2022;15 doi: 10.1161/CIRCGEN.121.003583. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 65. Pascat V., Zudina L., Maurin L., Ulrich A., Maina J.G., Demirkan A., Balkhiyarova Z., Pupko I., Sharhorodska Y., Pattou F., et al. Mechanistic Heterogeneity in Type 2 Diabetes and Hypertension Comorbidity Revealed with Partitioned Polygenic Scores. medRxiv. 2025 doi: 10.1101/2025.03.02.25323190. Preprint at. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 66. Jelinic M., Dona M.S., Hughes T.A.G., Farrugia G., Harper R., Hsu I., Haslem A., Tran V., Galvao H.B.F., Diep H., et al. High-resolution transcriptomic profiling of the aortic cellular landscape during hypertension reveals novel drivers of vascular fibrosis. bioRxiv. 2025 doi: 10.1101/2025.03.13.642936. Preprint at. [ DOI ] [ Google Scholar ] 67. Wagenseil J.E., Mecham R.P. Vascular extracellular matrix and arterial mechanics. Physiol. Rev. 2009;89:957–989. doi: 10.1152/physrev.00041.2008. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 68. Humphrey J.D., Schwartz M.A. Vascular Mechanobiology: Homeostasis, Adaptation, and Disease. Annu. Rev. Biomed. Eng. 2021;23:1–27. doi: 10.1146/annurev-bioeng-092419-060810. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 69. Rampino T., Dal Canton A. Peritoneal Dialysis and Epithelial-to-Mesenchymal Transition (2003) N. Engl. J. Med. 2003;348:2037–2039. doi: 10.1056/NEJM200305153482020. [ DOI ] [ PubMed ] [ Google Scholar ] 70. Wang J., Fan C.-N., Wei G.-X., Zhao X.-J., Liu K., Wang M.-L., Li L., Liu M., Zhao H.-Y. Recurrence of primary aldosteronism 20 years after surgery: a case report. J. Hum. Hypertens. 2025;39:237–239. doi: 10.1038/s41371-025-00994-x. [ DOI ] [ PubMed ] [ Google Scholar ] 71. Ferreira N.S., Tostes R.C., Paradis P., Schiffrin E.L. Aldosterone, Inflammation, Immune System, and Hypertension. Am. J. Hypertens. 2021;34:15–27. doi: 10.1093/ajh/hpaa137. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 72. Kurki M.I., Karjalainen J., Palta P., Sipilä T.P., Kristiansson K., Donner K.M., Reeve M.P., Laivuori H., Aavikko M., Kaunisto M.A., et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613:508–518. doi: 10.1038/s41586-022-05473-8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 73. Vuckovic D., Bao E.L., Akbari P., Lareau C.A., Mousas A., Jiang T., Chen M.-H., Raffield L.M., Tardaguila M., Huffman J.E., et al. The Polygenic and Monogenic Basis of Blood Traits and Diseases. Cell. 2020;182:1214–1231.e11. doi: 10.1016/j.cell.2020.08.008. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 74. Zhang F., Xia Y., Su J., Quan F., Zhou H., Li Q., Feng Q., Lin C., Wang D., Jiang Z. Neutrophil diversity and function in health and disease. Signal Transduct. Target. Ther. 2024;9:343–349. doi: 10.1038/s41392-024-02049-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 75. Herrero-Cervera A., Soehnlein O., Kenne E. Neutrophils in chronic inflammatory diseases. Cell. Mol. Immunol. 2022;19:177–191. doi: 10.1038/s41423-021-00832-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 76. Vaibhav K., Braun M., Alverson K., Khodadadi H., Kutiyanawalla A., Ward A., Banerjee C., Sparks T., Malik A., Rashid M.H., et al. Neutrophil extracellular traps exacerbate neurological deficits after traumatic brain injury. Sci. Adv. 2020;6 doi: 10.1126/sciadv.aax8847. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 77. Zenaro E., Pietronigro E., Della Bianca V., Piacentino G., Marongiu L., Budui S., Turano E., Rossi B., Angiari S., Dusi S., et al. Neutrophils promote Alzheimer’s disease–like pathology and cognitive decline via LFA-1 integrin. Nat. Med. 2015;21:880–886. doi: 10.1038/nm.3913. [ DOI ] [ PubMed ] [ Google Scholar ] 78. Trajanoska K., Bhérer C., Taliun D., Zhou S., Richards J.B., Mooser V. From target discovery to clinical drug development with human genetics. Nature. 2023;620:737–745. doi: 10.1038/s41586-023-06388-8. [ DOI ] [ PubMed ] [ Google Scholar ] 79. Minikel E.V., Painter J.L., Dong C.C., Nelson M.R. Refining the impact of genetic evidence on clinical success. Nature. 2024;629:624–629. doi: 10.1038/s41586-024-07316-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 80. Lumeng C.N., Bodzin J.L., Saltiel A.R. Obesity induces a phenotypic switch in adipose tissue macrophage polarization. J. Clin. Investig. 2007;117:175–184. doi: 10.1172/JCI29881. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 81. Jin L., Deng Z., Zhang J., Yang C., Liu J., Han W., Ye P., Si Y., Chen G. Mesenchymal stem cells promote type 2 macrophage polarization to ameliorate the myocardial injury caused by diabetic cardiomyopathy. J. Transl. Med. 2019;17:251. doi: 10.1186/s12967-019-1999-8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 82. Cuomo A.S.E., Nathan A., Raychaudhuri S., MacArthur D.G., Powell J.E. Single-cell genomics meets human genetics. Nat. Rev. Genet. 2023;24:535–549. doi: 10.1038/s41576-023-00599-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 83. Kanemaru K., Cranley J., Muraro D., Miranda A.M.A., Ho S.Y., Wilbrey-Clark A., Patrick Pett J., Polanski K., Richardson L., Litvinukova M., et al. Spatially resolved multiomics of human cardiac niches. Nature. 2023;619:801–810. doi: 10.1038/s41586-023-06311-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 84. Emani P.S., Liu J.J., Clarke D., Jensen M., Warrell J., Gupta C., Meng R., Lee C.Y., Xu S., Dursun C., et al. Single-cell genomics and regulatory networks for 388 human brains. Science. 2024;384:eadi5199. doi: 10.1126/science.adi5199. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 85. Richard D., Muthuirulan P., Young M., Yengo L., Vedantam S., Marouli E., Bartell E., GIANT Consortium. Hirschhorn J., Capellini T.D. Functional genomics of human skeletal development and the patterning of height heritability. Cell. 2025;188:15–32.e24. doi: 10.1016/j.cell.2024.10.040. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 86. Lee D.D., Seung H.S. Learning the parts of objects by non-negative matrix factorization. Nature. 1999;401:788–791. doi: 10.1038/44565. [ DOI ] [ PubMed ] [ Google Scholar ] 87. Wang W., Stephens M. Empirical Bayes Matrix Factorization. J. Mach. Learn. Res. 2021;22:120. [ PMC free article ] [ PubMed ] [ Google Scholar ] 88. Hou L., Xiong X., Park Y., Boix C., James B., Sun N., He L., Patel A., Zhang Z., Molinie B., et al. Multitissue H3K27ac profiling of GTEx samples links epigenomic variation to disease. Nat. Genet. 2023;55:1665–1676. doi: 10.1038/s41588-023-01509-5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 89. Hill W.G., Goddard M.E., Visscher P.M. Data and Theory Point to Mainly Additive Genetic Variance for Complex Traits. PLoS Genet. 2008;4 doi: 10.1371/journal.pgen.1000008. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 90. Mäki-Tanila A., Hill W.G. Influence of gene interaction on complex trait variation with multilocus models. Genetics. 2014;198:355–367. doi: 10.1534/genetics.114.165282. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 91. Zeng J., de Vlaming R., Wu Y., Robinson M.R., Lloyd-Jones L.R., Yengo L., Yap C.X., Xue A., Sidorenko J., McRae A.F., et al. Signatures of negative selection in the genetic architecture of human complex traits. Nat. Genet. 2018;50:746–753. doi: 10.1038/s41588-018-0101-4. [ DOI ] [ PubMed ] [ Google Scholar ] 92. Gazal S., Loh P.-R., Finucane H.K., Ganna A., Schoech A., Sunyaev S., Price A.L. Functional architecture of low-frequency variants highlights strength of negative selection across coding and non-coding annotations. Nat. Genet. 2018;50:1600–1607. doi: 10.1038/s41588-018-0231-8. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 93. Schoech A.P., Jordan D.M., Loh P.-R., Gazal S., O’Connor L.J., Balick D.J., Palamara P.F., Finucane H.K., Sunyaev S.R., Price A.L. Quantification of frequency-dependent genetic architectures in 25 UK Biobank traits reveals action of negative selection. Nat. Commun. 2019;10:790. doi: 10.1038/s41467-019-08424-6. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 94. O’Connor L.J., Schoech A.P., Hormozdiari F., Gazal S., Patterson N., Price A.L. Extreme Polygenicity of Complex Traits Is Explained by Negative Selection. Am. J. Hum. Genet. 2019;105:456–476. doi: 10.1016/j.ajhg.2019.07.003. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 95. Zhu X., Stephens M. Bayesian large-scale multiple regression with summary statistics from genome-wide association studies. Ann. Appl. Stat. 2017;11:1561–1592. doi: 10.1214/17-aoas1046. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Document S1. Figures S1–S15 mmc1.pdf (2.6MB, pdf) Table S1. Summary of GWAS diseases/traits, related to STAR Methods The column “Approx. genet. indep.” denotes whether a trait was part of the set of 55 approximately genetically independent traits selected for benchmarking analyses. The column “Criteria met” identifies the subset of 38 traits for which at least two robust clusters (see STAR Methods ) were identified consistently across all three partitioning methods: pleiotropic partitioning, epigenomic partitioning, and J-PEP. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. mmc2.xlsx (18.5KB, xlsx) Table S2. Simulation benchmarking: Frobenius norm error, EPA, PPA, and PEPA values, related to Figure 2 mmc3.xlsx (658KB, xlsx) Table S3. PEPA metric values across 165 diseases/traits for J-PEP, pleiotropic partitioning, and epigenomic partitioning, related to Figure 3 Missing values indicate cases where prediction was not performed because fewer than two clusters were identified in at least 5 out of the 22 chromosome-specific runs. Number of clusters is indicated for each trait-model pair (K_JPEP, K_EPI, and K_PLEIO). The column “Criteria met” identifies the subset of 38 traits for which at least two robust clusters (see STAR Methods ) were identified consistently across all three partitioning methods: pleiotropic partitioning, epigenomic partitioning, and J-PEP. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. Results for r g < 0.2 and r g < 0.9 are also shown below the default criterion of r g < 0.5. Results for the alternative J-PEP variant that inputs V tissue × (focal trait PIP) instead of the default V tissue are also reported. mmc4.xlsx (44.5KB, xlsx) Table S4. SNP-level correlation for pairs of J-PEP and pleiotropic or epigenomic partitioning clusters, related to Figures 5 and 6 SNP loadings ( W ) from J-PEP, pleiotropic partitioning, and epigenomic partitioning methods were merged by SNP. For each J-PEP cluster, we computed Pearson correlations between its W column and every cluster column from each comparison method. mmc5.xlsx (59.6KB, xlsx) Table S5. Projection correlation scores across 38 diseases/traits, related to STAR Methods Average across clusters for each focal trait. mmc6.xlsx (11.6KB, xlsx) Table S6. Epigenomic tracks used in bulk-level analyses, related to STAR Methods Number of tissue-specific epigenomic tracks and tissue-to-cluster scores for focal trait-model pairs are shown. mmc7.xlsx (349KB, xlsx) Table S7. Dominant tissues across focal trait-cluster pairs, related to STAR Methods Tissues with their proportion of times they appear to be the dominant tissue (tissue with maximum score) across focal trait-cluster pairs across 38 diseases/traits. mmc8.xlsx (11.3KB, xlsx) Table S8. Single-cell-derived cluster profiles across 38 diseases/traits, related to Figure 4 Single-cell-derived cluster profiles are shown sequentially and were obtained using the SNP-to-cluster membership matrix W derived from the respective partitioning method using bulk-level data. Note that “All_Metal_LDSC-CORR_Neff” denotes the trait T2D when analyzed with the default set of 82 auxiliary traits from this work, whereas “All_Metal_LDSC-CORR_Neff_comparison” denotes the trait T2D when analyzed with the set of auxiliary traits from Smith et al. 25 In both cases, we use the same summary statistics for T2D. mmc9.xlsx (137.2KB, xlsx) Table S9. Cell-type-to-tissue mapping for single-cell chromatin accessibility data, related to Figure 4 H tissue matrices obtained from single-cell data are sequentially shown for different focal traits. mmc10.xlsx (10.9KB, xlsx) Table S10. SNP-to-trait matrices for T2D, HTN, and NC analyses, related to Figures 5 and 6 V trait matrices are shown one after the other for T2D (as presented in the main text), HTN, and NC. mmc11.xlsx (5.6MB, xlsx) Table S11. GO enrichment results for clusters in T2D, HTN, and NC, related to Figures 5 and 6 Significant GO terms for T2D (as presented in the main text), HTN, and NC clusters. GO terms were considered significant if they included at least four overlapping genes and had a Benjamini-Hochberg adjusted p value below 0.05. mmc12.xlsx (11.4KB, xlsx) Document S2. Transparent peer review records for Kerner et al. mmc13.pdf (4.4MB, pdf) Document S3. Article plus supplemental information mmc14.pdf (11MB, pdf) Data Availability Statement All data used in this study and the corresponding access information are listed in the key resources table . This includes the jpepR package, developed as part of this work, which provides an environment to run J-PEP, as well as GWAS summary statistics for 165 traits and the corresponding J-PEP results. Links to the public software and algorithms used in the present study are listed in the key resources table . Articles from Cell Genomics are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (5.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 4034 · SHA-256 63d69d24f4c03d1f
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.