ConceptioArchiveZenodo (CERN)
Zenodo (CERN)open access

Population-scale single-nucleus multi-omic profiling of skeletal muscle reveals extensive context-specific genetic regulation

Varshney, Arushi · Zenodo (CERN)
Zenodo (CERN) · Papers · License: Open Access
Open Source ↗
eQTL, caQTL, snRNA, snATAC, GWAS

Population-scale single-nucleus multi-omic profiling of skeletal muscle reveals extensive context-specific genetic regulation | Zenodo Skip to main Communities My dashboard Log in Sign up Published April 20, 2026 | Version v6 Dataset Open Population-scale single-nucleus multi-omic profiling of skeletal muscle reveals extensive context-specific genetic regulation Authors/Creators Varshney, Arushi (Researcher) 1 Show affiliations 1. University of Michigan Description Data accompanying the manuscript "Population-scale skeletal muscle single-nucleus multi-omic profiling reveals extensive context specific genetic regulation", that is accepted for publication in the journal Nature Genetics. Note: For ATAC fragment files, e,caQTL full cis scan summary files, clustering objects, please see the CMDGA portal (https://cmdga.org/search/?searchTerm=stephen-parker%3AVarshney2024) For raw data including fastq files, please see dbGaP repo phs001048.v3.p1 Data in this repository includes: Filename: Description 1. genes_with_exon_counts_considered.txt : list of 8,666 genes for which exon-only counts were considered. See methods section "Adjusting RNA counts for overlapping gene annotations" in the manuscript. 2. nucleus_sample_cluster_map.tsv: nucleus-sample-cluster map with other QC info. # index: nucleus identified syntax <modality>.<batch>.NM.<10X channel>.<barcode> # UMAP_1, UMAP_2: UMAP coordinates for visualization # modality: rna or atac # batch: processing batch identifier # hqaa_umi: high quality autosomal alignments (HQAA) for atac nuclei, unique molecular identifier (UMI) for tna # fraction_mitochondrial: fraction of reads mapping to the mitochondrial genome # cohort: sample cohort # tss_enrichment: TSS enrichment for atac nuclei # coarse_cluster_name: cluster name 3. peaks.tar.gz: snATAC peak features including: # consensus-summits.bed: consensus summits along with the cell type that the summits was highest in. # narrow peaks in clusters # consensus summit feature (summit +- 150bp) identified in each cluster - these were used in GWAS enrichments. 4. snrna-cell-type-specific-genes.tsv: Normalized expression scores for genes in each cell-type cluster 5. eqtl_permute.tar.gz: Permutation scan eQTL in each cell-type cluster. Columns: # variant: syntax <chrom>:<hg38 pos>:<ref>:<alt> # effect_allele: effect allele (was the alt allele) # other_allele: non-effect allele # feature: gene name # featureCoordinates_tss: gene TSS # p-value: nominal p value # beta: slope/beta of the linear regression. Keyed on the alt allele # se: standard error of the slope # snp: SNP ID # strand: gene strand # n_variants_tested: number of variants tested for the gene # distance_var_pheno: distance of the variant with the gene TSS # n_effective_tests: number of effective tests # p_beta: beta distribution adjusted p value # qvalue: qvalue (Storey) 6. caqtl_permute.tar.gz: # Permutation scan caQTL in each cell-type cluster. Columns: # variant: syntax <chrom>:<hg38 pos>:<ref>:<alt> # effect_allele: effect allele (was the alt allele) # other_allele: non-effect allele # feature: peak feature coordinates # p-value: nominal p value # beta: slope/beta of the linear regression. Keyed on the alt allele # se: standard error of the slope # snp: SNP ID # n_variants_tested: number of variants tested for the gene # distance_var_pheno: distance of the variant with the gene TSS # n_effective_tests: number of effective tests # p_beta: beta distribution adjusted p value # qvalue: qvalue (Storey) 7. eqtl_credible_sets.tar.gz: # eQTL credible set. The file name denotes the egene and the signal hit id. Bed file columns: # 1: snp chromosome # 2: snp start # 3: snp end # 4: snp chrom_pos_ref_alt # 5: Bayes Factor # 6: PIP # 7: SNP rsid 8. caqtl_credible_sets.tar.gz: # caqtl credible set. The file name denotes the capeak and the signal hit id. Bed file columns: # 1: snp chromosome # 2: snp start # 3: snp end # 4: snp chrom_pos_ref_alt # 5: Bayes Factor # 6: PIP # 7: SNP rsid 9. cicero_all.tar.gz # Cicero coaccessibility results. Columns # Peak 1: Macs2 narrowpeak coordinate for peak 1 # Peak 2: Macs2 narrowpeak coordinate for peak 2 # coaccess: Cicero coaccessibility score 10. cicero_gene_tss.tar.gz: Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name. Columns # Cicero coaccessibility results between peak and genes. Macs2 narrow peaks in the TSS+1kb upstream region are assigned that gene name.Columns # Peak 1: Macs2 narrowpeak coordinate for peak 1 # gene_name: Assigned gene # Peak 2: Macs2 narrowpeak coordinate for peak 2 # coaccess: Cicero coaccessibility score ##  Peak1 is the narrowpeak in the TSS region, peak2 is the distal peak 11. consensus_to_narrowpeak_map.tar.gz # Cicero analysis used the full narrowpeak coordinates whereas caQTL scans used consensus summits - extended by 150 bp on either side. This tar provides files per cluster that map the consensus summit feature to narrowpeak coordinates in each file. 12. mash.tar.gz Mashr results for e/caQTL - lfsr, posterior means and posterior SD for each tested eSNP-eGene, caSNP-caPeak pair. 13. cellregmap.tar.gz: Cellregmap results for endothelial nucleus-level eQTL scans. ## Persistent genetic effect beta_g was calculated in a simple association model. ## An interaction model was fit to test for GxC effect. columns: # rho1, g2, e1, and eps2 are variance component measures outputs from CellRegMap corresponding to interaction, genetic, environment and residual variance components. # p_nominal: nominal p from cellRegMap # kind: model kind in CellRegMap - simple association or interaction # beta_g:  Persistent genetic effect # gene_name: gene name for eQTL or peak feature name for caQTL # context: context used either factors (continuous) or subclusters (discrete) # snp: index snp for which model is fit. This is the most significant identified snp from our standard e,caQTL scans. chrom-hg38pos-rsid 14. coloc-eqtl-caqtl.tsv: # Summary of eQTL-caQTL coloc in each cluster. Columns: # nsnps: Number of SNPs in the region # eqtl_hit: SNP with the highest Bayes factor in the SuSiE eQTL credible set # caqtl_hit: SNP with the highest Bayes factor in the SuSiE caQTL credible set # PP.H0.abf: Coloc posterior probability for no signal # PP.H1.abf: Coloc posterior probability for signal in dataset 1 # PP.H2.abf: Coloc posterior probability for signal in dataset 2 # PP.H3.abf: Coloc posterior probability for different signals in datasets 1 and 2 # PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2 # idx1: Index of the SuSiE credible set for dataset 1 # idx2: Index of the SuSiE credible set for dataset 2 # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates 15. cit-mrs-summary.tsv:  Summary from CIT and MR Steiger directionality tests. Columns: # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates # eqhit: SNP with the highest Bayes factor in the SuSiE eQTL credible set # cahit: SNP with the highest Bayes factor in the SuSiE caQTL credible set # p.cit_c_c-e: P value for CIT causal cahit-ca-to-e model # q.cit_c_c-e: q value for CIT causal cahit-ca-to-e model # p.cit_rc_c-e: P value for CIT reverse-causal eqhit-ca-to-e model # q.cit_rc_c-e: value for CIT reverse-causal eqhit-ca-to-e model # p.cit_c_e-c: P value for CIT causal eqhit-e-to-ca model # q.cit_c_e-c: q value for CIT causal eqhit-e-to-ca model # p.cit_rc_e-c: P value for CIT reverse-causal cahit-e-to-ca model # q.cit_rc_e-c: q value for CIT reverse-causal cahit-e-to-ca model # cit_direction: Direction inferred from CIT # correct_causal_direction--ca-to-e: MR Steiger directionality test - is ca-to-e direction correct? # correct_causal_direction--e-to-ca: MR Steiger directionality test - is e-to-ca direction correct? # sensitivity_ratio--ca-to-e: MR Steiger Sensitivity ratio for ca-to-e model # sensitivity_ratio--e-to-ca:  MR Steiger Sensitivity ratio for e-to-ca model # steiger_test--ca-to-e: MR Steiger directionality test P value for ca-to-e model # steiger_test--e-to-ca: MR Steiger directionality test P value for e-to-ca model # steiger_q--ca-to-e: MR Steiger directionality test q value for ca-to-e model # steiger_q--e-to-ca: MR Steiger directionality test q value for e-to-ca model # mrs_direction: Direction inferred from MR Steiger # direction: Direction inferred requiring consistent results between CIT and MR Steiger directionality test 16. coloc-gwas-eqtl.tsv and 17. coloc-gwas-caqtl.tsv # Summary of e/caQTL coloc with GWAS in each cluster. Columns: # nsnps: Number of SNPs in the region # gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set # eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set # caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set # PP.H0.abf: Coloc posterior probability for no signal # PP.H1.abf: Coloc posterior probability for signal in dataset 1 # PP.H2.abf: Coloc posterior probability for signal in dataset 2 # PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2 # PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2 # idx1: Index of the SuSiE credible set for dataset 1 # idx2: Index of the SuSiE credible set for dataset 2 # cluster: cluster name # egene: eGene name # capeak: caPeak coordinates # p12min: Min prior p12 where the PP H4 > 0.5. Lower this value, more robust is the colocalization # trait: GWAS trait name # gwas_locus: GWAS locus name for the coloc test - a 250kb left and right flanking genomic window on this SNP was considered for testing coloc between all pairs of GWAS/QTL signals identified in this region # traitname: Expanded GWAS trait name # variable_type: GWAS type # source: Source of GWAS - either UKBB or other study 18. sn_muscle_2023-code-main.zip - Zip of GitHub repo for code. 19. supplementary_tables.xlsx: Supplementary tables from the manuscript. Information included in sheets: 1. "marker_genes": Marker genes known from literature used to annotate clusters 2. "n_nuclei": n pass-QC nuclei per modality-sample-cluster 3. "snrna_GO_enrichment": GO term enrichment: matrix of cluster vs top 2 GO terms 4. "qtl_scan_info":  e/caQTL scan info cluster: cluster ntested_eqtl: N genes tested for eQTL nsig_eqtl: N significant (5% FDR) eGenes n_pheno_pcs_eqtl: N phenotype PCs considered for eQTL ratio_eqtl: Ratio of N eGenes/N genes tested nsig_caqtl:  N peaks tested for caQTL ntested_caqtl: N significant (5% FDR) caPeaks n_pheno_pcs_caqtl: N phenotype PCs considered for caQTL ratio_caqtl: Ratio of N caPeaks/N peaks tested nsamples_eqtl: N samples for eQTL nsamples_caqtl: N samples for caQTL 5. "gwas_trait_list": GWAS trait info trait: GWAS trait ID traitname: GWAS trait description variable_type: GWAS type. case/control (cc), continuous_irnt=continuous inverse-normal transformed source: GWAS source doi: GWAS study DOI 6. "traits_in_ldsc_baseline" - list of annotations included in the baseline model for LDSC 7. "gwas_enrichment_in_peaks" GWAS enrichment in cluster peaks (S-LDSC) 8. "gwas_enrichment_in_qtl_peaks" GWAS enrichment in QTL peaks (fGWAS) # fGWAS results comparing GWAS enrichment in type 1 annotations CI_lower_ln, estimate_ln, CI_upper_ln: natural log of lower confidence interval, estimate, and upper confidence interval trait: trait id traitname: trait name annotation: annotation sig: 1 if CIs don't overlap 0, otherwise 0 9. t2d_gwas_caqtl_coloc and 10. t2d_gwas_eqtl_coloc: Summary of e,caQTL coloc with T2D GWAS in each cluster, along with target gene nominations. Columns: nsnps: Number of SNPs in the region gwas_hit: SNP with the highest bayes factor in the SuSiE GWAS credible set eqtl_hit: SNP with the highest bayes factor in the SuSiE eQTL credible set caqtl_hit: SNP with the highest bayes factor in the SuSiE caQTL credible set PP.H0.abf: Coloc posterior probability for no signal PP.H1.abf: Coloc posterior probability for signal in dataset 1 PP.H2.abf: Coloc posterior probability for signal in dataset 2 PP.H3.abf: Coloc posterior probability for different signal in datasets 1 and 2 PP.H4.abf: Coloc posterior probability for shared signal in datasets 1 and 2 idx1: Index of the SuSiE credible set for dataset 1 idx2: Index of the SuSiE credible set for dataset 2 cluster: cluster name egene: eGene name capeak: caPeak coordinates p12min: Min prior p12 where the PP H4 > 0.5. Lower this value, more robust is the colocalization trait: GWAS trait id diamante_gwas_locus: GWAS signal from the DIAMANTE 2018 study. Some signals that our SuSiE runs identified were not present in the original study in which case this column is NA traitname: Expanded GWAS trait name capeak_in_tss: caPeak in TSS + 1kb upstream region of a gene gene_target_standard_cicero: caPeak coaccessible with TSS peak of a gene considering nuclei from all samples for co-accessibility gene_target_allelic_cicero:  caPeak coaccessible with TSS peak of a gene considering nuclei from samples homozygous for the caSNP allele associated with increased accessibility gwashit_nominal_egene: gwas_hit nominally associated with these genes nominated in the columns capeak_in_tss, gene_target_standard_cicero, and  gene_target_allelic_cicero 11. MPRA results for the C2CD4A locus 12. MPRA primers Files genes_with_exon_counts_considered.txt Files (3.4 GB) Name Size caqtl_credible_sets.tar.gz md5:eb0a644d5b2f0e46075ae9e8bf4e7d1d 94.0 MB Download caqtl_permute.tar.gz md5:674c51f70d16b4f9a5d308a72e99397d 108.6 MB Download cellregmap.tar.gz md5:0d72d6774eb98d68954d2090f7ff7324 8.1 MB Download cicero_all.tar.gz md5:ee968e0a48b8664ad4786a100fa4d083 2.8 GB Download cicero_gene_tss.tar.gz md5:39598ceab7ae1d6ac1eb4e8590511b91 92.5 MB Download cit-mrs-summary.tsv md5:700495275d355469c80ebd2dd6d56f87 4.3 MB Download coloc-eqtl-caqtl.tsv md5:fa6499115110e0b34efcd573c92c5500 2.2 MB Download coloc-gwas-caqtl.tsv md5:1fa299f5f90259f28510319e40432225 3.6 MB Download coloc-gwas-eqtl.tsv md5:5086a43ae303e5f22a12588c1834ac29 325.9 kB Download consensus_to_narrowpeak_map.tar.gz md5:3a38761c147728f4f4296517542b36ea 49.8 MB Download eqtl_credible_sets.tar.gz md5:d92556bf8a102acb5fe174e501864a0b 10.4 MB Download eqtl_permute.tar.gz md5:ea8018269ea8e25ac5e7614688fe0556 5.6 MB Download genes_with_exon_counts_considered.txt md5:de8629011f16c22cea75a3d716156ebc 158.1 kB Preview Download mash.tar.gz md5:bd1fad91a904125c8b54373c64c16916 34.2 MB Download nucleus_sample_cluster_map.tsv md5:0a09ab42e163093c650a849b9afdd8a2 60.7 MB Download peaks.tar.gz md5:1caa1665d5c51fbe8c7b879b5d78b7ee 102.2 MB Download README.txt md5:b22a30a6b424c6df49a6d8f6e4f3e668 12.7 kB Preview Download sn_muscle_2023-code-main.zip md5:be6f66d136949b00585dad4e49469b42 21.0 MB Preview Download snrna-cell-type-specific-genes.tsv md5:10658e7bbc33fbc608870932b214df39 271.1 kB Download supplementary_tables.xlsx md5:5f162c5f1110f0996b01965197e844e0 840.4 kB Download Additional details Dates Submitted 2023-11-08 914 Views 1K Downloads Show more details All versions This version Views Total views 914 125 Downloads Total downloads 1,044 370 Data volume Total data volume 156.7 GB 39.1 GB More info on how stats are collected.... Versions External resources Indexed in OpenAIRE Communities Keywords and subjects Keywords eQTL caQTL snRNA snATAC GWAS Details DOI DOI Badge DOI 10.5281/zenodo.19670721 Markdown [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.19670721.svg)](https://doi.org/10.5281/zenodo.19670721) reStructuredText .. image:: https://zenodo.org/badge/DOI/10.5281/zenodo.19670721.svg :target: https://doi.org/10.5281/zenodo.19670721 HTML <a href="https://doi.org/10.5281/zenodo.19670721"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.19670721.svg" alt="DOI"></a> Image URL https://zenodo.org/badge/DOI/10.5281/zenodo.19670721.svg Target URL https://doi.org/10.5281/zenodo.19670721 Resource type Dataset Publisher Zenodo Rights License Creative Commons Attribution 4.0 International The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Read more Citation Export Technical metadata Created April 22, 2026 Modified April 22, 2026 Jump up About About Policies Infrastructure Principles Projects Roadmap Contact Blog Blog Support Help FAQ Developers REST API OAI-PMH Contribute GitHub Donate Funded by Powered by CERN Data Centre & InvenioRDM Status Privacy policy Cookie policy Terms of Use This site uses cookies. Find out more on how we use cookies Accept all cookies Accept only essential cookies

Record · ID 134054 · SHA-256 5a1a35f026f5bed2
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.