ConceptioArchivearXiv CS
arXiv CSopen access

Deep and Probabilistic Models for Gene Regulatory Network Inference

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Deep and Probabilistic Models for Gene Regulatory Network Inference

arXiv:2607.16053v1 [stat.ML] 17 Jul 2026

by

Claudia Skok Gibbs

A dissertation submitted in partial fulfillment of the reqirements for the degree of Doctor of Philosophy Center for Data Science New York University May, 2026

Dr. Kyunghyun Cho

Dr. Richard Bonneau

© Claudia Skok Gibbs all rights reserved, 2026

Dedication

For my parents, who always believed in me, and for my husband, who carried me through.

iii

Acknowledgements Before this journey began, I often saw myself through the lens of my struggles with mathematics, and I was not always sure that I belonged in science. Today, I understand that a scientist is not defined by innate talent, but by the determination to keep learning, to grapple with difficult questions, to immerse oneself in complexity, and to find joy in the unknown. It is thanks to my advisors, colleagues, friends, and family that I have come to understand this. I would first like to thank my advisors Kyunghyun Cho and Richard Bonneau, for their encouragement, kindness, and support over the last five years. Kyunghyun, it has been a pleasure to learn from you and to have been given the freedom to explore the ideas which have always excited me the most. These ideas have benefited enormously from your expertise, creativity, and insight. Rich, I will always be grateful to you for helping launch my scientific career and fostering my love of biology, which ultimately led me to pursue this Ph.D. I have learned so much from both of you, and I am deeply grateful for your mentorship. I would also like to thank Camille Alexander-Norrell, without whom I would never have been able to wrangle the busy schedules of Kyunghyun and Rich. Thank you for always taking care of me like one of your own, and for nourishing me with lunch, snacks, jokes, and funny gossip. Your love and kindness have carried me through so much of this journey. To my thesis committee, Carlos Fernandez-Granda, Romain Lopez, and Anirvan Sengupta, thank you for your time, insight, and thoughtful feedback on this dissertation. There are so many incredible scientists I met along this journey who have also become my

iv

closest friends. I would especially like to thank Dongmin for her enthusiasm in our adventures together. Whether it is a lunch date, coffee in the park, or running marathons together, you are always the best company for any activity. I would also like to thank Dan, Harsh, Meet, Lavender, Angie, Maggie, and Omar for their friendship, support, and the insight they brought to so many projects. I would also like to thank my early collaborators, who taught me so much useful biology at the beginning of this journey. Chris Jackson, thank you for mentoring me from the very first day and teaching me how to approach a scientific problem. Andreas Tjarnberg, thank you for teaching me how to navigate the tangled mess of biological data and transform it into something clear and beautiful. Neset Ozel, thank you for your collaboration which exposed me to the fascinating world of developmental biology. Outside of science, I have been surrounded by so many incredible friends who have supported me throughout this journey and have listened to my endless rants about biology, math, and academia. I would like to thank Athena, Kyra, Laurel, Sarah, Cristiana, Aaron, Michelle, Mabelle, Harry, Charlotte, Helene, Romain, Kerda, Randall, Anzar, Katie, Matt, Zoe, Mario, Milly, Jacopo, Irina, and Tanya, for all of the joy and laughter which sustained me throughout this journey. I am also fortunate to be supported by several workout coaches, who gave me a physical outlet from the mental strain of this thesis. Thank you to Dale Elston, Eric Salvador, Jennifer Spina, and Mike Keohane for reminding me through exercise that I am always capable. I come from a large family that has supported and loved me from the very first day. To my sisters, Thoma, Victoria, and Sophie, and my brothers, Benjamin, and Sebastian, I am never alone because of the five of you. Thank you for your love, kindness, and for always being there whenever I needed someone to lean on. To their spouses, who have also become like siblings to me, Stuart, Chris, Max, Emma, and Melissa, thank you for loving my favorite people, and for becoming such an important part of our family’s wonderful chaos. To my uncles, David, and v

Michael, thank you for our adventures, but most of all, for always encouraging me in my pursuit of mathematics and science. Your belief in me and your confidence in my future helped me believe in myself. In the first year of my Ph.D., I made the best decision of my life, which gave me four additional family members who have made all the difference in this journey. I would first like to thank my mother-in-law, whom we lost in 2022, and whose absence is felt with irreparable sadness. Evelyne, thank you for teaching me that true strength means showing up even on days when it feels hardest, and for showing me that love and kindness are the greatest ambitions of all. Thank you to my father-in-law, Habib, for your love, kindness, and patience. I have learned so much from you, both in language and in life, and have benefited so much from your generosity. Thank you to my brother- and sister-in-law, Allan and Carla, for bringing so much laughter and fun to our adventures together. I cannot wait for the many more to come. I am so incredibly lucky to have parents who are not only my best friends and biggest supporters, but also my greatest source of motivation and inspiration. To my parents, Jane and Paul, your love, support, and encouragement have shaped every part of my life. Thank you for believing in me so fully, for supporting me through every stage of this Ph.D., and for always reminding me that I was capable of more than I sometimes believed. I love you both more than I could ever adequately put into words. Finally, to my husband, Yanis, for whom words will never do justice, in English or in French, I will never find enough ways to thank you for what your partnership has brought to my life and to this Ph.D. You have supported every part of this journey, from helping me through homework assignments to listening patiently to every research idea, frustration, and breakthrough. You picked me up in the moments I doubted myself, reminded me of my strength when I could not see it, and motivated me to give each day my best. I am endlessly grateful not only for everything you have done, but for the love, steadiness, and joy you bring to my life every day. Thank you for making this dream of mine possible. vi

Abstract Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection, while evaluation is constrained by incomplete reference networks and point-estimate outputs that lack uncertainty. GRN reconstruction also depends on prior knowledge to constrain TF–gene interactions, yet available priors are often assay-dependent and difficult to transfer across species and less-characterized systems. In this thesis, we develop two complementary frameworks that address these limitations. In the first, PMF-GRN casts GRN inference as a probabilistic graphical model optimized by variational inference, enabling principled model selection and uncertainty-aware edge estimates. In the second, GLM-Prior addresses the prior bottleneck by fine-tuning the pretrained Nucleotide Transformer to predict TF–target gene interactions directly from nucleotide sequence, while generalizing across yeast, mouse, and human settings. Together, PMF-GRN and GLM-Prior motivate a dual-stage view of GRN reconstruction in which sequence-derived priors provide a transferable starting scaffold and probabilistic inference refines regulatory estimates with quantified uncertainty under incomplete evaluation resources.

vii

Contents Dedication

iii

Acknowledgements

iv

Abstract

vii

List of Figures

xii

List of Tables

xxii

1 Introduction

1

2 Background

7

2.1

The Central Dogma of Biology . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

8

2.2

Gene Regulatory Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

10

2.3

Data Modalities for Gene Regulatory Networks . . . . . . . . . . . . . . . . . . .

12

2.3.1

Microarray Gene Expression Profiling . . . . . . . . . . . . . . . . . . . .

13

2.3.2

Bulk RNA Sequencing . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

15

2.3.3

Single-Cell RNA Sequencing . . . . . . . . . . . . . . . . . . . . . . . . . .

16

2.3.4

ATAC Sequencing for Chromatin Accessibility . . . . . . . . . . . . . . .

18

2.3.5

ChIP Sequencing for Transcription Factor Binding . . . . . . . . . . . . .

19

2.3.6

Transcription Factor Binding Motifs . . . . . . . . . . . . . . . . . . . . .

21

viii

2.4

Machine Learning Foundations for Gene Regulatory Network Inference . . . . . .

22

2.4.1

Problem Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

23

2.4.2

Challenges in Gene Regulatory Network Inference . . . . . . . . . . . . .

25

2.4.3

Gene Regulatory Network Inference Algorithms . . . . . . . . . . . . . .

30

2.4.4

Machine Learning and Statistical Foundations . . . . . . . . . . . . . . . .

40

2.4.5

Evaluating GRN Inference . . . . . . . . . . . . . . . . . . . . . . . . . . .

54

3 Probabilistic Matrix Factorization for Gene Regulatory Network Inference

59

3.1

Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

59

3.2

Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

60

3.3

Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

63

3.3.1

The PMF-GRN Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

63

3.3.2

Advantages of PMF-GRN . . . . . . . . . . . . . . . . . . . . . . . . . . .

67

3.3.3

PMF-GRN Recovers True Interactions in Simple Eukaryotes . . . . . . . .

69

3.3.4

PMF-GRN Provides Well-Calibrated Uncertainty Estimates . . . . . . . . .

75

3.3.5

PMF-GRN Integrates Single Cell Multi-Omic Data for GRN and TFA Inference in Human PBMCs . . . . . . . . . . . . . . . . . . . . . . . . . . .

76

Evaluating PMF-GRN with BEELINE Synthetic Data . . . . . . . . . . . .

81

Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

82

3.4.1

Model Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

82

3.4.2

Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

84

3.4.3

Computing Summary Statistics for the Posterior . . . . . . . . . . . . . .

86

3.4.4

Calculating AUPRC . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

86

3.4.5

Evaluating Calibration of Posterior Uncertainty . . . . . . . . . . . . . . .

87

3.4.6

Inference and Evaluation on Multiple Observations of W . . . . . . . . . .

87

3.4.7

Measuring the Impact of Prior Hyperparameters . . . . . . . . . . . . . .

87

3.3.6 3.4

ix

3.4.8

Cross-Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

88

3.4.9

Intersection over Union . . . . . . . . . . . . . . . . . . . . . . . . . . . .

88

3.4.10 Downsampling Expression . . . . . . . . . . . . . . . . . . . . . . . . . . .

89

3.4.11 Exploring the Effect of Cross-Validation Ratios on Hyperparamater Selection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

89

3.4.12 Datasets and Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . .

89

3.5

Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

91

3.6

Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

94

3.7

Retrospective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

95

3.8

Supplementary Material for Chapter 3 . . . . . . . . . . . . . . . . . . . . . . . . .

98

3.8.1

PMF-GRN Recovers True Interactions in Prokaryotes as Evaluated by CrossValidation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

98

3.8.2

Peripheral Blood Mononuclear Cells . . . . . . . . . . . . . . . . . . . . . 100

3.8.3

Inferelator, Scenic, and CellOracle Networks . . . . . . . . . . . . . . . . . 105

3.8.4

TF Target Gene Connectivity Matrix Generation . . . . . . . . . . . . . . 109

4 A Genomic Language Model Prior for Gene Regulatory Network Inference

110

4.1

Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110

4.2

Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111

4.3

Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 115 4.3.1

Dual-stage training pipeline overview . . . . . . . . . . . . . . . . . . . . 115

4.3.2

GLM-Prior’s generalization scales with data composition across species and cell types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 120

4.3.3

GLM-Prior supports successful species-transfer learning and multi-species training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 124

x

4.3.4

Integrating GLM-Prior into PMF-GRN provides contextual GRN inference and edge refinement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 128

4.3.5

Comparing prior construction strategies across yeast, mouse, and human cell lines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 130

4.3.6

Paired prior-GRN strategies reveal prior-dependent gains from expressionbased inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133

4.3.7

Cross-method comparison disentangles prior and GRN inference contributions to performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . 137

4.4

4.5

Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142 4.4.1

The GLM-Prior model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 142

4.4.2

GRN inference with PMF-GRN . . . . . . . . . . . . . . . . . . . . . . . . 148

Supplementary Material for Chapter 4 . . . . . . . . . . . . . . . . . . . . . . . . . 153 4.5.1

Single-Species Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 153

4.5.2

Transfer-Learning Experiments . . . . . . . . . . . . . . . . . . . . . . . . 160

4.5.3

Multi-Species Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . 161

4.5.4

Inferelator-Prior and CellOracle baseGRN . . . . . . . . . . . . . . . . . . 161

4.5.5

GRN Inference with PMF-GRN, the Inferelator, and CellOracle . . . . . . . 162

4.5.6

Cross-Method Comparison . . . . . . . . . . . . . . . . . . . . . . . . . . 163

5 Conclusion

166

Bibliography

170

xi

List of Figures 2.1

The central dogma of biology, illustrating the flow of genetic information from DNA to RNA via transcription and from RNA to protein via translation. . . . . . .

2.2

9

Schematic linking the central dogma to gene regulatory networks. A gene is transcribed into mRNA and translated into protein; when this protein is a transcription factor, it binds to DNA to regulate the transcription of a target gene, initiating a subsequent round of transcription and translation. This repeated feedback process is summarized as a directed graph, where nodes represent genes (or TFs) and directed edges represent regulatory influence from TFs to their target genes. . . .

xii

11

2.3

Schematic overview of data modalities commonly used to reconstruct gene regulatory networks. (A) Microarray workflow: labeled cDNA hybridizes to probes on a fixed array to measure transcript abundance via fluorescence intensity. (B) Bulk RNA-seq: population-level transcriptomes are sequenced to quantify gene expression across conditions or samples. (C) Single-cell RNA-seq: RNA from individual cells is barcoded, sequenced, and aggregated into a cell-by-gene count matrix. (D) ATAC-seq: a transposase preferentially inserts adapters into accessible chromatin, enabling the identification of open regulatory regions. (E) ChIP-seq: immunoprecipitation of DNA-bound TFs allows direct measurement of TF binding sites genome-wide. (F) Motif databases: position weight matrices (PWMs) derived from experimental binding data represent sequence-specific TF binding preferences and are used to scan genomes for candidate binding sites. . . . . . . .

2.4

14

Schematic overview of the input and output of GRN inference algorithms. The primary input is a gene expression matrix, and optional secondary input is a prior knowledge matrix. The selected GRN inference algorithm will produce a directed acyclic graph describing predicted releationships between TFs and their target genes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

2.5

23

Snapshot sampling and temporal mismatch between TF transcripts and targetgene expression. Schematic time course illustrating how TF mRNA abundance can be offset in time from the regulatory events that drive target gene transcription. At 𝑡 0 , TF mRNA is highly expressed and then declines over subsequent time points, while TF protein accumulates and becomes positioned for regulation. Following an activating signal at 𝑡 2 , the TF binds DNA at 𝑡 3 , and the downstream target gene exhibits increased transcript abundance only later at 𝑡 4 . . . . . . . . .

xiii

26

2.6

Key challenges in gene regulatory network inference. (A) A central challenge in GRN reconstruction is determining which transcription factors (TFs) regulate which target genes. (B) Regulation is often combinatorial, with multiple TFs acting cooperatively or competitively to control the same gene, while individual TFs may regulate many targets. (C) Regulatory interactions are also cell-type and context-specific, such that the active GRN can vary across cellular states or conditions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

2.7

28

Typical pipeline for constructing prior-knowledge matrix for downstream GRN inference. (A) Open chromatin regions are identified from ATAC-seq data to define candidate regulatory DNA accessible to transcription factor (TF) binding. (B) These accessible regions are scanned for known TF binding motifs to identify potential TF occupancy. (C) Motif-containing accessible regions are assigned to nearby genes, yielding candidate TF-target regulatory interactions. (D) Genomewide interactions are aggregated into a prior-knowledge matrix, in which rows correspond to target genes, columns correspond to TFs, and entries indicate putative regulatory relationships. This prior matrix can then be incorporated into GRN inference methods to constrain the search space toward biologically plausible edges. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

xiv

38

3.1

(A) PMF-GRN graphical model overview. Input single-cell gene expression 𝑊 is decomposed into several latent factors. Information obtained from chromatin accessibility data or genomics databases is incorporated into the prior distribution for 𝐴. (B) Input experimental data for PMF-GRN includes single-cell RNAseq gene expression data. Prior-known TF-target gene interactions can be obtained using chromatin accessibility in parallel with known TF motifs, or through databases or literature derived interactions.(C) Hyperparameter selection process is performed for optimal model selection. The provided prior-known network is split into a train and validation dataset. 80% of the prior-known information is used to infer a GRN, while the remaining 20% is used for validation by computing AUPRC. This process is repeated multiple times, using different hyperparameter configurations in order to determine the optimal hyperparameters for the GRN inference task at hand. Finally, using the optimal hyperparameters, a final network is inferred using the full prior and evaluated using an independent gold standard. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

xv

65

3.2

GRN inference in S. cerevisiae. (A) Consensus Network AUPR with a normal priorknowledge matrix (N): PMF-GRN (red) performance compared to Inferelator algorithms (AMuSR in yellow, BBSR in orange, StARS in green), SCENIC (blue), and CellOracle (purple). Dashed line represents the baseline if expression data is combined. Negative controls: no prior information (NP - black) and shuffled prior information (S - gray). (B) 5-Fold Cross-Validation Baseline: Each dot with low opacity represents one of the five experiments. Colored dots and lines depict the mean AUPR ± standard deviation for each GRN inference method. (C) GRNs inferred with increasing amounts of noise added to the prior. (D) Calibration results on S.cerevisiae (GSE125162 only) dataset. Posterior means are cumulatively placed in bins based on their posterior variances. AUPRC for each of these bins is computed against the gold standard (see Methods section for details) . . . . . .

3.3

71

(A) GRNs inferred by downsampling S. cerevisiae expression data. (B) Hyperparameter search performed on 4 different ratios of cross-validation. Dots represent validation AUPRC from hyperparameter search during cross-validation, triangle represents AUPRC from a GRN learned using the most optimal hyperparameters for each ratio. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

3.4

74

GRN and TFA inference in PBMC. (A) UMAP projection of predicted TFA for each annotated PBMC cell type. (B) Predicted IRF2 TFA demonstrates high activity in NK and CD8 T cells. (C) Heat-map dot-plot depicting TFA of selected immune TFs across annotated PBMC cell types. (D) GRN between IRF TFs and their targets. Pink edges indicate literature support for interaction. (E) Heat-map dot-plot indicates ten most highly active TFs for each PBMC cell type. (F) Violin plot demonstrates corresponding distribution of TFA profiles for ten most highly active TFs. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

xvi

77

3.5

PMF-GRN performance on BEELINE synthetic GRN data (A) PMF-GRN inference performance with half of the ground truth provided as prior network information and the remaining half provided as a gold standard for evaluation. Dashed lines are the expected baseline of a random predictor. (B) AUPRC ratio over the baseline random predictor for PMF-GRN in comparison to each of the GRN inference methods used in the original BEELINE benchmark. . . . . . . . . . . . . . . . . .

3.1

81

Results for GRNs learned in B. subtilis datasets B1 (GSE27219) and B2 (GSE67023) comparing "No Normalization" to "Min-Max Scaling". Colored dots represent the normal (N) GRN, with the line indicating the mean of the cross-validation experiments ± standard deviation. Negative controls are demonstrated by black dots for GRNs inferred with No Prior (NP) and grey dots for Shuffled Prior (S). . . . . .

3.2

99

PBMC GRN graphs for the family of TFs belonging to (A) SMAD, (B) STAT, (C) GATA, and (D) EGR. TFs for each graph are represented by orange nodes, while target genes are blue nodes. Pink regulatory edges represent interactions with supporting literature. Color scale for blue regulatory edges are scaled from light blue (less out-degree regulation) to darker blue (more out-degree regulation). . . . 101

3.3

UMAP of predicted PBMC TFA. Each UMAP highlights the specific location of activity for each TF considered from the immune PBMC TFs. The final UMAP serves as a reference to TFA annotated by cell-type. . . . . . . . . . . . . . . . . . 102

xvii

4.1

Schematic of the dual-stage training pipeline. (A) GLM-Prior is trained on concatenated TF binding motifs and gene sequences labeled as interacting or noninteracting. Input data is split into training and validation sets, where the training set is downsampled and placed into even class batches using a custom DataLoader. These input sequences are embedded and encoded by the transformer architecture, and passed through a classification head to predict the probability of a regulatory interaction. After training, inference is performed on all TF-gene pairs to generate a prior knowledge matrix. (B) Inferred prior knowledge is passed to PMF-GRN to provide a structural constraint to the probabilistic graphical model during inference. PMF-GRN infers transcription factor activity and a GRN for the input cell-line dataset. The resulting GRN is then evaluated using AUPRC and uncertainty calibration with an independent gold standard. . . . . . . . . . . . . . 117

4.2

GLM-Prior model performance across yeast, hESC, and HepG2 cell lines. (A) One-epoch hyperparameter sweep over class-weights and negative-class downsampling rates, evaluated by F1-score. (B) Validation metrics during 10 epoch training using the best hyperparameter configuration. (C) Test AUPRC on heldout gold standards (with chance performance shown for reference in gray). (D) Train and test split composition (number of genes and TFs). (E) Class composition (number of positive and negative labels) in train and test sets. . . . . . . . . . . . 121

4.3

GLM-Prior model performance across mESC, mDC, and mHSC cell lines. (A) One-epoch hyperparameter sweep over class-weights and negative-class downsampling rates, evaluated by F1-score. (B) Validation metrics during 10 epoch training using the best hyperparameter configuration. (C) Test AUPRC on heldout gold standards (with chance performance shown for reference in gray). (D) Train and test split composition (number of genes and TFs). (E) Class composition (number of positive and negative labels) in train and test sets. . . . . . . . . . . . 122 xviii

4.4

Transfer learning and multi-species training with GLM-Prior. (A) Schematic and AUPRC results of the species-transfer learning setup, showing three configurations: (𝑖) a model trained jointly on human and mouse data evaluated on yeast, (𝑖𝑖) a mouse-only model evaluated on human cell lines, and (𝑖𝑖𝑖) a human-only model evaluated on mouse cell lines. (B) Schematic and AUPRC results of the multi-species model trained sequentially on human, mouse and yeast data, evaluated across six cell lines. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125

4.5

Transfer learning and multi-species training with GLM-Prior continued. (A) Comparison of training composition (number of genes and TFs) for single-species, species-transfer, and multi-species models. (B) Class composition (positive and negative labels) in the corresponding training sets. (C) Normalized AUPRC across single-species, species-transfer, and multi-species models for each cell line. . . . . 126

4.6

Integration of GLM-Prior with PMF-GRN for downstream GRN inference. (A) AUPRC comparison of GLM-Prior alone versus integrated with PMF-GRN across six cell line contexts. (B) Edge flip analysis quantifying changes made to the prior knowledge interaction matrix during GRN inference. Blue bars indicate edges gained and red bars indicate edges removed by GRN inference. . . . . . . . . . . . 129

4.7

Comparison of prior construction strategies across species. (A) AUPRC of GLMPrior, Inferelator-Prior and CellOracle Prior across six cell lines, with chance performance (gray) provided for each context. (B) Normalized AUPRC values with the best-performing prior (green box) highlighted in each cell line. . . . . . . . . . 132

xix

4.8

Paired prior-GRN strategies reveal prior dependent gains from expression-based GRN inference across six cell lines. (A) AUPRC of each prior (GLM-Prior, InferelatorPrior, and CellOracle’s prior) and its corresponding GRN (PMF-GRN, Inferelator, CellOracle) across six cell lines, with gray dashed lines indicating the chance level in each species. (B) Change in AUPRC from prior to GRN (ΔAUPRC) for each paired method, quantifying the added value of expression-based inference. . . . . 134

4.9

Paired prior-GRN strategies reveal prior dependent gains from expression-based GRN inference across six cell lines continued. (A) Normalized AUPRC performance for priors and GRNs in each cell line. Green boxes highlight the best prior and blue boxes highlight the best GRN per cell line. (B) Edge flips between prior and GRN for each method demonstrating how GRN inference methods modulate their priors. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 136

4.10 Cross-method comparison disentangles prior and GRN inference contributions. (A) AUPRC for all nine combinations of prior and GRN inference methods across six cell lines. Gray dashed lines indicate chance performance, while light blue, light green, and light pink dashed lines represent baseline performance of GLMPrior, Inferelator-Prior, and CellOracle prior, respectively, across each cell line. . . 138 4.11 Cross-method comparison disentangles prior and GRN inference contributions continued. (A) Normalized AUPRC of all nine combinations of prior and GRN for directly comparing across method and cell line. Blue boxes indicate highest GRN performance per cell line. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139

xx

4.12 Sequence-homology analysis of yeast gene and TF splits using hashFrag. (A) Schematic of homology-aware partitioning, illustrating how highly similar sequences are identified and separated to reduce cross-split homology between training and test sets. (B) Histogram of per-test-gene maximum BLAST alignment score to any training gene, shown before (blue) and after (orange) hashFrag pruning of the training gene set. Blue dashed line indicates the similarity threshold used to define leakage. (C) Leakage curve for genes showing, for each similarity threshold 𝑡, the fraction of test genes whose maximum BLAST score to any training gene is ≥ 𝑡. Blue dashed line marks the leakage threshold. (D) AUPRC evaluation of GLM-Prior trained and tested on the hashFrag-defined gene partitions, with the gray dot indicating chance performance. (E) Schematic of homologyaware partitioning for TF sequences, analogous to panel A. (F) Histogram of pertest-TF maximum BLAST alignment score to any training TF, shown before (blue) and after (orange) hashFrag pruning of the training TF set. Blue dashed line indicates the similarity threshold. (G) Leakage curve for TFs showing, for each similarity threshold 𝑡, the fraction of test TFs whose maximum BLAST score to any training TF is ≥ 𝑡. Blue dashed line marks the leakage threshold. (H) AUPRC evaluation of GLM-Prior trained and tested on the hashFrag-defined TF partitions, with the gray dot indicating chance performance. . . . . . . . . . . . . . . . . . . 156 4.13 Performance Metrics. Additional performance metrics for GLM-Prior and PMFGRN across six cell lines. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164

xxi

List of Tables 3.1

AUPRCs achieved by PMF-GRN on the B. subtilis B1 dataset. Results are reported as the mean AUPRC across five cross-validation splits ± standard deviation. . . .

3.2

99

AUPRCs achieved by PMF-GRN on the B. subtilis B2 dataset. Results are reported as the mean AUPRC across five cross-validation splits ± standard deviation. . . . 100

3.3

Literature supported interactions for PBMC GRNs for STAT, GATA, and IRF immune TF families. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 103

3.4

Literature supported interactions for PBMC GRNs for SMAD and EGR immune TF families. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104

3.5

AUPRCs achieved by PMF-GRN, the Inferelator algorithms (AMuSR, BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets. . . . . . . . . . . . . . . . 105

3.6

AUPRCs achieved by PMF-GRN, the Inferelator algorithms (BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets using the gold standard for 5-fold cross validation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106

3.7

AUPRCs achieved by PMF-GRN, the Inferelator algorithms (BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets using increasing amounts of noise added to the prior-knowledge data. . . . . . . . . . . . . . . . . . . . . . . . . . . 106

3.8

Intersection over Union (IoU) scores achieved by PMF-GRN, the Inferelator algorithms (AMuSR, BBSR, and StARS), Scenic and CellOracle for GRNs learned on individual S. cerevisiae datasets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 xxii

3.9

AUPRCs achieved by PMF-GRN across 4 different downsample sizes (80%, 60%, 40%, and 20%), across 5 samples for each downsample size. . . . . . . . . . . . . . 107

3.10 AUPRCs achieved by PMF-GRN across 4 different cross-validation splits. 5 hyperparameters searches were performed for each cross-validation split. Full GRN was inferred using the hyperparameters for the best overall AUPRC per crossvalidation split. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 107 3.11 GRN inference method comparison table. Table includes input data, output, methodology and pipeline organization for each of the six GRN inference methods discussed in this work. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 108 4.1

Validation performance of the GLM-Prior model trained on yeast. Evaluation was conducted on a held-out validation set that assess positive and negative class contributions during training (top) and AUPRC against an independent gold standard for the final inferred prior-knowledge matrix (bottom). . . . . . . . . . . . . . . . 154

4.2

Validation performance of the GLM-Prior model trained on human, evaluated in hESC and HepG2. Evaluation is conducted on held-out sets using validation metrics that consider the contribution of the positive and negative classes and AUPRC against ChIP-seq-derived reference networks after inference of the priorknowledge matrix. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157

4.3

Validation performance of the GLM-Prior model trained on mouse, evaluated in mESC, mDC, and mHSC. Evaluation is conducted on held-out sets using validation metrics that consider the contribution of the positive and negative classes and AUPRC against ChIP-seq-derived reference networks after inference of the prior-knowledge matrix. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 159

xxiii

4.4

AUPRC scores for cross-species transfer learning. Each row indicates the species used to train the GLM-Prior model and the target species on which GRN inference was performed. Evaluation was conducted against species-specific ChIP-seq or curated reference datasets. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 160

4.5

AUPRC scores for multi-species model inference in human, mouse, and yeast. GRNs were inferred using a unified model trained jointly on all three species, and evaluated against cell-line specific reference networks. . . . . . . . . . . . . . . . 161

4.6

AUPRC of GLM-Prior, Inferelator-Prior and CellOracle baseGRN across six cell lines. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 162

4.7

AUPRC of three GRN inference methods (PMF-GRN, Inferelator 3.0, and CellOracle) across six cell lines. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163

4.8

Cross-method AUPRC comparison for six cell lines, combining three prior-knowledge constructions (CellOracle prior, Inferelator-Prior, GLM-Prior) with three GRN inference methods (PMF-GRN, Inferelator, CellOracle). . . . . . . . . . . . . . . . . 165

xxiv

1 | Introduction Gene regulatory networks (GRNs) provide a structured representation of transcriptional control by mapping regulatory relationships between transcription factors (TFs) and their target genes [Karlebach and Shamir 2008; He and Tan 2016]. These networks offer a useful lens for studying how coordinated gene expression programs arise in development [Özel et al. 2022], adapt to environmental signals [Jackson et al. 2020], and become disrupted in disease [Emad and Sinha 2021; Unger Avila et al. 2024]. A central challenge, however, is that GRNs cannot be directly measured at genome scale with standard sequencing assays. Instead, regulatory structure must be inferred from high-throughput genomic data that provide indirect, incomplete, and often noisy views of the underlying regulatory processes [Badia-i Mompel et al. 2023]. This thesis is motivated by the observation that the difficulty of GRN inference is driven as much by the limitations of available data as by the modeling choices made downstream. First, most experimental assays provide only indirect views of regulation. Expression measurements are noisy and high-dimensional, particularly in single-cell settings where sparsity, dropout, and sampling variability are substantial [Nesari et al. 2026]. Additionally, transcriptomic datasets provide precise snapshots taken at unknown points along the regulatory cascade [Gorin et al. 2022], causing contemporaneous TF-target gene correlations to be weak or misleading when regulatory effects are delayed. Regulation is also combinatorial, with genes integrating inputs from multiple TFs, and each TF influencing many potential target genes, expanding the space of plausible explanations for any

1

observed expressed pattern [Dubois-Chevalier et al. 2018]. Complementary modalities often help constrain this search space by providing mechanistic context beyond expression alone [Miraldi et al. 2019; Skok Gibbs et al. 2022]. For example, chromatin accessibility assays such as ATACseq can narrow attention to genomic regions that are permissive to regulation, and TF binding assays such as ChIP-seq can identify loci occupied by a specific TF under a particular condition. At the same time, these measurements remain context dependent and incomplete. Accessibility and occupancy capture regulatory potential and binding, but they do not themselves establish functional influence on transcription or the sign and magnitude of regulatory effects [Slattery et al. 2014; Cusanovich et al. 2014]. Limitations in available data for GRN inference also carries through to evaluation. Reference networks and gold standards are typically assembled from incomplete databases and contextspecific experimental evidence, capturing only a subset of true regulation in any given system [Hegde et al. 2025]. As a result, many true edges are missing from these reference networks, with unlabeled TF-gene pairs usually reflecting a lack of evidence or coverage, rather than providing evidence that an interaction does not occur [Kernfeld et al. 2024]. This partial observability of both regulation and evaluation presents GRN inference as an ill-posed inverse problem [Schäfer and Strimmer 2005], where many distinct regulatory explanations can be consistent with the same observed data. Beyond data limitations, existing GRN inference methods introduce practical constraints that further limit what can be inferred reliably in practice. Many existing algorithms tightly couple their statistical model to a specific inference procedure, often embedding assumptions that are tailored to a particular data modality or sampling regime [Hu et al. 2020]. As genomic technologies have shifted from microarrays to bulk RNA-seq and then to single-cell measurements, GRN inference methods have frequently required redesign to accommodate new measurement characteristics, rather than providing a stable modeling scaffold capable of absorbing new assumptions. Even within a fixed data regime, additional methodological choices further shape what can 2

be inferred. Model selection is often treated heuristically, with many approaches committing to a single algorithmic form or fixed set of hyperparameters without systematically comparing alternative modeling choices, even though different assumptions can yield qualitatively different networks from the same data. Finally, GRNs are commonly reported as point estimates without an accompanying measure of confidence, despite the fact that evaluation typically relies on incomplete and context-dependent references. Without uncertainty estimates, it becomes difficult to separate robust interactions from weakly identified edges, and to prioritize predictions in settings where labels are sparse, missing, or context-mismatched. In addition to inference and model selection limitations, the accuracy of GRN reconstruction is often determined by the availability and quality of prior knowledge. In realistic singlecell settings, expression data alone is insufficient to anchor regulatory structure, and meaningful results typically require incorporating prior evidence that breaks symmetries and restricts attention to biologically plausible edges. In current pipelines, priors are most commonly constructed by combining experimental and computational signals, for example by linking TF motifs to genes through regions of open chromatin, or by aggregating TF binding evidence where it exists [Skok Gibbs et al. 2022; Kamimoto et al. 2023; Van de Sande et al. 2020]. These strategies are powerful, but they also remain limited as they tend to emphasize promoter-proximal interactions and depend on cell-type specific assays that are unavailable in many contexts. In model organisms, curated regulatory databases provide an alternative source of prior edges, but these remain incomplete even across TFs and conditions. Together, these methodological limitations motivate the need for GRN inference frameworks that decouple model specification from the inference algorithm, support principled model selection across modeling assumptions and hyperparameters, and provide uncertainty-aware network estimates that can be interpreted and prioritized even when reference labels are incomplete or unavailable. These limitations also motivate the need for priors that do not depend on cell type-specific experimental assays, can capture regulatory logic beyond promoter-proximal in3

teractions, and support the transfer of regulatory information from well-studied organisms to less-characterized systems, so that downstream inference remains well constrained even when direct prior evidence is sparse or unavailable. The remainder of this thesis is organized to first establish the biological and methodological context for GRN inference, and then to develop two complementary approaches that address the central limitations outlined above. In Chapter 2, we first situate GRNs within the central dogma by emphasizing the feedback loop in which TF proteins return to the genome to shape transcription, motivating a network-based representation of regulation. We then survey the experimental modalities that make GRN reconstruction possible in practice, including transcriptomic assays and chromatin-based assays that provide observable and mechanistic context. Finally, we introduce the machine learning framing used throughout the remainder of this thesis, describing inference over latent regulatory structure under strong data limitations, the role of priors in anchoring interpretability and identifiability, and the consequence of incomplete reference networks for evaluation. Together, this background chapter provides the conceptual toolkit required to understand why GRN inference is challenging, what information different data modalities contribute, and how modeling assumptions determine what can be learned from modern genomics datasets. Chapter 3 then focuses on the inference problem itself and introduces a probabilistic framework designed to address methodological limitations that arise even when appropriate data and priors are available [Skok Gibbs et al. 2024]. The chapter develops a generative, latent-variable view of GRN inference in which TF activities and TF-gene influences are treated as unobserved quantities that must be inferred from noisy expression measurements. By separating model specification from the inference procedure, this framework makes it possible to modify distributional assumptions and incorporate additional biological structure without redesigning the optimization machinery used for inference. The probabilistic formulation also enables principled model selection by comparing alternative modeling assumptions and hyperparameters in a systematic 4

way, rather than relying on a single default algorithmic configuration. Finally, posterior uncertainty provides an explicit notion of confidence for inferred TF-gene interactions, supporting uncertainty-aware prioritization in settings where evaluation resources are incomplete or unavailable. Chapter 4 shifts the emphasis from inference to prior construction, motivated by the observation that the quality and availability of prior knowledge often determines whether GRN inference is usable in practice. This chapter introduces a sequence-based approach for constructing TF-gene prior networks that does not rely on cell type-specific experimental assays and can generalize beyond the limited contexts where curated priors exist [Skok Gibbs et al. 2025]. Instead of deriving priors primarily from promoter-proximal accessibility and motif heuristics, the approach learns regulatory sequence features directly from data using a transformer-based genomic language model, enabling prior construction at genome scale. A central goal of this chapter is to characterize how such priors generalize across settings, including single-species training, transfer learning between organisms, and multi-species training. Further, we demonstrate how stronger, more transferable priors improve downstream GRN reconstruction in complex mammalian systems where experimental priors are often noisy, incomplete, or unavailable. Together, Chapters 3 and 4 support a dual-stage perspective on GRN reconstruction in which prior construction and expression-based inference play distinct, complementary roles. Prior knowledge provides TF-specific structure that narrows the space of plausible regulatory explanations and anchors interpretability, while expression-based inference refines this scaffold in a context-dependent manner and quantifies uncertainty in the resulting network. This division of labor directly addresses the challenges outlined above, where noisy and incomplete measurements, imperfect and context-dependent evaluation resources, and methodological constraints have historically limited robustness and interpretability of GRN inference. The central argument developed across this thesis is that progress in GRN reconstruction depends jointly on inference frameworks that are flexible, principled, and uncertainty-aware, as well as on priors that can 5

capture regulatory logic beyond promoter-proximal signals and can transfer across biological systems where direct regulatory evidence is sparse or unavailable.

6

2 | Background This chapter lays the groundwork for the chapters that follow by introducing essential context on topics related to gene regulatory network inference. In Section 2.1, we describe the central dogma of biology and its natural connection to gene regulatory networks, highlighting how transcriptional regulation introduces feedback that links gene expression to protein function. This perspective motivates a network-based view of regulation which is formalized in Section 2.2 through the definition of gene regulatory networks. Following this, in Section 2.3, we introduce the major experimental data modalities that provide indirect or partial evidence of regulatory activity and form the empirical basis for gene regulatory network inference. In Section 2.4, we then define the GRN inference problem from a machine learning perspective, and outline key challenges that arise from both the properties of modern genomics datasets and the complexity of transcriptional regulation. Then, we review the major families of GRN inference algorithms and introduce the statistical foundations that will be used throughout this thesis, including latent variable modeling, probabilistic generative frameworks with variational inference, and sequence-based models for construction informative prior knowledge. Finally, we describe evaluation metrics used throughout this thesis to evaluate the predictive accuracy of inferred GRNs using available reference gold standards.

7

2.1

The Central Dogma of Biology

The central dogma of biology describes the flow of information within living systems, where genetic information encoded in the DNA is transcribed into RNA, and then translated into protein (Figure 2.1) [Crick 1970]. DNA is a complex double helix molecule consisting of four chemical bases, adenine (A), thymine (T), cytosine (C), and guanine (G), [Klug 2004], and contains the genetic instructions required for the development, function, and growth of all known organisms. Within this sequence, genes define specific, heritable regions of the genome that are transcribed to produce RNA molecules [Chaffey 2003]. RNA shares many structural similarities with DNA, but differs by using uracil (U) in place of thymine (T), giving it a nucleotide alphabet of A, U, C, and G. RNA is synthesized through a process called transcription, in which the DNA sequence of a gene is read by the enzyme RNA polymerase to produce a complementary RNA strand. To initiate transcription, the DNA double helix locally unwinds, allowing RNA polymerase to bind to the promoter (or start of a gene), and begin RNA synthesis [Borukhov and Nudler 2008]. As RNA polymerase moves along the DNA template, it assembles the RNA transcript by sequentially adding RNA bases that are complementary to the DNA bases [Abbondanzieri et al. 2005]. The resulting RNA molecule may function directly as a non-coding RNA, or more commonly, as messanger RNA (mRNA), which carries the gene’s information forward and acts as a template to build proteins [Cooper and Adams 2022]. Proteins are functional molecules responsible for executing most cellular processes. Unlike DNA and RNA, which are composed of four nucleotide bases, proteins are made up of chains of 20 different amino acids, giving rise to a more complex biochemical alphabet [Chaffey 2003]. Proteins synthesis occurs through a process called translation, in which the nucleotide sequence of an mRNA molecule is read in groups of three bases, known as codons [Cooper and Adams 2022]. Each codon corresponds to a specific amino acid, which is added sequentially to a growing 8

polypeptide chain. Once synthesized, this chain folds into a three-dimensional structure, with this structure largely determining the protein’s biological function [Sanvictores and Farci 2020].

Figure 2.1: The central dogma of biology, illustrating the flow of genetic information from DNA to RNA via transcription and from RNA to protein via translation.

Among the many types of proteins produced in the cell, transcription factors (TFs) play a central role in controlling gene regulation. TFs are DNA-binding proteins that recognize specific sequence motifs and modulate transcription by influencing the recruitment and activity of the transcriptional machinery, including RNA polymerase II, at gene regulatory regions [Latchman 1997; Lambert et al. 2018]. In eukaryotes, TFs act at gene promoters as well as at distal regulatory elements such as enhancers. TF-mediated regulation is typically combinatorial; a single gene integrates input from multiple TFs, and the regulatory outcome depends on the combination of TF partners, cofactors, and the broader cellular context [Mitsis et al. 2020]. This complexity allows cells to precisely control when transcription is initiated and to what extent each gene is expressed. In this way, the central dogma of biology, describing the flow of information from DNA to 9

RNA to protein via transcription and translation, serves as the backbone for gene expression, while gene regulation determines how this flow is modulated across space, time, and condition. Importantly, regulation introduces feedback that closes the loop; many proteins produced during translation, including TFs, return to bind DNA, and influence subsequent rounds of transcription. Because each TF can regulate multiple genes, and each gene can be regulated by multiple TFs, this many-to-many, combinatorial architecture gives rise to complex, distributed control systems. Feedback interactions between TFs and their target genes generate coordinated gene expression programs across the genome. This structure motivates a network-based perspective of regulation, formalized as gene regulatory networks (GRNs), which is the focus of the next section.

2.2

Gene Regulatory Networks

The feedback described in the previous section motivates a simple but powerful abstraction for transcriptional control. When a gene is transcribed and translated, the resulting protein can act back on the genome to influence the transcription of other genes. In particular, TF proteins bind specific regions of DNA and modulate the transcription of their target genes, thereby initiating new rounds of RNA production and protein synthesis. Gene regulatory networks (GRNs) encapsulate this process by representing the relationships between TFs and their target genes within a biological system (illustrated in Figure 2.2) [Delgado and Gómez-Vela 2019; Badia-i Mompel et al. 2023]. GRNs are commonly represented as directed graphs [Schlitt and Brazma 2007]. Formally, a GRN can be defined as a graph 𝐺 = (𝑉 , 𝐸), where 𝑉 = {1, . . . , 𝑝} is the set of nodes corresponding to genes (or gene products, such as TFs), and 𝐸 ⊆ 𝑉 × 𝑉 is the set of directed edges representing regulatory influences. For each node 𝑗 ∈ 𝑉 , the parent set 𝑃𝑎( 𝑗) = {𝑖 ∈ 𝑉 : (𝑖 → 𝑗) ∈ 𝐸} denotes the regulators of gene 𝑗. In this representation, an edge from TF 𝐴 to gene 𝐵 indicates that the activity of TF 𝐴 influences the expression of gene 𝐵. Importantly, regulation is not governed by

10

isolated one-to-one relationships. Individual TFs typically regulate many genes, individual genes integrate inputs from multiple TFs, and the sign and strength of regulation can depend on the cellular context [Hobert 2008]. The central goal of constructing a GRN is to determine which TFs are responsible for the activation or repression of each gene in a given biological setting, and to do so at genome scale [Babu et al. 2004; Marbach et al. 2012; Badia-i Mompel et al. 2023].

Figure 2.2: Schematic linking the central dogma to gene regulatory networks. A gene is transcribed into mRNA and translated into protein; when this protein is a transcription factor, it binds to DNA to regulate the transcription of a target gene, initiating a subsequent round of transcription and translation. This repeated feedback process is summarized as a directed graph, where nodes represent genes (or TFs) and directed edges represent regulatory influence from TFs to their target genes.

Reconstructing these relationships provides a systems-level view of regulatory programming. A GRN offers a mechanistic explanation for observed gene expression patterns and supports hypotheses about unobserved drivers of phenotypes, such as latent regulatory activities or upstream perturbations. By uncovering regulatory relationships between TFs and their target genes, GRNs can help clarify mechanisms that distinguish healthy from diseased cellular states and can inform downstream analyses relevant to drug discovery and targeted therapies [Kim et al. 2023]. GRNs also provide a framework for understanding developmental processes and cellular differentiation, where cell fate transitions reflect changes in regulatory programs over time [Allaway et al. 2021; Özel et al. 2022]. More broadly, GRNs can be used to study how regulatory mechanisms evolve 11

and diverge across organisms, and how biological systems respond to environmental stimuli or stressors through coordinated changes in transcriptional control [Jackson et al. 2020]. A central challenge of reconstructing GRNs is that the TF-target gene edges are not directly observable at scale with standard sequencing assays. While specific regulatory interactions can be validated experimentally, current technologies do not provide a complete, direct readout of TF-target gene relationships across all genes and conditions. Reconstructing a GRN therefore relies on integrating measurable genomic data to infer the most plausible regulatory landscape and underlying regulatory connections [Skok Gibbs et al. 2022, 2024; Kamimoto et al. 2023]. This motivates the next section, which describes the major data modalities and experimental settings that provide evidence for regulatory relationships and define the practical regimes in which GRNs can be inferred.

2.3

Data Modalities for Gene Regulatory Networks

The reconstruction of GRNs relies on integrating genomic data that provides indirect or partial evidence of regulatory activity. TF binding and target gene regulation are typically not observable at scale through direct measurement, requiring GRN inference to draw upon diverse data modalities that reflect the downstream consequences or correlates of regulation. These data modalities include gene expression profiles across conditions, chromatin accessibility landscapes, TF binding data, and DNA sequence features that encode regulatory specificity. Each data type captures a different aspect of the regulatory system, with its own strengths, biases, and assumptions. The order in which these technologies emerged, beginning with microarrays, followed by bulk RNA-seq, and more recently single cell RNA-seq, has shaped how transcriptional activity is measured and continues to influence the types of biological questions that can be addressed. This section describes the major experimental modalities used to characterize regulatory activity and provides context for the types of data available in modern genomics to reconstruct GRNs.

12

2.3.1

Microarray Gene Expression Profiling

DNA microarrays were one of the earliest genome-scale technologies for measuring gene expression, providing the first high-throughput view of transcriptional programs across conditions. In a microarray experiment, cellular RNA is extracted from a biological sample and reversetranscribed into fluorescently labeled complementary DNA (cDNA) (illustrated in Figure 2.3A). cDNA is then hybridized to an array containing thousands of short oligonucleotide probes, each designed to target a specific gene or transcript [Jaksik et al. 2015]. The strength of the hybridization between the cDNA and each probe is measured via fluorescence intensity, which serves as a proxy for transcript abundance. Since each probe set is predefined, microarrays can only measure transcripts that are already annotated and represented on the array [Dufva 2009; Blohm and Guiseppi-Elie 2001; Howbrook et al. 2003]. Microarrays are widely used in regulatory studies because they provide an efficient and costeffective snapshot of gene expression across diverse conditions, perturbations, and timepoints [Tarca et al. 2006]. Such measurements are often used to identify genes that change together across biological states, offering indirect evidence for shared regulation [Lee et al. 2004]. Computationally, microarray data is represented as a gene expression matrix 𝑊 ∈ R𝑁 ×𝑀 , where 𝑁 is the number of samples or experimental conditions, and 𝑀 is the number of genes (or transcripts) probed on the array [Smyth 2005]. Each entry 𝑊𝑛,𝑚 contains the normalized, often log-transformed fluorescence intensity for gene 𝑚 in sample 𝑛, reflecting its estimated expression level [Ritchie et al. 2015]. These matrices provide a standardized representation of expression dynamics across conditions and form the basis for many statistical analyses of transcriptional regulation. Several limitations affect the interpretability of microarray data [Russo et al. 2003; AbdullahSayani et al. 2006]. First, because probe design is limited to known transcripts, microarrays lack the ability to detect unannotated genes, novel isoforms, or transcript structure [Jaluria et al. 2007].

13

Figure 2.3: Schematic overview of data modalities commonly used to reconstruct gene regulatory networks. (A) Microarray workflow: labeled cDNA hybridizes to probes on a fixed array to measure transcript abundance via fluorescence intensity. (B) Bulk RNA-seq: population-level transcriptomes are sequenced to quantify gene expression across conditions or samples. (C) Single-cell RNA-seq: RNA from individual cells is barcoded, sequenced, and aggregated into a cell-by-gene count matrix. (D) ATAC-seq: a transposase preferentially inserts adapters into accessible chromatin, enabling the identification of open regulatory regions. (E) ChIP-seq: immunoprecipitation of DNA-bound TFs allows direct measurement of TF binding sites genome-wide. (F) Motif databases: position weight matrices (PWMs) derived from experimental binding data represent sequence-specific TF binding preferences and are used to scan genomes for candidate binding sites.

14

Second, cross-hybridization, where similar sequences partially bind non-target probes, can introduce noise and confound gene-specific quantification [Murphy 2002]. Third, the dynamic range of fluorescence measurements is limited [Nadon and Shoemaker 2002]. Highly expressed genes may saturate the signal, while weakly expressed genes may fall below the detection threshold, leading to distorted estimates at expression extremes. Finally, microarray data are prone to batch effects and platform-specific artifacts, which hinder integration across studies and require rigorous normalization [Forster et al. 2003]. Despite these challenges, microarrays provided the first genome-scale views of transcriptional activity and played a foundational role in early efforts to uncover regulatory structure from gene expression data.

2.3.2

Bulk RNA Seqencing

Bulk RNA-sequencing (RNA-seq) emerged as a sequencing-based alternative to microarrays, offering a more flexible and quantitative platform for transcriptome-wide expression profiling [Wang et al. 2009]. Unlike microarrays, which rely on pre-defined probes, RNA-seq does not require prior knowledge of transcript sequences, making it suitable for detecting novel genes, splice variants, and isoforms. In a typical RNA-seq workflow (Figure 2.3B), total RNA is extracted from a population of cells, converted into a library of cDNA, and sequenced using high-throughput sequencing technologies. The resulting reads are aligned to a reference genome or transcriptome, and geneor transcript-level abundances are estimated based on the number of aligned reads [Marguerat and Bähler 2010]. RNA-seq provides a high-resolution view of gene expression programs across experimental conditions, perturbations, or timepoints [Marguerat et al. 2008]. It offers a wider dynamic range than microarrays and can more accurately quantify both lowly and highly expressed transcripts, making it an adopted modality for studying transcriptional regulation and for generating datasets used in GRN analysis [Zhao et al. 2014; Rai et al. 2018]. Computationally, RNA-seq data is typically represented as a gene expression matrix 𝑊 ∈ 15

R𝑁 ×𝑀 , where 𝑁 is the number of samples and 𝑀 is the number of genes (or transcripts) included in the quantification. Each entry 𝑊𝑛,𝑚 denotes the normalized expression level of gene 𝑚 in sample 𝑛, commonly reported on a log scale after normalization procedures such as transcripts per million (TPM), counts per million (CPM), or variance-stabilizing transformations [Ghosh and Chan 2016]. This matrix serves as a standardized representation of transcript abundance across samples and is used as a primary input for downstream statistical and network-based analysis. Several limitations of bulk RNA-seq affect the interpretation and downstream modeling of expression measurements. First, the approach aggregates RNA from many cells within a sample, yielding an average expression profile that can obscure heterogeneity among individual cells [Hegenbarth et al. 2022]. As a result, gene expression differences arising from rare cell types or dynamic regulatory states may be masked. Second, RNA-seq data are subject to technical variation stemming from library preparation protocols, sequencing depth, and alignment biases [seq 2014; Degner et al. 2009]. These factors can introduce systematic shifts in expression estimates across samples, necessitating careful normalization and experimental design to ensure comparability. Despite these challenges, bulk RNA-seq has become a useful tool for transcriptomics and continues to assist large-scale efforts to characterize transcriptional programs in health and disease [Costa et al. 2013].

2.3.3

Single-Cell RNA Seqencing

Single-cell RNA-sequencing (scRNA-seq) was developed to measure gene expression at the resolution of individual cells, addressing a key limitation of bulk RNA-seq, which averages expression across heterogeneous populations [Saliba et al. 2014]. In complex tissues composed of diverse cell types and dynamic cellular states, such averaging can obscure important regulatory differences [Kolodziejczyk et al. 2015]. scRNA-seq preserves this heterogeneity by isolating and profiling the transcriptomes of thousands to millions of individual cells within a single experiment. In a typical scRNA-seq workflow (Figure 2.3C), single cells are encapsulated, barcoded, and lysed, after 16

which their RNA is reverse-transcribed into cDNA, amplified and sequenced [Slovin et al. 2021]. Barcodes enable each read to be assigned to its cell of origin, resulting in a cell-by-gene matrix that reflects transcriptional variation across individual cells. The level of resolution achieved by scRNA-seq has transformed the study of complex biological systems. scRNA-seq enables the identification of discrete cell types, the characterization of continuous developmental or differentiation trajectories, and the exploration of gene expression patterns specific to transient or rare cell states [Chen et al. 2019a]. These properties make scRNA-seq an especially valuable modality for studying dynamic and cell-type specific regulatory programs, and for generating data that can inform models of gene regulation at single-cell resolution. scRNA-seq data is typically represented as a matrix 𝑊 ∈ R𝑁 ×𝑀 , where 𝑁 is the number of individual cells and 𝑀 is the number of genes. Each entry 𝑊𝑛,𝑚 records the number of unique molecular identifiers (UMIs) detected for gene 𝑚 in cell 𝑛, serving as a count of RNA molecules [Luecken and Theis 2019]. Total UMI counts can vary substantially between cells, due to differences in RNA content, capture efficiency, or sequencing-depth, and can cause raw counts to not be directly comparable. Downstream analyses typically operate on normalized or transformed versions of 𝑊 , where methods such as library size normalization, log-transformation, or variance stabilization are used to mitigate technical variability and improve interpretability of gene expression differences across cells [Hafemeister and Satija 2019]. Several technical challenges affect scRNA-seq data. First, expression matrices are characteristically sparse and noisy, with many entries containing zeros due to limited RNA capture per cell, a phenomenon often referred to as dropout [Van Dijk et al. 2018; Qiu 2020]. This sparsity complicates downstream modeling and can obscure signals from lowly expressed genes. Second, technical variability across cells, including batch effects, differences in sequencing depth, and variation capture efficiency, can introduce systematic biases that must be carefully controlled [Tung et al. 2017]. Third, observed variability across cells may reflect a combination of true bi17

ological differences and technical noise, making it difficult to disentangle meaningful regulatory signals without appropriate processing and modeling strategies [Vallejos et al. 2017]. Despite these challenges, scRNA-seq has become a central tool in modern genomics and provides a rich modality for studying gene regulation at cellular resolution [Saliba et al. 2014].

2.3.4

ATAC Seqencing for Chromatin Accessibility

While transcriptomic technologies measure the consequence of gene regulation, chromatin accessibility assays provide insight into the regulatory landscape itself. The Assay for TransposaseAccessible Chromatin using sequencing (ATAC-seq) [Buenrostro et al. 2013] is a widely used method for identifying regions of open chromatin, which are often indicative of active regulatory elements such as promoters, enhancers, and transcription factor binding sites [Grandi et al. 2022]. Open chromatin is characterized by nucleosome depletion, allowing regulatory proteins to access the DNA. ATAC-seq leverages this property by using a hyperactive transposase enzyme (Tn5) to preferentially insert sequencing adapters into accessible regions of the genome. These tagged fragments are then PCR amplified and sequenced, enabling genome-wide mapping of chromatin accessibility with high resolution and minimal input material (Figure 2.3D) [Buenrostro et al. 2015a]. Following sequencing, ATAC-seq data undergoes alignment to a reference genome, after which regions of high read density are identified through a process known as peak calling [Yan et al. 2020]. Peaks represent genomic intervals where the transposase inserted more frequently, suggesting accessible chromatin. These peaks can be classified as open regulatory elements and annotated relative to nearby genes [Sun et al. 2019]. Importantly, accessibility at a given region varies across cell types and conditions, reflecting context-specific regulatory activity [Grandi et al. 2022; Corces et al. 2018]. Peaks overlapping gene promoters often indicate transcriptional readiness, while distal peaks may correspond to enhancers or other regulatory elements. Bulk ATAC-seq can be represented as a binary or quantitative matrix 𝐶 ∈ R𝑁 ×𝑃 , where 𝑁 18

is the number of samples and 𝑃 is the number of genomic regions (e.g., peaks) identified across the dataset. Each entry 𝐶𝑛,𝑝 indicates whether region 𝑝 is accessible in sample 𝑛, often quantified by read counts or normalized signal intensity (e.g., reads per kilobase of transcript per million mapped reads or counts per million). For downstream analyses, these regions can be further linked to genes by proximity (e.g., nearest transcription start site), overlapped with promoter annotations, or through regulatory annotations such as chromatin state maps or enhancer-gene linkage datasets [Nasser et al. 2021]. ATAC-seq provides important complementary information to expression-based modalities. While transcriptomic data reflect the outcome of the regulatory processes, chromatin accessibility captures upstream regulatory potential and TF binding opportunity. However, ATAC-seq still presents several limitations [Yan et al. 2020]. First, peak calling can be sensitive to sequencing depth and noise, particularly in regions with low accessibility or high background signal [Grandi et al. 2022]. Second, many accessible regions are distal to known genes, complicating assignment of regulatory influence. Third, bulk ATAC-seq represents an average over many cells, potentially obscuring cell type-specific regulatory elements [Buenrostro et al. 2015b]. While single cell ATAC-seq technologies have been developed to address this limitation, they introduce sparsity and increase noise [Ji et al. 2020]. Despite these challenges, ATAC-seq remains a core modality for probing the regulatory landscape and identifying candidate cis-regulatory elements genomewide.

2.3.5

ChIP Seqencing for Transcription Factor Binding

While chromatin accessibility assays such as ATAC-seq provide indirect evidence of regulatory potential, chromatin immunoprecipitation follow by sequencing (ChIP-seq) enables direct measurement of TF binding at specific genomic loci [Mardis 2007]. In a ChIP-seq experiment (Figure 2.3E), proteins are crosslinked to DNA in living cells, the chromatin is fragmented, and an antibody specific to the TF of interest is used to immunoprecipitate the protein-DNA complexes 19

[Park 2009]. After reversing the crosslinks, the co-precipitated DNA is purified, sequenced, and mapped to a reference genome. Regions enriched for sequencing reads reflect genomic loci bound by the targeted TF under the profiled conditions. Following sequencing, downstream analysis of ChIP-seq data involves peak calling to identify regions of significant enrichment compared to background, using control input samples or statistical models to correct for sequencing biases [Thomas et al. 2017]. These peaks represent putative binding sites of the immunoprecipitated protein and are often annotated to known regulatory elements such as gene promoters, enhancers, or other cis-regulatory features [Jeon et al. 2020]. ChIP-seq is therefore a direct method for characterizing TF binding landscapes genome-wide, under specific cellular and environmental conditions. ChIP-seq data can be represented as a binary or continuous signal track across the genome, with enriched regions (peaks) summarized into a matrix 𝑇 ∈ R𝐾×𝑃 , where 𝐾 is the number of TFs profiled and 𝑃 is the number of genomic loci or peaks identified across experiments. Each entry 𝑇𝑘,𝑝 indicates whether a TF 𝑘 binds to a region 𝑝, or reflects the binding strength at that region, depending on whether peaks or signal intensities are used. These peaks can then be linked to candidate target genes based on proximity or experimentally informed annotations, yielding TF-gene interaction datasets that are widely used in regulatory genomics. ChIP-seq offers high-resolution, TF-specific information that is widely valuable for mapping the regulatory wiring of the genome. However, there are several limitations to the ChIP-seq approach. First, it requires high-quality, TF-specific antibodies, which are often unavailable for all proteins or species, limiting general applicability [Tahara and Ozaki 2025]. Second, ChIP-seq typically profiles one TF at a time, making large-scale experiments resource-intensive [Park 2009]. Third, peak calling can be confounded by noise, indirect binding, or chromatin accessibility biases, and binding does not always equate to functional regulation [Nakato and Shirahige 2017]. Finally, ChIP-seq reflects binding in a particular cellular context and timepoint, making it difficult to generalize across dynamic or heterogeneous systems. Despite these challenges, ChIP-seq re20

mains the gold standard for identifying in vivo TF binding sites and has been essential for building reference regulatory maps in many model systems.

2.3.6

Transcription Factor Binding Motifs

TFs regulate gene expression by binding to specific DNA sequences in the genome. These sequences, known as TF binding motifs, are typically short (6-15bp) and degenerate, allowing variation at some positions while retaining overall binding specificity [Inukai et al. 2017]. Motifs are often enriched at cis-regulatory elements such as promoters and enhancers, where TFs bind to modulate transcriptional activity [Spitz and Furlong 2012]. Identifying the presence of TF binding motifs in regulatory regions provides a sequence-level signature of potential TF-DNA interactions and enables genome-wide scans for candidate binding sites [Grant et al. 2011]. Motifs are most commonly represented as position weight matrices (PWMs), which describe the base preference of a TF at each position in the motif [Stormo 2000]. A PWM is derived by aligning experimentally determined binding sites and computing the frequency of each nucleotide at each position, often normalized by background base frequencies. These matrices can then be used to scan DNA sequences and assign binding scores, indicating the likelihood that a given TF binds at a specific genomic locus [Grant et al. 2011]. Several large-scale databases provide curated collections of TF motifs derived from diverse experimental and computational sources. These databases, such as JASPAR [Sandelin et al. 2004; Rauluseviciute et al. 2024], TRANSFAC [Wingender et al. 1996], HOCOMOCO [Kulakovskiy et al. 2013; Vorontsov et al. 2024], and CisBP [Weirauch et al. 2014], serve as key resources for annotating potential TF binding sites in regulatory regions (Figure 2.3F). Annotations can be used to interpret chromatin accessibility and ChIP-seq peaks. While motifs alone do not confirm in vivo binding, they provide valuable sequence-level hypotheses about where and how transcriptional regulation may occur. Taken together, these data modalities offer complementary views of the regulatory landscape. 21

Transcriptomic technologies such as microarrays, bulk RNA-seq, and single-cell RNA-seq measure the downstream effects of gene regulation, capturing expression patterns that reflect underlying transcriptional programs. In contrast, chromatin-based assays like ATAC-seq and ChIP-seq provide more direct information about regulatory potential and TF binding activity, offering insight into the mechanisms that shape gene expression. Sequence-derived features, including TF motifs from databases like JASPAR, TRANSFAC, HOCOMOCO, and CisBP, capture the intrinsic specificity of TF-DNA interactions encoded in the genome itself. While each modality has unique strengths and limitations, they all serve as valuable sources of evidence for reconstructing GRNs. The next chapter describes computational methods that leverage these data types, both individually or in combination, to infer regulatory relationships and model the structure of transcriptional control.

2.4

Machine Learning Foundations for Gene Regulatory Network Inference

Machine learning provides a unifying framework for formalizing GRN inference as a problem of recovering latent regulatory structure from observed genomic data and for comparing the diverse algorithms used in practice. This section first defines the GRN inference task and describes the core data and modeling challenges that make it ill-posed in realistic settings. We then review major families of GRN inference methods, spanning unsupervised, supervised, and prior-informed approaches, and introduce the statistical foundations that underpin the methods developed in Chapters 3 and 4. Finally, we summarize evaluation metrics and uncertainty-based diagnostics used to benchmark inferred networks against available reference gold standards.

22

2.4.1

Problem Definition

GRN inference aims to reconstruct the regulatory relationships that govern transcriptional control within biological systems. Formally, a GRN is represented as a directed acyclic graph in which nodes correspond to genes (or gene products such as TFs) and directed edges represent regulatory influences from TFs to their target genes. An edge from TF 𝑖 to gene 𝑗 indicates that the activity of TF 𝑖 modulates the transcriptional output of gene 𝑗, either through activation or repression.

Figure 2.4: Schematic overview of the input and output of GRN inference algorithms. The primary input is a gene expression matrix, and optional secondary input is a prior knowledge matrix. The selected GRN inference algorithm will produce a directed acyclic graph describing predicted releationships between TFs and their target genes.

While a GRN describes direct regulatory interactions between TFs and their target genes, such relationships are not directly measurable by existing genomic assays. Instead, GRN inference algorithms operate on measurable genomic data that reflect the consequences or correlates of 23

regulation. The primary input to most GRN inference methods is a gene expression matrix 𝑊 ∈ R𝑁 ×𝑀 , where 𝑁 denotes the number of samples, conditions, or cells, and 𝑀 denotes the number of genes. Each entry of 𝑊 captures the observed expression level of a gene under a particular condition. In some settings, inference algorithms also incorporate prior knowledge in the form of an auxiliary matrix 𝐴 ∈ R𝐾×𝑀 , where 𝐾 corresponds to TFs and entries encode prior evidence that a TF may regulate a given gene. The output of GRN inference is an estimated regulatory network, typically described as a weighted, directed adjacency matrix (Figure 2.4). The central modeling goal of GRN inference is to estimate latent regulatory relationships that can best explain the observed data. Given expression measurements 𝑊 , and optionally prior information 𝐴, algorithms aim to infer which TFs regulate which genes, how strongly, and under what conditions. Crucially, TF-target gene relationships are latent variables; expression data alone provides indirect evidence of regulation, and chromatin-based assays capture regulatory potential rather than functional regulatory effects. As a result, GRN inference is a fundamentally ill-posed problem that requires statistical assumptions, regularization, or inductive biases to constrain the space of plausible solutions. Different data modalities contribute complementary information to this inference task. Expression-based measurements, including microarrays (Section 2.3.1), bulk RNA-seq (Section 2.3.2), and single-cell RNA-seq (Section 2.3.3), reflect the downstream outcomes of regulatory activity and provide the primary signal used to infer regulation and regulatory influence. In contrast, chromatin accessibility data such as ATAC-seq (Section 2.3.4), TF binding assays such as ChIP-seq (Section 2.3.5), and sequence-derived features such as TF binding motifs (Section 2.3.6), encode upstream regulatory potential and specificity, and are often used to construct prior knowledge or structural constraints on the inferred network. Machine learning methods for GRN inference differ not only in how they integrate these diverse data types, but also in the way they formulate the inference problem. Approaches can range from unsupervised co-expression analysis to probabilistic generative models and supervised classification frameworks. The following sections aim 24

to describe these modeling approaches and the statistical foundations that underpin them.

2.4.2

Challenges in Gene Regulatory Network Inference

Despite decades of research, GRN inference remains a fundamentally challenging problem, from both a data and a modeling perspective. High-throughput genomic technologies offer only indirect, incomplete views of the underlying regulatory architecture, and the complexity of transcriptional regulation itself poses deep statistical and computational challenges. This section outlines five core challenges that motivate the development of advanced machine learning approaches for inferring GRNs. 2.4.2.1

Noisy, High-Dimensional, and Sparse Expression Measurements

Gene expression matrices are challenging inputs for GRN inference, particularly scRNA-seq, where measurements are noisy, high-dimensional, and sparse. Technical variation, which includes capture efficiency, sequencing depth, and amplification bias, introduces substantial cell-tocell measurement noise that can obscure regulatory signal [Chu et al. 2022]. In addition, scRNAseq count matrices contain many zeros due to a combination of true biological non-expression and technical dropout, which complicates the distinction between absence of transcription and failure to detect transcripts [Xu et al. 2022]. Although bulk RNA-seq reduces sparsity by averaging across many cells, this averaging also obscures cell-to-cell variability that is often essential for identifying regulatory relationships [Li et al. 2019]. Together, these properties make the statistical problem of recovering accurate regulatory structure from expression matrices ill-conditioned. The dimensionality is large, the effective signal-to-noise ratio is low, and the data is dominated by zeros and sampling variability.

25

2.4.2.2

Snapshot Sampling and Temporal Mismatch

An additional limitation of expression data is that transcriptomic measurements provide a snapshot of RNA abundance at the moment of capture, rather than a direct record of regulatory events over time [Marr et al. 2016]. Regulatory programs frequently act upstream of the observed mRNA state, and their effects can be observed with delays that break simple contemporaneous associations between regulator TFs and target genes [Qiu et al. 2020]. In particular, TF activity can be transient and is often shaped by processes that are not reflected by TF mRNA abundance, including protein accumulation, nuclear localization, post-translational modifications, and regulated degradation [Brent 2016]. Post-transcriptional mechanisms such as mRNA stabilization and degradation further decouple mRNA levels from the timing and magnitude of upstream regulatory inputs.

Figure 2.5: Snapshot sampling and temporal mismatch between TF transcripts and target-gene expression. Schematic time course illustrating how TF mRNA abundance can be offset in time from the regulatory events that drive target gene transcription. At 𝑡 0 , TF mRNA is highly expressed and then declines over subsequent time points, while TF protein accumulates and becomes positioned for regulation. Following an activating signal at 𝑡 2 , the TF binds DNA at 𝑡 3 , and the downstream target gene exhibits increased transcript abundance only later at 𝑡 4 .

26

As illustrated schematically in Figure 2.5, profiling expression at different moments can capture qualitatively different snapshots along this regulatory cascade (for example, during TF transcription, after protein accumulation, during activation and DNA binding, or only after the target gene has been expressed), yet the expression matrix alone does not reveal which snapshot of the regulatory landscape has been captured [Sha et al. 2024]. This temporal ambiguity is especially problematic for causal interpretation, where observed correlations can reflect delayed responses, indirect pathways, or regulatory events that occurred prior to measurement, rather than direct TF-target regulation. 2.4.2.3

Limited Observability of Regulatory Mechanisms

Even when expression measurements are high quality, the underlying regulatory mechanisms that generate these profiles are only partially observable. Expression data provides indirect evidence of transcriptional outcomes, but it does not specify which regulators were active, which genes they targeted, or whether the net effect was activating or repressive. Additional assays can constrain the search space of plausible interactions [Kim et al. 2023], but they remain incomplete and context-dependent. For example, chromatin accessibility determined by ATAC-seq, and TF binding profiles defined by ChIP-seq identify regulatory potential and occupancy, but binding and accessibility alone do not establish functional relevance or quantify regulatory effect sizes [Slattery et al. 2014]. Moreover, gene regulation depends on molecular processes beyond TF-DNA binding, including cofactor recruitment [Inge et al. 2024], chromatin remodeling [Aoyagi and Archer 2008], enhancer-promoter communication through 3D genome organization [Chaumeil and Skok 2012], and other mechanisms that are not directly captured by standard assays. These layers create nonlinear and context-dependent relationships between regulatory inputs and transcriptional outputs. The same binding event can have different consequences across cell types, developmental stages, or signaling conditions. Limited observability introduces an irreducible ambiguity in GRN 27

inference, where measured data often lack the resolution and completeness needed to uniquely reconstruct regulatory interactions. This challenge motivates models that treat regulation as a latent process and methods that integrate multiple modalities to reduce uncertainty.

Figure 2.6: Key challenges in gene regulatory network inference. (A) A central challenge in GRN reconstruction is determining which transcription factors (TFs) regulate which target genes. (B) Regulation is often combinatorial, with multiple TFs acting cooperatively or competitively to control the same gene, while individual TFs may regulate many targets. (C) Regulatory interactions are also cell-type and context-specific, such that the active GRN can vary across cellular states or conditions.

2.4.2.4

Combinatorial and Context-Specific Regulation

Transcriptional regulation is inherently combinatorial (Figure 2.6). For example, multiple TFs can cooperatively or competitively regulate a single gene, and each TF may regulate many genes in different contexts [Dubois-Chevalier et al. 2018]. This many-to-many structure leads to regulatory networks that are sparse but structured, with modular or hierarchical organization [Ravasi et al. 2010]. Capturing this complex structure requires models that can represent conditional dependencies and interactions between regulators, rather than treating regulatory effects as independent or additive. 28

Complexity increases in dynamic or heterogeneous systems, where GRNs may shift over time or between cell states. In these settings, a single regulatory interaction may be active only in a specific cell type, condition, or developmental stage. When expression data is aggregated across diverse states, such context-specific patterns can be obscured, and static prior-knowledge networks may also fail to reflect the regulatory program operating in the sampled cells. Methods that infer GRNs must therefore balance generalization with specificity, and account for contextdependent regulatory when possible. 2.4.2.5

Incomplete Ground Truth for Evaluation

Evaluating the accuracy of predicted GRNs remains a major obstacle caused by the scarcity of reliable gold standard annotations. While experimentally validated TF-target gene relationships exist for select interactions in well-studied organisms like yeast or human, these datasets are incomplete, context-specific, and often biased toward particular pathways or cell types [Karamveer and Uzun 2024]. As a result, evaluation can only assess a small subset of the predicted network. Limited access to ground truth TF-target gene interactions complicates not only supervised model training, but also objective benchmarking and performance comparisons across inference methods. Without comprehensive labels, most studies resort to proxy metrics such as overlap with ChIP-seq peaks or motif enrichment near predicted targets [Kernfeld et al. 2024]. Alternatively, performance can be computed on a held-out set of experimentally validated interactions, assessing how well the model recovers known regulatory edges. While informative, these metrics are indirect and can reflect signal unrelated to the true regulatory activity. In cross-organism or cross-cell-type settings, the challenge is further confounded by the lack of transferable evaluation benchmarks. Some approaches attempt to address this gap by generating synthetic networks or simulating expression data [Pratapa et al. 2020], but these strategies risk introducing unrealistic assumptions or artifacts that do not reflect biological complexity. The lack of robust, standardized, and 29

generalizable evaluation frameworks remains a fundamental barrier to assessing GRN inference quality. 2.4.2.6

Computational Complexity and Scalability

GRN inference typically involves evaluating regulatory relationships between thousands of TFs and tens of thousands of genes, leading to a combinatorial number of possible interactions. Even simple modeling approaches that estimate pairwise associations (e.g., correlation or mutual information) must process millions of TF-gene pairs, while more complex methods that fit structured models (e.g., Bayesian networks, probabilistic matrix factorization, or graphical models) face high computational costs [Banf and Rhee 2017]. These challenges are compounded in single-cell datasets, where the number of observations (cells) regularly exceed 100, 000, and algorithms must scale to high-dimensional input matrices with efficient memory and runtime performance [Skok Gibbs et al. 2022]. Regularization techniques, dimensionality reduction, or batching strategies are often used to address this, but scalability remains a concern, particularly for iterative inference frameworks or models that incorporate multiple data modalities. Advances in approximate inference, sparse modeling, and distributed high performance computing have helped mitigate these constraints, enabling more expressive models to be applied at scale.

2.4.3

Gene Regulatory Network Inference Algorithms

GRN inference algorithms aim to reconstruct regulatory relationships between TFs and their target genes using high-throughput experimental datasets, such as gene expression data. Direct experimental validation for regulatory edges is expensive and incomplete, requiring computational approaches to generate candidate networks in order to study how regulatory programs vary across conditions, cell types and time. Existing GRN inference methods span a broad spectrum of assumptions and data requirements, ranging from approaches that rely only on statistical 30

structure in expression matrices to models that learn from curated regulatory interactions, or incorporate prior biological knowledge as constraints [Hegde et al. 2025]. In the following section, we first review unsupervised methods that infer networks directly from expression data, and then contrast them with supervised and semi-supervised approaches that use labeled regulatory edges to train predictive models. 2.4.3.1

Unsupervised Methods

Unsupervised methods for GRN inference aim to reconstruct network structure directly from gene expression data, without access to labeled regulatory interactions or experimentally validated edges. These methods rely on the statistical structure of the expression matrix itself and often assume that regulatory relationships can be uncovered through co-variation, conditional dependence, or sparse predictive modeling. Unsupervised approaches were first developed in the context of microarray experiments, where large numbers of samples across diverse conditions made statistical inference feasible. As high-throughput sequencing technologies evolved, many of these methods were adapted to bulk RNA-seq, and more recently single-cell RNA-seq, often with modifications to account for data sparsity and the lack of replicates. Information-Theoretic Approaches One of the earliest families of approaches were co-expression or association networks, which infer putative interactions based on pairwise relationships between gene expression profiles [Madhamshettiwar et al. 2012]. Relevance networks constructed using Pearson or Spearman correlation were among the first to be applied, providing a simple yet powerful mechanism to identify gene pairs with correlated activity. These networks are typically undirected and best suited to capturing gene modules rather than causal regulatory interactions. Mutual information (MI) was later introduced to capture nonlinear dependencies, leading to MI-based relevance networks [Butte and Kohane 1999] and more advanced formulations like

31

Weighted Gene Co-expression Network Analysis (WGCNA) [Zhang and Horvath 2005], which groups genes into co-expression modules with potential shared regulation. These methods are broadly applicable to microarray and bulk RNA-seq data, and have been adapted to scRNA-seq [Morabito et al. 2023] by averaging cells into pseudobulk profiles or applying data-smoothing techniques. A refinement of co-expression networks came with the development of information-theoretic approaches that attempt to distinguish direct from indirect interactions by incorporating statistical corrections and pruning strategies. One of the most influential of these is ARACNe [Margolin et al. 2006], which applies the Data Processing Inequality (DPI) to remove edges that can be explained by a shared intermediate regulatory, enriching for direct regulatory interactions in mutual information (MI) networks. Building on this idea, the Context Likelihood of Relatedness (CLR) algorithm extends MI by normalizing each gene-gene MI value against the empirical distribution of MI scores for each gene, yielding a context-aware z-score matrix that emphasizes unusually strong associations [Faith et al. 2007; Zhu et al. 2016]. In contrast, Minimum Redundancy Network (MRNET) adopts a feature selection framework, treating each gene as a response variable and iteratively selecting regulators based on both their MI with the target and their redundancy with already selected features [Meyer et al. 2007]. Alternatively, Conservative Causal Core Network (C3NET) takes a highly selective approach by retaining only the single most significant MI edge for each gene, thereby prioritizing high-confidence associations over network completeness [Altay and Emmert-Streib 2010]. Finally, Partial Information Decomposition and Context (PIDC), which was specifically designed for scRNA-seq data, decomposes MI into unique, redundant, and synergistic components to account for the complex dependencies and sparsity inherent in scRNA-seq [Chan et al. 2017]. These methods collectively offer improved specificity over simple MI networks, particularly in high-sample regimes, with PIDC retaining relevance for its adaptation to single-cell data.

32

Sparse Regression Approaches Another class of unsupervised methods formulates GRN inference as a sparse regression problem, treating gene expression as a predictive modeling task. In this framework, the expression of each gene is modeled as a response variable, while the expression of candidate regulators, typically TFs, serve as the predictor set. One of the earliest and most influential methods in this category is the Inferelator [Greenfield et al. 2013], which employs regularized regression to identify a parsimonious set of regulators per gene, often incorporating temporal or experimental design information to improve interpretability. This strategy has since been extended through the use of sparsity-inducing penalties such as LASSO or Elastic Net, enabling robust inference in highdimensional settings [Skok Gibbs et al. 2022; Miraldi et al. 2019]. Methods like TIGRESS [Haury et al. 2012] further stabilize inference by coupling regression with bootstrapping and stability selection, providing more reproducible edge estimates across data perturbations. Building on the same sparse modeling principle, ensemble-based methods such as GENIE3 [Huynh-Thu et al. 2010] and GRNBoost2 [Moerman et al. 2019] recast the problem as one of feature importance, using random forests or gradient-boosted trees to estimate the contribution of each TF to the expression of target genes. These approaches have gained popularity for single-cell applications due to their scalability, robustness to sparsity, and compatibility with datasets lacking explicit temporal or experimental labels. Conditional Dependence and Graphical Models Beyond direct association or regression, a separate set of methods aims to reconstruct regulatory relationships by estimating conditional dependencies among genes. These approaches are rooted in the principle that true regulatory interactions are best revealed by controlling for indirect effects, typically through partial correlation or graphical modeling techniques. Gaussian graphical models are a widely used formulation, where the inverse covariance (precision) matrix encodes the conditional dependence structure of the system. Estimation procedures such as the

33

Graphical Lasso [Friedman et al. 2008] improve sparsity constraints to recover interpretable and data-efficient networks. Complementary strategies based on shrinkage, as implemented in methods like GeneNet [Schäfer et al. 2006], provide stable estimates in high-dimensional but smallsample regimes. While these models are well suited to bulk expression data with large numbers of samples, their reliance on Gaussian assumptions and sensitivity to noise can limit their utility in single-cell contexts, unless mitigated by aggregation, denoising, or imputation techniques. Bayesian Networks Bayesian network learning introduces a probabilistic perspective to GRN inference by modeling the joint distribution over gene expression and identifying directed edges that represent conditional dependencies [Friedman et al. 2000]. These methods offer a principled framework for capturing regulatory directionality and uncertainty, with the added advantage of allowing incorporation of prior knowledge into the learning process [Werhli and Husmeier 2007]. However, Bayesian networks are computationally intensive and typically require discretizations of continuous expression values, which can obscure subtle expression changes. Dynamic Bayesian networks extend this approach to time-series data, modeling regulatory programs as evolving over discrete temporal intervals. These have proven particularly useful in developmental studies or perturbation experiments where time-course data is available and temporal ordering among genes is biologically meaningful [Husmeier 2003]. Trajectory-Aware Methods The advent of single-cell transcriptomics has spurred the development of trajectory-aware GRN inference methods that exploit the pseudo-temporal structure of differentiating or transitioning cells. These methods assume that cells can be ordered along a latent temporal axis and seek to uncover regulators that drive expression changes across this inferred trajectory. Methods such as SCODE [Matsumoto et al. 2017] model gene expression dynamics using linear ordinary differential equations, positing that observed expression levels arise from regulatory activity along pseu34

dotime. GRISLI [Aubin-Frankowski and Vert 2020] adopts a similar framework but incorporates sparse regression to infer dynamic, time-varying interactions. Methods like SINCERITIES [Papili Gao et al. 2018] leverage timestamped single-cell measurements, while LEAP [Specht and Li 2017] uses lagged correlation to infer putative regulatory delays between genes along the trajectory. These approaches are specifically tailored to the challenges and opportunities of single-cell data, including sparsity, asynchronous sampling, and nonlinear expression dynamics, and they represent a growing frontier in unsupervised GRN modeling. Summary Together, these methods highlight the breadth of statistical strategies available for unsupervised GRN inference using expression data alone. From simple pairwise correlations to sophisticated probabilistic models, each approach balances different trade-offs in scalability, interpretability, and robustness to noise. While unsupervised methods can reveal informative patterns in highthroughput data, they ultimately operate on the assumption that statistical dependencies reflect underlying regulatory mechanisms. In the absence of experimental validation or orthogonal priors, these networks should be interpreted with caution, particularly when used to derive biological hypotheses or inform downstream analyses. 2.4.3.2

Supervised and Semi-Supervised Methods

Supervised learning approaches for GRN inference leverage labeled examples of known regulatory interactions to train predictive models that generalize to unseen TF-gene pairs. In contrast to unsupervised methods, which infer regulatory relationships solely from expression patterns, supervised methods explicitly rely on curated gold standards. Examples of these curated gold standards, which guide the learning process, consist of interaction labels derived from ChIPbased experiments, genetic perturbations, or high-confidence database annotations. When sufficient labeled regulatory edges are available, supervised models can more accurately distinguish

35

true regulatory interactions from spurious associations by learning discriminative patterns that differentiate positive from negative examples. Within the space of supervised methods, SIRENE [Mordelet and Vert 2008] reformulates GRN inference as a series of local binary classification problems. For each TF, SIRENE trains a support vector machine (SVM) using expression-based features and known TF-target labels, with the assumption that co-regulated targets exhibit similar expression profiles. Another supervised approach that extends traditional machine learning to temporal expression profiles is dynGENIE3 [Huynh-Thu and Geurts 2018], which builds upon the random forest framework of GENIE3 [Huynh-Thu et al. 2010] by incorporating time-series information. Although GENIE3 is typically categorized as unsupervised learning, dynGENIE3 adds supervision through semi-parametric modeling of dynamic regulatory effects, integrating temporal gradients with ensemble regression to better capture time-dependent regulatory influences. In recent years, deep learning has been applied extensively to supervised GRN inference, particularly in contexts where sequence and chromatin features complement expression data. DeepIMAGER [Zhou et al. 2024] uses a convolutional neural network (CNN) architecture to convert co-expression patterns and TF binding features into an image-like format for classification, training on labeled regulatory pairs derived from scRNA-seq and ChIP-seq data. Similarly, DeepSEM [Shu et al. 2021] employs a deep structural equation model that jointly learns latent regulatory influences and gene expression dependencies, using supervision from curated TF-target gene interactions to refine its parameter estimates. Transformer-based models, which leverage self-attention mechanisms to capture long-range dependencies, have emerged as powerful supervised learners for GRN tasks. STGRNs [Xu et al. 2023] applies a transformer encoder to gene expression and positional features, learning representations that encapsulate regulatory context before classification. Parallel to this, GRNFormer [Hegde and Cheng 2025] leverages a variational graph transformer autoencoder to model both global and local regulatory patterns in single-cell data. In GRNFormer, subgraphs centered on 36

TFs are embedded using attention layers, and a variational decoder estimates a probabilistic adjacency matrix, trained with ground-truth regulatory data using a composite loss that balances prediction accuracy and latent regularization. Another supervised method is RSNET [Jiang and Zhang 2022], which combines mutual informationbased filtering with a sparse linear regression model. RSNET uses MI to identify candidate TFgene regulatory pairs, then applies a redundancy-silencing optimization that penalizes redundant regulators and enhances high-confidence edges. This approach improves GRN inference by prioritizing direct regulatory interactions and reducing spurious co-expression effects. Finally, alternative supervised strategies explore different problem formulations. AnomalGRN [Zhou et al. 2025] reframes GRN inference as a graph anomaly detection task, in which interacting gene pairs are treated as rare or abnormal nodes, and non-interacting pairs as normal nodes. The method reconstructs a graph over gene-pair nodes using expression-derived features, then applies graph neural networks with sparsification to identify meaningful regulatory interactions in the presence of extreme class imbalance and noisy connectivity. This perspective is particularly wellsuited to single-cell data, where dropout, sparsity, and the scarcity of positive regulatory links complicate standard supervised classification. Collectively, these supervised and semi-supervised methods illustrate the evolution of GRN inference from early binary classifiers to modern frameworks based on deep learning, graph neural networks, and alternative formulations such as anomaly detection. These methods demonstrate how supervised learning can integrate heterogeneous biological features, including expression, sequence, and chromatin context, while learning regulatory patterns from available labeled data. However, because such labels remain sparse outside of well-characterized organisms and experimental settings, the performance of these approaches is ultimately constrained by the scale and quality of known interactions. In the next section, we examine hybrid frameworks that incorporate prior knowledge directly into statistical models of expression, aiming to combine the fidelity of supervised learning with the scalability of unsupervised inference. 37

2.4.3.3

Prior-Informed Methods

An important class of GRN inference methods enhances expression-based modeling by incorporating prior biological knowledge. These priors, typically constructed from sequence motifs, chromatin accessibility, TF binding profiles, or curated regulatory databases, serve as structured constraints that guide or regularize the inference process (overview illustrated in Figure 2.7). Rather than treating every possible TF-gene pair as a candidate interaction, prior-informed models leverage biological evidence to define a subset of likely regulatory edges. This approach not only reduces the combinatorial search space, but also increases the biological plausibility and interpretability of the inferred networks.

Figure 2.7: Typical pipeline for constructing prior-knowledge matrix for downstream GRN inference. (A) Open chromatin regions are identified from ATAC-seq data to define candidate regulatory DNA accessible to transcription factor (TF) binding. (B) These accessible regions are scanned for known TF binding motifs to identify potential TF occupancy. (C) Motif-containing accessible regions are assigned to nearby genes, yielding candidate TF-target regulatory interactions. (D) Genome-wide interactions are aggregated into a prior-knowledge matrix, in which rows correspond to target genes, columns correspond to TFs, and entries indicate putative regulatory relationships. This prior matrix can then be incorporated into GRN inference methods to constrain the search space toward biologically plausible edges.

38

Of these prior knowledge constrained GRN inference algorithms, the Inferelator [Skok Gibbs et al. 2022; Greenfield et al. 2013; Arrieta-Ortiz et al. 2015; Miraldi et al. 2019] constructs a prior matrix by scanning for TF binding motifs within accessible chromatin regions, typically promoterproximal ATAC-seq peaks, and integrates this prior with gene expression using sparse regression models. The prior matrix represents binary or weighted confidence scores for potential regulatory edges and serves as a constraint during model fitting, encouraging the selection of edges that are supported by the data and consistent with known biology. Similarly, CellOracle [Kamimoto et al. 2023] builds motif-informed prior networks from promoter regions that intersect open chromatin. These priors are then integrated with single-cell expression data using linear models or dynamical simulations to estimate TF-gene relationships and predict perturbation outcomes. Additionally, pySCENIC [Van de Sande et al. 2020] provides another prior-constrained GRN inference approach, by incorporating motif-based priors as a posthoc filtering step in a hybrid inference pipeline. pySCENIC uses GENIE3 to infer co-expression modules from gene expression data. These candidate regulatory modules are then refined by pruning edges lacking motif support in the promoter regions of the predicted target genes. By constraining GRN inference to a biologically plausible sub-network, prior-informed methods improve statistical power, particularly in low-sample or high-noise regimes. They also enable transferability across datasets by anchoring inference in shared regulatory logic rather than dataset-specific expression correlations. Importantly, the effectiveness of these methods depends on the quality of the priors themselves. Mis-annotated or overly inclusive priors can bias the model or suppress the discovery of novel edges. Moreover, in organisms or cell types where chromatin or TF binding data are limited, constructing meaningful priors may be challenging, prompting recent efforts to generate priors directly from sequence using learned models, as is discussed in Chapter 4.

39

2.4.4

Machine Learning and Statistical Foundations

This section introduces the core statistical and machine learning ideas that underpin the GRN inference methods developed in Chapters 3 and 4. We begin with latent-variable representations of expression via matrix factorization and use identifiability to motivate why external information is needed to align latent dimensions with named TFs. We then place factorization within a probabilistic generative framework and describe variational inference as a practical approach for approximate posterior inference with quantified uncertainty. Finally, we introduce sequence language models and standard classification objectives as the tools used to construct sequencederived priors that provide TF-specific structure to constrain downstream GRN inference. 2.4.4.1

Latent Variables and Matrix Factorization

GRN inference is fundamentally a problem of reasoning about unobserved causes from noisy measurements. Gene expression assays provide an observed snapshot of mRNA transcript abundance, however, the regulatory quantities that generate these measurements, such as TF activity and TF-target gene influence, are not directly measured and may not be well-approximated by TF mRNA abundance alone. Latent variable models formalize this distinction by introducing unobserved variables that capture the underlying regulatory state, using them to explain the observed expression matrix. Let 𝑊 ∈ R𝑁 ×𝑀 denote an observed gene expression matrix, where 𝑁 is the number of samples or cells and 𝑀 is the number of genes. Matrix factorization assumes that the dominant structure in 𝑊 can be represented by a smaller number of latent factors 𝐾 ≪ 𝑚𝑖𝑛(𝑁 , 𝑀), yielding the approximation 𝑊 ≈ 𝑈𝑉 ⊤,

(2.1)

where 𝑈 ∈ R𝑁 ×𝐾 and 𝑉 ∈ R𝑀×𝐾 . The matrix 𝑈 assigns each sample 𝑛 a vector of 𝐾 latent factor values 𝑈𝑛,: , while the matrix 𝑉 assigns each gene 𝑚 a vector of factor loadings 𝑉𝑚,: . Under this 40

model, the predicted expression of gene 𝑚 in sample 𝑛 is a weighted sum of factor contributions,

(𝑊𝑝𝑟𝑒𝑑 )𝑛,𝑚 =

𝐾 ∑︁

𝑈𝑛,𝑘 𝑉𝑚,𝑘 .

(2.2)

𝑘=1

This decomposition separates sample-specific variation, captured by 𝑈 , from gene-specific response patterns, captured by 𝑉 , providing a compact representation of coordinated expression programs [Stein-O’Brien et al. 2018]. In the GRN setting, the latent factors are often interpreted in relation to TF-driven regulation. A common choice is to set 𝐾 to the number of TFs under study and interpret the columns of 𝑈 as latent TF activities across samples or cells, and the columns of 𝑉 as gene-specific responses to those activities. With this interpretation, each entry 𝑈𝑛,𝑘 represents the activity of TF 𝑘 in sample 𝑛, and each entry 𝑉𝑚,𝑘 represents the extent to which TF 𝑘 is associated with changes in expression of gene 𝑚. The matrix product 𝑈𝑉 ⊤ then aggregates the contributions of many TFs to each gene, consistent with combinatorial regulation. The latent variable framing is useful for two reasons. First, it permits drivers of regulation (TFs) to be inferred even when they are not directly observed, allowing the model to capture regulatory signals that may arise from post-transcriptional processes or upstream inputs not reflected in TF mRNA measurements. Second, it provides a natural route to reconstructing regulatory structure. The gene-by-factor loadings in 𝑉 can be interpreted as a candidate TF-gene interaction, while 𝑈 captures the cell- or condition-specific activity state for a TF, determining its contribution to a target gene’s observed expression. Matrix factorization alone, however, does not uniquely determine the latent factors. Multiple pairs (𝑈 , 𝑉 ) can yield the same product 𝑈𝑉 ⊤ , and additional constraints or side information are generally required to assign a stable interpretation to individual factors. This issue is particularly important in GRN inference, when factors are intended to correspond to specific TFs. The next subsection addresses this identifiability challenge and motivates the role of prior knowledge in

41

anchoring the latent space to interpretable regulatory components. 2.4.4.2

Identifiability and the Role of Priors

A central limitation of matrix factorization models is that the latent factors are not uniquely determined by the observed data [Papastamoulis and Ntzoufras 2022]. The approximation 𝑊 ≈ 𝑈𝑉 ⊤ specifies only the product of two matrices, not a unique choice of 𝑈 and 𝑉 . As a result, multiple distinct latent representations can explain the same expression matrix equally well. This non-uniqueness is referred to as an identifiability issue, and it becomes particularly consequential when the goal is to interpret latent dimensions as specific TFs. The simplest demonstration of non-identifiability arises from permutations of the latent dimensions. Let 𝑃 ∈ 𝑅 𝐾×𝐾 be any permutation matrix. Then 𝑈𝑉 ⊤ = (𝑈 𝑃)(𝑉 𝑃) ⊤,

(2.3)

since (𝑉 𝑃) ⊤ = 𝑃 ⊤𝑉 ⊤ and 𝑃𝑃 ⊤ = 𝐼 . This implies that the ordering of 𝐾 latent factors is arbitrary; swapping two columns of 𝑈 and the corresponding columns of 𝑉 leaves the reconstructed expression unchanged. More simply, without additional information, there is no principled way to determine which latent dimension corresponds to which TF. When 𝐾 is large, the number of equivalent permutations grows as 𝐾!, so a factorization aligned with TF identities is unlikely to emerge consistently from expression data alone. In GRN inference, interpretability requires more than recovering a low-rank approximation to 𝑊 . The latent dimensions are intended to correspond to named TFs, so the inferred regulatory structure must be aligned to those TF identities. Priors are essential for introducing external information that breaks the symmetry of the latent space by preferring solutions consistent with known or hypothesized TF-gene relationships. Conceptually, priors act as anchors biasing the model toward factorization solutions in which particular genes load onto particular TF-associated

42

dimensions, allowing a stable and interpretable mapping between latent factors and TF identities [Leung and Drton 2016]. Within probabilistic formulations of factor models, priors can be placed directly on the geneby-factor loadings 𝑉 , or on a structured representation of TF-gene interactions that governs 𝑉 . The key requirement is that the prior must encode TF-specific information, instead of generic regularization. For example, shrinkage priors that promote sparsity can improve stability and reduce overfitting. These priors, however, do not themselves resolve which factor corresponds to which TF. In contrast, a prior derived from a combination of TF binding profiles and chromatin accessibility, or curated regulatory interactions, supplies asymmetric information across TFs and genes, and therefore can distinguish one latent dimension from another. The role of priors is not only a technical fix, but reflects a broader principle in GRN inference. The space of candidate regulatory networks consistent with expression data is extremely large, and many networks can generate similar expression patterns under realistic noise levels and sampling constraints. Prior knowledge narrows this space by restricting attention to biologically plausible edges and by aligning latent factors with TF identities so they can be interpreted. Later sections will instantiate this idea in two complementary ways, the first by using priors to anchor latent-variable inference over regulatory structure (Chapter 3), while the second focuses on constructing higher-quality priors from sequence to better constrain downstream inference (Chapter 4). 2.4.4.3

Probabilistic Matrix Factorization and Generative Modeling

Matrix factorization provides a useful low-dimensional representation of gene expression, but it does not by itself specify how expression measurements are generated, how uncertainty should be quantified, or how prior biological knowledge should be incorporated in a principled way. Probabilistic matrix factorization [Mnih and Salakhutdinov 2007] addresses these limitations by reframing factorization as a generative model. Rather than fitting 𝑈 and 𝑉 solely by minimiz43

ing a reconstruction objective for 𝑊 ≈ 𝑈𝑉 ⊤ , it specifies a likelihood for the observed matrix 𝑊 together with priors over latent variables such as cell-specific activities and gene-specific regulatory effects. This perspective makes it possible to encode biological assumptions through prior distributions, to model measurement noise explicitly, and to infer a posterior distribution over unobserved regulatory structure rather than a single point estimate. As in the previous subsections, we let 𝑊 ∈ R𝑁 ×𝑀 denote an observed gene expression matrix, with 𝑁 cells and 𝑀 genes. In probabilistic matrix factorization, we introduce latent matrices 𝑈 ∈ R𝑁 ×𝐾 and 𝑉 ∈ R𝑀×𝐾 and define a likelihood for 𝑊 conditioned on these latent factors. The defining assumption is that the expected expression is given by the matrix product 𝑈𝑉 ⊤ , while the observations deviate from this expectation due to measurement noise. We can formulate this with equation (2.4),

𝑝 (𝑊 | 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 ) =

𝑁 Ö 𝑀 Ö

𝑝 (𝑊𝑛,𝑚 | (𝑈𝑉 ⊤ )𝑛,𝑚 , 𝜎𝑜𝑏𝑠 ),

(2.4)

𝑛=1 𝑚=1

where 𝜎𝑜𝑏𝑠 is an observation noise parameter and the conditional distribution 𝑝 (·) is chosen to reflect the expression measurement type. Here, the likelihood specifies an explicit stochastic relationship between the latent regulatory variables and observed expression, separating the biological signal captured by 𝑈𝑉 ⊤ from expression measurement variability. The probabilistic formulation is completed by specifying prior distributions over the latent variables, 𝑝 (𝑊 | 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 ) = 𝑝 (𝑈 )𝑝 (𝑉 )𝑝 (𝜎𝑜𝑏𝑠 ),

(2.5)

where the choice of priors determine both regularization and biological structure. TF-specific priors can incorporate external evidence about which TF-gene interactions are plausible, thereby anchoring the latent dimensions to interpretable TF identities as discussed in Section 2.4.4.2.

44

We can formulate the joint distribution over observed and latent variables with equation (2.6),

𝑝 (𝑊 , 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 ) = 𝑝 (𝑊 | 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 )𝑝 (𝑈 )𝑝 (𝑉 )𝑝 (𝜎𝑜𝑏𝑠 ),

(2.6)

which defines a complete generative story. Here, we first draw latent variables from their priors, then generate expression measurements from the likelihood. Under this model, GRN inference corresponds to computing the posterior distribution,

𝑝 (𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 | 𝑊 ),

(2.7)

and using the inferred regulatory component of V to characterize TF-gene relationships. In practice, exact posterior inference is intractable for most choices of priors and likelihoods, as we do not know the true data generating distribution to which 𝑊 belongs. This motivates the following section (Section 2.4.4.4), which describes approximate inference and optimization approaches. Overall, probabilistic matrix factorization provides a general statistical scaffold for GRN inference. It retains the dimensionality-reduction benefits of matrix factorization while enabling principled incorporation of prior knowledge, explicit modeling of noise, and posterior inference over regulatory structure with quantified uncertainty. 2.4.4.4

Variational Inference

The probabilistic matrix factorization model defined in the previous section specifies a generative story for how an observed expression matrix𝑊 could arise from latent variables such as 𝑈 , 𝑉 , and 𝜎𝑜𝑏𝑠 . Once a generative model is specified, the central computational task becomes posterior inference. Given the observed data 𝑊 , we would like to infer which latent configurations are plausible under the model, i.e., compute 𝑝 (𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 | 𝑊 ). This posterior is the mathematical object that encodes uncertainty about latent regulatory structure, and it is what enables uncertainty-aware GRN reconstruction rather than relying on a single point estimate. 45

For notational convenience, we collect all latent variables into a single variable 𝑧, where 𝑧 = 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 , so that the joint distribution defined by the generative model can be written completely as 𝑝 (𝑊 , 𝑧) and the posterior of interest becomes 𝑝 (𝑧 | 𝑊 ). In most reliable models, however, the posterior cannot be computed exactly. The obstacle is the normalizing constant in Bayes’ rule, the marginal likelihood (also referred to as the evidence), ∫ 𝑝 (𝑊 ) = 𝑝 (𝑊 , 𝑧)𝑑𝑧. This integral couples together all latent variables and is rarely tractable in high dimensions. Variational inference (VI) [Blei et al. 2017] addresses this by replacing exact inference with an optimization problem. Instead of computing 𝑝 (𝑧 | 𝑊 ) directly, we choose a simple family of distributions Q and search within that family for an approximation 𝑞(𝑧) that is close to the true posterior. KL Divergence In order to estimate a precise approximate posterior, VI typically uses the Kullback-Leibler (KL) divergence as a distance measurement between the true and approximate posterior distributions. The variational objective is  𝑞 ∗ (𝑧) = arg min 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) ,

(2.8)

𝑞(𝑧)∈Q

where  𝑞(𝑧) . 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) = E𝑧∼𝑞(𝑧) log 𝑝 (𝑧 | 𝑊 ) 



(2.9)

The formulation in (2.8) captures the intuition of VI where we select the best approximation to the posterior from a tractable family Q. The remaining problem is practical: the objective still involves 𝑝 (𝑧 | 𝑊 ), which depends on the intractable evidence 𝑝 (𝑊 ). Evidence Lower Bound The standard derivation of the evidence lower bound (ELBO) follows from rewriting the KL diver-

46

𝑝 (𝑊 ,𝑧)

gence in terms of quantities we can evaluate. The key identity is Bayes’ rule, 𝑝 (𝑧 | 𝑊 ) = 𝑝 (𝑊 ) . Substituting this into the definition of KL divergence (2.9) yields  𝑞(𝑧) 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) = E𝑧∼𝑞(𝑧) log 𝑝 (𝑊 , 𝑧)/𝑝 (𝑊 )   𝑞(𝑧) = E𝑧∼𝑞(𝑧) log + log 𝑝 (𝑊 ). 𝑝 (𝑊 , 𝑧) 



(2.10)

This step involves the main trick, in which we replace the posterior 𝑝 (𝑧 | 𝑊 ) with the joint distribution 𝑝 (𝑊 , 𝑧) and the evidence 𝑝 (𝑊 ). The joint 𝑝 (𝑊 , 𝑧) is exactly as defined in the previous section through the generative model (likelihood times priors), so it is accessible. The evidence log 𝑝 (𝑊 ) in (2.10) is still intractable, but it is also a constant with respect to 𝑞(𝑧). We can now collect all terms that depend on 𝑞 into a single objective:

  L (𝑞) = E𝑧∼𝑞(𝑧) log 𝑝 (𝑊 , 𝑧) − log 𝑞(𝑧) .

(2.11)

Using definition (2.11), the previous equation (2.10) becomes the following decomposition

 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) = −L (𝑞) + log 𝑝 (𝑊 ).

(2.12)

This relationship demonstrates how optimizing the ELBO solves the original variational problem [Jordan et al. 1999]. Since log 𝑝 (𝑊 ) does not depend on 𝑞(𝑧), minimizing the KL divergence is equivalent to maximizing L (𝑞):

 arg min 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) ≡ arg max L (𝑞). 𝑞(𝑧)∈Q

(2.13)

𝑞(𝑧)∈Q

The term, evidence lower bound, follows from one final rearrangement:

 log 𝑝 (𝑊 ) = L (𝑞) + 𝐷 KL 𝑞(𝑧) ∥ 𝑝 (𝑧 | 𝑊 ) .

47

(2.14)

Because the KL divergence is always nonnegative, this implies L (𝑞) ⩽ log 𝑝 (𝑊 ), so L (𝑞) is the lower bound on the evidence. Maximizing the ELBO therefore has a clear interpretation. It derives 𝑞(𝑧) toward the true posterior while also implicitly increasing the model’s ability to explain the observed data under the generative story. Expanding the joint 𝑝 (𝑊 , 𝑧) = 𝑝 (𝑊 | 𝑧)𝑝 (𝑧) yields a more interpretable form of the ELBO:

      L (𝑞) = E𝑧∼𝑞(𝑧) log 𝑝 (𝑊 | 𝑧) + E𝑧∼𝑞(𝑧) log 𝑝 (𝑧) − E𝑧∼𝑞(𝑧) log 𝑞(𝑧) .

(2.15)

The first term in 2.15 encourages latent variables that reconstruct the observed expression matrix well under the likelihood (data fit). The second term encourages latent variables that remain plausible under priors, while the final term is the entropy of 𝑞(𝑧), which discourages overly confident approximations when the posterior is uncertain. In the context of probabilistic matrix factorization for GRN inference, VI provides a practical route from a generative model to uncertainty-aware estimates of latent regulatory structure. Rather than attempting to compute 𝑝 (𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 | 𝑊 ) exactly, we optimize a tractable approximation 𝑞(𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 ) by maximizing the ELBO. This yields posterior summaries (e.g., means and variances) for regulatory parameters, enabling downstream GRN analyses that explicitly reflect uncertainty in inferred TF activities and TF-target gene relationships. 2.4.4.5

Seqence Language Models

Sequence-based priors naturally complement the latent-variable and probabilistic modeling framework described above by providing TF-specific structure that expression data alone cannot resolve. The earlier subsections highlighted two recurring challenges in GRN inference. First, expression measurements are noisy, partial snapshots of regulatory activity, and second, latent factor models are not identifiable without external information that anchors latent dimensions to specific TFs. Molecular sequence offers a direct source of such anchoring information. TF

48

binding and regulatory specificity are ultimately encoded in sequence, through properties of the TF sequence that shape its DNA-binding preferences, and through cis-regulatory DNA sequence that determines how genes respond to regulatory inputs. Sequence language models build on this premise by learning reusable representations of biological sequences from large corpora, enabling the construction of informative priors even when experimentally validated TF-gene interactions are sparse or incomplete. A sequence language model is a neural network trained to learn general-purpose representations of biological sequences from large-scale data. In biology, the alphabet consists of nucleotides A,T,C,G for DNA, A,C,G,U for RNA, or amino acids for proteins, and a sequences can be treated analogously to a sentence. The most widely used family of models for this purpose is based on the Transformer [Vaswani et al. 2017; Devlin et al. 2019], which maps an input sequence to contextual token representations. These representations capture dependencies between positions across the sequence, allowing the model to represent both local patterns, such as motifs, and broader context that can modulate binding and regulatory function. Pretraining provides the main advantage of sequence language models in settings where labeled biological data are limited. During pretraining, the model is optimized on large unlabeled sequence collections using a self-supervised objective, most commonly masked language modeling. Under this objective, a subset of tokens is hidden and the model learns to predict them from context, encouraging it to discover recurring sequence regularities. The resulting encoder can convert an input sequence into a latent representation, known as an embedding, that summarizes sequence content in a form useful for downstream tasks. Embeddings can be used in feature-based approaches, where the pretrained encoder is frozen and embeddings serve as inputs to a task-specific model, or in a fine-tuning approach, where the encoder is further trained end-to-end, allowing embedding representations to become specialized for a particular prediction problem. In GRN inference, a recurring goal is to score candidate TF-gene relationships in a way that 49

can be used as prior evidence. Sequence language models support this by producing TF- and gene-specific representations that can be combined into an interaction score. When repeated over many TF-gene pairs, these scores form a TF-by-gene matrix that can be interpreted as a prior network. This provides an additional mechanism for addressing the identifiability issue described earlier, where this sequence-derived prior can introduce TF-specific asymmetry across genes and therefore can help anchor latent dimensions to named TFs in downstream probabilistic inference. The role of sequence-derived priors is best understood as a way to constrain and structure uncertainty, rather than as a replacement for expression-based modeling. A probabilistic factor model operating on 𝑊 can quantify uncertainty about latent regulatory effects and cell-specific activities, but it requires a biologically meaningful inductive bias to map latent dimensions to TF identities. Sequence models supply this inductive bias by proposing which TF-gene edges are plausible apriori. Downstream inference can then refine, reweight, or prune these edges using the observed expression matrix, yielding a regulatory network that is both biologically grounded and adapted to the cellular context capture in 𝑊 . This division of labor mirrors the broader theme of this chapter: priors reduce the search space of plausible regulatory explanations, while probabilistic inference provides a mechanism to combine priors with noisy observations and propagate uncertainty into inferred regulatory structure. Sequence language models also introduce limitations and modeling choices that matter for GRN applications. First, sequence alone does not encode all determinants of regulation, including chromatin state, TF post-transcriptional modifications, cofactors, and 3D genome organization. Second, regulatory sequences can also be long, requiring practical implementations to require selecting specific regions (e.g., promoter windows, gene bodies, or putative regulatory elements), and choosing how to represent them. Further, a model pretrained on broad corpora may capture general biochemical constraints but still require task-specific fine-tuning to predict regulatory interactions accurately in a particular organism or cell type. These considerations motivate careful 50

dataset construction and evaluation when using language-model predictions as priors. Overall, sequence language models provide a principled way to transform raw biological sequence into informative priors for GRN inference. Sequence models complement probabilistic latent-variable models by supplying TF-specific structure that improves identifiability and interpretability, while enabling prior construction at genome scale even when direct experimental regulatory annotations are sparse. 2.4.4.6

Transformers and Encoders

Sequence language models in genomics are most commonly implemented using Transformer encoders [Ji et al. 2021; Dalla-Torre et al. 2024]. The starting point is a biological sequence (DNA, RNA, or protein), which is first broken into a sequence of discrete tokens. These tokens may be individual bases or amino acids, or short subsequences such as 𝑘-mers. After tokenization, the input can be written as a length-𝐿 sequence 𝑥 = (𝑥 1 . . . , 𝑥 𝐿 ), where each 𝑥𝑖 is a token identity drawn from a finite vocabulary. As token identities are just integers, the model first maps each token to a vector representation using an embedding layer. This produces a matrix of token embeddings with one vector per position in the sequence. A positional signal is then added so the model can distinguish token order, yielding an initial representation, 𝐻 (0) = Embed(𝑥) + PosEnc(𝑥).

(2.16)

Intuitively, 𝐻 (0) encodes both what each token is and where it occurs in the sequence. The Transformer encoder then applies a stack of 𝐿𝑒𝑛𝑐 layers that repeatedly update these token vectors by incorporating contexts from the rest of the sequence,   𝐻 (ℓ) = 𝑓ℓ 𝐻 (ℓ−1) ,

ℓ = 1, . . . , 𝐿𝑒𝑛𝑐 .

(2.17)

Each layer uses self-attention to decide, for every position 𝑖, which other positions in the se51

quence are most informative for updating its representation. In other words, the representation of token 𝑥, is refined by taking a learned, weighted combination of information from tokens throughout the sequence. Repeating this process across layers yields contextualized token embeddings 𝐻 (𝐿𝑒𝑛𝑐 ) ∈ R𝐿×𝑑 , where each row is a 𝑑-dimensional vector summarizing a position in the sequence in the context of its surrounding and distal sequence. Self-attention is particularly relevant in regulatory genomics because the regulatory impact of a motif often depends on its surrounding context and on other motifs located elsewhere in the sequence. For many downstream tasks, the encoder output must be converted into a fixed-dimensional representation. A common approach is to prepend a special classification token, often denoted <cls>, and use its final hidden state ℎ cls ∈ R𝑑 as a summary of the full input. Alternative pooling strategies include mean pooling over token embeddings or attention pooling. In Chapter 4, we leverage this encoder-to-representation mapping to construct sequence-derived features that can be used to score candidate TF-gene interactions and to define prior structure for downstream probabilistic inference. 2.4.4.7

Binary Classification and Cross-Entropy Loss

The encoder described above produces a fixed-dimensional representation of an input sequence (or sequence pair) that can be used as input to a downstream predictor. In the GRN setting, our goal is to assign a score to a candidate TF-gene pair and interpret this score as evidence for whether a regulatory interaction is present. This can be formalized as a binary classification problem. Given an input representation ℎ ∈ R𝑑 (for example, the pooled encoder embedding for a TF-gene pair), predict a label 𝑦 ∈ {0, 1} indicating the absence (𝑦 = 0) or presence (𝑦 = 1) of regulation. The simplest classifier maps the representation ℎ to a scalar logit 𝑠 ∈ R using a linear layer, 𝑠 = 𝑤 ⊤ℎ + 𝑏,

52

(2.18)

where 𝑤 ∈ R𝑑 and 𝑏 ∈ R are learnable parameters. The logit is converted into a probability using the logistic sigmoid, 𝑝 (𝑦 = 1|ℎ) = 𝜎 (𝑠) =

1 . 1 + exp(−𝑠)

(2.19)

Equivalently, one can predict a two-dimensional logit vector 𝑧 ∈ R⊭ with components (𝑧 +, 𝑧 − ) and apply a softmax, 𝑝 (𝑦 = 1|ℎ) =

exp(𝑧 + ) . exp(𝑧 + ) + exp(𝑧 − )

(2.20)

Both parameterizations produce a scalar probability 𝑝 ∈ [0, 1] that can be used as an interaction score. Model parameters are typically learned by minimizing the cross-entropy loss. For a single labeled example (ℎ, 𝑦) with predicted probability 𝑝 = 𝑝 (𝑦 = 1|ℎ), the binary cross-entropy loss is, L = −𝑦 log 𝑝 − (1 − 𝑦) log(1 − 𝑝).

(2.21)

In many TF-gene interaction datasets, positive regulatory edges are substantially rarer than negatives. A common adjustment is to use a class-weighted variant of cross-entropy to increase the penalty for mistakes on the minority class,

L = −𝑤 +𝑦 log 𝑝 − 𝑤 − (1 − 𝑦) log(1 − 𝑝),

(2.22)

where 𝑤 +, 𝑤 − ≥ 0 are user-specified weights. This loss retains the same probabilistic interpretations as standard cross-entropy while providing a simple mechanism to control the precisionrecall tradeoff under class imbalance. In summary, the components in this section define a coherent pipeline from data to regulatory structure. We begin by introducing latent variable models, which provide a compact representation of expression that separates cell-specific regulatory states from gene-specific responses. Building on this view, probabilistic matrix factorization reframes factorization as a generative 53

model, allowing us to quantify uncertainty and incorporate prior knowledge in a principled way. Since exact posterior inference under these models is typically intractable, we then turn to variational inference, which casts inference as an optimization problem and yields practical posterior approximations. Finally, we introduce sequence language models and standard binary classification objectives as an approach for prior knowledge construction. Here we allow learned sequence representations to be converted into TF-gene interaction scores and then assembled into a prior network that anchors downstream inference. Together, these tools support two main methodological threads of this thesis, combining uncertainty-aware probabilistic inference from noisy expression data, with the construction of informative priors to constrain and stabilize this inference.

2.4.5

Evaluating GRN Inference

Evaluating GRN inference methods requires comparing an inferred set of TF-target gene edges against a reference network that acts as a proxy ground truth. In practice, these reference networks are typically assembled from experimentally supported interactions curated in databases, aggregated from the literature, or derived from targeted assays such as ChIP-seq, perturbation experiments, or reporter measurements. Given an inferred network that assigns each candidate TF-gene pair either a binary prediction or continuous score, evaluation proceeds by aligning the predicted edges to the reference set and measuring how well high-confidence predictions recover known interactions while avoiding edges not supported by the reference. 2.4.5.1

Metrics

Most GRN inference methods produce ranked edge scores rather than a single hard network, allowing evaluation to be framed as a binary classification problem over all candidate TF-gene pairs. Let 𝑦 ∈ {0, 1} denote the reference label for a candidate edge, where 1 defines an edge as present in the reference network and 0 defines an edge is not present. Further, let 𝑠ˆ denote 54

the inferred score for that edge. Choosing a threshold 𝜏 converts scores into binary predictions 𝑦ˆ (𝜏) = I[ˆ𝑠 ≥ 𝜏]. For any fixed threshold, predictions can be summarized by the confusion matrix counts for true positives (TP), false positives (FP), false negative (FN), and true negatives (TN). These quantities form the basis of common evaluation metrics. GRN inference is typically a highly imbalanced prediction problem, where true regulatory edges are rare relative to the number of possible TF-gene pairs. Metrics that explicitly account for class imbalance and focus on the quality of positive predictions are therefore especially informative. We start by defining two useful metrics, precision and recall (2.23).

Precision =

TP , TP + FP

Recall =

TP . TP + FN

(2.23)

Precision measures how reliable predicted edges are by determining how many predicted positives are supported by the reference, while recall measures how completely the method recovers reference edges. To combine these two quantities into a single summary of performance, the F1 score 2.24 summarizes the precision-recall tradeoff as their harmonic mean, Precision · Recall Precision + Recall 2TP = . 2TP + FP + FN

F1 = 2 ·

(2.24) (2.25)

When methods output continuous scores, precision and recall vary with the threshold 𝜏. This motivates evaluation across thresholds using the precision-recall (PR) curve and its area, the area under the PR curve (AUPRC). AUPRC is widely used for GRN inference because it emphasizes performance on the positive class and remains informative under severe imbalance, where many other metrics may appear deceptively high. In addition to positive-class performance, it is sometimes useful to report class-specific error

55

rates. The negative analogs are,

Negative Precision =

TN , TN + FN

Negative Recall =

TN . TN + FP

(2.26)

These metrics (2.26) quantify, respectively, the reliability of predicted negatives and the ability to avoid false positive edges. A single threshold summary that uses all four confusion matrix entries is the Matthews correlation coefficient (MCC),

MCC = √︁

TP · TN − FP · FN . (TP + FP)(TP + FN)(TN + FP)(TN + FN)

(2.27)

MCC (2.27) can be interpreted as a correlation between predicted and true labels, and it is often preferred over accuracy in imbalanced settings because it penalizes trivial solutions, such as predicting all edges as negative. Another metric commonly used for GRN inference evaluation is the receiver operating characteristic (ROC) curve, which plots the true positive rate (TPR) against the false positive rate (FPR) as the threshold varies,

TPR =

TP , TP + FN

FPR =

FP . FP + TN

(2.28)

The area under the ROC curve (AUC-ROC) summaries ranking performance across thresholds and has a probabilistic interpretation. It represents the probability that a randomly chosen positive edge is scored higher than a randomly chosen negative edge. However, because the negative class is typically very large in GRN inference, AUC-ROC can remain high even when precision is low. This further motivates using AUPRC instead for a more discriminative metric when comparing methods in practice. An important consideration is that evaluating GRNs is a fundamentally challenging task. First, 56

reference networks are imperfect proxies for truth. Second, curated databases and literaturederived references are incomplete, with many true regulatory edges missing because they have not been tested or reported, creating apparent false positives when methods predict real interactions not present in the reference. Conversely, reference networks can contain false positives due to context mismatch, for example, evidence from a different cell type or condition. Further, negative edges are rarely experimentally established. Edges not present in a reference are often better viewed as unknown rather than truly absent. These issues are especially pronounced outside model organisms, where gold standards are sparse, uneven across TFs and genes, and biased towards well studied regulators. For these reasons, quantitative metrics such as AUPRC, F1 score (2.24), precision and recall (2.23), MCC (2.27), and AUROC are best interpreted as measures of agreement with an available reference, not definitive measures of biological truth. They remain essential for benchmarking, providing a consistent way to compare methods, training paradigms, and priors, while the limitations of reference networks motivate complementary analyses when drawing biological conclusions. 2.4.5.2

Uncertainty Quantification and Calibration

A key advantage of probabilistic GRN inference is that it can provide not only a point estimate for each TF-target gene interaction, but also the model’s uncertainty in that estimate. In this setting, uncertainty reflects how strongly the observed expression matrix 𝑊 constrains a candidate regulatory effect under the assumed generative model. Probabilistic matrix factorization infers a posterior distributions over latent variables and interaction parameters. Under variational inference, this posterior is approximated by a tractable distribution 𝑞(𝑧), which yields both posterior means and posterior variances. For each TF-gene interaction parameter, such as an element of the regulatory component of 𝑉 or a derived edge weight, we summarize the interaction by a point

57

estimate 𝜃ˆ𝑡,𝑔 = E𝑞 [𝜃 𝑡,𝑔 ] and quantify uncertainty with the corresponding posterior variance, 2 2 Var𝑞 (𝜃 𝑡,𝑔 ) = E𝑞 [𝜃 𝑡,𝑔 ] − E𝑞 [𝜃 𝑡,𝑔 ] .

(2.29)

Low posterior variance indicates that the interaction is consistently supported across plausible latent configurations under the model, while high variance indicates that the data permit a wider range of interaction strengths. To assess whether posterior uncertainty is meaningful for prioritization, calibration can be evaluated by testing whether low-uncertainty interactions are more likely to agree with a reference network than high uncertainty interactions. For example, TF-gene interactions can be ranked by their posterior variances to construct 10 cumulative bins corresponding to the lowest 10%, 20%, 30%, and so on of variance. For each bin, an overlap AUPRC can be computed using only those interactions that (𝑖) fall within the bin, and (𝑖𝑖) are present in the gold standard, ensuring that the evaluation is performed on edges that can be scored against the reference. Under good calibration, bins containing lower posterior variance interactions should exhibit higher overlap AUPRC than bins that include progressively more uncertain interactions, indicating that posterior uncertainty provides an informative ranking of interaction reliability.

58

3 | Probabilistic Matrix Factorization for Gene Regulatory Network Inference This chapter is a reprint of the published paper "PMF-GRN: a variational inference approach to single-cell gene regulatory network inference using probabilistic matrix factorization". Claudia Skok Gibbs, Omar Mahmood, Richard Bonneau, and Kyunghyun Cho. Genome Biology, 2024.

3.1

Abstract

Inferring gene regulatory networks (GRNs) from single-cell data is challenging due to heuristic limitations. Existing methods also lack estimates of uncertainty. Here we present Probabilistic Matrix Factorization for Gene Regulatory Network Inference (PMF-GRN). Using single-cell expression data, PMF-GRN infers latent factors capturing transcription factor activity and regulatory relationships. Using variational inference allows hyperparameter search for principled model selection and direct comparison to other generative models. We extensively test and benchmark our method using real single-cell datasets, and synthetic data. We show that PMF-GRN infers GRNs more accurately than current state-of-the-art single-cell GRN inference methods,

59

offering well-calibrated uncertainty estimates.

3.2

Introduction

An essential problem in systems biology is to extract information from genome wide sequencing data to unravel the mechanisms controlling cellular processes within heterogeneous populations [Hecker et al. 2009]. Gene regulatory networks (GRNs) that annotate regulatory relationships between transcription factors (TFs) and their target genes [Chai et al. 2014] have proven to be useful models for stratifying functional differences between cells [Karlebach and Shamir 2008; Äijö and Lähdesmäki 2009; Nachman et al. 2004; Burdziak et al. 2019] that can arise during normal development [Allaway et al. 2021], responses to environmental signals [Jackson et al. 2020] and dysregulation in the context of disease [Ciofani et al. 2012; Ji et al. 2019; Yosef et al. 2013]. GRNs cannot be directly measured with current sequencing technology. Instead, methods must be developed to piece together snapshots of transcriptional processes in order to reconstruct a cell’s regulatory landscape [Mercatelli et al. 2020]. Initial approaches to GRN inference relied on Microarray technology [Huynh-Thu et al. 2010; Wang et al. 2006; Chang et al. 2008], a hybridization-based method to measure the expression of thousands of genes simultaneously [Dufva 2009]. This technology was biased as it was limited to only those genes that were annotated at the time, which in turn presented challenges for inferring the complete regulatory landscape [Hecker et al. 2009]. Subsequently, the high-throughput sequencing method RNA-seq provided a genome wide readout of transcriptional output, allowing for the detection of novel transcripts [Wang et al. 2009] and thus improving GRN inference potential. More recently, singlecell RNA-seq technology has enabled the characterization of gene expression profiles within heterogeneous populations [Saliba et al. 2014], vastly increasing the potential for GRN inference algorithms [Akers and Murali 2021; Lähnemann et al. 2020]. In contrast to bulk RNA experiments (Microarray and RNA-seq) that average measurements of gene expression across heterogenous

60

cell populations, GRNs inferred from single-cell data have the advantage of unmasking biological signal in individual cells [Chen et al. 2019a]. Several matrix factorization approaches have been proposed to overcome the limitations of reconstructing GRNs from Microarray data [Ochs and Fertig 2012]. These include use of statistical techniques such as Singular Value Decomposition and Principal Component Analysis [Alter et al. 2000], Bayesian Decomposition [Moloshok et al. 2002], and Non-negative Matrix Factorization [Kim and Tidor 2003; Brunet et al. 2004; Gao and Church 2005]. More recently, matrix factorization approaches have been applied to integrative analysis of DNA methylation and miRNA expression data [Yang and Michailidis 2016], as well as single-cell RNA-seq and single-cell ATACseq data [Duren et al. 2018]. However, to the best of our knowledge, these matrix factorization approaches have not yet been used to infer GRNs from single-cell gene expression data. Meanwhile, several regression-based methods have been proposed to learn GRNs from single-cell RNA-seq and single-cell ATAC-seq to capture regulatory relationships at single-cell resolution [Hu et al. 2020]. So far, these integrative approaches to GRN inference have been successfully implemented using regularized regression [Skok Gibbs et al. 2022], self-organizing maps [Jansen et al. 2019], tree-based regression [Van de Sande et al. 2020], and Bayesian Ridge regression [Kamimoto et al. 2023]. Although regression-based methods for inferring GRNs from single-cell data are available, they still suffer from significant limitations [Äijö and Bonneau 2016]. Firstly, these methods are designed for specific input datasets, such as bulk or single-cell RNA-seq, causing issues when new data becomes available or new assumptions are required in the model. This can result in inaccurate predictions if the new data or assumptions are not well integrated into the existing model, leading to the need for a complete re-design of the algorithm, which can be costly and time-consuming. Additionally, these methods typically focus on inferring a single GRN that explains the available data, without performing hyperparameter search to determine the optimal model. This can lead to heuristic model selection, with no justification for the approach taken or 61

evidence that the best possible model has been selected. Conversely, hyperparameter search ensures the accuracy of the GRN inference algorithm by finding the optimal model that fits the data well while avoiding overfitting. Regression-based GRN inference algorithms that do not perform hyperparameter search may miss important data features or overemphasize irrelevant ones, leading to inaccurate or incomplete models. Moreover, these methods do not provide an indication of their uncertainty about the predictions that they make. Finally, several regression-based GRN inference algorithms struggle to scale optimally to the size of typical single-cell datasets, limiting inference to small subsets of data or requiring enormous amounts of computational time. In this study, we introduce PMF-GRN, a novel approach that uses probabilistic matrix factorization [Mnih and Salakhutdinov 2007] to infer gene regulatory networks from single-cell gene expression and chromatin accessibility information. This approach extends previous methods that applied matrix factorization for GRN inference with Microarray data, to address the current limitations in regression-based single-cell GRN inference. We implement our approach in a probabilistic setting with variational inference, which provides a flexible framework to incorporate new assumptions or biological data as required, without changing the way the GRN is inferred. We also use a principled hyperparameter selection process, which optimizes the parameters of our probabilistic model for automatic model selection. In this way, we replace heuristic model selection by comparing a variety of generative models and hyperparameter configurations before selecting the optimal parameters with which to infer a final GRN. Our probabilistic approach provides uncertainty estimates for each predicted regulatory interaction, serving as a proxy for the model confidence in each predicted interaction. Uncertainty estimates can be useful in the situation where there are limited validated interactions or a gold standard is incomplete. By using stochastic gradient descent (SGD), we perform GRN inference on a GPU, allowing us to easily scale to a large number of observations in a typical single-cell gene expression dataset. Unlike many existing methods, PMF-GRN is not limited by pre-defined organism restrictions, making it widely applicable for GRN inference. 62

To demonstrate the novelty and advantages of PMF-GRN, we apply our method to datasets from Sacchromyces cerevisiae, human Peripheral Blood Mononuclear Cells (PBMCs) and BEELINE. In our first experiment, we apply our method to two single cell gene expression datasets for the model organism S. cerevisiae. We evaluate our model’s performance in a normal inference setting, as well as with cross-validation and noisy data. To assess the accuracy of predicted regulatory interactions, we evaluate all regulatory predictions using Area Under the Precision Recall Curve (AUPRC) against database derived gold standards. Our findings show that the uncertainty estimates are well-calibrated for inferred TF-target gene interactions, as the accuracy of predictions increases when the associated uncertainty decreases. Here, in comparison to three state-of-theart regression-based methods for inferring single cell GRNs, namely the Inferelator [Skok Gibbs et al. 2022], Scenic [Van de Sande et al. 2020], and Cell Oracle [Kamimoto et al. 2023], our method demonstrates an overall improved performance in recovering the true underlying GRN. Additionally, we apply our method to a PBMC dataset and explore the inferred TFA profiles in the context of annotated cell types and specific immune TFs. We investigate regulatory edges in our inferred GRN and find compelling support for our predictions. Lastly, we benchmark our method using six synthetic datasets generated from BEELINE [Pratapa et al. 2020] and demonstrate consistent outperformance of PMF-GRN compared to the baseline.

3.3

Results

3.3.1

The PMF-GRN Model

The goal of our probabilistic matrix factorization approach is to decompose observed gene expression into latent factors, representing TF activity (TFA) and regulatory interactions between TFs and their target genes. These latent factors, which represent the underlying GRN, cannot be measured experimentally, unlike gene expression. We model an observed gene expression matrix

63

𝑊 ∈ R𝑁 ×𝑀 using a TFA matrix 𝑈 ∈ R𝑁>0×𝐾 , a TF-target gene interaction matrix 𝑉 ∈ R𝑀×𝐾 , observation noise 𝜎𝑜𝑏𝑠 ∈ (0, ∞) and sequencing depth 𝑑 ∈ (0, 1) 𝑁 , where 𝑁 is the number of cells, 𝑀 is the number of genes and 𝐾 is the number of TFs. We rewrite 𝑉 as the product of a matrix 𝐴 ∈ (0, 1) 𝑀×𝐾 , representing the degree of existence of an interaction, and a matrix 𝐵 ∈ R𝑀×𝐾 representing the interaction strength and its direction:

𝑉 = 𝐴 ⊙ 𝐵,

where ⊙ denotes element-wise multiplication. An overview of the graphical model is shown in Figure 3.1A. These latent variables are mutually independent a priori, i.e.,

𝑝 (𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) = 𝑝 (𝑈 )𝑝 (𝐴)𝑝 (𝐵)𝑝 (𝜎𝑜𝑏𝑠 )𝑝 (𝑑).

For the matrix 𝐴, prior hyperparameters represent an initial guess of the interaction between each TF and target gene which need to be provided by a user. These can be derived from genomic databases or obtained by analyzing other data types, such as the measurement of chromosomal accessibility, TF motif databases, and direct measurement of TF-binding along the chromosome, as shown in Figure 3.1B (see Methods section for details). The observations 𝑊 result from a matrix product 𝑈𝑉 ⊤ . We assume noisy observations by defining a distribution over the observations with the level of noise 𝜎𝑜𝑏𝑠 , i.e., 𝑝 (𝑊 |𝑈 , 𝑉 = 𝐴 ⊙ 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑). Given this generative model, we perform posterior inference over all the unobserved latent variables; 𝑈 , 𝐴, 𝐵, 𝑑 and 𝜎𝑜𝑏𝑠 , and use the posterior over 𝐴 to investigate TF-target gene interactions. Exact posterior inference with an arbitrary choice of prior and observation probability distributions is, however, intractable. We address this issue by using variational inference [Blei et al. 2017; Ranganath et al. 2014], where we approximate the true posterior distributions with 64

Figure 3.1: (A) PMF-GRN graphical model overview. Input single-cell gene expression 𝑊 is decomposed into several latent factors. Information obtained from chromatin accessibility data or genomics databases is incorporated into the prior distribution for 𝐴. (B) Input experimental data for PMF-GRN includes single-cell RNA-seq gene expression data. Prior-known TF-target gene interactions can be obtained using chromatin accessibility in parallel with known TF motifs, or through databases or literature derived interactions.(C) Hyperparameter selection process is performed for optimal model selection. The provided prior-known network is split into a train and validation dataset. 80% of the prior-known information is used to infer a GRN, while the remaining 20% is used for validation by computing AUPRC. This process is repeated multiple times, using different hyperparameter configurations in order to determine the optimal hyperparameters for the GRN inference task at hand. Finally, using the optimal hyperparameters, a final network is inferred using the full prior and evaluated using an independent gold standard.

65

tractable, approximate (variational) posterior distributions. We minimize the KL-divergence 𝐷 𝐾𝐿 (𝑞∥𝑝) between the two distributions with respect to the parameters of the variational distribution 𝑞, where 𝑝 is the true posterior distribution. This allows us to find an approximate posterior distribution 𝑞 that closely resembles 𝑝. This is equivalent to maximizing the evidence lower bound (ELBO) i.e. a lower bound to the marginal log likelihood of the observations 𝑊 :

log 𝑝 (𝑊 ) ≥ E𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑∼𝑞(𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑) [ log 𝑝 (𝑊 |𝑈 , 𝑉 = 𝐴 ⊙ 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) + log 𝑝 (𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) − log 𝑞(𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑)]

The mean and variance of the approximate posterior over each entry of 𝐴, obtained from maximizing the ELBO, are then used as the degree of existence of an interaction between a TF and a target gene and its uncertainty, respectively. It is important to note that matrix factorization based GRN inference is only identifiable up to a latent factor (column) permutation. In the absence of prior information, the probability that the user assigns TF names to the columns of 𝑈 and 𝑉 in the same order that the inference 1 , is essentially 0 for any reasonable value algorithm implicitly assigns TFs to these columns is 𝐾!

of 𝐾. Incorporating prior-knowledge of TF-target gene interactions into the prior distribution over 𝐴 is therefore essential in order to provide the inference algorithm with the information of which column corresponds to which TF. With this identifiability issue in mind, we design an inference procedure that can be used on any prior-knowledge edge matrices, described in Figure 3.1C. The first step is to randomly hold out prior information for some percentage of the genes in 𝑝 (𝐴) (we choose 20%) by leaving the rows corresponding to these genes in 𝐴 but setting the prior logistic normal means for all entries 66

in these rows to be the same low number. The second step is to carry out a hyperparameter search using this modified prior-knowledge matrix. The early stopping and model selection criteria are both the ‘validation’ AUPRC of the posterior point estimates of 𝐴, corresponding to the held out genes, against the entries for these genes in the full prior hyperparameter matrix. This step is motivated by the idea that inference using the selected hyperparameter configuration should yield a GRN whose columns correspond to the TF names that the user has assigned to these columns. The third step is to choose the hyperparameter configuration corresponding to the highest validation AUPRC and perform inference using this configuration with the full prior. An importance weighted estimate of the marginal log likelihood is used as the early stopping criterion for this step. The resulting approximate posterior provides the final posterior estimate of 𝐴.

3.3.2

Advantages of PMF-GRN

Existing methods almost always couple the description of the data generating process with the inference procedure used to obtain the final estimated GRN [Skok Gibbs et al. 2022; Kamimoto et al. 2023; Van de Sande et al. 2020]. Designing a new model thus requires designing a new inference procedure specifically for that model, which makes it difficult to compare results across different models due to the discrepancies in their associated inference algorithms. Furthermore, this ad hoc nature of model building and inference algorithm design often leads to the lack of a coherent objective function that can be used for proper hyperparameter search as well as model selection and comparison, as evident in [Skok Gibbs et al. 2022]. Heuristic model selection in available GRN inference methods presents the challenge of determining and selecting the optimal model in a given setting. The proposed PMF-GRN framework decouples the generative model from the inference procedure. Instead of requiring a new inference procedure for each generative model, it enables a single inference procedure through (stochastic) gradient descent with the ELBO objective func67

tion, across a diverse set of generative models. Inference can easily be performed in the same way for each model. Through this framework, it is possible to define the prior and likelihood distributions as desired with the following mild restrictions: we must be able to evaluate the joint distribution of the observations and the latent variables, the variational distribution and the gradient of the log of the variational distribution. The use of stochastic gradient descent in variational inference comes with a significant computational advantage. As each step of inference can be done with a small subset of observations, we can run GRN inference on a very large dataset without any constraint on the number of observations. This procedure is further sped up by using modern hardware, such as GPUs. Under this probabilistic framework, we carry out model selection, such as choosing distributions and their corresponding hyperparameters, in a principled and unified way. Hyperparameters can be tuned with regard to a predefined objective, such as the marginal likelihood of the data or the posterior predictive probability of held out parts of the observations. We can further compare and choose the best generative model using the same procedure. This framework allows us to encode any prior knowledge via the prior distributions of latent variables. For instance, we incorporate prior knowledge about TF-gene interactions as hyperparameters that govern the prior distribution over the matrix 𝐴. If prior knowledge about TFA is available, this can be similarly incorporated into the model via the hyperparameters of the prior distribution over 𝑈 . Because our approach is probabilistic by construction, inference also estimates uncertainty without any separate external mechanism. These uncertainty estimates can be used to assess the reliability of the predictions, i.e., more trust can be placed in interactions that are associated with less uncertainty. We verify this correlation between the degree of uncertainty and the accuracy of interactions in the experiments. Overall, the proposed approach of probabilistic matrix factorization for GRN inference is scalable, generalizable and aware of uncertainty, which makes its use much more advantageous com68

pared to most existing methods.

3.3.3

PMF-GRN Recovers True Interactions in Simple Eukaryotes

To evaluate PMF-GRN’s ability to infer informative and robust GRNs, we leverage two singlecell RNA-seq datasets from the model organism Saccharomyces cerevisiae [Jackson et al. 2020; Jariani et al. 2020]. This eukaryote, being relatively simple and extensively studied, provides a reliable gold standard [Tchourine et al. 2018] for assessing the performance of different GRN inference methods. We conduct three experiments to compare the performance of three stateof-the-art GRN inference methods, the Inferelator (AMuSR, BBSR, and StARS) [Skok Gibbs et al. 2022], SCENIC [Van de Sande et al. 2020], and CellOracle [Kamimoto et al. 2023]. Throughout these experiments, each method is provided with the exact same single-cell RNA-seq datasets (GSE125162 [Jackson et al. 2020]: N cells = 38, 225, GSE144820 [Jariani et al. 2020]: N cells = 6, 118, combined: N cells = 44, 343 by M genes = 6, 763), prior-knowledge (M genes = 6, 885 by K TFs = 220) and gold standard (M genes = 993 by K TFs = 98). In the first experiment, we infer GRNs for each of the two yeast datasets and average the posterior means of 𝐴 to simulate a "multi-task" GRN inference approach. Using AUPRC, we demonstrate that PMF-GRN outperforms AMuSR, StARS, and SCENIC, while performing competitively with BBSR and CellOracle (Figure 3.2A). We next combine the two expression datasets into one observation to test whether each method can discern the overall GRN accurately when data is not cleanly organized into tasks. This experiment reveals a substantial performance decrease for BBSR, indicating its dependence on organized gene expression tasks. This finding suggests potential challenges for BBSR in more complex organisms with less well-defined cell types or conditions. For benchmarking purposes we provide two negative controls for each method, a GRN inferred without prior information (No Prior), and a GRN inferred using shuffled prior information (Shuffled Prior). For all methods, these negative controls achieve an expected low AUPRC. It is essential to note that for CellOracle, an experiment with no prior information could not be 69

performed. This is due to the fact that by design, CellOracle cannot learn regulatory edges that are not included in the prior information. In our comparitive GRN inference analysis, we assess the number of edges predicted in common by each algorithm, on the individual S. cerevisiae datasets. We do so by computing the Intersection over Union (IoU) score, filtering each GRN to the top 25% of interactions to remove noisy predictions. Notably, PMF-GRN obtains an IoU score of 15.69%, outperforming alternative algorithms such as SCENIC (3.17%), AMuSR (12.46%), BBSR (14.56%), and StARS (11.78%). The superior performance of PMF-GRN can be attributed to an ability to discern meaningful regulatory interactions, thereby enriching the consensus among predictions. Importantly, our findings underscore a limitation of CellOracle, which achieves an IoU score of 30.28%. This algorithm, while proficient, can only ascertain edges present in the prior-knowledge matrix. Consequently, the two yeast GRNs inferred display high similarity, reflecting an inherent constraint. This characteristic imparts a degree of predictability to CellOracle, limiting its capacity to discover novel interactions beyond the established prior-knowledge. In contrast, PMF-GRNs IoU score is indicative of a more diverse and comprehensive set of common edges. This highlights PMF-GRNs capability to capture nuanced regulatory relationships as a robust and versitle tool for GRN inference. In a second experiment, we implement a 5-fold cross-validation approach to establish a baseline for each model. Cross-validation is crucial for evaluating the generalization ability of machine learning models like PMF-GRN, particularly in predicting TF-target gene interactions with limited data, a common scenario in experimental settings. To streamline the analysis, we combine the two S. cerevisiae single-cell RNA-seq datasets into a single observation matrix. The cross-validation process involves an 80% − 20% split of the gold standard, where a network is inferred using 80% as "prior-known information" and evaluated using the remaining 20%. This process is iterated five times with different random splits to yield meaningful results. We observe that PMF-GRN outperforms SCENIC and CellOracle, while achieving similar performance 70

Figure 3.2: GRN inference in S. cerevisiae. (A) Consensus Network AUPR with a normal prior-knowledge matrix (N): PMF-GRN (red) performance compared to Inferelator algorithms (AMuSR in yellow, BBSR in orange, StARS in green), SCENIC (blue), and CellOracle (purple). Dashed line represents the baseline if expression data is combined. Negative controls: no prior information (NP - black) and shuffled prior information (S - gray). (B) 5-Fold Cross-Validation Baseline: Each dot with low opacity represents one of the five experiments. Colored dots and lines depict the mean AUPR ± standard deviation for each GRN inference method. (C) GRNs inferred with increasing amounts of noise added to the prior. (D) Calibration results on S.cerevisiae (GSE125162 only) dataset. Posterior means are cumulatively placed in bins based on their posterior variances. AUPRC for each of these bins is computed against the gold standard (see Methods section for details)

71

to BBSR and StARS (Figure 3.2B). We note that for this experiment, we are unable to implement the AMuSR algorithm as it is a multi-task inference approach that requires more than one task (dataset). In a third experiment, we evaluate the robustness of each GRN inference method in the presence of noisy prior information. We conduct GRN inference with increasing levels of noise introduced into the prior knowledge. Specifically, the prior information begins with 1% non-zero edges, and we systematically introduce noise to observe the performance of each method. The noise levels are varied from zero noise (original prior, 1% non-zero edges), to 100% noise (resulting in 2% non-zero edges), 250% noise (3.5% non-zero edges), and 500% noise (6% non-zero edges). Our findings, illustrated in Figure 3.2C, reveal that as the noise in the prior information increases, PMF-GRNs AUPRC experiences a slow decline, mirroring the behavior observed in CellOracle. Notably, PMF-GRN consistently outperforms BBSR, StARS, and SCENIC under these noise conditions, showcasing its robustness in accurately inferring GRNs from noisy priors. These results underscore PMF-GRN as one of the most robust approaches in the face of noisy prior information, thereby emphasizing its utility in practical applications. To further emphasize PMF-GRN’s robustness in a diverse number of settings, we perform the following two experiments. In the first experiment, we examine the performance of PMF-GRN using different sizes of downsampled yeast expression (Figure 3.3A). The downsampling procedure involved reducing the expression data to sizes of 80%, 60%, 40%, and 20%, with each size undergoing random sampling five times to generate five distinct datasets per sample size. Remarkably, the AUPRC performance exhibits noteworthy stability across the downsampling variations. Despite the reduction in dataset size, PMF-GRN consistently demonstrates an ability to learn accurate GRNs as evidenced by the sustained AUPRC performance. These findings underscore the robustness of PMF-GRN, suggesting its reliability even under conditions of diminished dataset sizes, a critical consideration for practical applications where data availability may be limited. In a subsequent experiment, we explore the impact of different cross-validation split sizes on 72

hyperparameter tuning for PMF-GRN using the S. cerevisiae prior-knowledge (Figure 3.3B). Four distinct cross-validation splits, ranging from 80% training and 20% validation, to 20% training and 80% validation, were employed. For each split, we conducted a hyperparameter search across five samples, selecting the optimal hyperparameters based on the highest validation AUPRC. We then selected the best overall hyperparameters from each split to learn a GRN on the full dataset, in order to demonstrate the downstream effect of cross validation split choice on GRN inference. Surprisingly, our results revealed that the choice of cross-validation split size had a marginal impact on the overall performance of the inferred GRN. Specifically, the AUPRC values for the full GRN remained nearly unchanged regardless of whether an 80% train and 20% validation or 60% train and 40% validation split where employed. Even with more disparate splits, such as 40% train and 60% validation, or 20% train and 80% validation, the decrease in AUPRC was only minor. This implies that PMF-GRN exhibits robustness in hyperparameter selection, with the algorithm consistently converging to optimal settings across varying cross-validation scenarios. From our experiments on S. cerevisiae data, several key observations emerge. First, PMFGRN consistently outperforms the Inferelator in recovering true GRNs, surpassing two Inferelator algorithms (AMuSR and StARS) and performing similarly to BBSR. Notably, when expression data is not separated into tasks, PMF-GRN outperforms BBSR. In comparison to CellOracle, PMF-GRN demonstrates competitive performance during normal inference and significantly outperforms CellOracle in cross-validation. However, PMF-GRN, in contrast to CellOracle, is not constrained to predicting edges solely within the confines of the prior-knowledge matrix. Furthermore, PMFGRN consistently outperforms SCENIC across all experiments. A second key observation is that our approach addresses the high variance associated with heuristic model selection among different inference algorithms. When implementing the Inferelator on S. cerevisiae datasets under normal conditions, AUPRCs fall within the range of 0.2 to 0.4, showcasing significant variability without a priori information to guide algorithm selection. This diversity among Inferelator algorithms constitutes heuristic model selection, as one cannot 73

predict a priori which algorithm will perform better or discern the reasons behind their divergent performances. In contrast, our method offers reliable results grounded in a principled objective function, delivering competitive performance akin to the best-performing Inferelator algorithm (BBSR) and CellOracle. This underscores the importance of a consistent and robust approach in the face of uncertainty associated with heuristic model selection among disparate algorithms.

Figure 3.3: (A) GRNs inferred by downsampling S. cerevisiae expression data. (B) Hyperparameter search performed on 4 different ratios of cross-validation. Dots represent validation AUPRC from hyperparameter search during cross-validation, triangle represents AUPRC from a GRN learned using the most optimal hyperparameters for each ratio.

To underscore the identifiability issue and affirm the utility of prior-known information, we showcase PMF-GRN’s performance when prior information is unused (e.g., all prior logistic normal means of 𝐴 set to the same low number). This process is replicated for other GRN inference algorithms by providing an empty prior. Additionally, we assess PMF-GRN’s performance when 74

prior-known TF-target gene interaction hyperparameters are randomly shuffled before building the prior distribution for 𝐴. The results, along with those for the Inferelator and CellOracle, indicate the capability of these approaches to accommodate such prior information effectively.

3.3.4

PMF-GRN Provides Well-Calibrated Uncertainty Estimates

Through our inference procedure, we obtain a posterior variance for each element of 𝐴, in addition to the posterior mean. We interpret each variance as a proxy for the uncertainty associated with the corresponding posterior point estimate of the relationship between a TF and a gene. Due to our use of variational inference as the inference procedure, our uncertainty estimates are likely to be underestimates. However, these uncertainty estimates still provide useful information as to the confidence the model places in its point estimate for each interaction. We expect posterior estimates associated with lower variances (uncertainties) to be more reliable than those with higher variances. In order to determine whether this holds for our posterior estimates, we cumulatively bin the posterior means of 𝐴 according to their variances, from low to high. We then calculate the AUPRC for each bin as shown for the GSE125162 [Jackson et al. 2020] S.cerevisiae dataset in Figure 3.2D. We observe that the AUPRC decreases as the posterior variance increases. In other words, inferred interactions associated with lower uncertainty are more likely to be accurate than those associated with higher uncertainty. This is in line with our expectations as the more certain the model is about the degree of existence of a regulatory interaction, the more accurate it is likely to be, indicating that our model is well-calibrated.

75

3.3.5

PMF-GRN Integrates Single Cell Multi-Omic Data for GRN and TFA Inference in Human PBMCs

We next evaluate PMF-GRN’s ability to learn informative GRNs in a human cell line by focusing on Peripheral Blood Mononuclear Cells (PBMCs). PBMCs represent an essential component of the human immune system, and consist of Lymphocytes (CD4 and CD8 T cells, B cells, and Natural Killer cells), Monocytes and Dendritic cells. Unraveling the distinct regulatory landscape of PBMCs is an essential task to provide insight into how these immune cells interact, as well as coordinate to maintain homeostasis and respond effectively to infections. To infer an informative and comprehensive PBMC GRN, we harness information from a large, paired single cell RNA and ATAC-seq multi-omic dataset [Hao et al. 2021]. We adopt a priorknowledge matrix of TF-target gene interactions (M genes = 18, 557 by K TFs = 860) as previously constructed by [Tjärnberg et al. 2023] for GRN inference with this multi-omic dataset. In this work, the ATAC-seq data was used as a regulatory mask for ENCODE-derived TF ChIP-seq peaks. Regulatory associations were established through the Inferelator-Prior package based on the proximity of TFs to their potential target genes within 50kb upstream and 2kb downstream of the gene transcription start site. We integrate this prior knowledge with the raw expression profiles of 11, 909 PBMCs from a healthy donor to infer a global PBMC GRN and analyze the TFA profiles of eight annotated cell types and several families of immune TFs within this cell line. We first investigate whether our predicted TFA clusters into distinct cell-type groups, as annotated by [Hao et al. 2021]. Using UMAP dimensionality reduction, we are able to determine a near clear distinction between each cell type within PBMCs (Figure 3.4A). Interestingly, the TFA profiles for each of the T cell sub-types (CD4 T, CD8 T, and other T cells) are closely grouped together, suggesting that these cell types may have a similar lineage or TFA patterns, and may share common transcriptional programs or regulatory networks. We next explore the activity profiles of specific immune TF families, starting with the family 76

Figure 3.4: GRN and TFA inference in PBMC. (A) UMAP projection of predicted TFA for each annotated PBMC cell type. (B) Predicted IRF2 TFA demonstrates high activity in NK and CD8 T cells. (C) Heat-map dot-plot depicting TFA of selected immune TFs across annotated PBMC cell types. (D) GRN between IRF TFs and their targets. Pink edges indicate literature support for interaction. (E) Heat-map dot-plot indicates ten most highly active TFs for each PBMC cell type. (F) Violin plot demonstrates corresponding distribution of TFA profiles for ten most highly active TFs.

77

of TFs belonging to IRF. In PBMCs, IRF contributes to the activation of immune cells that modulate antiviral immunity. Notably, the UMAP projection for IRF2 indicates a high activity pattern within Natural Killer cells and CD8 T cells (Figure 3.4B). Indeed, IRF2 is essential for the development and maturation of natural killer cells [Persyn et al. 2022], and acts as a CD8 T cell nexus to translate signals from inflammatory tumor microenvironments [Lukhele et al. 2022]. In order to support our predicted TFA for the family of IRF TFs, we additionally investigate the regulatory interactions inferred by PMF-GRN (Figure 3.4C). To do this within a reasonable scale, we first threshold our predicted GRN interactions (described in detail in Methods). Within our thresholded GRN, we predict regulatory edges between IRF1 and the target genes B2M and BTN3A1. IF1 has been documented as a transactivator of B2M [Gobin et al. 2003], while BTN3A1, a defense-related gene, has been found to be upregulated via the IRF1 pathway [Pietz et al. 2017]. Further, we predict that IRF2 also regulates B2M. Supporting evidence demonstrates that IRF2 has been shown to directly bind to genes linked to the interferon response and MHC Class I antigen presentation, including B2M [Mercado et al. 2019]. Finally, we predict regulatory edges between IRF3 and GPR108, RNF5, and TRAF2. GPR108 has been shown to be a regulator of type I interferon responses by targeting IRF3 [Zhao et al. 2023]. Evidence supporting the interaction between IRF3 and RNF5 indicates that RNF5 has an inhibitory effect on the activation of IRF3 [Zhong et al. 2009]. Lastly, TRAF3 has been shown to be a critical component in the activation of IRF3 during the innate immune response to viral infections [Yu et al. 2023]. In addition to the IRF TFs, several other families of TFs, such as SMAD, STAT, GATA, and EGR, collectively play pivotal roles in PBMCs. These roles contribute to a wide spectrum of functions, including antiviral responses (IRF), fine-tuning immune responses (SMAD), immune cell development (GATA), immediate early responses to signals (EGR), and central regulation of T cells, B cells, and Natural Killer cells (STAT). Their coordinated activities orchestrate the complex interplay of immune cells, enabling PBMCs to effectively respond to diverse stimuli and maintain immune homeostasis. 78

Similarly to IRF, we also explore edges in our thresholded PBMC GRN for these immune TFs to identify regulatory edges supported by literature. Of the five families of immune TFs that we investigate, we find supporting literature for 60 regulatory edges predicted by PMFGRN. We provide these literature supported edges, along with their supporting references in Section Supplementary Material for Chapter 3 Table 3.3 and 3.4. Additionally, we provide a graph representation of each immune TF GRN in Figure 3.2. We next explore the TFA profiles of each of these immune TFs within the eight PBMC cell types. In Figure 3.4D, a heat-map dot-plot provides a visual representation of TFA for each immune TF family across the different PBMC cell types. In particular, we observe that within the IRF family, IRF1 is highly active in CD4 T cells. Previous studies have confirmed the pivotal role of IRF1 in CD4+ T cells, where it is essential for promoting the development of TH1 cells through the activation of the Il12rb1 gene [Kano et al. 2008]. Additionally, SMAD5 is predicted as highly active in B cells. SMAD5 is a key component of the TGF-𝛽 signaling pathway, and has been shown to play a crucial role in maintaining immune homeostasis in B cells [Malhotra and Kang 2013]. We provide a UMAP of the TFA profiles for each of these immune TFs Section Supplementary Material for Chapter 3 Figure 3.3. We further explore our predicted TFA profiles from our global PBMC GRN and calculate the ten most active TFs across the eight distinct cell types. For this experiment, we provide a heat-map dot-plot demonstrating the mean TFA value for each of the top TFs, as well as a corresponding violin plot depicting the distributions of these TFA profiles (Figure 3.4E and 3.4F). Visualizing these distinct activity profiles provides a concise and informative snapshot of the predominant TFs contributing significant transcriptional activity within each cell population. For example, within B cells we observe high activity for the TF PAX5. PAX5 is known to play a crucial role in B cell development by guiding the commitment of lymphoid progenitors to the B lymphocyte lineage while simultaneously repressing inappropriate genes and activating B lineage-specific genes [Cobaleda et al. 2007]. 79

For each annotated cell-type in the PBMC dataset, a set of marker genes were provided. From our ten most active TFs per cell-type analysis combined with their edges to target genes from our thresholded GRN, we find that several of these TFs are predicted to regulate marker genes. For example, within Dendritic cells, the marker gene HLA-DQA1 is predicted to be regulated by the TFs SMAD1 and RFX5; the marker gene HLA-DPA1 is predicted to be regulated by ZNF2 and RFX5; and the marker gene HLA-DRB1 is predicted to be regulated by RFX5. Within CD4 T cells, the marker gene LTB is predicted to be regulated by the TF ZNF436. Within Natural Killer cells, the marker gene PRF1 is predicted to be regulated by ZNF626. Finally, within B cells, the marker gene BANK1 is predicted to be regulated by the TFs ZNF792, EBF1, PAX8, and PAX5; and the marker gene HLA-DQA1 is predicted to be regulated by the TF SMAD1. From the predicted edges between a snapshot of highly active TFs and annotated marker genes, we find the following supporting evidence. The regulatory relationship between RFX5 and HLA-DQA1 involves the inability of RFX5 to bind to the proximal promoter region of HLADQA1, potentially due to DNA methylation, hindering the assembly of active regulatory regions [Majumder and Boss 2011]. Additionally, EBF1 orchestrates direct transcriptional regulation of BANK1, leading to the observed downregulation of BANK1 expression [Treiber et al. 2010]. Pairing the intensity (dot-plot) with the distribution (violin plot) of TFA offers a comprehensive view of the key TFs guiding our regulatory networks. This approach illuminates the variability in their activity levels across diverse immune cell populations, providing a nuanced understanding of the transcriptional dynamics in PBMCs. This information can be used to guide insights into the functional specialization and diversity of immune cells within PBMCs. Further, this comparison provides a sound starting point for exploring the commonalities and differences in the transcriptional regulation of various immune cell populations.

80

3.3.6

Evaluating PMF-GRN with BEELINE Synthetic Data

We next evaluated PMF-GRN using synthetic datasets curated from the BEELINE benchmark [Pratapa et al. 2020]. This benchmark provides six synthetic networks, linear (LI), linear long (LL), cycle (CY), bifurcating (BF), trifurcating (TF), and bifurcating converging (BFC). In repetitions of ten, expression datasets of increasing cell sizes (e.g., 𝑛 = 100, 200, 500, 2000 and 5000) were generated by sampling. Using these generated expression datasets, as well as the provided reference GRNs, we inferred 300 GRNs using PMF-GRN (Figure 3.5A). For each of the six synthetic datasets, PMF-GRN outperforms the BEELINE baseline, represented in Figure 3.5A with a black dashed line.

Figure 3.5: PMF-GRN performance on BEELINE synthetic GRN data (A) PMF-GRN inference performance with half of the ground truth provided as prior network information and the remaining half provided as a gold standard for evaluation. Dashed lines are the expected baseline of a random predictor. (B) AUPRC ratio over the baseline random predictor for PMF-GRN in comparison to each of the GRN inference methods used in the original BEELINE benchmark.

To further evaluate PMF-GRN, we calculate the AUPRC ratio of PMF-GRN over the baseline random predictor to compare to the similarly computed ratios in the original BEELINE paper (Figure 3.5B). We observe that for the linear, cycle, and bifurcating converging, PMF-GRN achieves

81

competitive AUPRC ratios in comparison to the original methods used in the BEELINE benchmark. Interestingly, PMF-GRN does not perform competitively on long linear. This could be due to a number of factors, such as the larger number of intermediate genes introducing additional complexity which PMF-GRN struggles to capture. Alternatively, the extended trajectory introduces a higher-dimensional space, which could present a challenge for our matrix factorization based approach to effectively decompose the data into meaningful latent factors. This presents an interesting avenue of consideration when developing future probabilistic matrix factorization approaches for GRN inference.

3.4

Methods

3.4.1

Model Details

We index cells, genes and TFs using 𝑛 ∈ {1, · · · , 𝑁 }, 𝑚 ∈ {1, · · · , 𝑀 } and 𝑘 ∈ {1, · · · , 𝐾 }, respectively. We treat each cell’s expression profile 𝑊𝑛 as a random variable, with local latent variables 𝑈𝑛 and 𝑑𝑛 , and global latent variables (that are shared among all cells) 𝜎𝑜𝑏𝑠 and 𝑉 = 𝐴 ⊙ 𝐵. We use the following likelihood for each of our observations: 2 𝑝𝑊 (𝑊𝑛 | 𝑈 , 𝑉 , 𝜎𝑜𝑏𝑠 , 𝑑) = N (𝑑𝑛 ∗ 𝑈𝑛𝑉 ⊤, 𝜎𝑜𝑏𝑠 ).

We assume that 𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 and 𝑑 are independent i.e.,

𝑝 (𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) = 𝑝𝑈 (𝑈 )𝑝𝐴 (𝐴)𝑝 𝐵 (𝐵)𝑝𝜎 (𝜎𝑜𝑏𝑠 )𝑝𝑑 (𝑑).

In addition to our i.i.d assumption over the rows of 𝑈 and 𝑑, we also assume that the entries of 𝑈𝑛 are mutually independent, and that all entries of 𝐴 and 𝐵 are mutually independent. We choose a lognormal distribution for our prior over 𝑈 and a logistic Normal distribution for our prior over

82

𝑑: 𝑝𝑈 (log(𝑈𝑛𝑘 )) = N (𝜇𝑢 , 𝜎𝑢2 ), 𝑝𝑑 (logit(𝑑𝑛 )) = N (0, 9) where 𝜇𝑢 ∈ R and 𝜎𝑢 ∈ R+ . We use a logistic Normal distribution for our prior over 𝐴, a Normal distribution for our prior over 𝐵 and a logistic Normal distribution for our prior over 𝜎𝑜𝑏𝑠 : 𝑝𝐴 (logit(𝐴𝑚𝑘 )) = N (logit(clip(𝐴¯𝑚𝑘 , 𝑎 max, 𝑎 min )), 𝜎𝑎2 ), 𝑝 𝐵 (𝐵𝑚𝑘 ) = N (0, 𝜎𝑏2 ). 𝑝𝜎 (log(𝜎𝑜𝑏𝑠 )) = N (0, 1), where 𝐴¯𝑚𝑘 ∈ {0, 1}, 𝑎 max, 𝑎 min ∈ (0, 1), 𝜎𝑎 , 𝜎𝑏 ∈ R>0,

and  clip(𝐴¯𝑚𝑘 , 𝑎 max, 𝑎 min ) = max min(𝐴¯𝑚𝑘 , 𝑎 max ), 𝑎 min . Here, 𝐴¯𝑚𝑘 is provided by a prior-knowledge pipeline also used by methods such as the Inferelator. The pipeline leverages ATAC-seq and TF binding motif data to provide binary initial guesses of gene-TF interactions. 𝑎 max and 𝑎 min are hyperparameters that determine how we clip these binary values before transforming them to the logit space. 83

For our approximate posterior distribution, we enforce independence as follows:

𝑞(𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) = 𝑞𝑈 (𝑈 )𝑞𝐴 (𝐴)𝑞𝐵 (𝐵)𝑞𝜎 (𝜎𝑜𝑏𝑠 )𝑞𝑑 (𝑑).

We impose the same independence assumptions on each approximate posterior as we do for its corresponding prior. Specifically, we use the following distributions: 𝑞𝑈 (log(𝑈𝑛𝑘 )) = N (𝑈˜𝑛𝑘 , 𝜎˜𝑈2𝑛𝑘 ) 𝑞𝑑 (logit(𝑑𝑛 )) = N (𝑑˜𝑛 , 𝜎˜𝑑2𝑛 ) 𝑞𝐴 (logit(𝐴𝑚𝑘 )) = N (𝐴˜𝑚𝑘 , 𝜎˜ 𝐴2 𝑚𝑘 ) 𝑞𝐵 (𝐵𝑚𝑘 ) = N (𝐵˜𝑚𝑘 , 𝜎˜ 𝐵2𝑚𝑘 ) ˜ 𝜎˜𝑜2 ), 𝑞𝜎 (log(𝜎𝑜𝑏𝑠 )) = N (𝑜, where the parameters on the right hand sides of the equations are called variational parameters: 𝑈˜𝑛𝑘 , 𝑑˜𝑛 , 𝐴˜𝑚𝑘 , 𝐵˜𝑚𝑘 , 𝑜˜ ∈ R and 𝜎˜𝑈𝑛𝑘 , 𝜎˜𝑑𝑛 , 𝜎˜ 𝐴𝑚𝑘 , 𝜎˜ 𝐵𝑚𝑘 , 𝜎˜𝑜 ∈ R+ . To avoid numerical issues during optimization, we place constraints on several of these variational parameters.

3.4.2

Inference

We perform inference on our model by optimizing the variational parameters to maximize the ELBo. In doing so, we minimise the KL-divergence between the true posterior and the variational posterior. In practice, to help with addressing the latent factor identifiability issue, we use a modified version of the ELBo where the prior and posterior terms are weighted by a constant

84

𝛽 ≥ 1 [Higgins et al. 2017]:

E𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑∼𝑞(𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑) [ log 𝑝 (𝑊 |𝑈 , 𝑉 = 𝐴 ⊙ 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) + 𝛽 (log 𝑝 (𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) − log 𝑞(𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑))]

Inference is carried out using the Adam optimizer with learning rate 0.1 and beta values of 0.9 and 0.99. We clip gradient norms at a value of 0.0001. We set 𝑎 min = 0.005, 𝑎 max = 0.995, 𝜎𝑏2 = 1 and 𝜇𝑢 = 0. We vary 𝜎𝑎 and 𝜎𝑢 as hyperparameters that control the strengths of the priors over 𝐴 and 𝑈 , respectively. We also vary 𝛽 as a hyperparameter. We choose a hyperparameter configuration using validation AUPRC as the objective function as well as the early stopping metric. We hold out hyperparameters for 𝑝 (𝐴) for a fraction of the genes. We do this by setting 𝐴¯𝑚𝑘 = 0 for 𝑚 corresponding to these genes for all 𝑘. During inference we regularly obtain posterior point estimates for these entries and measure the AUPRC against the original values of these entries as given in the full prior. This quantity is known as the validation AUPRC. Once we have picked the hyperparameter configuration corresponding to the best validation AUPRC, we perform inference with this model using the full prior without holding out any information. We use an importance weighted estimate of the marginal log likelihood as our early stopping criterion:  log 𝑝 (𝑊 ) = log E𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑∼𝑞(𝑈 ,𝐴,𝐵,𝜎𝑜𝑏𝑠 ,𝑑)



𝑝 (𝑊 |𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑)𝑝 (𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑) 𝑞(𝑈 , 𝐴, 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑)

 ,

Í where the expectation is computed using simple Monte Carlo and the log- -exp trick is used to avoid numerical issues.

85

3.4.3

Computing Summary Statistics for the Posterior

After training the model, we use 𝐴˜ and 𝜎˜ 𝐴 , the variational parameters of 𝑞(𝐴), to obtain a mean and a variance for each entry of 𝐴. Since 𝑞(𝐴) is logistic normal, it admits no closed form solution for the mean and variance. We therefore use Simple Monte Carlo i.e. we sample each entry of 𝐴 several times from its posterior distribution and then compute the sample mean and sample variance from these samples. We use each mean as a posterior point estimate of the probability of interaction between a TF and a gene, and its associated variance as a proxy for the uncertainty associated with this estimate.

3.4.4

Calculating AUPRC

The gold standards for the datasets used in this paper do not necessarily perfectly overlap with the genes and TFs that make up the rows and columns of 𝐴 as defined by the prior hyperparameters i.e. there may be genes and TFs in the gold standard with a recorded interaction or lack of interaction, that do not appear in our model at all because they are not present in the prior. The reverse is also true: the prior may contain genes and TFs that are not in the gold standard. For this reason, we compute the AUPRC using one of two methods: ‘keep all gold standard’ or ‘overlap’, which correspond to evaluating only interactions that are present in the gold standard or only interactions that are present in both the gold standard and the prior/posterior. We present results with ‘keep all gold standard’ AUPRC as the evaluation metric when comparing our model to the Inferelator in Figure 3.2. For our evaluation of uncertainty calibration (Figure 3.2D), we use the overlap AUPRC so that bins containing a lower number of posterior means do not have artificially deflated AUPRCs (see the Evaluating Calibration of Posterior Uncertainty part of the Methods Section for further information).

86

3.4.5

Evaluating Calibration of Posterior Uncertainty

We create 10 bins, corresponding to the lowest 10%, 20%, 30% and so on of posterior variances. We place the posterior point estimates of TF-gene interactions associated with these variances into these bins and then calculate the ‘overlap AUPRC’ for each bin using the corresponding gold standard. The AUPRC for each bin is calculated using those interactions that are in the gold standard and also in the bin. We use such a cumulative binning scheme because using a noncumulative scheme could result in some bins having very small numbers of posterior interactions that are present in the gold standard, which would lead to noisier estimates of the AUPRC.

3.4.6

Inference and Evaluation on Multiple Observations of W

The Inferelator method applies two scRNA-seq experiments separately on S. cerevisiae, with each resulting in a distinct model. These models are used to infer TF-gene interaction matrices, which are then sparsified. The final matrix is obtained by taking the intersection of the two matrices and retaining only the entries that are non-zero in both matrices. In our approach, we also train a separate model on each expression matrix, and obtain a posterior mean matrix for 𝐴 for each of them. To obtain the final posterior mean matrix for 𝐴, we average the posterior mean matrices from each model. While this approach works well, future research could focus on explicitly modeling separate expression matrices within the model, as mentioned in the Discussion section.

3.4.7

Measuring the Impact of Prior Hyperparameters

We evaluate the utility of each of the prior hyperparameter matrices used in our experiments. In Figures 3.2A and Section Supplementary Material for Chapter 3 Figure 3.1, we present with grey dots the AUPRCs achieved when performing inference using shuffled prior hyperparameters for 𝐴. This corresponds to randomly assigning to each row (gene) of 𝐴, the prior hyperparameters that correspond to a different row of 𝐴. Shuffling the hyperparameters should lead to worse 87

performance, as the posterior estimates should then also be shuffled, whereas the row/column labels for the posterior will remain unshuffled. For the ‘no prior’ setting, shown with black dots in the figures, we set 𝐴¯𝑚𝑘 = 0 ∀ 𝑚, 𝑘. The difference in AUPRC achieved using the unshuffled vs shuffled or no hyperparameters measures the usefulness of the provided hyperparameters for the inference task on the dataset in question.

3.4.8

Cross-Validation

For S. cerevisiae, we perform a five-fold cross validation experiment (Figure 3.2B). Cross-validation is performed by partitioning the gold standard into an 80% - 20% split, where 80% of the data represents prior-known information to be used as a prior for 𝑝 (𝐴), and the remaining 20% is treated as the gold standard for evaluation. This process is repeated five times to generate five random splits of the data in order to robustly evaluate GRN inference. It is important to note that PMF-GRN performs hyperparameter search before inferring a final GRN within each crossvalidation split. For each of the five partitioned cross-validation folds the 80%, or prior portion, is further split into 80% train and 20% test for hyperparameter search and evaluation. Once the optimal hyperparameters have been determined, the initial 80% split is treated as the training data, while the remaining 20%, which was not seen during hyperparameter selection, is used for evaluation.

3.4.9

Intersection over Union

Intersection over Union (IoU) scores were computed using the GRN learned by each algorithm for the two S. cerevisiae expression datasets. For each GRN, we calculate and retain the top 25% of predicted edges in order to obtain the best estimates for each algorithm and elimate noisier predictions. For each algorithm, we compute both the intersection and the union of the GRN interactions predicted from the two S. cerevisiae datasets. Dividing the Intersection by the Union

88

allows us to obtain a score indicating how similar the two inferred GRNs are for each algorithm.

3.4.10

Downsampling Expression

For S. cerevisiae, in repetitions of five, we randomly sample the S. cerevisiae expression matrix on the cell axis to obtain downsampled expression dataset sizes of 80%, 60%, 40%, and 20%. We perform a hyperparameter search, using an 80% training - 20% validation split of the prior-knowledge matrix, on each of these five expression matrices for each sample size. Using these hyperparameters, we infer GRNs for each repetition within each split to obtain our final downsampled GRNs.

3.4.11

Exploring the Effect of Cross-Validation Ratios on Hyperparamater Selection

To effectively explore the influence of cross-validation split size on obtaining optimal hyperparameters for GRN inference, we methodically separate our S. cerevisiae prior-knowledge into 4 different split sizes. These splits consist of 80% training - 20% validation, 60% training - 40% validation, 40% training - 60% validation, and 20% training - 80% validation. For each split size, we obtain 5 random training and validation splits to ensure robust results. We then perform hyperparameter search across each 5 random splits for each split size. Using the best overall hyperparameters for each split size, we infer a final GRN to demonstrate the impact each particular split had on obtaining the optimal hyperparameters for the final GRN.

3.4.12

Datasets and Preprocessing

We inferred each GRN using a single-cell RNA-seq expression matrix, a TF-target gene connectivity matrix, and a gold standard for bench-marking purposes. We modeled the single-cell expression matrices based on the raw UMI counts obtained from sequencing for the S. cerevisiae and PBMC datasets, which were therefore not normalized for the purpose of this work. For the 89

two B. subtilis datasets used in this work, we demonstrate the effect of different normalization and scaling techniques, and convert all data used to integers in order to create a single-cell-like dataset. We further obtained binary TF-gene matrices representing prior-known interactions, which served as prior hyperparameters over A, and were derived from the YEASTRACT and subtiwiki databases, as well as from [Tjärnberg et al. 2023] for PBMC. We acquired a gold standard for S. cerevisiae our datasets from independent work which is detailed below. 3.4.12.1

Saccharomyces cerevisiae

We used two raw UMI count expression matrices for the organism S. cerevisiae obtained from NCBI GEO (GSE125162 [Jackson et al. 2020] and GSE144820 [Jariani et al. 2020]). For this well studied organism, we employed the YEASTRACT [Monteiro et al. 2020; Teixeira et al. 2018] literature derived network of TF-target gene interactions to be used as a prior over 𝐴 in both S. cerevisiae networks. A gold standard for S. cerevisiae was additionally obtained from a previously defined network [Tchourine et al. 2018] and used for bench-marking our posterior network predictions.We note that the gold standard is roughly a reliable subset of the YEASTRACT prior. Additional interactions in the prior can still be considered to be true but have less supportive evidence than those in the gold standard. 3.4.12.2

Peripheral Blood Mononuclear Cells

We used a paired multi-omic single cell RNA-seq and ATAC-seq dataset for PBMC obtained from [Hao et al. 2021]. The single-cell expression matrix contained 11,909 cells. The prior-knowledge matrix was constructed using the ATAC-seq data from this multi-omic dataset, constructed and described in detail by [Tjärnberg et al. 2023]. The prior-knowledge matrix is 18, 557 genes by 860 TFs, and contains 0.5% non-zero edges. Due to the complex and dynamic nature of PBMCs, a gold standard is currently unavailable for this cell line. To evaluate our inferred network, we implement a 5-fold cross-validation pro90

cedure where our chromatin accessibility-based prior is split into 5 random sets, where 80% is used as prior knowledge and 20% is used as the gold standard for evaluation. We then took the intersection of the regulatory edges inferred across each of the 5 fold cross-validation experiments, and filtered to retain the highest quality edges, obtaining a prediction probability of 90% or higher. 3.4.12.3

BEELINE Synthetic Datasets

We used the BEELINE synthetic expression datasets [Pratapa et al. 2020] without modification. Reference GRNs were transformed into cross-tab matrices in order to use this information for prior-knowledge and gold standard evaluation. We used 50% of the reference GRN as the prior and the remaining 50% as the gold standard, as was similarly done in [Skok Gibbs et al. 2022].

3.5

Discussion

In this chapter, we introduce a robust framework for probabilistic matrix factorization, optimized through automatic variational inference, to infer GRNs from single-cell gene expression data. A distinctive feature of our approach is the decoupling of the data generation model from the inference procedure, providing unprecedented flexibility. This decoupling allows for modifications to the latent variables and their distributions, without altering the inference process. Such flexibility facilitates the seamless integration of diverse sequencing datasets and modeling assumptions. Unlike previous methods, our framework eliminates the need to define a new inference procedure for each specific dataset or biological context when building new models. PMF-GRN not only offers a flexible and unified approach to GRN inference but also provides a principled methodology for model selection and hyperparameter configuration. The use of a consistent objective function and inference procedure across all generative models streamlines the process of hyperparameter search, reducing ambiguity present in methods like the Inferelator.

91

By conducting hyperparameter search across different generative models, we identify configurations corresponding to optimal values of our objective function, minimizing the reliance on heuristic model selection. To validate the effectiveness of our approach, we applied PMF-GRN to infer GRNs from single cell S. cerevisiae gene expression, comparing results with state-of-the-art single cell GRN inference methods such as the Inferelator, SCENIC and CellOracle. Our method demonstrates competitive, if not superior, performance in terms of AUPRC, in each experiment performed. Here, PMF-GRN provides a stable and reliable inferred GRN without the need for heuristic model selection or data separation into tasks. Cross-validation experiments further support the robustness of PMF-GRN, BBSR, and StARS, indicating their ability to generalize well to new data without overfitting. In contrast, SCENIC and CellOracle exhibited poor performance during cross-validation, suggesting potential issues with generalizability. Notably, we assessed the robustness of each algorithm against increasing noise in the prior-knowledge, identifying PMF-GRN and CellOracle as the most resilient to noisy priors. This resilience ensures the reliability of inferred GRNs even in the presence of uncertain prior knowledge. Our model uniquely provides well-calibrated uncertainty estimates alongside point estimates for each interaction in the final GRN. The evaluation of uncertainty estimates demonstrated that as the posterior variance decreases, the AUPRC increases, indicating that the model is wellcalibrated. Biologists can leverage these uncertainty estimates for downstream experimental validation, placing more trust in estimates with lower posterior variance. Finally, the linear scalability of our models computational cost with the number of cells enables its application to single-cell RNA-seq datasets of any size. Our investigation into PMF-GRN’s application to human PBMCs provides insightful findings into the regulatory landscape of these essential immune cells. Leveraging a comprehensive multiomic dataset, we demonstrate that our approach integrates single cell RNA and well-curated prior 92

knowledge derived from ATAC-seq data. The resulting global PBMC GRN unveils distinct TFA profiles for eight annotated cell types and various immune TF families. Through UMAP dimensionality reduction, we observe clear clustering of TFA profiles. Focusing on the IRF family, we identify specific TF-target gene interactions supported by literature, shedding light on regulatory relationships critical for immune responses. Extension to other immune TF families reveals their orchestrated activities within PBMCs, contributing to antiviral responses, immune cell development, and the regulation of T cells, B cells, and Natural Killer cells. By exploring predicted edges between active TFs and marker genes, we establish connections between regulatory networks and cellular functions. The combined dot-plot and violin plot visualization strategy provides a nuanced understanding of TF activities, offering a valuable resource for deciphering the intricate transcriptional dynamics in PBMCs. This detailed exploration sets the stage for further investigations into the functional specialization and diversity of immune cells within the PBMC population, with implications for advancing our understanding of immune responses and disease mechanisms. In the context of synthetic datasets curated from the BEELINE benchmark, PMF-GRN demonstrates robust performance across various network structures. Outperforming the BEELINE baseline across different synthetic networks, PMF-GRN consistently achieves competitive AUPRC ratios compared to the original methods used in the BEELINE benchmark. Notably, PMF-GRN’s competitive performance is observed in linear, cycle, and bifurcating converging structures. However, challenges arise in the long linear structured synthetic data, suggesting potential limitations in capturing the complex dynamics of extended trajectories. Factors such as the increased number of intermediate genes and a higher-dimensional space may contribute to this limitation. This observation opens avenues for future development of probabilistic matrix factorization approaches, encouraging exploration of methods better suited for intricate network structures. The overall success of PMF-GRN in diverse synthetic network scenarios underscores its versatility and effectiveness in inferring GRNs, promising broad applicability in deciphering complex biological 93

systems and regulatory interactions.

3.6

Conclusion

In conclusion, the PMF-GRN framework provides a flexible and principled approach for inferring GRNs from single-cell gene expression data. By decoupling the model and inference procedure, PMF-GRN enables easy integration of new and various sequencing datasets as well as modeling assumptions without the need for defining a new inference procedure. Additionally, PMF-GRN provides a principled approach for model selection through hyperparameter search, reducing the need for heuristic model selection. Overall, PMF-GRN consistently yields high-performing competitive results compared to other state-of-the-art single cell GRN inference methods with a reliable gold standard, and is robust to cross validation, noisy priors and downsampling. Further, PMF-GRN provides well-calibrated uncertainty estimation, enabling a reliable set of results for downstream experimental validation. We envision many possible directions for future work to design a better algorithm for inferring GRNs under our framework. This framework could be extended to explicitly model multiple expression matrices and their batch effects. We could probabilistically model prior information for 𝐴 obtained from ATAC-seq and TF motif databases, and include this as part of the probabilistic model over which we carry out inference. Evaluating the posterior estimates of the direction of transcriptional regulation, provided by the matrix 𝐵, could provide a useful benchmark for the computational estimation of TF activation and repression. Research could also be carried out on improved self-supervised objectives for hyperparameter selection. Future work could also focus on how to use results from our framework to guide experimental wet-lab work. For example, the uncertainty quantification provided by our model could open up new research directions in active learning for GRN inference. Highly ranked, uncertain interactions could be experimentally tested and the results fed back into the prior hyperparameter

94

matrix for 𝐴. Inference with this updated matrix would ideally yield a better posterior GRN estimate. Posterior estimates of TFA provided by our model could be useful to wet lab scientists, as this quantity provides information about possible post-transcriptional modifications, which are currently challenging to measure experimentally. Most importantly, the study of GRN inference is far from complete. GRN inference approaches have thus far required new computational models and assumptions in order to keep up with relevant sequencing technologies. It is thus essential to develop a model that can be easily adapted to new biological datasets as they become available, without having to completely re-build each model. We have therefore proposed PMF-GRN as a modular, principled, probabilistic approach that can be easily adapted to both new and different biological data without having to design a new GRN inference method.

3.7

Retrospective

Perspective since publication in 2024

PMF-GRN was developed under the pragmatic premise that single-cell expression measurements are informative about regulation, yet they are only indirectly related to the regulatory mechanisms of interest. The PMF-GRN model therefore treats TF activity and TF-gene regulatory effects as latent quantities, and relies on prior knowledge to anchor these latent factors to biologically interpretable TF identities. With the benefit of subsequent work in this thesis, the strongest and most durable contribution of PMF-GRN is not a single likelihood choice or factorization variant, but the framing of GRN inference as a modular probabilistic program in which prior information, uncertainty, and model comparison are explicit, integral components of the modeling framework. The decoupling of the generative model from an automatic variational inference procedure remains the core design decision that enables extensibility, principled hyperparameter

95

selection, and uncertainty estimates at the level of individual TF-gene interactions. Several developments since PMF-GRN’s publication in 2024 have clarified how this framing fits into the broader trajectory of the field. Evaluation has shifted further toward perturbationgrounded benchmarks, motivated by the long recognized limitation that curated reference networks are incomplete, context-dependent, and often aggregate heterogeneous evidence. Resources such as CausalBench [Chevalley et al. 2025] explicitly foreground single-cell perturbation data as a basis for assessing network inference methods in settings that are closer to causal intervention than correlation-based reconstruction. This shift does not diminish the role of probabilistic modeling; rather, it strengthens the case for models that expose calibrated uncertainty and make their assumptions explicit, as perturbation benchmarks often highlight failure modes that are difficult to detect when evaluation relies only on incomplete reference networks. At the same time, GRN inference has continued moving toward multi-omic formulations in which chromatin accessibility is not merely an optional source of prior edges but part of the primary observational data [Yuan and Duren 2025]. Methods designed around paired RNA and ATAC measurements, including neural approaches that integrate atlas-scale external data, reflect an emerging default, where regulatory network reconstruction increasingly relies on combining transcriptional state with information about regulatory potential and binding. In retrospect, PMFGRN’s emphasis on priors derived from chromatin accessibility anticipated this direction, while also making a clear limitation that remains unresolved in many pipelines: accessibility and motif evidence constrain what regulation is plausible, but they do not uniquely determine functional influence on transcription. A probabilistic treatment of prior evidence, including uncertainty in peak-to-gene linking, motif matches, and cell-type specificity, is therefore an important next step for reconciling multi-omic evidence with expression-driven inference. A third development is the rapid growth of representation-learning approaches for cellular systems. Large cell models, trained on tens of millions of cells, are now explicitly positioned as tools for robust gene network prediction [Kalfon et al. 2025]. These cell models suggest that 96

a meaningful fraction of regulatory structure can be learned from scale even before introducing experiment-specific priors or mechanistic assumptions. This trend reframes the role of probabilistic GRN inference. One plausible synthesis is that foundation models provide strong, transferable representations or priors, while probabilistic inference provides the mechanism to (𝑖) adapt those priors to a specific dataset, and (𝑖𝑖) quantify what the data can and cannot resolve about regulatory structure. This perspective aligns naturally with the thesis-wide narrative that follows PMF-GRN: improving the quality and transferability of priors can shift the effective bottleneck in GRN reconstruction, while uncertainty estimates remain central for deciding which predicted interactions warrant experimental follow-up. These changes also sharpen several limitations of PMF-GRN that are important to acknowledge explicitly in this thesis. First, the interpretability of latent factors depends on the availability and quality of prior knowledge, and this dependence is not a defect so much as a statement of identifiability. Expression alone often under-specifies which TF is responsible for a regulatory program, especially when TFs are correlated, when regulation is combinatorial, or when regulatory signals precede the moment of capture. Second, the uncertainty estimates produced by variational inference should be treated as conservative in their scope: they are most reliable as a relative ranking of confidence across edges, while absolute calibration can be affected by posterior approximation. Finally, the BEELINE results suggest regimes where the modeling assumptions are strained (for example, extended trajectories with many intermediate genes), motivating extensions that more directly represent dynamics, time-lagged effects, or state transitions when such information is available. Viewed from the present, PMF-GRN serves as a foundation in two ways. Methodologically, it establishes a probabilistic scaffold for GRN inference with explicit priors, model comparison, and edge-wise uncertainty. Conceptually, it motivates this thesis’s later emphasis on prior construction and transfer-learning, where if expression data primarily refines, contextualizes, and prunes a regulatory scaffold, then learning that scaffold well, either from sequence and other 97

scalable modalities, becomes the central objective. The chapter that follows builds directly on this premise by focusing on how to construct stronger priors and how to integrate them into inference pipelines in ways that preserve interpretability and make uncertainty actionable.

3.8

Supplementary Material for Chapter 3

3.8.1

PMF-GRN Recovers True Interactions in Prokaryotes as Evaluated by Cross-Validation

To demonstrate GRN inference on a forth additional dataset, we carry out experiments using two microarray datasets for the prokaryote Bacillus Subtilis (B1 - GSE27219 [Nicolas et al. 2012] and B2 - GSE67023 [Arrieta-Ortiz et al. 2015]). Although PMF-GRN is not primarily designed to learn GRNs from microarray data, we show that it is still possible to learn informative GRNs with this data. For our B. subtilis experiments, we have access to prior-knowledge derived from the subtiwiki database [Michna et al. 2016; Zhu and Stülke 2018; Pedreira et al. 2022]. Here, we implement a 5 fold cross-validation approach by using five random splits of the subtiwiki database-derived information, where 80% is used as prior knowledge and 20% is used as the gold standard for evaluation. The two B. subtilis datasets were previously normalized after data collection as part of standard microarray processing. However, each dataset was normalized using different approaches (described in Methods). For B1, the expression data underwent no further normalization and was simply converted to integers to simulate single-cell-like data. For B2, the expression data was re-scaled and then converted to integers, in order to contain only positive integers resembling single-cell-like data. The results from our experiments are shown in Figure 3.1, and the numbers used to create this figure are given in Table 3.1 and 3.2. Using five repeats of cross-validation, we show the performance of GRNs inferred for the two B. subtilis datasets (B1 and B2). We remark

98

that the difference in performance between B1 and B2 is likely a result of the chosen microarray processing normalization. To further support this claim, we demonstrate GRN performance after re-scaling the data with min-max scaling (Figure 3.1). ‘No Prior’ and ‘Shuffled’ results are also shown in Figure 3.1 by black and gray dots respectively. Here, we are able to demonstrate that for B1 and B2, each GRN yields a better performance as compared to negative controls.

Figure 3.1: Results for GRNs learned in B. subtilis datasets B1 (GSE27219) and B2 (GSE67023) comparing "No Normalization" to "Min-Max Scaling". Colored dots represent the normal (N) GRN, with the line indicating the mean of the cross-validation experiments ± standard deviation. Negative controls are demonstrated by black dots for GRNs inferred with No Prior (NP) and grey dots for Shuffled Prior (S).

Method No Normalization Min-Max Scaling

B. subtilis Cross-Validation Dataset B1 Regular No Prior Shuffled 0.0509 ± 0.0273 0.0048 ± 0.0003 0.0050 ± 0.0028 0.1931 ± 0.0171 0.0042 ± 0.0003 0.0042 ± 0.0018

Table 3.1: AUPRCs achieved by PMF-GRN on the B. subtilis B1 dataset. Results are reported as the mean AUPRC across five cross-validation splits ± standard deviation.

99

Method No Normalization Min-Max Scaling

B. subtilis Cross-Validation Dataset B2 Regular No Prior Shuffled 0.2508 ± 0.0232 0.0062 ± 0.0006 0.0052 ± 0.0004 0.2886 ± 0.0312 0.0048 ± 0.0008 0.0038 ± 0.0005

Table 3.2: AUPRCs achieved by PMF-GRN on the B. subtilis B2 dataset. Results are reported as the mean AUPRC across five cross-validation splits ± standard deviation.

3.8.2

Peripheral Blood Mononuclear Cells

For GRN inference in PBMC, we perform a hyperparameter search using a 5 fold cross-validation split of the prior-knowledge matrix. Here, 80% of the prior-knowledge is used for training, while the remaining 20% is used for validation. For PBMC, we do not have access to a gold standard network. For this reason, to evaluate our inferred GRNs, we select the network for each of the 5 cross-validation splits that achieves the best training hyperparameters. We then take the intersection of each of these 5 optimal networks, and filter the predicted interactions by those obtaining predicted means (probability) of > 0.90 across every split. For TFA inference, we select the best overall hyperparameters from the 5 fold cross-validation hyperparemeter search to infer a single network. We do this in order to obtain a single TFA matrix where the entries of this matrix are not affected by averaging across multiple datasets. All predicted TFA values are non-zero by nature of matrix factorization. We thus apply an 𝑙1 regularization on the matrix with 𝜆 = 1 to push low scoring activity values to 0. UMAP projections were performed on this regularized TFA matrix using the scanpy package [Wolf et al. 2018] with the following parameters: n-neighbors= 10, n-pcs= 40. Additional UMAPs for each of the considered PBMC immune TFs are available in Figure 3.3. To create the GRN diagrams for PBMC, we used the Gephi Open Graph Viz Platform [Bastian et al. 2009]. Additional GRNs for each of the PBMC immune TF families are displayed in Figure 3.2. PBMC heat-map dotplots and violin-plot were created using the default scanpy parameters for these functions.

100

Figure 3.2: PBMC GRN graphs for the family of TFs belonging to (A) SMAD, (B) STAT, (C) GATA, and (D) EGR. TFs for each graph are represented by orange nodes, while target genes are blue nodes. Pink regulatory edges represent interactions with supporting literature. Color scale for blue regulatory edges are scaled from light blue (less out-degree regulation) to darker blue (more out-degree regulation).

101

Figure 3.3: UMAP of predicted PBMC TFA. Each UMAP highlights the specific location of activity for each TF considered from the immune PBMC TFs. The final UMAP serves as a reference to TFA annotated by cell-type.

102

TF family

TF

Gene

Citation

STAT1

IRF1 IRF9 KDM6B SPAG9

[Abou El Hassan et al. 2017] [Au-Yeung et al. 2013] [Johnstone et al. 2021] [Mayumi et al. 2021]

STAT3

IL24 IRF9 ADAM12 CRTC3 CXCL2 GOLPH3 IQGAP1 NEDD4L NRL PBX1 SLC9A8 YAP1 IRF9 ANXA4

[Kumari et al. 2013], [Andoh et al. 2009] [Edsbäcker et al. 2019] [Roy et al. 2017] [Kim et al. 2017] [Nguyen-Jackson et al. 2010] [Wu et al. 2018] [Wei and Lambert 2021] [Nie et al. 2022] [Keuthan et al. 2019] [Wei et al. 2018] [Liu et al. 2023a] [Shibata et al. 2020] [Edsbäcker et al. 2019] [Li et al. 2020b]

STAT5

AUTS2 EPAS1

[Nagel et al. 2016] [Pietz et al. 2017]

GATA1

ATP2B4 FOXO3 GATA2 LAPTM5 LYL1 PBX1

[Lessard et al. 2017] [Katsumura et al. 2017] [Gao et al. 2015] [Zhang et al. 2019] [Johnson et al. 2007] [Chakrabarti et al. 2016]

GATA2

CUX1

[Wu et al. 2020]

GATA3

ABCC3 EGFR ETS2 FOS FOSL2 IQGAP1 KLF6 RUNX2 ZNF462

[Kobayashi et al. 2016] [Kong et al. 2022] [Blumenthal et al. 1999] [Liu et al. 2023b] [Li et al. 2020a] [Yang et al. 2022] [Fitch et al. 2020] [Liao et al. 2017] [Hintze et al. 2017]

GATA4

EPAS1 HNF4A IL1R1 YAP1

[Arroyo et al. 2021] [San Roman et al. 2015] [Yu et al. 2021] [Khalid et al. 2021]

IRF1

B2M BTN3A1

[Gobin et al. 2003] [Pietz et al. 2017]

IRF2

B2M

[Mercado et al. 2019]

IRF3

GPR108 RNF5 TRAF2

[Zhao et al. 2023] [Zhong et al. 2009] [Yu et al. 2023]

STAT

GATA

IRF

Table 3.3: Literature supported interactions for PBMC GRNs for STAT, GATA, and IRF immune TF families.

103

TF family

TF

Gene

Citation

SMAD1

CHD7

[Liu et al. 2014]

SMAD3

LAPTM5 LIN28A MSL2 TRIB3

[Gao et al. 2019] [Jung et al. 2020] [Jiang et al. 2019] [Hua et al. 2011]

SMAD4

CEBPB CXXC5 GDF15 HNF4 ROCK2 ULK1 WWOX

[Hill 2016] [Wang et al. 2013] [Rochette et al. 2021] [Chen et al. 2019b] [Wang et al. 2017] [Trelford and Di Guglielmo 2021] [Hsu et al. 2017]

EGR1

JUND NEDD4L NME1

[Chen et al. 2010] [Liu et al. 2022] [Wong et al. 2021]

SMAD

EGR

Table 3.4: Literature supported interactions for PBMC GRNs for SMAD and EGR immune TF families.

3.8.2.1

Bacillus subtilis

We used two microarray datasets for B. subtilis, which we label as B1 (GSE27219) and B2 (GSE67023). Both B1 and B2 underwent different normalization as part of standard microarray processing, described in detail in [Nicolas et al. 2012] and [Arrieta-Ortiz et al. 2015]. For the experiment "No Normalization", B1 was simply converted to integers, while B2 contained negative numbers and had to be scaled and then converted to integers so that the data represented positive integers similar to single-cell data. To demonstrate the importance of scaling microarray data to place independently collected datasets on the same scale, we demonstrate how Min-Max Scaling improves inference in both B. subtilis datasets. For "Min-Max Scaling", both B1 and the positive scaled B2 dataset were subsequently normalized using the following logic. Using the observation axis, values were linearly transformed so that the minimum value was mapped to 0 and the maximum value was mapped to 1. Each value was then multiplied by 100 and converted to integers to produce the resulting expression matrix of scaled single-cell-like integers.

104

3.8.3

Inferelator, Scenic, and CellOracle Networks

3.8.3.1

Saccharomyces cerevisiae

Networks were inferred using the "multitask" workflow setting of the Inferelator for the same single-cell S.cerevisiae datasets described in [Skok Gibbs et al. 2022]. For each algorithm, BBSR, StARS, and AMuSR, the following parameters were used: gold_standard_filter_method = "keep_all_gold_standard", num_bootstraps=5. Aggregated multi-task networks were used for benchmarking, while single-task networks were disregarded for the purpose of this work. To make these networks directly comparable to PMF, we did not make use of normalization, count minimum, or meta-data options available within the Inferelator workflow. Networks inferred with Scenic and CellOracle used the same input files, with no additional parameters specified.

Method PMF-GRN AmUSR BBSR StARS SCENIC CellOracle

Prior Information Regular None Shuffled 0.375 0.014 0.023 0.223 0.024 0.019 0.402 0.022 0.018 0.186 0.028 0.017 0.014 0.014 0.014 0.383 N/A 0.013

Table 3.5: AUPRCs achieved by PMF-GRN, the Inferelator algorithms (AMuSR, BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets.

105

Method PMF-GRN BBSR StARS SCENIC CellOracle

Split 1 0.114 0.112 0.109 0.020 0.034

Cross Validation Split Split 2 Split 3 Split 4 0.096 0.086 0.1342 0.128 0.161 0.171 0.137 0.154 0.195 0.021 0.018 0.025 0.042 0.034 0.043

Split 5 0.118 0.139 0.151 0.021 0.034

Table 3.6: AUPRCs achieved by PMF-GRN, the Inferelator algorithms (BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets using the gold standard for 5-fold cross validation.

Method PMF-GRN BBSR StARS SCENIC CellOracle

No Noise 0.343 0.264 0.136 0.075 0.417

Noise Added 100% Noise 250% Noise 0.280 0.198 0.208 0.186 0.125 0.114 0.068 0.059 0.306 0.226

500% Noise 0.149 0.174 0.118 0.055 0.175

Table 3.7: AUPRCs achieved by PMF-GRN, the Inferelator algorithms (BBSR, and StARS), Scenic and CellOracle on S. cerevisiae datasets using increasing amounts of noise added to the prior-knowledge data.

Method PMF-GRN BBSR AMuSR StARS SCENIC CellOracle

Intersection over Union (IoU) IoU 15.69% 14.56% 12.46% 11.78% 3.17% 30.28%

Table 3.8: Intersection over Union (IoU) scores achieved by PMF-GRN, the Inferelator algorithms (AMuSR, BBSR, and StARS), Scenic and CellOracle for GRNs learned on individual S. cerevisiae datasets.

106

Expression size (%) 80% 60% 40% 20%

1 0.2838 0.2382 0.3157 0.2995

Sample Number 2 3 4 0.2947 0.3395 0.3325 0.3295 0.3252 0.2775 0.2723 0.3599 0.2858 0.3503 0.3229 0.2868

5 0.2774 0.3296 0.2939 0.3335

Table 3.9: AUPRCs achieved by PMF-GRN across 4 different downsample sizes (80%, 60%, 40%, and 20%), across 5 samples for each downsample size.

Cross Validation Split 80% train - 20% val 60% train - 40% val 40% train - 60% val 20% train - 80% val

1 0.0460 0.0479 0.0459 0.0401

2 0.0400 0.0378 0.0397 0.0434

Dataset 3 4 0.0463 0.0361 0.0463 0.0417 0.0375 0.0347 0.0403 0.0427

5 0.0408 0.0437 0.0381 0.0239

Full GRN 0.3337 0.3337 0.2909 0.2882

Table 3.10: AUPRCs achieved by PMF-GRN across 4 different cross-validation splits. 5 hyperparameters searches were performed for each cross-validation split. Full GRN was inferred using the hyperparameters for the best overall AUPRC per cross-validation split.

107

Method

Input

Output

Methodology

Pipeline

PMF-GRN

(1) one or more scRNA-seq datasets (2) priors constructed from genomic data such as ATAC-seq or ChIP-seq and TF motifs or literature database derived priors

(1) individual (and) combined GRN (2) TFA

probabilistic matrix factorization, variational inference

1) hyperparameter search using 80-20 split of input prior (2) GRN inference with optimal model parameters (3) evaluate GRNs with AUPRC, MCC and F1 scores and uncertainty calibration

SCENIC

(1) scRNA-seq (2) a list of TFs ranking databases for motif analysis

(1) GRN: loom file containing regulons (interactions between TFs and their target genes) (2) TFA: biological activity of regulons in a given cell

tree based regression models, gradient boosting machine regression, motif enrichment analysis, and gaussian mixture models for regulon activity binarization

(1-4) preprocessing steps (5) network inference (6) module generation (7) motif enrichment and TF regulon prediction (8) cellular enrichment (9) optional binarization of cellular regulon activity (10) clustering of cells based on regulon activity

CellOracle

(1) ATAC-seq + TF motifs (2) scRNA-seq

(1) GRN (2) cell-state transition vectors after gene perturbation

Bayesian Ridge regression

(1) base GRN construction with scATAC-seq or promoter databases (2) scRNA-seq preprocessing (3) context dependent GRN inference (4) network analysis (5) simulation of cell identity after TF perturbation (6) calculation of pseudotime gradient for perturbation scores

AMuSR

(1) two or more scRNA-seq datasets (2) one or more priors constructed from genomic data and TF motifs or literature database derived priors

(1) individual and combined GRNs (2) TFA

Regularized linear regression

(1) estimating TFA (2) learning regression parameters (3) model selection with bayesian information criterion (4) evaluate GRNs with AUPRC, MCC, F1 scores

BBSR

(1) one or more scRNA-seq datasets (2) one or more priors constructed from genomic data and TF motifs or literature database derived priors

(1) individual (and) combined GRN (2) TFA

Bayesian best subset regression

(1) estimate TFA (2) learn regression parameters (3) model selection (4) evaluate GRNs with AUPRC, MCC, F1 scores

StARS

(1) one or more scRNA-seq datasets (2) one or more priors constructed from genomic data and TF motifs or literature database derived priors

(1) individual (and) combined GRN (2) TFA

Least absolute shrinkage and selection operator combined with the Stability Approach to Regularization Selection

(1) estimate TFA (2) learn regression parameters (3) model selection (4) evaluate GRNs with AUPRC, MCC, F1 scores

Table 3.11: GRN inference method comparison table. Table includes input data, output, methodology and pipeline organization for each of the six GRN inference methods discussed in this work.

108

3.8.4

TF Target Gene Connectivity Matrix Generation

3.8.4.1

Saccharomyces cerevisiae

Datasets were obtained from [Skok Gibbs et al. 2022] without further modification. 3.8.4.2

Peripheral Blood Mononuclear Cells

Datasets were obtained from [Tjärnberg et al. 2023] without further modification. 3.8.4.3

BEELINE

Datasets were obtained from [Pratapa et al. 2020]. TF-target gene matrices were constructed by creating cross-tab matrices from the RefGRN. 50% of this matrix was used as prior-knowledge, the remaining 50% was used as a gold standard for evaluation. 3.8.4.4

Bacillus subtilis

A prior-known TF-target gene interactions matrix was obtained from the Subtiwiki database [Faria et al. 2016] from "regulations" (downloaded 07/21/22). Using the columns "regulator locus" and "gene locus" a cross-tab integer matrix was created, where 1 represents the existence of an interaction and 0 represents no interaction. This matrix was randomly split 5 times in 80%20% proportions along the gene axis to generate independent prior-known information and gold standard matrices.

109

4 | A Genomic Language Model Prior for Gene Regulatory Network Inference This chapter is adapted from the pre-print "GLM-Prior: a genomic language model for transferable sequence-derived priors in gene regulatory network inference". Claudia Skok Gibbs, Angelica Chen, Richard Bonneau, and Kyunghyun Cho. (2026).

4.1

Abstract

Gene regulatory network inference depends on high-quality prior knowledge, yet curated priors are often incomplete or unavailable across species and cell types. We present GLM-Prior, a genomic language model fine-tuned to predict transcription factor–target gene interactions directly from nucleotide sequence. We integrate GLM-Prior with PMF-GRN, a probabilistic matrix factorization model, to create a dual-stage pipeline that combines sequence-derived priors with single-cell gene expression data for GRN inference. Across six human, mouse, and yeast cell lines, GLM-Prior performance scales with positive label abundance and diverse transcription factor coverage, achieving strong accuracy in well-annotated mammalian contexts. We evaluate single-species, species-transfer learning, and multi-species training paradigms and show that

110

GLM-Prior generalizes to held-out gene and TF sequences, enabling experiment-agnostic prior construction in previously unprofiled contexts. Furthermore, comparisons to accessibility-based priors across multiple GRN inference methods show that GLM-Prior provides the most robust priors in mammalian cell lines. Together, our results demonstrate that prior construction, rather than the choice of GRN inference algorithm, is the primary determinant of GRN inference performance, and establish GLM-Prior as a framework for building high-quality, experiment-agnostic priors that can be deployed even in understudied or experimentally inaccessible systems.

4.2

Introduction

Gene regulatory networks (GRNs) map the transcriptional relationships between transcription factors (TFs) and their target genes, providing a framework for understanding cellular function and gene expression control in cells [Badia-i Mompel et al. 2023; Kim et al. 2023]. Constructing accurate GRNs requires integrating large-scale sequencing datasets, such as single-cell RNA-seq and ATAC-seq, to infer regulatory interactions that cannot be measured experimentally [Fiers et al. 2018]. However, the accuracy of GRN inference depends heavily on the quality of the prior knowledge used to guide the inference algorithm [Stock et al. 2025; McCalla et al. 2023]. In the context of GRNs, prior knowledge refers to an initial matrix of putative TF-gene regulatory interactions, which provides structural guidance to the inference algorithm by defining which TF-gene pairs are more likely to represent true regulatory edges. Recent methods have demonstrated that better prior knowledge, which captures biologically meaningful regulatory signals, leads to more accurate and informative GRN reconstruction, highlighting the critical role of prior knowledge in improving GRN inference quality [Skok Gibbs et al. 2022, 2024; Kamimoto et al. 2023; Van de Sande et al. 2020]. Prior knowledge for well-studied species is typically constructed using databases of experimentally validated TF-gene interactions [de Souza 2012; Teixeira et al. 2018]. For less-characterized

111

organisms, prior knowledge is inferred by combining structural genomic data with sequencing information [Mercatelli et al. 2020]. The Inferelator 3.0’s Inferelator-Prior performs TF motif enrichment within open chromatin regions to define edges between proximally bound TFs to target genes [Skok Gibbs et al. 2022]. CellOracle’s base GRN draws edges between TFs bound within promoter regions of accessible target genes [Kamimoto et al. 2023]. SCENIC’s CisTarget constructs prior knowledge by ranking enriched motifs in accessible regions using genome-wide motif scores and a hidden Markov model to predict TF-gene interactions [Van de Sande et al. 2020]. While these methods demonstrate the importance of prior knowledge in GRN inference, they also expose key limitations. Motif annotation remains incomplete [Inukai et al. 2017], chromatin accessibility data is noisy and cell-type specific [Buenrostro et al. 2015b], and existing methods often fail to capture long-range regulatory interactions [Loers and Vermeirssen 2024]. Further, these existing approaches cannot create prior knowledge matrices which generalize to other species. These limitations underscore the need for a more generalizable approach to constructing high-quality prior knowledge matrices, particularly for poorly characterized organisms [Stock et al. 2025; Barbosa et al. 2018]. A promising direction is to incorporate transformer-based genomic foundation models which learn regulatory logic directly from DNA sequence. Architectures such as the Nucleotide Transformer [Dalla-Torre et al. 2024] have demonstrated strong performance across genomic tasks, including predicting chromatin accessibility, gene expression and transcription factor binding [Avsec et al. 2021; Ji et al. 2021; Nguyen et al. 2023; Linder et al. 2025]. By learning conserved regulatory patterns directly from sequence data, transformer-based genomic models offer a compelling solution to the limitations of motif-based prior knowledge approaches [Consens et al. 2025]. By leveraging attention mechanisms, they effectively capture long-range dependencies in DNA sequences [Vaswani et al. 2017], enabling more accurate modeling of complex regulatory interactions. When trained on large, multi-species genomic datasets, these models encode both species-specific and cross-species features, supporting generalization across cell types [Avsec 112

et al. 2021; Cui et al. 2024] and organisms [Dalla-Torre et al. 2024; Zhou et al. 2023]. This opens the door to a powerful transfer learning paradigm: models trained on species with extensive regulatory data (e.g., H. Sapiens) can be adapted to predict interactions in related species (e.g., M. musculus), when curated priors are incomplete or unavailable. In this way, genomic language models offer the potential to expand prior knowledge construction to organisms and cell types for which no direct interaction data exist, enabling GRN inference in new and challenging contexts. In this work, we present GLM-Prior, a novel approach for constructing high-quality prior knowledge to improve downstream GRN inference. GLM-Prior is built by fine-tuning the 250 million parameter Nucleotide Transformer [Dalla-Torre et al. 2024], pretrained on the genomes of 850 species, into a sequence classification model for predicting TF-gene regulatory interactions. Unlike motif-based approaches that rely on the proximity of TF motifs to accessible promoter regions, GLM-Prior directly predicts interactions from paired nucleotide sequences of TF binding motifs and gene bodies. These sequences are jointly encoded by the transformer and passed through a classification head, which outputs a probability of regulatory interaction for each TFgene pair. To construct a binary prior knowledge matrix suitable for downstream GRN inference, we apply an optimized classification threshold that converts these probabilities into discrete binary interactions. As most TF-gene pairs are non-regulatory, the model is trained using classweighted loss and balanced sampling to mitigate class imbalance during training. Together with the Transformer’s attention mechanisms, which capture long-range dependencies and complex regulatory grammar, these features enable GLM-Prior to generalize across cell types and species more effectively than motif-based approaches. We investigate the generalizability of GLM-Prior by evaluating three training paradigms: (1) single-species models trained and evaluated on the same organism; (2) transfer learning, in which models trained in one species are used to predict regulatory interactions in an unseen species; and (3) multi-species training, where a single model is fine-tuned on data from multiple organisms. Across these settings, performance scales with abundance and diversity of positive class labels 113

and TF coverage in the training data. Human and mouse cell lines with rich label sets yield higher AUPRCs than yeast, where extreme class imbalance and sparse labels limit generalization. Transfer learning between evolutionarily related species (human and mouse) maintains or modestly improves performance relative to single-species models, demonstrating that GLM-Prior captures conserved regulatory sequence features that can be reused across lineages. In contrast, transfer learning between mammals and yeast does not exceed near-chance performance, reflecting a deeper divergence in regulatory grammar. Multi-species training on human, mouse, and yeast preserves accuracy comparable to the best single-species or transfer models in most cell lines, with only mild trade-offs in some human contexts, and produces a domain-stable prior that can be applied to new systems without species-specific retraining. To evaluate the utility of GLM-Prior, we incorporate it into a dual-stage GRN inference pipeline. In the first stage, GLM-Prior generates a prior knowledge matrix of TF-gene interactions from nucleotide sequence. In the second stage, this matrix is used to constrain PMF-GRN [Skok Gibbs et al. 2024], a probabilistic matrix factorization model that infers GRNs from single-cell gene expression data. This framework allows us to isolate and compare the contributions of prior construction and expression-based inference to overall GRN reconstruction performance. Building on this, we benchmark GLM-Prior against accessibility-based priors (InferelatorPrior [Skok Gibbs et al. 2022], and CellOracle’s baseGRN [Kamimoto et al. 2023]), and integrate each prior with it’s downstream GRN inference algorithm (PMF-GRN [Skok Gibbs et al. 2024], the Inferelator 3.0 [Skok Gibbs et al. 2022], and CellOracle [Kamimoto et al. 2023]) across yeast, mouse, and human. GLM-Prior consistently performs at or above chance in all six cell lines and provides the strongest prior in four of six contexts (hESC, mESC, mDC, and mHSC). In contrast, accessibility-based priors remain competitive in settings dominated by proximal regulatory logic, such as yeast, but struggle to generalize above chance performance in complex mammalian cell lines. When these priors are passed into downstream GRN inference, GRN inference typically yields modest gains when the prior is already strong, and larger improvements are observed 114

only when the prior is weak or inaccurate. Critically, the identity of the best-performing GRN in each cell line almost always matches the identity of the best prior, and no GRN inference method is uniformly superior across priors or species. Together, these results indicate that prior construction, rather than the choice of GRN inference algorithm, is the primary driver of GRN reconstruction performance. Together, our findings recast GRN inference as a problem primarily limited by the quality of prior knowledge rather than by expression-based modeling. GLM-Prior provides a scalable, species-aware framework for constructing priors, while downstream GRN methods act mainly to refine or contextualize this scaffold. In this work, we characterize GLM-Prior’s generalization across cell types and species, quantify its advantages over accessibility-based priors, and dissect how different GRN inference algorithms interact with these priors to shape the final inferred networks.

4.3

Results

4.3.1

Dual-stage training pipeline overview

We develop a dual-stage training pipeline to improve the accuracy and interpretability of GRN inference by integrating sequence-derived regulatory interactions with probabilistic modeling of single-cell expression data (Figure 4.1). The pipeline is designed to decouple the construction of prior knowledge from the expression-driven GRN inference step, allowing each component to be independently optimized and evaluated. In this framework, the first stage generates a biologically informed prior knowledge matrix by using a fine-tuned genomic language model (GLM-Prior), while the second stage uses this matrix as structural input to the probabilistic matrix factorization model (PMF-GRN [Skok Gibbs et al. 2024]) to a infer GRN.

115

4.3.1.1

Stage 1: Constructing seqence-informed prior knowledge with GLM-Prior

In the first stage of the pipeline, we construct a prior knowledge edge matrix representing predicted positive regulatory (1) and negative non-regulatory (0) interactions between TFs and their target genes (Figure 4.1A). To generate these predictions, we develop GLM-Prior, a transformerbased sequence classification model built by fine-tuning the 250 million parameter Nucleotide Transformer [Dalla-Torre et al. 2024] model, which was pretrained on the genomes of 850 different species. The primary goal of GLM-Prior is to leverage sequence information to accurately predict TF-gene regulatory interactions, in order to construct an informative prior knowledge matrix. GLM-Prior takes as input paired nucleotide sequences consisting of TF binding motifs and gene body regions, along with labels representing known positive or negative interactions. These sequences are concatenated and encoded into a shared representation using the transformer’s embedding and encoding layers. This representation is passed through a classification layer that outputs the probability of a true regulatory interaction. During training, the model evaluates a range of classification thresholds, selecting the threshold that maximizes the F1 score on the validation set. This optimal threshold is used to binarize the predicted probabilities into positive or negative interactions for the final prior knowledge matrix. Since GRN datasets are often highly imbalanced, with far fewer positive interactions compared to negative interactions, we apply several strategies to mitigate this imbalance during training. Specifically, we apply a downsampling rate to the negative class and use a class-weighted cross-entropy loss, tuning the negative class weight through hyperparameter search. Additionally, to further ensure stable training, we implement a custom DataLoader to create training batches with equal numbers of positive and negative examples, ensuring balanced gradient updates. After training, GLM-Prior performs inference across all provided TF-gene pairs, generating

116

Figure 4.1: Schematic of the dual-stage training pipeline. (A) GLM-Prior is trained on concatenated TF binding motifs and gene sequences labeled as interacting or non-interacting. Input data is split into training and validation sets, where the training set is downsampled and placed into even class batches using a custom DataLoader. These input sequences are embedded and encoded by the transformer architecture, and passed through a classification head to predict the probability of a regulatory interaction. After training, inference is performed on all TF-gene pairs to generate a prior knowledge matrix. (B) Inferred prior knowledge is passed to PMF-GRN to provide a structural constraint to the probabilistic graphical model during inference. PMF-GRN infers transcription factor activity and a GRN for the input cell-line dataset. The resulting GRN is then evaluated using AUPRC and uncertainty calibration with an independent gold standard.

117

a dense probability prior knowledge matrix. The predicted probabilities are then binarized using the optimal F1 score-derived classification threshold, generating a binary prior knowledge matrix that encodes regulatory interactions for a given organism. Although we can use the probabilities directly, we binarize the prior knowledge matrix to be in line with existing methods and for more direct comparison. We then either use this sequence-based prior knowledge as is and evaluate on an independent gold standard, or can pass it into the second stage of the pipeline, enabling a modular information transfer from model-driven sequence inference to data-driven GRN inference. 4.3.1.2

Stage 2: Refining GRN inference from single-cell data with PMF-GRN

In the second stage of the pipeline, the prior knowledge matrix predicted by GLM-Prior serves as a structured input to our previously developed model, PMF-GRN [Skok Gibbs et al. 2024] (Figure 4.1B). PMF-GRN leverages this sequence-informed prior knowledge to enhance the inference of GRNs from single-cell gene expression data. It does so by using probabilistic matrix factorization, which decomposes the observed gene expression into latent factors representing transcription factor activity (TFA) and TF-target gene regulatory relationships. Concretely, PMF-GRN models the expression matrix 𝑊 using the likelihood function

𝑝 (𝑊 |𝑈 , 𝑉 = 𝐴 ⊙ 𝐵, 𝜎𝑜𝑏𝑠 , 𝑑),

where E[𝑊 ] ≈ 𝑑 ⊙ 𝑈𝑉 ⊤ , with per-cell sequencing depth 𝑑 and observation noise 𝜎𝑜𝑏𝑠 . Here, 𝑈 is a nonnegative cells by TFs matrix of TF activities, and 𝑉 is a genes by TFs matrix, which is further factorized as 𝑉 = 𝐴 ⊙ 𝐵. 𝐴 ∈ (0, 1) is defined as the degree of existence of a TFtarget gene interaction and 𝐵 ∈ R is its signed effect. Each latent variable, 𝑈 , 𝐴, 𝐵, 𝑑,, 𝜎𝑜𝑏𝑠 is assumed independent a priori. In this formulation, the GLM-Prior predicted prior knowledge matrix provides the hyperparameters for the logistic-normal prior on 𝐴, anchoring factors to

118

specific TFs and mitigating column-label identifiability inherent to matrix factorization. Using variational inference, PMF-GRN estimates the posterior distribution of all latent variables. The posterior mean of 𝐴 defines the inferred GRN, while the posterior variance provides an edge-level uncertainty score. These uncertainty estimates enable confidence-weighted interpretation of TF-gene edges and are especially valuable when gold standard regulatory annotations are incomplete or unavailable. The combination of GLM-Prior and PMF-GRN is motivated by their complementary strengths. GLM-Prior uses nucleotide sequence to construct a biologically grounded scaffold of regulatory edges, capturing both local motif-driven interactions and potential long-range dependencies. This scaffold can be generated as a general prior matrix or tailored to specific cell lines or conditions. However, as GLM-Prior has not seen any experimental data, its predictions require refinement using real observations. PMF-GRN provides this refinement by integrating singlecell expression data within a principled generative framework, updating edge probabilities based on observed regulatory patterns while maintaining consistency with the sequence-derived prior. Moreover, like most GRN inference algorithms, PMF-GRN relies on prior knowledge to obtain a meaningful posterior GRN; without an informative prior, factor identifiability issues limit interpretability. GLM-Prior resolves this limitation by anchoring TF-gene relationships, while PMFGRN tailors the scaffold to the data, quantifies uncertainty, and produces context-specific networks. Together, this integration balances predictive power from large-scale genomic language models with the interpretability and flexibility of probabilistic modeling, yielding GRNs that are both grounded in sequence evidence and informed by experimental observations.

119

4.3.2

GLM-Prior’s generalization scales with data composition across species and cell types

We assess GLM-Prior by benchmarking across six cell lines spanning yeast, human, and mouse (Figure 4.2 and 4.3). For each organism, we create gene-TF sequence pairs by extracting nucleotide sequences for genes from GTF annotations and TF binding motifs from CisBP [Weirauch et al. 2014]. For each gene-TF pair, we gather available regulatory labels from the YEASTRACT database [Teixeira et al. 2018] for yeast, and the STRING [Mering et al. 2003; Szklarczyk et al. 2010, 2021, 2023] and TRRUST [Han et al. 2015, 2018] databases for human and mouse. To assess outof-distribution performance, we evaluate on held-out test sets constructed from well-established reference labels such as the literature curated gold standard from [Tchourine et al. 2018] for yeast, as well as cell-line specific reference networks from BEELINE [Pratapa et al. 2020] for human and mouse. Before training each cell-line specific model, we ensure strict independence between training and test sets at the level of gene-TF sequence pairs and their labels, eliminating sequence and label leakage. This design ensures that performance reflects the model’s ability to generalize to unseen sequences and correctly predict their regulatory interactions. To evaluate model performance within each cell line, we implement a multi-stage benchmarking procedure (Figure 4.2 and 4.3 A-E). In Figure 4.2A and 4.3A, we perform a one-epoch hyperparameter sweep over combinations of class weights in the loss function and downsampling rates for the negative class. We identify the configuration that yields the highest F1-score and use these parameters for full model training. In Figure 4.2B and 4.3B, we train the model for ten epochs using this optimal configuration and monitor performance on a held-out validation set. Following training, we apply the model to all gene-TF pairs, including those unseen during training, and evaluate our predictions against the corresponding held-out set’s gold standard labels (Figure 4.2C and 4.3C). To interpret what drives the resulting performance, we examine the number of genes and TFs represented in the training and test sets for each cell line (Figure 4.2D 120

and 4.3D), providing context for possible domain shifts between splits. Finally, we quantify the class composition of positive and negative examples in the training and test sets (Figure 4.2E and 4.3E), allowing us to relate generalization behavior to the underlying data balance.

Figure 4.2: GLM-Prior model performance across yeast, hESC, and HepG2 cell lines. (A) One-epoch hyperparameter sweep over class-weights and negative-class downsampling rates, evaluated by F1-score. (B) Validation metrics during 10 epoch training using the best hyperparameter configuration. (C) Test AUPRC on held-out gold standards (with chance performance shown for reference in gray). (D) Train and test split composition (number of genes and TFs). (E) Class composition (number of positive and negative labels) in train and test sets.

Across the six evaluated cell lines, GLM-Prior demonstrates variation in predictive performance, highlighting the influence of data composition, particularly the number of positive regulatory interactions, on generalization (Figure 4.2E and 4.3E). In yeast, GLM-Prior achieves a nearchance AUPRC of 0.02 compared to a 0.01 baseline, despite strong within-training metrics (Figure 121

4.2B). This limited generalization likely stems from the extreme class imbalance in the training data, where only 660 positive interactions are available among 724, 179 negatives. Although the test set contains a moderate number of positives (986), the small number of training examples and high-performing validation metrics suggest that the model overfits the limited regulatory patterns seen during training and fails to extrapolate to novel gene-TF combinations.

Figure 4.3: GLM-Prior model performance across mESC, mDC, and mHSC cell lines. (A) One-epoch hyperparameter sweep over class-weights and negative-class downsampling rates, evaluated by F1-score. (B) Validation metrics during 10 epoch training using the best hyperparameter configuration. (C) Test AUPRC on held-out gold standards (with chance performance shown for reference in gray). (D) Train and test split composition (number of genes and TFs). (E) Class composition (number of positive and negative labels) in train and test sets.

In human embryonic stem cells (hESC), performance improves modestly, with an AUPRC of 0.19 versus a 0.15 chance baseline. Here, GLM-Prior benefits from a larger pool of positive 122

training examples (5, 305) and a more balanced test set (53, 337 positives and 318, 730 negatives). However, TF coverage drops from 504 in training to 79 in testing, implying that while model generalizes across genes, reduced TF diversity in the test set limits broader generalization. In HepG2 hepatocytes, GLM-Prior achieves a stronger performance (AUPRC = 0.30 versus 0.25 chance) supported by both a higher number of positive training examples (8, 087) and a wellrepresented test set (56, 567 positives). Despite the test set containing fewer TFs (53 versus 527 in training), the model still generalizes effectively, indicating that it learns transferable sequence features that extend across TFs in this cell line. This performance, coupled with the relatively large and balanced training data, suggest that sufficient label diversity allows the model to learn regulatory sequence logic that is not confined to specific TF identities. In mouse embryonic stem cells (mESC), GLM-Prior achieves an AUPRC of 0.23 versus 0.20 chance, showing moderate but consistent generalization. The training set includes 3, 129 positives among 476, 260 negatives while the test set is considerably more balanced (65, 933 positives, 271, 015 negatives), covering 5, 711 genes unseen during training. The broad gene coverage supports cross-gene generalization, though limited positive training examples likely cap performance relative to the human cell lines. In mouse dendritic cells (mDC), performance remains at chance (AUPRC = 0.09, despite a substantial number of training examples 11, 551). The sharp reduction in TF coverage from 462 in training to only 29 in testing suggests that the test TFs may regulate distinct, cell-type specific programs absent from training. Finally, mouse hematopoietic stem cells (mHSC) exhibit the strongest generalization (AUPRC = 0.49 versus 0.37 chance). Although the training set contains a modest number of positives (2, 808), the test set is both large (182, 005 positives) and diverse (6, 777 genes, 72 TFs), providing a robust evaluation. The high AUPRC gain indicates that GLM-Prior captures sequence-level regulatory features that transfer effectively to unseen gene-TF pairs in this cell line. Together, these results demonstrate diverse composition of gene-TF relationships in training 123

data is a critical factor behind GLM-Prior’s generalization. In particular, generalization is most affected by the abundance and diversity of positive interactions and TFs in training. We interpret the above chance performance gains as the model’s ability to learn transferable sequence regulatory patterns when sufficient label diversity is available.

4.3.3

GLM-Prior supports successful species-transfer learning and multi-species training

Following our six cell line evaluation of GLM-Prior, we next task the model to predict prior knowledge in a species unseen during training (Figure 4.4 and 4.5). Building on the previous section, which established that generalization depends strongly on data composition, we now ask whether the model can transfer learned regulatory logic across species. Such cross-species transfer would represent a major advantage over existing approaches, enabling the construction of informative priors for understudied organisms using models trained on evolutionarily similar species. To test this, we perform two complementary experiments: A) transfer learning, where models trained in one species are applied to another, and B) multi-species training, where a single model is trained sequentially on human, mouse, and yeast data. In all cases, the held-out evaluation sets are fully disjoint from training data, and overlapping gene and TF identifiers are removed to prevent crossspecies leakage. In Figure 4.4A, we provide a schematic for the transfer learning setup. Three transfer learning configurations are evaluated, (𝑖) a model trained jointly on human and mouse data, evaluated on yeast, (𝑖𝑖) a mouse-only model evaluated on human cell lines (hESC and HepG2), and (𝑖𝑖𝑖) a human-only model evaluated on mouse cell lines (mESC, mDC, and mHSC). Across all six cell lines, we find that transfer learning maintains or modestly improves performance relative to single-species models. In yeast, performance remains near chance (AUPRC = 0.02 vs 0.01), consistent with deep evo-

124

lutionary divergence between eukaryotic and mammalian regulatory syntax. In contrast, crossspecies transfer between human and mouse demonstrates clear success, where hESC maintains an AUPRC = 0.19 and HepG2 improves to 0.31. Simiarly, human-trained models applied to mouse cell lines achieve stable performance comparable to single-species results. These results demonstrate that GLM-Prior captures conserved sequence-level features that transfer across closely related species, enabling accurate prior construction without species-specific retraining.

Figure 4.4: Transfer learning and multi-species training with GLM-Prior. (A) Schematic and AUPRC results of the species-transfer learning setup, showing three configurations: (𝑖) a model trained jointly on human and mouse data evaluated on yeast, (𝑖𝑖) a mouse-only model evaluated on human cell lines, and (𝑖𝑖𝑖) a human-only model evaluated on mouse cell lines. (B) Schematic and AUPRC results of the multi-species model trained sequentially on human, mouse and yeast data, evaluated across six cell lines.

In Figure 4.4B, we present the multi-species model trained sequentially on human, mouse, and yeast data and evaluated on all six cell lines. The multi-species model achieves comparable AUPRCs relative to both single-species and species-transfer settings (yeast = 0.02, hESC = 0.19, HepG2 = 0.29, mESC = 0.23, mDC = 0.09, and mHSC = 0.49). 125

Figure 4.5: Transfer learning and multi-species training with GLM-Prior continued. (A) Comparison of training composition (number of genes and TFs) for single-species, species-transfer, and multi-species models. (B) Class composition (positive and negative labels) in the corresponding training sets. (C) Normalized AUPRC across single-species, species-transfer, and multi-species models for each cell line.

126

To interpret these trends, Figure 4.5A-B summarize the training compositions and normalized AUPRC (Figure 4.5C) across single, transfer, and multi-species setups. Across each cell line, transfer learning consistently expands the number of genes, TFs, and positive labels seen during training. In some settings, such as with HepG2, this broader coverage allows for a modest normalized AUPRC increase (0.08 vs 0.07 single-species). In others, transfer learning maintains single-species performance. In contrast, mHSC shows a slight decrease in performance (0.14 species-transfer vs 0.19 single-species) despite larger training breadth. This drop suggests that the regulatory logic active in mouse HSCs is relatively cell-type and species-specific and therefore not well represented in the context-agnostic human interaction data used for species-transfer training, limiting the model’s ability to fully recover mHSC-specific edges. Notably, the multi-species model maintains similar performance to the single-species models with only slight performance drops in hESC and HepG2. Here, we demonstrate that multi-species training yields a robust, domain-stable prior that generalizes across lineages. This provides a practical route to build prior knowledge for poorly understood systems or cell lines when a welldiversified training set is desired, offering broad regulatory coverage without assuming a species specific match. Together, these results demonstrate that GLM-Prior achieves successful cross-species transfer, a property not attainable with accessibility-based prior knowledge approaches which depend on species-specific data. Applying a model trained in one organism directly to another enables scalable, sequence-based prior construction in under-characterized systems. In parallel, multispecies training yields a robust, lineage-agnostic prior that maintains stable performance across contexts. When the target cell line is poorly characterized or its closest species match is uncertain, the multi-species model presents a practical default for constructing reliable prior knowledge.

127

4.3.4

Integrating GLM-Prior into PMF-GRN provides contextual GRN inference and edge refinement

To evaluate how sequence-derived regulatory prior knowledge interacts with expression-based GRN inference, we next used each cell line-specific GLM-Prior matrix as input to PMF-GRN. Whereas Stage 1 (Section 4.3.1.1) directly predicts TF-gene interactions from sequence, Stage 2 (Section 4.3.1.2) treats these predictions as prior knowledge and refines the network using single cell expression data. For each cell line, we paired GLM-Prior with matched single cell datasets, for example, GSE125162 [Jackson et al. 2020] and GSE144820 [Jariani et al. 2020] for yeast, and the corresponding expression datasets provided in BEELINE [Pratapa et al. 2020] for the human and mouse cell lines. This experiment is designed to assess not only whether PMF-GRN improves predictive accuracy, but also how much the network must be altered when inference is initialized from an already information-rich prior. Across all datasets, integrating GLM-Prior into PMF-GRN results in equal or improved AUPRC relative to the prior alone (Figure 4.6A). The magnitude of improvement varies by species, where performance remains unchanged in yeast (0.02) and hESC (0.19), while clear gains are observed in HepG2 (0.30 prior → 0.38 GRN), mESC (0.23 prior → 0.27 GRN), and mHSC (0.49 prior → 0.52 GRN). These improvements occur despite PMF-GRN modifying only a small fraction of the GLM-Prior edge matrix (Figure 4.6B), with added edges ranging from 0.03% in yeast to 0.54% in mHSC, and edge removals remaining minimal across all contexts. Notably, modifications are strongly biased toward edge additions rather than deletions, indicating that PMF-GRN treats the GLM-Prior as a stable scaffold and makes targeted expansions rather than large-scale refinement. Interestingly, the extent of structural changes does not strictly track with performance improvement. For example, hESC and mESC exhibit tens of thousands of added or removed edges, yet show minimal differences in AUPRC. This mismatch suggests that many of the changes introduced by PMF-GRN occur outside the held-out evaluation portion of the network, which 128

only cover a subset of TF-gene interactions. Conversely, in yeast, where GLM-Prior is close to chance (AUPRC = 0.02) and the available training interactions are sparse, PMF-GRN fails to recover meaningful signal despite access to two large single-cell expression datasets, indicating that expression-driven inference alone cannot compensate for a weak or uninformative prior.

Figure 4.6: Integration of GLM-Prior with PMF-GRN for downstream GRN inference. (A) AUPRC comparison of GLM-Prior alone versus integrated with PMF-GRN across six cell line contexts. (B) Edge flip analysis quantifying changes made to the prior knowledge interaction matrix during GRN inference. Blue bars indicate edges gained and red bars indicate edges removed by GRN inference.

These findings indicate that when a sequence-derived prior already captures substantial regulatory signal, the role of expression-based inference shifts from discovery to contextual refinement. Rather than constructing the network from expression patterns alone, PMF-GRN behaves conservatively, preserving high-confidence structure, adding plausible missing edges, and per129

forming limited pruning. In this dual-stage framework, the strength of GLM-Prior largely determines the achievable performance ceiling, while PMF-GRN contributes selective improvements where transcriptomic evidence provides additional support; when the prior is poor, as in yeast, expression-based GRN inference has limited capacity to recover accurate regulatory structure.

4.3.5

Comparing prior construction strategies across yeast, mouse, and human cell lines

Next, we assess how GLM-Prior compares to widely used accessibility-based prior construction approaches, evaluating performance relative to CellOracle and Inferelator priors across six cell line contexts. Unlike GLM-Prior, which leverages sequence-level information to predict TF-gene regulatory edges, both CellOracle and Inferelator derive putative interactions from chromatin accessibility and motif enrichment. Importantly, these accessibility-based priors are constructed using ATAC-seq measured in the matched cell line, whereas GLM-Prior is trained on contextagnostic interaction resources (STRING and TRRUST) and evaluated without explicitly conditioning on cell-line specific chromatin state. This comparison therefore evaluates whether sequencederived priors can match or exceed the performance of accessibility-based approaches across diverse biological settings. Across all six cell lines, GLM-Prior performs at or above chance, exceeding the chance accuracy in five of the six contexts, demonstrating robust generalization across species and cell types (Figure 4.7A). In comparison, Inferelator-Prior performs above chance in only two out of six cell lines, and CellOracle’s Prior achieves above-chance performance in three out of six, indicating less consistent cross-context behavior. Yeast represents the main exception, where CellOracle produces the highest-performing prior (AUPRC 0.05 versus 0.03 for Inferelator-Prior and 0.02 for GLM-Prior), highlighting the strengths of accessibility-based priors in simple eukaryotic systems where regulatory interactions are predominantly promoter-proximal and well captured by open

130

chromatin-based assignment. In this setting, GLM-Prior is further disadvantaged by the limited number of positive regulatory interactions (660) available for model training (as seen in Figure ??D). A broader comparison across species using normalized AUPRC values further illustrates these patterns (4.7B), where normalized AUPRC is defined as (AUPRC − chance)/(1 − chance). While GLM-Prior provides the most consistent gains overall, CellOracle’s prior outperforms all other approaches in HepG2 (0.36 versus 0.30 for GLM-Prior and 0.24 for Inferelator-Prior). This comparatively strong performance in a mammalian cell line can be explained by alignment between the construction of CellOracle’s prior and the BEELINE reference network used for evaluation. BEELINE assembles reference networks using ChIP-seq supported TF-gene edges curated from ENCODE [de Souza 2012], ChIP-ATLAS [Zou et al. 2024] and ESCAPE [Xu et al. 2013] databases, where HepG2 is among the most densely profiled ENCODE cell line. As a result, the HepG2 reference likely contains a large and internally consistent set of TF occupancy-supported edges that are recoverable by accessibility-conditioned, proximity-based linking, precisely the evidence CellOracle prioritize when constructing its prior from motifs in accessible regions. By contrast, Inferelator-Prior follows a similar accessibility and motif-based strategy but imposes stronger sparsity which may remove a substantial number of true TF-gene edges in this densely profiled HepG2 context and thereby reduce its alignment with the BEELINE reference network. In contrast, GLM-Prior is the top-performing method in the four remaining cell lines: hESC, mESC, mDC, and mHSC. In hESC, all methods exceed chance, with GLM-Prior achieving the highest AUPRC (0.19 versus 0.17 for Inferelator-Prior and 0.16 for CellOracle’s prior). In mESC and mHSC, the advantage of GLM-Prior becomes more pronounced, as accessibility-based priors fall below chance while GLM-Prior remains predictive. These results suggest that in several mammalian contexts, sequence-based priors may better capture distal enhancer regulation and long-range TF-gene interactions that proximity-based methods miss. These comparisons indicate that accessibility-driven approaches such as CellOracle remain 131

Figure 4.7: Comparison of prior construction strategies across species. (A) AUPRC of GLM-Prior, Inferelator-Prior and CellOracle Prior across six cell lines, with chance performance (gray) provided for each context. (B) Normalized AUPRC values with the best-performing prior (green box) highlighted in each cell line.

132

highly effective in contexts dominated by proximal regulatory logic, whereas GLM-Prior provides a more robust and transferable strategy across diverse mammalian systems. The complementary performance profiles across species highlight that no single prior construction approach is universally optimal; instead, GLM-Prior offers a scalable, sequencing technology-independent alternative that is better suited to capturing complex long-range regulatory architecture and transferring information across closely related species and cell lines.

4.3.6

Paired prior-GRN strategies reveal prior-dependent gains from expression-based inference

Following our comparative analysis of prior construction approaches, we next asked how each complete prior-GRN pipeline performs across the six cell lines (Figure 4.8 and 4.9). GLM-Prior and PMF-GRN form our dual-stage pipeline, combining a finetuned sequence language model for prior construction with a probabilistic matrix factorization model for GRN inference from single cell expression. CellOracle constructs a prior ("CellOracle baseGRN") using motif enrichment within accessible promoter regions and pairs it with a Bayesian ridge regression model for GRN inference. Similarly, Inferelator-Prior uses chromatin accessibility and TF motif enrichment to build a prior edge matrix that is then refined using regularized regression during GRN inference. By comparing these three paired pipelines, GLM-Prior + PMF-GRN, Inferelator-Prior + Inferelator, and CellOracle prior + CellOracle, we can ask two related questions: (𝑖) how well each method converts its own prior into a high-performing GRN, and (𝑖𝑖) to what extent expressionbased GRN inference can compensate for poor priors versus refining already strong ones. To interpret these patterns, we consider the absolute AUPRCs in Figure 4.8A together with the corresponding ΔAUPRC values in Figure 4.8B. In yeast and mDC, all priors are at or only slightly above chance, and GRN inference yields minimal gains (ΔAUPRC ≤ 0.02), indicating that expression-based models struggle to recover meaningful signal when the starting network is

133

near-random. In hESC and mHSC, GLM-Prior already achieves above-chance performance, and PMF-GRN adds only small improvements, whereas Inferelator and CellOracle modestly refine their weaker priors. These contexts illustrate that when priors are already strong, GRN inference primarily fine-tunes edge weights rather than substantially reshaping the network.

Figure 4.8: Paired prior-GRN strategies reveal prior dependent gains from expression-based GRN inference across six cell lines. (A) AUPRC of each prior (GLM-Prior, Inferelator-Prior, and CellOracle’s prior) and its corresponding GRN (PMF-GRN, Inferelator, CellOracle) across six cell lines, with gray dashed lines indicating the chance level in each species. (B) Change in AUPRC from prior to GRN (ΔAUPRC) for each paired method, quantifying the added value of expression-based inference.

By contrast, HepG2 and mESC are the settings where GRN inference provides the largest benefits. Here, priors span a range from near- or slightly below-chance (Inferelator-Prior in HepG2; Inferelator-Prior and CellOracle prior in mESC) to moderately above-chance (GLM-Prior and CellOracle prior in HepG2; GLM-Prior in mESC), leaving partially specified structure that expression data can refine. In HepG2, GLM-Prior + PMF-GRN and Inferelator-Prior + Inferelator 134

show sizeable increases in AUPRC, while CellOracle begins from the strongest prior and improves only modestly. In mESC, regression-based inference on top of weak or near-chance priors yields the single largest observed ΔAUPRC for CellOracle and substantial gains for Inferelator, whereas PMF-GRN makes only a small but consistent improvement over an already above-chance GLM-Prior. Taken together, these six cell lines reveal a coherent pattern: the magnitude of GRNinference gains is strongly prior-dependent, with the largest boosts arising when priors are of intermediate quality (informative but incomplete), limited performance improvements when priors are very weak, and only modest refinements when priors are already strong. We next normalize AUPRC by the chance level in each species (as in Section 4.3.5) to directly compare performance above (or below) chance across methods and cell lines (Figure 4.9A). In this view, GLM-Prior provides the best prior in four of six cell lines (hESC, mESC, mDC, mHSC), while CellOracle prior yields the strongest priors in yeast and HepG2. We observe that in four mammalian cell lines (HepG2, mESC, mDC, and mHSC), Inferelator-Prior rarely exceeds random-chance performance and typically remains close to or below chance. Similarly, in all three mouse cell lines, CellOracle’s prior does not construct a prior above chance. Here, complex and long-range regulatory interactions may not be sufficiently captured by these proximitybased approaches, highlighting a key limitation of accessibility-based prior knowledge in complex mammalian cell lines. In addition, we find that the best GRNs largely track the best priors, with PMF-GRN producing the top-performing GRN in three of six cell lines, CellOracle in two, and Inferelator in one (Figure 4.9A). Notably, there are instances where GRN inference can partially compensate for a weaker prior. In mESC, CellOracle’s prior is below chance (normalized AUPRC = −0.04), yet its regression-based GRN achieves the highest normalized performance across all methods (0.21), outperforming GLM-Prior + PMF-GRN despite starting from an inferior prior. This behavior is clarified by examining how each GRN model modulates its prior in terms of edge flips (Figure 4.9B). 135

PMF-GRN primarily improves performance by adding edges to the prior, effectively expanding the regulatory network in a data-informed way while leaving most high-confidence edges intact. Inferelator both adds and removes edges, suggesting a balance between edge discovery and denoising. In contrast, CellOracle’s regression-based design only allows removal of edges: it can shrink edge weights to zero but cannot introduce new edges, discovered through expression data, that were absent from the accessibility-derived prior.

Figure 4.9: Paired prior-GRN strategies reveal prior dependent gains from expression-based GRN inference across six cell lines continued. (A) Normalized AUPRC performance for priors and GRNs in each cell line. Green boxes highlight the best prior and blue boxes highlight the best GRN per cell line. (B) Edge flips between prior and GRN for each method demonstrating how GRN inference methods modulate their priors.

The large performance gains observed for CellOracle in settings such as mESC therefore arise from aggressive pruning of false-positive edges in a dense, noisy prior, rather than from discovery 136

of novel regulatory interactions. This pruning behavior can make a weak but overly dense prior appear substantially stronger after GRN inference, yet also highlights an important limitation: when single-cell expression is used only to filter a fixed network, not to propose new edges, it cannot expand the regulatory scaffold beyond what is already encoded in the prior. Together, these results demonstrate that while GRN inference can substantially improve performance in some settings, particularly when priors are informative but incomplete, the overall ranking of methods is largely dictated by prior quality. Expression-based GRN inference models mostly refine and re-weight the structure supplied by the prior, rather than overturning it, underscoring the central role of prior construction in determining GRN reconstruction performance.

4.3.7

Cross-method comparison disentangles prior and GRN inference contributions to performance

To disentangle the contribution of prior construction and GRN inference from the GRN methods themselves, we next performed a full cross-comparison in which each GRN inference method is applied to each prior across all six cell lines (Figure 4.10 and 4.11). Specifically, we combine GLM-Prior, Inferelator-Prior, and CellOracle baseGRN with PMF-GRN, Inferelator, and CellOracle, yielding nine prior-GRN combinations per cell line. In Figure 4.10A, we show the resulting AUPRC values, along with species-specific chance levels (gray dashed line), and the baseline performance of the prior (colored dashed line) used for GRN inference. In Figure 4.11A, we report the same results normalized relative to chance, where values above zero indicate performance above random. We focus our interpretation on the normalized AUPRC values, which allows for easier comparison across species and methods, and refer to the raw AUPRCs in Figure 4.10A where relevant. Across species, the dominant pattern is that the identity of the best GRN for a given cell line is largely determined by the underlying prior, not by the choice of GRN inference method. In

137

Figure 4.10: Cross-method comparison disentangles prior and GRN inference contributions. (A) AUPRC for all nine combinations of prior and GRN inference methods across six cell lines. Gray dashed lines indicate chance performance, while light blue, light green, and light pink dashed lines represent baseline performance of GLM-Prior, Inferelator-Prior, and CellOracle prior, respectively, across each cell line.

138

yeast, for example, any GRN method paired with CellOracle’s prior attains the same normalized AUPRC (0.05), whereas all combinations using GLM-Prior or Inferelator-Prior remain at or below 0.03. Once CellOracle’s prior is selected, the choice among PMF-GRN, Inferelator, or CellOracle has no quantifiable impact, indicating that the prior itself, rather than the choice of downstream inference algorithm, is the bottleneck of GRN inference performance.

Figure 4.11: Cross-method comparison disentangles prior and GRN inference contributions continued. (A) Normalized AUPRC of all nine combinations of prior and GRN for directly comparing across method and cell line. Blue boxes indicate highest GRN performance per cell line.

Similar patterns appear in the mammalian cell lines. In hESC, GLM-Prior and InferelatorPrior yield the strongest priors (0.19 and 0.17 AUPRC, respectively), and nearly all GRNs inferred from these priors achieve competitive performance. In HepG2, CellOracle’s prior is best (AUPRC = 0.36), and the two highest performing GRNs are obtained by pairing this prior with PMFGRN or CellOracle, demonstrating again that performance is largely determined by the prior rather than the choice of inference algorithm. An analogous effect is observed in mHSC, where GLM-Prior combined with PMF-GRN yields the highest performing combination (AUPRC = 0.24), while all GRNs built on Inferelator-Prior or CellOracle’s prior plateau at 0.08 or lower. These results support a consistent conclusion, the quality and structure of the prior largely dictate the achievable performance, and the GRN inference methods primarily modulate performance within a range set by the prior. 139

The full cross-comparison design also allows us to ask whether any GRN inference method is inherently superior across priors and species. We find that across PMF-GRN, Inferelator, and CellOracle, the overarching evidences suggests no GRN inference method is superior. Each approach wins within different regimes defined by the combination of species and prior. PMF-GRN delivers the highest normalized performance in HepG2 when paired with CellOracle’s prior (0.20) and in mHSC when paired with GLM-Prior (0.24). PMF-GRN also performs well in yeast when paired with CellOracle’s prior, and mESC when paired with GLM-Prior or Inferelator-Prior. The Inferelator is best in hESC when paired with Inferelator-Prior (0.06) and in mDC when paired with CellOracle’s prior (0.02), and remains competitive in HepG2 across multiple priors. CellOracle is uniquely optimal in mESC when paired with CellOracle’s prior (0.21), substantially outperforming all other prior-GRN combinations in that species. The ranking of each GRN inference method thus changes with both prior and cell line, consistent with our claim that GRN methods are context-dependent refiners of information encoded in a prior using expression data, rather than the primary driver of performance with a single universally superior algorithm. The cross-method comparison further reveals characteristic interactions between each GRN inference method and the priors. PMF-GRN appears highly prior sensitive, where switching priors can produce large swings in performance within a given species. In mHSC, for instance, normalized performance for PMF-GRN ranges from −0.05 with CellOracle-Prior to 0.24 with GLM-Prior. In mESC, performance ranges from −0.04 with CellOracle Prior to 0.09 with GLMPrior. This pattern indicates that PMF-GRN is very efficient at exploiting high-quality priors, but struggles to rescue very weak ones. The Inferelator, by contrast, is comparatively robust to prior choice. Within each species, the variation in normalized AUPRC across priors is typically modest (often within the range of 0.02 to 0.04). In mHSC, all three priors yield identical normalized performance (0.08) with the Inferelator. This suggests that the Inferelator leans more heavily on expression-based regularized regression, making it less sensitive to exact prior structure, but also preventing it from reaching the highest performance in most cell lines. CellOracle exhibits 140

strong synergy with its own prior in specific settings. For example, in mESC, CellOracle’s prior + Cell Oracle perform well (0.21), significantly outperforming all other combinations, indicating that the regression-based pruning used by CellOracle is particularly well matched when the prior necessitates the removal of false positive edges. As each GRN inference method is applied to every prior, we can also ask whether cross-pairing ever allows a "strong" GRN inference method to elevate a weaker prior above combinations that use a stronger prior with a different inference method. Such reversals are rare. In mHSC, for example, no method applied to Inferelator-Prior or CellOracle prior achieves performance comparable to GLM-Prior + PMF-GRN. In HepG2, there is a modest example of beneficial cross-pairing, where CellOracle prior + PMF-GRN achieves an AUPRC of 0.20, outperforming the original CellOracle prior + CellOracle pairing (0.17). The performance gains here are likely due to PMF-GRN providing additional edges to the CellOracle’s prior using expression data. In mDC, CellOracle prior + Inferelator (0.02) is better than CellOracle prior + CellOracle (0) or CellOracle prior + PMF-GRN (0.02), although all values remain near chance. These exceptions are incremental, not complete reversals of the prior-driven ranking, and reinforce the view that GRN inference can fine-tune performance but typically cannot turn a weak prior into one that rivals the best priors. Overall, the cross-method comparison in Figure 4.10 and 4.11 supports three main conclusions. First, prior quality remains the primary determinant of GRN reconstruction performance, even when each GRN method is given access to each prior. Second, there is no universally better GRN inference algorithm: PMF-GRN, the Inferelator, and CellOracle perform best in different prior and species regimes, reflecting differences in how they use expression data to refine or reweight the prior. Third, GRN inference methods modulate, rather than overturn, the structure imposed by the prior. Changing the inference algorithm can yield meaningful gains in some contexts, particularly when paired with an appropriate prior, but rarely elevates a weak prior to the level of a strong one. These findings reinforce our claim that prior construction is the main driver of GRN inference performance, with expression-based GRN inference primarily serving a 141

prior-dependent refining role.

4.4

Methods

4.4.1

The GLM-Prior model

We developed GLM-Prior, a genomic language model fine-tuned to predict transcription factor (TF)-target gene regulatory interactions directly from sequence data. Specifically, we fine-tuned the 250 million parameter Nucleotide-Transformer model [Dalla-Torre et al. 2024], originally pretrained on the genomes of 850 species, to infer binary interaction matrices between TFs and their target genes using their nucleotide sequences. These predicted interactions serve as a priorknowledge for downstream GRN inference, introducing biologically grounded constraints that guide the structure of the inferred networks. The model uses a transformer encoder architecture to process two distinct sequence types, TF motif sequences from the CisBP database [Weirauch et al. 2014] and gene body sequences derived from genome annotations in a GTF file. We append a classification head to the transformer encoder to predict whether each TF-gene pair represents a true regulatory interaction (positive label) or non-regulatory interaction (negative label). Training labels are obtained from experimentally validated interaction databases such as STRING [Mering et al. 2003; Szklarczyk et al. 2010, 2021, 2023] and TRRUST [Han et al. 2015, 2018], supplying high confidence positive and negative examples of regulatory interactions. This design allows the model to leverage both sequence-level regulatory motifs encoded in the pre-trained transformer and database-derived TF-gene interactions to improve the predictive performance and generalization across genomic contexts. Each training example is constructed by concatenating the nucleotide sequence of a TF with that of its candidate target gene, separated by a special classification token (<cls>). This compos-

142

ite sequence is tokenized, during which a model-specific (<cls>) token is automatically prepended to the input. The tokenized sequence is passed through the pre-trained transformer, which produces contextual embeddings across the full input. The embedding corresponding to the prepended (<cls>) token is fed into the classification head composed of a dropout layer, a linear projection, a tanh activation, a second dropout, and a final linear layer that maps to two logits representing the binary interaction label. This setup allows the model to jointly learn representations of TF-binding specificity and gene regulatory potential in a unified sequence-to-interaction framework. 4.4.1.1

Dataset construction and preprocessing

To train the model, we constructed a dataset of TF-gene pairs where positive examples were derived from a curated database of experimentally validated interactions (YEASTRACT [Teixeira et al. 2018], STRING [Mering et al. 2003; Szklarczyk et al. 2010, 2021, 2023] and TRRUST [Han et al. 2015, 2018]), and negative examples were sampled from a combination of all remaining TF-gene sequence pairs. This dataset was split into a 99% training and 1% validation set. Given the substantial class imbalance between positive and negative samples, we implemented a downsampling strategy to reduce the number of negative samples in the training set. This preserved the diversity of negative samples, while preventing the model from overfitting to the negative class. Specifically, the number of retained negative samples after downsampling can be defined as,

𝑁𝑠𝑎𝑚𝑝𝑙𝑒𝑑 = ⌊𝑟 · 𝑁 − ⌋,

(4.1)

where 𝑟 ∈ (0, 1) is the downsampling rate, and 𝑁 − is the total number of negative examples in the dataset.

143

To further address class imbalance during training, we designed a custom DataLoader that performs even class batch sampling. Each batch was constructed to contain an equal number of positive and negative samples, ensuring a balanced signal during training. Specifically, each batch is defined as,

𝐵 = {(𝑥𝑖+, 𝑥 −𝑗 ) : 𝑥𝑖+ ∈ 𝑋 +, 𝑥 −𝑗 ∈ 𝑋 − },

(4.2)

where positive examples 𝑥𝑖+ were sampled uniformly with replacement from 𝑋 + (the positive class) and negative examples 𝑥 −𝑗 were cycled through without replacement:

𝑥𝑖+ ∼ 𝑋 +,

𝑥 −𝑗 ∼ 𝑋 −

with

𝑗 mod 𝑁 − .

(4.3)

This strategy ensures that each negative example was seen exactly once during training while positive examples were reused as needed to maintain class balance. Even-class batching was critical for stabilizing training dynamics and improving the model’s sensitivity to true positive interactions without inflating the false-negative rate. 4.4.1.2

Model Architecture and Classification

The model architecture builds on the pre-trained Nucleotide Transformer by retaining its embedding and transformer encoding layers. A classification head is applied to the final hidden state corresponding to the prepended (<cls>) token to compute logits for a binary classification task. This classification head outputs a two-dimensional logit vector 𝑧 ∈ R2 representing unnormalized scores for the positive and negative classes,

144

𝑧 = 𝑊 · ℎ + 𝑏,

(4.4)

where ℎ is the transformed hidden state of the (<cls>) token following the intermediate projection and non-linearity, and 𝑊 , 𝑏 are the weights and bias of the final linear layer. The components of 𝑧 are denoted 𝑧 + and 𝑧 − , corresponding to the logits for the positive and negative classes, respectively. Class probabilities are computed using a softmax function,

𝑃 (𝑦 = 1|𝑥𝑇 𝐹 , 𝑥𝑔𝑒𝑛𝑒 ) =

exp(𝑧 + ) , exp(𝑧 + ) + exp(𝑧 − )

(4.5)

Training used a class-weighted cross-entropy loss function to account for residual class imbalance after downsampling. The loss assigns a fixed weight of 1.0 to positive examples, while the negative class weight 𝑤 − is tuned through hyperparameter search to optimize the balance between precision and recall. The loss is computed as,

L = −𝑤 +𝑦 log 𝑝 − 𝑤 − (1 − 𝑦) log(1 − 𝑝),

where 𝑝 =

exp(𝑧 + ) . exp(𝑧 + ) + exp(𝑧 − )

(4.6)

Here, 𝑦 ∈ {0, 1} is the true label, 𝑧 + and 𝑧 − are the predicted logits for the positive and negative classes, respectively. 𝑤 − is the down-weighted negative class weight, varied during hyperparameter search, and the positive class weight 𝑤 + is fixed at 1.0. This approach maintains sensitivity to positive predictions while mitigating bias toward the negative class.

145

4.4.1.3

Training procedure

We trained the model on 4 H100 GPUs using PyTorch’s Distributed Data Parallel (DDP) framework [Li et al. 2020c] to enable efficient multi-GPU scaling. The training process was distributed across GPUs to accelerate computation and ensure consistent gradient updates. We used a perdevice batch size of 32 and set the gradient accumulation steps to 32, resulting in an effective batch size of 4096. The model was optimized using Adam with a learning rate of 10−5 . Training spanned 10 epochs, using optimal hyperparameters selected through a sweep over the negative class weight (𝑤 − ) and downsampling rate (see Appendix 4.5 for more details). We selected the configuration that achieved the highest F1 score on the validation set for final training. After training, the model was used to infer a prior-knowledge matrix of TF-gene regulatory interactions from a list of gene-TF sequence pairs without labels. This matrix then served as input for downstream GRN inference, where GRN inference can further tailor these language model derived interactions using cell-type, cell-line, or condition-specific expression data. 4.4.1.4

Performance and Evaluation

We evaluated model performance using standard binary classification metrics, with a focus on metrics that remain robust under class imbalance. Specifically, we report precision, recall, F1 score, area under the receiver operating characteristic curve (AUC-ROC), area under the precisionrecall curve (AUPRC), and Matthews correlation coefficient (MCC). Precision and recall were computed separately for the positive and negative classes to assess the model’s ability to minimize false positives and false negatives, respectively. Let 𝑇 𝑃, 𝐹 𝑃, 𝐹 𝑁 and 𝑇 𝑁 denote the true positives, false positives, false negatives and true negatives. Then,

Precision =

𝑇𝑃 , 𝑇𝑃 + 𝐹𝑃

146

Recall =

𝑇𝑃 . 𝑇𝑃 + 𝐹𝑁

(4.7)

The F1 score, which represents the harmonic mean of precision and recall, was used as the primary metric for model selection and hyperparameter optimization. It can be computed as,

𝐹1 = 2 ·

Precision · Recall . Precision + Recall

(4.8)

To account for class imbalance and provide a threshold-independent measure of performance, we also computed AUC-ROC and AUPRC. the ROC curve plots true positive rate (TPR) against false positive rate (FPR), defined as:

True Positive Rate (TPR) =

𝑇𝑃 , 𝑇𝑃 + 𝐹𝑁

False Positive Rate (FPR) =

𝐹𝑃 . 𝐹𝑃 + 𝑇 𝑁

(4.9)

While AUC-ROC captures the model’s general discrimintative ability, AUPRC is more informative in imbalanced settings, as it directly reflects the trade-off between precision and recall. We used AUPRC to benchmark model predictions against curated gold standard datasets of TF-gene interactions. To determine the optimal classification threshold, we performed a grid search over the predicted positive class probabilities. The threshold 𝑡 ∗ that maximized the F1 score on the validation set was selected for final inference,

𝑡 ∗ = arg max 𝐹 1 (𝑡)

(4.10)

Finally, we report the Matthews correlation coefficient (MCC), a balanced measure of classification quality that incorporates all four confusion matrix components,

147

MCC = √︁

𝑇𝑃 · 𝑇 𝑁 − 𝐹𝑃 · 𝐹𝑁 . (𝑇 𝑃 + 𝐹 𝑃)(𝑇 𝑃 + 𝐹 𝑁 )(𝑇 𝑁 + 𝐹 𝑃)(𝑇 𝑁 + 𝐹 𝑁 )

(4.11)

MCC ranges from −1 to +1, where +1 indicates perfect predictions, 0 indicates random predictions, and −1 indicates incorrect predictions. MCC remains informative even when classes are highly imbalanced, making it a useful complement to F1 and AUPRC.

4.4.2

GRN inference with PMF-GRN

We performed GRN inference using our previously published method, PMF-GRN (Probabilistic Matrix Factorization for Gene Regulatory Network inference) [Skok Gibbs et al. 2024] to infer the regulatory interactions between TFs and their target genes. The goal of PMF-GRN is to decompose an observed gene expression matrix into latent factors that represent TF activity and regulatory interactions between TFs and their target genes. These latent factors capture the underlying GRN structure, which cannot be measured directly from gene expression data alone. Further details regarding the PMF-GRN model and the inference strategy used to obtain GRNs can be found in [Skok Gibbs et al. 2022]. Using PMF-GRN, we perform inference independently on each single-cell dataset to obtain dataset-specific GRNs. These inferred networks are then combined post-inference using a simple averaging strategy to produce a consensus GRN,

𝑁

GRNConsensus =

1 ∑︁ GRN𝑖 , 𝑁 𝑖=1

(4.12)

where 𝑁 is the number of datasets and GRN𝑖 is the inferred network for dataset 𝑖. This consensus GRN captures regulatory interactions that are consistently inferred across datasets,

148

while preserving dataset-specific networks for context-specific analyses. To evaluate the relationship between posterior uncertainty and predictive accuracy, we rank all TF-gene interaction pairs by their posterior variance, constructing 10 cumulative bins corresponding to the variances in increments of 10%. For each bin 𝑘, we use the posterior point estimates of the interactions within the bin and compute the AUPRC using a gold standard set of validated TF-gene interactions G as the reference. All interactions in each bin are included in the evaluation, regardless of whether they appear in the gold standard. Let B𝑘 be the set of posterior point estimates in the 𝑘% of variances. Then the AUPRC for bin 𝑘 is given by:

AUPRC𝑘 = AUPRC(B𝑘 ; G),

(4.13)

where the AUPRC is computed by comparing the predicted scores in B𝑘 to labels derived from G. The cumulative binning strategy ensures that each bin contains a sufficient number of interactions for stable AUPRC estimation.

Discussion Accurately reconstructing GRNs remains a central challenge in genomics, particularly in complex systems where direct experimental measurements are incomplete or unavailable. In this work, we present GLM-Prior, a genomic language model fine-tuned to predict TF-gene regulatory interactions directly from DNA sequence, and integrate it with PMF-GRN, a probabilistic matrix factorization framework for expression-based GRN inference. This dual-stage design decouples prior construction from downstream GRN inference, allowing us to evaluate how different priors and GRN inference algorithms jointly shape GRN reconstruction across yeast, mouse, and human

149

cell lines. Our results first demonstrate that GLM-Prior’s performance is tightly linked to the composition of the training data. Across six cell lines spanning yeast, human, and mouse, predictive accuracy scales with the abundance and diversity of positive labels and TF coverage. Human and mouse cell lines with rich, well-annotated regulatory datasets support substantially higher AUPRCs than yeast, where extreme class imbalance and sparse labels limit generalization despite good within-training metrics. These trends highlight that even with powerful sequence models, generalization is fundamentally constrained by the available supervision. Here, we find that when only a small subset of true interactions are labeled, the model can capture some regulatory grammar but cannot fully resolves novel gene-TF combinations in held-out contexts. By systematically varying training paradigms, we further show that GLM-Prior generalizes in a biologically interpretable manner across species. Transfer learning between evolutionarily related species, such as human and mouse, maintains or modestly improves performance relative to single-species models, indicating that GLM-Prior captures conserved sequence-level features that can be reused across lineages. In contrast, models trained in mammals transfer poorly to yeast, where performance remains near chance, consistent with deeper divergence in regulatory grammar, TF binding preferences, and genome organization. Multi-species training on human, mouse, and yeast preserves accuracy comparable to the best single-species or transfer models in most cell lines, with only mild trade-offs in some human contexts, and yields a domain-stable prior that can be applied to new systems without species-specific retraining. Together, these findings establish GLM-Prior as a flexible framework for constructing priors across diverse data regimes, with strong performance in well-annotated human and mouse cell lines and more limited gains in sparsely labeled settings such as yeast. Furthermore, benchmarking GLM-Prior against widely used accessibility-based priors, including Inferelator-Prior and CellOracle’s baseGRN, across six cell lines demonstrates that GLMPrior consistently performs at or above chance and provides the strongest prior in four of six 150

contexts (hESC, mESC, mDC, and mHSC). At the same time, accessibility-based priors remain competitive and even superior in settings dominated by proximal regulatory logic, such as yeast, where promoter-focused chromatin accessibility and motif enrichment capture a large fraction of regulatory edges. In more complex mammalian systems, accessibility-based priors often fall below chance while GLM-Prior remains predictive, suggesting that sequence-based models may better capture distal enhancer regulation and long-range TF-gene coupling that are not easily recovered from proximity-based peak-gene assignment. These complementary performance profiles indicate that GLM-Prior provides the most robust and transferable prior knowledge across diverse mammalian contexts, while accessibility-based priors remain particularly effective in promoterdominated scenarios such as yeast, together capturing distinct and complementary facets of regulatory architecture. Integrating these priors with multiple GRN inference algorithms reveals a consistent pattern: the quality and structure of the prior largely determines the achievable performance, and GRN inference algorithms mostly modulate performance within a range set by the prior. First, when we use GLM-Prior as input to PMF-GRN, AUPRC is equal to or higher than the prior alone in all six cell lines, with modest gains when the prior is already strong and larger improvements when the prior is weaker. Edge-flip analyses in this setting show that PMF-GRN typically makes relatively small, targeted changes, treating the sequence-derived prior as a stable scaffold and adding or pruning a limited subset of edges rather than reconstructing the network from scratch. We then extend this analysis to the full cross-method setting, in which all three priors (GLM-Prior, Inferelator-Prior, and CellOracle’s baseGRN) are combined with all three GRN inference methods (PMF-GRN, Inferelator, and CellOracle). In this fully crossed design, the best-performing GRNs in each cell line almost always arise from combinations that use the strongest prior for that context, regardless of which inference algorithm is applied. At the same time, no GRN inference method is uniformly superior across priors or species. Each GRN algorithm performs best in different prior–species regimes and interacts with priors in characteristic ways, from highly prior-sensitive 151

behavior to more robust, expression-driven refinement. Together, these analyses indicate that GRN performance is primarily limited by prior construction, with inference algorithms providing prior-dependent gains rather than defining the overall performance ceiling. Our results recast GRN inference as a problem in which the central bottleneck is the construction of high-quality prior knowledge, rather than the specific choice of the downstream expression-based inference algorithm. GLM-Prior addresses this bottleneck by leveraging a transformerbased genomic language model to learn regulatory grammar directly from DNA, enabling priors that generalize across cell types, species, and training paradigms. Critically, because GLM-Prior operates directly on genome sequence, it removes the dependence on cell-type- and assay-specific measurements such as ATAC-seq, and thus makes it possible, in principle, to construct regulatory priors in understudied or experimentally inaccessible species where chromatin accessibility or expression data are sparse or unavailable. While accessibility-based priors remain valuable in contexts dominated by relatively simple, proximal promoter–gene regulation, they struggle to capture regulatory architecture in more complex mammalian cell lines and cannot be readily transferred to unseen species or cell types, in contrast to sequence-based priors such as GLMPrior. Several limitations of our study point to promising directions for future work. First, GRN benchmarks are constrained by incomplete and noisy interaction databases, in which many negative labels likely represent unknown or untested edges rather than true absences of regulation. This incompleteness affects both prior construction and evaluation, and motivates the development of benchmarks that better distinguish between confirmed negatives and unlabeled interactions. Further, as GLM-Prior downsamples negative examples from the input training data, we likely discard false negative edges during training and thus never learn these regulatory interactions using our sequence-based approach. Future work could include more sampling techniques to avoid losing access to interactions previously defined as negative example that could be in fact classified as positive examples during training. Second, our models are trained on a limited 152

set of cell lines and regulatory labels. Extending GLM-Prior to broader compendia of cell types, developmental stages, and species will be important for assessing how far sequence-based priors can be pushed as a general regulatory scaffold that other methods can refine. Third, while we consider multiple GRN inference algorithms, they share a common edge-centric view of network reconstruction. Exploring integration strategies that move beyond edge-wise scoring, such as jointly modeling regulatory modules, trajectories, or causally inferred dynamics [Kim et al. 2025], may reveal additional ways that expression data can complement strong sequence-derived priors. In summary, this work introduces GLM-Prior as a scalable, sequence-based framework for constructing regulatory priors and systematically characterizes how these priors interact with GRN inference methods across yeast, mouse, and human. By demonstrating that prior quality is the dominant determinant of GRN reconstruction performance, our study shifts the emphasis in GRN modeling toward learning robust, generalizable priors from sequence, and provides a foundation for future methods that integrate these priors with expression and other modalities in more flexible and biologically grounded ways.

4.5

Supplementary Material for Chapter 4

4.5.1

Single-Species Experiments

4.5.1.1

Yeast

To train our GLM-Prior model in yeast, we first obtained all 5, 999 gene body nucleotide sequences from the ENSEMBL S. cerevisiae (R64-1-1.UTR.gtf) genome. We obtained 212 TF sequences from the CisBP database [Weirauch et al. 2014] under S. cerevisiae. Next, to pair these gene and TF nucleotide sequences with interaction labels, we used the YEASTRACT database of interactions [Teixeira et al. 2018] (6, 885 genes by 220 TFs). From these 153

220 YEASTRACT TFs, 46 did not have sequences associated with them from CisBP. Due to this large portion of data loss (20%), we used the promoter regions of the target genes for each missing TF as a proxy for it’s binding sequence. We defined the promoter sequence following YEASTRACT’s definition of 1000bp upstream or downstream of the gene TSS (depending on the strand orientation). Training for yeast was conducted across 5, 893 genes and 123 TFs, with 660 positive labels and 724, 179 negative labels for these interactions derived from YEASTRACT. Validation Metrics for Yeast Single-Species GLM-Prior Metric Score Best F1 Score 0.76 ROC AUC 1.00 Best Classification Threshold 1.00 Positive Class Precision 0.89 Positive Class Recall 0.67 Negative Class Precision 1.00 Negative Class Recall 1.00 AUPRC (vs. held-out set [Tchourine et al. 2018]) 0.02 Table 4.1: Validation performance of the GLM-Prior model trained on yeast. Evaluation was conducted on a held-out validation set that assess positive and negative class contributions during training (top) and AUPRC against an independent gold standard for the final inferred prior-knowledge matrix (bottom).

A hyperparameter sweep over 1 epoch of training using different class-weights and downsampling rates for the negative class revealed [0.5, 1.0] to be the optimal class-weights and 0.4 to be the optimal negative class downsampling rate. These hyperparameters were used during final training over 10 epochs, obtaining the validation metrics on the held-out 1% of training data found in Table 4.1. Following training of the yeast single-species GLM-Prior model on YEASTRACT database labels, we ran inference on all genes and TFs, including those in a held-out set not seen during training. These held-out genes and TFs corresponded to those present in the gold standard from [Tchourine et al. 2018]. Our held-out set comprised 933 genes and 98 TFs, with 986 positive labels, and 95, 911 negative labels. After inference, the model predictions were evaluated using AUPRC 154

with the corresponding labels from the gold standard, with a chance baseline calculated by taking the positive rate of the held-out set. 4.5.1.2

Seqence homology experiments in yeast

Sequence homology refers to shared sequence similarity between loci, often introduced by duplication and other genome rearrangements. Because nucleotide language models are optimized to exploit recurring sequence patterns, high homology across train-test splits can, in principle, make held-out prediction easier by allowing the model to match familiar fragments rather than rely on fully generalizable features. For GLM-Prior, which is trained on labeled TF-sequence pairs, this motivates a simple check: whether genes or TFs are unusually similar to sequences seen during training, and whether controlling for that similarity changes performance at inference time. We therefore use hashFrag [Rafi et al. 2025], which leverages BLAST [Madden 2013] to quantify cross-split sequence similarity and generate homology-aware partitions (or prune highly similar training sequences), and then re-evaluate GLM-Prior on these homology-controlled splits to test whether sequence similarity contributes to our yeast performance. In Figure 4.12A we depict the homology-aware partitioning scheme. We then quantify leakage among genes by computing, for each test gene, its maximum BLAST score to any training gene (Figure 4.12B-C). Using an operational cutoff 𝑡 = 150, and holding the test set fixed (𝑛 = 978), pruning 346 homologous training genes reduces the fraction of "leaking" test genes from 24.8% to 7.8%. We next ask whether removing this homology affects model accuracy. Training on the hashFragdefined gene partitions yields a test AUPRC of 0.03, compared to 0.01 chance performance (Figure 4.12D). This magnitude is comparable to the yeast results (AUPRC = 0.02) reported in Section 4.3.2, indicating that sequence leakage is not responsible for our observed performance in yeast. We repeat the same analysis for TF sequences (Figure 4.12E-H). As with genes, hashFrag pruning decreases cross-split similarity, and model performance on the homology-clean TF splits 155

remains similar to the main yeast model, again motivating against leakage-driven gains.

Figure 4.12: Sequence-homology analysis of yeast gene and TF splits using hashFrag. (A) Schematic of homology-aware partitioning, illustrating how highly similar sequences are identified and separated to reduce cross-split homology between training and test sets. (B) Histogram of per-test-gene maximum BLAST alignment score to any training gene, shown before (blue) and after (orange) hashFrag pruning of the training gene set. Blue dashed line indicates the similarity threshold used to define leakage. (C) Leakage curve for genes showing, for each similarity threshold 𝑡, the fraction of test genes whose maximum BLAST score to any training gene is ≥ 𝑡. Blue dashed line marks the leakage threshold. (D) AUPRC evaluation of GLM-Prior trained and tested on the hashFrag-defined gene partitions, with the gray dot indicating chance performance. (E) Schematic of homology-aware partitioning for TF sequences, analogous to panel A. (F) Histogram of per-test-TF maximum BLAST alignment score to any training TF, shown before (blue) and after (orange) hashFrag pruning of the training TF set. Blue dashed line indicates the similarity threshold. (G) Leakage curve for TFs showing, for each similarity threshold 𝑡, the fraction of test TFs whose maximum BLAST score to any training TF is ≥ 𝑡. Blue dashed line marks the leakage threshold. (H) AUPRC evaluation of GLM-Prior trained and tested on the hashFrag-defined TF partitions, with the gray dot indicating chance performance.

Overall, these experiments show that while homology-aware partitioning is methodologically important and successfully reduces cross-split similarity, it does not materially affect our model performance, which is primarily dependent on access to sufficient positive training examples. Because aggressive pruning removes substantial numbers of sequences, and thus labels, from both training and evaluation, it can reduce statistical power without measurable benefit here. For 156

this reason, we include this analysis in the appendix to demonstrate its value, and motivate why we avoid homology pruning in the primary experiments to preserve sufficient data for training and evaluation. 4.5.1.3

Human cell lines (hESC & HepG2)

To train our GLM-Prior model in human, we first obtained hg38 gene body nucleotide sequences from ENSEMBL (RCh38.113.gtf). Due to the lengthy nature of human genes, and the inherent limitations of context length in large language models, we filtered our gene sequences to retain all sequences for a gene body ≤ 12, 000 nucleotides in length. This provided us with a list of 23, 533 genes. We obtained TF binding motif sequences from the CisBP database [Weirauch et al. 2014], under H. sapiens. Validation Metrics for Human Single-Species GLM-Prior Metric hESC HepG2 Best F1 Score 0.29 0.35 ROC AUC 0.85 0.83 Best Classification Threshold 0.94 0.97 Positive Class Precision 0.39 0.53 0.22 0.27 Positive Class Recall Negative Class Precision 0.99 0.99 Negative Class Recall 1.00 1.00 AUPRC (vs. reference network) 0.19 0.30 Table 4.2: Validation performance of the GLM-Prior model trained on human, evaluated in hESC and HepG2. Evaluation is conducted on held-out sets using validation metrics that consider the contribution of the positive and negative classes and AUPRC against ChIP-seq-derived reference networks after inference of the prior-knowledge matrix.

Training for hESC was conducted across 1, 801 genes and 504 TFs, with 5, 305 positive labels and 902, 399 negative labels for these interactions derived from STRING and TRRUST. Training for hESC was conducted across 2, 121 genes and 527 TFs, with 8, 087 positive labels and 1, 109, 680 negative labels for these interactions derived from STRING and TRRUST. 157

A hyperparameter sweep over 1 epoch of training using different class-weights and downsampling rates for the negative class revealed [0.8, 1.0] to be the optimal class-weights and 0.3 to be the optimal negative class downsampling rate in hESC, and [1.0, 1.0] to be the optimal classweights and 0.4 to be the optimal negative class downsampling rate in HepG2. These weights were used during final training and achieved the validation metrics found in Table 4.2. Following training of the human single-species GLM-Prior model on STRING and TRRUST database labels, we ran inference on all genes and TFs, including those in a held-out set not seen during training. These held-out genes and TFs corresponded to those present in the BEELINE hESC and HepG2 reference ChIP-seq networks, respectively. For hESC, our held-out set comprised 4, 773 genes and 79 TFs, with 58, 337 positive labels, and 318, 730 negative labels. For HepG2, our held-out set comprised 4, 338 genes and 53 TFs, with 56, 567 positive labels, and 173, 347 negative labels. After inference, the model predictions were evaluated using AUPRC with the corresponding labels from the hESC and HepG2 reference ChIP-seq networks, with a chance baseline calculated by taking the positive rate of the held-out set. 4.5.1.4

Mouse cell lines (mESC, mDC, & mHSC)

To train our GLM-Prior model in mouse, we first obtained the mm10 gene body nucleotide sequences from ENSEMBL (GRCm39.113.gtf). Similarly to the single-species human experiments, we again filtered the length of our mouse genes to retain all sequences for a gene body ≤ 12, 000 nucleotides in length. This provided us with a list of 37, 755 genes. We obtained TF binding motif sequences from the CisBP database [Weirauch et al. 2014], under M. musculus. Training for mESC was conducted across 1, 097 genes and 437 TFs, with 3, 129 positive labels and 476, 260 negative labels for these interactions derived from STRING and TRRUST. Training for mDC was conducted across 2, 837 genes and 462 TFs, with 11, 551 positive labels and 1, 299, 134 negative labels for these interactions derived from STRING and TRRUST. Training for mHSC was conducted across 1, 129 genes and 419 TFs, with 2, 808 positive labels and 470, 243 negative labels. 158

Validation Metrics for Mouse Single-Species GLM-Prior Metric mESC mDC mHSC Best F1 Score 0.19 0.20 0.19 ROC AUC 0.70 0.80 0.74 0.97 0.99 0.98 Best Classification Threshold Positive Class Precision 0.23 0.17 0.50 Positive Class Recall 0.16 0.19 0.12 0.99 0.99 0.99 Negative Class Precision Negative Class Recall 0.99 0.99 1.00 AUPRC (vs. reference network) 0.23 0.09 0.49 Table 4.3: Validation performance of the GLM-Prior model trained on mouse, evaluated in mESC, mDC, and mHSC. Evaluation is conducted on held-out sets using validation metrics that consider the contribution of the positive and negative classes and AUPRC against ChIP-seq-derived reference networks after inference of the prior-knowledge matrix.

A hyperparameter sweep over 1 epoch of training using different class-weights and downsampling rates for the negative class revealed [0.2, 1.0] to be the optimal class-weights and 0.5 to be the optimal negative class downsampling rate in mESC, and [0.1, 1.0] to be the optimal classweights and 0.4 to be the optimal negative class downsampling rate in mDC, and [0.1, 1.0] to be the optimal class-weights and 0.5 to be the optimal negative class downsampling rate in mHSC. These weights were used during final training, achieving the validation metrics shown in Table 4.3. Following training of the mouse single-species GLM-Prior model on STRING and TRRUST database labels, we ran inference on all genes and TFs, including those in a held-out set not seen during training. These held-out genes and TFs corresponded to those present in the BEELINE mESC, mDC and mHSC reference ChIP-seq networks, respectively. For mESC, our held-out set comprised 5, 711 genes and 59 TFs, with 65, 933 positive labels, and 271, 016 negative labels. For mDC, our held-out set comprised 3, 373 genes and 29 TFs, with 8, 708 positive labels, and 89, 109 negative labels. For mHSC, our held-out set comprised 6, 777 genes and 72 TFs, with 182, 005 positive labels, and 305, 939 negative labels. After inference, the model predictions were evaluated

159

using AUPRC with the corresponding labels from the mESC, mDC and mHSC reference ChIP-seq networks, with a chance baseline calculated by taking the positive rate of the held-out set.

4.5.2

Transfer-Learning Experiments

Transfer learning experiments were conducted by first ensuring that no overlap existed between the training set (human, mouse, or yeast), and the corresponding inference set with held-out evaluation labels. For the human and mouse model that transferring knowledge to yeast, no information was overlapping between training and inference a priori. AUPRC Results from Transfer Learning Training Dataset Inference Dataset AUPRC Human and Mouse Yeast 0.02 Mouse hESC 0.19 Mouse HepG2 0.31 Human mESC 0.23 Human mDC 0.09 Human mHSC 0.46 Table 4.4: AUPRC scores for cross-species transfer learning. Each row indicates the species used to train the GLM-Prior model and the target species on which GRN inference was performed. Evaluation was conducted against species-specific ChIP-seq or curated reference datasets.

For the mouse trained model which transferred knowledge to hESC and HepG2, two separate models were trained, one in which overlaps between the mouse genes, TFs and labels with hESC were removed, and the second in which the overlaps between mouse genes, TFs and labels with HepG2 were removed. For the human trained model which transferred knowledge to mESC, mDC, and mHSC, three separate models were trained. In the first model, overlaps between human genes, TFs, and labels with mESC were removed. In the second model, overlaps between human genes, TFs, and labels with mDC were removed. In the third model, overlaps between human genes, TFs, and labels with mHSC were removed. Training six separate models allowed us to ensure an optimal number of gene and TF se160

quences and their corresponding labels were seen during training, while maintaining complete independence from the inference set and corresponding evaluation performance. Results for the transfer learning experiments can be found in Table 4.4.

4.5.3

Multi-Species Experiments

To train a multi-species model, we consecutively trained GLM-Prior on human, mouse, and yeast. To ensure no sequence overlap between the genes, TFs, and labels used during training and inference, we removed all overlapping inference sets from each individual organisms training set. Here, we ensured there was no inter-species leakage, as well as intra-species leakage, between training and inference sets. Following the consecutive training of the multi-species model, we ran inference in yeast, hESC, HepG2, mESC, mDC, and mHSC held-out sets, respectively. Results for this multi-species model inference can be found in Table 4.5. AUPRC Results from Multi-Species Training Training Dataset Inference Dataset AUPRC Yeast 0.02 hESC 0.18 HepG2 0.29 Human + Mouse + Yeast mESC 0.23 mDC 0.09 mHSC 0.49 Table 4.5: AUPRC scores for multi-species model inference in human, mouse, and yeast. GRNs were inferred using a unified model trained jointly on all three species, and evaluated against cell-line specific reference networks.

4.5.4

Inferelator-Prior and CellOracle baseGRN

We generated prior knowledge for Inferelator-Prior [Skok Gibbs et al. 2022] and CellOracle’s baseGRN [Kamimoto et al. 2023], using each respective methods published Python software. 161

To generate prior knowledge with Inferelator-Prior and CellOracle baseGRN for each of the six cell lines, we obtained ATAC-seq datasets from the following accessions: Human embryonic stem cells: 4DNFIPGM38K4, Human HepG2 cells: ENCFF913MQB, Mouse embryonic stem cells: 4DNFIAEQI3RP, Mouse dendritic cells: ENCFF109TUH, and Mouse hematopoietic stem cells: ENCFF931CIR. In yeast, prior knowledge datasets were obtained from [Skok Gibbs et al. 2022] without further modification. These priors were constructed using motifs from CisBP, the same motifs used in GLM-Prior to ensure compatibility and fairness during evaluation.

Cell Line Yeast hESC HepG2 mESC mDC mHSC

GLM-Prior 0.02 0.19 0.30 0.23 0.09 0.49

AUPRC Inferelator-Prior CellOracle baseGRN 0.03 0.05 0.17 0.16 0.24 0.36 0.17 0.17 0.07 0.07 0.34 0.34

Table 4.6: AUPRC of GLM-Prior, Inferelator-Prior and CellOracle baseGRN across six cell lines.

Results for the Inferelator-Prior and CellOracle baseGRN prior knowledge for each six cell lines using the same reference sets described in Section 4.5.1 are provided in Table 4.6.

4.5.5

GRN Inference with PMF-GRN, the Inferelator, and CellOracle

GRN inference for each of the six cell lines was performed using PMF-GRN, the Inferelator 3.0, and CellOracle Python software, respectively. Single cell gene expression datasets were obtained from GSE125162 (38, 225 cells by 6, 763 genes) [Jackson et al. 2020] and GSE144820 (6, 118 cells by 6, 763 genes) [Jariani et al. 2020] without modification. The dataset was then combined by concatenating GSE125162 and GSE144820 on the cells axis (44, 343 cells by 6, 763 genes). The remaining hESC (758 cells by 17, 735 genes), HepG2 (425 cells by 11, 515 genes), mESC (421 cells by 18, 385 genes), mDC (383 cells by 7, 371 162

genes), and mHSC (2, 807 cells by 4, 762 genes) single-cell expression datasets were obtained from BEELINE [Pratapa et al. 2020].

Cell Line Yeast hESC HepG2 mESC mDC mHSC

PMF-GRN 0.02 0.19 0.38 0.27 0.10 0.52

AUPRC Inferelator 3.0 0.04 0.20 0.36 0.25 0.09 0.42

CellOracle 0.06 0.17 0.38 0.37 0.09 0.39

Table 4.7: AUPRC of three GRN inference methods (PMF-GRN, Inferelator 3.0, and CellOracle) across six cell lines.

Prior knowledge for each method is required to produce a GRN from single-cell expression data. For the experiments in this section, we use the specific cell line prior knowledge generated by the corresponding algorithm (described in Section 4.5.1 for GLM-Prior and Section 4.6 for Inferelator-Prior and CellOracle baseGRN. Results for these GRN inference experiments can be found in Table 4.7). Additional metrics for GLM-Prior and PMF-GRN, including F1 score, precision, recall, ROCAUC, positive class recall, positive class precision, negative class recall, and negative class precision, are provided in 4.13.

4.5.6

Cross-Method Comparison

In order to disentangle the performance contributions provided by each prior knowledge method (GLM-Prior, Inferelator-Prior, and CellOracle baseGRN) from their corresponding GRN inference method (PMF-GRN, Inferelator 3.0, and CellOracle), we provide a cross-method comparison in which each GRN inference algorithm is applied to each prior. No new datasets are introduced in this section. The results from the cross-method comparison can be found in Table 4.8.

163

Figure 4.13: Performance Metrics. Additional performance metrics for GLM-Prior and PMF-GRN across six cell lines.

164

Cell Line Yeast hESC HepG2 mESC mDC mHSC

CellOracle Prior PMF Inf CO 0.06 0.06 0.06 0.17 0.18 0.17 0.40 0.34 0.38 0.17 0.25 0.37 0.07 0.11 0.09 0.34 0.42 0.39

Inferelator-Prior PMF Inf CO 0.03 0.04 0.03 0.19 0.20 0.19 0.36 0.36 0.33 0.25 0.25 0.23 0.10 0.09 0.08 0.41 0.42 0.35

GLM-Prior PMF Inf CO 0.02 0.03 0.03 0.19 0.17 0.19 0.38 0.34 0.31 0.27 0.27 0.24 0.10 0.10 0.07 0.52 0.42 0.36

Table 4.8: Cross-method AUPRC comparison for six cell lines, combining three prior-knowledge constructions (CellOracle prior, Inferelator-Prior, GLM-Prior) with three GRN inference methods (PMF-GRN, Inferelator, CellOracle).

165

5 | Conclusion GRNs provide a mechanistic view of how TFs coordinate gene expression programs to sustain stable cell identities and mediate dynamic responses to signals and perturbations. Reconstructing these networks from genome-wide data is a central goal in modern genomics, yet remains challenging under realistic data and methodological constraints. In practice, several methodological limitations are especially consequential. GRN inference methods often tightly couple a chosen generative model to a specific inference procedure. As a result, methods developed for specific experimental modalities often require substantial redesign as new measurement technologies and modeling assumptions emerge. In addition, model selection is often treated heuristically, with many approaches committing to a single algorithmic form or fixed set of hyperparameters without systematic comparison to plausible alternatives, despite the fact that these choices can yield qualitatively different networks from the same dataset. Evaluation presents an additional obstacle that arises from three related limitations. First, GRN accuracy is typically assessed by comparing predicted edges to a reference network that is treated as a proxy for ground truth. Second, these reference networks are often incomplete, context dependent, or unavailable. Here, the absence of an edge in the reference frequently indicates missing evidence or limited coverage, rather than evidence that the interaction does not occur. Third, inferred networks are frequently reported only as point estimates, without an explicit measure of confidence to support interpretation when a predicted edge lacks a reference label. These limitations reduce the interpretability of agreement and disagreement with reference networks

166

and limit how reliably predictions can be prioritized and interpreted. Finally, GRN inference in practice depends critically on prior knowledge to anchor and constrain regulatory structure, yet commonly used priors often rely on cell type-specific experimental assays, emphasize promoter-proximal regulatory logic, and do not readily transfer to new species or less-characterized biological systems. Curated databases also remain biased toward non-model organisms, resulting in priors whose quality is highly variable across settings. This thesis aims to address these challenges by developing inference and prior construction frameworks that remain modular across data regimes, support principled model selection, provide uncertainty-aware regulatory estimates, and enable transferable prior knowledge to guide GRN reconstruction when direct regulatory evidence is sparse. In Chapter 3, we introduce PMF-GRN to address these methodological limitations by posing GRN inference as a probabilistic graphical model. In this framework, observed single cell gene expression data is decomposed into latent variables with explicitly defined distributions, providing a generative interpretation of regulatory structure. This probabilistic formulation decouples the generative model, which encodes assumptions about how TFs regulate genes, from the variational inference procedure used to fit the model, enabling straightforward adaptation to new datasets and modeling assumptions without redesigning the learning algorithm. Using the ELBO objective function, PMF-GRN supports principled model selection by comparing alternative generative models and hyperparameter configurations and selecting an optimal model under explicit performance criteria. Finally, posterior inference yields full distributions over TF-target gene interaction parameters, providing well-calibrated uncertainty estimates that serve as an interpretable measure of confidence in inferred edges, even when comprehensive ground-truth networks are incomplete or unavailable. Across single-cell datasets in yeast, human and mouse, we show that PMF-GRN recovers biologically meaningful regulatory structure, performs competitively with or better than stateof-the-art regression-based GRN inference methods, and scales efficiently across systems. At the 167

same time, these experiments highlight a practical limitation that remains even under a principled probabilistic framework. Accurate GRN recovery depends critically on the quality of the prior network used to specify candidate TF-target gene edges, with performance often constrained by the strength and correctness of the signal encoded in this prior. In Chapter 4, we introduce GLM-Prior to address this limitation by constructing prior knowledge directly from nucleotide sequence. GLM-Prior fine-tunes the 250 million parameter Nucleotide Transformer [Dalla-Torre et al. 2024], pretrained on the genomes of 850 species, to predict TF-target gene regulatory interactions as a supervised sequence-to-interaction problem. For each TF-gene pair, the model receives associated genic and TF-binding nucleotide sequences together with available interaction labels. These paired sequences are jointly encoded by the transformer and passed through a classification head to produce an interaction probability. Since positive regulatory interactions are sparse, GLM-prior is trained with class-weighted loss and balanced sampling to address extreme class imbalance. By leveraging representations learned across diverse genomes and the transformer’s attention mechanisms, GLM-Prior can transfer regulatory information to new cell types and species, capturing regulatory sequence dependencies that are difficult to recover from promoter-proximal, assay-dependent prior construction strategies. Across yeast, mouse, and human cell lines, we show that GLM-Prior produces sequencederived priors that outperform accessibility-based priors in several mammalian contexts. We also show that GLM-Prior generalizes across single-species, transfer learning, and multi-species training regimes, enabling prior construction in understudied systems. More broadly, we observe that when priors are informative, expression-based GRN inference primarily reweights and refines predicted TF-target gene edges rather than discovering them de novo. These results suggest that prior construction, rather than the choice of downstream GRN inference algorithm, is often the dominant bottleneck in GRN recovery. Taken together, Chapters 3 and 4 support a dual-stage perspective on GRN inference in which prior construction and expression-based inference play distinct, complementary roles. 168

Prior knowledge provides TF-specific structure that constrains the space of plausible regulatory explanations and anchor interpretability, while probabilistic inference refines this scaffold in a context-dependent manner and quantifies uncertainty in the resulting network. This division of labor directly targets the core constraints emphasized in this thesis, yielding GRN reconstruction pipelines that are more flexible across data regimes, less dependent on heuristic model selection, more interpretable under incomplete evaluation resources, and more transferable through priors learned by fine-tuning foundation models that generalize across species and cellular contexts. Several directions follow naturally from this work. On the prior construction side, an important next step is to move beyond DNA-only predictors toward foundation models that integrate multiple modalities along the central dogma, including DNA [Brixi et al. 2025], RNA, and protein [Hayes et al. 2025], with the goal of encoding regulatory specificity and context more directly. Another valuable avenue would be to develop a prior framework to jointly fine-tune multiple foundation models, including models that incorporate cell-type and tissue context from spatial transcriptomics [Tejada-Lapuerta et al. 2025] in order to construct priors that are explicitly conditioned on cellular contexts. These advances in prior modeling naturally motivate parallel progress on the inference side. A natural extension is to incorporate causal and temporal structure more explicitly, moving beyond static network reconstruction toward models that learn how regulatory interactions evolve over time. Such frameworks could support prediction of future regulatory states under perturbations [Kim et al. 2024], providing a more direct route to forecasting how drugs, signaling events, or disease progression reshapes regulatory programs. More broadly, this thesis motivates future GRN frameworks that couple context-aware, transferable priors with uncertainty aware and temporally grounded inference frameworks, yielding regulatory hypotheses that are both biologically grounded and practically testable at genome-scale.

169

Bibliography (2014). A comprehensive assessment of rna-seq accuracy, reproducibility and information content by the sequencing quality control consortium. Nature biotechnology, 32(9):903–914. Abbondanzieri, E. A., Greenleaf, W. J., Shaevitz, J. W., Landick, R., and Block, S. M. (2005). Direct observation of base-pair stepping by rna polymerase. Nature, 438(7067):460–465. Abdullah-Sayani, A., Bueno-de Mesquita, J. M., and Van De Vijver, M. J. (2006). Technology insight: tuning into the genetic orchestra using microarrays—limitations of dna microarrays in clinical practice. Nature clinical practice Oncology, 3(9):501–516. Abou El Hassan, M., Huang, K., Eswara, M. B., Xu, Z., Yu, T., Aubry, A., Ni, Z., Livne-Bar, I., Sangwan, M., Ahmad, M., et al. (2017). Properties of stat1 and irf1 enhancers and the influence of snps. BMC Molecular Biology, 18(1):1–19. Äijö, T. and Bonneau, R. (2016). Biophysically motivated regulatory network inference: progress and prospects. Human heredity, 81(2):62–77. Äijö, T. and Lähdesmäki, H. (2009). Learning gene regulatory networks from gene expression measurements using non-parametric molecular kinetics. Bioinformatics, 25(22):2937–2944. Akers, K. and Murali, T. (2021). Gene regulatory network inference in single-cell biology. Current Opinion in Systems Biology, 26:87–97.

170

Allaway, K. C., Gabitto, M. I., Wapinski, O., Saldi, G., Wang, C.-Y., Bandler, R. C., Wu, S. J., Bonneau, R., and Fishell, G. (2021). Genetic and epigenetic coordination of cortical interneuron development. Nature, 597(7878):693–697. Altay, G. and Emmert-Streib, F. (2010). Inferring the conservative causal core of gene regulatory networks. BMC systems biology, 4(1):132. Alter, O., Brown, P. O., and Botstein, D. (2000). Singular value decomposition for genome-wide expression data processing and modeling. Proceedings of the National Academy of Sciences, 97(18):10101–10106. Andoh, A., Shioya, M., Nishida, A., Bamba, S., Tsujikawa, T., Kim-Mitsuyama, S., and Fujiyama, Y. (2009). Expression of il-24, an activator of the jak1/stat3/socs3 cascade, is enhanced in inflammatory bowel disease. The Journal of Immunology, 183(1):687–695. Aoyagi, S. and Archer, T. K. (2008). Dynamics of coactivator recruitment and chromatin modifications during nuclear receptor mediated transcription. Molecular and cellular endocrinology, 280(1-2):1–5. Arrieta-Ortiz, M. L., Hafemeister, C., Bate, A. R., Chu, T., Greenfield, A., Shuster, B., Barry, S. N., Gallitto, M., Liu, B., Kacmarczyk, T., et al. (2015). An experimentally supported model of the bacillus subtilis global transcriptional regulatory network. Molecular systems biology, 11(11):839. Arroyo, N., Villamayor, L., Díaz, I., Carmona, R., Ramos-Rodríguez, M., Muñoz-Chápuli, R., Pasquali, L., Toscano, M. G., Martín, F., Cano, D. A., et al. (2021). Gata4 induces liver fibrosis regression by deactivating hepatic stellate cells. JCI insight, 6(23). Au-Yeung, N., Mandhana, R., and Horvath, C. M. (2013). Transcriptional regulation by stat1 and stat2 in the interferon jak-stat pathway. Jak-stat, 2(3):e23931. 171

Aubin-Frankowski, P.-C. and Vert, J.-P. (2020). Gene regulation inference from single-cell rna-seq data with linear differential equations and velocity inference. Bioinformatics, 36(18):4774–4780. Avsec, Ž., Agarwal, V., Visentin, D., Ledsam, J. R., Grabska-Barwinska, A., Taylor, K. R., Assael, Y., Jumper, J., Kohli, P., and Kelley, D. R. (2021). Effective gene expression prediction from sequence by integrating long-range interactions. Nature methods, 18(10):1196–1203. Babu, M. M., Luscombe, N. M., Aravind, L., Gerstein, M., and Teichmann, S. A. (2004). Structure and evolution of transcriptional regulatory networks. Current opinion in structural biology, 14(3):283–291. Badia-i Mompel, P., Wessels, L., Müller-Dott, S., Trimbour, R., Ramirez Flores, R. O., Argelaguet, R., and Saez-Rodriguez, J. (2023). Gene regulatory network inference in the era of single-cell multi-omics. Nature Reviews Genetics, 24(11):739–754. Banf, M. and Rhee, S. Y. (2017). Computational inference of gene regulatory networks: approaches, limitations and opportunities. Biochimica et Biophysica Acta (BBA)-Gene Regulatory Mechanisms, 1860(1):41–52. Barbosa, S., Niebel, B., Wolf, S., Mauch, K., and Takors, R. (2018). A guide to gene regulatory network inference for obtaining predictive solutions: Underlying assumptions and fundamental biological and data constraints. Biosystems, 174:37–48. Bastian, M., Heymann, S., and Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. In Proceedings of the international AAAI conference on web and social media, volume 3, pages 361–362. Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017). Variational inference: A review for statisticians. Journal of the American statistical Association, 112(518):859–877.

172

Blohm, D. H. and Guiseppi-Elie, A. (2001). New developments in microarray technology. Current Opinion in Biotechnology, 12(1):41–47. Blumenthal, S. G., Aichele, G., Wirth, T., Czernilofsky, A. P., Nordheim, A., and Dittmer, J. (1999). Regulation of the human interleukin-5 promoter by ets transcription factors: Ets1 and ets2, but not elf-1, cooperate with gata3 and htlv-i tax1. Journal of Biological Chemistry, 274(18):12910– 12916. Borukhov, S. and Nudler, E. (2008). Rna polymerase: the vehicle of transcription. Trends in microbiology, 16(3):126–134. Brent, M. R. (2016). Past roadblocks and new opportunities in transcription factor network mapping. Trends in Genetics, 32(11):736–750. Brixi, G., Durrant, M. G., Ku, J., Poli, M., Brockman, G., Chang, D., Gonzalez, G. A., King, S. H., Li, D. B., Merchant, A. T., et al. (2025). Genome modeling and design across all domains of life with evo 2. BioRxiv, pages 2025–02. Brunet, J.-P., Tamayo, P., Golub, T. R., and Mesirov, J. P. (2004). Metagenes and molecular pattern discovery using matrix factorization. Proceedings of the national academy of sciences, 101(12):4164–4169. Buenrostro, J. D., Giresi, P. G., Zaba, L. C., Chang, H. Y., and Greenleaf, W. J. (2013). Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, dna-binding proteins and nucleosome position. Nature methods, 10(12):1213–1218. Buenrostro, J. D., Wu, B., Chang, H. Y., and Greenleaf, W. J. (2015a). Atac-seq: a method for assaying chromatin accessibility genome-wide. Current protocols in molecular biology, 109(1):21–29. Buenrostro, J. D., Wu, B., Litzenburger, U. M., Ruff, D., Gonzales, M. L., Snyder, M. P., Chang, H. Y.,

173

and Greenleaf, W. J. (2015b). Single-cell chromatin accessibility reveals principles of regulatory variation. Nature, 523(7561):486–490. Burdziak, C., Azizi, E., Prabhakaran, S., and Pe’er, D. (2019). A nonparametric multi-view model for estimating cell type-specific gene regulatory networks. arXiv preprint arXiv:1902.08138. Butte, A. J. and Kohane, I. S. (1999). Mutual information relevance networks: functional genomic clustering using pairwise entropy measurements. In Biocomputing 2000, pages 418–429. World Scientific. Chaffey, N. (2003). Alberts, b., johnson, a., lewis, j., raff, m., roberts, k. and walter, p. molecular biology of the cell. 4th edn. Chai, L. E., Loh, S. K., Low, S. T., Mohamad, M. S., Deris, S., and Zakaria, Z. (2014). A review on the computational approaches for gene regulatory network construction. Computers in biology and medicine, 48:55–65. Chakrabarti, S., Kabra, M., Mandal, A. K., Senthil, S., and Kaur, I. (2016). The transcription factors pbx1 and gata1 are regulated by the mutation profiles of cyp1b1 in primary congenital glaucoma. Investigative Ophthalmology & Visual Science, 57(12):803–803. Chan, T. E., Stumpf, M. P., and Babtie, A. C. (2017). Gene regulatory network inference from single-cell data using multivariate information measures. Cell systems, 5(3):251–267. Chang, C., Ding, Z., Hung, Y. S., and Fung, P. C. W. (2008). Fast network component analysis (fastnca) for gene regulatory network reconstruction from microarray data. Bioinformatics, 24(11):1349–1358. Chaumeil, J. and Skok, J. A. (2012). The role of ctcf in regulating v (d) j recombination. Current opinion in immunology, 24(2):153–159.

174

Chen, H., Albergante, L., Hsu, J. Y., Lareau, C. A., Lo Bosco, G., Guan, J., Zhou, S., Gorban, A. N., Bauer, D. E., Aryee, M. J., et al. (2019a). Single-cell trajectories reconstruction, exploration and mapping of omics data with stream. Nature communications, 10(1):1903. Chen, L., Toke, N. H., Luo, S., Vasoya, R. P., Fullem, R. L., Parthasarathy, A., Perekatt, A. O., and Verzi, M. P. (2019b). A reinforcing hnf4–smad4 feed-forward module stabilizes enterocyte identity. Nature genetics, 51(5):777–785. Chen, L., Wang, S., Zhou, Y., Wu, X., Entin, I., Epstein, J., Yaccoby, S., Xiong, W., Barlogie, B., Shaughnessy Jr, J. D., et al. (2010). Identification of early growth response protein 1 (egr-1) as a novel target for jun-induced apoptosis in multiple myeloma. Blood, The Journal of the American Society of Hematology, 115(1):61–70. Chevalley, M., Roohani, Y. H., Mehrjou, A., Leskovec, J., and Schwab, P. (2025). A large-scale benchmark for network inference from single-cell perturbation data. Communications Biology, 8(1):412. Chu, S.-K., Zhao, S., Shyr, Y., and Liu, Q. (2022). Comprehensive evaluation of noise reduction methods for single-cell rna sequencing data. Briefings in bioinformatics, 23(2):bbab565. Ciofani, M., Madar, A., Galan, C., Sellars, M., Mace, K., Pauli, F., Agarwal, A., Huang, W., Parkurst, C. N., Muratet, M., et al. (2012). A validated regulatory network for th17 cell specification. Cell, 151(2):289–303. Cobaleda, C., Schebesta, A., Delogu, A., and Busslinger, M. (2007). Pax5: the guardian of b cell identity and function. Nature immunology, 8(5):463–470. Consens, M. E., Dufault, C., Wainberg, M., Forster, D., Karimzadeh, M., Goodarzi, H., Theis, F. J., Moses, A., and Wang, B. (2025). Transformers and genome language models. Nature Machine Intelligence, pages 1–17. 175

Cooper, G. and Adams, K. W. (2022). The cell: a molecular approach. Oxford University Press. Corces, M. R., Granja, J. M., Shams, S., Louie, B. H., Seoane, J. A., Zhou, W., Silva, T. C., Groeneveld, C., Wong, C. K., Cho, S. W., et al. (2018). The chromatin accessibility landscape of primary human cancers. Science, 362(6413):eaav1898. Costa, V., Aprile, M., Esposito, R., and Ciccodicola, A. (2013). Rna-seq and human complex diseases: recent accomplishments and future perspectives. European Journal of Human Genetics, 21(2):134–142. Crick, F. (1970). Central dogma of molecular biology. Nature, 227(5258):561–563. Cui, H., Wang, C., Maan, H., Pang, K., Luo, F., Duan, N., and Wang, B. (2024). scgpt: toward building a foundation model for single-cell multi-omics using generative ai. Nature Methods, 21(8):1470–1480. Cusanovich, D. A., Pavlovic, B., Pritchard, J. K., and Gilad, Y. (2014). The functional consequences of variation in transcription factor binding. PLoS genetics, 10(3):e1004226. Dalla-Torre, H., Gonzalez, L., Mendoza-Revilla, J., Lopez Carranza, N., Grzywaczewski, A. H., Oteri, F., Dallago, C., Trop, E., de Almeida, B. P., Sirelkhatim, H., et al. (2024). Nucleotide transformer: building and evaluating robust foundation models for human genomics. Nature Methods, pages 1–11. de Souza, N. (2012). The encode project. Nature methods, 9(11):1046–1046. Degner, J. F., Marioni, J. C., Pai, A. A., Pickrell, J. K., Nkadori, E., Gilad, Y., and Pritchard, J. K. (2009). Effect of read-mapping biases on detecting allele-specific expression from rnasequencing data. Bioinformatics, 25(24):3207–3212. Delgado, F. M. and Gómez-Vela, F. (2019). Computational methods for gene regulatory networks reconstruction and analysis: a review. Artificial intelligence in medicine, 95:133–145. 176

Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186. Dubois-Chevalier, J., Mazrooei, P., Lupien, M., Staels, B., Lefebvre, P., and Eeckhoute, J. (2018). Organizing combinatorial transcription factor recruitment at cis-regulatory modules. Transcription, 9(4):233–239. Dufva, M. (2009). Introduction to microarray technology. DNA Microarrays for Biomedical Research: Methods and Protocols, pages 1–22. Duren, Z., Chen, X., Zamanighomi, M., Zeng, W., Satpathy, A. T., Chang, H. Y., Wang, Y., and Wong, W. H. (2018). Integrative analysis of single-cell genomics data by coupled nonnegative matrix factorizations. Proceedings of the National Academy of Sciences, 115(30):7723–7728. Edsbäcker, E., Serviss, J. T., Kolosenko, I., Palm-Apergi, C., De Milito, A., and Tamm, K. P. (2019). Stat3 is activated in multicellular spheroids of colon carcinoma cells and mediates expression of irf9 and interferon stimulated genes. Scientific Reports, 9(1):536. Emad, A. and Sinha, S. (2021). Inference of phenotype-relevant transcriptional regulatory networks elucidates cancer type-specific regulatory mechanisms in a pan-cancer study. NPJ systems biology and applications, 7(1):9. Faith, J. J., Hayete, B., Thaden, J. T., Mogno, I., Wierzbowski, J., Cottarel, G., Kasif, S., Collins, J. J., and Gardner, T. S. (2007). Large-scale mapping and validation of escherichia coli transcriptional regulation from a compendium of expression profiles. PLoS biology, 5(1):e8. Faria, J. P., Overbeek, R., Taylor, R. C., Conrad, N., Vonstein, V., Goelzer, A., Fromion, V., Rocha, M., Rocha, I., and Henry, C. S. (2016). Reconstruction of the regulatory network for bacillus subtilis and reconciliation with gene expression data. Frontiers in Microbiology, 7:275. 177

Fiers, M. W., Minnoye, L., Aibar, S., Bravo González-Blas, C., Kalender Atak, Z., and Aerts, S. (2018). Mapping gene regulatory networks from single-cell omics data. Briefings in functional genomics, 17(4):246–254. Fitch, S. R., Kapeni, C., Tsitsopoulou, A., Wilson, N. K., Göttgens, B., de Bruijn, M. F., and Ottersbach, K. (2020). Gata3 targets runx1 in the embryonic haematopoietic stem cell niche. IUBMB life, 72(1):45–52. Forster, T., Roy, D., and Ghazal, P. (2003). Experiments using microarray technology: limitations and standard operating procedures. Journal of endocrinology, 178(2):195–204. Friedman, J., Hastie, T., and Tibshirani, R. (2008). Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3):432–441. Friedman, N., Linial, M., Nachman, I., and Pe’er, D. (2000). Using bayesian networks to analyze expression data. In Proceedings of the fourth annual international conference on Computational molecular biology, pages 127–135. Gao, J., Chen, Y.-H., and Peterson, L. C. (2015). Gata family transcriptional factors: emerging suspects in hematologic disorders. Experimental hematology & oncology, 4:1–7. Gao, Y., Chen, Q., and Yue, W. (2019). Laptm5 protein can regulate tgf-𝛽 mediated mapk and smad signaling pathways in ovarian cancer cell. Annals of Oncology, 30:v9. Gao, Y. and Church, G. (2005). Improving molecular cancer class discovery through sparse nonnegative matrix factorization. Bioinformatics, 21(21):3970–3975. Ghosh, S. and Chan, C.-K. K. (2016). Analysis of rna-seq data using tophat and cufflinks. In Plant Bioinformatics: Methods and Protocols, pages 339–361. Springer.

178

Gobin, S. J., Biesta, P., and Van den Elsen, P. J. (2003). Regulation of human 𝛽2-microglobulin transactivation in hematopoietic cells. Blood, The Journal of the American Society of Hematology, 101(8):3058–3064. Gorin, G., Fang, M., Chari, T., and Pachter, L. (2022). Rna velocity unraveled. PLoS computational biology, 18(9):e1010492. Grandi, F. C., Modi, H., Kampman, L., and Corces, M. R. (2022). Chromatin accessibility profiling by atac-seq. Nature protocols, 17(6):1518–1552. Grant, C. E., Bailey, T. L., and Noble, W. S. (2011). Fimo: scanning for occurrences of a given motif. Bioinformatics, 27(7):1017–1018. Greenfield, A., Hafemeister, C., and Bonneau, R. (2013). Robust data-driven incorporation of prior knowledge into the inference of dynamic regulatory networks. Bioinformatics, 29(8):1060–1067. Hafemeister, C. and Satija, R. (2019). Normalization and variance stabilization of single-cell rnaseq data using regularized negative binomial regression. Genome biology, 20(1):296. Han, H., Cho, J.-W., Lee, S., Yun, A., Kim, H., Bae, D., Yang, S., Kim, C. Y., Lee, M., Kim, E., et al. (2018). Trrust v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic acids research, 46(D1):D380–D386. Han, H., Shim, H., Shin, D., Shim, J. E., Ko, Y., Shin, J., Kim, H., Cho, A., Kim, E., Lee, T., et al. (2015). Trrust: a reference database of human transcriptional regulatory interactions. Scientific reports, 5(1):11432. Hao, Y., Hao, S., Andersen-Nissen, E., Mauck, W. M., Zheng, S., Butler, A., Lee, M. J., Wilk, A. J., Darby, C., Zager, M., et al. (2021). Integrated analysis of multimodal single-cell data. Cell, 184(13):3573–3587.

179

Haury, A.-C., Mordelet, F., Vera-Licona, P., and Vert, J.-P. (2012). Tigress: trustful inference of gene regulation using stability selection. BMC systems biology, 6(1):145. Hayes, T., Rao, R., Akin, H., Sofroniew, N. J., Oktay, D., Lin, Z., Verkuil, R., Tran, V. Q., Deaton, J., Wiggert, M., et al. (2025). Simulating 500 million years of evolution with a language model. Science, 387(6736):850–858. He, B. and Tan, K. (2016). Understanding transcriptional regulatory networks using computational models. Current opinion in genetics & development, 37:101–108. Hecker, M., Lambeck, S., Toepfer, S., Van Someren, E., and Guthke, R. (2009). Gene regulatory network inference: data integration in dynamic models—a review. Biosystems, 96(1):86–103. Hegde, A. and Cheng, J. (2025). Grnfomer: Accurate gene regulatory network inference using graph transformer. bioRxiv, pages 2025–01. Hegde, A., Nguyen, T., and Cheng, J. (2025). Machine learning methods for gene regulatory network inference. Briefings in Bioinformatics, 26(5):bbaf470. Hegenbarth, J.-C., Lezzoche, G., De Windt, L. J., and Stoll, M. (2022). Perspectives on bulk-tissue rna sequencing and single-cell rna sequencing for cardiac transcriptomics. Frontiers in Molecular Medicine, 2:839338. Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. (2017). beta-VAE: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations. Hill, C. S. (2016). Transcriptional control by the smads. Cold Spring Harbor perspectives in biology, 8(10):a022079.

180

Hintze, M., Prajapati, R. S., Tambalo, M., Christophorou, N. A., Anwar, M., Grocott, T., and Streit, A. (2017). Cell interactions, signals and transcriptional hierarchy governing placode progenitor induction. Development, 144(15):2810–2823. Hobert, O. (2008).

Gene regulation by transcription factors and micrornas.

Science,

319(5871):1785–1786. Howbrook, D. N., van der Valk, A. M., O’Shaughnessy, M. C., Sarker, D. K., Baker, S. C., and Lloyd, A. W. (2003). Developments in microarray technologies. Drug discovery today, 8(14):642–651. Hsu, L.-J., Hong, Q., Chen, S.-T., Kuo, H.-L., Schultz, L., Heath, J., Lin, S.-R., Lee, M.-H., Li, D.-Z., Li, Z.-L., et al. (2017). Hyaluronan activates hyal-2/wwox/smad4 signaling and causes bubbling cell death when the signaling complex is overexpressed. Oncotarget, 8(12):19137. Hu, X., Hu, Y., Wu, F., Leung, R. W. T., and Qin, J. (2020). Integration of single-cell multi-omics for gene regulatory network inference. Computational and structural biotechnology journal, 18:1925–1938. Hua, F., Mu, R., Liu, J., Xue, J., Wang, Z., Lin, H., Yang, H., Chen, X., and Hu, Z. (2011). Trb3 interacts with smad3 promoting tumor cell migration and invasion. Journal of cell science, 124(19):3235–3246. Husmeier, D. (2003). Sensitivity and specificity of inferring genetic regulatory interactions from microarray experiments with dynamic bayesian networks. Bioinformatics, 19(17):2271–2282. Huynh-Thu, V. A. and Geurts, P. (2018). dyngenie3: dynamical genie3 for the inference of gene networks from time series expression data. Scientific reports, 8(1):3384. Huynh-Thu, V. A., Irrthum, A., Wehenkel, L., and Geurts, P. (2010). Inferring regulatory networks from expression data using tree-based methods. PloS one, 5(9):e12776.

181

Inge, M. M., Miller, R., Hook, H., Bray, D., Keenan, J. L., Zhao, R., Gilmore, T. D., and Siggers, T. (2024). Rapid profiling of transcription factor–cofactor interaction networks reveals principles of epigenetic regulation. Nucleic Acids Research, 52(17):10276–10296. Inukai, S., Kock, K. H., and Bulyk, M. L. (2017). Transcription factor–dna binding: beyond binding site motifs. Current opinion in genetics & development, 43:110–119. Jackson, C. A., Castro, D. M., Saldi, G.-A., Bonneau, R., and Gresham, D. (2020). Gene regulatory network reconstruction using single-cell rna sequencing of barcoded genotypes in diverse environments. elife, 9:e51254. Jaksik, R., Iwanaszko, M., Rzeszowska-Wolny, J., and Kimmel, M. (2015). Microarray experiments and factors which affect their reliability. Biology direct, 10(1):46. Jaluria, P., Konstantopoulos, K., Betenbaugh, M., and Shiloach, J. (2007). A perspective on microarrays: current applications, pitfalls, and potential uses. Microbial cell factories, 6(1):4. Jansen, C., Ramirez, R. N., El-Ali, N. C., Gomez-Cabrero, D., Tegner, J., Merkenschlager, M., Conesa, A., and Mortazavi, A. (2019). Building gene regulatory networks from scatac-seq and scrna-seq using linked self organizing maps. PLoS computational biology, 15(11):e1006555. Jariani, A., Vermeersch, L., Cerulus, B., Perez-Samper, G., Voordeckers, K., Van Brussel, T., Thienpont, B., Lambrechts, D., and Verstrepen, K. J. (2020). A new protocol for single-cell rna-seq reveals stochastic gene expression during lag phase in budding yeast. elife, 9:e55320. Jeon, H., Lee, H., Kang, B., Jang, I., and Roh, T.-Y. (2020). Comparative analysis of commonly used peak calling programs for chip-seq analysis. Genomics & informatics, 18(4):e42. Ji, Y., Zhou, Z., Liu, H., and Davuluri, R. V. (2021). Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome. Bioinformatics, 37(15):2112–2120. 182

Ji, Z., He, L., Regev, A., and Struhl, K. (2019). Inflammatory regulatory network mediated by the joint action of nf-kb, stat3, and ap-1 factors is involved in many human cancers. Proceedings of the National Academy of Sciences, 116(19):9453–9462. Ji, Z., Zhou, W., Hou, W., and Ji, H. (2020). Single-cell atac-seq signal extraction and enhancement with scate. Genome biology, 21(1):161. Jiang, X., Tan, J., Wen, Y., Liu, W., Wu, S., Wang, L., Wangou, S., Liu, D., Du, C., Zhu, B., et al. (2019). Msi2-tgf-𝛽/tgf-𝛽 r1/smad3 positive feedback regulation in glioblastoma. Cancer chemotherapy and pharmacology, 84:415–425. Jiang, X. and Zhang, X. (2022). Rsnet: inferring gene regulatory networks by a redundancy silencing and network enhancement technique. BMC bioinformatics, 23(1):165. Johnson, K. D., Boyer, M. E., Kang, J.-A., Wickrema, A., Cantor, A. B., and Bresnick, E. H. (2007). Friend of gata-1–independent transcriptional repression: a novel mode of gata-1 function. Blood, The Journal of the American Society of Hematology, 109(12):5230–5233. Johnstone, A. L., Andrade, N. S., Barbier, E., Khomtchouk, B. B., Rienas, C. A., Lowe, K., Van Booven, D. J., Domi, E., Esanov, R., Vilca, S., et al. (2021). Dysregulation of the histone demethylase kdm6b in alcohol dependence is associated with epigenetic regulation of inflammatory signaling pathways. Addiction biology, 26(1):e12816. Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K. (1999). An introduction to variational methods for graphical models. Machine learning, 37(2):183–233. Jung, G.-S., Hwang, Y. J., Choi, J.-H., and Lee, K.-M. (2020). Lin28a attenuates tgf-𝛽-induced renal fibrosis. BMB reports, 53(11):594. Kalfon, J., Samaran, J., Peyré, G., and Cantini, L. (2025). scprint: pre-training on 50 million cells allows robust gene network predictions. Nature Communications, 16(1):3607. 183

Kamimoto, K., Stringa, B., Hoffmann, C. M., Jindal, K., Solnica-Krezel, L., and Morris, S. A. (2023). Dissecting cell identity via network inference and in silico gene perturbation. Nature, 614(7949):742–751. Kano, S.-i., Sato, K., Morishita, Y., Vollstedt, S., Kim, S., Bishop, K., Honda, K., Kubo, M., and Taniguchi, T. (2008). The contribution of transcription factor irf1 to the interferon-𝛾– interleukin 12 signaling axis and th1 versus th-17 differentiation of cd4+ t cells. Nature immunology, 9(1):34–41. Karamveer and Uzun, Y. (2024). Approaches for benchmarking single-cell gene regulatory network methods. Bioinformatics and Biology Insights, 18:11779322241287120. Karlebach, G. and Shamir, R. (2008). Modelling and analysis of gene regulatory networks. Nature reviews Molecular cell biology, 9(10):770–780. Katsumura, K. R., Bresnick, E. H., and Group, G. F. M. (2017). The gata factor revolution in hematology. Blood, The Journal of the American Society of Hematology, 129(15):2092–2102. Kernfeld, E., Keener, R., Cahan, P., and Battle, A. (2024). Transcriptome data are insufficient to control false discoveries in regulatory network inference. Cell systems, 15(8):709–724. Keuthan, C., Santiago, C., and Ash, J. D. (2019). Stat3 is a potential genetic modifier of photoreceptor gene expression during stress. Investigative Ophthalmology & Visual Science, 60(9):466–466. Khalid, A. B., Pence, J., Suthon, S., Lin, J., Miranda-Carboni, G. A., and Krum, S. A. (2021). Gata4 regulates mesenchymal stem cells via direct transcriptional regulation of the wnt signalosome. Bone, 144:115819. Kim, D., Tran, A., Kim, H. J., Lin, Y., Yang, J. Y. H., and Yang, P. (2023). Gene regulatory network reconstruction: harnessing the power of single-cell multi-omic data. NPJ Systems Biology and Applications, 9(1):51. 184

Kim, J.-H., Gibbs, C. S., Yun, S., Song, H. O., and Cho, K. (2024). Large-scale targeted cause discovery via learning from simulated data. arXiv preprint arXiv:2408.16218. Kim, J.-H., Gibbs, C. S., Yun, S., Song, H. O., and Cho, K. (2025). Large-scale targeted cause discovery via learning from simulated data. Transactions on Machine Learning Research. Kim, J.-H., Hedrick, S., Tsai, W.-W., Wiater, E., Le Lay, J., Kaestner, K. H., Leblanc, M., Loar, A., and Montminy, M. (2017). Creb coactivators crtc2 and crtc3 modulate bone marrow hematopoiesis. Proceedings of the National Academy of Sciences, 114(44):11739–11744. Kim, P. M. and Tidor, B. (2003). Subsystem identification through dimensionality reduction of large-scale gene expression data. Genome research, 13(7):1706–1718. Klug, A. (2004). The discovery of the dna double helix. Journal of molecular biology, 335(1):3–26. Kobayashi, M., Funayama, R., Ohnuma, S., Unno, M., and Nakayama, K. (2016). Wnt-𝛽-catenin signaling regulates abcc 3 (mrp 3) transporter expression in colorectal cancer. Cancer science, 107(12):1776–1784. Kolodziejczyk, A. A., Kim, J. K., Svensson, V., Marioni, J. C., and Teichmann, S. A. (2015). The technology and biology of single-cell rna sequencing. Molecular cell, 58(4):610–620. Kong, X., Wang, Q., Li, J., Li, M., Deng, F., and Li, C. (2022). Mammaglobin, gata-binding protein 3 (gata3), and epithelial growth factor receptor (egfr) expression in different breast cancer subtypes and their clinical significance. European Journal of Histochemistry: EJH, 66(2). Kulakovskiy, I. V., Medvedeva, Y. A., Schaefer, U., Kasianov, A. S., Vorontsov, I. E., Bajic, V. B., and Makeev, V. J. (2013). Hocomoco: a comprehensive collection of human transcription factor binding sites models. Nucleic acids research, 41(D1):D195–D202. Kumari, S., Bonnet, M. C., Ulvmar, M. H., Wolk, K., Karagianni, N., Witte, E., Uthoff-Hachenberg, C., Renauld, J.-C., Kollias, G., Toftgard, R., et al. (2013). Tumor necrosis factor receptor signaling 185

in keratinocytes triggers interleukin-24-dependent psoriasis-like skin inflammation in mice. Immunity, 39(5):899–911. Lähnemann, D., Köster, J., Szczurek, E., McCarthy, D. J., Hicks, S. C., Robinson, M. D., Vallejos, C. A., Campbell, K. R., Beerenwinkel, N., Mahfouz, A., et al. (2020). Eleven grand challenges in single-cell data science. Genome biology, 21(1):1–35. Lambert, S. A., Jolma, A., Campitelli, L. F., Das, P. K., Yin, Y., Albu, M., Chen, X., Taipale, J., Hughes, T. R., and Weirauch, M. T. (2018). The human transcription factors. Cell, 172(4):650–665. Latchman, D. S. (1997). Transcription factors: an overview. The international journal of biochemistry & cell biology, 29(12):1305–1312. Lee, H. K., Hsu, A. K., Sajdak, J., Qin, J., and Pavlidis, P. (2004). Coexpression analysis of human genes across many microarray data sets. Genome research, 14(6):1085–1094. Lessard, S., Gatof, E. S., Beaudoin, M., Schupp, P. G., Sher, F., Ali, A., Prehar, S., Kurita, R., Nakamura, Y., Baena, E., et al. (2017). An erythroid-specific atp2b4 enhancer mediates red blood cell hydration and malaria susceptibility. The Journal of clinical investigation, 127(8):3065–3074. Leung, D. and Drton, M. (2016). Order-invariant prior specification in bayesian factor analysis. Statistics & Probability Letters, 111:60–66. Li, G., Yang, Y., Van Buren, E., and Li, Y. (2019). Dropout imputation and batch effect correction for single-cell rna sequencing data. Journal of Bio-X Research, 2(04):169–177. Li, K., Wu, Y., Li, Y., Yu, Q., Tian, Z., Wei, H., and Qu, K. (2020a). Landscape and dynamics of the transcriptional regulatory network during natural killer cell differentiation. Genomics, Proteomics & Bioinformatics, 18(5):501–515. Li, L., Zhang, R., Liu, Y., and Zhang, G. (2020b). Anxa4 activates jak-stat3 signaling by interacting with anxa1 in basal-like breast cancer. DNA and Cell Biology, 39(9):1649–1656. 186

Li, S., Zhao, Y., Varma, R., Salpekar, O., Noordhuis, P., Li, T., Paszke, A., Smith, J., Vaughan, B., Damania, P., et al. (2020c). Pytorch distributed: Experiences on accelerating data parallel training. arXiv preprint arXiv:2006.15704. Liao, M.-H., Lin, P.-I., Ho, W.-P., Chan, W. P., Chen, T.-L., and Chen, R.-M. (2017). Participation of gata-3 in regulation of bone healing through transcriptional upregulation of bcl-xl expression. Experimental & Molecular Medicine, 49(11):e398–e398. Linder, J., Srivastava, D., Yuan, H., Agarwal, V., and Kelley, D. R. (2025). Predicting rna-seq coverage from dna sequence as a unifying model of gene regulation. Nature Genetics, pages 1–13. Liu, R., Liu, L., Bian, Y., Zhang, S., Wang, Y., Chen, H., Jiang, X., Li, G., Chen, Q., Xue, C., et al. (2022). The dual regulation effects of esr1/nedd4l on slc7a11 in breast cancer under ionizing radiation. Frontiers in Cell and Developmental Biology, 9:772380. Liu, W., Geng, C., Li, X., Li, Y., Song, S., and Wang, C. (2023a). Downregulation of slc9a8 promotes epithelial–mesenchymal transition and metastasis in colorectal cancer cells via the il6jak1/stat3 signaling pathway. Digestive Diseases and Sciences, 68(5):1873–1884. Liu, X., Bai, F., Wang, Y., Wang, C., Chan, H. L., Zheng, C., Fang, J., Zhu, W.-G., and Pei, X.H. (2023b). Loss of function of gata3 regulates fra1 and c-fos to activate emt and promote mammary tumorigenesis and metastasis. Cell Death & Disease, 14(6):370. Liu, Y., Harmelink, C., Peng, Y., Chen, Y., Wang, Q., and Jiao, K. (2014). Chd7 interacts with bmp rsmads to epigenetically regulate cardiogenesis in mice. Human molecular genetics, 23(8):2145– 2156. Loers, J. U. and Vermeirssen, V. (2024). A single-cell multimodal view on gene regulatory network inference from transcriptomics and chromatin accessibility data. Briefings in Bioinformatics, 25(5):bbae382. 187

Luecken, M. D. and Theis, F. J. (2019). Current best practices in single-cell rna-seq analysis: a tutorial. Molecular systems biology, 15(6):e8746. Lukhele, S., Abd Rabbo, D., Guo, M., Shen, J., Elsaesser, H. J., Quevedo, R., Carew, M., Gadalla, R., Snell, L. M., Mahesh, L., et al. (2022). The transcription factor irf2 drives interferon-mediated cd8+ t cell exhaustion to restrict anti-tumor immunity. Immunity, 55(12):2369–2385. Madden, T. (2013). The blast sequence analysis tool. The NCBI handbook, 2(5):425–436. Madhamshettiwar, P. B., Maetschke, S. R., Davis, M. J., Reverter, A., and Ragan, M. A. (2012). Gene regulatory network inference: evaluation and application to ovarian cancer allows the prioritization of drug targets. Genome medicine, 4(5):41. Majumder, P. and Boss, J. M. (2011). Dna methylation dysregulates and silences the hla-dq locus by altering chromatin architecture. Genes & Immunity, 12(4):291–299. Malhotra, N. and Kang, J. (2013). Smad regulatory networks construct a balanced immune system. Immunology, 139(1):1–10. Marbach, D., Costello, J. C., Küffner, R., Vega, N. M., Prill, R. J., Camacho, D. M., Allison, K. R., Kellis, M., Collins, J. J., et al. (2012). Wisdom of crowds for robust gene network inference. Nature methods, 9(8):796–804. Mardis, E. R. (2007). Chip-seq: welcome to the new frontier. Nature methods, 4(8):613–614. Margolin, A. A., Nemenman, I., Basso, K., Wiggins, C., Stolovitzky, G., Favera, R. D., and Califano, A. (2006). Aracne: an algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context. BMC bioinformatics, 7(Suppl 1):S7. Marguerat, S. and Bähler, J. (2010). Rna-seq: from technology to biology. Cellular and molecular life sciences, 67(4):569–579.

188

Marguerat, S., Wilhelm, B. T., and Bähler, J. (2008). Next-generation sequencing: applications beyond genomes. Marr, C., Zhou, J. X., and Huang, S. (2016). Single-cell gene expression profiling and cell state dynamics: collecting data, correlating data points and connecting the dots. Current opinion in biotechnology, 39:207–214. Matsumoto, H., Kiryu, H., Furusawa, C., Ko, M. S., Ko, S. B., Gouda, N., Hayashi, T., and Nikaido, I. (2017). Scode: an efficient regulatory network inference algorithm from single-cell rna-seq during differentiation. Bioinformatics, 33(15):2314–2321. Mayumi, A., Tomii, T., Kanayama, T., Mikami, T., Tanaka, K., Yoshida, H., Kato, I., Kawamura, M., Nakahata, T., Takita, J., et al. (2021). Activation of the stat1-bcl-2/mcl-1 axis in leukemic cells carrying a spag9-jak2 fusion. Blood, 138:4326. McCalla, S. G., Fotuhi Siahpirani, A., Li, J., Pyne, S., Stone, M., Periyasamy, V., Shin, J., and Roy, S. (2023). Identifying strengths and weaknesses of methods for computational network inference from single-cell rna-seq data. G3: Genes, Genomes, Genetics, 13(3):jkad004. Mercado, N., Schutzius, G., Kolter, C., Estoppey, D., Bergling, S., Roma, G., Gubser Keller, C., Nigsch, F., Salathe, A., Terranova, R., et al. (2019). Irf2 is a master regulator of human keratinocyte stem cell fate. Nature communications, 10(1):4676. Mercatelli, D., Scalambra, L., Triboli, L., Ray, F., and Giorgi, F. M. (2020). Gene regulatory network inference resources: A practical overview. Biochimica et Biophysica Acta (BBA)-Gene Regulatory Mechanisms, 1863(6):194430. Mering, C. v., Huynen, M., Jaeggi, D., Schmidt, S., Bork, P., and Snel, B. (2003). String: a database of predicted functional associations between proteins. Nucleic acids research, 31(1):258–261.

189

Meyer, P. E., Kontos, K., Lafitte, F., and Bontempi, G. (2007). Information-theoretic inference of large transcriptional regulatory networks. EURASIP journal on bioinformatics and systems biology, 2007(1):79879. Michna, R. H., Zhu, B., Mäder, U., and Stülke, J. (2016). Subti wiki 2.0—an integrated database for the model organism bacillus subtilis. Nucleic acids research, 44(D1):D654–D662. Miraldi, E. R., Pokrovskii, M., Watters, A., Castro, D. M., De Veaux, N., Hall, J. A., Lee, J.-Y., Ciofani, M., Madar, A., Carriero, N., et al. (2019). Leveraging chromatin accessibility for transcriptional regulatory network inference in t helper 17 cells. Genome research, 29(3):449–463. Mitsis, T., Efthimiadou, A., Bacopoulou, F., Vlachakis, D., Chrousos, G. P., and Eliopoulos, E. (2020). Transcription factors and evolution: An integral part of gene expression. World Academy of Sciences Journal, 2(1):3–8. Mnih, A. and Salakhutdinov, R. R. (2007). Probabilistic matrix factorization. Advances in neural information processing systems, 20. Moerman, T., Aibar Santos, S., Bravo González-Blas, C., Simm, J., Moreau, Y., Aerts, J., and Aerts, S. (2019). Grnboost2 and arboreto: efficient and scalable inference of gene regulatory networks. Bioinformatics, 35(12):2159–2161. Moloshok, T. D., Klevecz, R., Grant, J. D., Manion, F. J., Speier IV, W., and Ochs, M. F. (2002). Application of bayesian decomposition for analysing microarray data. Bioinformatics, 18(4):566–575. Monteiro, P. T., Oliveira, J., Pais, P., Antunes, M., Palma, M., Cavalheiro, M., Galocha, M., Godinho, C. P., Martins, L. C., Bourbon, N., et al. (2020). Yeastract+: a portal for cross-species comparative genomics of transcription regulation in yeasts. Nucleic acids research, 48(D1):D642–D649. Morabito, S., Reese, F., Rahimzadeh, N., Miyoshi, E., and Swarup, V. (2023). hdwgcna identifies co-expression networks in high-dimensional transcriptomics data. Cell reports methods, 3(6). 190

Mordelet, F. and Vert, J.-P. (2008). Sirene: supervised inference of regulatory networks. Bioinformatics, 24(16):i76–i82. Murphy, D. (2002).

Gene expression studies using microarrays: principles, problems, and

prospects. Advances in physiology education, 26(4):256–270. Nachman, I., Regev, A., and Friedman, N. (2004). Inferring quantitative models of regulatory networks from expression data. Bioinformatics, 20(suppl_1):i248–i256. Nadon, R. and Shoemaker, J. (2002). Statistical issues with microarrays: processing and analysis. TRENDS in Genetics, 18(5):265–271. Nagel, S., Pommerenke, C., Meyer, C., Kaufmann, M., Drexler, H. G., and MacLeod, R. A. (2016). Deregulation of polycomb repressor complex 1 modifier auts2 in t-cell leukemia. Oncotarget, 7(29):45398. Nakato, R. and Shirahige, K. (2017). Recent advances in chip-seq analysis: from quality management to whole-genome annotation. Briefings in bioinformatics, 18(2):279–290. Nasser, J., Bergman, D. T., Fulco, C. P., Guckelberger, P., Doughty, B. R., Patwardhan, T. A., Jones, T. R., Nguyen, T. H., Ulirsch, J. C., Lekschas, F., et al. (2021). Genome-wide enhancer maps link risk variants to disease genes. Nature, 593(7858):238–243. Nesari, A. M., MotieGhader, H., and Ghorbian, S. (2026). Advances and challenges in singlecell rna sequencing data analysis: a comprehensive review.

Briefings in Bioinformatics,

27(1):bbaf723. Nguyen, E., Poli, M., Faizi, M., Thomas, A., Wornow, M., Birch-Sykes, C., Massaroli, S., Patel, A., Rabideau, C., Bengio, Y., et al. (2023). Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. Advances in neural information processing systems, 36:43177– 43201. 191

Nguyen-Jackson, H., Panopoulos, A. D., Zhang, H., Li, H. S., and Watowich, S. S. (2010). Stat3 controls the neutrophil migratory response to cxcr2 ligands by direct activation of g-csf–induced cxcr2 expression and via modulation of cxcr2 signal transduction. Blood, The Journal of the American Society of Hematology, 115(16):3354–3363. Nicolas, P., Mäder, U., Dervyn, E., Rochat, T., Leduc, A., Pigeonneau, N., Bidnenko, E., Marchadier, E., Hoebeke, M., Aymerich, S., et al. (2012). Condition-dependent transcriptome reveals highlevel regulatory architecture in bacillus subtilis. Science, 335(6072):1103–1106. Nie, X.-H., Qiu, S., Xing, Y., Xu, J., Lu, B., Zhao, S.-F., Li, Y.-T., Su, Z.-Z., et al. (2022). Paeoniflorin regulates nedd4l/stat3 pathway to induce ferroptosis in human glioma cells. Journal of Oncology, 2022. Ochs, M. F. and Fertig, E. J. (2012). Matrix factorization for transcriptional regulatory network inference. In 2012 IEEE Symposium on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), pages 387–396. IEEE. Özel, M. N., Gibbs, C. S., Holguera, I., Soliman, M., Bonneau, R., and Desplan, C. (2022). Coordinated control of neuronal differentiation and wiring by sustained transcription factors. Science, 378(6626):eadd1884. Papastamoulis, P. and Ntzoufras, I. (2022). On the identifiability of bayesian factor analytic models. Statistics and Computing, 32(2):23. Papili Gao, N., Ud-Dean, S. M., Gandrillon, O., and Gunawan, R. (2018). Sincerities: inferring gene regulatory networks from time-stamped single cell transcriptional expression profiles. Bioinformatics, 34(2):258–266. Park, P. J. (2009). Chip–seq: advantages and challenges of a maturing technology. Nature reviews genetics, 10(10):669–680. 192

Pedreira, T., Elfmann, C., and Stülke, J. (2022). The current state of subti wiki, the database for the model organism bacillus subtilis. Nucleic Acids Research, 50(D1):D875–D882. Persyn, E., Wahlen, S., Kiekens, L., Van Loocke, W., Siwe, H., Van Ammel, E., De Vos, Z., Van Nieuwerburgh, F., Matthys, P., Taghon, T., et al. (2022). Irf2 is required for development and functional maturation of human nk cells. Frontiers in immunology, 13:1038821. Pietz, G., De, R., Hedberg, M., Sjöberg, V., Sandström, O., Hernell, O., Hammarström, S., and Hammarström, M.-L. (2017). Immunopathology of childhood celiac disease—key role of intestinal epithelial cells. PLoS One, 12(9):e0185025. Pratapa, A., Jalihal, A. P., Law, J. N., Bharadwaj, A., and Murali, T. (2020). Benchmarking algorithms for gene regulatory network inference from single-cell transcriptomic data. Nature methods, 17(2):147–154. Qiu, P. (2020). Embracing the dropouts in single-cell rna-seq analysis. Nature communications, 11(1):1169. Qiu, X., Rahimzamani, A., Wang, L., Ren, B., Mao, Q., Durham, T., McFaline-Figueroa, J. L., Saunders, L., Trapnell, C., and Kannan, S. (2020). Inferring causal gene regulatory networks from coupled single-cell expression dynamics using scribe. Cell systems, 10(3):265–274. Rafi, A. M., Kiyota, B., Yachie, N., and de Boer, C. (2025). Detecting and avoiding homology-based data leakage in genome-trained sequence models. bioRxiv, pages 2025–01. Rai, M. F., Tycksen, E. D., Sandell, L. J., and Brophy, R. H. (2018). Advantages of rna-seq compared to rna microarrays for transcriptome profiling of anterior cruciate ligament tears. Journal of Orthopaedic Research®, 36(1):484–497. Ranganath, R., Gerrish, S., and Blei, D. (2014). Black box variational inference. In Artificial intelligence and statistics, pages 814–822. PMLR. 193

Rauluseviciute, I., Riudavets-Puig, R., Blanc-Mathieu, R., Castro-Mondragon, J. A., Ferenc, K., Kumar, V., Lemma, R. B., Lucas, J., Chèneby, J., Baranasic, D., et al. (2024). Jaspar 2024: 20th anniversary of the open-access database of transcription factor binding profiles. Nucleic acids research, 52(D1):D174–D182. Ravasi, T., Suzuki, H., Cannistraci, C. V., Katayama, S., Bajic, V. B., Tan, K., Akalin, A., Schmeier, S., Kanamori-Katayama, M., Bertin, N., et al. (2010). An atlas of combinatorial transcriptional regulation in mouse and man. Cell, 140(5):744–752. Ritchie, M. E., Phipson, B., Wu, D., Hu, Y., Law, C. W., Shi, W., and Smyth, G. K. (2015). limma powers differential expression analyses for rna-sequencing and microarray studies. Nucleic acids research, 43(7):e47–e47. Rochette, L., Dogon, G., Zeller, M., Cottin, Y., and Vergely, C. (2021). Gdf15 and cardiac cells: current concepts and new insights. International Journal of Molecular Sciences, 22(16):8889. Roy, R., Dagher, A., Butterfield, C., and Moses, M. A. (2017). Adam12 is a novel regulator of tumor angiogenesis via stat3 signaling. Molecular Cancer Research, 15(11):1608–1622. Russo, G., Zegar, C., and Giordano, A. (2003). Advantages and limitations of microarray technology in human cancer. Oncogene, 22(42):6497–6507. Saliba, A.-E., Westermann, A. J., Gorski, S. A., and Vogel, J. (2014). Single-cell rna-seq: advances and future challenges. Nucleic acids research, 42(14):8845–8860. San Roman, A. K., Aronson, B. E., Krasinski, S. D., Shivdasani, R. A., and Verzi, M. P. (2015). Transcription factors gata4 and hnf4a control distinct aspects of intestinal homeostasis in conjunction with transcription factor cdx2. Journal of Biological Chemistry, 290(3):1850–1860. Sandelin, A., Alkema, W., Engström, P., Wasserman, W. W., and Lenhard, B. (2004). Jaspar: an

194

open-access database for eukaryotic transcription factor binding profiles. Nucleic acids research, 32(suppl_1):D91–D94. Sanvictores, T. and Farci, F. (2020). Biochemistry, primary protein structure. Schäfer, J., Opgen-Rhein, R., and Strimmer, K. (2006). Reverse engineering genetic networks using the genenet package. The Newsletter of the R Project Volume 6/5, December 2006, 6(9):50. Schäfer, J. and Strimmer, K. (2005). An empirical bayes approach to inferring large-scale gene association networks. Bioinformatics, 21(6):754–764. Schlitt, T. and Brazma, A. (2007). Current approaches to gene regulatory network modelling. BMC bioinformatics, 8(Suppl 6):S9. Sha, Y., Qiu, Y., Zhou, P., and Nie, Q. (2024). Reconstructing growth and dynamic trajectories from single-cell transcriptomics data. Nature Machine Intelligence, 6(1):25–39. Shibata, M., Ooki, A., Inokawa, Y., Sadhukhan, P., Ugurlu, M. T., Izumchenko, E., Munari, E., Bogina, G., Rudin, C. M., Gabrielson, E., et al. (2020). Concurrent targeting of potential cancer stem cells regulating pathways sensitizes lung adenocarcinoma to standard chemotherapy. Molecular cancer therapeutics, 19(10):2175–2185. Shu, H., Zhou, J., Lian, Q., Li, H., Zhao, D., Zeng, J., and Ma, J. (2021). Modeling gene regulatory networks using neural network architectures. Nature Computational Science, 1(7):491–501. Skok Gibbs, C., Chen, A., Bonneau, R., and Cho, K. (2025). Glm-prior: a genomic language model for transferable sequence-derived priors in gene regulatory network inference. bioRxiv, pages 2025–06. Skok Gibbs, C., Jackson, C. A., Saldi, G.-A., Tjärnberg, A., Shah, A., Watters, A., De Veaux, N., Tchourine, K., Yi, R., Hamamsy, T., et al. (2022). High-performance single-cell gene regulatory network inference at scale: the inferelator 3.0. Bioinformatics, 38(9):2519–2528. 195

Skok Gibbs, C., Mahmood, O., Bonneau, R., and Cho, K. (2024). Pmf-grn: a variational inference approach to single-cell gene regulatory network inference using probabilistic matrix factorization. Genome biology, 25(1):88. Slattery, M., Zhou, T., Yang, L., Machado, A. C. D., Gordân, R., and Rohs, R. (2014). Absence of a simple code: how transcription factors read the genome. Trends in biochemical sciences, 39(9):381–399. Slovin, S., Carissimo, A., Panariello, F., Grimaldi, A., Bouché, V., Gambardella, G., and Cacchiarelli, D. (2021). Single-cell rna sequencing analysis: a step-by-step overview. RNA bioinformatics, pages 343–365. Smyth, G. K. (2005). Limma: linear models for microarray data. In Bioinformatics and computational biology solutions using R and Bioconductor, pages 397–420. Springer. Specht, A. T. and Li, J. (2017). Leap: constructing gene co-expression networks for single-cell rna-sequencing data using pseudotime ordering. Bioinformatics, 33(5):764–766. Spitz, F. and Furlong, E. E. (2012). Transcription factors: from enhancer binding to developmental control. Nature reviews genetics, 13(9):613–626. Stein-O’Brien, G. L., Arora, R., Culhane, A. C., Favorov, A. V., Garmire, L. X., Greene, C. S., Goff, L. A., Li, Y., Ngom, A., Ochs, M. F., et al. (2018). Enter the matrix: factorization uncovers knowledge from omics. Trends in Genetics, 34(10):790–805. Stock, M., Losert, C., Zambon, M., Popp, N., Lubatti, G., Hörmanseder, E., Heinig, M., and Scialdone, A. (2025). Leveraging prior knowledge to infer gene regulatory networks from single-cell rna-sequencing data. Molecular Systems Biology, pages 1–17. Stormo, G. D. (2000). Dna binding sites: representation and discovery. Bioinformatics, 16(1):16–23.

196

Sun, Y., Miao, N., and Sun, T. (2019). Detect accessible chromatin using atac-sequencing, from principle to applications. Hereditas, 156(1):29. Szklarczyk, D., Franceschini, A., Kuhn, M., Simonovic, M., Roth, A., Minguez, P., Doerks, T., Stark, M., Muller, J., Bork, P., et al. (2010). The string database in 2011: functional interaction networks of proteins, globally integrated and scored. Nucleic acids research, 39(suppl_1):D561–D568. Szklarczyk, D., Gable, A. L., Nastou, K. C., Lyon, D., Kirsch, R., Pyysalo, S., Doncheva, N. T., Legeay, M., Fang, T., Bork, P., et al. (2021). The string database in 2021: customizable protein– protein networks, and functional characterization of user-uploaded gene/measurement sets. Nucleic acids research, 49(D1):D605–D612. Szklarczyk, D., Kirsch, R., Koutrouli, M., Nastou, K., Mehryary, F., Hachilif, R., Gable, A. L., Fang, T., Doncheva, N. T., Pyysalo, S., et al. (2023). The string database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic acids research, 51(D1):D638–D646. Tahara, S. and Ozaki, H. (2025). Unmeasured human transcription factor chip-seq data shape functional genomics and demand strategic prioritization. Briefings in Functional Genomics, 24:elaf016. Tarca, A. L., Romero, R., and Draghici, S. (2006). Analysis of microarray experiments of gene expression profiling. American journal of obstetrics and gynecology, 195(2):373–388. Tchourine, K., Vogel, C., and Bonneau, R. (2018). Condition-specific modeling of biophysical parameters advances inference of regulatory networks. Cell reports, 23(2):376–388. Teixeira, M. C., Monteiro, P. T., Palma, M., Costa, C., Godinho, C. P., Pais, P., Cavalheiro, M., Antunes, M., Lemos, A., Pedreira, T., et al. (2018). Yeastract: an upgraded database for the analysis of transcription regulatory networks in saccharomyces cerevisiae. Nucleic acids research, 46(D1):D348–D353. 197

Tejada-Lapuerta, A., Schaar, A. C., Gutgesell, R., Palla, G., Halle, L., Minaeva, M., Vornholz, L., Dony, L., Drummer, F., Richter, T., et al. (2025). Nicheformer: a foundation model for singlecell and spatial omics. Nature methods, pages 1–14. Thomas, R., Thomas, S., Holloway, A. K., and Pollard, K. S. (2017). Features that define the best chip-seq peak calling algorithms. Briefings in bioinformatics, 18(3):441–450. Tjärnberg, A., Beheler-Amass, M., Jackson, C. A., Christiaen, L., Gresham, D. J., and Bonneau, R. (2023). Structure primed embedding on the transcription factor manifold enables transparent model architectures for gene regulatory network and latent activity inference. bioRxiv, pages 2023–02. Treiber, T., Mandel, E. M., Pott, S., Györy, I., Firner, S., Liu, E. T., and Grosschedl, R. (2010). Early b cell factor 1 regulates b cell gene networks by activation, repression, and transcriptionindependent poising of chromatin. Immunity, 32(5):714–725. Trelford, C. B. and Di Guglielmo, G. M. (2021). Canonical and non-canonical tgf𝛽 signaling activate autophagy in an ulk1-dependent manner. Frontiers in Cell and Developmental Biology, 9:712124. Tung, P.-Y., Blischak, J. D., Hsiao, C. J., Knowles, D. A., Burnett, J. E., Pritchard, J. K., and Gilad, Y. (2017). Batch effects and the effective design of single-cell gene expression studies. Scientific reports, 7(1):39921. Unger Avila, P., Padvitski, T., Leote, A. C., Chen, H., Saez-Rodriguez, J., Kann, M., and Beyer, A. (2024). Gene regulatory networks in disease and ageing. Nature Reviews Nephrology, 20(9):616– 633. Vallejos, C. A., Risso, D., Scialdone, A., Dudoit, S., and Marioni, J. C. (2017). Normalizing singlecell rna sequencing data: challenges and opportunities. Nature methods, 14(6):565–571. 198

Van de Sande, B., Flerin, C., Davie, K., De Waegeneer, M., Hulselmans, G., Aibar, S., Seurinck, R., Saelens, W., Cannoodt, R., Rouchon, Q., et al. (2020). A scalable scenic workflow for single-cell gene regulatory network analysis. Nature protocols, 15(7):2247–2276. Van Dijk, D., Sharma, R., Nainys, J., Yim, K., Kathail, P., Carr, A. J., Burdziak, C., Moon, K. R., Chaffer, C. L., Pattabiraman, D., et al. (2018). Recovering gene interactions from single-cell data using data diffusion. Cell, 174(3):716–729. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30. Vorontsov, I. E., Eliseeva, I. A., Zinkevich, A., Nikonov, M., Abramov, S., Boytsov, A., Kamenets, V., Kasianova, A., Kolmykov, S., Yevshin, I. S., et al. (2024). Hocomoco in 2024: a rebuild of the curated collection of binding models for human and mouse transcription factors. Nucleic Acids Research, 52(D1):D154–D163. Wang, X., Liao, P., Fan, X., Wan, Y., Wang, Y., Li, Y., Jiang, Z., Ye, X., Mo, X., Ocorr, K., et al. (2013). Cxxc5 associates with smads to mediate tnf-𝛼 induced apoptosis. Current Molecular Medicine, 13(8):1385–1396. Wang, Y., Jiang, L., Mo, X., Lan, Y., Yang, X., Liu, X., Zhang, J., Zhu, L., Liu, J., and Wu, X. (2017). Megakaryocytic smad4 regulates platelet function through syk and rock2 expression. Molecular pharmacology, 92(3):285–296. Wang, Y., Joshi, T., Zhang, X.-S., Xu, D., and Chen, L. (2006). Inferring gene regulatory networks from multiple microarray datasets. Bioinformatics, 22(19):2413–2420. Wang, Z., Gerstein, M., and Snyder, M. (2009). Rna-seq: a revolutionary tool for transcriptomics. Nature reviews genetics, 10(1):57–63. 199

Wei, T. and Lambert, P. F. (2021). Role of iqgap1 in carcinogenesis. Cancers, 13(16):3940. Wei, X., Yu, L., and Li, Y. (2018). Pbx1 promotes the cell proliferation via jak2/stat3 signaling in clear cell renal carcinoma. Biochemical and biophysical research communications, 500(3):650– 657. Weirauch, M. T., Yang, A., Albu, M., Cote, A. G., Montenegro-Montero, A., Drewe, P., Najafabadi, H. S., Lambert, S. A., Mann, I., Cook, K., et al. (2014). Determination and inference of eukaryotic transcription factor sequence specificity. Cell, 158(6):1431–1443. Werhli, A. V. and Husmeier, D. (2007). Reconstructing gene regulatory networks with bayesian networks by combining expression data with multiple sources of prior knowledge. Statistical Applications in Genetics & Molecular Biology, 6(1). Wingender, E., Dietze, P., Karas, H., and Knüppel, R. (1996). Transfac: a database on transcription factors and their dna binding sites. Nucleic acids research, 24(1):238–241. Wolf, F. A., Angerer, P., and Theis, F. J. (2018). Scanpy: large-scale single-cell gene expression data analysis. Genome biology, 19:1–5. Wong, K. M., Song, J., and Wong, Y. H. (2021). Ctcf and egr1 suppress breast cancer cell migration through transcriptional control of nm23-h1. Scientific Reports, 11(1):491. Wu, S., Fu, J., Dong, Y., Yi, Q., Lu, D., Wang, W., Qi, Y., Yu, R., and Zhou, X. (2018). Golph3 promotes glioma progression via facilitating jak2–stat3 pathway activation. Journal of neuro-oncology, 139:269–279. Wu, W., Xu, N., Zhou, X., Liu, L., Tan, Y., Luo, J., Huang, J., Qin, J., Wang, J., Li, Z., et al. (2020). Integrative genomic analysis reveals cancer-associated gene mutations in chronic myeloid leukemia patients with resistance or intolerance to tyrosine kinase inhibitor. OncoTargets and therapy, pages 8581–8591. 200

Xu, H., Baroukh, C., Dannenfelser, R., Chen, E. Y., Tan, C. M., Kou, Y., Kim, Y. E., Lemischka, I. R., and Ma’ayan, A. (2013). Escape: database for integrating high-content published data collected from human and mouse embryonic stem cells. Database, 2013:bat045. Xu, J., Cui, L., Zhuang, J., Meng, Y., Bing, P., He, B., Tian, G., Pui, C. K., Wu, T., Wang, B., et al. (2022). Evaluating the performance of dropout imputation and clustering methods for singlecell rna sequencing data. Computers in Biology and Medicine, 146:105697. Xu, J., Zhang, A., Liu, F., and Zhang, X. (2023). Stgrns: an interpretable transformer-based method for inferring gene regulatory networks from single-cell transcriptomic data. Bioinformatics, 39(4):btad165. Yan, F., Powell, D. R., Curtis, D. J., and Wong, N. C. (2020). From reads to insight: a hitchhiker’s guide to atac-seq data analysis. Genome biology, 21(1):22. Yang, X., Wang, C., Lin, Y., and Zhang, P. (2022). Identification of crucial hub genes and differential t cell infiltration in idiopathic pulmonary arterial hypertension using bioinformatics strategies. Frontiers in Molecular Biosciences, 9:800888. Yang, Z. and Michailidis, G. (2016). A non-negative matrix factorization method for detecting modules in heterogeneous omics multi-modal data. Bioinformatics, 32(1):1–8. Yosef, N., Shalek, A. K., Gaublomme, J. T., Jin, H., Lee, Y., Awasthi, A., Wu, C., Karwacz, K., Xiao, S., Jorgolli, M., et al. (2013). Dynamic regulatory network controlling th17 cell differentiation. Nature, 496(7446):461–468. Yu, J.-H., Moon, E.-Y., Kim, J., and Koo, J. H. (2023). Identification of small gtpases that phosphorylate irf3 through tbk1 activation using an active mutant library screen. Biomolecules & Therapeutics, 31(1):48.

201

Yu, T.-Y., Chen, X.-X., Liu, Q.-W., Ma, F.-F., Huang, H.-L., Zhou, L., and Zhang, W. (2021). Loss of gata4 c-terminus by p. s335x mutation modulates coronary artery vascular smooth muscle cell phenotype. Mediators of Inflammation, 2021. Yuan, Q. and Duren, Z. (2025). Inferring gene regulatory networks from single-cell multiome data using atlas-scale external data. Nature Biotechnology, 43(2):247–257. Zhang, B. and Horvath, S. (2005). A general framework for weighted gene co-expression network analysis. Statistical applications in genetics and molecular biology, 4(1). Zhang, Z., Parker, M. P., Graw, S., Novikova, L. V., Fedosyuk, H., Fontes, J. D., Koestler, D. C., Peterson, K. R., and Slawson, C. (2019). O-glcnac homeostasis contributes to cell fate decisions during hematopoiesis. Journal of Biological Chemistry, 294(4):1363–1379. Zhao, M., Zhang, Y., Qiang, L., Lu, Z., Zhao, Z., Fu, Y., Wu, B., Chai, Q., Ge, P., Lei, Z., et al. (2023). A golgi-resident gpr108 cooperates with e3 ubiquitin ligase smurf1 to suppress antiviral innate immunity. Cell Reports, 42(6). Zhao, S., Fung-Leung, W.-P., Bittner, A., Ngo, K., and Liu, X. (2014). Comparison of rna-seq and microarray in transcriptome profiling of activated t cells. PloS one, 9(1):e78644. Zhong, B., Zhang, L., Lei, C., Li, Y., Mao, A.-P., Yang, Y., Wang, Y.-Y., Zhang, X.-L., and Shu, H.-B. (2009). The ubiquitin ligase rnf5 regulates antiviral responses by mediating degradation of the adaptor protein mita. Immunity, 30(3):397–407. Zhou, X., Pan, J., Chen, L., Zhang, S., and Chen, Y. (2024). Deepimager: deeply analyzing gene regulatory networks from scrna-seq data. Biomolecules, 14(7):766. Zhou, Z., Ji, Y., Li, W., Dutta, P., Davuluri, R., and Liu, H. (2023). Dnabert-2: Efficient foundation model and benchmark for multi-species genome. arXiv preprint arXiv:2306.15006.

202

Zhou, Z., Wei, J., Liu, M., Zhuo, L., Fu, X., and Zou, Q. (2025). Anomalgrn: deciphering single-cell gene regulation network with graph anomaly detection. BMC biology, 23(1):73. Zhu, B. and Stülke, J. (2018). Subti wiki in 2018: from genes and proteins to functional network annotation of the model organism bacillus subtilis. Nucleic acids research, 46(D1):D743–D748. Zhu, F., Panwar, B., and Guan, Y. (2016). Algorithms for modeling global and context-specific functional relationship networks. Briefings in bioinformatics, 17(4):686–695. Zou, Z., Ohta, T., and Oki, S. (2024). Chip-atlas 3.0: a data-mining suite to explore chromosome architecture together with large-scale regulome data. Nucleic Acids Research, 52(W1):W45– W53.

203

Record · ID 381738 · SHA-256 f5f2ca414faa4ff2
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.