ConceptioArchivearXiv CS
arXiv CSopen access

Canopy: A Heterograph Foundation Model for Metabolic Engineering

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Canopy: A Heterograph Foundation Model for Metabolic Engineering

Jake Bowden 1 Laurence Legon 1 Satnam Surae 1

arXiv:2607.06224v1 [cs.LG] 7 Jul 2026

Abstract

ticals, and specialty chemicals. The design–build–test– learn (DBTL) cycle that underpins strain engineering is slow: a single round of genetic modification, fermentation, and analytical characterisation can take weeks to months (Opgenorth et al., 2019), and the combinatorial space of candidate modifications grows exponentially in the number of target genes. Computational tools that prioritise genetic interventions before wet-lab experiments would shorten this loop.

Designing microbial strains that produce highvalue chemicals at commercially viable titers remains a central challenge in metabolic engineering. Existing computational approaches either rely on stoichiometric constraint-based models that cannot learn from experimental data, or apply tabular machine learning to hand-crafted features that discard the relational structure of biological knowledge. We present C ANOPY, a heterogeneous graph foundation model that integrates ten public and proprietary data sources into a unified knowledge graph (KG) of 6.9 M nodes across 13 types and 34 edge types, covering genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments. Node features are encoded through domain-specific foundation models (ESM-2 for protein sequences, MoLFormer for chemical SMILES, and PubMedBERT for biomedical text), yielding a multi-modal representation within a single graph. We pretrain a Heterogeneous Graph Transformer (HGT) augmented with SignNet positional encodings, Jumping Knowledge aggregation, and virtual nodes using four self-supervised objectives (link prediction, masked node modelling, distance prediction, and contrastive experiment clustering), balanced via learned homoscedastic uncertainty weighting. On the downstream task of fermentation titer prediction, frozen C ANOPY embeddings achieve R2 = 0.41 with a lightweight probe, outperforming tabular baselines (best R2 = 0.24) and homogeneous GNN variants.

The dominant computational paradigm for strain design is constraint-based modelling via genome-scale metabolic models (GEMs). Flux balance analysis (FBA) and its extensions predict steady-state fluxes under stoichiometric and thermodynamic constraints, enabling in silico knockout analysis through tools such as OptKnock (Burgard et al., 2003) and StrainDesign (Schneider et al., 2022). While GEMs encode mechanistic knowledge of metabolism, they cannot incorporate experimental measurements of titer, rate, and yield; they ignore regulatory, expression-level, and environmental effects; and they treat each organism in isolation without using cross-organism transfer. Machine learning offers a complementary approach. Previous work has applied random forests, gradient boosting, and neural networks to predict fermentation titer from features extracted from strain descriptions (Oyetunde et al., 2019; Czajka et al., 2021). These tabular approaches discard the relational structure that links genes to proteins, proteins to reactions, reactions to pathways, and pathways to production phenotypes. Graph neural networks have been applied to metabolic networks for gene essentiality (Hasibi et al., 2024) and site-of-metabolism prediction (Porokhin et al., 2023), but these operate on single-organism reaction graphs rather than on cross-organism knowledge graphs. Meanwhile, the broader ML community has advanced heterogeneous graph transformers (Hu et al., 2020b), selfsupervised graph pretraining (Hu et al., 2020a), and domainspecific foundation models for proteins (Lin et al., 2023) and small molecules (Ross et al., 2021). Biomedical knowledge graphs such as PrimeKG (Chandak et al., 2023) have enabled GNN-based drug repurposing, but no analogous resource or foundation model exists for the metabolic engineering domain, which sits at the intersection of microbial genetics, enzyme biochemistry, and fermentation science.

1. Introduction The bioeconomy depends on engineered microorganisms that convert renewable feedstocks into fuels, pharmaceu1

Twig Bio, London, United Kingdom. Correspondence to: Jake Bowden <[email protected]>, Satnam Surae <[email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).

1

Canopy: A Heterograph Foundation Model for Metabolic Engineering

We introduce C ANOPY, a heterogeneous graph foundation model for metabolic engineering. Our contributions are:

yields more accurately than purely stoichiometric models, but their parameterisation is organism-specific and does not transfer across strains. Tabular ML on hand-crafted features remains dominant for titer prediction (Oyetunde et al., 2019; Czajka et al., 2021; Radivojević et al., 2020); graph methods so far have been single-organism (Hasibi et al., 2024; Xin et al., 2024) or relied on shallow KG embeddings (Song et al., 2026; Gema et al., 2024). Diffusion- and flowmatching-based generative models are seeing rapid uptake for protein backbones and enzyme active sites (Li et al., 2026; Ahern et al., 2026), but target single molecules rather than whole-strain phenotypes. C ANOPY predicts titer from a learned cross-organism representation that fuses sequence, chemistry, and KG context. The same frozen embeddings can condition a downstream Bayesian-optimisation loop (Cheng et al., 2023) or generative model.

• A metabolic-engineering knowledge graph integrating ten data sources (MetaNetX, GO, UniRef, InterPro, NCBI Taxonomy and Genomics, KEGG, an in-house laboratory information management system (LIMS) of unpublished DBTL records, literature-mined experiments, and transcriptomics) via BioCypher (Lobentanzer et al., 2023) into a unified heterogeneous graph with 13 node types and 34 edge types. • Multi-modal feature encoding using a schema-driven dispatch system that routes protein sequences to ESM2 (650M), SMILES strings to MoLFormer-XL, free text to PubMedBERT, and numeric features to normalised scalars within a single graph. • An augmented Heterogeneous Graph Transformer combining HGTConv with SignNet positional encodings, random walk structural encodings, per-type feedforward networks (FFNs) with learnable residual scaling, Jumping Knowledge aggregation, and virtual nodes.

Benchmarks for ML in metabolic engineering. A useful benchmark for ML-driven metabolic engineering needs three things: experimental titer measurements, crossorganism coverage, and multi-omic context for each strain. No existing resource provides all three. SimDBTL (van Lent et al., 2023) offers consistent DBTL splits but its data are simulated rather than experimental. Therapeutics Data Commons (Huang et al., 2021) standardises prediction splits for therapeutic rather than fermentation tasks. Text-mining catalogues of over 15,000 strain-design publications (MárquezZavala et al., 2025) surface raw records without a held-out evaluation. C ANOPY’s 4,791 fermentation experiments with a deterministic MD5-hashed 5x CV split address all three criteria.

• Four-objective self-supervised pretraining with learned uncertainty weighting: link prediction, masked node modelling, distance prediction, and a contrastive experiment-pair loss. • Downstream evaluation on fermentation titer prediction, showing that learned graph representations outperform tabular baselines and provide the conditioning substrate for downstream generative strain-design pipelines.

2. Related Work Foundation models across biological scales. Domainspecific foundation models now span genomes (Nguyen et al., 2024; Brixi et al., 2025), protein sequence (Lin et al., 2023; Hayes et al., 2025) and structure (Su et al., 2024; Abramson et al., 2024), molecules (Ross et al., 2021), and single-cell transcriptomics (Cui et al., 2024; Hao et al., 2024). Each captures a single modality at a single scale; none are jointly trained over the relational structure that links genes to enzymes to metabolites to fermentation outcomes. C ANOPY embeds frozen ESM-2, MoLFormer, and PubMedBERT features inside a heterogeneous KG, treating these models as feature encoders within a graph that crosses scales.

Biological knowledge graphs and heterogeneous GNNs. Large biomedical KGs (PrimeKG, Chandak et al., 2023; Hetionet, Himmelstein et al., 2017; SPOKE, Nelson et al., 2019) and graph foundation models built on them (Huang et al., 2024; Hu et al., 2026) target human disease, not metabolic engineering. We use BioCypher (Lobentanzer et al., 2023) to integrate ten metabolic-engineering adapters and build on the Heterogeneous Graph Transformer (Hu et al., 2020b) with SignNet (Lim et al., 2022), random-walk PE (Dwivedi et al., 2022), Jumping Knowledge (Xu et al., 2018), and virtual nodes (Gilmer et al., 2017); pretraining objectives draw on Hu et al. (2020a) and the graph-FM surveys of Wang et al. (2025); Mao et al. (2024). The closest cross-organism transfer is IKT4Meta (Xin et al., 2024); C ANOPY extends this from two-organism alignment to a 13-node-type, multi-modal KG with fermentation-scale supervision.

Predictive and generative models for strain, pathway, and enzyme design. Kinetic GEMs such as k-ecoli457 (Khodayari & Maranas, 2016) extend FBA with parameterised rate equations fit to fluxomic data and predict product 2

Canopy: A Heterograph Foundation Model for Metabolic Engineering child of

Trscr. study

Taxa

has result

Trscr. result child of

child of

used in

Chassis

contains

base for Strain

used in

measures

Gene

encodes

modification

Pathway

contains

has

UniRef Cluster

described by

catalyses

has domain

Rxn

Gene Ontol.

InterPro Domain

participates in

produces

Exp.

Metab.

Genomics

Proteomics

Metabolomics

Experiment

Ontology

Figure 1. C ANOPY heterograph schema. Each node type has a distinct shape and is coloured by data domain (genomics, proteomics, metabolomics, experiment, ontology); each edge is typed by relation. Strain-to-gene modification edges include knockouts, knockins, and knockdowns.

3. Method

Property schema. Each node property carries a mandatory prefix that declares its role: sys for system keys used in deduplication (not exported as features), feat for numeric values, text for free text routed to the text encoder, seq for amino acid sequences routed to the protein encoder, and smi for SMILES routed to the chemical encoder. This dispatch ensures graph construction, embedding, and training share a single source of truth.

3.1. Knowledge Graph Construction C ANOPY’s knowledge graph is constructed using BioCypher (Lobentanzer et al., 2023). We implement ten adapter modules that ingest data from complementary sources and yield nodes and edges in a unified schema. Node types. The graph contains 13 node types spanning five biological scales: (i) molecular metabolites (InChIKey, SMILES, formula) and reactions; (ii) protein UniRef90 clusters together with their InterPro domains; (iii) genomic genes, chassis organisms, and engineered strains; (iv) functional annotations from the Gene Ontology, metabolic pathways, and NCBI taxonomy; and (v) experimental fermentation runs alongside transcriptomic experiments with pergene expression measurements.

Data sources. Molecular data is drawn from MetaNetX (Moretti et al., 2021) and KEGG (Kanehisa & Goto, 2000). Protein data comes from UniRef90 (Suzek et al., 2015) with sequences from UniProt (The UniProt Consortium, 2023) and domain annotations from InterPro (Paysan-Lafosse et al., 2023). Genomic data includes NCBI gene records (Brown et al., 2015) and taxonomy (Schoch et al., 2020). Functional annotations come from the Gene Ontology (Ashburner et al., 2000). Experimental data is sourced from a quality-controlled literature corpus and a proprietary experimental database; transcriptomic data provides per-gene expression linked to experimental conditions.

Edge types. Thirty-four edge types capture relationships across these scales: catalytic associations linking proteins to the reactions they catalyse, pathway membership, genetic modifications (knockouts, knockins, and overexpression edits applied to a strain), functional annotations, ontological structure, experiment tracking, and gene-expression links between transcriptomic measurements and the genes they quantify.

3.2. Multi-Modal Feature Encoding C ANOPY employs three pretrained foundation models as frozen extractors. Protein sequences (seq ) are encoded by ESM-2 (Lin et al., 2023) (esm2 t33 650M UR50D) and 3

Canopy: A Heterograph Foundation Model for Metabolic Engineering Table 1. Knowledge graph statistics by node type. The graph contains 11.2 M relationships and 17.7 M node properties in total. Domain

Node type

Genomics

Taxon Chassis Genomic Gene

143 26 168,913

Proteomics

UniRef Cluster InterPro Domain

150,914 51,489

Metabolomics

Metabolite Reaction

1,495,667 83,796

Experiment

Strain Pathway Experiment Transcriptomic

6,910 1,872 4,791 4,860,266

Ontology

GO Term (BP/MF/CC) Total

xv

Eig. PE

RWPE

Input projection. For each node type t, a learnable linear projection Wt maps xv to a hidden representation of dimension d = dh · H, with dh the per-head dimension and H the number of heads.

Count

Positional encodings. We augment node features with two complementary signals. The k smallest non-trivial Laplacian eigenvectors are computed via randomised SVD and processed through SignNet (Lim et al., 2022): SignNet(e) = MLP(e) + MLP(−e).

Random walk structural encodings of length ℓ are computed per subgraph sample and projected per-type. Both encodings are added element-wise to the projected features.

38,739

Transformer blocks. The model stacks L blocks: pertype pre-norm; HGTConv with type-dependent multi-head attention; GELU; a residual connection scaled by a learnable S CALE L AYER initialised at 0.1; a per-type FFN with expansion r and dropout; and a second scaled residual. The small initial scale stabilises training of deep networks by letting the residual stream dominate early. On top of the stack, we apply Jumping Knowledge (Xu et al., 2018) with max-pooling across all L layers per type. A virtual node type is added with bidirectional edges to all other nodes, reducing effective graph diameter and providing a shortcut path between otherwise disconnected components. The final representations are layer-normalised.

6,863,526

node embτ (t)

10-d

16-d

sign netτ (t)

(1)

+

(0)

hv

rw embτ (t)

Figure 2. Input feature construction. Each node type τ (v) has its own node emb, sign net (consuming a 10-d Laplacian eigenvector positional encoding), and rw emb (consuming a 16-d randomwalk positional encoding); the three streams are summed to form (0) the per-node residual stream hv .

3.4. Multi-Task Self-Supervised Pretraining We pretrain C ANOPY with four complementary objectives. Hold-out integrity. The same hash-keyed Experimentnode split used at probe time (Section 4.1) gates every pretraining objective that touches an Experiment node: edges incident to any val/test Experiment are excluded from the link-prediction label set, masked-node-modelling ignores held-out Experiment rows when sampling its mask, and the contrastive Experiment-clustering loss filters heldout Experiments out of its pair pool before sampling. Heldout Experiment nodes remain in the message-passing graph so probe-time message passing matches the pretrain-time graph the model was conditioned on; only supervision signal involving their identity, features, or neighbour structure is removed.

mean-pooled across residues. Chemical structures (smi ) are encoded by MoLFormer-XL (Ross et al., 2021). Biomedical text (text ) is encoded by S-PubMedBERT, applied to GO definitions, gene names and summaries, pathway names, and experiment descriptions. Numeric features (feat ) are z-score normalised using per-type statistics; categorical integers are one-hot encoded. All features for a node are concatenated into xv ; per-type input projection handles dimensional heterogeneity. 3.3. Model Architecture

Link prediction. Thirty percent of edges are deterministically reserved as supervision labels via an MD5 hash of edge endpoints, ensuring consistent splits across sampling runs (subject to the hold-out filter above). A dot-product predictor scores edges, ŷuv = σ(z⊤ u zv ), trained with binary cross-entropy. For each edge type we draw negatives at a 1:1 ratio with positives by uniformly sampling type-constrained

C ANOPY’s encoder is a Heterogeneous Graph Transformer that processes typed nodes and edges natively, extending HGTConv (Hu et al., 2020b) with several modern components. Figures 2–4 give a schematic overview of the full architecture, including the per-type and per-relation parameterisation that distinguishes HGT from a homogeneous GNN. 4

Canopy: A Heterograph Foundation Model for Metabolic Engineering Attention residual

K-Linτ (s)

√ Ks WϕATT Q⊤ t / d · µ⟨τ (s),ϕ,τ (t)⟩

softmaxs∈N (t)

Q-Linτ (t)

t

(att)

× α[τ ]

V -Linτ (s)

MSG

X

Vs Wϕ

αs→t ms→t

A-Linτ (t)

LayerNorm Linear GELU Dropout Linear

+

s

t′

+ (ffn)

× α[τ ]

(ℓ+1)

ht

FFN residual

Figure 3. One HGT layer. The dst-side query Qτ (t) is type-specific; keys Kτ (s) and values Vτ (s) are projected per source type, while attention logits and messages are scaled by per-relation matrices WϕATT , WϕMSG and a learnable relation prior µ⟨τ (s),ϕ,τ (t)⟩ . Attention is softmaxed over the destination’s neighbourhood and aggregated; a per-type output projection A-Linτ (t) feeds an FFN sub-block. Both attention and FFN use learnable scaled residual connections (blue).

t′

{h(ℓ) v }

Jumping Knowledge

LayerNorm

loss is

hout v

L=

X Li i

Figure 4. Readout. After L stacked HGT layers, JumpingKnowl(ℓ) edge max-pools the per-layer representations {hv } (combined (in) with a learnable input-residual scale α[τ ] ) and a final per-type LayerNorm produces the encoder output hout v .

2σi2

+ log σi .

(2)

log σi is clamped at −4.0 to prevent collapse; tasks with noisier supervision automatically receive lower weight. 3.5. Scalable Subgraph Sampling Training on the full graph is infeasible; C ANOPY uses a Neo4j-backed streaming sampler that constructs mini-batch subgraphs in parallel worker processes and writes each to disk as a PyTorch Geometric (Fey & Lenssen, 2019) HeteroData object. The sampler supports three strategies (simple BFS, batched random walks, and multi-anchor); we use multi-anchor by default. Seeds are split 50/50 between yield-bearing Experiment nodes (so every probe target appears in some subgraph) and inverse-degree-weighted random nodes (so peripheral nodes are not crowded out by hubs).

(src, dst) pairs from the corresponding node pools within the sampled subgraph and rejecting pairs that are already edges. Masked node modelling. We mask 15% of node features per batch. Per-type linear decoders reconstruct masked features under MSE, which acts as a regulariser against representation collapse. Distance prediction. For each subgraph, 200 random node pairs are sampled and their shortest-path distances computed via BFS on the undirected graph before virtualnode edges are added (capped at 5 hops); virtual nodes would otherwise collapse the diameter to two and trivialise the target. A distance head (the Hadamard product of source and target embeddings followed by a two-layer MLP) regresses these distances under MSE, supervising pairwise embedding distances against the graph metric.

For each Experiment seed batch the multi-anchor strategy resolves four BFS anchors: the seed Experiment; its S TRAIN (via USES STRAIN); parent TAXON plus capped sibling strains (via BELONGS TO, capped at 10 per species); and adjacent M ETABOLIC PATHWAYs (via HAS PATHWAY from the Strain). It also resolves four directed metapath anchors that walk fixed causal chains the undirected BFS under-covers: Experiment→M ETABOLITE (target compound), Pathway→Reaction→U NIREF C LUSTER (pathway enzymes), Strain→G ENOMIC G ENE→U NIREF C LUSTER (genetic edits to encoded proteins), and the reverse-direction chain Metabolite→Reaction→U NIREF C LUSTER→G ENOMIC G ENE (target enzymes). Each anchor is allocated a fraction of a global node budget Nmax (default 1,000); unused budget from inactive anchors is proportionally redistributed. Ontological edges (SUBCLASS OF, PART OF, REGULATES) and transcriptomic measurement edges are excluded from

Contrastive experiment-pair loss. Experiment node embeddings are contrasted within each batch with a target similarity inversely proportional to graph distance, sij = 1 − dij /dmax , trained with temperature-scaled cosine similarity (τ =0.07). Loss combination. The four losses are combined via homoscedastic uncertainty weighting (Kendall et al., 2018). Each task has a learnable log-variance log σi , and the total 5

Canopy: A Heterograph Foundation Model for Metabolic Engineering Link prediction σ(MLP(zu , zv )) → BCE

Masked node model mnm head[τ ] (zv ) → MSE on xv L=

{zv }v∈V encoder output

X

Li + log σi 2σi2

i

Kendall & Gal, 2018 (uncertainty weighting)

Graph distance MLP(zi ⊙ zj ) → MSE (shortest-path)

Contrastive Exp↔Exp cos(zi , zj ) → MSE on 1 − d/dmax

Figure 5. Pretraining objective. Four parallel heads consume the encoder output {zv }v∈V : link prediction (BCE), per-type masked node modelling (MSE), graph-distance regression (MSE), and an Experiment↔Experiment contrastive loss (MSE on cosine similarity). Their weighted sum, using Kendall et al. (2018) uncertainty weighting, forms the training objective.

BFS expansion to prevent hub-dominated subgraphs but are included in the final subgraph if both endpoints are present. The multi-anchor configuration rebalances the GenomicGene fraction from 78% to 15% and increases Experiment-node coverage roughly 13× relative to a naive 5-hop BFS. The 30% link-supervision split is computed deterministically via MD5 hashing of (src, rel, dst), so an edge belongs to the same set regardless of which subgraph contains it; edges incident to held-out Experiment nodes are kept in the message-passing graph but excluded from supervision. Continuous scalar features are z-score normalised per node type, and random-walk positional encodings are added to each subgraph at materialisation.

virtual-node edges, bringing the per-batch typed-tuple count above 100. Sampling. We sample 10,000 subgraphs with the multianchor expansion (k=3, Nmax =1000) of Section 3.5. Pretraining uses the deterministic MD5 edge-supervision split described in Section 3.5; the downstream titer probe uses a 5x CV Experiment-node split (Experiment IDs hashed and bucketed) so test experiments are unseen during both pretrain and probe fitting. Model configurations.

We evaluate three scales:

Table 2. Model configurations used in this paper.

3.6. Downstream Tasks The prediction target is the measured product titer of each fermentation Experiment: the experimentally reported concentration of the target compound from a physical microbial culture. Titers are drawn from two sources: published metabolic-engineering studies (literature-mined) and unpublished in-house design-build-test-learn runs recorded in our LIMS. After pretraining, we train lightweight probes on frozen experiment-node embeddings: a linear probe and a two-layer MLP probe for titer regression. We report R2 , RMSE, and Spearman ρ for regression, and AUROC and F1 for binary classification (above/below median titer). This protocol isolates representation quality from probe capacity.

Config

Hidden

Heads

Layers

FFN

Params

Demo 500M 3B

64 128 256

4 4 8

6 6 12

4 4 4

∼80M ∼0.5B ∼3B

Training. AdamW (β1 =0.9, β2 =0.999) with cosine annealing and linear warmup; bfloat16 mixed precision; FSDP for distributed training across Intel Data Center GPU Max accelerators (Dawn). Gradient clipping at ℓ2 norm 1.0. Hyperparameters tuned via a two-tier Optuna (Akiba et al., 2019) sweep: Tier 1 (1,000 trials at 500M, ∼500 XPUhours, search over h ∈ {64, 128, 192, 256}, L ∈ {4, 6, 8}); Tier 2 (200 trials at 3B, ∼800 XPU-hours, search narrowed from Tier 1). A single model trains to convergence in 7 XPUhours at 500M (7 h on one Max GPU) and 160 XPU-hours at 3B (40 h on four); the ∼500 and ∼800 XPU-hour budgets above amortise over the 1,000 and 200 median-pruned trials of each tier. Reported parameter counts are for the trainable HGT only: the ESM-2, MoLFormer, and PubMedBERT encoders are frozen and their node features precomputed, so pretraining does not backpropagate through them.

4. Experiments 4.1. Setup Knowledge graph statistics. C ANOPY’s graph comprises 6.9 M nodes across 13 types and 11.2 M edges across 34 typed relations (Table 1, Section 3.1). At training time we additionally materialise reverse edges, self-loops, and 6

Canopy: A Heterograph Foundation Model for Metabolic Engineering Table 4. Pretraining-objective ablation (probe R2 , 500M scale, 5k samples).

Table 3. Titer prediction on the held-out test split (n=410 unique experiments). All methods share the same MD5-hashed split. The graph baselines separate backbone (homogeneous SAGEConv vs. heterogeneous HGT) from C ANOPY’s SignNet/JK/virtual-node augmentations (SN/JK/VN); C ANOPY is the HGT backbone with those augmentations. Method

R2 ↑

AUROC↑

Ridge (raw conditions) XGBoost (raw conditions) MLP (raw conditions) GraphSAGE (vanilla) GraphSAGE + SN/JK/VN HGT (vanilla)

0.086 0.236 0.078 0.241 0.334 0.308

0.698 0.726 0.691 0.713 0.734 0.711

C ANOPY (Demo) C ANOPY (500M) C ANOPY (3B)

0.301 0.380 0.413

0.703 0.805 0.820

Link

MNM

Dist

Contrast

R2 ↑ / ∆pp

✓ ✓

✓ ✓ ✓

✓ ✓ ✓ ✓

0.359 −1.6 −0.6 −4.7 −7.6

✓ ✓ ✓

✓ ✓

Table 5. Architectural ablation (500M, 5k samples; probe R2 ). Variant Full C ANOPY − SignNet PE − JK aggregation − Virtual nodes − All three

Baselines. Ridge, MLP (two hidden layers, 128→64), and XGBoost (500 trees, depth 6) trained directly on the 429-dim raw experiment-condition feature vector (the same input that feeds the Experiment node in the graph), using the identical MD5-hashed train/test split as the C ANOPY probe. We further compare against graph baselines sharing the same probe and split: a homogeneous GraphSAGE backbone (evaluated both in a vanilla form and with C ANOPY’s SignNet, Jumping Knowledge, and virtual-node augmentations) and a vanilla HGT backbone without those augmentations. Together with C ANOPY itself (the HGT backbone with the augmentations), these separate the contribution of the heterogeneous backbone from that of the augmentations.

R2 ↑ / ∆pp 0.359 −2.7 −3.5 −10.3 −11.1

port headline numbers. Mitigations beyond JK and scaled residuals (DropEdge, GraphNorm, or deeper feature fusion) remain unexplored and are a natural next ablation. Task weighting. The four pretraining losses operate on very different scales: BCE for link prediction, MSE for masked-node and distance, and contrastive cosine for experiment pairs. Under flat (equal) weighting the highest-scale loss dominates the gradient and the probe collapses: removing the learned weighting drops R2 by 9.18 (∆R2 =−9.18 relative to the learned scheme), well below the constantmean predictor (Table 7). The homoscedastic uncertainty scheme of Kendall et al. (2018) attains R2 =0.359. The learned log σi values converge to per-task scales that balance gradient contributions automatically and adapt as the relative difficulty of each task shifts during training, removing the manual sweep over fixed per-task weights that would otherwise be needed.

4.2. Main Results 4.3. Ablations Unless stated otherwise, ablations are run at the 500M scale defined in Section 4.1 with a 5k-sample budget; the runs in Table 3 use the full 10k-sample budget. Absolute R2 values in the ablations therefore sit below the headline numbers and should be read as relative comparisons.

5. Discussion Each of the four pretraining objectives contributes (Table 4), with no single task dominating. Learned uncertainty weighting (Table 7) eliminates manual loss balancing and adapts as training progresses.

We ablate the four self-supervised tasks (Table 4), the three architectural augmentations (SignNet PE, Jumping Knowledge, and virtual nodes; Table 5), and network depth (Table 6). For depth, we compare the default L=6 architecture against an iso-parameter shallow variant (L=2, h=224) at 500M scale. The shallow model reaches a higher peak probe R2 , runs ∼2× faster per epoch, and, unlike the deeper variant, does not exhibit probe degradation through 20 epochs. We attribute this to a lower oversmoothing burden once HGTConv attention has enough per-head capacity (dh =56); Jumping Knowledge alone is insufficient at L=6 on this graph. We retain the deeper default in the headline configuration: the gap is small (∆R2 =0.011) and the deeper topology matches the 3B configuration from which we re-

The 3B model improves R2 from 0.380 to 0.413 over the 500M model, a modest gain relative to the 6× parameter increase. With 4,791 total literature-mined fermentation records, the regime is likely data-bound rather than capacitybound at this scale; substantiating scaling claims will require more experimental data, particularly across the long tail of compounds and organisms. In deployment, titer prediction is used to rank candidate designs for the wet lab rather than to replace measurement; in this triage setting the meaningful gain is in regression, from the best tabular 7

Canopy: A Heterograph Foundation Model for Metabolic Engineering Table 6. Depth ablation (iso-parameter, 500M scale, 5k samples). 2

Variant

R ↑ / ∆pp

AUROC / ∆pp

0.3703 −1.0

0.7881 −1.1

L=2, h=224 (shallow) L=6, h=128 (deep)

gesting orders of magnitude more protein sequences, and adding predicted protein structures as a complementary modality. On the data side, we are expanding the KG with DNA parts (promoters, RBSs, terminators, CDS variants) and metagenomic sequences from environmental and engineered communities, broadening coverage well beyond the current cultured-organism backbone. On the application side, the same frozen embeddings are being extended from titer prediction to chassis selection and de novo pathway design, both of which reuse C ANOPY’s heterogeneous representation but condition on different query node types. Substantiating the “foundation model” framing also requires breadth in downstream evaluation; beyond titer prediction, we plan probes on chassis selection, gene essentiality, reaction prediction, and cross-organism transfer (e.g., pretraining on E. coli experiments and evaluating on yeast) to test whether the same frozen embeddings transfer across distinct metabolic-engineering tasks. Finally, we are pairing the learned representation with generative strain-design pipelines: a flow-matching model conditioned on C ANOPY embeddings proposes candidate multi-gene interventions, which are then scored by the frozen titer probe inside a Bayesian-optimisation loop, turning the predictive oracle into a closed-loop generative design system.

Table 7. Task-weighting comparison (500M, 5k samples). Weighting

R2 ↑

∆R2

Learned (Kendall) Flat (equal)

0.359 —

— −9.18

baseline (R2 =0.24) to C ANOPY’s cross-organism representation (R2 =0.41). We report AUROC for completeness but read it with caution: fermentation volume (feat volume) correlates with titer (Pearson 0.15) and on its own separates above/below-median titer at AUROC 0.65, so the binary task is partly trivialised and all methods sit in a narrow AUROC band regardless of regression skill. Because each fermentation costs weeks of bench time, even a modest improvement in how designs are ranked can reduce the number of physical builds per successful design, though we do not yet quantify this end-to-end. Limitations. First, experimental data is sparse relative to the KG, with uneven distribution across compounds and organisms. Second, the model is predictive rather than mechanistic; attention analysis offers some interpretability but does not replace mechanistic modelling. Third, our baselines do not include a graph-free pooled-encoder control, e.g., concatenating frozen ESM-2 and MoLFormer embeddings of a strain’s genes and target compound and applying an MLP; such a probe would isolate the contribution of graph structure from the underlying foundation-model encoders and is a planned ablation. Fourth, release scope is constrained for this submission: the 4,791-experiment literature-mined benchmark and its split files will be released in a forthcoming publication, while the in-house LIMS records and trained model weights are not released with this workshop paper, with an archival release of the full pipeline planned to accompany a later journal submission. Fifth, we do not yet isolate the contribution of individual data sources. A node-type feature-masking ablation, which zeros proteomic (UniRef/InterPro), genomic, or transcriptomic features at probe time, would reveal which modalities are influential versus redundant. A per-fold check would test whether the experiment-level split induces feature-distribution shift. Both are planned and inform the data-expansion priorities below.

Broader impact. Accelerating strain design supports a shift from petrochemical to fermentative production of chemicals, fuels, and therapeutics.

6. Conclusion Frozen C ANOPY embeddings outperform tabular and homogeneous graph baselines on titer prediction, providing a cross-organism representation that can be reused as the conditioning substrate for downstream generative strain-design pipelines.

Impact Statement This paper presents work whose goal is to advance the field of machine learning for biological design. There are many potential societal consequences of our work, including positive impact through greener bioproduction.

Acknowledgements The authors acknowledge the use of resources provided by the Dawn National AI Research Resource (AIRR). Dawn is operated by the University of Cambridge and is funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) via UK Research and Innovation, the Science and Technology Facilities Council [ST/Z000890/1], Dell Technologies and Intel.

Future work. Several extensions are already underway. On the representation side, we are scaling node features with larger and more sophisticated embedding models, in8

Canopy: A Heterograph Foundation Model for Metabolic Engineering

References

Lu, A. X., Mehta, R., Mofrad, M. R. K., Ng, M. Y., Pannu, J., Ré, C., Schmok, J. C., John, J. S., Sullivan, J., Zhu, K., Zynda, G., Balsam, D., Collison, P., Costa, A. B., Hernandez-Boussard, T., Ho, E., Liu, M.-Y., McGrath, T., Powell, K., Burke, D. P., Goodarzi, H., Hsu, P. D., and Hie, B. L. Genome modeling and design across all domains of life with Evo 2, February 2025. URL https://www.biorxiv.org/ content/10.1101/2025.02.18.638918v1. Pages: 2025.02.18.638918 Section: New Results.

Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A. J., Bambrick, J., Bodenstein, S. W., Evans, D. A., Hung, C.-C., O’Neill, M., Reiman, D., Tunyasuvunakool, K., Wu, Z., Žemgulytė, A., Arvaniti, E., Beattie, C., Bertolli, O., Bridgland, A., Cherepanov, A., Congreve, M., Cowen-Rivers, A. I., Cowie, A., Figurnov, M., Fuchs, F. B., Gladman, H., Jain, R., Khan, Y. A., Low, C. M. R., Perlin, K., Potapenko, A., Savy, P., Singh, S., Stecula, A., Thillaisundaram, A., Tong, C., Yakneen, S., Zhong, E. D., Zielinski, M., Žı́dek, A., Bapst, V., Kohli, P., Jaderberg, M., Hassabis, D., and Jumper, J. M. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 630(8016): 493–500, June 2024. ISSN 1476-4687. doi: 10.1038/ s41586-024-07487-w. URL https://www.nature. com/articles/s41586-024-07487-w.

Brown, G. R., Hem, V., Katz, K. S., Ovetsky, M., Wallin, C., Ermolaeva, O., Tolstoy, I., Tatusova, T., Pruitt, K. D., Maglott, D. R., and Murphy, T. D. Gene: a gene-centered information resource at NCBI. Nucleic Acids Research, 43(Database issue):D36–42, January 2015. ISSN 13624962. doi: 10.1093/nar/gku1055. URL https://doi. org/10.1093/nar/gku1055. Burgard, A. P., Pharkya, P., and Maranas, C. D. Optknock: A bilevel programming framework for identifying gene knockout strategies for microbial strain optimization. Biotechnology and Bioengineering, 84(6): 647–657, 2003. ISSN 1097-0290. doi: 10.1002/bit. 10803. URL https://onlinelibrary.wiley. com/doi/abs/10.1002/bit.10803.

Ahern, W., Yim, J., Tischer, D., Salike, S., Woodbury, S. M., Kim, D., Kalvet, I., Kipnis, Y., Coventry, B., Altae-Tran, H. R., Bauer, M. S., Barzilay, R., Jaakkola, T. S., Krishna, R., and Baker, D. Atom-level enzyme active site scaffolding using RFdiffusion2. Nature Methods, 23(1): 96–105, January 2026. ISSN 1548-7105. doi: 10.1038/ s41592-025-02975-x. URL https://www.nature. com/articles/s41592-025-02975-x.

Chandak, P., Huang, K., and Zitnik, M. Building a knowledge graph to enable precision medicine. Scientific Data, 10(1):67, February 2023. ISSN 2052-4463. doi: 10.1038/ s41597-023-01960-3. URL https://www.nature. com/articles/s41597-023-01960-3.

Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’19, pp. 2623–2631, New York, NY, USA, July 2019. Association for Computing Machinery. ISBN 978-1-4503-6201-6. doi: 10.1145/3292500.3330701. URL https://dl.acm. org/doi/10.1145/3292500.3330701.

Cheng, Y., Bi, X., Xu, Y., Liu, Y., Li, J., Du, G., Lv, X., and Liu, L. Machine learning for metabolic pathway optimization: A review. Computational and Structural Biotechnology Journal, 21:2381–2393, January 2023. ISSN 20010370. doi: 10.1016/j.csbj.2023.03.045. URL https: //doi.org/10.1016/j.csbj.2023.03.045.

Ashburner, M., Ball, C. A., Blake, J. A., Botstein, D., Butler, H., Cherry, J. M., Davis, A. P., Dolinski, K., Dwight, S. S., Eppig, J. T., Harris, M. A., Hill, D. P., Issel-Tarver, L., Kasarskis, A., Lewis, S., Matese, J. C., Richardson, J. E., Ringwald, M., Rubin, G. M., and Sherlock, G. Gene Ontology: tool for the unification of biology. Nature Genetics, 25(1):25–29, May 2000. ISSN 1546-1718. doi: 10.1038/75556. URL https://www.nature.com/ articles/ng0500_25.

Cui, H., Wang, C., Maan, H., Pang, K., Luo, F., Duan, N., and Wang, B. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods, 21(8):1470–1480, August 2024. ISSN 1548-7105. doi: 10.1038/ s41592-024-02201-0. URL https://www.nature. com/articles/s41592-024-02201-0.

Brixi, G., Durrant, M. G., Ku, J., Poli, M., Brockman, G., Chang, D., Gonzalez, G. A., King, S. H., Li, D. B., Merchant, A. T., Naghipourfar, M., Nguyen, E., Ricci-Tam, C., Romero, D. W., Sun, G., Taghibakshi, A., Vorontsov, A., Yang, B., Deng, M., Gorton, L., Nguyen, N., Wang, N. K., Adams, E., Baccus, S. A., Dillmann, S., Ermon, S., Guo, D., Ilango, R., Janik, K.,

Czajka, J. J., Oyetunde, T., and Tang, Y. J. Integrated knowledge mining, genome-scale modeling, and machine learning for predicting Yarrowia lipolytica bioproduction. Metabolic Engineering, 67:227–236, September 2021. ISSN 1096-7176. doi: 10.1016/j.ymben.2021.07. 003. URL https://www.sciencedirect.com/ science/article/pii/S1096717621001130. 9

Canopy: A Heterograph Foundation Model for Metabolic Engineering

Dwivedi, V. P., Luu, A. T., Laurent, T., Bengio, Y., and Bresson, X. Graph Neural Networks with Learnable Structural and Positional Representations, February 2022. URL http://arxiv.org/abs/2110. 07875. arXiv:2110.07875 [cs].

and Baranzini, S. E. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife, 6:e26726, September 2017. ISSN 2050-084X. doi: 10.7554/eLife.26726. URL https://doi.org/10. 7554/eLife.26726.

Fey, M. and Lenssen, J. E. Fast Graph Representation Learning with PyTorch Geometric, April 2019. URL http:// arxiv.org/abs/1903.02428. arXiv:1903.02428 [cs].

Hu, E. Y., Oleshko, S., Firmani, S., Cheng, H., Zhu, Z., Ulmer, M., Arnold, M., Colomé-Tatché, M., Tang, J., Xhonneux, S., and Marsico, A. Enhancing link prediction in biomedical knowledge graphs with BioPathNet. Nature Biomedical Engineering, pp. 1– 23, January 2026. ISSN 2157-846X. doi: 10.1038/ s41551-025-01598-z. URL https://www.nature. com/articles/s41551-025-01598-z.

Gema, A. P., Grabarczyk, D., De Wulf, W., Borole, P., Alfaro, J. A., Minervini, P., Vergari, A., and Rajan, A. Knowledge graph embeddings in the biomedical domain: are they useful? A look at link prediction, rule learning, and downstream polypharmacy tasks. Bioinformatics Advances, 4(1):vbae097, January 2024. ISSN 2635-0041. doi: 10.1093/bioadv/vbae097. URL https://doi. org/10.1093/bioadv/vbae097.

Hu, W., Liu, B., Gomes, J., Zitnik, M., Liang, P., Pande, V., and Leskovec, J. Strategies for Pre-training Graph Neural Networks, February 2020a. URL http://arxiv. org/abs/1905.12265. arXiv:1905.12265 [cs]. Hu, Z., Dong, Y., Wang, K., and Sun, Y. Heterogeneous graph transformer. CoRR, abs/2003.01332, 2020b. URL https://arxiv.org/abs/2003.01332.

Gilmer, J., Schoenholz, S. S., Riley, P. F., Vinyals, O., and Dahl, G. E. Neural Message Passing for Quantum Chemistry. In Proceedings of the 34th International Conference on Machine Learning, pp. 1263–1272. PMLR, July 2017. URL https://proceedings.mlr.press/v70/ gilmer17a.html. Hao, M., Gong, J., Zeng, X., Liu, C., Guo, Y., Cheng, X., Wang, T., Ma, J., Zhang, X., and Song, L. Large-scale foundation model on singlecell transcriptomics. Nature Methods, 21(8):1481– 1491, August 2024. ISSN 1548-7105. doi: 10.1038/ s41592-024-02305-7. URL https://www.nature. com/articles/s41592-024-02305-7. Hasibi, R., Michoel, T., and Oyarzún, D. A. Integration of graph neural networks and genomescale metabolic models for predicting gene essentiality. npj Systems Biology and Applications, 10(1): 24, March 2024. ISSN 2056-7189. doi: 10.1038/ s41540-024-00348-2. URL https://www.nature. com/articles/s41540-024-00348-2. Hayes, T., Rao, R., Akin, H., Sofroniew, N. J., Oktay, D., Lin, Z., Verkuil, R., Tran, V. Q., Deaton, J., Wiggert, M., Badkundri, R., Shafkat, I., Gong, J., Derry, A., Molina, R. S., Thomas, N., Khan, Y. A., Mishra, C., Kim, C., Bartie, L. J., Nemeth, M., Hsu, P. D., Sercu, T., Candido, S., and Rives, A. Simulating 500 million years of evolution with a language model. Science, 387(6736):850–858, February 2025. doi: 10.1126/ science.ads0018. URL https://www.science. org/doi/10.1126/science.ads0018. Himmelstein, D. S., Lizee, A., Hessler, C., Brueggeman, L., Chen, S. L., Hadley, D., Green, A., Khankhanian, P., 10

Huang, K., Fu, T., Gao, W., Zhao, Y., Roohani, Y. H., Leskovec, J., Coley, C. W., Xiao, C., Sun, J., and Zitnik, M. Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development. June 2021. URL https://openreview. net/forum?id=8nvgnORnoWr. Huang, K., Chandak, P., Wang, Q., Havaldar, S., Vaid, A., Leskovec, J., Nadkarni, G. N., Glicksberg, B. S., Gehlenborg, N., and Zitnik, M. A foundation model for clinician-centered drug repurposing. Nature Medicine, 30(12):3601–3613, December 2024. ISSN 1546-170X. doi: 10.1038/ s41591-024-03233-x. URL https://www.nature. com/articles/s41591-024-03233-x. Kanehisa, M. and Goto, S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research, 28(1): 27–30, January 2000. ISSN 0305-1048. doi: 10.1093/nar/ 28.1.27. URL https://doi.org/10.1093/nar/ 28.1.27. Kendall, A., Gal, Y., and Cipolla, R. Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics, April 2018. URL http://arxiv. org/abs/1705.07115. arXiv:1705.07115 [cs]. Khodayari, A. and Maranas, C. D. A genome-scale Escherichia coli kinetic metabolic model k-ecoli457 satisfying flux data for multiple mutant strains. Nature Communications, 7(1):13806, 2016. doi: 10.1038/ ncomms13806. URL https://www.nature.com/ articles/ncomms13806.

Canopy: A Heterograph Foundation Model for Metabolic Engineering

Li, Z., Zeng, Z., Lin, X., Fang, F., Qu, Y., Xu, Z., Liu, Z., Ning, X., Wei, T., Liu, G., Tong, H., and He, J. Flow matching meets biology and life science: a survey. npj Artificial Intelligence, 2(1):17, January 2026. ISSN 3005-1460. doi: 10.1038/ s44387-025-00066-y. URL https://www.nature. com/articles/s44387-025-00066-y. Lim, D., Robinson, J., Zhao, L., Smidt, T., Sra, S., Maron, H., and Jegelka, S. Sign and Basis Invariant Networks for Spectral Graph Representation Learning, September 2022. URL http://arxiv.org/abs/2202. 13013. arXiv:2202.13013 [cs]. Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., and Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, March 2023. doi: 10.1126/ science.ade2574. URL https://www.science. org/doi/10.1126/science.ade2574. Lobentanzer, S., Aloy, P., Baumbach, J., Bohar, B., Charoentong, P., Danhauser, K., Doğan, T., Dreo, J., Dunham, I., Fernandez-Torras, A., Gyori, B. M., Hartung, M., Hoyt, C. T., Klein, C., Korcsmaros, T., Maier, A., Mann, M., Ochoa, D., Pareja-Lorente, E., Popp, F., Preusse, M., Probul, N., Schwikowski, B., Sen, B., Strauss, M. T., Turei, D., Ulusoy, E., Wodke, J. A. H., and SaezRodriguez, J. Democratising Knowledge Representation with BioCypher, January 2023. URL http://arxiv. org/abs/2212.13543. arXiv:2212.13543 [q-bio]. Mao, H., Chen, Z., Tang, W., Zhao, J., Ma, Y., Zhao, T., Shah, N., Galkin, M., and Tang, J. Position: Graph Foundation Models are Already Here, May 2024. URL http://arxiv.org/abs/2402. 02216. arXiv:2402.02216 [cs]. Moretti, S., Tran, V. D. T., Mehl, F., Ibberson, M., and Pagni, M. Metanetx/mnxref: unified namespace for metabolites and biochemical reactions in the context of metabolic models. Nucleic Acids Research, 49(D1):D570–D574, 01 2021. ISSN 0305-1048. doi: 10.1093/nar/gkaa992. URL https://doi.org/10.1093/nar/gkaa992.

ate knowledge-based biologically meaningful machinereadable embeddings. Nature Communications, 10(1): 3045, July 2019. ISSN 2041-1723. doi: 10.1038/ s41467-019-11069-0. URL https://www.nature. com/articles/s41467-019-11069-0. Nguyen, E., Poli, M., Durrant, M. G., Kang, B., Katrekar, D., Li, D. B., Bartie, L. J., Thomas, A. W., King, S. H., Brixi, G., Sullivan, J., Ng, M. Y., Lewis, A., Lou, A., Ermon, S., Baccus, S. A., Hernandez-Boussard, T., Ré, C., Hsu, P. D., and Hie, B. L. Sequence modeling and design from molecular to genome scale with evo. Science, 386(6723):eado9336, 2024. doi: 10.1126/ science.ado9336. URL https://www.science. org/doi/abs/10.1126/science.ado9336. Opgenorth, P., Costello, Z., Okada, T., Goyal, G., Chen, Y., Gin, J., Benites, V., de Raad, M., Northen, T. R., Deng, K., Deutsch, S., Baidoo, E. E. K., Petzold, C. J., Hillson, N. J., Garcia Martin, H., and Beller, H. R. Lessons from two design–build–test–learn cycles of dodecanol production in escherichia coli aided by machine learning. ACS Synthetic Biology, 8(6):1337–1351, 2019. doi: 10.1021/acssynbio.9b00020. PMID: 31072100. Oyetunde, T., Liu, D., Martin, H. G., and Tang, Y. J. Machine learning framework for assessment of microbial factory performance. PLOS ONE, 14 (1):e0210558, January 2019. ISSN 1932-6203. doi: 10.1371/journal.pone.0210558. URL https: //journals.plos.org/plosone/article? id=10.1371/journal.pone.0210558. Paysan-Lafosse, T., Blum, M., Chuguransky, S., Grego, T., Pinto, B. L., Salazar, G., Bileschi, M., Bork, P., Bridge, A., Colwell, L., Gough, J., Haft, D., Letunić, I., MarchlerBauer, A., Mi, H., Natale, D., Orengo, C., Pandurangan, A., Rivoire, C., Sigrist, C. J. A., Sillitoe, I., Thanki, N., Thomas, P. D., Tosatto, S. C. E., Wu, C., and Bateman, A. InterPro in 2022. Nucleic Acids Research, 51(D1):D418– D427, January 2023. ISSN 0305-1048. doi: 10.1093/ nar/gkac993. URL https://doi.org/10.1093/ nar/gkac993. Porokhin, V., Liu, L.-P., and Hassoun, S. Using graph neural networks for site-of-metabolism prediction and its applications to ranking promiscuous enzymatic products. Bioinformatics, 39(3):btad089, March 2023. ISSN 1367-4811. doi: 10.1093/ bioinformatics/btad089. URL https://doi.org/ 10.1093/bioinformatics/btad089.

Márquez-Zavala, E., Bartolomeu, F. D., and Machado, D. A database of over 15.000 strain design publications reveals a conserved set of metabolic engineering targets across microbial hosts and products, December 2025. URL https://www.biorxiv.org/content/10. 64898/2025.12.15.694291v1. ISSN: 2692-8205 Pages: 2025.12.15.694291 Section: New Results.

Radivojević, T., Costello, Z., Workman, K., and Garcia Martin, H. A machine learning Automated Recommendation Tool for synthetic biology. Nature Communications, 11(1):4879, September 2020. ISSN 2041-1723. doi: 10.1038/

Nelson, C. A., Butte, A. J., and Baranzini, S. E. Integrating biomedical research and electronic health records to cre11

Canopy: A Heterograph Foundation Model for Metabolic Engineering

s41467-020-18008-4. URL https://www.nature. com/articles/s41467-020-18008-4. Ross, J., Belgodere, B., Chenthamarakshan, V., Padhi, I., Mroueh, Y., and Das, P. Large-Scale Chemical Language Representations Capture Molecular Structure and Properties, June 2021. URL https://arxiv.org/abs/ 2106.09553v3.

neering. ACS Synthetic Biology, 12(9):2588–2599, August 2023. ISSN 2161-5063. doi: 10.1021/acssynbio. 3c00186. URL https://pmc.ncbi.nlm.nih. gov/articles/PMC10510747/. Wang, Z., Liu, Z., Ma, T., Li, J., Zhang, Z., Fu, X., Li, Y., Yuan, Z., Song, W., Ma, Y., Zeng, Q., Chen, X., Zhao, J., Li, J., Jiang, M., Lio, P., Chawla, N., Zhang, C., and Ye, Y. Graph Foundation Models: A Comprehensive Survey, May 2025. URL http://arxiv.org/abs/2505. 15116. arXiv:2505.15116 [cs].

Schneider, P., Bekiaris, P. S., von Kamp, A., and Klamt, S. StrainDesign: a comprehensive Python package for computational design of metabolic networks. Bioinformatics, 38(21):4981–4983, November 2022. ISSN 1367-4811. doi: 10.1093/bioinformatics/ btac632. URL https://doi.org/10.1093/ bioinformatics/btac632.

Xin, K., Wang, Q., Chen, J., Yu, P., Zhao, H., and Ji, H. Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production, October 2024. URL http://arxiv.org/ abs/2410.18475. arXiv:2410.18475 [cs].

Schoch, C. L., Ciufo, S., Domrachev, M., Hotton, C. L., Kannan, S., Khovanskaya, R., Leipe, D., Mcveigh, R., O’Neill, K., Robbertse, B., Sharma, S., Soussov, V., Sullivan, J. P., Sun, L., Turner, S., and KarschMizrachi, I. NCBI Taxonomy: a comprehensive update on curation, resources and tools. Database, 2020: baaa062, January 2020. ISSN 1758-0463. doi: 10. 1093/database/baaa062. URL https://doi.org/ 10.1093/database/baaa062.

Xu, K., Li, C., Tian, Y., Sonobe, T., Kawarabayashi, K.i., and Jegelka, S. Representation Learning on Graphs with Jumping Knowledge Networks. In Proceedings of the 35th International Conference on Machine Learning, pp. 5453–5462. PMLR, July 2018. URL https:// proceedings.mlr.press/v80/xu18c.html.

Song, T., Yin, L., Han, Z., and Xu, Z. Improving Enzyme Prediction with Chemical Reaction Equations by Hypergraph-Enhanced Knowledge Graph Embeddings, January 2026. URL http://arxiv.org/ abs/2601.05330. arXiv:2601.05330 [cs]. Su, J., Han, C., Zhou, Y., Shan, J., Zhou, X., and Yuan, F. SaProt: Protein Language Modeling with Structure-aware Vocabulary, April 2024. URL https://www.biorxiv.org/content/ 10.1101/2023.10.01.560349v5. Pages: 2023.10.01.560349 Section: New Results. Suzek, B. E., Wang, Y., Huang, H., McGarvey, P. B., Wu, C. H., and the UniProt Consortium. UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics, 31(6):926–932, March 2015. ISSN 1367-4803. doi: 10.1093/bioinformatics/btu739. URL https://doi. org/10.1093/bioinformatics/btu739. The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research, 51 (D1):D523–D531, January 2023. ISSN 0305-1048. doi: 10.1093/nar/gkac1052. URL https://doi.org/10. 1093/nar/gkac1052. van Lent, P., Schmitz, J., and Abeel, T. Simulated Design–Build–Test–Learn Cycles for Consistent Comparison of Machine Learning Methods in Metabolic Engi12

Canopy: A Heterograph Foundation Model for Metabolic Engineering

A. Knowledge graph construction details C ANOPY’s graph is assembled with BioCypher (Lobentanzer et al., 2023), which separates what the graph contains from how it is collected. A single declarative schema (YAML, keyed by Biolink-aligned node and edge types) fixes the allowed types, their identifier namespaces, and their property contracts. Ten adapter modules each emit typed node and edge streams against that schema; BioCypher rejects labels that are not declared, collapses sub-classes onto the declared parent type, and writes neo4j-admin-importable CSVs. Adapters can therefore be added, swapped, or version-pinned without touching graph code, and the same schema drives both build-time validation and downstream sampling. The ten adapters span public reference resources (sequence, structure, pathway, ontology), a literature-mined fermentation corpus, and one in-house LIMS adapter contributing unpublished DBTL records (Table 8). Table 8. BioCypher adapters used to construct C ANOPY’s KG. Adapter

Source

Nodes / edges produced

Genomic

Gene Ontology

UniProt REST + curated chassis list UniRef90 .tsv, MetaNetX cross-refs protein2ipr.dat.gz (bulk) or UniProt batch API GO OBO release (obonet)

MetaNetX

MNXref (Moretti et al., 2021)

Taxonomy Transcriptomic

NCBI Taxonomy Public RNA-seq compendia

Experimental scrape

Literature-mined records

Pathway enrichment

KEGG REST

LIMS (in-house)

Internal experiment registry

GenomicGene; ENCODES Chassis, (Gene→UnirefCluster), HAS GENE (Chassis→Gene). CATALYZED BY UnirefCluster; (Reaction→UnirefCluster), HAS GO TERM. HAS DOMAIN InterProDomain; (UnirefCluster→Domain). (BP/MF/CC); ontology hierarchy edges GOTerm (SUBCLASS OF, PART OF, REGULATES). Metabolite, Reaction; HAS PARTICIPANT, HAS PRODUCT. Taxon; BELONGS TO (Strain→Taxon). Transcriptomic measurement nodes; MEASURED BY TRANSCRIPTOMIC, DERIVED FROM TRANSCRIPTOMIC, USES STRAIN TRANSCRIPTOMIC edges. Experiment, Strain, MetabolicPathway; USES STRAIN, HAS PATHWAY, TARGETS COMPOUND, MEASURES COMPOUND, HAS KNOCKOUT, HAS KNOCKIN, HAS OVEREXPRESSION. enriches MetabolicPathway nodes with KEGG metadata; HAS STEP (Pathway→Reaction). supplements Experiment / Strain with unpublished DBTL records and the edit-edge types above.

UniRef InterPro

fermentation

Schema harmonisation. BioCypher’s ontology layer maps every adapter’s emitted labels to Biolink superclasses. Only types declared in the project schema are retained at build time, and sub-class collapse prevents label fragmentation (e.g. “BiologicalProcess” and “MolecularFunction” are both rolled up under GOTerm). Edges referencing undeclared node types are dropped before import and logged. Identifier normalisation and deduplication. Cross-source identifiers follow a fixed precedence: MetaNetX MNX IDs for metabolites and reactions, UniRef90 cluster IDs for proteins, NCBI gene IDs for genes, NCBI Taxonomy IDs for taxa, GO IDs for ontology terms, and InterPro IDs for domains. Organism names in the literature-mined corpus are normalised through a curated synonym table that maps frequently-confused taxonomic labels (e.g. Pichia pastoris → Komagataella phaffii, Clostridium thermocellum → Acetivibrio thermocellus) to the canonical chassis list. Duplicate metabolite and reaction nodes from MetaNetX cross-references are merged on MNX ID; Uniref→Reaction and Uniref→GOTerm mappings are resolved against the MetaNetX curated reaction model and the GO release pinned at build time. Edges with malformed or missing endpoints are dropped at adapter time and logged. Build pipeline. Adapters write CSVs to a staging directory, BioCypher emits the matching neo4j-admin headers and edge files, and the resulting graph is loaded into Neo4j 5 via neo4j-admin import. The final graph contains 6.86 M nodes and 11.2 M schema-typed edges across 13 node types and 34 typed relations (Table 1). 13

Canopy: A Heterograph Foundation Model for Metabolic Engineering

B. Hyperparameter sweep details Hyperparameters are selected with a staged Optuna sweep that narrows the search as model scale increases (Section 4.1). Tier 1 explores broadly at the 500M scale on a single XPU; Tier 2 refines around the Tier 1 optimum at the 3B operating point under FSDP across four XPUs, optionally with gradient checkpointing. All trials optimise the held-out titer probe R2 at its best epoch, and both studies use Optuna’s TPESampler with median pruning. Tables 9 and 10 list the Tier 1 and Tier 2 search spaces in full. Trials whose hidden channels is not divisible by heads are pruned at sample time. Table 9. Tier 1 search space (500M scale, single XPU). 1,000 trials, ∼500 XPU-hours total compute budget, 80 epochs per trial, probe evaluated every 5 epochs. Parameter

Range / set

Sampling

Architecture hidden channels num layers heads ffn expansion dropout

{64, 128, 192, 256} [2, 8] [2, 8] {2, 4} [0.05, 0.40]

categorical integer integer categorical float

Optimisation lr weight decay warmup epochs cosine end offset batch size accumulate grad batches

[10−4 , 5×10−3 ] [10−4 , 10−1 ] [1, 10] [1, 30] {8, 16, 32, 64, 128, 256} {1, 2, 4}

log-uniform log-uniform integer integer categorical categorical

Probe geometry probe hidden dim probe num layers probe dropout

{32, 64, 128, 256} [1, 4] [0.0, 0.3]

categorical integer float

Table 10. Tier 2 search space (3B scale, FSDP/4). 200 trials, 30 epochs per trial, probe evaluated every 5 epochs. Initial trials warm-started from the top-10 Tier-1 configurations. Parameter

Range / set

Sampling

Architecture hidden channels num layers heads ffn expansion dropout

{192, 256, 384} [8, 12] [4, 8] {2, 4} [0.05, 0.30]

categorical integer integer categorical float

Optimisation lr weight decay warmup epochs cosine end offset batch size accumulate grad batches

[5×10−5 , 3×10−3 ] [10−4 , 10−1 ] [2, 10] [1, 15] {32, 64, 128} {1, 2, 4, 8}

log-uniform log-uniform integer integer categorical categorical

The Tier-1 winner combines L=6, h=128 (matching the 500M row of Table 2) with bs=256, lr=2×10−3 , four warmup epochs, a cosine end offset of 6, and all four pretraining losses active. This configuration is used for the 500M headline run and seeded the Tier-2 search at 3B scale.

14

Record · ID 346507 · SHA-256 1fcd7e3ae3417376
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.