Conceptio › Archive › arXiv CS
arXiv CSopen access

Dynamic language model representations for multi-objective reaction optimisation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Dynamic language model representations for multi-objective reaction optimisation

arXiv:2609.11790v1 [cs.LG] 10 Sep 2026

Joshua W. Sin * 1 2 David Ming Segura * 1 2 3 Bojana Ranković * 2 3 Siu Lun Chau 4 Marius D. R. Lutz 5 Andrea Anelli 5 Ryan P. Burwood 6 Kurt Püntener 1 Maximilian J. Notheis 1 Raphael Bigler 1 Philippe Schwaller 2 3

Abstract

prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of highthroughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.

Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied

1. Introduction Chemical reaction optimisation is essential to chemical synthesis, and optimising reaction conditions requires navigating vast combinatorial spaces of reaction components such as ligands, catalysts, and bases, while balancing multiple objectives such as yield, selectivity, and cost [1–5]. This challenge is particularly acute in pharmaceutical process and preclinical chemistry, where identifying robust conditions for active pharmaceutical ingredient (API) synthesis entails additional environmental, health, and safety considerations [6–9], and in academic settings where enabling novel transformations often requires extensive exploration of many diverse reaction parameters [10, 11]. Machine learning approaches have proved remarkably effective, with Bayesian optimisation emerging as the leading framework for data-efficient reaction optimisation. It has been applied in low-data regimes [12–19] and, more recently, in parallel high-throughput settings [20, 21]. Despite these algorithmic advances, how best to represent reactions for model-driven optimisation remains an open and consequential challenge.

*

Equal contribution 1 Process Chemistry & Catalysis, Synthetic Molecules Technical Development, F. Hoffmann-La Roche AG, Basel, Switzerland 2 Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland 3 National Centre of Competence in Research (NCCR) Catalysis, EPFL, Lausanne, Switzerland 4 Epistemic Intelligence & Computation Lab, College of Computing & Data Science, Nanyang Technological University, Singapore 5 Roche Pharma Research and Early Development (pRED), F. Hoffmann-La Roche AG, Basel, Switzerland 6 Solid State Sciences, Synthetic Molecules Technical Development, F. Hoffmann-La Roche AG, Basel, Switzerland. Correspondence to: Joshua W. Sin <wing [email protected]>, David Ming Segura <[email protected]>, Bojana Ranković <[email protected]>, Philippe Schwaller <[email protected]>.

Molecular descriptors have been the dominant featurisation strategy in reaction optimisation and prediction (Figure 1a). Pre-computed descriptor libraries such as Kraken [25] and the COSMO-RS [26] database provide readily available steric and electronic properties for established compound classes and have been widely adopted [15, 16, 20, 21, 27– 30]. Alternatively, tailored and often laborious feature en-

Preliminary work. Under review.

1

Dynamic language model representations for multi-objective reaction optimisation

a

Approaches for featurising chemical entities in model-driven reaction optimisation Chemical descriptors

Ph P Ph

θ O

Pd

P Ph Ph

Conformer property

Conformer ensemble properties

One-hot encoding

COSMO σ surface properties

...

Max value

Ligand 1

Base 3

Boltzmann

1

Minimum

0

Solvent 1

0

0

1

0

0

1

0

0

0

1

0

0

0

0

0

0

0

0

...

...

Bite angle NBO charges Vbur …

Precursor 2

• •

Informative chemical properties Can be tailored to select relevant features for each specific system

• •

Inflexible across entity classes (e.g., mono- vs. bisphosphines) Laborious feature engineering

• •

Straightforward to compute Applicable to any chemical entity

• •

No incorporated chemical knowledge Sparse and high-dimensional

Static representations which do not adapt to the optimisation task as more data is collected

b

This approach: trainable large language model embeddings for reaction optimisation Joint LLM-GP training improves representations at each iteration

• •

Execute reactions experimentally

Reaction conditions:

Textual reaction data

Trainable LLM module (LoRA)

Projection heads for each objective

Independent Gaussian processes

. . .

. . .

•

•

Dynamic reaction representation

Straightforward to compute Generally applicable and unified across entity classes (mono-/ bisphosphines, etc.) Task-specific joint LLM-GP finetuning tailors representation to specific optimisation problem Dynamically improving representations as more data is collected

Acquisition function selects next experiment

Figure 1. Featurisation approaches for model-driven reaction optimisation. a, Conventional approaches for representing chemical entities. Left: molecular descriptors (e.g., Sterimol parameters, buried volume, and bite angle) incorporate physicochemical properties derived from density functional theory (DFT) [22] or extended tight-binding (xTB) [23], enabling chemical similarity-aware representations. However, descriptor sets are inflexible across entity classes, require domain-specific expertise for feature engineering, and can be computationally expensive to generate. Right: one-hot encoding assigns a binary indicator vector to each categorical component (e.g., different ligands, bases, and precursors), yielding sparse, high-dimensional representations that encode no chemical similarity. b, This approach: trainable large language model (LLM) embeddings for reaction optimisation. Textual descriptions of reaction conditions are encoded by a LoRA-fine-tuned [24] language model into dense, continuous embeddings. These embeddings are passed through objective-specific projection heads into independent Gaussian process (GP) surrogate models, and the entire architecture (LLM-GP) is trained jointly. The architecture is embedded within a multi-objective Bayesian optimisation cycle, with the reaction representations improving at each iteration, organising representations by reaction performance. This yields a unified, task-adaptive representation without requiring descriptor computation or domain-specific feature engineering.

gineering workflows, involving computationally expensive density functional theory (DFT) or semi-empirical calculations, conformer ensemble analysis, and reaction-specific mechanistic interrogation, can yield informative descriptors [22, 23, 31–37].

parameters for monodentate phosphines, for example, are only partially applicable to bidentate ligands, and reactions involving chemically heterogeneous components such as inorganic bases and organometallic additives may lack any meaningful shared descriptor space, making feature engineering increasingly prohibitive as chemical diversity grows. One-hot encoding offers a simpler alternative that is universally applicable and has shown competitive performance in some settings [12, 13, 38]. However, it produces sparse, high-dimensional representations that encode no chemical similarity, and its dimensionality scales unfavourably with the number of reaction components [20, 38] (Figure 1a).

Despite their widespread use, descriptor-based approaches face fundamental limitations. Selecting which descriptors to compute relies on a priori chemical intuition about which molecular properties govern reactivity, a process that is biased by existing mechanistic understanding and must often be revisited for each new reaction system. More critically, descriptors are often not transferable across chemically distinct reaction components. Steric and electronic

Advances in large language models (LLMs) have shown 2

Dynamic language model representations for multi-objective reaction optimisation

strong performance across chemical tasks [39–47] and demonstrated that their learned embeddings can serve as dense, informative representations for predictive modelling [48–51]. Recent work has further shown that LLM representations can be adapted to specific optimisation tasks through joint fine-tuning with Gaussian process surrogates, demonstrated for single-objective, sequential optimisation [52]. Crucially, LLMs can encode arbitrarily complex chemical information into fixed-length representations directly from text, offering a naturally unified featurisation method across component classes. These important advances notwithstanding, a general framework for multi-objective reaction optimisation, applicable across chemically heterogeneous reaction components and validated across experimental regimes from low-data to highthroughput settings, has yet to be realised.

design spaces. Neither required descriptor computation or feature engineering, indicating that effective multi-objective optimisation is achievable in reaction spaces where constructing a shared descriptor representation would itself be a substantial undertaking.

2. Computational results We benchmarked the LLM-GP framework against established and readily available featurisation methods in reaction optimisation: DFT descriptors from the Kraken [25] and COSMO RS [26] descriptor libraries, and one-hot encoding baselines. Descriptor coverage is uneven across component classes: Kraken supplies descriptors for the monophosphine ligands and COSMO-RS for the solvents, but no comparable library exists for the bases and precursors, which were therefore represented by one-hot encoding within the descriptor baseline. All three representations were embedded in the same multi-objective Bayesian optimisation loop, using the same acquisition function (qLogNParEGO) and an identical experimental budget (see Methods).

Here, we introduce Alice, a multi-objective chemical reaction optimisation framework that uses trainable LLM embeddings as a unified reaction representation (Figure 1b). Textual descriptions of reaction conditions are encoded by a low-rank adaptation (LoRA)-fine-tuned language model and passed through reaction objective-specific projection heads into independent Gaussian process surrogates, with the architecture trained end-to-end via joint marginal loglikelihood maximisation. This jointly trained architecture, hereafter referred to as the LLM-GP framework, produces dynamic embeddings that are refined at each optimisation iteration, adapting the reaction representation to the optimisation task at hand. We first benchmark the LLM-GP framework retrospectively against established descriptor libraries and one-hot encoding baselines across sequential low-data optimisation of nickel- and palladiumcatalysed cross-couplings and large-batch 96-well plate high-throughput experimentation campaigns, converging on high-performing conditions in fewer experiments across all settings.

A set of reaction conditions is Pareto optimal when no other condition improves one objective without degrading another; these points define the Pareto front (Figure 2a). Optimisation performance was measured by hypervolume, the volume of objective space enclosed by the Pareto-optimal conditions identified so far. A larger hypervolume means the conditions found are both closer to the best achievable trade-offs and better spread across them, and we report it as a percentage of the value attained by the true Pareto front of each dataset. We compared methods by the mean number of iterations required to reach 90% and 95% of the maximum hypervolume, with standard error across 20 random seeds, as these thresholds represent practically relevant levels of convergence. For each dataset the optimisation budget was set by the number of iterations needed for at least one method to reach 95% of the maximum hypervolume on average. We additionally report the percentage of optimisation runs reaching each hypervolume threshold within the allocated experimental budget as a measure of robustness; these results are provided in Extended Data Section A. Among the pre-trained language models and pooling strategies evaluated, T5-base [53] with mean pooling consistently yielded the best optimisation performance and was used throughout (full ablation studies across model architectures and pooling strategies are provided in Extended Data Section B).

We then apply the framework prospectively to two wet-lab campaigns executed on an automated high-throughput platform, each chosen for a form of chemical heterogeneity for which descriptor libraries are not readily available. In a palladium-catalysed cyanation, a single representation spans mono- and bidentate phosphine catalysts alongside additives ranging from elemental zinc to organic and inorganic bases; within this space the framework identified conditions using a lower-hazard cyanide source that gave >99% conversion and >99% selectivity, and 94% isolated yield on gram scale. In an asymmetric dynamic kinetic hydrogenation, the framework balanced conversion, diastereomeric excess (de), and enantiomeric excess (ee) across chiral catalysts spanning different metals, donor sets, and coordination geometries, delivering the target syn alcohol in 84% isolated yield and 99.6% ee on scale-up. Both campaigns used only two rounds of 96 experiments, sampling under 3% of their respective

2.1. Sequential low-data reaction optimisation We first evaluated our approach in sequential low-data optimisation regimes, where the primary goal is sample efficiency: minimising the number of experiments required to identify high-performing conditions (Figure 2b). We used two open-source multi-objective reaction datasets, each 3

Dynamic language model representations for multi-objective reaction optimisation a

Metrics for benchmarking multi-objective optimisation algorithms Quantifying multi-objective optimisation performance

Measuring convergence to practical performance thresholds

Pareto optimal points enclose the hypervolume

95%

Hypervolume [%]

Reaction objective 2

90%

Dominated solution space

Iterations required to reach 95% convergence

Iterations of optimisation

Reaction objective 1

Benchmarking sequential low-data reaction optimisation

Iterations to reach threshold

b

Varied reaction parameters Ligand, Ni-precursor, Base Solvent, Co-solvent, Temperature

40

90% Hypervolume threshold

95% Hypervolume threshold

90% Hypervolume threshold

95% Hypervolume threshold

30 20 10 0

384 experiments to optimise over in a chemical space of 88,000 combinations

Iterations to reach threshold

100

Varied reaction parameters Ligand, Pd-precursor, Base Solvent, Co-solvent

80 60 40 20 0

404 experiments to optimise over in a chemical space of 59,040 combinations

LLM-GP

Benchmarking 96-well plate HTE reaction optimisation

4 Iterations to reach threshold

c

Varied reaction parameters Ligand, Pd-precursor, Base Solvent

DFT Descriptor databases

90% Hypervolume threshold

One-hot encoding

95% Hypervolume threshold

3 2 1 0

Virtual benchmark of a 21,648 experiment chemical space to optimise over

Figure 2. Benchmarking multi-objective reaction optimisation across experimental regimes. a, Metrics for evaluating multi-objective optimisation. Left: the hypervolume indicator quantifies the volume of objective space dominated by the current Pareto-optimal set, providing a scalar measure of multi-objective performance. Right: convergence is assessed by tracking the normalised hypervolume over optimisation iterations; the number of iterations required to reach practical performance thresholds (90% and 95% of maximum hypervolume) serves as a comparative metric across featurisation methods. The percentage of optimisation runs (initialised with different random seeds) reaching each threshold is reported in Extended Data Section A. b, Sequential low-data reaction optimisation benchmarks. Two multi-objective case studies from published reaction datasets are evaluated: a nickel-catalysed Suzuki coupling [20] and a palladiumcatalysed Suzuki coupling [21]. Iterations required to reach 90% and 95% hypervolume thresholds are compared for the LLM-GP framework, DFT descriptor libraries (Kraken [25], COSMO-RS [26]), and one-hot encoding, repeated across 20 random seeds (see Methods for more details). Runs not reaching a threshold within the experimental budget were assigned the maximum budget. c, 96-well plate HTE reaction optimisation benchmark. A palladium-catalysed sulfonamide coupling [21] (virtual benchmark of 21,648 experiments) from a published dataset is optimised in large parallel batches of 96 experiments.

comprising experimental yield and selectivity as the optimisation objectives and presenting large combinatorial search spaces with diverse categorical reaction components: a nickel-catalysed Suzuki coupling (384 experiments from a space of 88,000 combinations, varying ligand, Ni-precursor, base, solvent, co-solvent, and temperature) [20], and a

palladium-catalysed Suzuki coupling (404 experiments from 59,040 combinations, varying ligand, Pd-precursor, base, solvent, and co-solvent) [21]. Each campaign was initialised with 5 experiments selected using Sobol sampling to ensure diverse coverage of initial points across the search space (see Methods), after which the Bayesian optimisation loop 4

Dynamic language model representations for multi-objective reaction optimisation

selected one candidate per iteration. Iteration counts below refer to these optimisation iterations and exclude the initial experiments. On the nickel-catalysed Suzuki coupling, the LLM-GP framework reached the 90% hypervolume threshold in approximately 13 optimisation iterations on average, roughly half the number required by both the DFT descriptor-based (∼26 iterations) and one-hot encoding (∼24 iterations) baselines (Figure 2b). At the more stringent 95% threshold, the gap narrowed, though the framework retained a consistent advantage, converging in approximately 24 iterations compared to ∼28 for DFT descriptors and ∼30 for one-hot encoding. On the palladium-catalysed Suzuki coupling benchmark, the framework achieved both the 90% and 95% thresholds in approximately 65 iterations on average, while the descriptor-based and one-hot encoding baselines required approximately 80 iterations each to reach the same performance levels (Figure 2b). Across both case studies, the LLM-GP framework consistently reached practical convergence thresholds in fewer experiments than either baseline, improving sample efficiency. Over 20 independent runs initialised with different random seeds, the framework also achieved a higher proportion of successful optimisations reaching both convergence thresholds within the allocated experimental budget, indicating more robust convergence (see Extended Data Section A).

mately one week of experimental and analytical time, this reduction translates directly into meaningful savings in time and resources. These results demonstrate that our approach offers consistent advantages across both sequential low-data regimes, where sample efficiency is crucial, and parallel large-batch settings increasingly adopted in pharmaceutical process development and academic high-throughput screening [55–57].

Building on our retrospective benchmarks across conventional Buchwald–Hartwig and Suzuki–Miyaura datasets, we next sought to apply our framework prospectively to wet-lab reaction optimisation campaigns, moving beyond standard reaction spaces. Here, we targeted transformations containing chemically heterogeneous components that traditional reaction representations struggle to capture. In each campaign, we used our approach to identify high-performing conditions within the search space, iteratively generating ML predictions and conducting the suggested experiments with our HTE robotic platform (see Supplementary Information Section 1 for more details). All experimental data and characterisation of any isolated products are reported in the Supplementary Information.

2.2. Highly parallel HTE reaction optimisation

3.1. Case study 1: Palladium-catalysed cyanation

Modern high-throughput experimentation (HTE) platforms have combined parallel screening with data-driven optimisation [20, 54], enabling the exploration of large reaction spaces in compressed experimental timescales. We next sought to evaluate our approach in this batched optimisation setting, where reducing the number of HTE plate iterations directly translates to savings in experimental time. We used an open-source palladium-catalysed sulfonamide coupling dataset (a Buchwald–Hartwig-type C–N coupling with 21,648 virtual experiments, varying ligand, Pd-precursor, base, and solvent) [21] as a benchmark, optimising yield and selectivity in parallel batches of 96 experiments to simulate 96-well plate HTE campaigns (Figure 2c). Each campaign was initialised with a Sobol-sampled plate of 96 experiments, after which each iteration selected 96 conditions at once, so one optimisation iteration corresponds to one physical plate. The LLM-GP framework reached practical convergence thresholds in approximately 1.2 plate optimisation iterations beyond initialisation, whereas both the DFT descriptor-based and one-hot encoding baselines required approximately double (∼2.4) to achieve the same thresholds (Figure 2c). The framework also achieved a higher proportion of successful optimisation runs reaching both convergence thresholds, and results for 24- and 48-well plate batch sizes show the same trend (Extended Data Section A). Given that each 96-well plate campaign typically requires approxi-

Aryl nitriles are prevalent across pharmaceuticals, agrochemicals, and functional materials [58, 59], serving as compact, metabolically stable hydrogen-bond acceptors and hydroxyl and carboxyl isosteres [60]. Among synthetic routes accessing these motifs, the transition-metal-catalysed cyanation of aryl halides remains one of the most widely adopted and functionally tolerant strategies [61]. While substantial methodology development has produced a broad repertoire of candidate catalysts, additives, and cyanide sources, this very diversity introduces three intersecting sources of chemical heterogeneity that challenge traditional modelling approaches. First, both monodentate and bidentate phosphine ligands are competitive across cyanation substrate classes, yet descriptor libraries developed for monophosphines are not directly transferable to bisphosphines (and vice versa) thus fragmenting the feature space. Second, productive cyanation conditions often require additives spanning chemically disparate classes such as elemental reductants (e.g., Zn), organic bases (e.g., NEt3 ), and inorganic salts (e.g., KOAc), which resist straightforward unification within standard chemical feature spaces. Finally, the choice of cyanide source carries practical implications beyond reactivity. Classical sources such as NaCN and Zn(CN)2 are highly toxic, whereas safer but less reactive alternatives such as K4 [Fe(CN)6 ] are increasingly preferred in industrial process settings despite their poor organic solubility, which

3. Prospective experimental case studies

5

Dynamic language model representations for multi-objective reaction optimisation b

Cyanation reaction condition optimisation

O

Catalyst Cyanide source, additive

O

F

Solvent, co-solvent Temperature

Br

Conversion (%)

O

CN

1

Count

F

Objectives

Reaction parameters

O

P P Pd e.g., EPhos Pd G3

Pd

e.g., [Pd(DPPF)Cl2]

...

Most toxic

N

Preference for more benign CN sources

100

80

N

Diverse additives (5) Zn

NaCN

K2CO3 NEt3

...

Zn(CN)2

Pd Least toxic

Diverse ligand and catalyst motifs

60

40

20

▪ ▪ ▪

e.g., Pd PEPPSI IPent

Plate 2 (Optimisation)

10 0

Reaction search space of 30,000 conditions Cyanide sources (3)

Plate 1 (Initialisation)

20

Selectivity (%)

2

Mono- and bidentate Pd catalysts (40)

P

Reaction conditions identified for lower-hazard cyanide source

Selectivity [%]

a

Solvents (10) Co-solvents (2) Temperatures (3)

0 0

Unifying complex chemical spaces without feature engineering

Reaction conditions:

c

Reaction condition: Catalyst: [Pd(allyl) (Xantphos)]Cl Cyanide source: NaCN Additive: Zn Solvent: DMC Co-solvent: None Temperature: 75

Textual description of reaction conditions

20

40

60

Conversion [%]

80

100 0

Scale-up of best conditions identified from optimisation

1

[Pd(Xantphos)(allyl)]Cl (1 mol%) K4[Fe(CN)6], DMC/H2O (2:1), 75oC

(5.0 g scale)

Our LLM-GP optimisation framework

Predictions executed on robotic platform

20

Count

2 94% isolated yield

milligram-scale

gram-scale

Lower-hazard, greener conditions identified in <1% search space

Figure 3. Case study 1: Palladium-catalysed cyanation. a, Search space for the cyanation of 1 to 2. Excluding reaction conditions with temperatures above the solvent boiling point results in 30,000 conditions (full search space in Extended Data Figure 7). Conversion (%) and selectivity (%) were maximised. Each reaction condition was encoded as a textual description, embedded by the joint LLM-GP framework, and model selected conditions were executed on an automated high-throughput experimentation (HTE) platform in 96-well HTE plates (see Supplementary Information Section 2.1). b, Reaction objective distributions for HTE optimisation rounds, Plate 1 (Initialisation) and Plate 2 (Optimisation), with marginal histograms. Only conditions using K4 [Fe(CN)6 ], the least toxic cyanide source, are shown on the plot. The step line traces the Pareto front among these conditions after Plate 1. c, Scale-up of the best conditions identified by optimisation, delivering 2 in 94% isolated yield at gram scale (see Supplementary Information Section 2.2).

necessitates biphasic reaction media [62, 63]. Together, these intersecting dimensions of heterogeneity in cyanation reactions pose a distinct challenge for conventional descriptor frameworks, highlighting the value of more flexible, unified representation strategies.

ecution without manual descriptor curation. The full search space is provided in Extended Data Figure 7. Optimisation was initialised with a Sobol-sampled first plate of 96 experiments to probe the global design space. While the first round identified several high-performing conditions, the top hits relied exclusively on toxic cyanide sources (NaCN and Zn(CN)2 ), while K4 [Fe(CN)6 ] yielded only mediocre outcomes with either poor conversion or selectivity (Figure 3b). To test whether high-performing conditions could be identified using a lower-hazard cyanide source, we fixed the cyanide source for round 2 to K4 [Fe(CN)6 ] and re-optimised within this constrained domain. The second optimisation round successfully met this target, identifying multiple conditions achieving >99% conversion and >99% selectivity (Figure 3b). Notably, high-performing hits spanned both monodentate ([Pd(tBuXPhos)(allyl)]OTf) and bidentate ([Pd(Xantphos)(allyl)]Cl) ligands, exemplifying the benefit of jointly searching over both ligand families within a unified representation. To confirm that these milligram-scale HTE hits translate to preparative scale-up, we scaled the lead conditions ([Pd(Xantphos)(allyl)]Cl with K4 [Fe(CN)6 ] · 3 H2 O in DMC/H2 O) to gram scale (Figure 3c). The reaction proceeded cleanly, delivering the target

We assembled a design space for the cyanation of methyl 3-bromo-5-fluorobenzoate, comprising 40 commercially available palladium catalysts spanning both mono- and bidentate ligand classes (e.g., [Pd(tBuXPhos)(allyl)]OTf, [Pd(DPPF)Cl2 ]), three cyanide sources of varying toxicity (NaCN: median lethal dose (mg/kg) LD50 = 3.6, Zn(CN)2 : LD50 = 54, K4 [Fe(CN)6 ]: LD50 = 3613), five additive options (Zn, KOAc, NEt3 , K2 CO3 , or none), ten solvents, two co-solvent options, and three temperatures (Figure 3a). Conditions in which the reaction temperature exceeded the solvent boiling point were removed, giving a final design space of 30,000 conditions. While constructing a bespoke descriptor representation across these heterogeneous components is in principle feasible, doing so would require laborious, trial-and-error feature engineering with no guarantee of success; our approach bypasses this bottleneck by providing a unified, out-of-the-box featurisation directly from textual descriptions, enabling immediate campaign ex6

Dynamic language model representations for multi-objective reaction optimisation a

Asymmetric hydrogenation reaction condition search space Target syn-pair

OH Reaction parameters

O

Catalyst base (mol%)

O Ph

Ph

Ph

Ph

3a

Objectives

Ph

Ph

Conversion (%)

4a

Solvent Temperature

3b

OH Ph

Ph

Diastereomeric excess (syn) (%)

4b

OH

OH Ph

Ph

Enantiomeric excess (syn) (%)

Ph

Ph

Base-catalysed epimerisation 4c

4d

Reaction search space of 8,064 conditions Bases (9)

Chiral iridium and ruthenium catalysts (32)

P * N

Ir

...

N

N N

*

Ru

P P

Ru

N N

K2CO3

*

...

*

e.g., (R,R)-Ts-DPEN

P

O

KOH

...

Na2CO3

Ru

N

e.g., (R)-DM-BINAP

KO

Challenging to featurise coordination environments

b

Solvents (7) OH

N

NaOH

e.g., (R)-DM-PPhos, (R,R)-DPEN

e.g., (R)-DTB-SpiroPAP-3-Me *

P

▪ ▪

N

...

c

Reaction objectives across iterations

HO

Base mol% (2) Temperatures (2)

Scale-up of best conditions identified from optimisation

Plate 1 (Initialisation)

P(DTB)2 Cl H Ir H

Plate 2 (Optimisation) 100

HN

N

Sole active ligand scaffold for syn selectivity

40 20

Conversion [%]

80 60

[IrClH2(S)-DTB-SpiroPAP-3-Me] (1 mol%) 3

tAmOH, DBU (50 mol%), H2 (20 bar), 60oC

(2.0 g scale)

4a 84% isolated yield 99.6% ee

0 100 50 100

50

0 0

de (syn) [%]

50

50 100

ee (syn) [%]

milligram-scale

gram-scale

100

Figure 4. Case study 2: Asymmetric ketone hydrogenation. a, Search space for the dynamic kinetic hydrogenation of 3 to the target syn-(1S, 2R) alcohol 4a, one of four accessible stereoisomers (4a–4d) arising from base-catalysed epimerisation between 3a and 3b. Conversion (%), de (syn) (%), and ee (syn) (%) were maximised simultaneously across 32 chiral iridium and ruthenium catalysts, 9 bases, 7 solvents, 2 base loadings and 2 temperatures, giving 8,064 conditions (full search space in Extended Data Figure 8). b, Reaction objectives across HTE optimisation rounds, Plate 1 (Initialisation) and Plate 2 (Optimisation). Positive de and ee denote excess of the target syn diastereomer and (1S, 2R) enantiomer respectively. Negative values denote the corresponding anti diastereomer or (1R, 2S) enantiomer, respectively. Syn-selective outcomes increased from 5 of 96 conditions in Plate 1 to 59 of 96 in Plate 2. c, Scale-up of the best conditions identified by optimisation, delivering 4a in 84% isolated yield and 99.6% ee at gram scale (see Supplementary Information Section 3.2). [IrClH2 ((R/S) – DTB – SpiroPAP – 3-Me)] was the only catalyst pair in the campaign selective for the target syn isomer.

ticular, the dynamic kinetic hydrogenation of α-substituted aryl ketones, in which two contiguous stereocentres are set in a single step to yield up to four potential stereoisomers [65, 66], exemplifies a class of multi-objective optimisation problems that require balancing trade-offs across a Pareto front defined by conversion, diastereomeric excess (de), and enantiomeric excess (ee). Productive catalysts include chiral iridium and ruthenium complexes drawn from a structurally diverse pool of ligand motifs, spanning bidentate phosphines, mixed phosphine–nitrogen donors (P, N and P, N, N), and a variety of associated ancillary ligands and counter-ions. Representing this coordination diversity in a unified descriptor framework is non-trivial; physical featurisation schemes engineered to quantify the

aryl nitrile with 94% isolated yield (see Supplementary Information Section 2). This scalable, low-hazard protocol was identified in just two iterative campaign rounds (one week per round) totalling 192 experiments, evaluating less than 1% of the 30,000-condition design space. 3.2. Case study 2: Asymmetric ketone hydrogenation Having established the approach on reaction systems with multi-component heterogeneity, we next sought to evaluate its capacity to navigate stereoselective transformations. Catalytic asymmetric hydrogenation of prochiral ketones is among the most widely deployed enantioselective methods in pharmaceutical and agrochemical synthesis [64]. In par7

Dynamic language model representations for multi-objective reaction optimisation

(50 mol%) in tAmOH at 60 ◦ C, delivered the target syn alcohol with the desired enantiomer in 84% isolated yield with 99.6% enantiomeric excess (see Supplementary Information Section 3). This scalable protocol was identified in two iterative campaign rounds (one week per round) using 192 experiments out of 8,064, corresponding to approximately 2.4% of the design space.

3D chiral pocket of one ligand geometry or metal centre fail to generalise across fundamentally disparate coordination spheres. Jointly balancing three reaction objectives across this heterogeneous chiral space compounds the modelling challenge. We assembled a design space of 8,064 conditions for the asymmetric dynamic kinetic hydrogenation of 1,2-diphenylpropan-1-one towards the syn-(1S, 2R)-1,2diphenylpropan-1-ol diastereomer, comprising 32 chiral iridium and ruthenium catalysts, 9 organic and inorganic bases, 7 solvents, varied base loading (mol%), and temperature (Figure 4a). The full search space is provided in Extended Data Figure 8.

4. Outlook This work demonstrates that dynamic language model representations provide a general route to multi-objective reaction optimisation, applicable across chemically diverse reaction systems. By encoding reaction components directly from textual descriptions, the framework searches jointly over chemical entities of different classes (e.g., ligand families, additives, and catalysts) within a single representation, without descriptors being selected or computed for each new system. Applied prospectively to two wet-lab campaigns, it identified conditions that translated directly to preparative scale within two rounds of high-throughput experimentation. Engineered descriptors are chemically interpretable, but their selection presupposes an understanding of which molecular properties govern reactivity in the system at hand. Learning the representation dynamically from text removes this requirement, allowing optimisation to begin before such understanding is established, where mechanistic insight could emerge from the conditions identified rather than being needed to find them. By removing expert-guided featurisation as a prerequisite, we anticipate this approach will extend model-guided optimisation to a wider range of chemistry, including biocatalysis and enzymatic reactions, and ultimately enable more efficient chemical processes.

Optimisation was initialised with a Sobol-sampled first HTE plate of 96 experiments spanning the full design space (Figure 4b). The first round of experiments revealed that the undesired anti diastereomer was favoured across the majority of conditions evaluated, yielding multiple hits with high anti selectivity (−99% de (syn)). In contrast, accessing the target syn diastereomer was more challenging: only 5 out of 96 experiments were syn selective, and out of the 32 chiral catalysts tested, only a single catalyst pair, [IrClH2 ((R/S) – DTB – SpiroPAP – 3-Me)], demonstrated activity towards the syn isomer. The top hit for the desired isomer in Plate 1 achieved quantitative conversion but moderate stereoselectivity (73.5% de (syn) and 86.7% ee (syn)) (Figure 4b). In achiral media, enantiomeric catalyst pairs yield mirror-image products, so each training observation was augmented with its reflected counterpart under a sign inversion of the enantiomeric excess (see Supplementary Information Section 3.3). Guided by the updated LLM representations, the framework refocused optimisation towards syn-producing regions, increasing syn-selective outcomes from 5/96 to 59/96 conditions in Plate 2 (Figure 4b). The Pareto front was expanded towards higher stereoselectivity, with the most stereoselective conditions identified reaching 87.1% de (syn) and 93.2% ee (syn). The most balanced lead hit achieved >99% conversion, 80.5% de (syn), and 89.7% ee (syn). Identifying the productive catalyst family did not by itself resolve the optimisation problem: across the reaction conditions employing [IrClH2 ((R/S) – DTB – SpiroPAP – 3-Me)], diastereoselectivity spanned -67% to +87.1% de (syn), and in 12 instances its sign inverted upon changing temperature or base loading alone, consistent with the dynamic kinetic nature of the transformation. Systematically removing the highest-performing conditions from the Plate 1 training data did not prevent the model from recovering high-performing syn-selective conditions, indicating that the outcome did not depend on a small number of fortunate initial hits (see Supplementary Information Section 3.4). Translating the lead conditions identified from optimisation to preparative gram scale proceeded smoothly (Figure 4c): the lead conditions, [IrClH2 ((S) – DTB – SpiroPAP – 3-Me)] with DBU

Methods Metrics for multi-objective optimisation problems We consider the simultaneous maximisation of M objective functions ym : X → R, m = 1, . . . , M (e.g., reaction yield and selectivity), over a finite chemical design space X = {x1 , . . . , xN }, where each xi is a vector of reaction conditions. The aim is to identify the Pareto optimal set, also referred to as the Pareto front. We say that x′ dominates x (written x′ ≻ x) if ym (x′ ) ≥ ym (x) for all m = 1, . . . , M and yj (x′ ) > yj (x) for some j ∈ {1, . . . , M }. The Pareto optimal set is then defined as:  P∗ = x ∈ X

∄ x′ ∈ X : x′ ≻ x

(1)

We seek to approximate P ∗ using as few experimental iterations or evaluations as possible with multi-objective Bayesian optimisation. The quality of a candidate Pareto front P ⊆ RM , obtained from the current set of observa8

Dynamic language model representations for multi-objective reaction optimisation

tions, is measured by its dominated hypervolume (Figure 2a) relative to a reference point r ∈ RM , chosen such that rm < pm for all p ∈ P and m = 1, . . . , M :

pairs, for example: Reaction condition: ligand: {ligand name} solvent: {solvent name} precursor: {precursor name} base: {base name}

 HV(P, r) = Vol y ∈ RM | ∃ p ∈ P : rm ≤ ym ≤ pm ∀ m



(2)

with one pair per parameter varied in the design space. A pre-trained language model hϕ maps the tokenised prompt to a sequence of embedding vectors Hi = hϕ (xi ) ∈ RL×demb , where L is the sequence length and demb is the model’s hidden dimension. A pooling operator P OOL : RL×demb → Rdemb aggregates the token-level representations into a fixed-dimensional embedding ei = P OOL(Hi ) ∈ Rdemb to produce a sequence-level representation. We adopted pooling strategies based on each model’s architecture. For encoder-decoder models, only the encoder stack was used, and pooling was applied over its hidden states. The T5 encoder-decoder models (T5-base [53], T5-small, and T5-chem [72]) used mean pooling, which averages hidden states over non-padded tokens. We used last-token pooling for the decoder-only model Qwen2.5 [73], which takes the hidden state at the last non-padding position, as causal attention ensures that only this token has attended to the full input sequence [74]. Mean pooling was also evaluated for Qwen2.5. For the BART-base encoder-decoder model [75], we used mean pooling following the T5 models, and additionally evaluated CLS pooling as BART inherits a <s> classification token. All models were accessed via the Hugging Face Transformers library [76]. Ablation studies over pre-trained language model architectures and pooling strategies across all benchmark datasets are presented in Extended Data Section B.

where Vol(·) denotes the Lebesgue measure. For all benchmark datasets, the optimisation objectives are percentages bounded on [0, 100], and we set r = (0, 0). The same reference point was used for all featurisation methods. Hypervolume is a widely used quality indicator for multi-objective optimisation, rewarding both convergence to the true Pareto front and diversity along it [13, 17, 67, 68]. Importantly, it is monotonic and Pareto-compliant [69, 70], meaning that if one candidate Pareto front strictly dominates another, it will achieve a strictly higher hypervolume. In this work, we report the normalised hypervolume percentage, defined as HV(P, r)/HV(P ∗ , r) × 100%, to evaluate the quality of candidate Pareto fronts P identified by optimisation algorithms relative to the true Pareto front P ∗ . Representing reaction conditions We benchmarked three featurisation approaches for representing reaction conditions. One-hot encoding. As a universally applicable and inexpensive baseline that has shown competitive performance in prior work [12, 13, 38], each unique categorical variable (e.g., XPhos, dioxane, PhMe) was represented as a binary indicator vector, and all component vectors were concatenated to form the input representation (Figure 1a).

Overview of optimisation loop

Molecular descriptors. Widely adopted as the primary featurisation strategy in reaction optimisation, molecular descriptors from established descriptor libraries were included as a chemically informative baseline [15, 16, 20, 21, 27–30]. Monophosphine ligands were represented using 190 DFT descriptors from the Kraken library [25], with Principal Component Analysis (PCA) [71] applied to retain components explaining 99% of the variance, reducing dimensionality while preserving information. Solvents were parameterised using four COSMOtherm-derived DFT descriptors from the COSMO-RS database [26]. The remaining categorical variables (e.g., bases and precursors), for which comparable descriptor libraries are not readily available, were represented using one-hot encoding.

We initialised our Bayesian optimisation workflows using low-discrepancy Sobol sequences [77] to provide broad initial coverage of the design space, which were evaluated to form the initial training data. At each subsequent iteration of the optimisation loop, the surrogate model was fitted to all currently observed data, a batch of candidates was selected by optimising the acquisition function over the remaining design space, and the selected experiments were evaluated and added to the training set. We used the qLogNParEGO acquisition function [68, 78, 79], a log acquisition function that conditions on noisy baseline observations and applies Chebyshev scalarisation to reduce the multi-objective problem into single-objective subproblems. For batch acquisition, candidates were selected in a greedy sequential fashion, with each selection conditioned on previously chosen pending points.

Large language model representations. Natural language provides a flexible modality for encoding chemically heterogeneous reaction components without requiring componentclass-specific featurisation. Each reaction condition xi ∈ X was represented as a textual prompt of component-value

For surrogate models using fixed input representations (onehot encoding and molecular descriptors), we used an opti9

Dynamic language model representations for multi-objective reaction optimisation

where the marginal log-likelihood for objective m is:  1 ⊤ −1 (m) log p(ym | Z , θm ) = − ỹm Km ỹm 2  + log |Km | + n log 2π

mised Gaussian process (GP) configuration adapted from EDBO+ [13] and Minerva [20] following standard marginal likelihood optimisation. For our approach using an LLMbased deep kernel surrogate with learned input representations, the GP hyperparameters and language model parameters were optimised jointly as described below. All computational experiments were repeated over 20 random seeds to report statistical variation.

(6)

where n is the number of observations and ỹm = ym −µm 1 denotes the targets centred by the constant mean function. GP targets are standardised to zero mean and unit variance for training. Gradients from all M objectives are backpropagated through the kernel and projection heads to the shared LLM parameters ϕ, where each projection head gϕm receives gradients only from its corresponding objective. This enables the LLM to learn representations that are jointly optimised for GP predictive performance across all objectives, while each projection head specialises for its target objective.

LLM-based deep kernel Gaussian process surrogate The LLM first maps tokenised reaction condition prompts to pooled embeddings ei ∈ Rdemb . For each objective m = 1, . . . , M , a trainable projection head gϕm transforms the language model embedding ei to an objective-specific (m) representation zi = gϕm (ei ). A Gaussian process (GP) then models the mapping from each projected representation to its corresponding objective value:

Parameter-efficient fine-tuning of language models (m)

ym (xi ) = fm (zi

The language model embeddings of reaction conditions evolve across BO iterations as the LLM is fine-tuned on accumulating observations. To do so efficiently, we apply parameter-efficient fine-tuning (PEFT) with Low-Rank Adaptation (LoRA) [24]. Rather than updating all of the pre-trained language model’s parameters, we use LoRA to update only a small subset. For a pre-trained weight matrix W0 ∈ Rd×k , LoRA injects trainable low-rank decompositions such that the weight update takes the form:

) + ϵm ,

fm ∼ GP(µm , kθm ),

(3)

2 ϵm ∼ N (0, σm )

where fm is the latent function for objective m, modelled as a Gaussian process with constant mean function µm and kernel function kθm , and ϵm is Gaussian observation noise 2 with variance σm . We use the Matérn-5/2 kernel: √ ′

2 kθm (z, z ) = σf,m

1+

5ρ

ℓm

5ρ2 + 2 3ℓm

√

! exp −

5ρ

ℓm

W = W0 + ∆W, (7) α ∆W = BA, with B ∈ Rd×r , A ∈ Rr×k r where r ≪ min(d, k) is the rank and α is a scaling hyperparameter. We use LoRA to target a subset of linear projection matrices within the pre-trained language models, using rank r = 4 and α = 16.

! ,

where ρ = ∥z − z′ ∥2 (4)

Computational implementation 2 where σf,m is the signal variance, and ℓm is the lengthscale. Together with the constant mean function µm and 2 noise variance σm , these constitute the GP hyperparameters 2 2 θm = {µm , σf,m , ℓm , σm }. The kernel matrix for objec2 tive m is given by Km = kθm (Z(m) , Z(m) ) + σm I, where Z(m) collects the projected embeddings of all observed reaction conditions. The M GPs are queried jointly by the multi-objective acquisition function. The language model hϕ , projection heads {gϕm }M m=1 , and GP hyperparameters {θm }M are jointly optimised by minimising the sum of m=1 negative marginal log-likelihoods:

L=−

M X

log p(ym | Z(m) , θm ),

All models, Bayesian optimisation workflows, and computational analyses were implemented in Python 3.10. We used PyTorch [80] (v2.8.0), GPyTorch [81] (v1.14), and BoTorch [82] (v0.13.0) to construct our multi-objective Bayesian optimisation algorithms. Pretrained language models were loaded using Hugging Face transformers [76] (v4.51.2), with Low-Rank Adaptation (LoRA) implemented using peft (v0.15.1) for parameter-efficient fine-tuning. We used PyTorch Lightning [83] (v2.5.4) for reproducible random seed management and workflow control. Scikit-learn [84] (v1.7.1) and NumPy [85] (v1.26.4) were utilised for machine learning pipelines and numerical data processing, including Principal Component Analysis (PCA) for descriptor dimensionality reduction. Experiment tracking

(5)

m=1

10

Dynamic language model representations for multi-objective reaction optimisation

and logging were conducted using Weights & Biases (wandb, v0.21.0) [86]. Matplotlib [87] (v3.10.1) and seaborn [88] (v0.13.2) were used for plotting and visualising all computational results presented in this work. All computations were run on the Roche HPC cluster, using NVIDIA A100 Tensor Core and Blackwell GPUs.

M.D.R.L., K.P., M.J.N., R.B. and P.S. S.L.C., A.A., R.P.B., K.P., M.J.N., R.B. and P.S. provided supervision. J.W.S. wrote the paper with input from all authors.

Competing interests J.W.S., D.M.S., M.D.R.L., A.A., R.P.B., K.P., M.J.N. and R.B. declare potential financial and non-financial conflicts of interest as full employees of F. Hoffmann-La Roche Ltd. The other authors declare no competing interests.

Experimental procedures High-throughput experimentation procedures, analytical methods, and scale-up syntheses and characterisation for the palladium-catalysed cyanation and asymmetric ketone hydrogenation reactions are provided in the Supplementary Information.

Data availability All benchmark datasets, reaction condition search spaces, and high-throughput experimentation (HTE) experimental data generated in this study are included in the manuscript, Supplementary Information, and on the accompanying public GitHub repository.

Code availability The custom code developed for this work is implemented in Python and is made available in a public GitHub repository under the Apache 2.0 licence: https://github.com/schwallergroup/alice

Acknowledgements J.W.S., R.P.B., K.P. and R.B. thank Roche and its Technology Innovation and Science (TIS) initiative for financial support. We thank the Global Internship Program in Innovation & Sustainability (IP2TIS) 2025 (https://careers. roche.com/global/en/ip2tis-program) for financial support of this project. D.M.S., B.R., and P.S. acknowledge support from NCCR Catalysis (grant no. 225147), a National Centre of Competence in Research funded by the Swiss National Science Foundation. D.M.S. was funded by the Swiss National Science Foundation (SNSF) [226509].

Author contributions J.W.S., B.R., R.B. and P.S. conceived the project. J.W.S., D.M.S. and B.R. developed the computational workflow with input from S.L.C., A.A., R.P.B. and P.S. J.W.S. and D.M.S. performed the computations and trained the models. J.W.S., M.J.N. and R.B. carried out the experimental work. J.W.S., K.P., M.J.N. and R.B. designed and planned the experimental case studies, with input from M.D.R.L. J.W.S., D.M.S. and B.R. analysed the results with help from 11

Dynamic language model representations for multi-objective reaction optimisation

References

[8] Hui Zhao, Anne K. Ravn, Michael C. Haibach, Keary M. Engle, and Carin C. C. Johansson Seechurn. Diversification of pharmaceutical manufacturing processes: Taking the plunge into the non-pgm catalyst pool. ACS Catalysis, 14(13):9708–9733, 2024. ISSN 2155-5435. doi: 10.1021/ acscatal.4c01809. URL http://dx.doi.org/ 10.1021/acscatal.4c01809.

[1] Connor J. Taylor, Alexander Pomberger, Kobi C. Felton, Rachel Grainger, Magda Barecka, Thomas W. Chamberlain, Richard A. Bourne, Christopher N. Johnson, and Alexei A. Lapkin. A Brief Introduction to Chemical Reaction Optimization. Chemical Reviews, 123(6):3089–3126, 2023. ISSN 0009-2665, 1520-6890. doi: 10.1021/acs.chemrev. 2c00798. URL https://pubs.acs.org/doi/ 10.1021/acs.chemrev.2c00798. [2] Jonas Düker, Lukas Hebing, Samuel Leweke, Rachel L. Nicholls, Maximilian Lübbesmeyer, Giulio Volpin, Burkhard König, and Julius Hillenbrand. Data science-assisted workflow for reaction optimization in process chemistry. Organic Process Research & Development, 30(1):175–188, January 2026. ISSN 1520-586X. doi: 10.1021/acs. oprd.5c00384. URL http://dx.doi.org/10. 1021/acs.oprd.5c00384. [3] Matthew Ball, Dragos Horvath, Thierry Kogej, Mikhail Kabeshov, and Alexandre Varnek. Predicting reaction conditions: a data-driven perspective. Chemical Science, 16(38):17523–17541, 2025. ISSN 2041-6539. doi: 10.1039/d5sc03045e. URL http: //dx.doi.org/10.1039/D5SC03045E.

[9] Coby J. Clarke, Wei-Chien Tu, Oliver Levers, Andreas Bröhl, and Jason P. Hallett. Green and sustainable solvents in chemical processes. Chemical Reviews, 118(2):747–800, January 2018. ISSN 1520-6890. doi: 10.1021/acs.chemrev.7b00571. URL http://dx. doi.org/10.1021/acs.chemrev.7b00571. [10] Benjamin J. S. Rowsell, Harry M. O’Brien, Gayathri Athavan, Patrick R. Daley-Dee, Johannes Krieger, Emma Richards, Karl Heaton, Ian J. S. Fairlamb, and Robin B. Bedford. The iron-catalysed suzuki coupling of aryl chlorides. Nature Catalysis, 7(11): 1186–1198, October 2024. ISSN 2520-1158. doi: 10.1038/s41929-024-01234-0. URL http://dx. doi.org/10.1038/s41929-024-01234-0. [11] Connor P. Delaney, Eva Lin, Qinan Huang, Isaac F. Yu, Guodong Rao, Lizhi Tao, Ana Jed, Serena M. Fantasia, Kurt A. Püntener, R. David Britt, and John F. Hartwig. Cross-coupling by a noncanonical mechanism involving the addition of aryl halide to cu(ii). Science, 381(6662):1079–1085, 2023. ISSN 10959203. doi: 10.1126/science.adi9226. URL http:// dx.doi.org/10.1126/science.adi9226.

[4] David F. Nippa, Alexander J. Boddy, Kenneth Atz, Uwe Grether, Hayley Binch, and Rainer E. Martin. Accelerating compound synthesis in drug discovery: the role of digitalisation and automation. RSC Medicinal Chemistry, 16(12):5753–5764, 2025. ISSN 26328682. doi: 10.1039/d5md00672d. URL http: //dx.doi.org/10.1039/D5MD00672D.

[12] Benjamin J. Shields, Jason Stevens, Jun Li, Marvin Parasram, Farhan Damani, Jesus I. Martinez Alvarado, Jacob M. Janey, Ryan P. Adams, and Abigail G. Doyle. Bayesian reaction optimization as a tool for chemical synthesis. Nature, 590(7844):89–96, 2021. ISSN 0028-0836, 1476-4687. doi: 10.1038/s41586-021-03213-y. URL https://www.nature.com/articles/ s41586-021-03213-y.

[5] Alexandre A. Schoepfer, Jan Weinreich, Ruben Laplaza, Jerome Waser, and Clemence Corminboeuf. Cost-informed bayesian reaction optimization. Digital Discovery, 3(11):2289–2297, 2024. ISSN 2635098X. doi: 10.1039/d4dd00225c. URL http: //dx.doi.org/10.1039/D4DD00225C. [6] Tony Y. Zhang. Process Chemistry: The Science, Business, Logic, and Logistics. Chemical Reviews, 106(7):2583–2595, 2006. ISSN 0009-2665. doi: 10.1021/cr040677v. URL https://doi.org/10. 1021/cr040677v. [7] Michael F. Lipton and Anthony G. M. Barrett. Introduction: Process chemistry. Chemical Reviews, 106(7):2581–2582, 2006. ISSN 1520-6890. doi: 10.1021/cr068400d. URL http://dx.doi.org/ 10.1021/cr068400d.

[13] Jose Antonio Garrido Torres, Sii Hong Lau, Pranay Anchuri, Jason M. Stevens, Jose E. Tabora, Jun Li, Alina Borovika, Ryan P. Adams, and Abigail G. Doyle. A Multi-Objective Active Learning Platform and Web App for Reaction Optimization. Journal of the American Chemical Society, 144(43):19999– 20007, 2022. ISSN 0002-7863, 1520-5126. doi: 10.1021/jacs.2c08592. URL https://pubs.acs. org/doi/10.1021/jacs.2c08592. [14] Elena Braconi and Edouard Godineau. Bayesian optimization as a sustainable strategy for early-stage

12

Dynamic language model representations for multi-objective reaction optimisation

process development? a case study of cu-catalyzed c–n coupling of sterically hindered pyrazines. ACS Sustainable Chemistry & Engineering, 11(28): 10545–10554, 2023. ISSN 2168-0485. doi: 10.1021/ acssuschemeng.3c02455. URL http://dx.doi. org/10.1021/acssuschemeng.3c02455. [15] Natalie P. Romer, Daniel S. Min, Jason Y. Wang, Richard C. Walroth, Kyle A. Mack, Lauren E. Sirois, Francis Gosselin, Daniel Zell, Abigail G. Doyle, and Matthew S. Sigman. Data science guided multiobjective optimization of a stereoconvergent nickelcatalyzed reduction of enol tosylates to access trisubstituted alkenes. ACS Catalysis, 14(7):4699–4708, March 2024. ISSN 2155-5435. doi: 10.1021/ acscatal.4c00650. URL http://dx.doi.org/ 10.1021/acscatal.4c00650.

[20] Joshua W. Sin, Siu Lun Chau, Ryan P. Burwood, Kurt Püntener, Raphael Bigler, and Philippe Schwaller. Highly parallel optimisation of chemical reactions through automation and machine intelligence. Nature Communications, 16(1):6464, 2025. ISSN 2041-1723. doi: 10.1038/s41467-025-61803-0. URL https://www.nature.com/articles/ s41467-025-61803-0. [21] Rémi Schlama, Joshua W. Sin, Ryan P. Burwood, Kurt Püntener, Raphael Bigler, and Philippe Schwaller. Swarm intelligence for chemical reaction optimization. Chem, page 103035, April 2026. ISSN 2451-9294. doi: 10.1016/j.chempr.2026. 103035. URL http://dx.doi.org/10.1016/ j.chempr.2026.103035. [22] W. Kohn, A. D. Becke, and R. G. Parr. Density functional theory of electronic structure. The Journal of Physical Chemistry, 100(31):12974–12980, January 1996. ISSN 1541-5740. doi: 10.1021/jp960669l. URL http://dx.doi.org/10.1021/jp960669l.

[16] Jamie A. Cadge, Cedric Lozano, Morgan T. Merriman, Paul Oblad, Matthew S. Sigman, and Sarah E. Reisman. A data science-guided approach for the development of nickel-catalyzed homo-diels–alder reactions. Journal of the American Chemical Society, 147(34):31175–31186, August 2025. ISSN 1520-5126. doi: 10.1021/jacs.5c09948. URL http: //dx.doi.org/10.1021/jacs.5c09948. [17] Jiyizhe Zhang, Naoto Sugisawa, Kobi C. Felton, Shinichiro Fuse, and Alexei A. Lapkin. Multiobjective bayesian optimisation using q -noisy expected hypervolume improvement ( q nehvi) for the schotten–baumann reaction. Reaction Chemistry & Engineering, 9(3):706–712, 2024. ISSN 2058-9883. doi: 10.1039/d3re00502j. URL http://dx.doi. org/10.1039/D3RE00502J. [18] Connor J. Taylor, Kobi C. Felton, Daniel Wigh, Mohammed I. Jeraal, Rachel Grainger, Gianni Chessari, Christopher N. Johnson, and Alexei A. Lapkin. Accelerated chemical reaction optimization using multi-task learning. ACS Central Science, 9(5): 957–968, April 2023. ISSN 2374-7951. doi: 10.1021/ acscentsci.3c00050. URL http://dx.doi.org/ 10.1021/acscentsci.3c00050.

[23] Christoph Bannwarth, Sebastian Ehlert, and Stefan Grimme. Gfn2-xtb—an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions. Journal of Chemical Theory and Computation, 15(3): 1652–1671, February 2019. ISSN 1549-9626. doi: 10.1021/acs.jctc.8b01176. URL http://dx.doi. org/10.1021/acs.jctc.8b01176. [24] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021. URL https://arxiv.org/ abs/2106.09685. [25] Tobias Gensch, Gabriel dos Passos Gomes, Pascal Friederich, Ellyn Peters, Théophile Gaudin, Robert Pollice, Kjell Jorner, AkshatKumar Nigam, Michael Lindner-D’Addario, Matthew S. Sigman, and Alán Aspuru-Guzik. A Comprehensive Discovery Platform for Organophosphorus Ligands for Catalysis. Journal of the American Chemical Society, 144(3):1205– 1217, 2022. ISSN 0002-7863. doi: 10.1021/jacs. 1c09718. URL https://doi.org/10.1021/ jacs.1c09718.

[19] John H. Dunlap, Jeffrey G. Ethier, Amelia A. Putnam-Neeb, Sanjay Iyer, Shao-Xiong Lennon Luo, Haosheng Feng, Jose Antonio Garrido Torres, Abigail G. Doyle, Timothy M. Swager, Richard A. Vaia, Peter Mirau, Christopher A. Crouse, and Luke A. Baldwin. Continuous flow synthesis of pyridinium salts accelerated by multi-objective bayesian optimization with active learning. Chemical Science, 14 (30):8061–8069, 2023. ISSN 2041-6539. doi: 10. 1039/d3sc01303k. URL http://dx.doi.org/ 10.1039/D3SC01303K.

[26] Laurianne Moity, Morgan Durand, Adrien Benazzouz, Christel Pierlot, Valérie Molinier, and Jean-Marie Aubry. Panorama of sustainable solvents using the COSMO-RS approach. Green Chemistry, 14(4):1132–1145, 2012. ISSN 1463-9270. doi: 10.1039/C2GC16515E. URL 13

Dynamic language model representations for multi-objective reaction optimisation

American Chemical Society, 145(1):110–121, December 2022. ISSN 1520-5126. doi: 10.1021/ jacs.2c08513. URL http://dx.doi.org/10. 1021/jacs.2c08513.

https://pubs.rsc.org/en/content/ articlelanding/2012/gc/c2gc16515e. [27] Derek M. Dalton, Richard C. Walroth, Caroline Rouget-Virbel, Kyle A. Mack, and F. Dean Toste. Utopia Point Bayesian Optimization Finds ConditionDependent Selectivity for N-Methyl Pyrazole Condensation. Journal of the American Chemical Society, 146(23):15779–15786, 2024. ISSN 0002-7863. doi: 10.1021/jacs.4c01616. URL https://doi.org/ 10.1021/jacs.4c01616.

[33] Lucas W. Souza, Nathan D. Ricke, Braden C. Chaffin, Mike E. Fortunato, Shutian Jiang, Cihan Soylu, Thomas C. Caya, Sii Hong Lau, Katherine A. Wieser, Abigail G. Doyle, and Kian L. Tan. Applying active learning toward building a generalizable model for ni-photoredox cross-electrophile coupling of aryl and alkyl bromides. Journal of the American Chemical Society, 147(22):18747–18759, May 2025. ISSN 1520-5126. doi: 10.1021/jacs.5c02218. URL http: //dx.doi.org/10.1021/jacs.5c02218.

[28] Jamie A. Cadge, Sierra D. Hart, Richard C. Walroth, Kyle A. Mack, and Matthew S. Sigman. Bisphosphine ligand conformer selection to enhance descriptor database representation: improving statistical modelling outcomes. Chemical Science, 16(43): 20473–20485, 2025. ISSN 2041-6539. doi: 10. 1039/d5sc04691b. URL http://dx.doi.org/ 10.1039/D5SC04691B. [29] Alexander S. Shved, Blake E. Ocampo, Elena S. Burlova, Casey L. Olen, N. Ian Rinehart, and Scott E. Denmark. molli: A general purpose python toolkit for combinatorial small molecule library generation, manipulation, and feature extraction. Journal of Chemical Information and Modeling, 64(21):8083–8090, October 2024. ISSN 1549-960X. doi: 10.1021/acs. jcim.4c00424. URL http://dx.doi.org/10. 1021/acs.jcim.4c00424. [30] Melodie Christensen, Lars P. E. Yunker, Folarin Adedeji, Florian Häse, Loı̈c M. Roch, Tobias Gensch, Gabriel dos Passos Gomes, Tara Zepel, Matthew S. Sigman, Alán Aspuru-Guzik, and Jason E. Hein. Data-science driven autonomous process optimization. Communications Chemistry, 4 (1), August 2021. ISSN 2399-3669. doi: 10. 1038/s42004-021-00550-x. URL http://dx.doi. org/10.1038/s42004-021-00550-x. [31] Mohammad H. Samha, Lucas J. Karas, David B. Vogt, Emmanuel C. Odogwu, Jennifer Elward, Jennifer M. Crawford, Janelle E. Steves, and Matthew S. Sigman. Predicting success in cu-catalyzed c–n coupling reactions using data science. Science Advances, 10(3), January 2024. ISSN 2375-2548. doi: 10.1126/sciadv.adn3478. URL http://dx.doi. org/10.1126/sciadv.adn3478. [32] Jordan J. Dotson, Lucy van Dijk, Jacob C. Timmerman, Samantha Grosslight, Richard C. Walroth, Francis Gosselin, Kurt Püntener, Kyle A. Mack, and Matthew S. Sigman. Data-driven multi-objective optimization tactics for catalytic asymmetric reactions using bisphosphine ligands. Journal of the

[34] Shivaani S. Gandhi, Giselle Z. Brown, Santeri Aikonen, Jordan S. Compton, Paulo Neves, Jesus I. Martinez Alvarado, Iulia I. Strambeanu, Kristi A. Leonard, and Abigail G. Doyle. Data science-driven discovery of optimal conditions and a condition-selection model for the chan–lam coupling of primary sulfonamides. ACS Catalysis, 15(3):2292–2304, January 2025. ISSN 2155-5435. doi: 10.1021/ acscatal.4c07972. URL http://dx.doi.org/ 10.1021/acscatal.4c07972. [35] N. Ian Rinehart, Rakesh K. Saunthwal, Joël Wellauer, Andrew F. Zahrt, Lukas Schlemper, Alexander S. Shved, Raphael Bigler, Serena Fantasia, and Scott E. Denmark. A machine-learning tool to predict substrateadaptive conditions for pd-catalyzed c–n couplings. Science, 381(6661):965–972, 2023. ISSN 1095-9203. doi: 10.1126/science.adg2114. URL http://dx. doi.org/10.1126/science.adg2114. [36] Arnau Call, Andrea Palone, Jordan P. Liles, Natalie P. Romer, Jacquelyne A. Read, Josep M. Luis, Matthew S. Sigman, Massimo Bietti, and Miquel Costas. Understanding catalytic enantioselective c–h bond oxidation at nonactivated methylenes through predictive statistical modeling analysis. ACS Catalysis, 15(3):2110–2123, January 2025. ISSN 2155-5435. doi: 10.1021/acscatal.4c05659. URL http://dx. doi.org/10.1021/acscatal.4c05659. [37] Blake E. Ocampo, Bilal Altundas, Matthew J. Bock, Sara Feiz, and Scott E. Denmark. Data-driven prediction of enantioselectivity for the sharpless asymmetric dihydroxylation: Model development and experimental validation. ACS Central Science, 11(9): 1640–1650, 2025. ISSN 2374-7951. doi: 10.1021/ acscentsci.5c00900. URL http://dx.doi.org/ 10.1021/acscentsci.5c00900. [38] Bojana Ranković, Ryan-Rhys Griffiths, Henry B. Moss, and Philippe Schwaller. Bayesian optimisa-

14

Dynamic language model representations for multi-objective reaction optimisation

tion for additive screening and yield improvements – beyond one-hot encoding. Digital Discovery, 3 (4):654–666, 2024. ISSN 2635-098X. doi: 10. 1039/d3dd00096f. URL http://dx.doi.org/ 10.1039/D3DD00096F. [39] Islambek Ashyrmamatov, Su Ji Gwak, Su-Young Jin, Ikhyeong Jun, Umit V. Ucak, Jay-Yoon Lee, and Juyong Lee. A survey on large language models in biology and chemistry. Experimental & Molecular Medicine, April 2026. ISSN 2092-6413. doi: 10.1038/s12276-025-01583-1. URL http://dx. doi.org/10.1038/s12276-025-01583-1. [40] Alexandru Oarga, Matthew Hart, Andres M. Bran, Magdalena Lederbauer, and Philippe Schwaller. Scientific knowledge graph and ontology generation using open large language models. Digital Discovery, 5(3):1269–1279, 2026. ISSN 2635-098X. doi: 10. 1039/d5dd00275c. URL http://dx.doi.org/ 10.1039/D5DD00275C. [41] Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. Augmenting large language models with chemistry tools. Nature Machine Intelligence, 6(5):525–535, May 2024. ISSN 2522-5839. doi: 10.1038/ s42256-024-00832-8. URL http://dx.doi. org/10.1038/s42256-024-00832-8. [42] Yoel Zimmermann, Adib Bazgir, Alexander AlFeghali, Mehrad Ansari, Joshua Bocarsly, L Catherine Brinson, Yuan Chiang, Defne Circi, Min-Hsueh Chiu, Nathan Daelman, Matthew L Evans, Abhijeet S Gangan, Janine George, Hassan Harb, Ghazal Khalighinejad, Sartaaj Takrim Khan, Sascha Klawohn, Magdalena Lederbauer, Soroush Mahjoubi, Bernadette Mohr, Seyed Mohamad Moosavi, Aakash Naik, Aleyna Beste Ozhan, Dieter Plessers, Aritra Roy, Fabian Schöppach, Philippe Schwaller, Carla Terboven, Katharina Ueltzen, Yue Wu, Shang Zhu, Jan Janssen, Calvin Li, Ian Foster, and Ben Blaiszik. 32 examples of llm applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery. Machine Learning: Science and Technology, 6(3):030701, 2025. ISSN 2632-2153. doi: 10.1088/2632-2153/ ae011a. URL http://dx.doi.org/10.1088/ 2632-2153/ae011a.

[44] Nawaf Alampara, Anagha Aneesh, Martiño Rı́osGarcı́a, Adrian Mirza, Mara Schilling-Wilhelmi, Ali Asghar Aghajani, Meiling Sun, Gordan Prastalo, and Kevin Maik Jablonka. General-purpose models for the chemical sciences: Llms and beyond. Chemical Reviews, 126(4):2484–2549, February 2026. ISSN 1520-6890. doi: 10.1021/acs. chemrev.5c00583. URL http://dx.doi.org/ 10.1021/acs.chemrev.5c00583. [45] Andres M. Bran, Théo A. Neukomm, Daniel Armstrong, Zlatko Jončev, and Philippe Schwaller. Chemical reasoning in llms unlocks strategy-aware synthesis planning and reaction mechanism elucidation. Matter, page 102812, April 2026. ISSN 2590-2385. doi: 10.1016/j.matt.2026.102812. URL http://dx. doi.org/10.1016/j.matt.2026.102812. [46] Daniel Armstrong, Zlatko Jončev, Andres M Bran, and Philippe Schwaller. Synthstrategy: Extracting and formalizing latent strategic insights from llms in organic chemistry, 2025. URL https://arxiv. org/abs/2512.01507. [47] Ewa Wieczorek, Joshua W. Sin, Sara Tanovic, Matthew T. O. Holland, Liam Wilbraham, Victor Sebastián-Pérez, Anthony Bradley, Dominik Miketa, Paul E. Brennan, and Fernanda Duarte. Transfer learning for heterocycle retrosynthesis. Journal of Chemical Information and Modeling, 65(15): 7851–7861, 2025. ISSN 1549-960X. doi: 10.1021/ acs.jcim.4c02041. URL http://dx.doi.org/ 10.1021/acs.jcim.4c02041. [48] Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das. Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4(12):1256–1264, December 2022. ISSN 2522-5839. doi: 10.1038/ s42256-022-00580-7. URL http://dx.doi. org/10.1038/s42256-022-00580-7. [49] Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: Largescale self-supervised pretraining for molecular property prediction, 2020. URL https://arxiv.org/abs/2010.09885.

[43] Jieyu Lu and Yingkai Zhang. Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling, 62(6):1376–1387, March 2022. ISSN 1549-960X. doi: 10.1021/acs.jcim.1c01467. URL http://dx.doi. org/10.1021/acs.jcim.1c01467. 15

[50] Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, and Berend Smit. Leveraging large language models for predictive chemistry. Nature Machine Intelligence, 6(2):161–169, February 2024. ISSN 2522-5839. doi: 10.1038/ s42256-023-00788-1. URL http://dx.doi. org/10.1038/s42256-023-00788-1.

Dynamic language model representations for multi-objective reaction optimisation

[51] David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković, and Philippe Schwaller. Bi-semantic chemical embedder for joint representation learning of smiles and natural language, 2026. URL https: //arxiv.org/abs/2608.03855. [52] Bojana Ranković, Ryan-Rhys Griffiths, and Philippe Schwaller. Large language models as uncertaintycalibrated optimizers for experimental discovery, 2025. URL https://arxiv.org/abs/2504. 06265. [53] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/ 20-074.html. [54] Vincent Porte, Luca Hepp, Philipp Kollmus, Shizhao Lu, Eloisa Serrano, Daniela Blanco, and Marco Santagostino. Frugal sampling strategies for navigating complex reaction spaces. Organic Process Research & Development, April 2026. ISSN 1520-586X. doi: 10.1021/acs.oprd.6c00027. URL http://dx.doi. org/10.1021/acs.oprd.6c00027. [55] Julian Götz, Moritz K. Jackl, Chalupat Jindakun, Alexander N. Marziale, Jérôme André, Daniel J. Gosling, Clayton Springer, Marco Palmieri, Marcel Reck, Alexandre Luneau, Cara E. Brocklehurst, and Jeffrey W. Bode. High-throughput synthesis provides data for predicting molecular properties and reaction success. Science Advances, 9(43), October 2023. ISSN 2375-2548. doi: 10.1126/sciadv.adj2314. URL http: //dx.doi.org/10.1126/sciadv.adj2314.

[58] Mohan Neetha, C. M. A. Afsina, Thaipparambil Aneeja, and Gopinathan Anilkumar. Recent advances and prospects in the palladium-catalyzed cyanation of aryl halides. RSC Advances, 10(56): 33683–33699, 2020. ISSN 2046-2069. doi: 10. 1039/d0ra05960a. URL http://dx.doi.org/ 10.1039/d0ra05960a. [59] Daniel T. Cohen and Stephen L. Buchwald. Mild palladium-catalyzed cyanation of (hetero)aryl halides and triflates in aqueous media. Organic Letters, 17 (2):202–205, January 2015. ISSN 1523-7052. doi: 10.1021/ol5032359. URL http://dx.doi.org/ 10.1021/ol5032359. [60] Fraser F. Fleming, Lihua Yao, P. C. Ravikumar, Lee Funk, and Brian C. Shook. Nitrile-containing pharmaceuticals: Efficacious roles of the nitrile pharmacophore. Journal of Medicinal Chemistry, 53(22): 7902–7917, August 2010. ISSN 1520-4804. doi: 10.1021/jm100762r. URL http://dx.doi.org/ 10.1021/jm100762r. [61] Pazhamalai Anbarasan, Thomas Schareina, and Matthias Beller. Recent developments and perspectives in palladium-catalyzed cyanation of aryl halides: synthesis of benzonitriles. Chemical Society Reviews, 40(10):5049, 2011. ISSN 1460-4744. doi: 10.1039/c1cs15004a. URL http://dx.doi.org/ 10.1039/c1cs15004a. [62] Alexander M. Nauth and Till Opatz. Non-toxic cyanide sources and cyanating agents. Organic & Biomolecular Chemistry, 17(1):11–23, 2019. ISSN 1477-0539. doi: 10.1039/c8ob02140f. URL http: //dx.doi.org/10.1039/c8ob02140f.

[56] Julian Götz, Euan Richards, Iain A. Stepek, Yu Takahashi, Yi-Lin Huang, Louis Bertschi, Bertran Rubi, and Jeffrey W. Bode. Predicting three-component reaction outcomes from 40, 000 miniaturized reactant combinations. Science Advances, 11(22), May 2025. ISSN 2375-2548. doi: 10.1126/ sciadv.adw6047. URL http://dx.doi.org/10. 1126/sciadv.adw6047. [57] Babak Mahjour, Rui Zhang, Yuning Shen, Andrew McGrath, Ruheng Zhao, Osama G. Mohamed, Yingfu Lin, Zirong Zhang, James L. Douthwaite, Ashootosh Tripathi, and Tim Cernak. Rapid planning and analysis of high-throughput experiment arrays for reaction discovery. Nature Communications, 14(1), 2023. ISSN 2041-1723. doi: 10. 1038/s41467-023-39531-0. URL http://dx.doi. org/10.1038/s41467-023-39531-0. 16

[63] Nicolas A. Wilson, William M. Palmer, Meredith K. Slimp, Eric M. Simmons, Matthew V. Joannou, Jennifer Albaneze-Walker, Jacob M. Ganley, and Doug E. Frantz. Ni-catalyzed cyanation of (hetero)aryl electrophiles using the nontoxic cyanating reagent K4 [Fe(CN)6 ]. ACS Catalysis, 15(8): 6459–6465, April 2025. ISSN 2155-5435. doi: 10.1021/acscatal.5c00158. URL http://dx.doi. org/10.1021/acscatal.5c00158. [64] H.U. Blaser, F. Spindler, and M. Studer. Enantioselective catalysis in fine chemicals production. Applied Catalysis A: General, 221(1-2):119–143, November 2001. ISSN 0926-860X. doi: 10.1016/ s0926-860x(01)00801-8. URL http://dx.doi. org/10.1016/S0926-860X(01)00801-8. [65] R. Noyori, T. Ikeda, T. Ohkuma, M. Widhalm, M. Kitamura, H. Takaya, S. Akutagawa, N. Sayo, T. Saito,

Dynamic language model representations for multi-objective reaction optimisation

T. Taketomi, and H. Kumobayashi. Stereoselective hydrogenation via dynamic kinetic resolution. Journal of the American Chemical Society, 111(25):9134–9135, December 1989. ISSN 1520-5126. doi: 10.1021/ ja00207a038. URL http://dx.doi.org/10. 1021/ja00207a038. [66] Taiga Yurino, Ryo Nishihara, Toshihisa Yasuda, Shuangli Yang, Noriyuki Utsumi, Takeaki Katayama, Noriyoshi Arai, and Takeshi Ohkuma. Asymmetric hydrogenation of α-alkyl-substituted β-keto esters and amides through dynamic kinetic resolution. Organic Letters, 26(14):2872–2876, January 2024. ISSN 1523-7052. doi: 10.1021/acs. orglett.3c04036. URL http://dx.doi.org/10. 1021/acs.orglett.3c04036.

[72] Dimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther, Teodoro Laino, and Matteo Manica. Unifying molecular and textual representations via multi-task language modelling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 6140–6157. PMLR, 2023. URL https://proceedings.mlr. press/v202/christofidellis23a.html.

[73] An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, [67] Samuel Daulton, Maximilian Balandat, and Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Eytan Bakshy. Differentiable Expected HyperWang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao volume Improvement for Parallel Multi-Objective Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, XiaoBayesian Optimization. In Advances in Neuhuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, ral Information Processing Systems, volume 33, Xuancheng Ren, Yang Fan, Yang Yao, Yichang Zhang, pages 9851–9864. Curran Associates, Inc., 2020. Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru URL https://proceedings.neurips. Zhang, and Zhihao Fan. Qwen2 technical report. arXiv cc/paper_files/paper/2020/hash/ 6fec24eac8f18ed793f5eaad3dd7977c-Abstract. preprint arXiv:2407.10671, 2024. html. [74] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by [68] Samuel Daulton, Maximilian Balandat, and Eytan generative pre-training. 2018. Bakshy. Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume Improvement, 2021. URL http://arxiv.org/ abs/2105.08195. [69] Andreia P. Guerreiro, Carlos M. Fonseca, and Luı́s Paquete. The Hypervolume Indicator: Problems and Algorithms. ACM Computing Surveys, 54(6):1–42, 2022. ISSN 0360-0300, 1557-7341. doi: 10.1145/ 3453474. URL http://arxiv.org/abs/2005. 00515.

[75] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. CoRR, abs/1910.13461, 2019. URL http: //arxiv.org/abs/1910.13461.

[70] Charles Audet, Jean Bigeon, Dominique Cartier, Sébastien Le Digabel, and Ludovic Salomon. Performance indicators in multiobjective optimization. European Journal of Operational Research, 292(2): 397–422, 2021. ISSN 0377-2217. doi: 10.1016/ j.ejor.2020.11.016. URL http://dx.doi.org/ 10.1016/j.ejor.2020.11.016.

[76] Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Huggingface’s transformers: State-of-the-art natural language processing, 2019. URL https://arxiv.org/ abs/1910.03771.

[71] Andrzej Maćkiewicz and Waldemar Ratajczak. Principal components analysis (pca). Computers & Geosciences, 19(3):303–342, March 1993. ISSN 0098-3004. doi: 10.1016/0098-3004(93) 90090-r. URL http://dx.doi.org/10.1016/ 0098-3004(93)90090-R.

[77] Sebastian Burhenne, Dirk Jacob, and Gregor P Henze. Sampling based on Sobol’ sequences for Monte Carlo techniques applied to building simulations. Proceedings of Building Simulation 2011: 12th Conference of International Building Performance Simulation Association, pages 1816–1823, 2011. 17

Dynamic language model representations for multi-objective reaction optimisation

[78] Sebastian Ament, Samuel Daulton, David Eriksson, Maximilian Balandat, and Eytan Bakshy. Unexpected improvements to expected improvement for bayesian optimization, 2023. URL https://arxiv.org/ abs/2310.20708.

Akshay Kulkarni, Shunta Komatsu, Martin.B, JeanBaptiste SCHIRATTI, Hadrien Mary, Donal Byrne, Cristobal Eyzaguirre, cinjon, and Anton Bakhtin. PyTorchLightning/pytorch-lightning: 0.7.6 release, 2020. URL https://zenodo.org/records/ 3828935.

[79] Simone Pilon, Elia Savino, Oliver M. Bayley, Michael Vanzella, Miguel Claros, Petros Siasiaridis, Junsong Liu, Florian Lukas, Matteo Damian, Vasilis Tseliou, Niccolò Intini, Aidan Slattery, Jesus SanJoséOrduna, Tim den Hartog, Ron A. H. Peters, Andrea F. G. Gargano, Francesco G. Mutti, and Timothy Noël. A flexible and affordable self-driving laboratory for automated reaction optimization. Nature Synthesis, April 2026. ISSN 2731-0582. doi: 10.1038/s44160-026-01053-0. URL http://dx. doi.org/10.1038/s44160-026-01053-0. [80] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, highperformance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, number 721, pages 8026– 8037. Curran Associates Inc., 2019.

[84] Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res., 12:2825–2830, 2011. ISSN 1532-4435. [85] Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Rı́o, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant. Array programming with NumPy. Nature, 585 (7825):357–362, September 2020. doi: 10.1038/ s41586-020-2649-2. URL https://doi.org/ 10.1038/s41586-020-2649-2. [86] Lukas Biewald et al. Experiment tracking with weights and biases. 2020.

[81] Jacob R. Gardner, Geoff Pleiss, David Bindel, Kilian Q. Weinberger, and Andrew Gordon Wilson. GPyTorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, pages 7587–7597. Curran Associates Inc., 2018.

[87] John D. Hunter. Matplotlib: A 2D Graphics Environment. Computing in Science & Engineering, 9(3):90– 95, 2007. ISSN 1558-366X. doi: 10.1109/MCSE.2007. 55. URL https://ieeexplore.ieee.org/ document/4160265.

[88] Michael L. Waskom. Seaborn: Statistical data vi[82] Maximilian Balandat, Brian Karrer, Daniel Jiang, sualization. Journal of Open Source Software, 6 Samuel Daulton, Ben Letham, Andrew G Wilson, and (60):3021, 2021. ISSN 2475-9066. doi: 10.21105/ Eytan Bakshy. BoTorch: A Framework for Efficient joss.03021. URL https://joss.theoj.org/ Monte-Carlo Bayesian Optimization. In Advances papers/10.21105/joss.03021. in Neural Information Processing Systems, volume 33, pages 21524–21538. Curran Associates, Inc., 2020. URL https://proceedings.neurips. cc/paper_files/paper/2020/hash/ f5b1b89d98b7286673128a5fb112cb9a-Abstract. html. [83] William Falcon, Jirka Borovec, Adrian Wälchli, Nic Eggert, Justus Schock, Jeremy Jordan, Nicki Skafte, Ir1dXD, Vadim Bereznyuk, Ethan Harris, Tullie Murrell, Peter Yu, Sebastian Præsius, Travis Addair, Jacob Zhong, Dmitry Lipin, So Uchida, Shreyas Bapat, Hendrik Schröter, Boris Dayma, Alexey Karnachev, 18

Dynamic language model representations for multi-objective reaction optimisation

A. Supplemental main text benchmarking studies Nickel-catalysed Suzuki coupling dataset 90% Hypervolume threshold

95% Hypervolume threshold Iterations to reach threshold

Iterations to reach threshold

40

Palladium-catalysed Suzuki coupling dataset

30 20 10 0

90% Hypervolume threshold

90% Hypervolume threshold

95% Hypervolume threshold

90% Hypervolume threshold

95% Hypervolume threshold

80 60 40 20 0

95% Hypervolume threshold 100%

% of seeds achieving threshold

100%

% of seeds achieving threshold

100

80% 60% 40% 20% 0%

LLM-GP

80% 60% 40% 20% 0%

DFT Descriptor databases

One-hot encoding

Extended Data Figure 1. Sequential low-data optimisation benchmarks comparing different representation methods (the LLM-GP framework, DFT descriptor databases, and one-hot encoding) on the nickel-catalysed and palladium-catalysed Suzuki coupling datasets. Top panels: mean number of optimisation iterations required to reach 90% and 95% hypervolume thresholds (lower is better, error bars indicate standard error across 20 random seeds). Bottom panels: percentage of optimisation runs (initialised with different random seeds) reaching each threshold within the allocated experimental budget (higher is better). Batch size 24 HTE optimisation

Batch size 48 HTE optimisation

4

2

0

90% Hypervolume threshold

40% 20% 0%

3 2 1 0

90% Hypervolume threshold

90% Hypervolume threshold

95% Hypervolume threshold

90% Hypervolume threshold

95% Hypervolume threshold

3 2 1 0

100%

80% 60% 40% 20% 0%

LLM-GP

4

95% Hypervolume threshold

100%

% of seeds achieving threshold

% of seeds achieving threshold

60%

Batch size 96 HTE optimisation

95% Hypervolume threshold

4

95% Hypervolume threshold

100% 80%

90% Hypervolume threshold

% of seeds achieving threshold

6

5

Iterations to reach threshold

95% Hypervolume threshold Iterations to reach threshold

Iterations to reach threshold

90% Hypervolume threshold

DFT Descriptor databases

80% 60% 40% 20% 0%

One-hot encoding

Extended Data Figure 2. High-throughput experimentation (HTE) optimisation benchmarks on the palladium-catalysed sulfonamide coupling dataset, comparing the LLM-GP framework, DFT descriptor databases, and one-hot encoding across 24-well (left), 48-well (middle), and 96-well (right) plate batch sizes. Top panels: mean number of plate iterations required to reach 90% and 95% hypervolume thresholds (lower is better, error bars indicate standard error across 20 random seeds). Bottom panels: percentage of optimisation runs (initialised with different random seeds) reaching each threshold within the allocated experimental budget (higher is better).

19

Dynamic language model representations for multi-objective reaction optimisation

B. Influence of different pre-trained models and pooling strategies Nickel-catalysed Suzuki coupling dataset 95% Hypervolume threshold

Iterations to reach threshold

90% Hypervolume threshold

Palladium-catalysed Suzuki coupling dataset 95% Hypervolume threshold

Iterations to reach threshold

90% Hypervolume threshold

Extended Data Figure 3. Ablation studies over pre-trained language model architectures and pooling strategies on the nickel-catalysed (top) and palladium-catalysed (bottom) Suzuki coupling datasets. Mean number of optimisation iterations required to reach 90% and 95% hypervolume thresholds are compared across T5-chem (mean pooling), BART-base (mean and CLS pooling), T5-base and T5-small (mean pooling), and Qwen2.5-0.5B (mean and last token pooling). Error bars indicate standard error across 20 random seeds. T5-base with mean pooling consistently achieved practical hypervolume thresholds within the fewest iterations across all datasets and thresholds.

20

Dynamic language model representations for multi-objective reaction optimisation

Nickel-catalysed Suzuki coupling dataset 90% Hypervolume threshold

95% Hypervolume threshold

Palladium-catalysed Suzuki coupling dataset 90% Hypervolume threshold

95% Hypervolume threshold

Extended Data Figure 4. Percentage of optimisation runs (initialised with 20 different random seeds) reaching 90% and 95% hypervolume thresholds within the allocated experimental budget, across pre-trained language model architectures and pooling strategies on the nickel-catalysed (top) and palladium-catalysed (bottom) Suzuki coupling datasets. Models and pooling strategies are as described in the Methods.

21

Dynamic language model representations for multi-objective reaction optimisation

Batch size 24 HTE optimisation 95% Hypervolume threshold

Iterations to reach threshold

90% Hypervolume threshold

Iterations to reach threshold

Batch size 48 HTE optimisation 90% Hypervolume threshold

95% Hypervolume threshold

Iterations to reach threshold

Batch size 96 HTE optimisation 90% Hypervolume threshold

95% Hypervolume threshold

Extended Data Figure 5. Ablation studies over pre-trained language model architectures and pooling strategies for high-throughput experimentation (HTE) optimisation on the palladium-catalysed sulfonamide coupling dataset, across 24-well (top), 48-well (middle), and 96-well (bottom) plate batch sizes. The bar charts show the mean number of plate iterations required to reach 90% and 95% hypervolume thresholds. Models and pooling strategies are as described in the Methods. Error bars indicate standard error across 20 random seeds.

22

Dynamic language model representations for multi-objective reaction optimisation

Batch size 24 HTE optimisation 90% Hypervolume threshold

95% Hypervolume threshold

Batch size 48 HTE optimisation 90% Hypervolume threshold

95% Hypervolume threshold

Batch size 96 HTE optimisation 90% Hypervolume threshold

95% Hypervolume threshold

Extended Data Figure 6. Percentage of optimisation runs (initialised with 20 different random seeds) reaching 90% and 95% hypervolume thresholds within the allocated experimental budget, across pre-trained language model architectures and pooling strategies for highthroughput experimentation (HTE) optimisation benchmarks on the palladium-catalysed sulfonamide coupling dataset, across 24-well (top), 48-well (middle), and 96-well (bottom) plate batch sizes. Models and pooling strategies are as described in the Methods.

23

Dynamic language model representations for multi-objective reaction optimisation

C. Design spaces for prospective optimisation campaigns Pd catalyst (0.01 eq) Cyanide source (1.5 - 3.0 CN− eq) Additive (0.5 - 1.5 eq)

O F

O

O F

Solvent (240 µL) / Co-solvent (120 µL) T (oC), 20 h

Br

19.0 µL

O CN

30,000 reaction condition combinations Mono- and bidentate Palladium catalysts (40)

Ph

Ph Ph

Ph Cl P Pd P Ph Cl Ph Ph

P

Ph

Ph Ph P P Ph Pd Ph Cl2

PdCl2

Fe P Ph

Ph

Cl Pd Cl

P

N

MeO

tBu

OMe tBu P Pd

OTf Cy

iPr

iPr

iPr

iPr

Cy OMs P Pd

OMs

tBu tBu P

OMs

Pd

H 2N

P

Pd

H 2N

2

H 2N

N

Me2N iPr

iPr

O

[PdCl2(PPh3)2]

MeO

Cy

[PdCl2(DPPF)]

OMe Cy OMs P Pd iPr

iPr

[PdCl2(DPPP)]

Ph Ph P O

H 2N

Ph Cl

Pd

P P Ph

P Ph Ph

iPr

Et

MeO

P Et

Et

Pd

PdCl2

Fe

Et Cl Pd Cl Et N

H 2N

P tBu

Ad

[PdCl2(DTBPF)]

Pd PEPPSI IPent

Ph Ph P Cl Pd

Cy

Cl

O Cy OMs P Pd iPr

iPr

P Ph

MeO

Cy

H 2N

EPhos Pd G3

HN

O

tBu

tBu P H

Pd Cl

Ph

Cl Pd

Cl

Ad

OMs

Ad P

Pd

OMs Pd

NMe2

Me2N

Ph

HN

H

tBu P

Ph

HN

Fe

Ph

MeCN

O

CPME

tBu tBu P tBu

Pd H 2N

RuPhos Pd G3

OMs

O

O

Pd

Ph

P

HN

MsO

Pd

[Pd(allyl)(tBuXPhos)]OTf

O

NH2

tBu

tBu OMs

Cy

Pd

Cy

P

Pd Fe

tBu tBu

OMs P

P

Pd

OMs

Pd

Cy P

H 2N

MeCgPPh Pd G3

PtBu3 Pd G4

tBu

H 2N

HN

H 2N

tBu

Cy

Ph Ph

PCy3 Pd G4

DTBPF Pd G3

O

O

H 2N

Cy P

Cl Pd MeO H 2N

Cy

O Cy OMs P Pd

O

tBuOH

NBP

H 2N

tBu

GPhos Pd G3

sSPhos Pd G2

iPr

iPr H 2N

HN

iPr

Me4tBuXPhos Pd G3

Additives (5)

O

K4[Fe(CN)6].3H2O

Zn(CN)2

NaCN

O

DMC

tBuXPhos Pd G4

Temperatures (3)

none

Zn

50

PhMe 75

KOAc

Co-solvents (2) O

tBu OMs P Pd

O N

O

none EtOAc

tBu OMs P Pd iPr

iPr iPr

iPr

NaSO3

tBu

Cyanide sources (3)

OH

O

TrixiePhos Pd G3

iPr

O N

OTf

OMs

Cy P O

tBu OMs Pd

QPhos Pd G3

PAd3 Pd G3

N

DMF

Cy

iPr

Xantphos Pd G3

2

O

iPr O

H 2N

P Ph Ph

iPr

iPr

PhCPhos Pd G3

Solvents (10)

MeTHF

iPr

tBu P Pd

tBu

Cl

OMs

Ph P NMe2

Me2N

CPhos Pd G4

tBu

Ad Cl

[HtBuXPhos]2[Pd2Cl6]

O

Ph Pd Ph

O

[Pd(crotyl)Cl(RuPhos)]

[Pd(allyl)Cl((R)-BINAP)]

Ph

iPr

N-Xantphos Pd G3

H 2N

Cy P

Cy

BrettPhos Pd G4

Cl

Cl

iPr

iPr

H 2N

P Ph Ph

P P

AdBrettPhos Pd G3

HN

O

iPr

iPr

Ph Ph OMs P Pd

Ph

Ph Ph OMs P Pd

Cy Cl P Pd

Cy

MorDalphos Pd G3

AmPhos Pd G3

iPr

OMe Cy OMs P Pd iPr

iPr

Ph

[PdCl2(DPEPhos)]

Cl

iPr

iPr

O

iPr O

[PdCl2(XantPhos)]

tBu

Cl

cataCXium A Pd G3

Cl

P Ph Ph

OMe Ad OMs P Pd iPr

iPr

Pd

O

iPr Cl

[Pd(cinnamyl)Cl(IPr)]

tBu

tBu

OMs P

iPr Pd

[PdCl2((rac)-BINAP)]

N

N

Pd Cl Ph

XPhos Pd G3

Ph Ph P

N

N

Cl

Et

Et

Et

Ph

Ph

[Pd(allyl)(Xantphos)]Cl

BrettPhos Pd G3

[Pd(allyl)(tBuBrettPhos)]OTf

[PdCl2(AmPhos)]

water

K2CO3

N

100

γ-valerolactone

Extended Data Figure 7. Reaction condition search space for the prospective Pd-catalysed cyanation reaction comprising 40 mono- and bidentate Pd catalysts, 10 solvents, 3 cyanide sources, 2 co-solvents, 5 additives, and 3 temperatures.

24

Dynamic language model representations for multi-objective reaction optimisation

Chiral catalyst (1 mol%) Base (20/50 mol%)

O

OH

Solvent (200 µL), H2 (20 bar) T (oC), 20 h 20 mg

8064 reaction conditions combinations Chiral iridium and ruthenium catalysts (32) Ph

Me O Ph

Ir

Ph

Me

O

N

CF3

Ph

O

B

PPh2

Ph CF3

Ir

Cl

Cl CF3

Ph

O

N

Ph Ph

Ph Ph

P

P

B

PPh2

Ru

Ru CF3

4

P

4

iPr

Cl

[Ir(COD)((R,R)-Ph-UBAPHOX)]BARF

Ph Ph P Cl Ru

H2 N

Ph Ph P Cl Ru

H2 N

P Cl Ph Ph

N H2

P Cl Ph Ph

N H2

Ph Ph

[RuCl2((S)-BINAP)((R,R)-DPEN)]

[RuCl(p-cymene)((R)-BINAP)]Cl

[RuCl2((R)-BINAP)((R,R)-DPEN)]

P

P Ru

P

P

iPr

Cl

tBu

tBu

P

MeO MeO

Ru iPr

Cl

tBu

tBu

Ru Cl N

Ph

Ph

tBu

H

O P

O

Cl

H2 N

O iPr

Ru

O

P

OMe

N H2

Cl

O

O

P

O

P

Cl

H2 N

iPr

P

Ru Cl

Cl

H2 N

P

MeO MeO

tBu

P

Cl

P

tBu

H2 N

Cl

Cl

P

N Ph

Cl

N H2

P

MeO MeO

Ph

OMe

Ph Ph P Cl Ru

H2 N

P Cl Ph Ph

N H2

Ph Ph

P

[RuCl2((S)-DM-BINAP)((S)-DAIPEN)]

Cl

H2 N

P

Ph Ph P Cl Ru

H2 N

P Cl Ph Ph

N H2

H2 N

Ph Ph

P Cl Ph Ph

N H2

[RuCl2((R)-BINAP)((S,S)-DPEN)]

Cl

OH

Cl Cl

H2 N

Ph

Ph

Ph

P

Ph

P

Cl

H2 N

Ph

Ru N H2

H2 N

Cl

N H2

Ph

[RuCl2((S)-DM-BINAP)((S,S)-DPEN)]

P

Cl

H2 N

Ph

Ru N H2

P

P OMe

Cl

N H2

Ph

OMe

iPr OMe

[RuCl((R)-DM-BINAP)((R)-DAIPENA)]

Ph Ph P Cl Ru

H2 N

P Cl Ph Ph

N H2

KOH

OMe

OMe

[RuCl2((S)-BINAP)((S)-DAIPEN)]

Na2CO3

K2CO3

NaOtBu

O

[RuCl2((R)-DM-BINAP)((R,R)-DPEN)]

iPr

[RuCl2((R)-BINAP)((R)-DAIPEN)]

Bases (9) none

O

Ru Cl NH2

O

Ru N H2

[RuCl((S)-DM-BINAP)((S)-DAIPENA)]

Ph Ph P Cl Ru

Solvents (7) OH

N

[RuCl(p-cymene)((R,R)-Ts-DPEN)]

[RuCl2((R)-DM-PPhos)((R,R)-DPEN)]

Ru P

OMe

[RuCl2((S)-BINAP)((S,S)-DPEN)]

OH

H

Ru P

OMe

OH

Ph

S

OMe

OMe

[RuCl2((R)-DM-BINAP)((R)-DAIPEN)]

[RuCl2((R)-DM-SEGPHOS)((R)-DAIPEN)]

Ph

[RuCl(p-cymene)((S,S)-Ts-DPEN)]

O

OMe

OMe

Ph

OMe

O

[RuCl2((S)-DM-SEGPHOS)((S)-DAIPEN)]

Ph

N

[RuCl2((S)-DM-PPhos)((S,S)-DPEN)]

OMe

N H2

Cl

H2 N

Ru P

iPr

Ru

OMe

N H2

Ru Cl N

O

OMe

[Ru(OAc)2((S)-DTB-MeOBIPHEP)]

iPr

Ru

OMe

N H2

Ru Cl NH2

OMe

N

Ru(OAc)2 P

tBu

N

[RuCl((R,R)-Teth-Ts-DPEN)]

[RuCl((S,S)-Teth-Ts-DPEN)]

tBu

tBu

tBu

[Ru(OAc)2((R)-DTB-MeOBIPHEP)]

S

OMe

tBu

tBu

[RuCl(p-cymene)((S)-DM-BINAP)]Cl

[IrClH2((S)-DTB-SpiroPAP-3-Me)]

O

N

N

[RuCl(p-cymene)((R)-DM-BINAP)]Cl

iPr

N

tBu

P

MeO MeO

Ru(OAc)2 P

S O

Ph Ph

[Ru(OAc)2((S)-MeOBIPHEP)]

tBu

S

iPr O

Ru(OAc)2 P

Ph Ph

tBu

[IrClH2((R)-DTB-SpiroPAP-3-Me)]

P

MeO MeO

Ru(OAc)2 P

[Ru(OAc)2((R)-MeOBIPHEP)]

Cl

Cl

O

N

Ph Ph

P

MeO MeO

Ph

P(DTB)2 Cl H Ir H HN

O

[RuCl(p-cymene)((S)-BINAP)]Cl

Ph Ph Ph

N

iPr

P Cl Ph Ph

Ph Ph

[Ir(COD)((S,S)-Ph-UBAPHOX)]BARF

P(DTB)2 Cl H Ir H HN

KOtBu

NaOH

Base mol% (2)

Temperatures (2)

20

30

50

60

N

MeOH

EtOH

iPrOH

tAmOH

THF

MeTHF

PhMe

N

N

Extended Data Figure 8. Reaction condition search space for the prospective asymmetric ketone hydrogenation comprising 32 chiral catalysts, 7 solvents, 9 bases, 2 base mol%, and 2 temperatures.

25

Supplementary Information Dynamic language model representations for multi-objective reaction optimisation Joshua W. Sin1,2† , David Ming Segura1,2,3† , Bojana Ranković2,3† , Siu Lun Chau4 , Marius D. R. Lutz5 , Andrea Anelli5 , Ryan P. Burwood6 , Kurt Püntener1 , Maximilian J. Notheis1 , Raphael Bigler1 , Philippe Schwaller2,3* 1

Process Chemistry & Catalysis, Synthetic Molecules Technical Development, F. Hoffmann-La Roche AG, Basel, Switzerland 2 Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland 3 National Centre of Competence in Research (NCCR) Catalysis, EPFL, Lausanne, Switzerland 4 Epistemic Intelligence & Computation Lab, College of Computing & Data Science, Nanyang Technological University, Singapore 5 Roche Pharma Research and Early Development (pRED), F. Hoffmann-La Roche AG, Basel, Switzerland 6 Solid State Sciences, Synthetic Molecules Technical Development, F. Hoffmann-La Roche AG, Basel, Switzerland †These authors contributed equally to this work. *Corresponding authors. Email: [email protected]

Contents 1 High-throughput experimentation (HTE) platform

2

2 Prospective application: Palladium-catalysed cyanation 2.1 HTE experimental procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Scale-up and synthesis of methyl 3-cyano-5-fluorobenzoate . . . . . . . . . . . . . . . .

4 4 4

3 Prospective application: Asymmetric ketone hydrogenation 3.1 HTE experimental procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2 Scale-up and synthesis of (1S, 2R)-1,2-diphenylpropan-1-ol . . . . . . . . . . . . . . . . 3.3 Enantiomeric data augmentation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4 Model sensitivity to initial training data . . . . . . . . . . . . . . . . . . . . . . . . . .

7 7 7 10 10

1

1

High-throughput experimentation (HTE) platform

We conducted all high-throughput experiments using a parallel experimentation platform customdesigned by UnchainedLabs (Supplementary Figure 1). This system comprises two interconnected Big Kahuna platforms, one of which is further integrated with a LiCONiC LiCotel system. The entire setup is encased within an LC Technology Solutions glove box featuring dual circulation systems, with one dedicated to solid dispensing and another to reaction execution. LC-MS analysis was performed using an ACQUITY UPLC I-Class system with QDa from Waters. For Chiral HPLC analyses, a Chiralcel OJ-3 column from Daicel was used, with ethanol and heptane as the mobile phase. Solid components were dispensed with a Mettler balance in vial dispense mode, using SV hoppers for precursors and ligands and SV hoppers (≤ 15 mg) or 10 mL classic hoppers (>15 mg) for solid additives. For target dispense quantities < 0.4 mg, materials were dispensed as coated ChemBeads. Following automated solid dispensing, liquid components (substrates, liquid additives, and solvents) were manually added using an Eppendorf Multipette E3 single channel pipette (4987000010). Reactions were performed in standard 96-position parallel synthesis reaction blocks (Analytical Sales and Services, SKU: 96960), with V&P Scientific super tumble stir discs (VP 721F-1) for stirring, or in a Screening Pressure Reactor (SPR) from UnchainedLabs. Data analysis was performed using the HTE OS workflow described by Wuitschik et al. [1] HTE OS is integrated with a SpotFire application that enables tagging of LC-MS and HPLC signals into categories (e.g., “limiting SM”, “other SM”, “solvent”, “ignore peak”). During analysis, all peaks were tagged accordingly, with peaks corresponding to ligands and precatalyst components designated as “ignore peak”. For LCAP calculations, signals tagged as “other SM”, “solvent”, or “ignore peak” were excluded from consideration.

2

h-Output Reaction Screening System Facility

4

1

2 Screening LC/MS LiCotel

highly performant UPLC with short standard methods to measure up to 1,000 samples/d

stores up to 4,500 compounds (e.g., catalysts, substrates) 2

3

Solid Dosing

Reaction Execution

Big Kahuna with two balances for precise (submg) solid dosing to vials or plates

Big Kahuna for accurate liquid dispensing and parallel reaction execution (-20 to 150 °C)

Supplementary Figure 1: Schematic overview of the high-throughput experimentation (HTE) setup.

3

2

Prospective application: Palladium-catalysed cyanation

2.1

HTE experimental procedure

For all HTE plates, the following experimental procedure was performed. Palladium catalysts (1 mol%), cyanide sources, and solid additives were dispensed into 1 mL vials with stirring disks in a 96-well plate. Methyl 3-bromo-5-fluorobenzoate (19.0 µL, 129 µmol), solvents (240 µL), and cosolvent (120 µL) were added to the vials, followed by the addition of liquid additive NEt3 (26.9 µL, 1.50 eq) where applicable. The plate was sealed and stirred (400 rpm) at the specified temperature for 20 h. 4/1 MeCN/H2 O (400 µL) was added to each well and shaken for 20 min at room temperature. Samples were taken and analysed by Liquid Chromatography-Mass Spectrometry (LC-MS).

2.2

Scale-up and synthesis of methyl 3-cyano-5-fluorobenzoate O F

O

Br

[Pd(allyl)(Xantphos)]Cl (1 mol%) K4Fe(CN)6.3H2O (0.5 eq) DMC (40 mL) / H2O (20mL) 75oC, 4 h

O F

O

CN

A nitrogen-flushed 100 mL Easymax flask with overhead stirring was charged with methyl 3-bromo-5fluorobenzoate (3.17 mL, 21.5 mmol) and DMC (40 mL). In a glovebox, a solution of K4 Fe(CN)6 · 3 H2 O (4.53 g, 10.7 mmol, 0.50 eq) in water (20 mL) was prepared and added to the Easymax flask under argon. The mixture was heated to 75 ◦ C. Once the temperature was reached, [Pd(allyl)(Xantphos)]Cl (163.4 mg, 215 µmol, 1.0 mol%) was added and the mixture was stirred at 75 °C for 4 h. The reaction was transferred with degassed 25% aq. NaCl solution (100 mL) to a separation funnel and extracted three times with degassed EtOAc (3 × 100 mL). The combined organic phases were washed with sat. aq. NaCl solution, dried over Na2 SO4 , filtered, and the solvent was removed using a rotary evaporator. The crude product was purified by flash column chromatography on silica gel (EtOAc:heptane = 0:100 to 30:70) over 20 min, 80 g, only UV active fractions collected), affording the product as a white solid. Yield: 3.61 g (94%). 1 H NMR (400 MHz, CDCl ): δ 8.18 – 8.13 (m, 1H, ArH ), 8.00 – 7.95 (m, 1H, ArH ), 7.58 – 7.53 (m, 3

1H, ArH ), 3.98 (s, 3H, OCH 3 ). 13 C{1 H} NMR (101 MHz, CDCl ): δ 164.0 (d, 4 J 1 3 F,C = 2.9 Hz, C O2 Me), 162.1 (d, JF,C = 252.4 Hz, arom.), 133.8 (d, 3 JF,C = 7.7 Hz, arom.), 129.2 (d, 4 JF,C = 3.7 Hz, arom.), 123.1 (d, 2 JF,C = 24.6 Hz, arom.), 121.4 (d, 2 JF,C = 23.1 Hz, arom.), 116.7 (d, 4 JF,C = 2.9 Hz, C N), 114.3 (d, 3 JF,C = 8.8

Hz, arom.), 53.0 (OC H3 ).

4

1

H NMR (400 MHz, CDCl3): δ 8.18 – 8.13 (m, 1H, ArH), 8.00 – 7.95 (m, 1H, ArH), 7.58 – 7.53 (m, 1H, ArH), 3.98 (s, 3H, OCH3). S%JE:d

JLe# JLe#

!""#$%"C"CC'EF%$"*"HC"#,C-.LMNO .3E4R4-4RS7N H83V:

S%"ENad

S%*ENad

e

#

JL%%

"L%%

"L%% "L%%

fLee fLee fLe# fLe# fLef fLef fL,f fL,f fL,, fL,, fL,, fL,,

#L"$ #L"$

S%$ENd

f

H

,

$

J

*

"

.;<=>?@A8_;>Ca8Ecc=d

Supplementary Figure 2: 1 H NMR (400 MHz, CDCl3 ) of methyl 3-cyano-5-fluorobenzoate C{1H} NMR (101 MHz, CDCl3): δ 164.0 (d, 4JF,C = 2.9 Hz, CO2Me), 162.1 (d, 1JF,C = 252.4 Hz, arom.), 133.8 (d, 3JF,C = 7.7 Hz, arom.), 129.2 (d, 4JF,C = 3.7 Hz, arom.), 123.1 (d, 2JF,C = 24.6 Hz, arom.), 121.4 (d, 2JF,C = 23.1 Hz, arom.), 116.7 (d, 4JF,C = 2.9 Hz, CN), 114.3 (d, 3JF,C = 8.8 Hz, arom.), 53.0 (OCH3). 13

5

S%8c;e

,fL%' ,fL%'

!""#$%"C'CCEF*%$"'"HC"#,C-.LMNO .3F4R4-4RS7N 8V.:; S%,cNe S%$cNe S%JcNe S%HcNe S%#cNe

"'8L"8 "'8L"8 "'fL'% "'fL'% "''L8H "''L8H "'"L$J "'"L'$ "'"L'$

S%fcNe

S%'cNe

"#%

"H%

"$%

""HLH$ ""HLH$ ""$Lf$ ""$Lf$ ""$L', ""$L',

"ffL#f "ffL#f "ffLJ,

"HfL8$ "HfL8$ "HfLf' "HfLf' "H%L#" "H%L#"

S%"cNe

"'%

"%%

#%

H%

$%

.<=>?@A_VC<?aEVcdd>e

Supplementary Figure 3: 13 C{1 H} NMR (101 MHz, CDCl3 ) of methyl 3-cyano-5-fluorobenzoate

6

3

Prospective application: Asymmetric ketone hydrogenation

3.1

HTE experimental procedure

For all HTE plates, the following experimental procedure was performed inside a nitrogen filled glovebox. Iridium and ruthenium catalysts (1 mol%) and solid bases were dispensed into 1 mL vials in a 96-well plate. Stock solutions of 1,2-diphenylpropan-1-one (20 mg, 95.1 µmol in solvent (200 µL)) were prepared and added to the vials, followed by addition of liquid bases, where applicable. The plate was sealed inside the glovebox, placed in a screening pressure reactor (SPR) and shaken overnight at the indicated temperature under 20 bar H2 . The solvent was removed using a GeneVac and the residue redissolved in EtOH (200 µL). Samples were taken and analysed by chiral HPLC.

3.2

Scale-up and synthesis of (1S, 2R)-1,2-diphenylpropan-1-ol

O

[Ir(Cl)H2((S)-DTB-SpiroPAP-3-Me)] (1 mol%) DBU (0.5 eq)

OH

tAmOH (19.2 eq), H2 (20 bar) 60oC, 4 h

In an argon filled glovebox, [Ir(Cl)H2 ((S) – DTB – SpiroPAP – 3-Me)] (93.1 mg, 95.1 µmol, 0.01 eq) and 1,2-diphenylpropan-1-one (2000 mg, 9.51 mmol, 1.0 eq) were added to a 50 mL autoclave. Degassed tAmOH (16.1 g, 20 mL, 19.2 eq) and DBU (724.0 mg, 711.2 µL, 4.76 mmol, 0.5 eq) were subsequently added to the reaction mixture. The autoclave was then sealed, pressurised with Ar (5 bar) and removed from the glove box. The autoclave was connected to a H2 line, purged five times with H2 (10 bar) and stirred at 60 °C and 500 rpm under 20 bar of H2 . Reaction progress was monitored via H2 uptake, reaching completion within 4 h. The autoclave was allowed to cool to room temperature and carefully depressurised. The deep yellow reaction mixture was transferred with EtOAc (130 mL) to a separation funnel and washed sequentially with 1 M aqueous HCl solution (2 * 50 mL), water (50 mL) and saturated aq. NaCl solution (50 mL). The organic phase was dried over Na2 SO4 , filtered, and concentrated under reduced pressure to afford a lightly brown oil. This crude oil was purified by flash column chromatography, eluting with EtOAc/heptane (18%, v/v). Solvent removal under reduced pressure yielded 1,2-diphenylpropan-1-ol as a transparent, oily solid (1.99 g). The crude solid was then dissolved in pentane (10 mL) and allowed to crystallise by slow evaporation over 48 h at room temperature. The resulting crystal bearing oily residue on the surface was transferred into a 15 mL glass vial. Ice-cold pentane (3 mL) was added to submerge the crystal. The vial was gently swirled for 3–5 seconds, after which the supernatant was drawn off. The crystal was rinsed a second time with ice-cold pentane (3 mL) and immediately decanted. The purified crystal was transferred onto lint-free filter paper and was allowed to dry to constant mass at room temperature, affording the product. Yield: 1687.89 mg (83.6%). 1H

NMR (400 MHz, DMSO – d6 ): δ 7.22 – 7.07 (m, 10H, ArH ), 5.31 (d, 3 JH,H’ = 4.8 Hz, 1H, (H)COH ), 4.64 (dd, 3 JH,H’ = 4.6 Hz, 3 JH,H’ = 6.4 Hz, 1H, (H )COH), 2.96 (p, 3 JH,H’ = 6.9 Hz, 1H, (H )CCH3 ), 1.24 (d, 3 JH,H’ = 7.0 Hz, 3H, (H)CCH 3 ).

7

13 C{1 H} NMR (101 MHz, DMSO – d ): 6

δ 145.41 (arom.), 145.18 (arom.), 128.52 (arom.), 128.23 (arom.), 127.95 (arom.), 126.99 (arom.), 126.90 (arom.), 126.25 (arom.), 77.49 (C OH), 47.53 (C CH3 ), 16.97 (CC H3 )

V119360_1__ELN042419_432_1_final_sample_dmso.jdx DMSO-d6 16 H's

7.20

7.17

5.32

2.97

5.32 5.31

7.14 7.10 7.09 7.14 7.12 7.09 7.10

Water 4.64

10.1

1.0

7.0

6.5

6.0

5.5

3.00 2.98 2.97 2.95 2.93

4.66 4.65 4.64 4.63

7.24 7.23 7.23

7.21 7.20 7.16

1.26 1.24

1.25

1.0

5.0

DMSO-d6

1.0

4.5

4.0

3.5

3.0

3.0

2.5

2.0

1.5

Chemical Shift (ppm)

Supplementary Figure 4: 1 H NMR (400 MHz, DMSO – d6 ) of (1S, 2R)-1,2-diphenylpropan-1-ol

8

V119360_2__ELN042419_432_1_final_sample_dmso_13C.jdx DMSO-d6 11 C's

128.52 128.23 127.95 126.99 126.9

128.52 128.52 128.23 128.23 127.95 127.95 126.99 126.99

126.25

DMSO-d6

77.49

16.97 16.97

47.53 47.53

16.97

145.41 145.41 145.18 145.18

145.18

47.53

77.49 77.49

126.90 126.90 126.25 126.25

145.41

152

144

136

128

120

112

104

96

88

80

72

64

56

48

40

32

24

16 Chemical Shift (ppm)

Supplementary Figure 5: 13 C{1 H} NMR (101 MHz, DMSO – d6 ) of (1S, 2R)-1,2-diphenylpropan1-ol

9

3.3

Enantiomeric data augmentation

All solvents and bases in the design space are achiral, and all 32 chiral catalysts are present as enantiomeric pairs. A catalyst and its enantiomer are therefore expected to yield mirror-image product distributions under otherwise identical conditions. Conversion and diastereomeric excess are unchanged, while the enantiomeric excess changes sign. We exploited this symmetry as a data augmentation strategy. For each observed reaction, a reflected counterpart was added to the training data in which the catalyst was replaced by its enantiomer, the sign of ee (syn) inverted, and conversion and de (syn) retained. Augmented observations were used to fit the surrogate models only. The acquisition function was evaluated over the full design space excluding conditions already run experimentally.

3.4

Model sensitivity to initial training data

The first initialisation plate in the HTE reaction optimisation campaign contained only five conditions producing the target syn diastereomer. To assess how strongly the round-2 optimisation outcome depended on these observations, we repeated the optimisation with the top n conditions removed from the Plate 1 training data, for n = 1, 3, and 5, ranking conditions by de (syn). In each case the surrogate models were refitted on the reduced dataset and the acquisition function was evaluated over the full design space, excluding conditions already queried. No new experiments were performed: suggested conditions for which experimental data were already available were compared against the best condition remaining in the ablated training set. Because only suggestions already measured in round 2 could be evaluated, the values below are lower bounds on the performance recoverable at each ablation level. With the single most diastereoselective condition removed (n = 1), the best remaining training observation reached 47.9% de (syn). The model recovered conditions delivering 99.4% conversion, 75.4% de (syn), and 89.5% ee (syn), improving stereoselectivity. With the top three removed (n = 3), the best remaining training observation reached 23.5% de (syn), and the model recovered conditions at 64.2% de (syn), again with higher enantioselectivity. Removing all five conditions (n = 5) that were syn-selective left no positive diastereoselectivity examples in the training data, yet the model still recovered conditions reaching 83.4% de (syn) and 93.3% ee (syn). Even with no syn-selective examples to learn from, the model identified conditions delivering high syn selectivity.

10

References [1]

G. Wuitschik, V. Jost, T. Schindler, M. Jakubik, Organic Process Research & Development 2024, 28, 2875–2884, DOI 10.1021/acs.oprd.4c00160.

11

Record · ID 673506 · SHA-256 7c7e0ad5697e777c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.