Conceptio › Archive › arXiv CS
arXiv CSopen access

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory Michael M. Craig, Riley J. Hickman, Yingshan Ma, Rémi Piché-Taillefer, Christine Allen, Pauric Bannigan1 1 Intrepid Labs, Toronto, Canada

arXiv:2609.19099v1 [cs.LG] 16 Sep 2026

Abstract Self-emulsifying drug delivery systems (SEDDS) can improve the oral bioavailability of poorly soluble drugs, but identifying high-performing formulations remains experimentally intensive. We present Andromeda 2, an agentic system that reasons over structured in-house experimental evidence and invokes computational and experimental tools to design and execute successive formulation batches. Using a miniaturized automated laboratory at a matched budget, we benchmark it against Andromeda 1, a probabilistic optimization model deployed across dozens of live development projects, and a wet-lab design-of-experiments (DoE) campaign. For paclitaxel, Andromeda 2 achieved a 50% high-performance hit rate versus 17% for Andromeda 1 and 2% for DoE, and identified 12 formulations meeting all four target product profile (TPP) objectives versus 6 and 0, respectively. Median AUC10–240 was 70.1, 12.0, and 3.5 mg·min/mL, while maximum AUC was comparable between Andromeda 2 and Andromeda 1. A selected full-TPP formulation achieved an apparent effective paclitaxel loading of 19 ± 5% w/w at the first FaSSIF measurement, approximately 3.3-fold higher than the 5.7% w/w loading reported for a published paclitaxel S-SEDDS. A controlled ablation showed that access to structured in-house experimental evidence increased mean AUC by 34%.

1

Introduction

(API), early proposals may therefore make limited use of the broader experimental knowledge accumulated across previous formulation programs. This motivates a different question: can an autonomous system reason over a laboratory’s existing experimental evidence and use that knowledge to allocate scarce wet-lab experiments more productively?

Self-emulsifying drug delivery systems (SEDDS) are an established strategy for improving the oral delivery of poorly soluble drugs by enhancing apparent solubility and promoting or sustaining supersaturation following gastrointestinal dilution [1, 2]. Their development, however, remains experimentally intensive, requiring simultaneous optimization Why oral paclitaxel matters of oil, surfactant, cosolvent, drug loading, Paclitaxel is a clinically important but excepand, where needed, precipitation inhibitors tionally challenging candidate for oral lipidwithin strict compositional and experimental based formulation. Its very poor aqueous solconstraints. Since exhaustive exploration ubility requires formulations that achieve therof such a large design space is impractical, apeutically relevant drug loading while mainconventional workflows typically concentrate taining the drug in a solubilized or supersaturated state following gastrointestinal dilution experimental effort within selected regions and minimizing precipitation. Oral exposure through excipient screening, phase diagram is further limited by biological barriers, includconstruction, and design-of-experiments ing intestinal P-glycoprotein efflux and first(DoE)-guided formulation campaigns [3, 4]. pass metabolism, although these are outside Sequential optimization methods can imthe scope of the present study [9]. Nevertheless, prove experimental efficiency by using the reDHP107, an oral lipid-based formulation of paclitaxel, has demonstrated noninferior efficacy sults of each batch to determine what should relative to intravenous paclitaxel in Phase III be tested next, and autonomous laboratories clinical evaluation [10, 11], establishing oral paincreasingly couple such algorithms to iterative design–make–test–learn workflows [5–8]. clitaxel as a clinically relevant formulation objective. However, these optimization strategies typically initialize each experimental campaign largely de novo, learning the formulationWe introduce Andromeda 2, an evidenceperformance landscape primarily through ex- grounded agentic system for autonomous forperiments conducted during that campaign. mulation design. The system reasons over For a new active pharmaceutical ingredient structured in-house experimental evidence to1

Agentic Formulation Development

gether with API, excipient, assay, and TPP jectives, so maximum AUC is not equivalent context; proposes and critiques candidate for- to a balanced formulation. Andromeda 2’s mulations; and invokes computational and ex- best formulation maintained apparent solubiperimental tools to construct executable for- lized paclitaxel through 240 min and reached mulation batches. Deterministic feasibility an AUC ∼18× the unformulated control drug checks enforce laboratory and compositional (Figure 2C). constraints before robotic execution. We evaluate Andromeda 2 against Andromeda 1 2.2 Structured in-house evidence and a physically executed DoE campaign usimproves Andromeda 2 ing the same formulation space, target prodperformance uct profile, assay, automated laboratory, and matched experimental budget. Andromeda To isolate the contribution of Intrepid’s 1 has been used across dozens of development structured in-house experimental evidence, we programs spanning diverse APIs and formu- compared full-evidence Andromeda 2 with lation modalities, making it an ideal point an otherwise matched reduced-evidence conof comparison for assessing the agentic sys- figuration on paclitaxel (Figure 3). Access to tem. A reduced-evidence ablation separately historical in-house experimental evidence was tests the contribution of accumulated in-house withheld in the reduced-evidence condition, experimental evidence. The central question while the formulation objective, executable deis whether evidence-grounded autonomous de- sign space, TPP, feasibility constraints, batch sign can use a fixed experimental budget size, and wet-lab workflow were held conmore effectively than conventional optimiza- stant. Both configurations continued to retion strategies, identifying high-performing for- ceive the paclitaxel measurements generated mulations more consistently. during their respective campaigns after each batch. Across the full campaign, mean AUC10–240 was 62.0 mg·min/mL with full evidence versus 2.1 Andromeda 2 improves 46.2 mg·min/mL with the evidence withheld, paclitaxel formulation corresponding to a 34% higher mean AUC. performance at a matched The high-AUC hit rate was similarly higher budget with full evidence at 50% versus 31%. The We compared Andromeda 2 with An- batch-resolved trajectories suggest that this dromeda 1 and a physically executed DoE difference was not simply due to a stronger inicampaign using the same miniaturized labora- tial batch.

2

Results

The reduced-evidence configuration tory and a matched budget of 96 unique formulations per strategy. Andromeda 2, An- matched or out-performed full-evidence dromeda 1, and experimental DoE attained Andromeda 2 in early batches but submedian AUC10–240 values of 70.1, 12.0, and sequently regressed, whereas full-evidence 3.5 mg·min/mL, respectively, and high-AUC Andromeda 2 improved across successive hit rates of 50%, 17%, and 2% (Figure 1; Ta- batches and sustained its performance. This ble 1; Figure 2A). The maximum observed trajectory is consistent with historical eviAUC was similar for Andromeda 2 and An- dence supporting more reliable interpretation dromeda 1 (157.7 vs 156.4 mg·min/mL). A and use of newly generated experimental descriptive formulation-level Mann–Whitney results, rather than simply providing an initial The reduced-evidence system comparison is reported in the Supplementary advantage. nevertheless continued to generate valid, exMaterial (Section A.1). Andromeda 2 produced 12 formulations perimentally executable SEDDS formulations, that met all four TPP criteria, compared with indicating that the ablation primarily affected 6 for Andromeda 1 and 0 for experimental candidate selection rather than formulation DoE (Figure 2B; Figure 5). The highest-AUC feasibility. formulation met only two of the four TPP ob2

Agentic Formulation Development 175

AUC10-240 (FaSSIF) (mg.min/mL)

150

125

100

75

med 70.1

50

25 med 12.0

API control med 3.5

0 ANDROMEDA 2

ANDROMEDA 1

Experimental DoE

Figure 1. Paclitaxel formulation performance at a matched experimental budget. Formulation-level distributions of the area under the curve (AUC) of apparently solubilized paclitaxel concentration in FaSSIF from 10 to 240 min achieved using the evidence-grounded agentic platform (Andromeda 2), the optimization framework (Andromeda 1) and a design of experiment (DoE) approach. Each strategy evaluated 96 unique formulations. Andromeda 2 produced an upward-shifted AUC distribution and a greater frequency of higher performing formulations, while both Andromeda platforms generated formulations with higher AUC values than experimental DoE. Table 1. Summary of paclitaxel formulation-selection performance at a matched 96-formulation experimental budget.

Arm

N

Best AUC

Median AUC

High-AUC hits n (%)

Formulations to first full TPP

Andromeda 2 Andromeda 1 Experimental DoE Reduced-evidence Andromeda 2

96 96 96 96

157.7 156.4 119.7 123.3

70.1 12.0 3.5 46.1

48 (50%) 16 (17%) 2 (2%) 30 (31%)

12 39 No pass 53

AUC denotes AUC10–240 in FaSSIF. High-AUC hits are formulations reaching the pooled upper-quartile AUC threshold. “Formulations to first full TPP” denotes the number of formulations screened before the first formulation satisfying all four TPP criteria was identified. Experimental DoE identified no full-TPP formulation within the 96formulation campaign.

2.3

Benchmarking Andromeda 2 against published oral paclitaxel formulations

measurement (10 min) was 19±5% w/w across three replicates. This is approximately 3.3-fold higher than the 5.7% w/w paclitaxel loading reported for the S-SEDDS formulation of Gao et al. [12], and more than an order of magnitude higher than the loadings reported for selected clinically evaluated lipid-based oral paclitaxel formulations. However, these values provide formulation context rather than a direct quantitative ranking because loading and measurement methods differ across studies.

Published oral paclitaxel formulations provide useful context for the formulations identified by Andromeda 2, although differences in formulation type, experimental endpoint, and study setting preclude direct head-tohead comparison. In particular, the FaSSIF AUC10–240 measured here should not be interpreted as oral bioavailability or systemic exposure. A compact literature comparison is reported in the Supplementary Material (Ta3 Discussion ble 2). For the selected formulation meeting all four This study shows that evidence-grounded TPP objectives, the apparent effective pacli- agentic design can substantially improve the taxel loading estimated from the first FaSSIF experimental efficiency of SEDDS formulation 3

Agentic Formulation Development A

B Cumulative formulations meeting all TPP criteria

Cumulative best AUC10 − 240 (mg⋅min/mL)

160 140 120 100 80 60 40 20

ANDROMEDA 2 ANDROMEDA 1 Experimental DoE Simulated DoE

API control

0 1

16

32

48

64

80

96

12

12

10

8

6

6

4

2

0

0 1

16

32

Formulations screened

48

64

80

96

Formulations screened

C ANDROMEDA 2 (AUC 158) ANDROMEDA 1 (AUC 156) Experimental DoE (AUC 120) API control (AUC 9)

Concentration (mg/mL)

0.8

0.6

0.4

0.2

0.0 0

50

100

150

200

250

FaSSIF time (min)

Figure 2. Matched-budget benchmark of paclitaxel formulation performance. All strategies were compared at an experimental budget of 96 formulations. (A) Cumulative-best area under the curve (AUC) of apparently solubilized paclitaxel concentration in FaSSIF from 10 to 240 min as a function of the number of formulations evaluated. Simulated DoE trajectories (grey dotted lines) and the unformulated-drug control (black dotted line) are included for reference. (B) Cumulative number of formulations satisfying all four prespecified target product profile (TPP) criteria. At 96 formulations, Andromeda 2 identified 12 formulations that met all TPP objectives, compared with 6 for Andromeda 1 and 0 for experimental DoE. (C) In vitro concentration–time profiles in FaSSIF for the highest AUC formulation from each experimental strategy and the unformulated-drug control. The highest-AUC Andromeda 2 formulation maintained a high apparently solubilized paclitaxel concentration through 240 min, whereas the highest AUC Andromeda 1 formulation reached a comparable peak concentration but declined after 120 min, consistent with precipitation.

development. At a matched budget of 96 pacli- is the coupling of autonomous experimentataxel formulations per strategy, Andromeda tion with evidence-grounded agentic decision2 achieved a 50% high-AUC hit rate, compared making to identify high-performing formulawith 17% for Andromeda 1 and 2% for the tions more frequently within a fixed wet-lab physically executed DoE campaign, and iden- budget than either probabilistic optimization tified 12 formulations meeting all four TPP or DoE. criteria, compared with 6 and 0, respectively. The evidence-ablation and composition analNotably, Andromeda 2 and Andromeda 1 yses provide complementary insight into how reached comparable maximum AUC values, in- the Andromeda 2 advantage emerged. Acdicating that the agentic system’s main advan- cess to structured in-house experimental evitage was not identifying a higher isolated opti- dence increased mean paclitaxel AUC by 34% mum, but producing substantially more high- and the high-AUC hit rate from 31% to 50%. performing formulations across the experimen- In addition, Andromeda 2 achieved a 49% hit tal campaign. These findings extend a growing rate within the high-performing Type IIIB/IV literature on self-driving laboratories and data- region, compared with 29% for Andromeda driven optimization in chemistry and materi- 1, indicating that its advantage was not exals science [5–8], as well as machine-learning plained solely by identifying a favourable reapproaches to drug formulation design [13– gion of composition space. More importantly, 15]. Here, the distinguishing contribution the reduced-evidence configuration remained 4

Agentic Formulation Development Full-evidence mean 62 vs 46 mg⋅min/mL (34% higher with in-house evidence)

ANDROMEDA 2 (full evidence) ANDROMEDA 2 (evidence removed)

160 140

AUC10 − 240 (mg⋅min/mL)

120 100 80 60 40 20 0 1

2

3 4 Wet-lab batch

5

6

Figure 3. Effect of in-house evidence on Andromeda 2 performance across paclitaxel batches. Per-formulation area under the curve of apparently solubilized paclitaxel concentration in FaSSIF from 10 to 240 min is shown across six sequential wet-lab batches for Andromeda 2 with access to its full structured in-house evidence and for the reduced-evidence ablation in which this evidence was withheld. Both configurations received the experimental results generated during their respective campaigns. Points represent individual formulations, and lines connect batch means. The full-evidence configuration achieved a higher campaign-wide mean AUC than the reduced-evidence configuration (62.0 versus 46.2 mg·min/mL; 34% higher), showed more consistent improvement across batches, and maintained its gains in later batches. The reduced- evidence configuration matched or exceeded the full-evidence configuration in some early batches but did not maintain this performance.

multiple lead candidate formulations rather than a single isolated composition. Vitamin E TPGS was prominent among these high-performing compositions, consistent with its established paclitaxel-solubilizing properties. Vitamin E TPGS also exhibits independent effects on P-glycoprotein-mediated paclitaxel transport, which may prove advantageous in future in vivo studies. Since intestinal transport was not measured in the present study, these results support only the physicochemical performance of the formulations in FaSSIF; nevertheless, identifying formulation chemistry relevant to both paclitaxel solubilization and a known biological barrier to oral delivery is encouraging. The apparent effective paclitaxel loading of formulations meeting all four TPP criteria also provides a useful point of comparison with previously reported oral lipid-based paclitaxel systems (Table 2), although differences in experimental endpoints preclude direct ranking.

competitive in early batches but subsequently regressed, whereas full-evidence Andromeda 2 improved across successive batches and sustained its performance. Both configurations received the experimental results generated during their respective campaigns; the divergent trajectories are therefore consistent with historical evidence supporting more effective interpretation and use of newly generated measurements, rather than simply strengthening the initial proposal. The ablation does not by itself attribute the full Andromeda 2– Andromeda 1 performance difference to historical evidence, but demonstrates that access to this evidence materially contributes to Andromeda 2 performance. The reducedevidence system continued to generate valid, experimentally executable formulations, further suggesting that the principal effect of the ablation was on candidate selection rather than formulation feasibility. This finding is consistent with the evidence-ablation effect we previously observed for clofazimine [16]. From a formulation science perspective, Andromeda 2 concentrated experimental effort within high-performing regions of the formulation design space while identifying

Several limitations define the scope of these findings. The FaSSIF assay measures the ability of a formulation to maintain paclitaxel in an apparently solubilized state and does not capture digestion, intestinal permeability or 5

Agentic Formulation Development

state simulated intestinal fluid (FaSSIF) [17]. Apparent solubilized-drug concentrations were quantified by HPLC analysis of sampled aliquots at 10, 30, 60, 120 and 240 min after the intestinal-fluid transfer. The reported endpoint was the apparent concentration of solubilized drug in the FaSSIF stage. This endpoint captures the ability of a formulation to disperse and maintain drug in an apparently solubilized state under biorelevant conditions. It was used as the principal high-throughput measure of formulation performance and not interpreted as a surrogate for absorbed dose or oral bioavailability. Drug loading was specified as a formulationdesign input rather than measured in the undiluted preconcentrate. An apparent effective loading was therefore estimated retrospectively by multiplying the nominal loading by the maximum fraction of the nominal API dose observed in the dispersed phase by HPLC analysis. This operational estimate is used only for comparison with literature formulations and is not a direct measurement of preconcentrate drug content.

metabolism, or in vivo performance; digestioncoupled testing and ultimately pharmacokinetic studies will therefore be required to determine whether the observed formulation’s advantages translate to oral exposure. In addition, each strategy was evaluated in a single independently initialized campaign, limiting campaign-level statistical inference and motivating replicate autonomous campaigns to establish reproducibility. The experimental DoE comparator also explored the shared composition space directly, without the expert prescreening and prioritization that may accompany conventional formulation development, and should be interpreted in that context. Finally, prospective evaluation across chemically distinct APIs is needed to establish the generality of the evidence-grounded advantage. Together, these studies would test whether the improved allocation of experimental effort observed here persists across campaigns, APIs, and increasingly translational endpoints. More broadly, these findings point toward a model of formulation development in which experimental data are not consumed within individual programs but accumulated as reusable evidence that improves subsequent autonomous decision-making. If this advantage generalizes across APIs and formulation modalities, evidence-grounded autonomous laboratories could make formulation development progressively more efficient as institutional experimental knowledge grows.

4

Methods

4.1

Autonomous SEDDS laboratory

4.2

Andromeda 1: probabilistic optimization

Andromeda 1 is a sequential probabilisticoptimization platform and the non-agentic comparator in this study. Candidate formulations are drawn from a predefined, constrained composition space and evaluated against the same TPP used by Andromeda 2. At each iteration, probabilistic models relate composition to measured outcomes and guide selection of the next 16-formulation batch, balancing promising regions of design space against formulations whose performance remains uncertain. Newly measured dissolution, droplet size, PDI and HPLC data are incorporated before the next cycle. The matched paclitaxel campaign comprised six batches (96 formulations).

All campaigns used the same miniaturized design–build–test platform. Each wet-lab batch comprised 16 SEDDS preconcentrates prepared by liquid-handling robots and evaluated by automated dispersion/dissolution, dynamic light scattering (droplet size and polydispersity index [PDI]), HPLC drug quantification, and structured data capture. All experimental arms used the same operators, instruments, assay timing, data-processing workflow, and dissolution protocol. Formulations were evaluated in a two-stage biorelevant dissolution assay: an initial gastricfluid stage followed by transfer into fasted-

4.3

Andromeda 2: agentic formulation design

Andromeda 2 is an evidence-grounded agentic system that orchestrates models and other computational capabilities as tools (Figure 4). Proposal, critique, review, planning, 6

Agentic Formulation Development

4.4

and constraint-checking operate over a shared formulation context. The system can invoke the probabilistic models used by Andromeda 1, together with additional models, to construct 16-formulation batches with scientific rationale and platform-compatible composition records.

Evidence grounding and ablation

A key distinction between the two Andromeda versions is the evidence available at the initiation of an experimental campaign. Andromeda 1 does not have access to historical formulation data. Instead, each campaign begins with an initial design batch, afA deterministic feasibility engine then enter which the planner learns the formulation– forces the hard constraints required for wetperformance landscape from measurements lab execution, including valid excipient segenerated during that campaign and uses lections, compatibility with the permitted these observations to guide subsequent batches composition grid, sum-to-one compositional within a probabilistic-optimization framework. constraints, uniqueness relative to previously In contrast, Andromeda 2 is grounded from tested formulations, intra-batch distinctivethe outset in Intrepid’s structured in-house forness, and robotic readiness. This separation mulation evidence, comprising prior composidistinguishes scientific search from experimentions, measured dissolution, dispersion, size tal implementation: the agentic layer directs and PDI outcomes, excipient behaviour, and exploration, while the engineering layer enplatform-specific measurement history. Imporsures that submitted compositions are exetantly, this historical evidence contained no pacutable. clitaxel formulations, such that Andromeda A formulation scientist reviewed each pro- 2 could draw on broader formulation knowlposed batch before execution as a safety and edge without access to prior experimental refeasibility check that the batch was scien- sults for the specific API evaluated here. It tifically reasonable and experimentally exe- can reason over this accumulated evidence durcutable. This was not a redesign step: the ing proposal, critique, planning, and tool use, reviewer did not author, substitute, or re-tune while also incorporating observations generated during the ongoing campaign. the proposed compositions. To evaluate transfer to previously unseen APIs, all historical formulation data for the With no measurements yet available from the current campaign, the first batch explored target API were excluded at the initiation of several first-principles hypotheses about which each campaign. The systems were not trained excipient combinations could satisfy the TPP. on in-house experimental formulation data for Subsequent batches received all preceding mea- paclitaxel. After each wet-lab batch, newly surements, and the system preferentially ex- generated target-API measurements were replored perturbations of high-performing for- turned to the system to guide subsequent mulations while retaining limited broader ex- batch selection. The campaigns therefore asploration. No model retraining occurred sess whether knowledge accumulated across Inbetween batches; sequential adaptation was trepid’s broader laboratory experience can supachieved by updating the agent’s working con- port formulation design for a new API, foltext with newly generated experimental evi- lowed by sequential adaptation to measurements generated during the campaign. dence. To isolate the contribution of this accuBoth platforms are fully in-house. Language mulated evidence, paclitaxel was also run in models, probabilistic models and supporting a reduced-evidence Andromeda 2 configuramodels are developed, trained and operated tion that withheld historical in-house experon Intrepid infrastructure, using Intrepid ex- imental evidence while holding constant the perimental data exclusively. Neither platform formulation objective, TPP, executable dedepends on external commercial model appli- sign space, feasibility constraints, batch size cation programming interfaces. and wet-lab workflow. Both configurations received the paclitaxel measurements generated 7

Agentic Formulation Development BATCH GENERATION WORKFLOW

SYSTEM PROMPT

Paclitaxel • example formulation • AUC₁₀–₂₄₀ = 94.2

SCIENCE

DECISION

what informs formulation proposals what it optimizes and verifies

EXECUTION

GENERATE A BATCH

how proposals are generated

SEDDS principles

TPP objectives

Agent role

Surfactant HLB; oil reservoir capacity; cosolvent leaching; drug loading; criteria for identifying failed emulsions.

AUC ≥ target; C₂₄₀ ≥ threshold; f₂₄₀ ≥ floor; C₂₄₀/C₃₀ ≥ minimum. Rank valid formulations only.

SEDDS formulation scientist, responsible for every proposed formulation.

Design space

Verification tools

Search strategy

Allowed components: soybean, sesame or Captex oil; TPGS; PEG 400 or Phosal. Compositions use 0.05 increments and sum to 1.

0.05 composition increments; mass balance; allowed components; novelty and duplicate checks.

Batch-level search strategy: exploit valid winners, adjust 1–2 levers, retain informative boundary probes.

Cross-API evidence

Measured memory

Output requirements

Historical non-paclitaxel SEDDS formulations, provided only as in-context examples.

This campaign: measured paclitaxel formulations from earlier batches and failed emulsions excluded.

Structured formulation record for each proposal; formulationspecific rationale.

1 GENERATE – exact model rationale “Neighbor of S32 (drug 0.20, soybean_oil 0.25, peg_400 0.05, vit_e_tpgs 0.50): ~ changed oil from 0.25 to 0.35 (+0.10) and surf from 0.50 to 0.40 (-0.10) while keeping drug and peg_sum 0.20+0.35+0.05+0.40=1.00; pushes oil level toward upper bound while fixing peg; hypothesis: higher oil may increase AUC but risks size; expects higher late AUC, size possibly larger, PDI similar; unique formulation.”

2 VERIFY

3 ASSAY

All checks passed

FaSSIF 10–240 min Dissolution + dispersion

grid • mass balance novelty • duplicates

4 TPP ATTRIBUTES

4/4

soybean 35% • TPGS 40% • PEG 400 5% • drug 20% size 198 nm • PDI 0.072

AUC

C₂₄₀

f₂₄₀

C₂₄₀/C₃₀

≥ target

≥ threshold

≥ floor

≥ minimum

KEY TO ACRONYMS SEDDS = self-emulsifying drug delivery system

TPP

target product profile

C₂₄₀

= concentration at 240 min (mg/mL)

HLB

= hydrophilic-lipophilic balance

AUC₁₀–₂₄₀ = area under dissolution curve

f₂₄₀

= dose fraction dissolved at 240 min

TPGS

= d-α-tocopheryl polyethylene glycol 1000 succinate

=

from 10–240 min (FaSSIF; mg·min/mL)

C₂₄₀/C₃₀

= precipitation resistance

PEG 400= polyethylene glycol 400

Figure 4. System prompt structure and representative batch-generation workflow for Andromeda 2. The system prompt organizes formulation knowledge, optimization objectives, and execution requirements into science, decision, and execution components. During batch generation, Andromeda 2 uses this context to propose candidate formulations with model-generated rationales, verifies compositional and feasibility constraints, and submits valid formulations for experimental evaluation. Measured formulation performance is then evaluated against the predefined target product profile (TPP) attributes, including dissolution AUC, late-time concentration, fraction dissolved, and precipitation resistance. The example shown illustrates a representative formulation from a generated batch rather than the complete batch-level output.

during their respective campaigns after each Simulated non-adaptive and sequential batch. adaptive DoE strategies were additionally evaluated against a validated in-silico emulator of the paclitaxel formulation landscape. Emula4.5 Comparators and paclitaxel tor validation and those results are reported design task in the Supplementary Material (Section A.4, Paclitaxel was selected as a hard-API Table 3). stress test (high molecular weight and very low aqueous solubility) and evaluated with 4.6 Metrics and statistical analysis Andromeda 2, Andromeda 1, reducedevidence Andromeda 2, and a physically exThe primary endpoint is trapezoidal ecuted experimental DoE approach. Each AUC10–240 in FaSSIF. Cumulative-best AUC arm received a matched budget of six 16- versus formulations screened describes how formulation batches (96 unique formulations). rapidly high-performing formulations appear. The experimental DoE was designed by a formulation scientist using commercial statistical software supporting custom and responsesurface designs over the same composition parameters and hard constraints, and executed end-to-end on the same platform and FaSSIF assay. It explored the shared composition grid directly rather than beginning from a separate excipient solubility pre-screen.

Mean and median AUC describe the distribution of performance, which is distinct from peak AUC: a method may identify a single exceptional formulation while producing substantially lower typical candidates. A highAUC hit is a formulation whose AUC10–240 reaches the pooled upper quartile of all measured paclitaxel formulations. Because that threshold is relative, comparisons are 8

Agentic Formulation Development

also reported against the pool-independent TPP. Formulation usability is summarized with a four-attribute TPP (Figure 5): a minimum dissolution AUC, a late-time concentration floor C240 , a late-time fraction-dissolved floor f240 (which normalizes for drug loading), and a precipitation-resistance ratio C240 /C30 ≥ 0.85. A formulation meets the full TPP only if all four objectives are satisfied. Thresholds are API-specific (Figure 5B). Peak AUC, the number of objectives met, and full-TPP passage are not interchangeable. Pairwise distributional comparisons use twosided Mann–Whitney U tests. Hit rates are reported with 95% confidence intervals calculated using percentile bootstrapping. We lead with effect sizes (medians, hit rates, and counts of formulations meeting all TPP objectives). All formulation-level tests and bootstrap intervals are treated as descriptive because observations within adaptive campaigns are not statistically independent; no campaign-level inferential testing was performed.

9

Agentic Formulation Development

1.0

A

B

Concentration (illustrative)

2/3: late-time level & dose fraction (C240, f240 floors)

0.8 Objective

0.6

0.4 1: AUC10 − 240 (area under the curve)

0.2

4: precipitation ratio C240/C30 ≥ 0.85

Ensures

Paclitaxel

1 AUC10 − 240

Solubilization

≥8

2 C240

Sustained level

≥ 0.04

3 f240

Dose fraction dissolved

≥ 10%

4 C240/C30

Precipitation resistance

≥ 0.85

AUC in mg⋅min/mL; C in mg/mL. Full-TPP pass = all four criteria.

Passes TPP (sustained) Fails TPP (precipitation)

0.0

0

50

100 150 FaSSIF time (min)

200

250

Figure 5. Target product profile (TPP) definition. (A) An illustrative formulation that meets the TPP (sustained) and one that does not (precipitating), annotated with the four TPP objectives. (B) Acceptance objectives for paclitaxel. A formulation meets the full TPP only if all four objectives are satisfied.

10

Agentic Formulation Development

References [1] Colin W. Pouton. Formulation of poorly water-soluble drugs for oral administration: physicochemical and physiological issues and the lipid formulation classification system. European Journal of Pharmaceutical Sciences, 29(3–4):278– 287, 2006. doi: 10.1016/j.ejps.2006.04. 016. URL https://doi.org/10.1016/j. ejps.2006.04.016. [2] Christopher J. H. Porter, Natalie L. Trevaskis, and William N. Charman. Lipids and lipid-based formulations: optimizing the oral delivery of lipophilic drugs. Nature Reviews Drug Discovery, 6(3):231– 248, 2007. doi: 10.1038/nrd2197. URL https://doi.org/10.1038/nrd2197. [3] Hywel D. Williams, Natalie L. Trevaskis, Susan A. Charman, Ravi M. Shanker, William N. Charman, Colin W. Pouton, and Christopher J. H. Porter. Strategies to address low drug solubility in discovery and development. Pharmacological Reviews, 65(1):315–499, 2013. doi: 10.1124/pr.112.005660. URL https:// doi.org/10.1124/pr.112.005660. [4] Huiling Mu, Rene Holm, and Anette Müllertz. Lipid-based formulations for oral administration of poorly watersoluble drugs. International Journal of Pharmaceutics, 453(1):215–224, 2013. doi: 10.1016/j.ijpharm.2013.03. 054. URL https://doi.org/10.1016/j. ijpharm.2013.03.054.

as a tool for chemical synthesis. Nature, 590(7844):89–96, 2021. doi: 10.1038/ s41586-021-03213-y. URL https://doi. org/10.1038/s41586-021-03213-y. [7] Benjamin P. MacLeod, Fraser G. L. Parlane, Thomas D. Morrissey, Florian Häse, Loı̈c M. Roch, Kevan E. Dettelbach, Raphaell Moreira, Lars P. E. Yunker, Michael B. Rooney, Joseph R. Deeth, Veronica Lai, Gary J. Ng, Henry Situ, Ruihan H. Zhang, Michael S. Elliott, Ted H. Haley, David J. Dvorak, Alán Aspuru-Guzik, Jason E. Hein, and Curtis P. Berlinguette. Self-driving laboratory for accelerated discovery of thinfilm materials. Science Advances, 6(20): eaaz8867, 2020. doi: 10.1126/sciadv. aaz8867. URL https://doi.org/10. 1126/sciadv.aaz8867. [8] Benjamin Burger, Phillip M. Maffettone, Vladimir V. Gusev, Catherine M. Aitchison, Yang Bai, Xiaoyan Wang, Xiaobo Li, Ben M. Alston, Buyi Li, Rob Clowes, Nicola Rankin, Brandon Harris, Reiner Sebastian Sprick, and Andrew I. Cooper. A mobile robotic chemist. Nature, 583(7815):237–241, 2020. doi: 10. 1038/s41586-020-2442-2. URL https:// doi.org/10.1038/s41586-020-2442-2. [9] Alex Sparreboom, Judith van Asperen, Ulrich Mayer, Alfred H. Schinkel, Johan W. Smit, Dirk K. F. Meijer, Piet Borst, Willem J. Nooijen, Jos H. Beijnen, and Olaf van Tellingen. Limited oral bioavailability and active epithelial excretion of paclitaxel (Taxol) caused by Pglycoprotein in the intestine. Proceedings of the National Academy of Sciences, 94 (5):2031–2035, 1997. doi: 10.1073/pnas. 94.5.2031. URL https://doi.org/10. 1073/pnas.94.5.2031.

[5] Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P. Adams, and Nando de Freitas. Taking the human out of the loop: a review of Bayesian optimization. Proceedings of the IEEE, 104(1):148–175, 2016. doi: 10.1109/JPROC.2015.2494218. URL https://doi.org/10.1109/JPROC. [10] Y.-K. Kang, M.-H. Ryu, S. H. Park, J. G. Kim, J. W. Kim, S.-H. Cho, Y.2015.2494218. I. Park, S. R. Park, S. Y. Rha, M. J. Kang, J. Y. Cho, S. Y. Kang, S. Y. [6] Benjamin J. Shields, Jason Stevens, Jun Roh, B.-Y. Ryoo, B.-H. Nam, Y.-W. Li, Marvin Parasram, Farhan Damani, Jo, K.-E. Yoon, and S. C. Oh. EffiJesus I. Martı́nez Alvarado, Jacob M. cacy and safety findings from DREAM: Janey, Ryan P. Adams, and Abigail G. a phase III study of DHP107 (oral paDoyle. Bayesian reaction optimization 11

Agentic Formulation Development

clitaxel) versus i.v. paclitaxel in patients with advanced gastric cancer after failure of first-line chemotherapy. Annals of Oncology, 29(5):1220–1226, 2018. doi: 10.1093/annonc/mdy055. URL https:// doi.org/10.1093/annonc/mdy055.

Natsuda Navamajiti, Apolonia Gardner, Rosanna M. Zhang, Tina Esfandiary, Johanna L’Heureux, Thomas von Erlach, Elena M. Smekalova, Dominique Leboeuf, Kaitlyn Hess, Aaron Lopes, Jaimie Rogner, Joy Collins, Siddartha M. Tamang, Keiko Ishida, Paul Chamberlain, DongSoo Yun, Abigail Lytton-Jean, Christian K. Soule, Jaime H. Cheah, Alison M. Hayward, Robert Langer, and Giovanni Traverso. Computationally guided high-throughput design of self-assembling drug nanoparticles. Nature Nanotechnology, 16(6):725–733, 2021. doi: 10.1038/ s41565-021-00870-y. URL https://doi. org/10.1038/s41565-021-00870-y.

[11] B. Xu, H. Jeong, T. Sun, J. Sohn, Q. Zhang, S. Kostic, X. Wang, Z. Tong, S. Wang, J. Wang, W. Li, K. S. Lee, Y. W. Moon, M. J. Kang, X. Hu, T. Y. Kim, D. Milenković, J. H. Seo, J. H. Kim, J. Lee, and S.-B. Kim. OPTIMAL: a multinational phase III study of oral paclitaxel (DHP107) versus intravenous weekly paclitaxel in HER2negative recurrent or metastatic breast cancer. Annals of Oncology, 37(7):936– [16] Michael Craig, Pauric Bannigan, Gary 945, 2026. doi: 10.1016/j.annonc.2026.03. Tom, Yingshan Ma, Remi Piche-Taillefer, 002. URL https://doi.org/10.1016/j. Christine Allen, and Riley Hickman. annonc.2026.03.002. Agentic systems for sample-efficient drug formulation design. In ICML 2026 [12] Ping Gao, Bobby D. Rush, William P. Workshop on AI for Science, 2026. Pfund, Tiehua Huang, Juliane M. Bauer, URL https://openreview.net/forum? Walter Morozowich, Ming-Shang Kuo, id=SyPcOsoGgp. Oral presentation. and Michael J. Hageman. Development of a supersaturable SEDDS (S-SEDDS) [17] E. Galia, E. Nicolaides, D. Hörter, formulation of paclitaxel with improved R. Löbenberg, C. Reppas, and J. B. Dressoral bioavailability. Journal of Pharmaman. Evaluation of various dissolution ceutical Sciences, 92(12):2386–2398, 2003. media for predicting in vivo performance doi: 10.1002/jps.10511. URL https:// of class I and II drugs. Pharmaceutidoi.org/10.1002/jps.10511. cal Research, 15(5):698–705, 1998. doi:

10.1023/A:1011910801212. URL https: [13] Pauric Bannigan, Matteo Aldeghi, Ze//doi.org/10.1023/A:1011910801212. qing Bao, Florian Häse, Alán AspuruGuzik, and Christine Allen. Machine [18] S. A. Veltkamp, B. Thijssen, J. S. Garlearning directed drug formulation develrigue, G. Lambert, F. Lallemand, F. Binopment. Advanced Drug Delivery Reviews, lich, A. D. R. Huitema, B. Nuijen, A. Nol, 175:113806, 2021. doi: 10.1016/j.addr. J. H. Beijnen, and J. H. M. Schellens. 2021.05.016. URL https://doi.org/10. A novel self-microemulsifying formulation 1016/j.addr.2021.05.016. of paclitaxel for oral administration to patients with advanced cancer. British [14] Pauric Bannigan, Zeqing Bao, Riley J. Journal of Cancer, 95(6):729–734, 2006. Hickman, Matteo Aldeghi, Florian Häse, doi: 10.1038/sj.bjc.6603312. URL https: Alán Aspuru-Guzik, and Christine Allen. //doi.org/10.1038/sj.bjc.6603312. Machine learning models to accelerate the design of polymeric long-acting in[19] Jung Wan Hong, In-Hyun Lee, jectables. Nature Communications, 14:35, Young Hak Kwak, Young Taek Park, 2023. doi: 10.1038/s41467-022-35343-w. Ha Chin Sung, Ick Chan Kwon, and URL https://doi.org/10.1038/ Hesson Chung. Efficacy and tissue s41467-022-35343-w. distribution of DHP107, an oral paclitaxel formulation. Molecular Cancer [15] Daniel Reker, Yulia Rybakova, Ameya R. Therapeutics, 6(12):3239–3247, 2007. doi: Kirtane, Ruonan Cao, Jee Won Yang, 12

Agentic Formulation Development

10.1158/1535-7163.MCT-07-0261. URL https://doi.org/10.1158/1535-7163. MCT-07-0261.

13

Agentic Formulation Development

A

Supplementary Material

A.1

Descriptive Mann–Whitney comparison of Andromeda 2 and Andromeda 1

A formulation-level Mann–Whitney comparison between Andromeda 2 and Andromeda 1 yielded p = 9.6 × 10−8 . This value is treated as descriptive rather than inferential because observations within each adaptive campaign are not independent: selection of later formulations depends on the measured performance of formulations tested in earlier batches. The nominal sample size of 96 formulations per strategy therefore overstates the effective sample size. Interpretation of the matched-budget result in the main text therefore rests primarily on the magnitude of the distributional shift and the observed hit rates.

A.2

Formulation composition and the LFCS taxonomy

Formulation development decisions depend not only on the final formulations identified but also on how the method employed allocates its experimental budget across composition space. To characterize these exploration patterns, we classified each formulation using the Lipid Formulation Classification System (LFCS), a widely used compositional framework for lipid-based formulations [1]. The LFCS distinguishes formulations according to their relative proportions of oil, water-insoluble surfactant (typically HLB less than 12), water-soluble surfactant (typically HLB greater than 12), and hydrophilic cosolvent (Figure 6A). Within the design space, formulations were classified as Type II, comprising oil and waterinsoluble surfactant without water-soluble components; Type IIIA, containing substantial oil together with water-soluble surfactant and optionally cosolvent; Type IIIB, containing less oil and a greater proportion of water-soluble components; or Type IV, an oil-free system based predominantly on water-soluble surfactants and cosolvents (Figure 6A). No Type I oil-only formulations were included in the evaluated design space. For each method, we report the distribution of formulations across LFCS types, the hit rate within the productive region, and the evolution of hit rate across experimental batches. Across the tested paclitaxel formulations, the high-performance hit rate varied markedly by LFCS type: 0% for Type II, 6% for Type IIIA, 35% for Type IIIB, and 32% for Type IV (Figure 6B). Thus, within the evaluated design space, high-performing formulations were concentrated in the Type IIIB/IV region, characterized by water-soluble surfactants and low or no oil content. Andromeda 2 directed most of its experimental budget toward this region and largely avoided the unproductive Type II systems, whereas the experimental DoE allocated a substantial fraction of its budget to Type II formulations. The advantage of Andromeda 2 was not attributable solely to identifying the productive LFCS region. Among Type IIIB/IV formulations, Andromeda 2 achieved a 49% hit rate (bootstrap 95% CI [0.39, 0.59]), compared with 29% for Andromeda 1 and 3% for experimental DoE (Figure 6C). Andromeda 2 therefore both concentrated its search within the most productive region and selected higher-performing candidates within that region. The composition trajectory also demonstrated adaptation over the course of the campaign (Figure 6D). Andromeda 2 began with broad chemical exploration sampling 12 distinct oilsurfactant families in its first batch. In subsequent batches it increasingly concentrated on productive Type IIIB/IV formulations as experimental measurements accumulated and performulation AUC increased. DoE showed little corresponding improvement across batches. Although Andromeda 2 sampled fewer distinct oil-surfactant families overall than experimental DoE (16 versus 90), it generated substantially more high-performing formulations. This focused search did not collapse onto a single composition family: Andromeda 2 retained several productive families, providing chemically distinct alternatives for subsequent formulation development.

14

Agentic Formulation Development

18

A

B IV (oil-free)

IIIB (oil-lean)

IIIA (oil-rich)

40

16

35% (n=207) 32% (n=62)

35

High-AUC hit rate (%)

Surfactant HLB

14

12

10 Type II (water-insoluble)

8

30 25 20 15 10

6% (n=71)

6 5

ANDROMEDA 2 ANDROMEDA 1 Experimental DoE

4 0.0

0.2

0.4

0.6

0.8

0% (n=44)

0

1.0

Type II

Type IIIA

Type IIIB

Type IV

Oil fraction of excipients

C

D

1.0

160

ANDROMEDA 2 ANDROMEDA 1 Experimental DoE

120

AUC10− − 240 (mg⋅min/mL)

Hit rate in productive region (IIIB/IV)

140 0.8 49% n=95

0.6 29% n=55

0.4

0.2

3% n=36

100 80 60 40 20 0

0.0

ANDROMEDA 2

ANDROMEDA 1

Experimental DoE

1

2

3

4

5

6

Wet-lab batch

Figure 6. Formulation composition using the LFCS taxonomy (paclitaxel). (A) Composition space (oil fraction versus surfactant HLB) with LFCS Type regions shaded and formulations coloured by method. (B) High-AUC hit rate by LFCS Type; the water-soluble-surfactant Type IIIB/IV are productive. (C) Hit rate within the productive region by method, with bootstrap 95% confidence intervals. (D) Batch progression: each point is a formulation (batch on the x-axis, AUC on the y-axis), with per-method batch medians.

15

Agentic Formulation Development

A.3

Literature comparison of oral paclitaxel formulations

The table below is offered as context rather than as a ranking of formulations. Endpoints differ across studies, and the FaSSIF AUC used in the main text should not be equated with oral bioavailability or systemic exposure. Table 2. Selected published lipid-based oral paclitaxel formulations compared with Andromeda 2. Endpoints differ across studies and are shown for context rather than as direct head-to-head comparisons. Formulation

PTX w/w)

Andromeda 2

Up to 25 (nominal)a

S-

5.7

Veltkamp et al. [18], SMEOF#3 DHP107 [11, 19]

1.6c

Gao et al. [12], SEDDS

loading

(%

Formulation / translational result 12/96 full-TPP passes; high apparent solubilized PTX maintained in FaSSIF through 240 min. Precipitation-inhibiting polymer improved maintenance of PTX after dilution; subsequent rat PK demonstrated oral exposure. Clinical oral PTX formulation; evaluated in cancer patients with cyclosporine A. Clinically validated lipid-based oral PTX formulationb ; Phase III development without co-administered P-gp inhibitor.

∼1c

a

Nominal drug loading is a formulation-design input in the present study and should not be interpreted as directly measured drug content or as effective solubilized loading. b DHP107 is included as a translational lipid-based oral paclitaxel benchmark and should not be described as a conventional SEDDS. c Published loadings that were not originally reported as % w/w are converted here for comparison. Gao et al. reported 57 mg/g paclitaxel, which is 5.7% w/w. Veltkamp et al. reported 1.6% w/v (160 mg in 10 mL); their composition table sums to 100 g per 100 mL, so 1.6% w/v is numerically equivalent to 1.6% w/w under that accounting. DHP107 is reported as 10 mg/mL (approximately 1% w/v); the % w/w value assumes a formulation density of ∼1 g/mL and should be treated as approximate.

A.4

In silico emulator validation

The simulated DoE strategies shown in Figure 2A were evaluated using an in-silico emulator of the paclitaxel formulation landscape trained on in-house data. On held-out formulations the emulator recovered the measured performance ordering (Spearman 0.89) and enriched topquartile formulations 3.02-fold relative to random selection, against a ceiling of 4.0 for perfect selection (Table 3). Under these conditions, neither non-adaptive nor sequential adaptive DoE reached the upper-tail AUC values measured for the Andromeda strategies. This is consistent with the experimental DoE campaign, which likewise identified no formulations in that range. These simulated arms are reported as directional context rather than as a matched comparator. The emulator was validated on random held-out splits; its ranking accuracy is lower for oil–surfactant composition families not represented in that data, and the simulated result should be weighted accordingly. Table 3. In silico emulator validation for paclitaxel: held-out ranking skill, error, and top-quartile enrichment. MAE in mg·min/mL. Top-quartile enrichment is relative to a chance baseline of 1.0 (ceiling 4.0 for perfect selection).

API Paclitaxel

Test Spearman

Test MAE

Top-Q enrich.

0.89

12.5

3.02

16

Record · ID 965402 · SHA-256 1e1b376903cb511e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.