ConceptioArchivearXiv CS
arXiv CSopen access

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

A TWO-YEAR COMMUNITY ROADMAP

arXiv:2607.12113v1 [cs.DC] 13 Jul 2026

Toward Trustworthy Autonomous Science

Technical Report ORNL/TM-2026/4663 · July 2026 Licensed under CC BY 4.0

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP Disclaimer. This report describes products of research sponsored by the U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research under Contract No. DE-SCL0000175, “A Testbed for Multi-Agent Autonomous Science: From Lab Bench to Supercomputer”. This report was prepared as an account of work sponsored by agencies of the United States Government. Neither the United States Government nor any agency thereof, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. License. This report is made available under a Creative Commons Attribution 4.0 International Public license (https:// creativecommons.org/licenses/by/4.0).

Preferred citation R. Ferreira da Silva, M. Abolhasani, P. Beaucage, L. Biven, M. Bussmann, K. Chard, R. Coffee, S. DeWitt, S. Dolas, C. Eckert, D. Elbert, I. T. Foster, T. Ghosal, A. Giannakou, T. Gibbs, L. Hamilton, G. Lockwood, T. Mayer, B. Mintz, R. Nazikian, S. Nimer, A. Randles, W. Shin, S.R. Sukumar, F. Suter, M. Taheri, M. Taufer, D. Vrabie, “Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap”, Technical Report, ORNL/TM-2026/4663, July 2026. @techreport{autonomousscience2026roadmap, author = {Ferreira da Silva, Rafael and Abolhasani, Milad and and Beaucage, Peter and Biven, Laura and Bussmann, Michael and Chard, Kyle and Coffee, Ryan and DeWitt, Stephen and Dolas, Sagar and Eckert, C. and Elbert, David and Foster, Ian T. and Ghosal, Tirthankar and Giannakou, A. and Gibbs, Tom and Hamilton, Leslie and Lockwood, Glenn and Mayer, Theresa and Mintz, Benjamin and Nazikian, Raffi and Nimer, Salahudin and Randles, Amanda and Shin, Woong and Sukumar, Sreenivas Rangan and Suter, Fr\’ed\’eric and Taheri, Mitra and Taufer, Michela and Vrabie, Draguna}, title = {{Toward Trustworthy Autonomous Science: A Two−Year Community Roadmap}}, year = {2026}, number = {ORNL/TM−2026/4663}, institution = {Oak Ridge National Laboratory} }

2

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

Executive Summary

A

YEAR AGO, we argued that autonomous laboratories remained isolated islands, and we proposed a

grassroots network, the Autonomous Interconnected Science Lab Ecosystem (AISLE), organized around five critical dimensions for connecting them. The field has since moved faster than that roadmap anticipated. Multi-agent systems have produced experimentally validated hypotheses; self-driving laboratories have grown markedly more interoperable, increasingly coordinated by shared orchestration software; AI agents, built on reasoning-trained models and scientific foundation models, have grown markedly more capable of multi-step reasoning and tool use; and a national mobilization, the Genesis Mission, has placed autonomous experimentation at the center of U.S. federal science strategy. Most strikingly, industry has become a primary actor, as well-capitalized entrants build robotic discovery laboratories at a pace that public programs cannot match. Progress has encountered a sobering counter-current. A published correction to a flagship autonomousdiscovery result retracted its novelty claims and removed a training-data leak; a growing body of benchmarks shows that agents that rival experts on closed-ended questions still complete only a small fraction of openended research; and fabricated citations have surfaced even in papers accepted at leading venues. Tellingly, no system has yet made and experimentally self-validated a genuinely novel discovery from end to end. We read this not as a transient growing pain but as the defining tension of the field.

Producing a candidate discovery is no longer the hard part. Verifying it is. This asymmetry, more than raw model capability, is what limits autonomous science today. Accordingly, this report organizes the roadmap around seven dimensions, the five we revisit and the two we elevate from cross-cutting afterthoughts to first-class dimensions. • Instrument and cyberinfrastructure integration. Vendor-neutral access to increasingly interoperable self-driving laboratories, with AI inference moving onto the instruments themselves. • Agent-driven data management. Provenance and quality enforced at the point of capture, where scarce data are made AI-ready by construction. • Agent-driven orchestration. Reasoning agents that plan and coordinate specialized scientific methods, and that explain why a result holds. • Interoperable agent interfaces. The consolidating MCP and A2A stack, now joined by an emerging skills layer. • Education and workforce development. The judgment to resist the homogenizing pull of automation. ★ Trust, verification, and reproducibility. (elevated) Verification and validation across the lifecycle as a first-class requirement. ★ Safety, security, integrity, and governance. (elevated) Screening and accountability where digital design meets physical execution. For each dimension, we assess the original milestones (M1–M14), classifying each one as achieved, partially achieved, reframed, or open, and we add four new milestones (M15–M18) for the two elevated dimensions. We scope the path forward to a two-year horizon, with the first year concentrating on interfaces, protocol adoption, and the scaffolding of verification, and the second targeting federation, zero-trust coordination, and governance. Throughout, we position the grassroots network as the interoperability fabric that allows national programs, international initiatives, and commercial platforms to connect rather than re-silo. 3

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

Table of Contents 1

Introduction

5

2 A Brief State of Autonomous Science 2.1 Shift 1: From Isolated Demonstrations to Validated Discoveries . . . . . . . . . . . . . . . . 2.2 Shift 2: Self-Driving Laboratories Enter a Second Generation . . . . . . . . . . . . . . . . . 2.3 Shift 3: The Interface Layer Begins to Consolidate . . . . . . . . . . . . . . . . . . . . . . 2.4 Shift 4: Reasoning and Foundation Models as a New Substrate . . . . . . . . . . . . . . . . 2.5 Shift 5: Benchmarks Expose the Capability-Reliability Gap . . . . . . . . . . . . . . . . . . 2.6 Shift 6: Trust and Governance Move to the Foreground . . . . . . . . . . . . . . . . . . . . 2.7 Shift 7: Industry Becomes a Primary Actor . . . . . . . . . . . . . . . . . . . . . . . . . . 2.8 Relation to Prior Roadmaps and Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . .

6 6 8 8 8 9 9 9 9

3

10 11 12 13 14 16 16 18

Critical Dimensions of the Roadmap 3.1 Instrument and Cyberinfrastructure Integration . . . . . . . . . . . . . . . . . . . . . . . . 3.2 Agent-Driven Data Management . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3 Agent-Driven Autonomous Orchestration . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4 Interoperable Agent Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5 Education and Workforce Development . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6 Trust, Verification, and Reproducibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7 Safety, Security, Integrity, and Governance . . . . . . . . . . . . . . . . . . . . . . . . . . .

4 The National and Global Ecosystem

19

5

20

Conclusion and Revised Roadmap

References

21

4

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

1

Introduction

The scientific discovery process is being reshaped by the convergence of automation, robotics, machine learning (ML), and artificial intelligence (AI). Modern instruments generate data at rates that outpace human analysis, and the cadence of human decision-making is increasingly mismatched with the speed at which experiments can be planned, executed, and interpreted [1]. Autonomous science addresses this mismatch by closing the loop between hypothesis, experiment, and analysis, with the goal of compressing discovery cycles that once took years or decades into months or weeks. One year ago, we argued that this promise was being held back by fragmentation, in that autonomous laboratories operated as isolated islands, unable to communicate across institutional or disciplinary boundaries [1]. To address this, we proposed the Autonomous Interconnected Science Lab Ecosystem (AISLE), a grassroots network organized around five critical dimensions: (1) instrument and cyberinfrastructure integration, (2) agent-driven data management with FAIR compliance [2], (3) agent-driven autonomous orchestration, (4) interoperable agent interfaces, and (5) education and workforce development. For each dimension, we surveyed the state of the art, identified open challenges, and proposed a set of milestones (M1–M14) to guide community efforts. The intervening year has been unusually eventful, and on balance, it has moved faster than the original roadmap anticipated. Multi-agent systems built on large language models (LLMs) have generated hypotheses that were subsequently validated in the laboratory, from SARS-CoV-2 nanobody design [3] to drug-repurposing candidates in oncology and ophthalmology [4, 5]. The interface “plumbing” that the original roadmap called for has begun to consolidate, with the Model Context Protocol (MCP) [6] and the Agent2Agent (A2A) protocol [7] emerging as complementary standards, and with science-specific middleware layering them onto high-performance computing (HPC) and experimental facilities [8, 9]. Most consequentially for a roadmap aimed at national-scale coordination, the launch of the Genesis Mission [10] has placed robotic laboratories and autonomous experimentation at the center of U.S. federal science strategy, alongside parallel efforts in other regions [11, 12]. Progress has been accompanied by a sobering counter-current. A published correction to a widely cited autonomous materials-discovery result walked back its central novelty claims, re-characterizing “novel” compounds as novel to a prediction platform rather than to science, and removing one compound that had leaked from the training data [13]. A growing body of benchmarks shows that agents which approach expert performance on closed-ended scientific questions still complete only a small fraction of open-ended, end-to-end research tasks [14, 15, 16], and fabricated citations have begun to appear even in papers accepted at leading venues [17]. We claim that these are not transient growing pains but the defining tension of the field. We can now generate candidate discoveries faster than we can verify them. Consequently, an updated roadmap cannot treat trust, verification, safety, and governance as cross-cutting concerns folded into other dimensions; they must become first-class dimensions of the roadmap itself. In this report, we present an updated community roadmap for interconnected autonomous science, one year after AISLE, scoped deliberately to a two-year horizon. Given how quickly the field is moving, we favor a two-year roadmap over a longer-range vision, so that the milestones we propose remain concrete and accountable rather than speculative. We group them into targets for the first year and targets for the second. This work makes the following contributions: 1. We characterize how the landscape of autonomous science has changed over the past year, organized around seven shifts that bear directly on the original roadmap (Section 2).

5

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

2. We assess progress against the original AISLE milestones (M1–M14), classifying each as achieved, partially achieved, reframed, or open, and we refine the five original dimensions accordingly (Section 3). 3. We elevate two concerns to first-class dimensions of the roadmap, namely trust and verification (Section 3.6), and safety, security, and governance (Section 3.7), and we propose milestones for each. 4. We position the grassroots AISLE network within the new federal and international landscape (Section 4), arguing that top-down mobilization and bottom-up coordination are complementary rather than redundant. Figure 1 shows the resulting structure in two complementary views. Figure 1(a) renders the architecture as a closed discovery loop in which federated AI agents, drawing on a shared data fabric, hypothesize and design, execute on self-driving laboratories and instruments, capture and curate data, and analyze and verify results, with MCP and A2A as the layers that connect agents to tools and to one another, all enclosed by the two dimensions we elevate to first-class, trust and verification on the inside, and safety, security, and governance on the outside, above an education and workforce foundation. Figure 1(b) recasts the same roadmap as a two-year trajectory that climbs the autonomy ladder from tool to analyst to scientist, grouping representative milestones into successive phases so that trustworthy autonomy rises to meet, rather than outrun, frontier model capability. The remainder of this report is structured as follows. Section 2 characterizes the current state of autonomous science. Section 3 develops the seven dimensions of the roadmap, the five we revisit, and the two we elevate. Section 4 situates the roadmap within national and global initiatives. Finally, Section 5 concludes the report and outlines a revised milestone roadmap.

2

A Brief State of Autonomous Science

Before revisiting the roadmap, we characterize how the landscape has changed since the creation of the AISLE grassroots network. We organize the discussion around seven shifts and adopt two lenses that the community converged on during the past year. The first is an autonomy ladder that distinguishes degrees of agency. Recent surveys have largely abandoned the binary vision of automated versus not automated in favor of graded taxonomies of autonomy. A representative formulation distinguishes three levels, tool, analyst, and scientist, according to whether the system executes well-defined tasks under direct supervision, conducts analysis within human-set boundaries, or formulates hypotheses and proposes new lines of inquiry on its own [18, 19]. The practical value of such a ladder is that it locates a given system, and a given roadmap milestone, at a specific rung rather than asserting wholesale autonomy. The second is a persistent capability-reliability gap that separates what systems can propose from what can be trusted, that is, the gap between performance on closed-ended scientific questions and performance on open-ended research. We claim that this gap, rather than raw model capability, is the binding constraint on autonomous science today, and we return to it throughout the report.

2.1

Shift 1: From Isolated Demonstrations to Validated Discoveries

A year ago, the strongest claims for autonomous discovery rested on self-driving laboratories optimizing within narrow design spaces. Since then, multi-agent LLM systems have produced hypotheses that were subsequently validated experimentally. A virtual laboratory of AI agents designed SARS-CoV-2 nanobodies

6

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

Safety, Security & Governance Trust, Verification & Reproducibility iterative discovery loop

Execute on SDLs & Instruments

Hypothesize & Design

Analyze & Verify

Capture & Curate Data MCP

A2A

Federated AI Agents + Data Fabric Education & Workforce Development

(a) Interconnected, closed-loop architecture

Closing the capability–reliability gap over two years frontier model capability

autonomy ↑

trustworthy autonomy Scientist

• Federated data mesh (M6) • Verification & validation (M15) • Zero-trust comms (M11)

Analyst

Tool

• Self-discovering networks (M12) • Agent identity & governance (M18) • National framework (M4)

• Vendor-agnostic interfaces (M1) • MCP/A2A adoption (M10) • AI metadata (M5) 0–8 mo

8–16 mo

16–24 mo

time

(b) From assistance to trustworthy autonomy

Figure 1: The updated AISLE roadmap. (a) A closed discovery loop, in which federated AI agents over a distributed data fabric hypothesize and design, execute on self-driving laboratories and instruments, capture and curate data, then analyze and verify, is coordinated by two protocol layers, MCP for agent-to-tool access and A2A for agent-to-agent collaboration across institutions, and is enclosed by the two dimensions this paper elevates to first-class, namely trust and verification and safety, security, and governance, over an education and workforce foundation. (b) Across a two-year horizon, the roadmap climbs an autonomy ladder from tool to analyst to scientist, with representative milestones per phase, so that trustworthy autonomy rises to close the gap with frontier model capability.

7

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

that were synthesized and shown to bind variant targets [3], an AI co-scientist generated drug-repurposing and target-discovery hypotheses confirmed in vitro by collaborating laboratories [4], and a multi-agent system proposed a therapeutic candidate for an ophthalmic indication that was validated in follow-up assays [5]. These are genuine advances. In every case, however, the physical experiments were executed by humans, the problems were human-selected, and at least one celebrated result amounted to re-deriving a mechanism that a human laboratory had already established but not yet published. At the time of writing, there is no verified instance of an agent autonomously making and experimentally self-validating a genuinely novel discovery end to end. The best-funded recent efforts each miss a different aspect of autonomous discovery, as the systems that reason most autonomously still rely on humans to run the physical experiments [5], the systems that synthesize most autonomously have had their novelty claims contested [13], and the systems that close the full loop without human intervention do so only in computational settings with no wet-lab validation [20]. We revisit both the agents that produce such results and the verification their claims still demand in Section 3.

2.2

Shift 2: Self-Driving Laboratories Enter a Second Generation

The self-driving laboratory (SDL) community has begun to describe its own trajectory as a move from a first generation of narrow, hand-tuned, poorly interoperable platforms toward what it calls a second generation, or SDL 2.0, that is interoperable, orchestrated, safe, and capable of hypothesis generation [21, 22]. The concrete advance so far is interoperability, with the other properties still largely aspirational. By orchestrated, the community means coordinated by a software layer that sequences instruments, robots, and computational steps into managed, restartable campaigns rather than hand-scripted one-off runs. This shift directly concerns the first three dimensions of the original roadmap, which we revisit in Section 3, where we detail the orchestration frameworks and robotic platforms that make it concrete.

2.3

Shift 3: The Interface Layer Begins to Consolidate

When the original roadmap was written, agents reached tools and one another through point-to-point and proprietary interfaces, and we described the requirements in generic terms. Within a year, two complementary standards have emerged. The Model Context Protocol (MCP) connects a single agent to many tools and data sources, a vertical concern [6], while the Agent2Agent (A2A) protocol connects agents to one another across vendor and institutional boundaries, a horizontal concern [7]. For science specifically, federated-agent middleware and protocol adapters now expose HPC, data-movement, and instrument services to LLM agents through these interfaces [8, 9]. We treat this consolidation as the most concrete update to the original roadmap, even as the pattern by which agents consume these protocols keeps changing, and we develop it in Section 3.4.

2.4

Shift 4: Reasoning and Foundation Models as a New Substrate

The underlying models have also changed. Reasoning-oriented LLMs trained with reinforcement learning now approach expert performance on graduate-level scientific question answering [23], and domain foundation models have produced experimentally corroborated designs in materials and structural biology [24, 25]. These models are the substrate on which orchestration (Section 3.3) is increasingly built. They also sharpen the reliability question because their fluency makes unsupported outputs harder, not easier, to detect.

8

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

2.5

Shift 5: Benchmarks Expose the Capability-Reliability Gap

A wave of benchmarks has made the second framing lens quantitative. Agents that perform well on isolated, closed-ended tasks complete only a small fraction of open-ended, end-to-end research tasks [14, 15], and research-grade coding benchmarks curated by scientists remain largely unsolved by frontier models [16]. Reproducibility-focused agent benchmarks report accuracies that, in some settings, fall below random guessing [26]. Deployment evidence points the same way, as the first systematic study of agents in production finds that teams keep them deliberately bounded and human-supervised, with most running only a handful of steps before a human intervenes and the majority still gated by human evaluation, and with reliability named as the top obstacle [27]. The lesson for the roadmap is that progress should be measured against open-ended, verifiable tasks rather than against exam-style proxies that are rapidly saturating.

2.6

Shift 6: Trust and Governance Move to the Foreground

The past year also supplied concrete evidence that trust and governance can no longer be treated as secondary. A published correction to a flagship autonomous-discovery result re-characterized its novelty claims and removed a training-data leak [13], fabricated citations appeared in papers at leading venues [17], analyses found that journal AI-disclosure policies have done little to curb undisclosed AI-assisted writing [28], and the convergence of capable models with remotely accessible laboratories raised dual-use concerns that the existing policy apparatus does not address [29]. These developments motivate the two dimensions we add to the roadmap in Sections 3.6 and 3.7.

2.7

Shift 7: Industry Becomes a Primary Actor

Finally, the past year changed who builds autonomous science. A field that had been driven largely by academic and national-laboratory prototypes acquired well-capitalized industrial entrants, and the change is large enough to bear on every dimension of the roadmap. Startups founded by senior industry researchers and dedicated to autonomous discovery raised sums that dwarf typical academic budgets, with Lila Sciences assembling roughly half a billion dollars to build robotic “AI science factories” [30] and Periodic Labs raising a $300 million seed round to pursue closed-loop materials discovery [31]. The major cloud and chip vendors moved in parallel, shipping agentic research platforms [32] and scientific foundation models [24], and positioning large-scale compute and physics-based simulation as the substrate for autonomous experimentation [33]. Commercial cloud laboratories matured into a credible delivery model for remote, programmatic experimentation, a development we take up in Section 3.1. We read this influx as a double-edged development. It supplies capital, engineering, and infrastructure that the academic community cannot match, yet the distance between the capability these entrants claim and the evidence they have published is itself an instance of the capability-reliability gap, since several of the best-funded autonomous-discovery efforts have released little peer-reviewed validation as of this writing [30, 31]. A roadmap for interconnected science must therefore treat industry as a primary actor while asking of it the same open interfaces, provenance, and verification that it asks of public laboratories.

2.8

Relation to Prior Roadmaps and Surveys

The past year also produced a wave of surveys and roadmaps for autonomous science, including taxonomies of autonomy [18], broad framings of agentic discovery [19], and domain-oriented catalogs of agentic systems [34]. These works are valuable as maps of the literature, and we draw on them for the two lenses

9

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

Table 1: Status of the original AISLE milestones (M1–M14) and proposed new milestones (M15–M18), each with a target year (Y1 or Y2) on the two-year roadmap. Statuses are provisional and subject to confirmation against the cited evidence. #

Milestone (abbreviated)

Status

By

M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14

Common instrument interfaces, HAL End-to-end cross-institution workflows Open compute fabric, fault tolerance, digital twins Scalable national instrument framework AI-driven metadata and annotation Federated data mesh, FAIR governance Near-real-time processing and AI provenance Hierarchical LLM orchestration, verification Cross-facility knowledge integration Standardized cross-vendor agent interfaces Zero-trust, sub-second agent coordination Self-discovering agent networks National education consortium Virtual labs, human-AI assessment

Partial Partial Open Open Partial Open Partial Partial Open Reframed Open Partial Open Open

Y1 Y2 Y2 Y2 Y1 Y2 Y2 Y1 Y2 Y1 Y2 Y2 Y1 Y2

M15 M16 M17 M18

Verification and validation across the lifecycle Reproducibility and efficiency benchmark Screening at digital-physical interface Cross-institution agent identity, governance

New New New New

Y1 Y1 Y1 Y2

above. Our contribution is different in kind. Rather than surveying the field anew, we update a specific, milestone-bearing roadmap [1, 35] against one year of evidence, score its milestones, and revise its structure where the evidence demands. We believe that this form of accountable revision, in which a community commits to milestones and then publicly assesses them, is a useful complement to survey-style framings and one that the field currently lacks. We intend it as a recurring practice rather than a one-off, revisiting and rescoring these milestones as the field advances so that the roadmap stays honest about what has and has not been achieved.

3

Critical Dimensions of the Roadmap

We organize the roadmap around seven dimensions. Rather than enumerate them in isolation, Figure 1 situates them within a single closed-loop architecture, where the discovery-loop stages carry orchestration, instruments, and data, the coordinating protocol layers carry the agent interfaces, the enclosing bands carry trust and governance, and the foundation carries education and workforce development. The first five revisit and update the dimensions of the original AISLE roadmap [1] in light of the past year, while the last two, trust and verification (Section 3.6) and safety, security, and governance (Section 3.7), are elevated here from cross-cutting concerns to first-class dimensions. For each dimension, we summarize the state of the art, identify the challenges that remain, state research priorities, and record the status of the associated milestones in the scorecard of Table 1. The subsections that follow justify these entries dimension by dimension, and we interpret the overall pattern in the conclusion (Section 5).

10

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

3.1

Instrument and Cyberinfrastructure Integration

The first dimension concerns how autonomous agents orchestrate diverse experimental equipment and computational resources across institutional boundaries. Increasingly, the two are inseparable, as instruments, edge devices, HPC, cloud, and digital twins form a single distributed cyber-physical system in which AI models, simulation, and laboratory automation execute as one coordinated workflow whose reliability determines the trustworthiness of the result. Over the past year, the defining development has been the community’s move toward the second-generation self-driving laboratories (SDLs) described in Section 2, more interoperable and orchestrated than the narrow, hand-tuned systems that dominated earlier demonstrations [21, 22]. Brief state of the art. Orchestration software has matured from bespoke scripts toward reusable frameworks, such as MADSci for modular discovery campaigns [36] and ChemOS 2.0 for chemical SDLs [37], while mobile and multi-robot platforms now approach human throughput on real synthesis tasks [38]. Beneath the agent layer, device-control standards such as SiLA 2 provide vendor-neutral, typed instrument communication [39], and standards bodies have begun to target the interfaces that a modular autonomous laboratory requires [40]. At the facility scale, programs that connect instruments, robotic laboratories, and HPC into shared workflows have continued to grow [41]. A complementary commercial route to instrument access has matured in parallel, namely cloud laboratories that expose physical instruments to remote programmatic control, so that an experiment specified as code is executed by robots and technicians at a central facility [42]. This model has reached academia through the first university cloud lab [43], and it has already been driven by a language-model agent that learned the facility’s scripting language from documentation and carried out cross-coupling reactions end to end [44]. National programs have begun to move this model from isolated facilities toward a networked resource, most concretely a federal test-bed for a network of programmable cloud laboratories linked by shared networking and data standards (Section 4) [45]. A further shift concerns where computation happens relative to the instrument. As the data rates of upgraded light sources, particle detectors, and radio arrays began to outpace the store-then-analyze pipeline, reaching exabyte-scale annual volumes at the largest facilities [46, 47], machine-learning inference has moved toward the instrument edge so that detector data are processed as they stream rather than after they are written to storage [48]. Demonstrations include real-time ptychographic reconstruction from a streaming detector on an edge accelerator [49] and machine-learning triggers that decide within microseconds which events to retain [50]. Challenges. The heterogeneity that motivated the original roadmap persists, as most instruments still ship without a standard control interface, retrofitting drivers is labor intensive, and the majority of deployed SDLs operate at the lower rungs of the autonomy ladder of Section 2. The past year added a sharper concern at the instrument level, namely that conclusions drawn from automated characterization can be wrong in ways that human inspection would catch, as the corrected analysis of an autonomous materials laboratory made clear [13]. Organizational barriers compound the technical ones, since intellectual-property and liability questions arise the moment an instrument in one institution is driven by an agent in another. A deeper constraint is that the hardware has advanced more slowly than the software around it. Robotics, sample handling, and reconfigurable experimental setups remain comparatively rigid, and where the apparatus cannot be recomposed under program control, the reach of an agentic workflow is bounded by a largely fixed experiment, which optimization and campaign-driven science can accommodate but open-ended discovery cannot. Research priorities. We reaffirm the need for vendor-agnostic hardware abstraction layers and self-describing

11

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

instruments that expose their capabilities semantically, and we add the validation of automated characterization as a priority. In effect, this extends the MCP-style capability discovery of the agent interface layer (Section 3.4) down to the instrument and its data-acquisition system, so that an agent can discover and drive a detector as readily as any other software tool. This capability sits above the typed device-control interfaces that standards such as SiLA 2 already provide [39]. We also add modular, agile, and reconfigurable experimental hardware as a priority in its own right, since vendor-neutral software interfaces deliver little if the apparatus beneath them cannot be recomposed. Recent work pursues exactly this for the structural refinement that the corrected materials result called into question, building a rubric-bounded agent for Rietveld analysis of X-ray diffraction whose scoring rewards credible fits and flags out-of-scope cases rather than forcing a pattern fit [51]. Physics-aware digital twins should be used to test autonomous workflows before they touch physical instruments, as in self-driving laboratories that validate a workcell in simulation before any physical run [52], and instrument-level outputs should carry the provenance needed to audit downstream claims (Section 3.6). We further prioritize real-time, edge-side inference and data reduction so that analysis and on-the-fly experiment steering keep pace with instruments whose output exceeds what centralized pipelines can absorb, and so that raw streams are turned into analysis-ready data at the point of acquisition (Section 3.2) [48]. Above the individual instrument, the cyberinfrastructure half of this dimension calls for an open compute fabric that spans edge, HPC, cloud, and heterogeneous accelerators, exposing vendor-neutral interfaces for model execution, accelerator scheduling, workflow portability, and provenance capture, much as MCP and A2A do for tools and agents (Section 3.4), so that workflows move across platforms while hardware and software ecosystems evolve independently beneath them [53]. The status of the original milestones M1 through M4 is summarized in Table 1, and the national framework envisioned by M4 is now partly scaffolded by federal mobilization [10].

3.2

Agent-Driven Data Management

The second dimension concerns the shift from centralized repositories to distributed systems in which autonomous agents curate, validate, and federate scientific data, enforcing FAIR principles at the point of capture [1, 2]. The past year reframed this dimension around trust. As autonomous campaigns generate data that humans never inspect, provenance and quality assessment become prerequisites for credible discovery rather than after-the-fact bookkeeping. Brief state of the art. Federated data services coordinate movement and processing across many facilities [54], and pass-by-reference systems allow large datasets to be shared without duplication in distributed agent workflows [55]. Provenance models such as PROV-O provide a vocabulary for traceability [56], and standards bodies have begun to target the data and knowledge-management gaps specific to autonomous laboratories [40]. What remains missing is the autonomous enforcement of these capabilities, i.e., agents that curate and annotate data as it is produced rather than leaving it to downstream human stewardship. Challenges. Beyond the format and schema heterogeneity identified previously, four issues have moved to the foreground. First, autonomous agents must negotiate schema evolution when they encounter new experiment types, without manual intervention. Second, data quality is now a verification problem because contaminated or leaked data can propagate silently through AI-driven decision chains, as illustrated by a training-data leak in a flagship autonomous-discovery result [13]. Agents, therefore, require mechanisms to assess reliability from the experimental context, rather than treating all data as equally trustworthy. Third, automation multiplies the volume of raw data, yet the high-quality, well-characterized data needed to train reliable models remains scarce, and how that data is represented can determine whether a model learns the underlying physics at 12

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

all [57, 58]. Much of what separates reusable data from raw output is well-curated metadata, the theoretical and experimental context, conditions, and provenance that make a measurement findable, interpretable, and trustworthy [2], and capturing it at the moment of acquisition rather than reconstructing it afterward is what keeps data quality tractable as volume grows. Fourth, the move to online, in-transit data reduction (Section 3.1) sharpens the problem because signals that are filtered or summarized at the instrument edge can never be re-examined. The FAIR lifecycle and its provenance must therefore be preserved through reduction, recording not only what was kept but also what was removed and why, lest reproducibility be quietly lost as the data is reduced. Research priorities. We prioritize federated data-mesh architectures in which each laboratory maintains a node with standardized interfaces and global discovery indices, support for both explicit and implicit schemas so that agents can infer structure from heterogeneous sources, and the embedding of provenance frameworks [56] into instrument middleware so that every autonomous decision is traceable across facilities and timescales. We also prioritize making data AI-ready by construction, with readiness levels that track a dataset from raw capture to a form fit for training [59], and community standards for autonomous-laboratory data that play, for machine consumption, the role that FAIR played for sharing [2, 60]. Data curated to this standard are not only an audit trail but also the substrate for surrogate models that compress expensive experiments and carry materials knowledge toward manufacturing timescales [61, 62]. More broadly, as autonomous agents become the dominant consumers of these data, the systems that serve them are better designed to be agent-first from the outset, anticipating the high-volume, exploratory access patterns of agents rather than retrofitting interfaces built for human analysts [63]. The status of milestones M5 through M7 is summarized in Table 1.

3.3

Agent-Driven Autonomous Orchestration

The third dimension is the cognitive core of interconnected autonomous laboratories, namely the agents that navigate scientific decision spaces while remaining aligned with physical and domain knowledge [1]. The substrate for this dimension changed markedly over the past year, as reasoning-oriented LLMs and domain foundation models matured into orchestrators that coordinate specialized methods such as Bayesian optimization, uncertainty quantification, and reinforcement learning. Brief state of the art. Multi-agent systems that generate, debate, and evolve hypotheses have produced experimentally validated results [4], and agentic tree-search systems now carry an idea from conception to a written manuscript [20]. The models beneath them have improved on two fronts. Reasoning-trained LLMs approach expert performance on graduate-level science questions [23], and domain foundation models produce experimentally corroborated designs in materials and structural biology [24, 25]. A complementary line of work casts the LLM as an orchestrator of established tools rather than a replacement for them, for example, by exposing expert-designed chemistry tools to a language model [64]. Challenges. The probabilistic nature of LLM-based agents remains in tension with the determinism that reproducible science assumes, and it is unclear how to guarantee reproducible outcomes or grounding in physical law, when an orchestrator is non-deterministic, higher-latency, and difficult to verify. The capabilityreliability gap of Section 2 makes this concrete, since strong performance on closed-ended tasks has not carried over to open-ended research. The orchestrator is therefore best understood as one component of the ecosystem, rather than a replacement for the established methods it coordinates. A recurring response is to keep those methods in the loop, combining data-driven learning with physical and chemical constraints and

13

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

checking an agent’s proposals against known laws before they reach an instrument [34]. In this view, the probabilistic reasoning of an LLM is an asset for exploration and a liability for commitment, and the design problem is to route each to where it belongs. Research priorities. We reaffirm three thrusts. The first is hierarchical architectures in which LLM-based agents orchestrate established scientific methods abstracted as actuators. The second is a verification and validation infrastructure that uses digital twins and formal or symbolic methods to enforce physics-based constraints as hard boundaries, which we develop in Section 3.6. The third is distributed, real-time knowledge integration that keeps agents grounded across facilities and long campaigns. Grounding of this kind rests on a primitive the original roadmap did not name, namely durable agent memory. We mean policy-bound, provenance-rich, scoped persistence, spanning the working, episodic, semantic, and procedural forms long studied for cognitive agents, that lets an agent carry context, decisions, and lessons across a long campaign instead of rebuilding them inside a prompt, and that is distinct from both the scientific data of Section 3.2 and the model weights beneath it [65]. Early systems realize this idea by managing a small in-context working set against external stores, much as an operating system pages physical memory against a larger virtual address space [66]. Cutting across all three, the construction of an agent is worth treating as a workflow stage in its own right, so that a scientist authors a durable contract, a version-controlled rubric, a graded curriculum, and a curated knowledge base, that bounds an automated builder and survives the churn of the underlying models. This has been demonstrated for the orchestration of a crystallographic refinement tool [51]. We also hold that an orchestrator should yield understanding, not only predictions, since an explanation of why a design or result holds is what makes an autonomous finding actionable, and interpretability of this kind is increasingly expected of scientific machine learning [67, 68]. A further fragility is dependence on the models themselves, since most deployed agents rely on a small set of proprietary frontier models [27], an uneven foundation for reproducible science and one not equally available across regions. Open-weight models accompanied by clear provenance of their training data are therefore of growing interest, because reproducibility requires knowing what a model learned from, and because scientific sovereignty requires not depending on a single vendor. Once agents depend on them throughout a campaign, models are best treated as persistent scientific infrastructure, versioned, evaluated, and retired with the discipline applied to instruments and scientific software rather than swapped silently beneath a running workflow. The status of milestones M8 and M9 is summarized in Table 1.

3.4

Interoperable Agent Interfaces

The fourth dimension concerns the interfaces and standards through which autonomous agents reach the tools, data, and capabilities they depend on, and through which they coordinate with one another across institutional and disciplinary boundaries. This dimension has changed more than any other since the original roadmap, which described its requirements in generic terms because no widely adopted standard yet existed [1]. Within a year, a recognizable two-axis stack has emerged. The Model Context Protocol (MCP) standardizes how a single agent connects to many tools and data sources, while the Agent2Agent (A2A) protocol standardizes how agents discover, authenticate, and delegate to one another across vendor and institutional boundaries [6, 7]. We treat these as complementary, vertical and horizontal respectively, rather than competing [69]. More recently, a third layer has begun to form above these protocols, namely agent skills: composable bundles of instructions, scripts, and resources that an agent discovers and loads on demand. While MCP and A2A govern how agents reach tools and one another, skills capture what an agent needs to know to carry out a task [70].

14

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

Brief state of the art. Both protocols moved under neutral governance within months of each other, which lowers the risk of relying on a single vendor for betting infrastructure [7, 69]. For science specifically, federated-agent middleware deploys stateful agents across HPC, data, and experimental resources [8], and thin adapters now expose facility services, such as data movement and remote function execution, to LLMbased agents through MCP [9]. These layers sit above the instrument-control standards of Section 3.1 [39], so that a typed instrument interface can be wrapped for agent access without discarding existing control software. The way agents consume these interfaces is itself in flux. Loading every tool definition into the context window and routing each intermediate result back through the model scales poorly, in both token cost and error surface, once an agent faces dozens of tools. A growing practice instead has the agent write code that calls the tools and discovers their definitions on demand, which repositions the protocol as a substrate beneath code execution rather than a direct call interface [71]. A further step, pursued in industry though likely beyond our two-year horizon for science, is a runtime that surfaces tools and agents dynamically, exposing each step of a workflow only to the capabilities relevant to its goal and folding discovery, scoping, and least-privilege access into the interface layer itself. We read this churn not as the demise of any one protocol but as a reason to standardize the interface while leaving the calling pattern free to evolve, which is the layered stance we take in the priorities below. Challenges. The most-cited blocker is identity and delegated authorization at the facility scale because agents must act on a scientist’s behalf across institutions without holding overly broad, long-lived credentials [9]. A second mismatch is structural, as the request-response tool model fits poorly with the stateful, long-running jobs that characterize scientific campaigns [8, 9]. Semantic interoperability across domains, the provenance of agent-to-agent decisions, and the tension between protocol fragmentation and convergence all remain open. The security posture of these protocols also trails their adoption, as national-security guidance warns that placing tool descriptions, control flow, and data in one shared context blurs trust boundaries, overlapping contexts can leak state across tasks, and unverified dynamic tool discovery widens the reach of a compromised component [72]. Research priorities. We prioritize layered protocol architectures that separate networking, message formatting, semantic interpretation, and coordination, agent identity with attribute-based access control designed for multi-institutional collaboration, provenance-carrying messages that record which agent did what with which tool on which data, and self-describing capability discovery validated on multi-facility testbeds. On identity in particular, the same runtime that scopes which tools a step may see can also scope its authority, issuing signed delegations on demand and limiting their lifetime to a single invocation rather than to the agent as a whole, a pattern commercial systems already implement and one that facility identity-and-access infrastructure will need to grow to support. Recording the prompts, decisions, and tool calls exchanged between agents is a solved problem in agent-observability products, so the open question for science is not whether to log these interactions but what a scientifically adequate record must contain, one that ties each decision to the data, tools, and model versions behind it and serves reproducibility and cross-facility reuse rather than operational monitoring alone. We develop that provenance model, and its use in verification, in Section 3.6 [73]. In the same spirit, and because no widely used skill library yet encodes scientific procedures, we see an opportunity to package validated experimental and analysis workflows as portable, self-describing skills that move across laboratories, which would give capability discovery concrete content to advertise and reuse [70]. The original milestone M10 named specific transport protocols, and we reframe it around the MCP and A2A consolidation. The status of M10 through M12 is summarized in Table 1.

15

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

3.5

Education and Workforce Development

The fifth dimension concerns preparing researchers for environments increasingly shaped by autonomous systems, where competencies span AI/ML methods, workflow thinking, human-machine collaboration, and ethical reasoning [1]. The past year added empirical urgency, as a large-scale analysis of AI-engaged researchers found that they publish more and are cited more, even as the range of topics the community collectively studies contracts [74]. Workforce development must therefore cultivate not only fluency with autonomous tools but also the judgment to resist their homogenizing pull. Brief state of the art. Scientific education still largely treats AI/ML as supplementary rather than integral, which leaves competency gaps for autonomous-laboratory settings. National and international programs have begun to fund workforce development as an explicit component of their AI-for-science strategies [10, 11], but systematic curriculum redesign remains limited across institutions. Challenges. The central tension is balancing automation with foundational understanding since over-reliance risks producing scientists who cannot critically evaluate automated results. Compounding this, current assessment methods do not measure the ability to collaborate with AI, many educators lack the relevant expertise, and access to autonomous-laboratory infrastructure for hands-on training is uneven, which raises equity concerns. Research priorities. We prioritize modular curricula that integrate AI/ML competencies with scientific reasoning across disciplines, immersive virtual environments that simulate autonomous laboratories where physical access is limited, and assessment methodologies that evaluate trust calibration and the interpretation of AI decisions, with ethical reasoning integrated throughout. As agentic workflows take over routine tasks and extend what a single researcher can attempt, the design of the human-machine interface becomes a research problem in its own right. Interfaces must let a scientist supervise, interrogate, and override an agent, calibrate trust in its output, and be augmented rather than sidelined by it. The status of milestones M13 and M14 is summarized in Table 1.

3.6

Trust, Verification, and Reproducibility

We elevate trust, verification, and reproducibility from a cross-cutting concern, as it appeared in the original roadmap, to a first-class dimension. The motivation is the capability-reliability gap introduced in Section 2. The ability of agents to propose hypotheses, designs, and analyses has outrun the infrastructure needed to verify them. The past year supplied concrete evidence of the cost of this gap, from the corrected materials-discovery result we examine below [13], to fabricated citations surfacing in papers accepted at leading venues [17], to reproducibility benchmarks on which agents perform near or below chance [26]. We claim that an interconnected ecosystem without commensurate verification infrastructure would amplify, not contain, these failure modes. Brief state of the art. Evaluation has shifted from closed-ended question answering toward reproduction and end-to-end research, with benchmarks that ask agents to reproduce published computational results [26] or to carry a task through the full discovery process [14]. Early verification mechanisms are appearing, including uncertainty quantification, neurosymbolic constraints, and provenance that classifies each claim by its epistemic source, but they are not yet standard components of autonomous workflows. Provenance for agentic workflows, in particular, is beginning to mature, with recent work extending the W3C PROV standard over the Model Context Protocol to capture agent prompts, responses, and decisions as near-real-time,

16

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

end-to-end workflow provenance, so that an erroneous output can be traced as it propagates from one agent to the next across edge, cloud, and HPC resources [73]. Challenges. Verification in this setting is hard for four reasons. First, the non-determinism of LLM-based orchestration is in tension with run-to-run reproducibility. Second, fluent outputs make hallucinated results and citations harder to detect, not easier. Third, distinguishing genuine novelty from novelty relative to a model or database requires clean train and test separation that current pipelines do not guarantee [13]. The same difficulty recurred for a generative materials model whose experimentally highlighted compound was argued to be isostructural with a phase known for decades and already present in the model’s training distribution [75], which shows that the problem is not confined to a single laboratory or method. Fourth, verifying long campaigns distributed across facilities and agents is qualitatively harder than checking a single result. A cautionary case. The corrected materials-discovery result illustrates the failure mode in miniature. An autonomous laboratory reported dozens of newly synthesized compounds, and subsequent analysis showed that many were ordered versions of already-known disordered phases, that the automated structural refinement was below the quality a human expert would accept, and that one reported compound had leaked from the training data [13, 76]. None of these errors required a flaw in the robotics. They followed from treating an automated pipeline’s output as a discovery without independent verification. We read this not as an indictment of autonomous laboratories, but as evidence that verification must be designed in from the start, especially once results from one laboratory begin to feed the agents of another. Research priorities. We prioritize verification-and-validation architectures spanning the experiment lifecycle, in which physics-based and logic-based constraints act as hard boundaries rather than soft preferences, end-to-end provenance that links each claim to the data and tool executions that produced it, automated verification of citations and reported results at the point of submission [17], and reproducibility metrics for autonomous workflows that make run-to-run variation measurable. We use the two terms deliberately. Verification asks whether a workflow was built and executed correctly, and validation asks whether its result corresponds to physical reality, the second being the harder problem in the physical sciences, where a simulation is a poor substitute for an experiment and only measurement can settle the question. It is also why the strongest claims of autonomous discovery have so far come from mathematics and computation, where a result can be checked by a machine, rather than from physics, chemistry, or biology, where it cannot. Because an agent can reason soundly from a mistaken picture of the world, a high-consequence action should additionally be checked against independent measurements that the agent does not control. Two further practices sharpen this agenda. First, the natural unit of reproducibility for an agentic workflow is the run itself, an auditable record of the plan, the tool calls, the data and model versions, the human approvals, and the evaluation outcomes that produced a result, rather than the result alone [73]. That record must reach the computational execution as well, what we call AI provenance, capturing model identities and versions, inference parameters, retrieved sources and prompt context, uncertainty estimates, and the accelerators and execution environments behind a result, so that a finding can be reproduced as models, data, and computing environments evolve. Second, evaluation must judge the process and not only the answer, since tool choice, recovery from failure, and respect for budget and policy all bear on whether a result can be trusted [77]. As autonomous laboratories move toward continuous operation, efficiency itself becomes a measure of scientific productivity, so community benchmarks should report not only discovery rate and accuracy but systems-level costs, from energy per validated experiment and compute per optimization cycle to data-movement overhead, end-to-end latency, and the overhead of verification, enabling reproducible comparison across heterogeneous 17

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

computing environments [78]. We propose two milestones, M15 and M16, summarized in Table 1.

3.7

Safety, Security, Integrity, and Governance

We add safety, security, integrity, and governance as the seventh dimension of the roadmap. The original roadmap noted intellectual-property and liability concerns in passing, and the past year made a broader set of governance questions unavoidable. The convergence of capable models with remotely accessible laboratories makes it easier for non-experts to run sophisticated experiments, which raises dual-use concerns that current biosecurity policy does not address [29, 79]. Human interaction with self-driving labs and their LLM interfaces continues to have unexpected and unintended actions resulting in, for example, disclosure of personal information [80]. In parallel, the integrity of the scholarly record has come under strain, as analyses find that journal disclosure policies have done little to curb undisclosed AI-assisted writing [28], and as community consensus holds that AI systems cannot be accountable authors. Brief state of the art. Technology-and-policy reviews have begun to map the governance landscape for autonomous laboratories [29], work on AI models and agents has started to catalog capabilities of concern [79, 81], and national strategy now frames autonomous experimentation as a security as well as a scientific priority [10]. Publisher norms and disclosure requirements exist, but evidence indicates that they are widely unobserved [28], and screening at the boundary between digital design and physical execution remains largely voluntary. Analyses of commercial cloud laboratories have made the threat concrete, observing that an experiment can be designed on a laptop in one jurisdiction and executed by robots in another, and have proposed know-your-customer screening together with a cloud-lab security consortium modeled on the existing self-regulation of commercial gene synthesis [82]. Challenges. Dual-use risk concentrates where capable agents meet cloud and self-driving laboratories because the same automation that lowers the barrier for legitimate researchers also lowers it for misuse [29, 79]. Security, safety, and ethical concerns can arise at the human interface with agents and self-driving laboratories when there is physical proximity between humans and equipment or dangerous materials; when humanderived data is accessible to agents; or when there is the potential for misalignment of objectives and ethical norms [81, 83]. Accountability is diffuse when agents act across institutions, complicating intellectual property, liability for cross-institutional failures, and the question of who is answerable for an autonomous decision. As campaigns increasingly span public and commercial partners, unresolved questions of data ownership and value capture become a near-term barrier rather than a distant one, and they need to be settled alongside the technical interfaces, not after them [29]. Throughout, governance lags capability, and a networked ecosystem widens the gap by increasing both reach and speed. These are precisely the conditions under which a small number of failures, or a single misuse, can erode public trust in autonomous science as a whole. Research priorities. We prioritize human-in-the-loop guardrails with reliable override, screening, and audit at the digital-to-physical interface for networked SDLs, accountable agent identity and provenance building on Sections 3.4 and 3.6, governance frameworks that preserve institutional autonomy while enabling collaboration, and disclosure-by-construction, in which AI-assisted contributions are recorded automatically rather than self-reported. Trust in a distributed autonomous system also needs roots below the application layer, in trusted execution environments, cryptographically verifiable model and agent identities, signed and immutable workflow and provenance records, and hardware-rooted attestation, which together let one institution rely on another’s computation without surrendering control of it. Underlying these safeguards is a

18

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

simple principle, that the autonomy granted to an agent should match the cost, risk, and reversibility of the action it would take, and should be raised only as auditable evidence of reliability accumulates rather than asserted at the outset [77, 27]. We propose two milestones, M17 and M18, summarized in Table 1.

4

The National and Global Ecosystem

The original roadmap argued for a grassroots, bottom-up network on the premise that no coordinated national program existed to connect autonomous laboratories. That premise has partly changed. The launch of the Genesis Mission has mobilized U.S. national laboratories around an integrated platform that couples high-performance computing, scientific foundation models, datasets, and automated laboratory systems, with explicit near-term objectives for robotic laboratories and autonomous experimentation [10]. In the months since launch, that mandate has acquired concrete form, as the program named twenty-six national science and technology challenges, one of which, achieving AI-driven autonomous laboratories, targets the same automated experimentation that the present roadmap addresses [84]. Related efforts, such as the Trillion Parameter Consortium [85], the European strategy for AI in science [11], and the Acceleration Consortium [12], constitute a dense and growing institutional landscape. That landscape now has a substantial private layer as well since the Genesis Mission enlisted two dozen industrial partners, among them chip and cloud vendors and venture-funded autonomous-discovery startups [86], and since those same vendors and startups are building platforms and laboratories of their own (Section 2). Within the United States, the mobilization reaches well beyond the U.S. Department of Energy. The National Science Foundation has established a Directorate for Technology, Innovation, and Partnerships and a milestone-funded X-Labs initiative for breakthrough science, with early topics that include scientific instrumentation [87, 88], the Defense Advanced Research Projects Agency treating AI-driven autonomous experimentation as a national-security capability [89], and the National Institute of Standards and Technology anchoring the measurement science and data standards that autonomous laboratories require, building on the Materials Genome Initiative [60, 62]. A concrete near-term pathway is a DOE call for robotics and automation testbeds for autonomous scientific discovery, framed as reusable community infrastructure for the national laboratories and their industry partners [90]. Closest of all to this roadmap’s own premise, a National Science Foundation test-bed program aims to build a network of programmable cloud laboratories, remotely accessible autonomous facilities to be linked by computational networking and shared data and AI standards [45]. This multi-agency picture strengthens, rather than weakens, the case for a coordinating fabric because each program risks building its own island unless interfaces and standards are held in common. The European setting shows how much this division of labor depends on local conditions. Public compute for large-scale AI is concentrated in a few sites, such as the EuroHPC exascale system JUPITER at Jülich and the national centres of the Gauss Centre for Supercomputing, while compute at the laboratories themselves is more modest, and high-speed networking between laboratories and compute sites is not yet in place [91]. Federation across facilities, which much of this roadmap presumes, is therefore gated by infrastructure that is unevenly available. Access to the most widely used proprietary models is likewise not guaranteed in every region, and European alternatives remain sparse, which is one more reason that open-weight models with transparent provenance (Section 3.3) matter beyond reproducibility alone. We include this less as a digression than as a reminder that a roadmap written largely against U.S. mobilization must be read and adapted against the compute, data, and network realities of each region that adopts it. We argue that top-down mobilization and bottom-up coordination are complementary rather than redundant. Large programs supply what a grassroots network cannot, including compute at scale, foundation

19

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

models, a security framing, and funding at scale [10, 85]. A grassroots network, in turn, supplies what large programs and commercial vendors tend to underweight, namely vendor-neutral interfaces and crossinstitutional standards that prevent each new program or proprietary platform from becoming another silo, together with the deliberate inclusion of resource-constrained institutions through portable, low-footprint laboratory modules. In this division of labor, an interconnected network such as AISLE is best understood not as an alternative to federal mobilization but as the interoperability fabric and community-standards layer that allows these initiatives to connect to one another and to the broader research community, rather than to interconnect only internally. We therefore see the roadmap of this paper as a contribution that is independent of but synergistic with the national and global programs now taking shape. A practical consequence is that the milestones we revise below should be read as community commitments that any of these programs can adopt, instrument, and report against, rather than as the agenda of a single institution. As concrete evidence that such adoption is feasible, the Genesis Mission consortium has organized its own work into groups for robotics and automation, data integration and standards, model development and validation, and computing infrastructure, which align respectively with the instrument, data, trust, and orchestration dimensions of Section 3 [92].

5

Conclusion and Revised Roadmap

In this paper, we presented an updated community roadmap for interconnected autonomous science, one year after the original AISLE roadmap [1]. We characterized seven shifts in the landscape, namely validated discoveries, a second generation of self-driving laboratories, a consolidating interface layer, reasoning and foundation models as a new substrate, benchmarks that expose a capability-reliability gap, the move of trust and governance to the foreground, and the entry of industry as a primary actor. We used two lenses, an autonomy ladder and that same gap, to interpret them. We refined the five original dimensions in light of this evidence, and we elevated two concerns, trust and verification (Section 3.6) and safety, security, and governance (Section 3.7), to first-class dimensions of the roadmap. Our milestone-by-milestone assessment is collected in the scorecard of Table 1, presented with the dimensions in Section 3. The pattern is consistent in that the dimensions closest to raw model capability have advanced the fastest, while those that require crossinstitutional infrastructure, verification, and governance remain largely open. Orchestration (Section 3.3) and agent interfaces (Section 3.4) moved the furthest; the former carried by reasoning and foundation models and the latter by the consolidation of agent protocols. Data management, education, and the federated aspects of instrument integration moved the least because they depend on coordination that no single model improvement can supply. The two dimensions we add, trust and governance, do not appear on the original scorecard precisely because the past year revealed them to be prerequisites rather than refinements. Read as a two-year plan, the first year concentrates on interfaces, protocol adoption, AI-driven metadata, and the scaffolding of verification, while the second year targets federation, zero-trust coordination, cross-facility knowledge integration, and governance. We deliberately keep the horizon short because, at the current pace of change, a longer-range plan would be obsolete before it could be acted upon. In future work, we plan to convene the community around this two-year roadmap to develop reference implementations that exercise the consolidated interface layer (Section 3.4) across multiple facilities and to prototype the verification, validation, and governance mechanisms that the past year has shown to be prerequisites, rather than refinements, for trustworthy autonomous discovery.

20

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

References [1] R. Ferreira da Silva, M. Abolhasani, D. A. Antonopoulos, L. Biven, R. Coffee, I. T. Foster, L. Hamilton, S. Jha, T. Mayer, B. Mintz et al., “A grassroots network and community roadmap for interconnected autonomous science laboratories for accelerated discovery,” in Workshop Proceedings of the 54th International Conference on Parallel Processing, 2025, pp. 142–150. [2] M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne et al., “The FAIR guiding principles for scientific data management and stewardship,” Scientific data, vol. 3, no. 1, pp. 1–9, 2016. [3] K. Swanson, W. Wu, N. L. Bulaong, J. E. Pak, and J. Zou, “The virtual lab of ai agents designs new sars-cov-2 nanobodies,” Nature, vol. 646, no. 8085, pp. 716–723, 2025. [4] J. Gottweis, W.-H. Weng, A. Daryin, T. Tu, A. Palepu, P. Sirkovic, A. Myaskovsky, F. Weissenberger, K. Rong, R. Tanno et al., “Towards an AI co-scientist,” arXiv preprint arXiv:2502.18864, 2025. [5] A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, J. M. Laurent, M. T. Razzak, A. D. White, M. M. Hinks, and S. G. Rodriques, “Robin: A multi-agent system for automating scientific discovery,” arXiv preprint arXiv:2505.13400, 2025. [6] Anthropic, “Introducing the model context protocol,” https://www.anthropic.com/news/model-context-protocol, 2024. [7] Linux Foundation, “Agent2agent (A2A) protocol,” https://a2a-protocol.org/, 2025. [8] A. Kamatar, J. G. Pauloski, Y. Babuji, R. Chard, M. Sakarvadia, D. Babnigg, K. Chard, and I. Foster, “Empowering scientific workflows with federated agents,” arXiv preprint arXiv:2505.05428, 2025. [9] H. Pan, R. Chard, R. Mello, C. Grams, T. He, A. Brace, O. P. Skelly, W. Engler, H. Holbrook, S. Y. Oh et al., “Experiences with model context protocol servers for science and high performance computing,” arXiv preprint arXiv:2508.18489, 2025. [10] The White House, “Launching the genesis mission,” Executive Order 14363, https://www.whitehouse.gov/ presidential-actions/2025/11/launching-the-genesis-mission/, 2025. [11] European Commission, “A european strategy for AI in science and the RAISE initiative,” https:// research-and-innovation.ec.europa.eu/, 2025. [12] Acceleration Consortium, “The acceleration consortium,” https://acceleration.utoronto.ca/, 2023. [13] N. J. Szymanski, B. Rendy, Y. Fei, R. E. Kumar, T. He, D. Milsted, M. J. McDermott, M. Gallant, E. D. Cubuk, A. Merchant et al., “Author correction: An autonomous laboratory for the accelerated synthesis of inorganic materials,” Nature, vol. 650, no. 8100, pp. 1–1, 2026. [14] J. Bragg, M. D’Arcy, N. Balepur, D. Bareket, B. Dalvi, S. Feldman, D. Haddad, J. D. Hwang, P. Jansen, V. Kishore et al., “Astabench: Rigorous benchmarking of ai agents with a scientific research suite,” arXiv preprint arXiv:2510.21652, 2025. [15] Z. Chen, S. Chen, Y. Ning, Q. Zhang, B. Wang, B. Yu, Y. Li, Z. Liao, C. Wei, Z. Lu et al., “Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 96 934–96 990. [16] M. Tian, L. Gao, S. D. Zhang, X. Chen, C. Fan, X. Guo, R. Haas, P. Ji, K. Krongchon, Y. Li et al., “Scicode: A research coding benchmark curated by scientists,” Advances in Neural Information Processing Systems, vol. 37, pp. 30 624–30 650, 2024. [17] S. Ansari, “Compound deception in elite peer review: A failure mode taxonomy of 100 fabricated citations at neurips 2025,” arXiv preprint arXiv:2602.05930, 2026. [18] T. Zheng, Z. Deng, H. T. Tsang, W. Wang, J. Bai, Z. Wang, and Y. Song, “From automation to autonomy: A survey on large language models in scientific discovery,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 17 744–17 761. [19] J. Wei, Y. Yang, X. Zhang, Y. Chen, X. Zhuang, Z. Gao, D. Zhou, G. Wang, Z. Gao, J. Cao et al., “From ai for science to agentic science: A survey on autonomous scientific discovery,” arXiv preprint arXiv:2508.14111, 2025. [20] Y. Yamada, R. T. Lange, C. Lu, S. Hu, C. Lu, J. Foerster, J. Clune, and D. Ha, “The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search,” arXiv preprint arXiv:2504.08066, 2025.

21

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

[21] H. Lee, H. J. Yoo, H. S. Jang, B. Park, Y. J. Park, and S. S. Han, “Toward self-driving laboratory 2.0 for chemistry and materials discovery,” Materials Horizons, vol. 13, no. 10, pp. 4712–4739, 2026. [22] G. Tom, S. P. Schmid, S. G. Baird, Y. Cao, K. Darvish, H. Hao, S. Lo, S. Pablo-Garcı́a, E. M. Rajaonson, M. Skreta et al., “Self-driving laboratories for chemistry and materials science,” Chemical Reviews, vol. 124, no. 16, pp. 9633–9732, 2024. [23] D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025. [24] C. Zeni, R. Pinsler, D. Zügner, A. Fowler, M. Horton, X. Fu, Z. Wang, A. Shysheya, J. Crabbé, S. Ueda et al., “A generative model for inorganic materials design,” Nature, vol. 639, no. 8055, pp. 624–632, 2025. [25] J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Ballard, J. Bambrick et al., “Accurate structure prediction of biomolecular interactions with alphafold 3,” Nature, vol. 630, no. 8016, pp. 493–500, 2024. [26] Z. S. Siegel et al., “CORE-Bench: Fostering the credibility of published research through a computational reproducibility agent benchmark,” 2024. [27] M. Z. Pan et al., “Measuring agents in production,” 2025. [28] Y. He and Y. Bu, “Academic journals’ ai policies fail to curb the surge in ai-assisted academic writing,” Proceedings of the National Academy of Sciences, vol. 123, no. 9, p. e2526734123, 2026. [29] A. V. Tobias and A. Wahab, “Autonomous ‘self-driving’laboratories: a review of technology and policy implications,” Royal Society Open Science, vol. 12, no. 7, p. 250646, 2025. [30] Lila Sciences, “Lila sciences raises $235m series a to advance scientific superintelligence and autonomous laboratories,” https://www.lila.ai/, 2025. [31] Andreessen Horowitz, “Investing in Periodic Labs,” https://a16z.com/announcement/investing-in-periodic-labs/, 2025. [32] Microsoft, “Transforming R&D with agentic AI: Introducing Microsoft Discovery,” https://azure.microsoft.com/ en-us/blog/transforming-rd-with-agentic-ai-introducing-microsoft-discovery/, 2025. [33] NVIDIA, “NVIDIA and U.S. government to boost AI infrastructure and r&d investments,” https://blogs.nvidia. com/blog/nvidia-us-government-to-boost-ai-infrastructure-and-rd-investments/, 2025. [34] M. Gridach, J. Nanavati, K. Z. E. Abidine, L. Mendes, and C. Mack, “Agentic ai for scientific discovery: A survey of progress, challenges, and future directions,” arXiv preprint arXiv:2503.08979, 2025. [35] R. Ferreira da Silva, R. Moore II, B. Mintz, R. Advincula, A. Alnajjar, L. Baldwin, C. A. Bridges, R. Coffee, E. Deelman, C. Engelmann et al., “Shaping the future of self-driving autonomous laboratories workshop,” Oak Ridge National Laboratory, Tech. Rep. ORNL/TM-2024/3714, 2024. [36] Argonne AD-SDL, “The modular autonomous discovery for science (MADSci) framework,” https://github.com/ AD-SDL/MADSci, 2025. [37] M. Sim, M. G. Vakili, F. Strieth-Kalthoff, H. Hao, R. J. Hickman, S. Miret, S. Pablo-Garcı́a, and A. Aspuru-Guzik, “Chemos 2.0: An orchestration architecture for chemical self-driving laboratories,” Matter, vol. 7, no. 9, pp. 2959–2977, 2024. [38] E. J. Brass, S. Veeramani, Z. Zhou, H. Fakhruldeen, J. S. Manzano, R. Clowes, I. Akpinar, M. R. Ward, J. W. Ward, and A. I. Cooper, “A mobile robotic process chemist,” Digital Discovery, vol. 5, no. 3, pp. 1363–1371, 2026. [39] SiLA Consortium, “SiLA 2: Standardization in lab automation,” https://sila-standard.com/, 2024. [40] National Institute of Standards and Technology, “Development of standards to support a modular and autonomous laboratory ecosystem,” https://www.nist.gov/programs-projects/ development-standards-support-modular-and-autonomous-laboratory-ecosystem, 2025. [41] Oak Ridge National Laboratory, “INTERSECT: Interconnected science ecosystem,” https://www.ornl.gov/ intersect, 2025. [42] Emerald Cloud Lab, “Emerald cloud lab and the Symbolic Lab Language,” https://www.emeraldcloudlab.com/, 2024. [43] Carnegie Mellon University, “Carnegie mellon to build first-of-its-kind university cloud lab,” https://www.cmu. edu/news/stories/archives/2021/august/first-academic-cloud-lab.html, 2021.

22

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

[44] D. A. Boiko, R. MacKnight, B. Kline, and G. Gomes, “Autonomous chemical research with large language models,” Nature, vol. 624, no. 7992, pp. 570–578, 2023. doi: 10.1038/s41586-023-06792-0 [45] National Science Foundation, “Test bed: Toward a network of programmable cloud laboratories (PCL test bed),” NSF 25-541, https://www.nsf.gov/funding/opportunities/ pcl-test-bed-test-bed-toward-network-programmable-cloud-laboratories, 2025. [46] A. M. M. Scaife, “Big telescope, big data: towards exascale with the Square Kilometre Array,” Philosophical Transactions of the Royal Society A, vol. 378, p. 20190060, 2020. doi: 10.1098/rsta.2019.0060 [47] HEP Software Foundation, “A roadmap for HEP software and computing R&D for the 2020s,” Computing and Software for Big Science, vol. 3, p. 7, 2019. doi: 10.1007/s41781-018-0018-8 [48] A. M. Deiana, N. Tran, J. Agar, M. Blott, G. Di Guglielmo, J. Duarte, P. Harris, S. Hauck, M. Liu, M. S. Neubauer et al., “Applications and techniques for fast machine learning in science,” Frontiers in Big Data, vol. 5, p. 787421, 2022. doi: 10.3389/fdata.2022.787421 [49] A. V. Babu, T. Zhou, S. Kandel, T. Bicer, Z. Liu, W. Judge, D. J. Ching, Y. Jiang, S. Veseli, S. Henke et al., “Deep learning at the edge enables real-time streaming ptychographic imaging,” Nature Communications, vol. 14, p. 7059, 2023. doi: 10.1038/s41467-023-41496-z [50] Z. Jiang, B. Carlson, A. Deiana, J. Eastlack, S. Hauck, S.-C. Hsu, R. Narayan, S. Parajuli, D. Yin, and B. Zuo, “Machine learning evaluation in the Global Event Processor FPGA for the ATLAS trigger upgrade,” Journal of Instrumentation, vol. 19, p. P05031, 2024. doi: 10.1088/1748-0221/19/05/P05031 [51] W. Shin, C. A. Bridges, M. T. McDonnell, and R. Ferreira da Silva, “Fantastic scientific agents and how to build them: AgentBuild for Rietveld refinement,” arXiv preprint arXiv:2606.12834, 2026. [52] S. K. Moore, “This self-driving lab uses a digital twin to speed discovery,” IEEE Spectrum, https://spectrum.ieee. org/autonomous-lab-argonne-polybot, 2025. [53] U.S. Department of Energy, Office of Science, “Integrated research infrastructure architecture blueprint activity: Final report 2023,” U.S. Department of Energy, Tech. Rep., 2023. [54] R. Chard, J. Pruyne, K. McKee, J. Bryan, B. Raumann, R. Ananthakrishnan, K. Chard, and I. T. Foster, “Globus automation services: Research process automation across the space–time continuum,” Future Generation Computer Systems, vol. 142, pp. 393–409, 2023. [55] J. G. Pauloski, K. Rydzy, V. Hayot-Sasson, I. Foster, and K. Chard, “Accelerating Python applications with Dask and ProxyStore,” arXiv preprint arXiv:2410.12092, 2024. [56] T. Lebo et al., “PROV-O: The PROV ontology,” W3C Recommendation, https://www.w3.org/TR/prov-o, 2013. [57] P. Xu, X. Ji, M. Li, and W. Lu, “Small data machine learning in materials science,” npj Computational Materials, vol. 9, p. 42, 2023. doi: 10.1038/s41524-023-01000-z [58] M. Haghighatlari, J. Li, F. Heidar-Zadeh, Y. Liu, X. Guan, and T. Head-Gordon, “Learning to make chemical predictions: The interplay of feature representation, data, and machine learning methods,” Chem, vol. 6, no. 7, pp. 1527–1542, 2020. doi: 10.1016/j.chempr.2020.05.014 [59] W. Brewer, P. Widener, V. Anantharaj, F. Wang, T. Beck, A. Shankar, and S. Oral, “Data readiness for scientific AI at scale,” 2025. [60] H. Joress, Z. Trautt, A. McDannald, B. DeCost, A. G. Kusne, and F. Tavazza, “Driving U.S. innovation in materials and manufacturing using AI and autonomous labs,” National Institute of Standards and Technology, Tech. Rep. NIST Special Publication 1320, 2024. [61] D. P. Tabor, L. M. Roch, S. K. Saikin, C. Kreisbeck, D. Sheberla, J. H. Montoya, S. Dwaraknath, M. Aykol, C. Ortiz, H. Tribukait et al., “Accelerating the discovery of materials for clean energy in the era of smart automation,” Nature Reviews Materials, vol. 3, pp. 5–20, 2018. doi: 10.1038/s41578-018-0005-z [62] National Science and Technology Council, “Materials genome initiative strategic plan,” https://www.mgi.gov, 2021. [63] S. Liu, S. Ponnapalli, S. Shankar, S. Zeighami, A. Zhu, S. Agarwal, R. Chen, S. Suwito, S. Yuan, I. Stoica, M. Zaharia, A. Cheung, N. Crooks, J. E. Gonzalez, and A. G. Parameswaran, “Supporting our AI overlords: Redesigning data systems to be agent-first,” 2025. [64] A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller, “Augmenting large language models with chemistry tools,” Nature machine intelligence, vol. 6, no. 5, pp. 525–535, 2024.

23

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

[65] T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths, “Cognitive architectures for language agents,” Transactions on Machine Learning Research, 2024, arXiv:2309.02427. [66] C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez, “MemGPT: Towards LLMs as operating systems,” 2023. [67] R. Roscher, B. Bohn, M. F. Duarte, and J. Garcke, “Explainable machine learning for scientific insights and discoveries,” IEEE Access, vol. 8, pp. 42 200–42 216, 2020. doi: 10.1109/ACCESS.2020.2976199 [68] F. Oviedo, J. L. Ferres, T. Buonassisi, and K. T. Butler, “Interpretable and explainable machine learning for materials science and chemistry,” Accounts of Materials Research, vol. 3, no. 6, pp. 597–607, 2022. doi: 10.1021/accountsmr.1c00244 [69] A. Ehtesham, A. Singh, G. K. Gupta, and S. Kumar, “A survey of agent interoperability protocols: MCP, ACP, A2A, and ANP,” 2025. [70] Anthropic, “Equipping agents for the real world with Agent Skills,” https://www.anthropic.com/engineering/ equipping-agents-for-the-real-world-with-agent-skills, 2025. [71] ——, “Code execution with MCP: Building more efficient agents,” https://www.anthropic.com/engineering/ code-execution-with-mcp, 2025. [72] National Security Agency, “Model context protocol (MCP),” Cybersecurity Information Sheet U/OO/6030316-26, https://www.nsa.gov/Portals/75/documents/Cybersecurity/CSI MCP SECURITY.pdf, 2026. [73] R. Souza, A. Gueroudji, S. DeWitt, D. Rosendo, T. Ghosal, R. Ross, P. Balaprakash, and R. F. Da Silva, “PROVAGENT: Unified provenance for tracking ai agent interactions in agentic workflows,” in 2025 IEEE International Conference on eScience (eScience). IEEE, 2025, pp. 467–473. [74] Q. Hao, F. Xu, Y. Li, and J. Evans, “Artificial intelligence tools expand scientists’ impact but contract science’s focus,” Nature, 2025. doi: 10.1038/s41586-025-09922-y [75] M. Juelsholt et al., “Continued challenges in high-throughput materials predictions: MatterGen predicts compounds from the training dataset,” Materials Horizons, 2026. doi: 10.1039/D6MH00268D [76] N. J. Szymanski et al., “An autonomous laboratory for the accelerated synthesis of novel materials,” Nature, vol. 624, no. 7990, pp. 86–91, 2023. doi: 10.1038/s41586-023-06734-w [77] National Institute of Standards and Technology, “Artificial intelligence risk management framework (AI RMF 1.0),” NIST AI 100-1, https://www.nist.gov/itl/ai-risk-management-framework, 2023. [78] R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni, “Green AI,” Communications of the ACM, vol. 63, no. 12, pp. 54–63, 2020. doi: 10.1145/3381831 [79] J. Pannu, D. Bloomfield, R. MacKnight, M. S. Hanke, A. Zhu, G. Gomes, A. Cicero, and T. V. Inglesby, “Dual-use capabilities of concern of biological ai models,” PLoS computational biology, vol. 21, no. 5, p. e1012975, 2025. [80] O. G. S. Project, “Llm02:2025 – sensitive information disclosure,” https://genai.owasp.org/llmrisk/ llm022025-sensitive-information-disclosure/, 2025, oWASP Top 10 for LLM Applications. [81] N. Shapira, C. Wendler, A. Yen, G. Sarti, K. Pal, O. Floody, A. Belfki, A. Loftus, A. R. Jannali, N. Prakash, J. Cui, G. Rogers, J. Brinkmann, C. Rager, A. Zur, M. Ripa, A. Sankaranarayanan, D. Atkinson, R. Gandikota, J. Fiotto-Kaufman, E. Hwang, H. Orgad, P. S. Sahil, N. Taglicht, T. Shabtay, A. Ambus, N. Alon, S. Oron, A. Gordon-Tapiero, Y. Kaplan, V. Shwartz, T. R. Shaham, C. Riedl, R. Mirsky, M. Sap, D. Manheim, T. Ullman, and D. Bau, “Agents of chaos,” 2026. [Online]. Available: https://arxiv.org/abs/2602.20021 [82] Y.-C. J. Lee and B. Del Castello, “Documenting cloud labs and examining how remotely operated automated laboratories could enable bad actors,” RAND Corporation, Tech. Rep. PE-A3851-1, 2024. [83] I. C. Ngong, K. Murugesan, S. Kadhe, J. D. Weisz, A. Dhurandhar, and K. N. Ramamurthy, “Agentscope: Evaluating contextual privacy across agentic workflows,” 2026. [Online]. Available: https://arxiv.org/abs/2603.04902 [84] U.S. Department of Energy, “Energy department announces 26 Genesis Mission science and technology challenges,” https://www.energy.gov/articles/ energy-department-announces-26-genesis-mission-science-and-technology-challenges, 2026. [85] Trillion Parameter Consortium, “Trillion parameter consortium (TPC),” https://www.anl.gov/cels/ trillion-parameter-consortium, 2024.

24

TOWARD TRUSTWORTHY AUTONOMOUS SCIENCE: A TWO-YEAR COMMUNITY ROADMAP

[86] U.S. Department of Energy, “Energy department announces collaboration agreements with 24 organizations to advance the Genesis Mission,” https://www.energy.gov/articles/ energy-department-announces-collaboration-agreements-24-organizations-advance-genesis, 2025. [87] National Science Foundation, “Directorate for technology, innovation and partnerships (TIP),” https://www.nsf. gov/tip/about-tip, 2026. [88] ——, “NSF X-Labs,” https://www.nsf.gov/funding/initiatives/nsf-x-labs, 2026. [89] Defense Advanced Research Projects Agency, “Biological technologies office (BTO),” https://www.darpa.mil/ about/offices/bto, 2026. [90] U.S. Department of Energy, Office of Science, “Robotics and automation testbeds for autonomous scientific discovery,” DOE National Laboratory Announcement LAB 26-3601, https://science.osti.gov/grants/Lab-Announcements/ Open, 2026. [91] EuroHPC Joint Undertaking, “JUPITER: Launching europe’s exascale era,” https://www.eurohpc-ju.europa.eu/ jupiter-launching-europes-exascale-era-2025-09-05 en, 2025. [92] U.S. Department of Energy, “Genesis Mission collaboration,” https://www.energy.gov/undersecretaryforscience/ genesis-mission/genesis-mission-collaboration, 2026.

25

Record · ID 366232 · SHA-256 b8d86a1f6e291344
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.