ConceptioArchivearXiv CS
arXiv CSopen access

Hephaestus: Toward a Cybersecurity AI Scientist

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

Hephaestus: Toward a Cybersecurity AI Scientist∗

arXiv:2606.29981v1 [cs.CR] 29 Jun 2026

Jiaqi Li, Yang Zhao, Wen Lu, Lvyang Zhang, and Lidong Zhai Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China [email protected] June 26, 2026

Abstract Cyber offense is moving to machine speed; cyber research itself is not. Existing AI scientist systems make end-to-end research automation increasingly plausible, but they target relatively stable scientific domains. We argue that AI-native cybersecurity is a different kind of scientific object. Its recurring units of study are security events and interaction traces, not static assets; its model and tool substrate is non-stationary, not steady-state; and credible evaluation depends on digital twins, cyber ranges, and auditable evidence rather than on a single benchmark score. We call this object the Cybersecurity AI Scientist. A practical realization is a modular, rolespecialized multi-agent research system that coordinates problem framing, threat modeling, tool generation, controlled experimentation, evaluation, governance, and scientific reporting, and that anchors its concrete objectives in a four-zeros frame spanning risk, trust, incident, and energy dimensions. As a representative agenda we focus on AI-native defense, where steady-state perimeters give way to resilient agent legions and the classical category of terminal security is itself being deconstructed into agent security. This paper defines the object, separates it from any single organizational realization, and offers an architecture and an agenda on which later systems, benchmarks, and empirical programs can be built.

1

Introduction

Two trajectories have begun to converge on cybersecurity. On one side, autonomous AI agents are already pushing into operational territory: language-model agents exploit real one-day vulnerabilities, and team-of-agent systems improve performance on harder zero-day scenarios [13, 36]. On the other side, AI scientist systems are pushing research itself toward automation, with end-to-end pipelines that generate ideas, write code, run experiments, draft papers, and conduct simulated review [24, 34], and with domain scientist platforms in biology and biomedicine that organize multi-agent hypothesis generation, experiment planning, and analysis [3, 16, 17, 27]. These two trajectories have so far progressed in parallel. AI scientist work has matured on relatively stable scientific objects, where the system under study does not adapt to the act of being studied. Cybersecurity agent work, in turn, has concentrated on task execution and benchmark performance, ranging from autonomous penetration testing and exploitation to evaluation of cyberexpert capability and the security of the agents themselves [19–21, 23, 31]. Cybersecurity has therefore become a frontier arena for AI capability and control without yet being framed as a ∗

Hephaestus, the divine artisan of Homeric tradition, is the maker of both spear and shield. The name marks a research system that produces both offensive and defensive scientific artifacts under one design.

1

scientist-level research object in its own right. Recent institutional efforts, including Google’s Big Sleep and DARPA’s AI Cyber Challenge, make clear that autonomous cyber reasoning and AI-native defense are no longer hypothetical [9, 10, 29]. We argue that this gap is not bridged by a simple domain transfer. The cybersecurity research substrate is unusually unforgiving for generic AI scientist pipelines. Adversaries adapt to the very capabilities used to study them. Dual-use pressure is structural rather than peripheral. Model platforms, guardrails, and tool affordances drift on a timescale shorter than the research loop. And the validity of conclusions depends on environments, digital twins, and evidence chains that are themselves part of the method. What is needed is not a renamed agent pipeline but a cybersecurity-specific research object. We call this object the Cybersecurity AI Scientist. It denotes the kind of research capability that becomes necessary when scientific-problem shift and research-paradigm shift in cybersecurity are taken seriously at the same time. At the systems level, one practical realization is a modular, role-specialized multi-agent research system that coordinates problem framing, literature and threat analysis, tool generation, digital-twin experimentation, evaluation, governance-aware analysis, and scientific reporting through specialized agents, reusable capability packages, and bounded tools. We adopt and extend the end-to-end scientific pipeline of The AI Scientist line of work [24, 34]: we keep the closed-loop ideation-execution-reporting structure, but we replace its assumption of a steady computational substrate with an adversarial, non-stationary, multi-model substrate, and we make governance and environment design first-class parts of the method. Contributions. 1. We define the Cybersecurity AI Scientist as a distinct research object rather than a domain-specific relabeling of generic AI scientist systems. 2. We characterize the coupled problem shift and research-paradigm shift induced by AI-native security conditions, and we anchor concrete objectives in a four-zeros frame spanning risk, trust, incident, and energy dimensions. 3. We give a framework architecture and research pipeline, and present a modular multi-agent research system organized around specialized roles and reusable capability packages as one plausible organizational realization. 4. We identify AI-native defense, including resilient agent legions and the deconstruction of terminal security into agent security, as a representative forward-looking research agenda. Figure 1 and Table 3 summarize the coupled shifts and the resulting design pressures.

2

Related Work

Our work sits at the intersection of AI for research, autonomous AI scientist systems, domain-specific scientific discovery systems, and cybersecurity agents and benchmarks. The closest prior literature contributes important building blocks, but does not yet treat cybersecurity as a scientist-level research object.

2.1

AI for research and AI scientist surveys

Recent surveys organize the broader landscape of AI for scientific research and AI scientist systems, covering hypothesis discovery, experiment planning, scientific writing, peer review, and the end2

to-end research pipeline [6, 25, 28]. These works clarify how AI is reshaping research as a whole, but they remain domain-agnostic and do not engage with cybersecurity’s adversarial and dual-use constraints.

2.2

General autonomous scientist systems

The AI Scientist demonstrated a closed-loop pipeline from ideation to paper drafting and simulated review, primarily in machine-learning subfields. The AI Scientist-v2 advances autonomy through agentic tree search and lower template dependence [24, 34]. These systems show that end-to-end AI-driven research is feasible, but they are calibrated for stable computational settings rather than for adversarial research substrates.

2.3

Domain-specific scientist systems

Co-Scientist frames scientific collaboration as a multi-agent hypothesis-generation and proposaldevelopment system for biomedical discovery, and Robin extends the trajectory toward hypothesis generation, experimentation, and analysis in biology [16, 17]. PaperQA2 shows strong scientific literature-agent performance on realistic synthesis tasks, and ERA targets expert-level empirical software generation for computational discovery [3, 27]. These works validate domain scientist systems, but their scientific objects are largely non-adversarial.

2.4

Cybersecurity agents, benchmarks, and evaluation

A separate line has rapidly advanced cybersecurity task automation. PentestGPT and HackSynth opened autonomous penetration testing, and follow-up work showed that LLM agents can exploit real one-day vulnerabilities and that agent teams can improve zero-day exploitation [11, 13, 26, 36]. More recent systems push orchestration, planning, tooling, and red-team coordination further [8, 12, 19, 23], while CVE-Bench, CyberGym, and multi-step attack benchmarks extend evaluation toward realistic scenarios at scale [14, 22, 33, 37]. In parallel, evaluation and safety research has consolidated: CyberSecEval 2/3 and related work broaden cybersecurity evaluation, the Digital Cybersecurity Expert and the Cyber Defense Benchmark study role-aligned capability and openended defense, OpenSec targets calibration under adversarial evidence, ExploitGym benchmarks weaponized exploits at scale, and dedicated work studies offensive-security benchmarking practices and the security of agent architectures themselves [4, 5, 7, 18, 20, 30–32].

2.5

Frontier cyber-capable models and operational systems

Three more recent developments mark a qualitative shift and deserve separate treatment. First, frontier models have crossed a capability threshold: Anthropic disclosed Claude Mythos Preview under Project Glasswing, a frontier model whose offensive cyber capability was strong enough to warrant restricted release through a vetted partner program, and subsequent reporting attributes large-scale defect discovery in widely deployed software to that model [1, 2]. Second, vertical foundation models for security have begun to appear, with the Llama-3.1-based Foundation Security LLM offering an explicit cybersecurity-specialized base [35]. Third, operational AI-assisted security has moved from prototype to deployment: Microsoft’s Guided Response architecture for Security Copilot documents an in-production AI system built around incident triage, action recommendation, and similar-incident retrieval over a large real-world incident corpus [15]. These are not isolated benchmarks; they are evidence that frontier cyber-capable models, vertical foundation models, and

3

Line

Typical research object

Main output

Environment assumption

Governance pressure

What it still misses here

AI Scientist

General scientific workflow

Ideas, code, experiments, papers

Relatively stable computational setting

Usually external or weakly embedded

Domain Scientist

Biology, biomed, computational discovery

Hypotheses, plans, analyses

Domain-rich but mostly non-adversarial setting

Safety matters, but usually not cyber dual-use

Cyber agents and benchmarks

Exploitation, pentest, threat hunting, cyber-role capability

Task success, benchmark scores, trajectories

Operational or benchmark environment

High, but usually task-level

Cybersecurity AI Scientist

Cybersecurity as an adversarial scientific object

Research designs, tools, traces, evaluations, governance analyses, papers

Digital twin, cyber range, and evidence pipeline as part of method

Native and internal to workflow

Adversarial cyber object, dual-use-native workflow, digital-twin validation AI-versus-AI co-evolution, mutable cyber substrate, bounded cyber permissions Question-topaper research system, scientist-level role-capabilityartifact stack The target object of this paper

Table 1: Adjacent research lines that the Cybersecurity AI Scientist distinguishes itself from. operational AI security systems are all crossing into program-scale reality at the same time, which is precisely the regime a Cybersecurity AI Scientist must be designed for.

2.6

Gap

Taken together, this body of work shows that cyber-agent research is already rich and methodologically dense, while AI scientist research is already mature for stable domains. What is still missing is the joint object. Existing cyber work centers on task execution, benchmarking, or agent safety; existing AI scientist work centers on stable, non-adversarial objects. Neither line treats cybersecurity as an organized scientific process in its own right. We position the Cybersecurity AI Scientist exactly in this gap.

3

The Coupled Shifts: Problem and Paradigm

Classical cybersecurity research was largely organized around human attackers, human defenders, and human analysts, with relatively stable software tools and human-paced adaptation loops. A substantial portion of the field was therefore built around human bottlenecks: expert scarcity, slow tool development, manual experiment design, and delayed evidence synthesis. AI-native conditions disturb this picture from two directions at once.

3.1

Problem shift

As autonomous systems increasingly participate in offensive and defensive processes, an important subset of cybersecurity problems becomes AI-versus-AI rather than human-versus-human. Some 4

questions about analyst behavior or workflow efficiency collapse into engineering questions about orchestration and policy tuning. At the same time, new scientific questions emerge: how defensive heuristics accumulate in agent societies, how adversarial adaptation behaves under machine-speed iteration, how mutable model ecosystems reshape risk, and what counts as reliable evidence when both sides are partially autonomous. The natural unit of study moves with these questions. Security events, interaction traces, and dynamic scenarios become a more honest research handle than fixed assets and one-shot tasks. Problem definitions increasingly need to capture adaptation, coordination, containment, and evidence generation, not only task success [6, 7, 21].

3.2

Paradigm shift

The same conditions change how research is organized. Cybersecurity research moves from a humanled mode, through human-AI collaboration, toward AI-led execution under human agenda setting, where small numbers of senior researchers define questions, constraints, and acceptance criteria while large portions of execution are delegated to coordinated research agents [6, 28]. The transition is sharper in cybersecurity than in most other domains because the substrate itself is unstable: model access policies change, tool affordances shift, guardrails evolve, evaluation environments drift, and the threat landscape mutates in response to newly available capabilities. A method that works today may become unusable tomorrow; a workflow that currently spans multiple stages may collapse into one. The substrate is non-stationary rather than steady-state, and the pipeline must be re-calibrated continuously. The paradigm shift is also rarely a single-model story. One model may be stronger at exploitation planning, another at code synthesis, another at long-context evidence handling, and another at operating safely under stronger policy constraints. Existing evidence already shows that different models exhibit different role-aligned knowledge gaps and different end-to-end penetration-testing behavior [12, 31]. A credible Cybersecurity AI Scientist therefore cannot be conceived as a monolithic model endpoint. It is better understood as a calibrated orchestration layer over multiple models, tools, and evaluators whose complementary strengths must be measured and routed rather than assumed away. Governance moves with this same logic: in many scientific domains governance appears mainly as an external review layer, but in AI-native security research it becomes an internal design constraint, where permission scopes, dual-use containment, experiment isolation, and release boundaries belong inside the research method rather than around it.

4

Defining the Cybersecurity AI Scientist

We define the Cybersecurity AI Scientist as the cybersecurity-specific AI-native research object and framework that emerges once scientific-problem shift and research-paradigm shift are taken seriously together. It names the kind of research capability that becomes necessary when the object of study is adaptive, high-stakes, and tightly coupled to rapidly evolving technical infrastructure. The definition is deliberately separated from any particular organizational realization. One plausible and tractable realization is a modular multi-agent research system, in which reusable capabilities appear as bounded operators or procedural packages. Other realizations are possible; the concept does not depend on a single implementation idiom. This separation matters. The scientific object answers what must be studied and what kind of research system must exist. The organizational realization answers how such a system may be instantiated. The goal is therefore not to build a single omnipotent agent that performs every research task, and not to merely create a more capable offensive or defensive operator. It is to

5

Figure 1: Cybersecurity AI Scientist is motivated by two coupled shifts. The scientific object changes, and the organization of research changes with it. organize cybersecurity research as a coordinated system of specialized roles, each operating over reusable capability packages, explicit tools, bounded permissions, and structured outputs. Under this definition, the Cybersecurity AI Scientist differs from generic AI scientist systems in at least four ways. First, it is designed for adversarial and dual-use conditions rather than stable experimental settings. Second, it must operate over a non-stationary substrate, interfacing with cyber ranges, digital twins, mutable platforms, and evidence-sensitive evaluation. Third, it should coordinate heterogeneous models and evaluators rather than presume that a single model can stably cover the workflow. Fourth, it must support not only empirical execution but also governance-aware reasoning, threat framing, and strategic hypothesis development. It is also distinct from pure attack or defense automation: a penetration-testing agent, a vulnerability-exploitation system, or a threat-hunting assistant may be useful components inside a larger framework, but none of them by itself constitutes a Cybersecurity AI Scientist. The defining feature is not task specialization but research-system organization—the ability to move from question definition to experimental design, tool assembly, controlled execution, evaluation, interpretation, and scientific reporting as a coherent process.

4.1

Objectives: a four-zeros frame

Concept formation is not enough; a Cybersecurity AI Scientist also needs a clear sense of what it is for. We anchor its objectives in a four-zeros frame that spans the principal dimensions along which AI-native security risk accumulates. Each zero names a distinct kind of failure that this research system is expected to study, reduce, and ultimately bound. Together they keep the framework from collapsing into either pure technology language or pure governance language. Table 2 summarizes the frame.

6

Dimension

Question type

Target failure mode

Primary focus

Implication for the research system

Risk

Scientific

Zero hidden defects (in systems)

Machine system

Trust

Technical

Zero implicit trust (against human error)

Human

Incident

Operational

Zero incidents (from operational fault)

Event

Energy

Ecological

Zero loss (in organizational outcomes)

Enterprise / organization

Event-driven inquiry into how flaws accumulate, propagate, and persist under AI-mediated change Calibration, role-aligned evaluation, and assistance designs that do not silently shift control Scenario-based research, replayable environments, evidence pipelines, and live-fire validation Long-horizon agenda continuity, strategic framing, and institutional governance research

/

Table 2: The four-zeros frame for the Cybersecurity AI Scientist. The four dimensions span scientific, technical, operational, and ecological failure modes, and they jointly fix the research targets that the framework must serve.

5

Framework Architecture

A systems realization of the framework can be read both as a layered architecture and as a research pipeline. The pipeline begins with agenda setting—strategic questions, scenario definitions, and governance constraints—and then moves through problem framing, literature and threat analysis, toolsmithing, controlled execution, evaluation, and scientific reporting. Across this flow sits a role layer of specialized agents: problem-framing agents, literature agents, threat-modeling agents, toolsmithing agents, experiment managers, evaluators, and reporting agents. These agents do not operate as free-form generalists. They invoke bounded procedures that encode reusable capabilities and stable interaction patterns. Beneath the role layer sits a capability-packaging layer. Reusable capabilities are packaged as bounded operators or procedural modules with explicit input–output contracts, operational assumptions, permission scopes, and reusable routines. Three properties follow. Coordination becomes more modular and auditable, because behavior is exposed as named, contracted units rather than buried in agent policies. Partial reuse across projects becomes natural, because capabilities can be lifted without collapsing all behavior into one monolithic agent. And capability packages give governance a clean interface, because permissions can be scoped at the package level and exposure can be controlled selectively. The next layer is the tool and runtime layer, where concrete execution happens: code generation, experiment orchestration, model routing, data handling, judge and evaluator pipelines, and runtime control logic. For cybersecurity, this layer must be coupled with an environment layer of digital twins, cyber ranges, sandboxes, replay systems, and evidence-capture pipelines. The environment is not a passive container. It is part of the research method, because the validity of cyber findings depends heavily on containment, observability, replayability, and scenario realism. Weak environments produce misleading conclusions even when nominal benchmark outcomes look strong. The framework terminates in an artifact layer whose outputs include not only task completions but also research questions, threat models, tools, configurations, experiment traces, evaluation summaries, governance analyses, and paper-ready narrative. The Cybersecurity AI Scientist is not defined by a single score or benchmark outcome; its value lies in structured research production across the full lifecycle of inquiry. Read as design implications, four capability classes recur across these layers and any serious realization will need all of them: agenda continuity, calibration and 7

Design pressure

Why a generic AI scientist pipeline is insufficient

Required response in Cybersecurity AI Scientist

Adaptive adversaries

The object of study reacts to new capabilities and evaluation setups Static tasks underdescribe how cyber research is actually triggered and validated Models, guardrails, tools, and workflows can change during the research loop No single model is reliable across planning, coding, evidence handling, and safety-constrained operation Critical infrastructure, enterprise systems, platforms, and services impose different defensive constraints Useful research capabilities can be repurposed offensively Credible cyber claims depend on replayability, observability, and controlled environments

Threat-aware framing, iterative red-blue evaluation, and bounded deployment Event-driven problem formulation and trace-centered artifacts Modular orchestration, rapid reconfiguration, and recalibration Model routing, evaluator diversity, and role-specialized coordination

Security events as units of study Non-stationary substrate Multi-model capability asymmetry Heterogeneous fense targets

de-

Dual-use pressure Evidence-intensive evaluation

Target-aware scenario design and domain-specific evaluation pathways Permission scopes, containment, auditability, and release boundaries Digital twins, cyber ranges, evidence capture, and governance-linked validation

Table 3: Design pressures that separate the Cybersecurity AI Scientist from a straightforward domain transfer of generic AI scientist pipelines. control, high-integrity environments, and rapid tool formation.

6

Representative Agendas under the Four-Zeros Frame

The four-zeros frame introduced in Section 4 is not a decorative summary; it is the spine on which concrete research agendas hang. Each axis names a class of failure the Cybersecurity AI Scientist must study, organize, and ultimately bound. Each axis also corresponds to a forward-looking research agenda that is already in motion in the wider community and is sharper as a Cybersecurity AI Scientist question than as a single benchmark or single-agent problem. We walk the four axes in turn and then return to what binds them.

6.1

Risk axis: Hidden-defect discovery under AI-mediated change

Recent months have made it clear that frontier and vertical-cyber models can find real defects at a scale that the earlier benchmark literature was not designed to capture. Anthropic’s Claude Mythos Preview, released under Project Glasswing, was withheld from general access specifically because its offensive capability was strong enough to require a vetted partner program; reported deployment has attributed large-scale defect discovery in widely deployed software, including very long-lived flaws, to that model [1, 2]. Google’s Big Sleep and DARPA’s AIxCC make the same point at program scale [10, 29], and the emergence of vertical foundation models for security shows that cyber-specialized substrates are now first-class [35]. Benchmarks have moved with this capability shift: CyberGym evaluates agents over more than a thousand real-world vulnerabilities in nearly two hundred open-source projects, with frontier models already reporting tens of percent single-trial success and standalone discovery of fresh zero-days, while multi-agent web-pentest systems push end-to-end exploitation [8, 32, 33, 37]. For a Cybersecurity AI Scientist, the research question on this axis is not “can model M find defect X?” It is how to organize an event-driven research loop in which model capability, scenario

8

Figure 2: One operational view of the Cybersecurity AI Scientist. The framework is not a single agent loop, but a governance-constrained research pipeline with role specialization, model routing, controlled environments, and evidence-producing outputs. construction, evidence handling, and governance are coordinated across many defect classes and a long research horizon. Hidden-defect discovery becomes an ongoing scientific process rather than a sequence of one-shot capability demonstrations.

6.2

Trust axis: Calibrated assistance, multi-model routing, and role-aligned evaluation

The Trust axis concerns the conditions under which human operators can rely on a partly autonomous research and defense system without silently surrendering control. The most explicit operational evidence here is Microsoft’s Guided Response architecture for Security Copilot, which documents a production AI system organized around incident summarization, action recommendation, similarincident retrieval, and grading—and which makes calibration against analyst behavior a first-class concern [15]. Research-side work on role-aligned cyber capability shows that different models exhibit different blind spots and different end-to-end pentest behaviors [12, 31], which is exactly why the framework treats multi-model orchestration as native rather than optional. The Cybersecurity AI Scientist question here is when AI assistance starts to shift control away from operators in ways that the operators cannot detect, and how the research system itself should be designed to surface that shift rather than absorb it.

6.3

Incident axis: Resilient agent legions and the deconstruction of terminal security

The Incident axis is where the previous version of this paper’s flagship example lives, and we keep it as a worked agenda because it makes the framework concrete. Traditional defense models assume relatively stable perimeters, department-shaped responsibility, and human-paced repair loops, and many enterprise defenses are still optimized for steady-state business environments. Those assumptions weaken once both offense and defense are mediated by autonomous or semiautonomous agents, once attack surfaces shift faster than patch cycles, and once the very category

9

of an “application” begins to dissolve into transient agent-generated behavior. In this regime, the natural research direction is not to harden a small number of centralized defensive systems with fixed roles. It is to study what we will call resilient agent legions: highredundancy populations of defensive agents distributed across network boundaries, observability layers, coordination channels, and recovery functions. Such legions communicate peer-to-peer rather than only upward, exchange defensive heuristics rather than only logs, and accumulate bounded defensive experience over time. Each agent carries an event-and-defense capsule—a packaged unit that binds a class of security events to the defensive routines that should respond to them, and that the agent itself is allowed to update as it learns. Diversity across the population, organized along a layered or fractal structure rather than a single template, is part of the design: it raises the cost of any single adaptive attack and lets the population specialize where local conditions actually differ. This reframing carries a sharper claim. Classical terminal security assumed a relatively stable endpoint host that could be protected by a fixed defensive product. As applications themselves become ephemeral, agent-generated, and behavior-defined, the terminal stops being the load-bearing unit. What used to be terminal security is being deconstructed into agent security: the question is no longer how to protect a fixed endpoint, but how to organize, equip, and govern a bounded population of defensive agents whose collective behavior produces the protective effect. For mass-consumer settings, this argues for population-level monitoring rather than per-instance prevention, because the protected object will not sit still. For critical infrastructure, this argues for engineered legions whose composition, redundancy, and update behavior are themselves objects of study. The Cybersecurity AI Scientist question on this axis is no longer “can an agent detect X?” but “how should a bounded population of defensive agents be designed, organized, evaluated, and governed under adaptive AI-mediated attack?” Figure 3 sketches the shift.

Figure 3: From terminal-anchored defense to resilient agent legions. Each legion carries an eventand-defense capsule and updates it through bounded peer interaction; the protected object is no longer a fixed terminal but a governed population of defensive agents.

6.4

Energy axis: Strategy, ethics, and thought experiments

The Energy axis governs organizational-level outcomes across long horizons. A Cybersecurity AI Scientist is not only a faster experimental engine; the same shift that makes AI-native research possible also raises a class of questions that cannot be settled inside any individual experiment: how research agendas should persist as model platforms turn over, how dual-use containment should be

10

designed when the offensive and defensive uses of a capability are difficult to separate at the level of code, and how a partly autonomous research system should be held accountable for what it chooses to study. These are strategic and ethical questions in a strict sense, not engineering details. The framework treats such questions as native to the research method rather than as a reviewing layer placed around it. In practice this means three things. First, agenda continuity is itself an object of design: the research system should be able to carry questions, constraints, and acceptance criteria across substrate changes, so that conclusions are not silently reset every time the platform shifts. Second, ethics and dual-use review should be expressible as constraints the system can reason with at planning time, not only as post-hoc audits of finished work. Third, a portion of the research surface should be reserved for thought experiments and counterfactual studies that are difficult or unsafe to run as live experiments—questions about long-horizon misuse, about the boundary between research and operations, and about institutional placement of partly autonomous research. These complement, rather than replace, empirical work, and they belong inside the framework precisely because they shape what the empirical work is allowed to ask. This is also where the Cybersecurity AI Scientist diverges most clearly from a faster lab notebook: a lab notebook is judged by what it records; a research system of the kind we describe is judged at least as much by what it declines to ask, and by how it justifies that choice.

6.5

From four agendas to one research object

The four axes are not parallel case studies. Risk gives the system its sharpest empirical target—defect discovery against an adaptive adversary at frontier capability. Trust forces it to remain calibrated against the human operators and downstream decisions it touches. Incident anchors it in operational reality, where the protected object itself is shifting from terminals to bounded populations of agents. Energy holds it accountable for long-horizon organizational and ethical consequences that no single experiment can settle. A system that pursues any one of these axes well but ignores the others is a strong cyber agent, a strong assistant, a strong defense platform, or a strong governance tool—but it is not yet a Cybersecurity AI Scientist. The defining move is to hold all four together inside one research-system organization, with the four-zeros frame as the spine and the multi-agent architecture of Section 5 as one tractable realization.

7

Open Challenges and Boundaries

Several challenges remain open. Defense targets are heterogeneous: what works for one class of organizational system, critical infrastructure asset, or digital platform may fail for another, and broad-transfer claims will require much stronger empirical programs than a framework paper can supply. The dual-use problem is intrinsic rather than peripheral: many capabilities required for high-quality security research are also offensively repurposable, which is why containment, release control, auditability, and permission design are central rather than ornamental. The experimental substrate is unusually demanding: cybersecurity research often needs digital twins, replayable environments, sandboxing, and high-integrity evidence capture, and poorly designed environments can produce misleading conclusions even when nominal benchmark outcomes look strong. A further open question is institutional placement. A Cybersecurity AI Scientist may ultimately need to serve not only technical research execution but also strategic planning, governance exploration, educational cultivation, and thought-experiment-driven safety design. Whether these functions should be integrated into one system or separated into interoperable subsystems is not resolved here. We treat this as a productive open question rather than a deficiency: a framework paper that closed it prematurely would foreclose realizations we cannot yet evaluate. 11

8

Conclusion

AI-native conditions are changing not only cyber operations but cybersecurity research itself. Existing AI scientist systems make end-to-end research automation plausible. Cybersecurity agents, frontier cyber-capable models, vertical foundation models, and operational AI-security systems show that security tasks now impose distinctive pressure on capability, control, benchmarking, and governance at the same time. At their intersection, simple domain transfer is no longer a sufficient framing. Cybersecurity should be treated as a distinct scientific object for AI-driven research, one shaped by adaptive adversaries, non-stationary infrastructures, heterogeneous defense targets, and evidence-intensive validation. We define the Cybersecurity AI Scientist on that basis, anchor its objectives in a four-zeros frame spanning Risk, Trust, Incident, and Energy, and separate the object from any single organizational realization. One practical realization is a modular multi-agent research system built around bounded capability packages, explicit tools, controlled environments, and continuous calibration. Across the four axes, representative agendas—from hidden-defect discovery at frontier capability, through calibrated assistance and multi-model routing, to resilient agent legions and the deconstruction of terminal security into agent security, and on to strategic and ethical thought experiments—show why no single one of them is sufficient. The contribution of this paper is to fix the category, the design pressures, and an initial architecture on which later systems, benchmarks, and empirical programs can be built.

References [1] Anthropic. Project glasswing: Securing critical software for the ai era. https://www.anthropi c.com/glasswing, 2026. Anthropic announcement, April 7, 2026; introduces Claude Mythos Preview as the frontier model powering Project Glasswing. [2] Anthropic. Expanding project glasswing. https://www.anthropic.com/news/expanding-p roject-glasswing, 2026. Anthropic news, 2026; expansion of Claude Mythos Preview access to critical-infrastructure partners. [3] Eser Aygün, Anastasiya Belyaeva, Gheorghe Comanici, Marc Coram, Hao Cui, Jake Garrison, Renee Johnston Anton Kast, Cory Y. McLean, Peter Norgaard, Zahra Shamsi, David Smalling, James Thompson, Subhashini Venugopalan, Brian P. Williams, Chujun He, Sarah Martinson, Martyna Plomecka, Lai Wei, Yuchen Zhou, Qian-Ze Zhu, Matthew Abraham, Erica Brand, Anna Bulanova, Jeffrey A. Cardille, Chris Co, Scott Ellsworth, Grace Joseph, Malcolm Kane, Ryan Krueger, Johan Kartiwa, Dan Liebling, Jan-Matthis Lueckmann, Paul Raccuglia, Wang Xuefei, Katherine Chou, James Manyika, Yossi Matias, John C. Platt, Lizzie Dorfman, Shibl Mourad, and Michael P. Brenner. An ai system to help scientists write expert-level empirical software, 2025. URL https://arxiv.org/abs/2509.06503. [4] Jarrod Barnes. Opensec: Measuring incident response agent calibration under adversarial evidence, 2026. URL https://arxiv.org/abs/2601.21083. [5] Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, David Molnar, Spencer Whitman, and Joshua Saxe. Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models, 2024. URL https://arxiv.org/abs/2404.13161. [6] Qiguang Chen, Mingda Yang, Libo Qin, Jinhao Liu, Zheng Yan, Jiannan Guan, Dengyun Peng, Yiyan Ji, Hanjing Li, Mengkang Hu, Yimeng Zhang, Yihao Liang, Yuhang Zhou, Jiaqi Wang, 12

Zhi Chen, and Wanxiang Che. Ai4research: A survey of artificial intelligence for scientific research, 2025. URL https://arxiv.org/abs/2507.01903. [7] Alankrit Chona, Igor Kozlov, and Ambuj Kumar. Cyber defense benchmark: Agentic threat hunting evaluation for llms in secops, 2026. URL https://arxiv.org/abs/2604.19533. [8] Isaac David et al. Multi-agent penetration testing ai for the web, 2025. URL https://arxiv. org/abs/2508.20816. [9] Defense Advanced Research Projects Agency. Darpa ai cyber challenge aims to secure nation’s most critical software. https://www.darpa.mil/news/2023/ai-cyber-challenge-software, 2023. DARPA News, August 9, 2023. [10] Defense Advanced Research Projects Agency. Ai cyber challenge marks pivotal inflection point for cyber defense. https://www.darpa.mil/news/2025/aixcc-results, 2025. DARPA News, August 8, 2025. [11] Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. Pentestgpt: An llm-empowered automatic penetration testing tool, 2023. URL https://arxiv.org/abs/2308.06782. [12] Gelei Deng, Yi Liu, Yuekang Li, Ruozhao Yang, Xiaofei Xie, Jie Zhang, Han Qiu, and Tianwei Zhang. What makes a good llm agent for real-world penetration testing?, 2026. URL https://arxiv.org/abs/2602.17622. [13] Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. Llm agents can autonomously exploit one-day vulnerabilities, 2024. URL https://arxiv.org/abs/2404.08144. [14] Linus Folkerts, Will Payne, Simon Inman, Philippos Giavridis, Joe Skinner, Sam Deverett, James Aung, Ekin Zorer, Michael Schmatz, Mahmoud Ghanem, John Wilkinson, Alan Steer, Vy Hong, and Jessica Wang. Measuring ai agents’ progress on multi-step cyber attack scenarios, 2026. URL https://arxiv.org/abs/2603.11214. [15] Scott Freitas, Jovan Kalajdjieski, Amir Gharib, and Rob McCann. Ai-driven guided response for security operation centers with microsoft copilot for security, 2024. URL https://arxiv. org/abs/2407.09017. [16] Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J. Szostkiewicz, Dmytro Shved, Gavin J. Gyimesi, Jon M. Laurent, Samantha M. Wright, Muhammed T. Razzak, Andrew D. White, Silvia C. Finnemann, Michaela M. Hinks, and Samuel G. Rodriques. A multiagent system for automating scientific discovery. Nature, 2026. doi: 10.1038/s41586-026-10652-y. URL http://dx.doi.org/10.1038/s41586-026-10652-y. [17] Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad Tomašev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R. D. Costa, José R. Penadés, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan. Accelerating 13

scientific discovery with co-scientist. Nature, 2026. doi: 10.1038/s41586-026-10644-y. URL http://dx.doi.org/10.1038/s41586-026-10644-y. [18] Andreas Happe and Jürgen Cito. Benchmarking practices in llm-driven offensive security: Testbeds, metrics, and experiment design, 2025. URL https://arxiv.org/abs/2504.10112. [19] Pengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli, Daniel Fabian, Burak Gokturk, Jiliang Tang, Chen-Yu Lee, Tomas Pfister, and Long T. Le. Co-redteam: Orchestrated security discovery and exploitation with llm agents, 2026. URL https://arxiv.org/abs/2602.02164. [20] Yifeng He, Ethan Wang, Yuyang Rong, Zifei Cheng, and Hao Chen. Security of ai agents, 2024. URL https://arxiv.org/abs/2406.08689. [21] Sahaya Jestus Lazer, Kshitiz Aryal, Maanak Gupta, and Elisa Bertino. A survey of agentic ai and cybersecurity: Challenges, opportunities and use-case prototypes, 2026. URL https: //arxiv.org/abs/2601.05293. [22] Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, Anna Wu, Arnold Tianyi Yang, Neil Perry, Andy Zou, Matt Fredrikson, J. Zico Kolter, Percy Liang, Dan Boneh, and Daniel E. Ho. Comparing ai agents to cybersecurity professionals in real-world penetration testing, 2025. URL https://arxiv.org/abs/2512.09882. [23] Hanzhi Liu, Chaofan Shou, Xiaonan Liu, Hongbo Wen, Yanju Chen, Ryan Jingyang Fang, and Yu Feng. Synthesizing multi-agent harnesses for vulnerability discovery, 2026. URL https://arxiv.org/abs/2604.20801. [24] Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery, 2024. URL https: //arxiv.org/abs/2408.06292. [25] Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. Llm4sr: A survey on large language models for scientific research, 2025. URL https://arxiv.org/abs/2501.04306. [26] Lajos Muzsai, David Imolai, and András Lukács. Hacksynth: Llm agent and evaluation framework for autonomous penetration testing, 2024. URL https://arxiv.org/abs/2412.0 1778. [27] Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White. Language agents achieve superhuman synthesis of scientific knowledge, 2024. URL https://arxiv.org/ abs/2409.13740. [28] Guiyao Tie, Pan Zhou, and Lichao Sun. A survey of ai scientists, 2025. URL https: //arxiv.org/abs/2510.23045. [29] Kent Walker. A summer of security: empowering cyber defenders with ai. https://blog.goo gle/innovation-and-ai/technology/safety-security/cybersecurity-updates-summe r-2025/, 2025. Google Blog, July 15, 2025. [30] Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, Vlad Ionescu, Yue Li, and Joshua Saxe. Cyberseceval 3: Advancing the evaluation of cybersecurity risks and capabilities in large language models, 2024. URL https://arxiv.org/abs/2408.01605. 14

[31] Dawei Wang, Geng Zhou, Xianglong Li, Yu Bai, Li Chen, Ting Qin, Jian Sun, and Dan Li. The digital cybersecurity expert: How far have we come?, 2025. URL https://arxiv.org/ab s/2504.11783. [32] Zhun Wang, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, Eric Wallace, Elie Bursztein, Luca Invernizzi, Kurt Thomas, Yan Shoshitaishvili, Wenbo Guo, Jingxuan He, Thorsten Holz, and Dawn Song. Exploitgym: Can ai agents turn security vulnerabilities into real attacks?, 2026. URL https://arxiv.org/abs/2605.11086. [33] Zhun Wang et al. Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2025. URL https://arxiv.org/abs/2506.02548. [34] Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search, 2025. URL https://arxiv.org/abs/2504.08066. [35] Zhuoran Yang et al. Llama-3.1-foundationai-securityllm-reasoning-8b technical report, 2026. URL https://arxiv.org/abs/2601.21051. [36] Yuxuan Zhu, Antony Kellermann, Akul Gupta, Philip Li, Richard Fang, Rohan Bindu, and Daniel Kang. Teams of llm agents can exploit zero-day vulnerabilities, 2024. URL https: //arxiv.org/abs/2406.01637. [37] Yuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, Jet Geronimo, Avi Dhir, Sudhit Rao, Kaicheng Yu, Twm Stone, and Daniel Kang. Cve-bench: A benchmark for ai agents’ ability to exploit real-world web application vulnerabilities, 2025. URL https://arxiv.org/abs/2503 .17332.

15

Record · ID 321750 · SHA-256 afe032b6627dc048
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.