Feedback-Driven Execution for LLM-Based Binary Analysis XiangRui Zhang
Qiang Li
Haining Wang
BeiJing JiaoTong· University China
BeiJing JiaoTong· University China
Virginia Tech USA
arXiv:2604.15136v1 [cs.CR] 16 Apr 2026
Abstract Binary analysis increasingly relies on large language models (LLMs) to perform semantic reasoning over complex program behaviors. However, existing approaches largely adopt a one-pass execution paradigm, where reasoning operates over a fixed program representation constructed by static analysis tools. This formulation limits the ability to adapt exploration based on intermediate results and makes it difficult to sustain long-horizon, multi-path analysis under constrained context. We present FORGE, a system that rethinks LLM-based analysis as a feedback-driven execution process. FORGE interleaves reasoning and tool interaction through a reasoning–action–observation loop, enabling incremental exploration and evidence construction. To address the instability of long-horizon reasoning, we introduce a Dynamic Forest of Agents (FoA), a decomposed execution model that dynamically coordinates parallel exploration while bounding per-agent context. We evaluate FORGE on 3,457 real-world firmware binaries. FORGE identifies 1,274 vulnerabilities across 591 unique binaries, achieving 72.3% precision while covering a broader range of vulnerability types than prior approaches. These results demonstrate that structuring LLMbased analysis as a decomposed, feedback-driven execution system enables both scalable reasoning and high-quality outcomes in long-horizon tasks. Keywords: Binary analysis, Large language models, LLM agents, Execution Model, Vulnerability discovery ACM Reference Format: XiangRui Zhang, Qiang Li, and Haining Wang. 2026. FeedbackDriven Execution for LLM-Based Binary Analysis. In Proceedings of XXX . ACM, New York, NY, USA, 17 pages. https://doi.org/XXXX XXX.XXXXXXX Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. XXX, XXX © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-X-XXX-XXXX-X/XXX/XX https://doi.org/XXXXXXX.XXXXXXX
OnePass Binary
Tool
Representation
Analysis
Tool
FeedbackDriven Agent
local observation
Binary
Figure 1. Execution model: one-pass vs. feedback-driven.
1
Introduction
Binary vulnerability analysis is a representative instance of long-horizon program analysis tasks, especially for closedsource firmware and commercial binaries deployed in realworld systems [1, 2, 7, 14]. Despite decades of research, existing approaches largely follow a one-pass analysis paradigm: static analysis tools first construct a global program representation, and vulnerability detection operates over this fixed view using rule-based matching or symbolic reasoning [12, 29, 39]. Recent LLM-based methods [5, 15, 21] also adopt this formulation, where reasoning is performed over precomputed slices of disassembly or decompiled code. From a systems perspective, this paradigm corresponds to a monolithic execution model, where analysis is performed over a fixed, precomputed representation without feedback from intermediate results. This limitation is particularly pronounced in binary analysis. Unlike source code, binaries lack highlevel semantic information such as types, variable names, and structured control flow. As a result, analysis must incrementally reconstruct semantics from low-level artifacts (e.g., assembly instructions, indirect calls, and partial dataflows), making it inherently difficult to construct a complete and reliable global representation in a single pass. Notably, this formulation introduces a fundamental limitation: reasoning is decoupled from analysis. Once the representation is constructed, exploration strategies cannot adapt based on intermediate results, nor refine hypotheses dynamically. This limitation is fundamentally an execution issue: the system lacks a mechanism to incorporate runtime feedback into subsequent analysis decisions. In contrast, real-world binary analysis is inherently feedback-driven. Expert analysts continuously interleave reasoning with analysis operations, refining their understanding step by step through repeated tool interactions. Figure 1 contrasts the traditional one-pass
XXX, XXX, XXX
execution model with a feedback-driven alternative. This difference is not merely a matter of analysis strategy, but reflects two fundamentally different execution models: one operates over a fixed representation, while the other evolves through iterative interaction with the target binary. In the feedback-driven model, reasoning, tool invocation, and observation form a closed loop, where intermediate results directly influence subsequent decisions. Such a mechanism is essential in binary analysis, where semantic information is incomplete and must be incrementally reconstructed from low-level evidence. Recent advances in Large Language Models (LLMs) make such feedback-driven execution practically realizable [41, 43– 45]. By interacting with analysis tools in a loop, LLM agents can incrementally explore binaries, interpret program behavior, and construct execution traces over time. However, scaling this execution paradigm introduces new challenges. Binary analysis often requires long sequences of dependent reasoning steps (depth) and simultaneous exploration of multiple candidate paths (breadth). From a systems perspective, this leads to challenges in managing execution depth, branching factor, and intermediate state under limited context capacity. Under the constraints of LLMs, these properties lead to context degradation, error accumulation, and instability in reasoning traces. Thus, the central challenge is to design an execution model that maintains stable and coherent reasoning processes across long-horizon, multi-path analysis. To address these challenges inherent to binary analysis, we propose FORGE, an end-to-end LLM-driven system designed to stabilize iterative binary analysis. At its core, FORGE introduces a Dynamic Forest of Agents (FoA), which recursively decomposes analysis tasks and dynamically instantiates agents for parallel exploration. FoA adapts its structure according to the evolving analysis state, limiting per-agent reasoning horizons while preserving global exploration coverage. This design transforms analysis from a monolithic execution into a decomposed and feedback-driven system. It prevents long reasoning chains from collapsing and mitigates context drift across divergent paths, enabling stable reasoning over complex binaries. Building on this stabilized execution model, FORGE further enables a unified discovery–verification workflow. Intermediate results—such as call-chain fragments, symbolic values, and taint flows—are continuously recorded during analysis, forming structured evidence chains. These chains can be directly reused in a verification stage, allowing vulnerability discovery and validation to be performed within the same framework. We evaluate FORGE on 3,457 real-world firmware binaries and compare it with state-of-the-art approaches, including Mango [13], SaTC [3], LATTE [21], and SWE [16, 44]. FORGE identifies 1,274 vulnerabilities across 591 unique binaries, achieving 72.3% precision and covering a broader range of
XiangRui Zhang, Qiang Li, and Haining Wang
vulnerability types. These results demonstrate that stabilizing iterative analysis leads to both improved detection performance and more consistent vulnerability discovery across complex binaries. They further suggest that controlling execution structure is a first-order factor in scaling LLM-based analysis systems. Our main contributions are summarized as follows: • Execution formulation. We reformulate LLM-based analysis as a feedback-driven execution problem under partial observability, highlighting the limitations of one-pass reasoning paradigms. • Decomposed execution model. We propose FoA, a dynamic execution model that stabilizes long-horizon reasoning through recursive decomposition and parallel exploration under a bounded context. • Execution–validation integration. We show how structured execution naturally produces replayable evidence, as a unified discovery–verification framework. • End-to-end system and evaluation. We implement FORGE 1 and demonstrate its effectiveness and scalability on large-scale real-world experiments. Roadmap. Section 2 reviews the background and technical challenges. Section 3 details task decomposition, agent generation, and coordination mechanisms. Section 4 presents how FORGE unifies vulnerability discovery and validation under the same FoA architecture. Section 5 describes the experimental results. Section 6 discusses the limitations, and Section 7 reviews related work. Finally, Section 8 concludes the paper.
2
Background
2.1
Binary Vulnerability Detection: One-Pass vs. Iterative Paradigms
Binary vulnerability detection has traditionally been formulated as a one-pass analysis problem, which can be viewed as a static computation model. In this paradigm, static analysis tools—such as disassemblers, control-flow graph (CFG) builders, and symbolic execution engines—are executed once to construct a global program representation, over which vulnerability detection is performed. This formulation assumes that a sufficiently complete representation can be constructed upfront, and that reasoning can be applied independently of the analysis process itself. This abstraction underlies a wide range of existing approaches, including pattern-based tools (e.g., Flawfinder [39], CWE-checker [12]) and symbolic execution systems (e.g., Karonte [29]). Recent LLM-based methods [5, 15, 21] also follow this formulation by performing reasoning over precomputed code representations. In contrast, binary analysis can be more accurately modeled as a sequential decision process under partial observability. 1 Code and Artifacts are available at https://github.com/bjtu-SecurityLab/
FORGE
Feedback-Driven Execution for LLM-Based Binary Analysis
Rather than operating on a fixed representation, the analysis process must iteratively select actions (e.g., disassembling functions, resolving dependencies), observe partial program state, and update its reasoning accordingly. Each decision influences subsequent exploration, and the program representation itself is progressively constructed during execution. This distinction fundamentally changes the nature of the problem: analysis is no longer a one-pass computation over a static structure, but an adaptive process that interleaves reasoning and analysis under incomplete information. This shift introduces new system-level requirements, as the execution process must support sequential decision-making, maintain intermediate state across steps, and dynamically adapt exploration based on partial observations.
XXX, XXX, XXX
Source read()
step 1
step 2
step 3
context lost
attention drift
2.2
2.3
intermediate results step 1 Source
LLM-Driven Iterative Analysis
Challenges in Scaling Iterative Analysis
Unlike one-pass approaches, the iterative execution model introduces distinct structural properties. (1) Partial observability. The agent only observes a small portion of the program at each step and must decide what to explore next without a global view. (2) Sequential dependency. Each decision influences subsequent exploration, making early reasoning steps critical to the overall analysis outcome. (3) Trace-based execution. The analysis process produces a sequence of reasoning steps and tool invocations, forming an explicit execution trace that records how conclusions are derived. These properties, while enabling flexible and adaptive analysis, also introduce fundamental challenges when scaling to large binaries. Reasoning depth. Even moderately sized binaries can contain millions of instructions with complex control and data
... ...
step N
error accumulation
Sink strcpy()
Figure 2. Reasoning Depth: Long analysis chains suffer from context degradation, error accumulation, and attention drift.
step 2 step 3
An execution model aligned with the iterative paradigm can be realized through LLM-based agents. An LLM agent forms a reasoning–action–observation loop. At each step, the agent selects an analysis action (e.g., disassembling a function, tracing a call), receives a localized program fragment, and updates its reasoning before deciding the next action. This process may involve hundreds of iterations, gradually constructing an understanding of the binary. Within this iterative paradigm, reasoning and static analysis are tightly interleaved, leading to two key advantages. (1) Semantic-aware reasoning. By incrementally analyzing relevant code regions and refining hypotheses, the agent can interpret program behavior beyond fixed patterns. It captures implicit relationships (e.g., sanitization logic and control dependencies) and dynamically adapts its analysis strategy. (2) Evidence-preserving analysis. Each step in the reasoning–action loop produces concrete artifacts, including disassembly snippets, addresses, call chains, and intermediate states. These artifacts form structured evidence chains that can be replayed to validate detected vulnerabilities.
new information
intermediate results
Source
step 4
new information
step 6 step 7
Sink
Sink
step 8 step 9
step 5
Sink
Context Overload
Figure 3. Reasoning Breadth: Simultaneous analysis of multiple source-sink combinations leads to context overload and selection bias.
dependencies. As the analysis depth increases, earlier information is easily lost, errors accumulate, and the system may lose focus on the original taint source. Single-agent methods (e.g., ReAct [47]) degrade rapidly beyond a handful of steps, while multi-agent frameworks with static workflows (e.g., AutoGen [40]) cannot dynamically adjust to variable-length reasoning chains. Figure 2 illustrates this degradation. Reasoning Breadth. Large binaries often contain multiple sources, sinks, and propagation paths that must be analyzed concurrently. Tracking all relevant paths without omissions or confusion is challenging for a single LLM. Static multiagent frameworks lack dynamic coordination, making simultaneous path analysis prone to overload or selection bias. Figure 3 illustrates this combinatorial challenge. The central challenge is therefore not enabling semantic reasoning or evidence generation—both are naturally supported in iterative analysis—but maintaining stable and coherent execution processes across long-horizon and multipath analysis. This requires a system design that supports dynamic exploration while preserving coherence across both depth and breadth.
3
System Design
This section presents the design of FORGE as an execution model for long-horizon binary analysis. FORGE is built upon the Forest of Agents (FoA) model, which organizes analysis as a dynamically expanding set of interacting agents.
XXX, XXX, XXX
XiangRui Zhang, Qiang Li, and Haining Wang Taint Tracing Task
Tree of Agent
Src f1
root
f2
f1
f3 A1
... h1
... h2
... h3
Sink Sink/Source
Taint Tracing Task Src
A2
... h4
A3
........
....
root
.... A1
....
h1
A4
Tree of Agent
....
A2
A4
Sink
Sink LLM-based Agent
callee-caller
create
assign task to agent
Figure 4. FoA as a dynamic execution model. Each tree corresponds to a source-rooted analysis process. Nodes represent agents, and edges denote task decomposition and propagation. 3.1
FoA Execution Model
FoA models binary analysis as a dynamically constructed computation structure in the form of a forest of agent-executed task trees: 𝐹𝑜𝐴 = {T1, T2, ..., T𝑛 }, where each tree T corresponds to an analysis process rooted at a specific source. Each node in a tree is an agent 𝐴𝑖 (𝑇𝑖 ), where 𝑇𝑖 is an exploration task assigned to the agent. An agent represents the minimal unit of reasoning and execution in FORGE, encapsulating a localized analysis task over a bounded program context. Agents are strictly organized in parent–child relationships, ensuring that the FoA maintains a well-defined hierarchical structure. Collectively, these trees form a dynamically evolving execution FoA. FoA as a Runtime Structure. FoA is not constructed upfront; instead, it is incrementally materialized as the analysis proceeds. At any point during execution, only a subset of the full exploration structure is instantiated, corresponding to the portions of the analysis that have been explored. This implies that FoA is inherently partial: unexplored branches remain implicit, and only those required for resolving the analysis objective are realized at runtime. As a result, FoA can be viewed as a partially materialized computation graph whose shape is dynamically determined during execution. As such, the FoA represents not only the exploration structure but also the evolving execution state of the analysis. As illustrated in Figure 4, FoA organizes agents to collaboratively accomplish binary vulnerability detection. Each tree represents a taint tracing task for a source, while agents perform semantic reasoning, produce intermediate evidence, and make path pruning decisions. This process naturally forms a branching structure that captures alternative execution paths and data dependencies. Different branches may stop expanding at different depths depending on local analysis decisions, resulting in an irregular, demand-driven structure.
Execution Semantics. FoA execution can be understood as the composition of two tightly coupled processes over this dynamically constructed structure. (1) Forward (structure materialization). The forward process incrementally constructs the FoA by expanding agents into sub-tasks. It defines the exploration space of the analysis, determining which program paths and data flows are examined. (2) Backward (evidence propagation). In parallel, the backward process propagates analysis results along the constructed structure. Each agent continuously produces structured evidence fragments during its reasoning process, including code locations, taint states, and semantic interpretations. These fragments are returned to parent agents and recursively aggregated. A complete evidence chain emerges when evidence from descendant agents is successively integrated along a path to the root, yielding a coherent explanation of a vulnerability. Together, these two processes define the execution semantics of FoA: the forward process determines the reachable computation space, while the backward process defines how partial results are composed into semantically valid outcomes. 3.2
Task Decomposition and Agent Generation
FoA constructs its analysis through a forward execution process, which progressively expands from a source task into a hierarchy of subtasks and agents. Rather than predefining a fixed execution graph, the system dynamically builds an exploration structure at runtime. Starting from an initial task 𝑇0 , FoA performs recursive decomposition and delegation: 𝑇0 → {𝑇1,𝑇2, . . . } → · · · → {𝑇𝑖( 𝑗 ) } ↓ ↓ ↓ ↓ 𝐴0 → {𝐴1, 𝐴2, . . . } → · · · → {𝐴𝑖( 𝑗 ) } The upper row captures recursive task decomposition, while the lower row represents the corresponding runtime agent structure. Each task 𝑇𝑖 is not executed directly; instead, it is realized through a delegation operation that instantiates
Feedback-Driven Execution for LLM-Based Binary Analysis
BaseLLM
JSONOutput
Base Class
Derived Class
BaseAgent Derived Class
XXX, XXX, XXX
New
dynamic execution
pre define ... ... Configuration
output schema memory type
task
tool
Figure 5. Dynamic Agent Generation Process an agent 𝐴𝑖 . Formally, delegation is performed via a toolmediated operation: 𝐴 𝑗 = Delegate(𝑇 𝑗 ) where the delegation operator encapsulates agent creation and task passing. Thus, FoA execution alternates between two tightly coupled steps: (1) task decomposition, which produces new tasks, and (2) delegation, which maps tasks to runtime agents. Task Decomposition. An exploration task is the unit of recursive taint tracing. We represent a task as 𝑇 (𝑓 , 𝑒, 𝑠, 𝑜) where 𝑓 is the current function (with its disassembly address); 𝑒 is the taint entry (the register/stack slot/argument in 𝑓 that receives the taint); 𝑠 is the taint source (the origin expression or call that produced the taint); 𝑜 is the analysis objective (a concise, machine-readable description of what to prove about the propagation, e.g., “can value in 𝑒 reach an argument of system without sanitization”). An exploration task is recursively decomposed into a set of subtasks: 𝑇 (𝑓 , 𝑒, 𝑠, 𝑜) → {𝑇 (𝑓1, 𝑒 1, 𝑠 1, 𝑜 1 ), . . . ,𝑇 (𝑓𝑛 , 𝑒𝑛 , 𝑠𝑛 , 𝑜𝑛 )} where each child task corresponds to a refined analysis subproblem derived from the current task context (e.g., following a data dependency or exploring a control-flow branch). The decomposition is determined by the agent’s reasoning process, which is instantiated using an LLM but constrained by the task structure. Subtasks may represent sequential steps, parallel branches, or alternative exploration paths. Delegation as Agent Generation. FoA treats delegation as a first-class runtime primitive that defines how computation is dynamically expanded and distributed. Each subtask𝑇𝑖 is passed as input to a delegation operator, which instantiates delegate
a new agent: 𝑇𝑖 −−−−−−→ 𝐴𝑖 . One possible instantiation of this abstraction is to realize delegation through a tool-mediated interface, which triggers the creation of a new agent responsible for executing the task. As illustrated in Figure 5, AgentTool acts as a bridge between agent reasoning and runtime instantiation: subtasks are first produced by the agent’s reasoning process, and then materialized as new agents via tool invocation. Each agent is equipped with AgentTool, enabling it to create new agents as needed. The generated agent 𝐴𝑖 is
responsible for executing 𝑇𝑖 independently, including invoking tools, interacting with the environment, and producing intermediate results. This abstraction elevates delegation from an implementation detail to a first-class runtime primitive, defining how computation is dynamically distributed across agents in the FoA. For example, a concrete implementation (AgentTool) may support both parallel and sequential delegation strategies: 1 Parallel Generation (1→n). When the LLM identifies multiple callee functions requiring exploration, the current agent creates multiple child agents, each responsible for one distinct subtask (𝑓𝑖 , 𝑒𝑖 , 𝑠𝑖 , 𝑜𝑖 ). These child agents can execute concurrently, enabling true runtime parallelism and efficient handling of reasoning breadth across multiple paths. 2 Sequential Generation (1→1). When the LLM identifies a single critical callee requiring deep tracing, the agent creates one child agent to continue exploration along that specific path. This ensures that long reasoning chains are preserved without overloading any single agent. 3.3 Agent Execution and Evidence Formation Agent Execution. FORGE models agent execution as a constrained execution interface between the reasoning process and the target binary. In this design, the binary serves as the external environment, and tools provide the only observable channel through which agents can query program state and behavior. Each agent step consists of invoking a tool with structured parameters derived from its current task context (e.g., querying disassembly, control-flow relations, or data dependencies), and incorporating the returned results into its reasoning process. Thus, tool invocations define the operational semantics of agent–environment interaction, while remaining orthogonal to the FoA structure itself. Evidence Formation. Rather than exposing raw tool outputs, agents produce structured evidence as execution artifacts. These artifacts constitute a persistent representation of the execution trace, enabling replay and verification. An evidence object summarizes the essential information required to describe a propagation step, including: code location (e.g., function address), relevant snippet or instruction context, taint-related variables or registers, and a concise semantic interpretation of the propagation. For example: {"next_function": "sub_401B20", "evidence": {"addr": 0x401B20, "snippet": "...", "vars": ..., "note": "propagated via R0"}} This abstraction decouples low-level tool outputs from higherlevel reasoning, enabling consistent aggregation and interpretation across independently executed agents. This evidence serves two purposes: (1) it compresses lowlevel tool outputs into semantically meaningful units, and (2) it provides a uniform interface for cross-agent communication and aggregation.
XXX, XXX, XXX
3.4
Runtime Coordination and Aggregation
FoA execution is governed by three tightly coupled runtime mechanisms: (1) forward expansion constraints, (2) hierarchical aggregation, and (3) decentralized coordination. Together, these mechanisms ensure that the dynamically constructed FoA remains bounded, consistent, and semantically coherent. Forward Expansion Constraints. FoA expansion is not centrally scheduled; instead, it is driven by local agent decisions under structural constraints. First, expansion preserves the hierarchical structure defined by FoA: tasks can only be delegated along parent–child relationships. New agents can only be generated through parent–child delegation, i.e., 𝑇𝑖 can only produce subtasks {𝑇 𝑗 } that are delegated to child agents {𝐴 𝑗 }. Second, each agent operates within a strictly bounded scope defined by its task 𝑇 (𝑓 , 𝑒, 𝑠, 𝑜). The parent agent passes only task-relevant context to the child, including the target function, taint entry, and analysis objective. This ensures that reasoning remains localized and avoids uncontrolled context growth. Third, expansion decisions are fully delegated to the LLM. At each step, an agent determines whether to: (i) generate new subtasks for further exploration, or (ii) stop expansion for the current path. Termination is therefore not a separate phase, but a natural outcome of the expansion process. An agent stops generating new tasks when the LLM determines that the current path is either sufficiently resolved (e.g., reaching a vulnerabilityrelevant endpoint) or no longer worth exploring. This design enables adaptive, demand-driven exploration while preventing unbounded growth. Hierarchical Aggregation Model. While forward expansion constructs the FoA, execution results are propagated in the reverse direction through hierarchical aggregation. Each agent produces a structured result upon completion of its assigned task. This result is not merely a tool return; instead, it is a runtime-level output that encapsulates: (i) the local analysis outcome, and (ii) a structured evidence fragment derived from tool interactions. When a child agent finishes execution, its result is returned to the parent agent as part of the runtime control flow (rather than as a direct tool response). The parent agent incorporates this result into its local reasoning state, potentially combining multiple child results. This recursive returnand-integration process forms a hierarchical aggregation structure: leaf agents produce atomic evidence fragments, which are progressively merged along the tree edges. A complete evidence chain is formed when results from leaf agents are successively integrated up to the root, yielding a coherent explanation of the full propagation path. Importantly, aggregation is tightly coupled with execution: a parent agent may suspend its own reasoning until child results are available, making aggregation an inherent part of the runtime rather than a post-processing step. Thus, aggregation defines how distributed execution results are
XiangRui Zhang, Qiang Li, and Haining Wang
composed, forming a core part of the system’s execution semantics rather than a passive collection mechanism. Decentralized Coordination. FoA does not rely on a centralized planner or global scheduler; instead, execution control is fully embedded in the runtime structure itself. Instead, coordination emerges from local interactions between agents. Each agent independently decides how to proceed based on its local context and received evidence. Parent agents control the lifecycle of their children through delegation and aggregation, while child agents operate independently within their assigned scope. This decentralized design enables: (i) adaptive exploration across multiple paths, (ii) parallel execution of independent branches, and (iii) robustness to partial failures or inconsistent intermediate results. Overall, FoA enforces a key execution invariant: reasoning is localized, while execution is globally compositional through hierarchical structure.
4
FoA Instantiation: Vulnerability Discovery and Validation
We demonstrate how FoA can be instantiated to support a multi-phase vulnerability analysis workflow, consisting of discovery and validation. While these phases serve distinct purposes, they are both realized as executions of the same FoA model, operating over shared task abstractions and evidence structures. Figure 6 illustrates a closed-loop execution process: discovery → vulnerability hypotheses (evidence chains) → validation → verified vulnerabilities (evidence chains). These phases correspond to different execution modes over the same FoA structure. The two modes share identical execution semantics—including task decomposition, delegation, and hierarchical aggregation—and differ primarily in their inputs and task specifications. Discovery operates over source-initialized tasks to explore the program space, while validation operates over previously generated evidence chains to perform constrained re-execution. An illustrative 5-turn example of this process is provided in Listings D–D.5 in the appendix. 4.1
Vulnerability Discovery
In the discovery phase, root agents are initialized from a set of high-risk sources obtained via root enumeration of the binary. Each source serves as the root of a task-specific exploration tree. Each agent operates under a discoveryoriented task specification (Listing B.1), which prioritizes: • Evidence-based analysis: All findings must be grounded in concrete evidence extracted via the tools’ outputs. • Taint identification: Agents identify and validate all externally controllable taint entries. • Delegated tracing: New agents perform the actual taint propagation.
Feedback-Driven Execution for LLM-Based Binary Analysis Evidence Evidence Evidence Evidence Evidence Evidence
vuln. verified vuln. ... ...
... ...
Evidence Chain 1 .... Evidence Chain N
vuln. ... ...
... ...
binary
source 1 ..... source K
XXX, XXX, XXX
Evidence verified Evidence vuln.
Figure 6. FoA for Vulnerability Discovery and Validation.
• Focused task execution: Each agent works strictly within its assigned task to ensure comprehensive coverage of exploitable paths. The output of the discovery phase is not a raw list of function names or warnings, but a structured evidence chain that records concrete propagation steps and semantic reasoning results. Each chain records concrete taint propagation steps, call-chain reasoning traces, and semantic explanations of exploitability. An example evidence chain is shown in Listing C.1. Thus, discovery produces both candidate vulnerabilities and the structured evidence artifacts required for subsequent execution in the validation phase. 4.2
Vulnerability Validation
Unlike discovery, which explores from identified sources, validation is formulated as an evidence-constrained execution process. It begins with a structured evidence chain generated during discovery and re-executes it within the FoA framework by verifying each propagation step under the same execution semantics. The validation agent replays and verifies this chain within the binary by confirming the existence of each propagation step and verifying exploitability under the defined threat model. This process follows a validationoriented task specification (Listing B.2), which constrains execution to the given evidence structure while verifying its correctness. Each validation task thus operates as an evidence-driven replay, where execution is constrained by the structure of the input evidence chain. Each validation task produces one of two outcomes— (1) a verified, address-corrected propagation path if all evidence is consistent, or (2) a rejection with explicit reasoning (e.g., the path does not exist, data is sanitized, or the sink is unreachable). By grounding the validation process in previously recorded artifacts, FoA achieves replayable verification without requiring re-discovery or manual intervention. After validation, the verified vulnerability and its associated evidence chains are aggregated into structured reports (List C.2).
5
Evaluation
In this section, we evaluate FORGE’s effectiveness in realworld vulnerability detection tasks. Specifically, we aim to answer the following research questions: RQ1 (Effectiveness of Execution Model): Does FoA improve long-horizon vulnerability analysis compared to existing paradigms, including static analysis tools and prior LLM-based approaches? RQ2 (Mechanism Analysis): What are the key system mechanisms in FoA that contribute to its effectiveness in complex vulnerability analysis tasks? RQ3 (Efficiency and Scalability): Is FoA efficient and scalable for real-world vulnerability analysis in terms of cost, resource usage, and workload complexity? 5.1
Experimental Settings
For the LLM backend, we use DeepSeek-v3 for vulnerability detection. To ensure deterministic outputs, we fix the temperature parameter at 0. Each agent is allowed up to 30 iterative reasoning steps per task instance, after which the process is terminated if no solution is found. Datasets. We evaluate FORGE on 3,457 binaries extracted from real-world firmware images of four major IoT vendors (NETGEAR, D-Link, TP-Link, and Tenda), originally collected from the Karonte dataset [29]. This dataset—also used by Mango [13] and SaTC [3]—covers diverse device families and architectures commonly found in consumer IoT products. Each binary satisfies both of the following conditions: (1) it contains at least one known dangerous function (sink); and (2) it contains at least one taint source from five predefined entry types: network, environment variable (including NVRAM and ENV), file, command-line argument (argv), and other. This ensures that all analyzed binaries are both security-relevant and comparable across different approaches. Baselines. We compare FORGE with four baselines. (1) Operation Mango [13] is a SOTA tool for finding binary vulnerabilities via taint tracing and symbolic execution. (2) SaTC [3] is a static analysis tool that leverages shared input keywords across firmware to detect bugs/vulnerabilities in embedded systems. (3) LATTE [21] is an LLM-powered static binary taint analysis system that integrates LLMs for taint analysis and vulnerability discovery in binaries and source code. (4) SWE Agent [44] is a framework that interleaves reasoning and tool actions in multiple steps, enabling dynamic decision-making and context tracking for software repositories. For fair comparison, we equip SWE Agent with the same set of tools and prompts as our system. Metrics. We evaluate the performance of FORGE using two key metrics that directly correspond to its two phases: discovery and validation. (1) Underlying Vulnerability (Discovery Phase). This refers to any finding automatically identified by the system during the discovery phase—i.e., a potential
XXX, XXX, XXX
XiangRui Zhang, Qiang Li, and Haining Wang
Table 1. Vulnerability Discovery Comparison between Operation Mango and FORGE on the Karonte Dataset. For Operation Mango, Underlying Vulns = CI + BoF; for FORGE, Underlying Vulns = CI + BoF + Other Vulns.
Underlying Vulns Netgear Tenda D-link TP-link Total
1447 391 85 94 2017
Operation Mango Affected CI Vulns Binaries 173 21 36 14 244
339 154 8 9 510
FORGE BoF Vulns
Underlying Vulns
Affected Binaries
CI Vulns
BoF Vulns
Other Vulns
1108 237 77 85 1507
880 185 159 50 1274
429 61 74 27 591
457 100 85 25 667
232 42 30 25 329
191 43 44 0 278
Table 2. Comparison of vulnerability detection between SaTC and FORGE.
CI BoF Affected Unique Binaries
SaTC
FORGE
144 44 131
667 329 591
unsanitized data flow from a plausible untrusted source to a critical sink. These findings represent candidate vulnerabilities inferred through LLM-based semantic reasoning and evidence collection. (2) Verified Vulnerability (Validation Phase). This refers to underlying vulnerabilities that were further confirmed as exploitable during the validation phase through concrete evidence replay. When full manual labeling is unavailable, we estimate the total number of verified vulnerabilities by applying the measured precision from our random manual sample to the total underlying count. 5.2
Table 3. Vulnerability detection prediction in FORGE by type Vulnerability Type BoF CI Others
Underlying Vulns
Verified Vulns
Precision (%)
150 150 50
89 121 34
59.3 80.6 68.0
Table 4. Comparison of estimated verified vulnerabilities for Operation Mango and FORGE.
Total
Mango Est. Valid (%)
Total
FORGE Est. Valid (%)
BoF CI Others
1507 510 N/A
618 (41.0%) 245 (48.1%) N/A
329 667 278
195 (59.3%) 538 (80.6%) 189 (68.0%)
Total
2017
∼863 (42.7%)
1274
≈922 (72.3%)
Est. Valid (%) indicates the estimated precision, i.e., the proportion of verified vulnerabilities among all underlying vulnerabilities.
Effectiveness (RQ1)
We evaluate whether the FoA execution model improves both the effectiveness and practicality of vulnerability analysis. We compare FoA against traditional static analysis tools and prior LLM-based approaches in terms of vulnerability detection effectiveness. Comparison with Static Analysis Paradigms. Table 1 summarizes the number of detected underlying vulnerabilities (“Vulns”), affected binaries, buffer overflows (“BoF”), and command injections (“CI”) across vendors for both FORGE and Operation Mango. Table 2 shows complementary results from SaTC, which focuses on different aspects of vulnerability analysis. Although FORGE reports fewer underlying vulnerabilities than Mango, this difference reflects a fundamental distinction in execution behavior. Traditional tools such as Mango rely on large-scale path enumeration, which tends to produce many syntactic candidates, a substantial fraction of which cannot be validated. In practice, the raw count of underlying alerts is less meaningful: a high number of alerts often corresponds to a large fraction of false
positives, which in turn imposes a substantial verification burden on human analysts. This difference becomes more evident when considering validation outcomes. The validation phase evaluates the exploitability of candidate vulnerabilities identified in discovery. We randomly sampled 150 buffer overflow (BoF), 150 command injection (CI), and 50 other cases, and verified each vulnerability through independent review by two annotators. A vulnerability was labeled as verified only if both confirmed an evidence-backed dataflow from an attacker-controlled input to a sensitive sink. Table 3 shows that FORGE achieves an overall precision of 72.3%, significantly higher than the estimated 42.7% for Mango. As shown in Table 4, this results in approximately 922 verified vulnerabilities for FORGE, compared to ∼863 for Mango. Despite generating fewer raw alerts, FORGE achieves comparable or higher numbers of verified vulnerabilities, demonstrating that FoA improves the quality of exploration rather than merely increasing its volume. This behavior is consistent
Feedback-Driven Execution for LLM-Based Binary Analysis
Table 5. Comparison of unique vulnerability types discovered by FORGE and Mango. Vulns Type
CWE-ID (Num.)
Operation Mango
2
CWE-78 (510), CWE-120 (1507)
FORGE
6+
CWE-22 (25),CWE-73 (61), CWE-78 (667),CWE-120 (329), CWE-134 (109), CWE-200 (73)
with the FoA design, where exploration and validation are tightly coupled. Comparison with LLM-based Paradigms. We further compare FORGE with representative LLM-based approaches. LATTE detects 219 vulnerabilities (94 unique), while SWE detects only 82 underlying vulnerabilities (11 verified). LATTE follows a static pipeline where LLMs are used only for postprocessing, limiting flexibility and coverage. SWE adopts a single-agent, sequential reasoning model, which struggles to handle long and complex dataflow chains. In contrast, FoA enables dynamic multi-agent generation and parallel exploration across multiple reasoning paths. This allows FORGE to scale to complex binaries and significantly improve both coverage and verified vulnerability yield. The key difference lies not in the use of LLMs themselves, but in how reasoning is structured and executed. These results suggest that simply incorporating LLMs is insufficient; instead, the execution model—particularly parallelization, decomposition, and evidence-driven validation—plays a critical role in effective analysis. Generalization and Robustness. FoA also improves generalization and robustness for binary vulnerability detection. Table 5 shows that FORGE supports a broader range of vulnerability types (6+ categories) compared to Mango’s two primary types. Extending Mango to handle additional vulnerability categories requires manually designing new detection rules, integrating them into its static analysis pipeline, and performing additional engineering efforts. This broader coverage comes from agents inspecting dynamic intermediate artifacts (assembly, pseudocode, xrefs, strings) and using semantic understanding to judge diverse patterns. This is consistent with FoA’s ability to adapt to diverse patterns without relying on manually crafted rules. Such flexibility is enabled by its compositional execution structure, where agents dynamically interpret intermediate artifacts rather than relying on fixed rule templates. FORGE also identifies substantially more unique vulnerable binaries (591 vs. 244) and achieves a more balanced distribution across vendors. Furthermore, Table 6 shows that FORGE successfully analyzes binaries where Mango fails due to path explosion, memory exhaustion, or timeouts, discovering additional vulnerabilities in these challenging cases. For
XXX, XXX, XXX
Table 6. Vulnerability discovery by FORGE in binaries where Operation Mango failed. Vendor
Angr Errors
OOM Kill
Time out
Affected Binaries
Underlying Vulns
Netgear Tenda D-link TP-link Total
151 13 67 68 299
43 5 17 37 102
0 3 0 1 4
57 10 14 4 85
124 25 23 5 179
Context Collapse
FoA Execution
Stable LongHorizon Reasoning
Search Explosion
Semantic Pruning
Focused Exploration
Unverified Alerts
Validation Pipeline
High-Precision Output
Figure 7. FoA addresses three fundamental failure modes in long-horizon vulnerability analysis.
example, FORGE detects 429 affected binaries for Netgear (2.4× that of Mango) and 61 for Tenda (nearly 3×). In contrast, Mango’s vulnerability findings are heavily skewed toward a small subset of binaries—mainly those from Netgear and Tenda—while its detection results for other vendors such as D-Link and TP-Link remain sparse. This demonstrates that FoA improves not only detection capability but also robustness in real-world scenarios. Summary. Overall, the results consistently show that FoA improves long-horizon vulnerability analysis across multiple dimensions. Rather than increasing the number of raw alerts, FoA restructures the execution process to produce higherquality, verifiable vulnerabilities, while achieving broader coverage and better robustness than both traditional static analysis tools and prior LLM-based approaches. 5.3
Mechanism Analysis (RQ2)
To understand why FoA improves vulnerability analysis, we analyze the system through the lens of three fundamental failure modes in long-horizon reasoning: (1) context collapse in linear execution, (2) search explosion due to uncontrolled branching, and (3) unverified alerts caused by the lack of systematic validation. As illustrated in Figure 7, FORGE addresses these challenges through three corresponding mechanisms: FoA execution model, LLM-guided semantic pruning, and a discovery–validation pipeline. We evaluate these mechanisms using controlled ablations on a sampled subset of 500 binaries from the Karonte dataset, under identical infrastructure and LLM settings. Each experiment is repeated 5 times to ensure stable and comparable results.
XXX, XXX, XXX
XiangRui Zhang, Qiang Li, and Haining Wang
Table 7. Ablation: Single Agent, Sequential-only Generation and FORGE.
Single Agent Sequential-only Generation FORGE
Vul.
Verified Vulns
8.4 45.8 172.0
1.3 22.3 136.0
Execution Structure: Mitigating Context Collapse. We first examine how FoA addresses context collapse, a common failure mode in long-horizon reasoning where the LLM fails to maintain extended dependency chains. We compare FORGE with two degraded variants: a single-agent system and a sequential-only execution model. As shown in Table 7, the single-agent variant detects only 8.4 underlying vulnerabilities and 1.3 verified vulnerabilities, while sequential-only execution improves to 45.8 and 22.3, respectively, but remains far below FORGE (172.0 / 136.0). These results indicate that linear reasoning structures cannot sustain deep and branching dataflow analysis. In the single-agent setting, the LLM must maintain the entire reasoning state within a single context, leading to rapid degradation as analysis depth increases. Sequential execution partially alleviates this issue but still enforces strict serialization, causing early context loss and limiting coverage. In contrast, FoA decomposes reasoning into dynamically generated agents and enables parallel exploration across multiple paths. This design localizes reasoning contexts and preserves intermediate evidence, effectively mitigating context collapse and enabling scalable long-horizon analysis. Search Control: Mitigating Search Explosion. We next analyze how FoA addresses search explosion, where uncontrolled branching leads to excessive and unproductive exploration. Table 8 compares full FORGE with a NoPrune variant that enforces exhaustive exploration. While NoPrune increases the number of agents (40.4 vs. 34.3) and reasoning steps (503.1 vs. 464.1), it reduces verified vulnerabilities from 136.0 to 98.7. This demonstrates that naive expansion of the search space is counterproductive: without filtering, the system explores many semantically irrelevant paths, diluting computational resources and increasing noise. FoA addresses this issue through LLM-guided semantic pruning, which selectively expands only meaningful branches. Although LLMs exhibit some implicit filtering ability, such filtering is coarse-grained and inconsistent. The explicit pruning mechanism provides fine-grained control over exploration, ensuring efficient and focused reasoning. The degradation in verified vulnerabilities under NoPrune directly attributes performance gains to the pruning mechanism, confirming its causal role in controlling search explosion. Validation Pipeline: Mitigating Unverified Alerts. We finally analyze how FoA addresses unverified alerts, a common issue where systems produce large numbers of
Table 8. Pruning ablation between full FORGE and NoPrune.
FORGE NoPrune Failure rate
Vul.
Verified
Avg agents
Avg steps
172.2 167.4
136.0 98.7
34.3 40.4
464.1 503.1
Full: 0.00%
NoPrune: 0.00%
candidate vulnerabilities without confirming exploitability. FoA adopts a two-stage pipeline consisting of discovery and validation. The discovery stage generates candidate vulnerabilities, while the validation stage verifies exploitability through evidence-backed analysis. The effect of this design is reflected in the gap between underlying and verified vulnerabilities reported in RQ1, which quantifies the number of spurious candidates filtered out during validation. Without this stage, the system would resemble traditional static analysis tools that produce many unverifiable alerts. The validation mechanism thus ensures that reported vulnerabilities are actionable, significantly improving precision. This separation between discovery and validation stages provides an explicit mechanism for eliminating false positives, linking the observed precision improvements in RQ1 to the validation pipeline design. Summary. Overall, FoA’s effectiveness can be understood as addressing three failure modes in long-horizon vulnerability analysis: context collapse, search explosion, and unverified alerts. Through controlled ablations, we show that each mechanism contributes causally to performance improvements by isolating and mitigating these failure modes. In short, FoA transforms vulnerability analysis from linear reasoning into a structured and robust process. 5.4
Efficiency and Scalability (RQ3)
We evaluate whether FORGE is efficient and scalable for real-world vulnerability analysis from three perspectives: (1) end-to-end analysis cost per binary, (2) cost normalized by useful outputs (verified vulnerabilities), and (3) scalability under varying workload complexity. End-to-End Efficiency. Table 9 summarizes the average per-binary cost of FORGE. On average, FORGE completes analysis within 43.8 minutes per binary, consuming 1.61M tokens (1.29M for discovery and 0.32M for validation). Each run involves 33.8 agents and 464 reasoning steps. These results indicate that FORGE maintains moderate per-instance cost despite extensive exploration. Notably, the large number of agents does not lead to prohibitive overhead, suggesting that the lightweight agent abstraction enables parallel exploration without significantly increasing marginal cost. Cost per Verified Vulnerability. Since different systems produce substantially different numbers of findings, raw execution cost is not directly comparable. We therefore normalize cost by the number of verified vulnerabilities. As
Feedback-Driven Execution for LLM-Based Binary Analysis
XXX, XXX, XXX
Table 9. The overview of time cost and token usage in FORGE. per binary Value Average time Average tokens (discovery / validation) Average agents Average reasoning steps
43.8 min 1.29 / 0.32 (M) 33.79 464.25
Table 10. Per-verified vulnerability cost and time. Variant
Time/vul.
Tokens/vul.
Single Agent FORGE
357.1 min 140.2 min
12.5 M 4.71 M
6
1.0
CDF
0.8 0.6 0.4 0.2
Agents Steps
0.0 0
500
1000
1500
2000
Value (Steps / Agents)
2500
Scalability in Large-Scale Deployment. We further evaluate FORGE on 3,457 real-world binaries across multiple vendors. FORGE successfully analyzes all binaries and identifies a substantially larger number of vulnerabilities compared to prior systems, while maintaining manageable per-instance cost. This demonstrates that FORGE can sustain both high throughput and high yield in large-scale settings, a key requirement for practical deployment. Summary. FORGE achieves efficiency–effectiveness tradeoffs, reducing cost per verified vulnerability while scaling to complex and large-scale analysis workloads.
3000
Figure 8. CDF of reasoning steps and the number of agents created per in the discovery phase.
shown in Table 10, FORGE requires 140.2 minutes and 4.71M tokens per verified vulnerability, compared to 357.1 minutes and 12.5M tokens for the single-agent baseline. This corresponds to an approximate 2.5× improvement in both time and token efficiency. This gain stems from FORGE’s two-stage execution model: the discovery phase aggressively explores candidate paths, while the validation phase filters and confirms only high-confidence findings. As a result, computational resources are concentrated on actionable outputs rather than redundant reasoning trajectories. Scalability with Analysis Complexity. To understand how FORGE scales with workload complexity, we analyze the distribution of reasoning steps and agent usage in Figure 8. Both exhibit a long-tail distribution: most binaries require moderate effort, while a small fraction demands significantly deeper exploration. Importantly, resource consumption (tokens and time cost) correlates primarily with reasoning depth and path complexity, rather than the number of instantiated agents. This behavior suggests that FORGE scales with intrinsic task complexity rather than suffering from systemic inefficiencies, making it suitable for heterogeneous real-world workloads.
Discussion and Limitations
Stability and Variance in LLM-driven Execution. A potential concern is the stability of FORGE across runs and inputs. While we reduce randomness by setting temperature to zero and enforcing evidence-based reasoning constraints, the system may still exhibit variability due to the inherent non-determinism of LLMs. Importantly, this variability manifests primarily at the reasoning path level (e.g., different exploration trajectories), rather than completely invalid outputs. Our empirical results show that, despite such variations, the overall effectiveness remains stable across repeated runs and diverse binary sets. This suggests that FoA mitigates—but does not fully eliminate—LLM-induced variance by structuring reasoning into constrained and verifiable steps. Residual Risk of Hallucination. Although FoA enforces evidence-backed reasoning and validation, hallucination cannot be fully eliminated. In particular, errors may still arise when intermediate representations (e.g., decompiled code or dataflow traces) are ambiguous or incomplete. To mitigate this, FORGE requires structured outputs, enforces schema validation, and applies independent verification for sampled cases (Table 3). However, these safeguards rely on the availability and correctness of supporting evidence. As a result, hallucination remains a residual risk, especially for complex binaries with unclear semantics. Dependence on Underlying Analysis Tools. FORGE relies on external analysis tools (currently Radare2Tool) to provide intermediate artifacts such as disassembly, control flow, and cross-references. Limitations of these tools—such as inaccurate function boundary recovery, unresolved indirect calls, or aliasing issues—directly affect the quality of downstream reasoning. While FoA partially mitigates these issues by reasoning over multiple intermediate representations and refining hypotheses iteratively, it cannot recover information that is fundamentally missing or incorrect. Therefore, the effectiveness of FORGE is bounded by the fidelity of the underlying analysis tools. Future work may explore integrating complementary techniques such as symbolic execution or fuzzing to improve analysis completeness.
XXX, XXX, XXX
Incomplete Coverage. Despite improved exploration efficiency, FORGE does not guarantee full path coverage. Binary analysis inherently suffers from path explosion, complex control flows, and indirect jumps. FoA alleviates this issue through structured exploration and pruning (RQ2), but may still miss vulnerabilities located in rarely explored or highly obfuscated paths. This limitation reflects a fundamental trade-off between exploration completeness and computational tractability. Improving coverage likely requires hybrid approaches that combine static, dynamic, and learning-based techniques. Scope of Mechanism Validation. Finally, while our ablation study (RQ2) isolates key mechanisms—execution structure, pruning, and validation—it is conducted on a sampled subset of binaries. Although the results are consistent and repeated across runs, they may not capture all edge cases in large-scale real-world deployments. Extending mechanismlevel validation to broader datasets and more diverse environments remains an important direction for future work.
7
Related Work
Binary Analysis. Static binary analysis includes rule-based detection (e.g., Flawfinder [39], CWE-checker [12]) and symbolic execution (e.g., Mango [13], SaTC [3], Karonte [29]). These approaches improve coverage and precision through techniques such as data-flow analysis, constraint solving, and path exploration. Dynamic approaches, including fuzzing [11, 19, 30, 49] and concolic execution [8, 28, 33, 48], complement static analysis by exploring runtime behaviors, but rely heavily on execution environments and input generation strategies. In practice, both paradigms require substantial engineering effort to balance coverage and efficiency, especially in complex firmware settings. Despite these advances, existing binary analysis techniques largely operate under a one-pass or weakly adaptive paradigm, where exploration strategies are fixed or only locally adjusted, limiting the ability to incorporate intermediate results into global decision making and leading to path explosion, limited semantic adaptability, and high engineering cost. Systems for Multi-Path and Tool-Mediated Execution. Several works study how to scale multi-path exploration and tool–computation loops under resource constraints [6, 22, 31, 34, 41, 43, 45, 47]. These systems typically treat complex tasks as iterative processes that interleave reasoning, execution, and feedback, enabling adaptive decision making during execution. For example, S2E [6] demonstrates selective symbolic execution with execution multiplexing, while more recent LLM-based systems emphasize closedloop interaction through tool use and environment feedback. At larger scales, LLM systems are integrated into infrastructure and pipelines, such as scheduling in heterogeneous environments [22] and debugging workflows [34], highlighting the role of resource-aware execution. However, these
XiangRui Zhang, Qiang Li, and Haining Wang
systems primarily focus on scalability, interaction patterns, or resource management, and do not provide a structured execution model for coordinating long-horizon reasoning, exploration, and validation in semantically complex domains such as binary analysis. LLM Agents and LLM-assisted Code Analysis. LLMs demonstrate strong reasoning capabilities, enhanced by prompting strategies such as task decomposition, chainof-thought, tree-of-thought, and self-reflection [20, 32, 37, 38, 46]. LLM agents extend these capabilities into toolaugmented systems [24, 36, 42], enabling iterative reasoning with external tools and environments. These approaches have been applied to software engineering [23] and security tasks [4, 10, 26, 27], including bug reproduction [10], vulnerability discovery [18, 25], and test generation [9, 17]. Recent systems such as LLMSAN [35] and LATTE [21] further explore applying LLMs to program and binary analysis. However, existing LLM-based approaches primarily improve reasoning quality at the prompt or pipeline level, while their execution remains largely linear or weakly structured, making them prone to context collapse, uncontrolled exploration, and lack of systematic validation in long-horizon tasks. In contrast, FoA introduces a structured execution model that unifies reasoning, exploration, and validation, systematically addressing context collapse, search explosion, and unverified alerts in long-horizon vulnerability analysis.
8
Conclusion
We proposed FORGE, an LLM-driven system that rethinks binary vulnerability analysis as an iterative and structured execution process rather than a one-pass pipeline. By tightly interleaving reasoning and analysis through a reasoning–action–observation loop, FORGE enables adaptive exploration and incremental construction of semantic evidence. To support scalable long-horizon analysis, we introduce a FoA execution model, which decomposes tasks and coordinates parallel exploration while maintaining bounded reasoning contexts. We evaluate FORGE on 3,457 real-world binaries and show that it identifies 1,274 vulnerabilities across 591 unique binaries with 72.3% precision, covering a broader range of vulnerability types than prior approaches. Overall, our results demonstrate that restructuring vulnerability analysis as a coordinated, multi-agent execution process—rather than improving individual analysis components—provides a principled way to address context collapse, search explosion, and unverified alerts, enabling more effective and scalable binary vulnerability detection.
Feedback-Driven Execution for LLM-Based Binary Analysis
References [1] Marcel Busch, Aravind Machiry, Chad Spensky, Giovanni Vigna, Christopher Kruegel, and Mathias Payer. 2023. Teezz: Fuzzing trusted applications on cots android devices. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 1204–1219. [2] Daming D Chen, Maverick Woo, David Brumley, and Manuel Egele. 2016. Towards automated dynamic analysis for linux-based embedded firmware.. In NDSS. [3] Libo Chen, Yanhao Wang, Quanpu Cai, Yunfan Zhan, Hong Hu, Jiaqi Linghu, Qinsheng Hou, Chao Zhang, Haixin Duan, and Zhi Xue. 2021. Sharing more and checking less: Leveraging common input keywords to detect bugs in embedded systems. In 30th USENIX Security Symposium (USENIX Security 21). 303–319. [4] Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128 (2023). [5] Xiang Chen, Chengfeng Ye, Anshunkang Zhou, and Charles Zhang. 2025. ClearAgent: Agentic Binary Analysis for Effective Vulnerability Detection. In Proceedings of the 1st ACM SIGPLAN International Workshop on Language Models and Programming Languages (LMPL 2025), co-located with ICFP/SPLASH 2025. ACM, Singapore, 130–137. doi:10.1145/3759425.3763397 [6] Vitaly Chipounov, Volodymyr Kuznetsov, and George Candea. 2011. S2E: A Platform for In-Vivo Multi-Path Analysis of Software Systems. In Proceedings of the 16th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS XVI). ACM, New York, NY, USA, 265–278. doi:10.1145/1950365.1950396 [7] Jake Christensen, Ionut Mugurel Anghel, Rob Taglang, Mihai Chiroiu, and Radu Sion. 2020. {DECAF}: Automatic, adaptive de-bloating and hardening of {COTS} firmware. In 29th USENIX Security Symposium (USENIX Security 20). 1713–1730. [8] Emilio Coppa, Heng Yin, and Camil Demetrescu. 2022. Symfusion: hybrid instrumentation for concolic execution. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering. 1–12. [9] Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023. Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis. 423–435. [10] Sidong Feng and Chunyang Chen. 2023. Prompting Is All Your Need: Automated Android Bug Replay with Large Language Models. arXiv preprint arXiv:2306.01987 (2023). [11] Xiaotao Feng, Ruoxi Sun, Xiaogang Zhu, Minhui Xue, Sheng Wen, Dongxi Liu, Surya Nepal, and Yang Xiang. 2021. Snipuzz: Black-box fuzzing of iot firmware via message snippet inference. In Proceedings of the 2021 ACM SIGSAC conference on computer and communications security. 337–350. [12] Fraunhofer FKIE. accessible 2024. Detect common bug classes formally known as Common Weakness Enumerations (CWEs). [13] Wil Gibbs, Arvind S Raj, Jayakrishna Menon Vadayath, Hui Jun Tay, Justin Miller, Akshay Ajayan, Zion Leonahenahe Basque, Audrey Dutcher, Fangzhou Dong, Xavier Maso, et al. 2024. Operation mango: Scalable discovery of {Taint-Style} vulnerabilities in binary firmware services. In 33rd USENIX Security Symposium (USENIX Security 24). 7123–7139. [14] HyungSeok Han, JeongOh Kyea, Yonghwi Jin, Jinoh Kang, Brian Pak, and Insu Yun. 2023. Queryx: Symbolic query on decompiled code for finding bugs in COTS binaries. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 3279–3295. [15] Nasir Hussain, Haohan Chen, Chanh Tran, Philip Huang, Zhuohao Li, Pravir Chugh, William Chen, Ashish Kundu, and Yuan Tian. 2025. VulBinLLM: LLM-powered Vulnerability Detection for Stripped Binaries. arXiv:2505.22010 [cs.SE] https://arxiv.org/abs/2505.22010
XXX, XXX, XXX [16] Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770 (2023). [17] Caroline Lemieux, Jeevana Priya Inala, Shuvendu K Lahiri, and Siddhartha Sen. 2023. Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 919–931. [18] Haonan Li, Yu Hao, Yizhuo Zhai, and Zhiyun Qian. 2023. The Hitchhiker’s Guide to Program Analysis: A Journey with Large Language Models. arXiv preprint arXiv:2308.00245 (2023). [19] Wenqiang Li, Jiameng Shi, Fengjun Li, Jingqiang Lin, Wei Wang, and Le Guan. 2022. 𝜇AFL: non-intrusive feedback-driven fuzzing for microcontroller firmware. In Proceedings of the 44th International Conference on Software Engineering. 1–12. [20] Hao Liu, Carmelo Sferrazza, and Pieter Abbeel. 2023. Languages are rewards: Hindsight finetuning using human feedback. arXiv preprint arXiv:2302.02676 (2023). [21] Puzhuo Liu, Chengnian Sun, Yaowen Zheng, Xuan Feng, Chuan Qin, Yuncheng Wang, Zhenyang Xu, Zhi Li, Peng Di, Yu Jiang, et al. 2025. Llm-powered static binary taint analysis. ACM Transactions on Software Engineering and Methodology 34, 3 (2025), 1–36. [22] Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, and Rashmi Vinayak. 2025. Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’25). ACM, New York, NY, USA. doi:10.1145/3669940.3707215 [23] Microsoft Org. 2023. Copilot: The AI developer tool. [24] Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology. 1–22. [25] Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt. 2023. Examining zero-shot vulnerability repair with large language models. In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2339–2356. [26] Hammond Pearce, Benjamin Tan, Prashanth Krishnamurthy, Farshad Khorrami, Ramesh Karri, and Brendan Dolan-Gavitt. 2022. Pop Quiz! Can a Large Language Model Help With Reverse Engineering? arXiv preprint arXiv:2202.01142 (2022). [27] Kexin Pei, David Bieber, Kensen Shi, Charles Sutton, and Pengcheng Yin. 2023. Can large language models reason about program invariants?. In International Conference on Machine Learning. PMLR, 27496– 27520. [28] Sebastian Poeplau and Aurélien Francillon. 2020. Symbolic execution with {SymCC}: Don’t interpret, compile!. In 29th USENIX Security Symposium (USENIX Security 20). 181–198. [29] Nilo Redini, Aravind Machiry, Ruoyu Wang, Chad Spensky, Andrea Continella, Yan Shoshitaishvili, Christopher Kruegel, and Giovanni Vigna. 2020. Karonte: Detecting insecure multi-binary interactions in embedded firmware. In 2020 IEEE Symposium on Security and Privacy (SP). IEEE, 1544–1561. [30] Tobias Scharnowski, Nils Bars, Moritz Schloegel, Eric Gustafson, Marius Muench, Giovanni Vigna, Christopher Kruegel, Thorsten Holz, and Ali Abbasi. 2022. Fuzzware: Using precise {MMIO} modeling for effective firmware fuzzing. In 31st USENIX Security Symposium (USENIX Security 22). 1239–1256. [31] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems 36 (2023), 68539–68551.
XXX, XXX, XXX [32] Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In Thirty-seventh Conference on Neural Information Processing Systems. [33] Nick Stephens, John Grosen, Christopher Salls, Andrew Dutcher, Ruoyu Wang, Jacopo Corbetta, Yan Shoshitaishvili, Christopher Kruegel, and Giovanni Vigna. 2016. Driller: Augmenting fuzzing through selective symbolic execution.. In NDSS, Vol. 16. 1–16. [34] Bogdan Alexandru Stoica, Utsav Sethi, Yiming Su, Cyrus Zhou, Shan Lu, Jonathan Mace, Madanlal Musuvathi, and Suman Nath. 2024. If At First You Don’t Succeed, Try, Try, Again. . . ? Insights and LLMInformed Tooling for Detecting Retry Bugs in Software Systems. In Proceedings of the 30th ACM Symposium on Operating Systems Principles (SOSP ’24). ACM, New York, NY, USA. doi:10.1145/3694715.3695971 [35] Chengpeng Wang, Wuqi Zhang, Zian Su, Xiangzhe Xu, and Xiangyu Zhang. 2024. Sanitizing large language models in bug detection with data-flow. In Findings of the Association for Computational Linguistics: EMNLP 2024. 3790–3805. [36] Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345. [37] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Selfconsistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171 (2022). [38] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (2022), 24824–24837. [39] David A. Wheeler. accessible 2024. Flawfinder is a simple program that scans C/C++ source code and reports potential security flaws. [40] Yizhou Wu, Diyi Yang, Kexin Wang, Vincent Y. F. Tan, Canwen Xu, Zhenyu Yang, Xiang Li, Xiaoxiao Tan, et al. 2023. AutoGen: Enabling next-gen LLM applications via multi-agent conversation framework. arXiv preprint arXiv:2309.12307 (2023). [41] Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. 2024. OSCopilot: Towards Generalist Computer Agents with Self-Improvement. arXiv:2402.07456 [cs.AI] https://arxiv.org/abs/2402.07456 [42] Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864 (2023). [43] Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972 [cs.AI] https://arxiv.org/abs/2404.07972 [44] John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024. Swe-agent: Agentcomputer interfaces enable automated software engineering. Advances in Neural Information Processing Systems 37 (2024), 50528–50652. [45] John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. 2023. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. arXiv:2306.14898 [cs.CL] https://arxiv.org/ abs/2306.14898 [46] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601 (2023). [47] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and
XiangRui Zhang, Qiang Li, and Haining Wang acting in language models. arXiv preprint arXiv:2210.03629 (2022). [48] Insu Yun, Sangho Lee, Meng Xu, Yeongjin Jang, and Taesoo Kim. 2018. {QSYM}: A practical concolic execution engine tailored for hybrid fuzzing. In 27th USENIX Security Symposium (USENIX Security 18). 745–761. [49] Yaowen Zheng, Ali Davanian, Heng Yin, Chengyu Song, Hongsong Zhu, and Limin Sun. 2019. {FIRM-AFL}:{High-Throughput} greybox fuzzing of {IoT} firmware via augmented process emulation. In 28th USENIX Security Symposium (USENIX Security 19). 1099–1114.
A
System Implementation Table 11. Tools Available to Agents. Tool
Description
Radare2Tool
Provides Radare2 commands for binary analysis including disassembly, decompilation, and control flow analysis.
GrighraTool
Provides Grighra commands for reverse engineering and analysis.
To evaluate the feasibility of FORGE, we implemented a prototype system in Python. The system consists of the following core components: • Tools: Each agent is equipped with the tools summarized in Table 11. In particular, radare2 serves as a lightweight, scriptable reverse engineering framework that provides essential functionalities for binary analysis. It enables the agent to extract raw assembly code, cross-references (xrefs), pseudocode (when available), and relevant metadata such as strings or function information from binaries. • Output Schema: Each agent follows a JSON schema (thought, action, action_input, status). Here, thought represents the agent’s reasoning or decision rationale; action denotes the selected tool (from Table 11); action_input specifies the tool parameters; and status records the execution result or any associated error message. • Response Parsing: Structured JSON responses are extracted from the LLM’s output using regex matching, schema validation, and a bounded feedback–retry mechanism to ensure both syntactic and semantic robustness. • Memory: The conversation history is maintained as a sequence of role–content message pairs (system, user, assistant, tool, error, and parse error). This memory design provides a persistent context for agent reasoning and decision-making. All modules of FORGE are implemented from scratch without relying on existing frameworks. The source code is publicly available at: https://github.com/bjtu-SecurityLab/FORG E.
Feedback-Driven Execution for LLM-Based Binary Analysis
B
XXX, XXX, XXX
C
Prompt
This section details the system prompts used to guide the LLM-based agents for both the vulnerability discovery and validation phases. B.1
The validation phase takes a structured path from the discovery phase as its input. This process is crucial for eliminating false positives and confirming the exploitability of a potential vulnerability. The following example demonstrates how the system validates the CI vulnerability discovered in our analysis.
System Prompt for Vulnerability Discovery
The following prompt is used to initialize all agents during the discovery phase. It establishes the core principles of operation for finding new vulnerabilities. 1
You are a professional binary security analyst. Your mission is to comprehensively analyze the specified binary file, identify all externally controllable taint sources, delegate tracing to function-level agents, and report all exploitable paths.
2
**Core Principles:** - **Evidence-Based:** All analyses MUST be grounded in concrete evidence from the `r2` tool. No speculation. 5 - **Taint Identification:** Autonomously identify and validate all genuinely controllable external variables (e.g., HTTP params, NVRAM, IPC). The threat model assumes an attacker has network access and valid user credentials. Do not trace unexploitable taints. 6 - **Delegated Tracing:** You DO NOT perform deep taint tracing yourself. Your role is to delegate the tracing of specific data flows to function-level agents, ensuring a complete path from source to sink is reconstructed. 7 - **Focused Analysis:** Concentrate strictly on your assigned task. Your final output must be a comprehensive list of all evidence-based, exploitable paths. DO NOT provide fix suggestions or subjective commentary. 3 4
System prompt for the discovery phase.
C.1
1
{
2 3 4 5
6
7
8 9 10
11
}
"type": "CWE-78", "identifier": ["dev_name"], "propagation": [ "Source: HTTP query parameter 'dev_name' received from client request via QUERY_STRING environment variable", "Step: uVar1 = sym.get_querry_var(puVar3[-0x204], 0 x1aef0) --> User input from 'dev_name' parameter is retrieved at 0xa730", "Step: sym.imp.snprintf(puVar3 + -0x808 + -4, 0x800, 0x1af04, puVar3[-3]) --> Tainted user input is formatted into command string 'cfm post netctrl 51?op=3 ,string_info=%s' at 0xa7ac", "Sink: sym.imp.system(puVar3 + -0x808 + -4) --> Tainted command string is passed to system() at 0xa7c0" ], "reason": "process_datamanage_usbeject function retrieves user-controlled dev_name parameter via get_querry_var, directly formats it into a system command using snprintf, and executes via system().", "file_path": "../../karonte/Tenda/../_US_AC18V1.0BR_V15 .03.05.05../squashfs-root/usr/bin/app_data_center"
System Prompt for Vulnerability Validation A vulnerability and its Evidence Chain
During the validation phase, agents are initialized with a specific taint propagation path and guided by the following prompt to ensure rigorous verification. 1
Evidence Chain
The validation process is initiated with an evidence chain (JSON) generated during discovery.
12
B.2
Example for Vulnerability Validation
C.2
You are a binary call chain validation agent. Your SOLE mission is to strictly verify a given call chain provided as a clue in the specified binary.
After the verification, the agent produces a verified evidence chain (JSON). This report not only confirms the original finding but also enriches it with the missing authentication vulnerability.
2
**Verification Requirements:** - **Evidence-Only:** Base all judgments exclusively on evidence from `r2`. No guessing. 5 - **Path Verification:** Confirm if a reproducible propagation path exists from the specified source to the sink. Verify if the taint is genuinely exploitable under the defined threat model. 6 - **Success Case:** If verified, output the complete, evidence-backed propagation path with corrected addresses. 7 - **Failure Case:** If not verified, clearly state the reason (e.g., path does not exist, data is sanitized, sink is not reached). 3 4
1 2 3 4 5 6
7
System prompt for the validation phase.
Verified Evidence Chain
{
"accuracy": "accurate", "vulnerability": true, "propagation": [ "Source: An unauthenticated HTTP request is made to the path containing 'usbeject'.", "Step: At 0x9e6c, the dispatcher `do_request_process ` matches the 'usbeject' string (located at 0x1ad18) and, without any authentication check, calls ` process_datamanage_usbeject` at 0x9e8c.", "Step: Inside `process_datamanage_usbeject`, `sym. get_querry_var` is called at 0xa730 to retrieve the ' dev_name' string from the query parameters.",
XXX, XXX, XXX
8
9 10 11
12
}
XiangRui Zhang, Qiang Li, and Haining Wang
"Step: At 0xa7ac, the retrieved string is formatted into a command buffer using `sym.imp.snprintf` with format string 'cfm post netctrl 51?op=3,string_info=%s' (located at 0x1af04).", "Sink: At 0xa7c0, the unsanitized command buffer is passed directly to `sym.imp.system` for execution." ], "reason": "Verification successful. The data flow from the external source 'dev_name' to the `system` sink is confirmed. The propagation path is correct. Furthermore , analysis of the calling function `do_request_process` revealed that this entire code path is accessible without authentication, escalating the vulnerability from a standard Command Injection to an Unauthenticated Remote Code Execution. A PoC can be constructed by injecting shell metacharacters (e.g., `; reboot`) into the 'dev_name' parameter of a request to the 'usbeject' endpoint."
5
6 7 8
}
Turn 2: Root Agent identifies the dispatcher and vulnerable function. D.3
1
{
A vulnerability and its Verified Evidence Chain.
Example of Vulnerability Discovery
D.1
Turn 1: Identify the source
The system initiates the analysis by identifying a high-risk hypothesis and delegating it.
3 4 5 6
1
{
2
3 4 5
6 7 8
}
7
"thought": "I've identified the binary 'app_data_center' which likely handles web requests. A common vulnerability class is command injection from web parameters. I will create a task-level agent to investigate all data flows from HTTP parameters to command execution sinks like `system()`.", "action": "AgentTool", "action_input": { "task": "Analyze binary 'app_data_center'. Trace data flows from all HTTP query parameters to command execution sinks." }, "status": "continue"
8 9 10 11 12 13 14 15 16
D.4
Turn 2: Root Agent – Function-level Delegation
The Root Agent finds a potential entry point and dispatcher, then delegates deeper analysis. 1 2
3 4
{
"thought": "Using `r2`, I've identified ` do_request_process` as a central request dispatcher. It compares the URL path to several strings. One path, ' usbeject' (at 0x1ad18), leads to a call to ` process_datamanage_usbeject` at 0x9e8c. Unlike other paths, this one lacks a call to `authorization_check`. This is highly suspicious. I will delegate the analysis of `process_datamanage_usbeject` to a function-level agent, noting the lack of authentication.", "action": "AgentTool", "action_input": {
}
"thought": "I've analyzed `process_datamanage_usbeject`. The function retrieves the 'dev_name' parameter via ` get_querry_var` at 0xa730. This value is used in an ` snprintf` call at 0xa7ac, and the resulting buffer is passed directly to `system()` at 0xa7c0. This is a clear, unsanitized path from source to sink. I will report this path segment.", "action": "finish", "action_input": { "final_response": { "status": "SINK_REACHED", "path_segment": [ "Source: 'dev_name' parameter retrieved via `get_querry_var` at 0xa730.", "Step: `snprintf` at 0xa7ac formats the value into a command.", "Sink: `system` at 0xa7c0 executes the command." ], "reason_snippet": "Unsanitized user input ' dev_name' flows directly into a `system` call." } }, "status": "complete"
Turn 3: Child Agent identifies the sink and prepares its local report.
Turn 1: The initial analysis task. D.2
Turn 3: Child Agent – Local Discovery and Reporting
The Child Agent performs the focused analysis and finds the direct command injection chain. 2
D
"task": "Analyze `process_datamanage_usbeject`. Trace all external inputs within this function to potential sinks. Context: The entry to this function appears to be unauthenticated." }, "status": "continue"
Turn 4: Root Agent – Aggregating and Relaying Upwards
The Root Agent receives the report, combines it with its own crucial context (the missing authentication), and reports the full, severe vulnerability. 1 2
3 4
{
"thought": "My child agent has confirmed a command injection sink. When combined with my earlier discovery that the 'usbeject' path is unauthenticated, the full picture emerges: an unauthenticated remote command execution vulnerability. I will prepend my findings about the dispatcher to the child's path segment and report this complete path to the root agent.", "action": "finish", "action_input": {
Feedback-Driven Execution for LLM-Based Binary Analysis
5 6 7 8 9
10 11 12 13 14
15 16 17 18
}
"final_response": { "status": "PATH_COMPLETE", "full_path": [ "Source: Unauthenticated request to ' usbeject' endpoint handled by `do_request_process`.", "Step: No authentication check performed before calling `process_datamanage_usbeject` at 0x9e8c .", "Step: 'dev_name' parameter retrieved via ` get_querry_var` at 0xa730.", "Step: `snprintf` at 0xa7ac formats the value into a command.", "Sink: `system` at 0xa7c0 executes the command." ], "reason_snippet": "A missing authentication check allows an unsanitized user parameter to reach a ` system` call." } }, "status": "complete"
Turn 4: Root Agent combines its context with the child’s report. D.5
Turn 5: Final Report Assembly and Storage
Finally, the system receives the complete path, assembles the final, user-facing report, and assigns a critical risk score. 1 2
3 4 5 6 7 8 9 10 11 12 13 14 15 16
{
"thought": "I have received a complete, evidence-backed vulnerability path for an Unauthenticated RCE. This is a critical finding. I will assemble the final structured report with all details from the agent chain and assign a risk score reflecting the maximum severity.", "action": "finish", "action_input": { "final_response": { "type": "CWE-78", "additional_weaknesses": ["CWE-862"], "identifier": ["dev_name"], "propagation": [ "Source: HTTP query parameter 'dev_name' on the unauthenticated '/usbeject' endpoint.", "Step: `do_request_process` dispatches to ` process_datamanage_usbeject` without authentication.", "Step: At 0xa730, `get_querry_var` retrieves the user-supplied 'dev_name' string.", "Step: At 0xa7ac, `snprintf` formats this string into a command.", "Sink: At 0xa7c0, the malicious command string is executed by `system`." ], "reason": "The function ` process_datamanage_usbeject` contains a command injection vulnerability via the 'dev_name' parameter. Crucially, the calling function `do_request_process` fails to perform an authentication check for the ' usbeject' endpoint, making the vulnerability accessible without credentials. This combination allows for Unauthenticated Remote Code Execution.",
XXX, XXX, XXX
17 18 19
20 21 22 23
}
"risk_score": 9.0, "confidence": 9.0, "file_path": "../../karonte/Tenda/../_US_AC18V1 .0BR_V15.03.05.05../squashfs-root/usr/bin/ app_data_center" } }, "status": "complete"
Turn 5: Root agent assembles the final, structured report for storage.