FirmPilot: Evidence-Guided Multi-Agent Environment Recovery for IoT Firmware Rehosting Yanbing Shen, Fan Zhang, and Haitao Xu
arXiv:2607.14903v1 [cs.SE] 16 Jul 2026
Zhejiang University
Abstract—Firmware rehosting executes firmware images in emulated environments such as QEMU to enable scalable dynamic analysis of Internet of Things (IoT) devices. In practice, rehosting pipelines remain fragile across diverse real-world firmware images, as reaching an externally observable execution state depends on tightly coupled artifacts spanning boot scripts, persistent configuration (e.g., NVRAM-like key–value state), and network setup. Template-driven frameworks often fail to accommodate long-tail vendor conventions, while unconstrained use of large language models (LLMs) risks unsupported modifications and irreproducible executions. We introduce F IRM P ILOT, an evidence-guided multi-agent framework for environment recovery in firmware rehosting. F IRM P ILOT reformulates rehosting as iterative environment reconstruction in which a search agent grounds decisions through similarity-based retrieval, a planner coordinates executionaccepted transitions, and specialized agents recover filesystem/init artifacts, persistent state, and network exposure. Through repeated execution and evidence-grounded artifact deltas, the system resolves cross-layer dependencies across boot, state, and networking that otherwise prevent firmware executions from reaching a stable, externally reachable state in emulation. Evaluated on the large-scale, real-world LFwC firmware corpus, F IRM P ILOT improves web-service reachability over FirmAE from 25.49% to 52.39% and network reachability from 39.30% to 71.93%. The resulting rehosts raise the average number of detected services per firmware from 0.86 to 1.62 and support downstream analysis workflows, including RouterSploit interaction and protocol-aware fuzzing over recovered service surfaces. The evaluation shows that evidence- and feedbackgrounded agent coordination improves rehosting success, service recovery, and downstream utility in automated firmware rehosting. Index Terms—Firmware Rehosting, IoT security, Emulation, Multi-Agent System, Retrieval-Augmented Generation
I. I NTRODUCTION IoT security analysis often starts from firmware images rather than source code or live devices. Analysts need to know which services actually start, which persistent configuration state gates them, and whether security tools can interact with the resulting runtime. Whole-system firmware rehosting answers this need by unpacking a firmware image, reconstructing a bootable environment, and executing it in QEMU so that network-facing behavior can be probed at scale [1], [2], [12]. The practical barrier is that rehosting remains brittle on heterogeneous firmware corpora [4], [13], [14]. A failed run is rarely caused by one missing command. Boot scripts, vendor wrappers, NVRAM-like key–value state, interface naming, and service launch order interact with each other; a change that makes the kernel boot may still leave the web stack disabled or bound to an unreachable interface. The available feedback
is also partial because serial logs, reachability probes, and service snapshots expose symptoms rather than a complete specification of the original device environment. This setting is well suited to LLM-based agents, which can interpret logs, scripts, and vendor-specific conventions. However, treating an LLM as an unconstrained operator is a poor fit for rehosting because edits affect boot-critical artifacts, success depends on repeated execution, and uncontrolled generation can make results irreproducible. A useful agentic system must therefore make model influence bounded and auditable. Decisions should be grounded in observed or retrieved evidence, actions should pass through typed interfaces, and claims should be validated by fixed probes rather than by model explanations. We present F IRM P ILOT, an evidence-guided multi-agent framework for firmware rehosting. F IRM P ILOT treats rehosting as evidence-indexed environment reconstruction. The system indexes runtime observations and retrieved evidence, maps them to typed transitions over boot artifacts, persistent state, and network exposure, and accepts a transition only after reexecution and probing. A S EARCH agent retrieves firmwarespecific evidence, a P LAN agent schedules environmentrecovery actions, and F ILE, NVRAM, and N ETWORK agents materialize structured deltas through bounded action interfaces. This design turns model reasoning into a constrained control plane for state and exposure recovery. Beyond basic reachability, F IRM P ILOT targets securityworkflow fidelity, namely service-facing, protocol-consistent behavior that lets standard tools execute and leave auditable artifacts in a controlled offline environment. This shifts evaluation from root-page response to whether recovered firmware executions can sustain service discovery, protocolaware fuzzing, and RouterSploit interactions. We evaluate F IRM P ILOT on 10,033 images from the largescale, real-world LFwC firmware corpus and on the 1,122image public FirmAE benchmark. On the LFwC corpus, F IRM P ILOT improves web-service reachability over FirmAE from 25.49% to 52.39% and network reachability from 39.30% to 71.93%; on the public benchmark, where FirmAE already performs strongly, F IRM P ILOT raises Web reachability from 79.4% to 82.2%. Full-corpus ablations attribute the gain to planning, state synthesis, network exposure recovery, retrieval, and filesystem/init recovery, while a Claude Code baseline reaches only 5.43% web-service reachability on the same corpus. Across all 5,256 web-reachable LFwC rehosts, downstream analyses enumerate services, validate HTTP/HTTPS
TABLE I R EPRESENTATIVE F IRMWARE E MULATION S YSTEMS FOR I OT S ECURITY A NALYSIS System
Fidelity
Emulation
Analysis
Automation
Firmadyne [1], FirmAE [2] FirmGuide [3], PANDaWAN [4] Jetset [5], Greenhouse [6], FIRMWELL [7] Avatar2 [8] HALucinator [9], P2IM [10], Fuzzware [11] F IRM P ILOT (this work)
System and services System and services Service and dependencies System and peripherals Peripherals and MMIO System and services
OS emulation OS emulation User-space Hybrid Model emulation OS emulation
Dynamic Dynamic Dynamic Dynamic Fuzzing Dynamic
Pipeline Diagnosis Dependency-aware Assisted Models Agentic
fingerprints, support manual validation of 193 RouterSploit findings, and enable protocol-aware fuzzing. We make the following contributions. • We formulate firmware rehosting as evidence-indexed environment reconstruction under partial observability, where progress depends on resolving coupled boot, state, and network dependencies through execution-accepted transitions. • We design F IRM P ILOT , an evidence-guided multi-agent environment-recovery loop with typed action interfaces for filesystem/init recovery, NVRAM state synthesis, and network exposure recovery. • We conduct a large-scale evaluation on 10,033 LFwC images and the public FirmAE benchmark, combining full-corpus ablations, a general-purpose coding-agent baseline, and downstream workflow analyses to show that F IRM P ILOT improves rehosting success and produces 5,256 web-reachable executions that support service discovery, manually validated RouterSploit findings, and protocol-aware fuzzing. II. BACKGROUND A. Firmware Emulation Firmware emulation executes firmware code in a controlled environment so analysts can observe runtime behavior without physical devices. IoT-security systems range from wholesystem QEMU rehosting to user-space service and dependency rehosting, hybrid execution, and peripheral/MMIO modeling; surveys show that these choices trade off scale, automation, and fidelity [13], [14]. Table I positions representative systems. Whole-system pipelines remain the primary option for scalable service-facing analysis, but they are brittle on heterogeneous corpora because success depends on coupled init, persistent-state, and networking assumptions [13], [14]. Guidance systems improve observability, but they do not close the loop by converting evidence into bounded environment transitions [3], [4]. F IRM P ILOT therefore targets whole-system rehosting as evidence-indexed environment reconstruction. B. Rehosting Failure Modes Typical pipelines perform extraction, architecture detection, image construction, init selection, networking setup, and service probing. Each stage can fail when the emulator diverges from the device environment and runtime feedback is incomplete. Common failure patterns are as follows. • Boot and Init Discrepancies. Incorrect init selection, divergent init systems, or vendor wrappers that assume device-specific mounts and environment.
Hardware No No No Yes No No
Missing Persistent Configuration State. Absent NVRAMbacked values for interface naming, addressing, service toggles, and authentication settings, preventing services from starting even when the kernel boots. • Network Configuration Discrepancies. Inconsistent interface naming, bridging, or routes, leading to reachability without usable service exposure. • Cross-Binary Service Dependencies. Multi-process service stacks that fail under incomplete state or partial startup, leading to fragile partial progress [15]. These failures often cascade as a filesystem transition changes init progress, new logs reveal missing state keys, constrained defaults enable interface setup, and services then become observable under probing. Static one-shot pipelines do not exploit this feedback. User-space rehosting [5]–[7], hybrid execution [8], [16], and peripheral/MMIO modeling [9], [10] improve fidelity in complementary settings, but whole-system service analysis still needs scalable environment recovery. •
C. LLM Agent Foundations Large language models can interpret heterogeneous logs and scripts, but rehosting decisions must be grounded in verifiable signals and bounded in what they can change. We use retrievalaugmented prompting [17] and role-specific prompts for multistep decisions [18]–[20]. Unlike security agents for penetrationtesting orchestration, fuzzing guidance, and program/binary analysis [21]–[24], our agents act on boot-critical artifacts under strict budgets, so each role has a bounded action surface. III. M OTIVATION AND D ESIGN C HALLENGES Whole-system rehosting is the most direct path to scaling service-facing security workflows on heterogeneous IoT firmware, yet progress depends on iterative environment reconstruction across init, persistent state, networking, and services. A practical agentic system must interpret partial runtime evidence, apply small artifact deltas under fixed budgets, and avoid drift into unrealistic device state. This yields three design challenges. • C1–Limited Domain Knowledge in Firmware Emulation. General-purpose models have uneven coverage of firmware rehosting toolchains and device-specific boot conventions, so plausible suggestions can still conflict with emulator constraints and lead to fragile fixes. • C2–Inference Under Partial and Noisy Observability. Agents must infer the next environment transition from incomplete, vendor-specific logs, probes, and filesystem evidence while avoiding unsupported causal claims.
NVRAM Agent
IoT Firmware Binary File System Kernel
NVRAM Table
Key- Values
NVRAM State Recovery ①Report Failure&Logs
Emulation Environment (QEMU)
Iterative Repair Cycle
③Dispatch Repair Task
Planner Agent ②Consult Knowledge Base(RAG)
④Generate Patch
File Agent
Init/Service Patching
Patch Generation
Network Agent
⑤Apply Patch &Reboot
Knowledge Base
Search Agent
Enable Interfaces
Fig. 1. System Architecture of F IRM P ILOT
C3–Constrained and Convergent Environment Actuation. Actions can mutate boot scripts, configuration files, and emulator arguments, so deltas must be typed, and the loop must converge under strict budgets without drifting toward unrealistic state injection. These challenges motivate a separation between semantic inference and environment transition. In firmware rehosting, the unknown object is not a source-level patch but an execution environment whose state is only partially observable through logs and probes. F IRM P ILOT therefore treats LLM outputs as hypotheses over a constrained transition space, while state changes are performed by typed operators and accepted through execution evidence. This turns environment recovery into an evidence-conditioned search problem in which the model helps identify likely missing dependencies, the system exposes only the corresponding action surface, and progress is attributed to observed state changes, not model rationales. •
Each action agent a ∈ pt models a bounded function from context to an artifact delta, a(ct ) = δt,a . Applying these deltas yields the next artifact state. M xt+1 ← xt ⊕ δt,a . a∈pt
The L operator ⊕ applies a delta to the current artifacts, and a∈pt denotes sequential composition in plan order. The agent set is A = {S EARCH, P LAN, F ILE, NVRAM, N ETWORK},
where S EARCH and P LAN acquire evidence and schedule actions, while F ILE, NVRAM, and N ETWORK update xt through filesystem/init deltas, NVRAM overlays, and network artifacts. Model outputs are treated as candidate environment transitions rather than direct authority over the firmware image. A transition becomes part of the run state only after it is emitted through a typed interface, rendered by deterministic tooling, IV. S YSTEM D ESIGN and validated by the next execution and fixed probes. The loop stops when Web reachability is achieved, no admissible action A. Overview Figure 1 illustrates F IRM P ILOT as an iterative environment- remains under unchanged evidence, or the outer budget is recovery loop over a shared context. Each iteration executes exhausted. The context records artifact versions, observations, the current artifacts in QEMU, collects serial logs and probe retrieved support, and accepted deltas as an execution-indexed results ( 1 ), retrieves firmware-specific evidence ( 2 ), plans a causal trace, preserving attribution between evidence, action, dependency-aware transition schedule ( 3 ), and asks specialized and outcome even when progress is non-monotonic. agents to emit typed deltas ( 4 ). The orchestrator applies C. Search Agent for Evidence Acquisition accepted deltas, reboots the firmware, and reruns fixed probes ( 5 ). The loop addresses C1 through retrieval, C2 through execution-grounded planning, and C3 through budgeted, typed, and replayable actuation. B. Environment-Recovery Loop Formalization We model F IRM P ILOT as a bounded transition system. At iteration t, artifact state xt contains filesystem overlays, launch configuration, NVRAM overrides, and network exposure scripts. Executing xt in QEMU yields observations ot such as serial logs, reachability results, and service snapshots. The orchestrator extracts evidence Et and maintains ct = ⟨x≤t , o≤t , E≤t ⟩; S EARCH augments this context and P LAN selects an ordered environment-recovery plan pt .
Fig. 2. Design of the S EARCH Agent
Role and Evidence Flow. The S EARCH agent gives the loop a controlled path to outside knowledge. Its query packet is built from firmware metadata, vendor/model/board identifiers,
architecture, recent serial-log windows, probe summaries, missing binary or library names, NVRAM key names, service binaries, port/banner observations, and compact failure tokens. Retrieval targets Web and GitHub-visible evidence such as vendor documentation, public code repositories, support/forum threads, and prior snippets cached by provenance. Retrieved resources enter the shared context only as cited snippets with query, URL, timestamp, source type, and similarity score. Search keeps a bounded candidate set and admits only the highest-ranked cited snippets to the shared context. It is readonly and cannot change files, state keys, or network parameters. This makes external evidence available to the planner without turning retrieval into an environment-transition mechanism [17]. D. Plan Agent for Orchestration and Scheduling
express them as typed startup deltas and may abstain when evidence is insufficient. F IRM P ILOT validates paths, normalizes arguments, rejects shell metacharacters and destructive utilities, and defaults to existing on-image init scripts when external evidence is insufficient.
Fig. 4. Design of the F ILE Agent
F. NVRAM Agent for Persistent State Synthesis
Fig. 3. Design of the P LAN Agent
Responsibilities and Feedback. The P LAN agent is the loop control point. It consumes artifact readiness, init and service candidates, probe outcomes, recent error signatures, retrieved snippets, port signals, and accepted deltas, then emits an ordered schedule over F ILE, NVRAM, and N ETWORK with a failure class and stop/retry condition. The schedule is explicit because persistent-state recovery can enable service startup, and exposure recovery is meaningful only after a daemon binds or a likely bind address appears. Actions execute through typed interfaces. F ILE emits entrypoint and startup deltas, NVRAM emits key/value/source/guard tuples, and N ETWORK emits IP, interface, bridge/VLAN, port, and QEMU-argument tuples. The planner never writes artifacts directly; the next plan observes updated logs and probe state after emulation-relevant mutations. Table II defines the role contract. S EARCH and P LAN reason over evidence and ordering, while F ILE, NVRAM, and N ETWORK enact typed transitions over boot artifacts, persistent state, and exposure state. E. File Agent for Boot-Environment Recovery Objective and Artifacts. The F ILE agent recovers boottime filesystem artifacts so static firmware contents reach a runnable user space. It maintains candidate init entrypoints and service-launch specifications, emitting minimal filesystem deltas plus replayable records of inferred init paths, startup commands, and evaluated variants. Discovery and Service Selection. Figure 4 separates bootenvironment inference from artifact mutation. The agent ranks init candidates from filesystem enumeration, common init locations, kernel command-line hints, and retrieved evidence. The LLM may add vendor-specific hypotheses, but it must
Fig. 5. Design of the NVRAM Agent
Purpose and Motivation. The NVRAM agent reconstructs persistent configuration dependencies required for service startup. Embedded services consult NVRAM-like stores for interface naming, access control, feature flags, and managementservice settings; missing keys can keep daemons disabled even after boot and network recovery. Reconstructed state is materialized as an external key–value overlay rather than a root-filesystem mutation. Grounded State Synthesis. Figure 5 shows the staterecovery boundary. The agent does not apply a universal NVRAM template; it treats persistent state as a perfirmware inference problem. It builds candidates from runtime NVRAM reads, default files, configuration strings, and model/board/serial/MAC/LAN/Web evidence, then asks the LLM to resolve only prioritized keys not already determined by firmware evidence. Each proposal must be expressed as a key, value, source, and guard. The action interface preserves identity values, rejects malformed fragments and broad credential-series artifacts, and drops empty entries before producing the runtime overlay. State recovery runs when execution evidence shows missing NVRAM devices, invalid flash configuration, pre-bind crashes, or missing product identity; the transition is accepted only when the next execution supports it. G. Network Agent for Connectivity and Exposure Recovery Purpose and Design Goal. The N ETWORK agent recovers the exposure layer needed for firmware-started services to become reachable, including guest-side connectivity and host-side forwarding/TAP state. It maps noisy logs and vendor-specific
TABLE II AGENT ROLES AND B OUNDED ACTION I NTERFACES IN F IRM P ILOT Agent
Evidence used
Bounded action
Execution guard
S EARCH
Serial/probe logs; firmware metadata; service/failure tokens Probe state; artifact readiness; recent errors; retrieved snippets Root filesystem; init candidates; service binaries; retrieval context Access traces; configuration files; missing-key evidence Boot logs; listener/interface evidence; probe outcomes
Bounded Web/GitHub queries; top-5 cited snippets Agent sequence; failure class; stop/retry state
Provenance logging; retrieval cache only; no artifact mutation. Budget and dependency checks; no direct filesystem, state, or network mutation. Path checks; command filtering; replayable startup deltas. Value filtering; malformed entries removed; overlays isolated from rootfs edits. Consistency checks; rendering from accepted exposure deltas; fixed reachability probes.
P LAN F ILE NVRAM N ETWORK
Entrypoint; service command; filesystem/startup delta Key, value, source, and guard overlays IP, interface, bridge/VLAN, port, and QEMU-argument tuples
Fig. 6. Design of the N ETWORK Agent
interface conventions into a structured network representation, then selects the smallest evidence-consistent update that improves reachability while preserving application behavior. Inference, Exposure, and Validation. Figure 6 separates exposure inference from environment mutation. The agent reasons over interfaces, addresses, VLANs, bridges, MAC transitions, and service bind addresses using serial logs, probes, NVRAM LAN evidence, and the FirmAE-compatible seed. It then emits a typed exposure delta rather than a broad forwarding rule. The action interface materializes only the selected IP/interface, bridge/VLAN, port, and QEMU-argument tuples. DHCP-like guests use user-mode forwarding for observed bind ports, TAP mode preserves guest subnets and exported probe IPs, and ARM runs keep one active interface with additional probe IPs for multi-address stacks. Updates are accepted only when re-execution confirms both network reachability and a live service endpoint under the fixed probes. V. E XPERIMENTAL E VALUATION We evaluate F IRM P ILOT as an environment-aware firmware rehosting system. The experiments ask four questions about end-to-end effectiveness, mechanism contribution, domain specificity, and downstream workflow evidence. • RQ1 Rehosting effectiveness. Does F IRM P ILOT improve network and web-service reachability over a widely used automated rehosting baseline on a large firmware corpus? • RQ2 Mechanisms. Which environment-recovery mechanisms account for the observed improvement, and are their effects measurable at full-corpus scale? • RQ3 Domain specificity. Can a general-purpose coding agent substitute for a firmware-specific recovery workflow? • RQ4 Downstream security workflow support. Do webreachable F IRM P ILOT rehosts support downstream security workflows beyond reachability?
Dataset and baselines. The primary large-scale dataset is LFwC [25], a public Linux-firmware corpus built for reproducible firmware vulnerability research with documented acquisition metadata, unpacking checks, deduplication, content identification, and ground-truth annotations. Its metadata snapshot contains 10,913 records; 880 vendor or archival links were no longer reachable, leaving 10,033 locally executable images. These 10,033 images are the fixed denominator for all main LFwC comparisons, ablations, and general-agent baselines. We compare F IRM P ILOT against FirmAE [2], a QEMU-based rehosting pipeline from the Firmadyne lineage [1] that produces comparable disk images, run scripts, logs, and probes. We also evaluate on the public FirmAE benchmark, which contains 1,122 images across 8 vendors and ARM-LE/MIPSBE/MIPS-LE architectures. This established benchmark is closely aligned with FirmAE’s original templates and provides a comparability point, while LFwC supplies the main corpus. Comparison and ablation protocol. The evaluation separates three sources of evidence, namely a deterministic rehosting baseline, an internal role ablation, and a general-purpose codingagent baseline. FirmAE represents the template-driven deterministic alternative, with fixed image construction, boot inference, NVRAM defaults, network heuristics, QEMU execution, and the same probes, but no retrieval, planning, or typed agent actions. RQ1 measures the deterministic-to-F IRM P ILOT gap under identical predicates. RQ2 keeps F IRM P ILOT’s runnable substrate while replacing one adaptive role at a time with its deterministic fallback. In this setting, no-Search uses only local evidence, no-Plan follows a fixed dependency order, no-File uses baseline-style init and launch candidates, no-NVRAM retains FirmAE-compatible defaults, and no-Network retains deterministic network inference. RQ3 replaces the specialized action space with a general coding agent under the same dataset, timeout, input evidence, and success predicates. Execution policy and metrics. All RQ1–RQ3 systems use a QEMU-based execution path [12], fixed probes, and identical success predicates. All agents use DeepSeek V4 Flash as the LLM backend throughout the evaluation. Consistent with FirmAE’s Docker evaluation workflow, which uses a 2,400second per-firmware emulation-result check, each RQ1–RQ3 image receives the same 2,400-second wall-clock budget for rehosting and probing; FirmAE, F IRM P ILOT, and Claude Code all use this cap, and F IRM P ILOT runs at most three planner rounds within it. Network reachable requires a standardized
TABLE III V ENDOR -L EVEL P ING AND W EB -S ERVICE R EACHABILITY ON THE P UBLIC F IRM AE B ENCHMARK AND LF W C Public FirmAE benchmark Images
Ping FirmAE
Netgear D-Link TP-Link Trendnet ASUS Linksys Belkin Zyxel Overall
375 262 148 118 107 55 37 20
351 (93.6%) 249 (95.0%) 148 (100.0%) 101 (85.6%) 63 (58.9%) 48 (87.3%) 30 (81.1%) 18 (90.0%)
Web service F IRM P ILOT
FirmAE
Vendor
Ping
Images
F IRM P ILOT
356 (94.9%) 336 (89.6%) 345 (92.0%) Netgear 249 (95.0%) 231 (88.2%) 236 (90.1%) D-Link 148 (100.0%) 113 (76.4%) 119 (80.4%) ASUS 101 (85.6%) 73 (61.9%) 78 (66.1%) Ubiquiti 64 (59.8%) 62 (57.9%) 64 (59.8%) TP-Link 49 (89.1%) 44 (80.0%) 48 (87.3%) Trendnet 31 (83.8%) 22 (59.5%) 22 (59.5%) AVM 18 (90.0%) 10 (50.0%) 10 (50.0%) Linksys EnGenius
1,122 1,008 (89.8%) 1,016 (90.6%) 891 (79.4%) 922 (82.2%) Overall
ICMP response. Web-service reachable requires a completed HTTP(S) transaction to an exported guest-IP candidate, using bounded GET requests to /. HTTPS candidates complete the TLS handshake before request delivery, and the probe records status, redirects, authentication challenges, banners, and fingerprints when present. Success includes root pages, redirects, authentication gates, and other parseable management responses; status 200 is not required. Each request uses a twosecond timeout inside the image budget. The predicates separate network liveness from service readiness. ICMP measures packet exchange, while Web success requires a reachable and parseable management endpoint. This keeps Ping/Web columns directly comparable across FirmAE, F IRM P ILOT, and the general-agent baseline. RQ4 evaluates whether the recovered rehosts can support downstream analysis workflows beyond basic reachability. The audit reruns service discovery with nmap -O -sV, retries incomplete scans with nmap -Pn -n -sV, and credits valid scans only when complete XML contains auditable service records. HTTP/HTTPS evidence requires an endpoint, service label, protocol, and available banner/fingerprint fields from nmap or confirmation probes. RouterSploit and protocol-aware fuzzing then operate on this recovered surface. Counts use the 10,033-image denominator unless a table states otherwise. Evaluation records. All reported counts are computed from per-image execution records rather than ad hoc log inspection. Each record binds the firmware identifier, execution budget, generated artifacts, serial output, probe transcript, and final reachability labels. RQ4 records additionally attach the servicediscovery, RouterSploit, and fuzzing artifacts credited in the downstream workflow. This common record structure keeps the LFwC, ablation, and general-agent comparisons aligned at the image level.
Web service
FirmAE
F IRM P ILOT
2,551 1,337 (52.41%) 2,254 (88.36%) 1,909 175 (9.17%) 726 (38.03%) 1,645 871 (52.95%) 1,310 (79.64%) 1,394 312 (22.38%) 897 (64.35%) 1,163 723 (62.17%) 1,129 (97.08%) 752 379 (50.40%) 565 (75.13%) 301 43 (14.29%) 83 (27.57%) 173 6 (3.47%) 110 (63.58%) 145 97 (66.90%) 143 (98.62%)
FirmAE
F IRM P ILOT
879 (34.46%) 1,679 (65.82%) 130 (6.81%) 574 (30.07%) 757 (46.02%) 1,161 (70.58%) 13 (0.93%) 396 (28.41%) 468 (40.24%) 817 (70.25%) 268 (35.64%) 381 (50.66%) 16 (5.32%) 39 (12.96%) 0 (0.00%) 92 (53.18%) 26 (17.93%) 117 (80.69%)
10,033 3,943 (39.30%) 7,217 (71.93%) 2,557 (25.49%) 5,256 (52.39%) 86
80 Success rate (%)
Vendor
LFwC
60
Ping
74
Web
70 58 52
51
54
51
40 20 0
MIPS BE
MIPS LE
ARM LE
Other/ Rare
Fig. 7. LFwC Ping/Web Success by Firmware Architecture
corpus. F IRM P ILOT increases these numbers to 7,217 networkreachable images and 5,256 web-service-reachable images, corresponding to 71.93% and 52.39%. This is an absolute improvement of 3,274 network-reachable images and 2,699 web-service-reachable images over FirmAE; Table III reports the public-benchmark and LFwC vendor results under the same Ping/Web-service columns. Benchmark and long-tail effects. On the public FirmAE benchmark, where existing templates already cover many cases, F IRM P ILOT preserves the high reachability level expected on this established setting and raises Web-service reachability from 79.4% to 82.2%. The LFwC results expose a larger gap under heterogeneous, long-tail firmware images. For Netgear, ASUS, TP-Link, Linksys, and EnGenius, F IRM P ILOT reaches webservice rates above 50%. For D-Link, Ubiquiti, and AVM, the absolute rates remain lower, but F IRM P ILOT still substantially improves over FirmAE, especially on Ubiquiti where FirmAE reaches only 13 web services while F IRM P ILOT reaches 396. This contrast shows that environment recovery is most valuable when template coverage is sparse and reachability depends on coordinating boot, persistent state, and network exposure. Architecture split. Figure 7 complements the vendor view with an endian-aware primary firmware architecture split of A. RQ1 Overall Rehosting Effectiveness the same 10,033 LFwC images. F IRM P ILOT reaches Web Overall reachability. RQ1 tests the central rehosting claim services on 1,588 MIPS-BE, 942 MIPS-LE, and 2,192 ARMthat a rehost becomes useful for interactive analysis only when LE images, while making 2,309, 1,389, and 2,951 images the firmware exposes an externally reachable service, not merely network-reachable in the same groups. These router-class when the kernel boots or a network stack responds. Under architecture groups account for 4,722 of the 5,256 Web-service identical probes on the full 10,033-image LFwC set, FirmAE successes; PowerPC, x86, unresolved-endian, and rare groups makes 3,943 images network reachable and 2,557 images web- are aggregated into the long tail. service reachable, corresponding to 39.30% and 25.49% of the Residual failure modes. The final-state distribution in
5,256 (52.39%)
Web-service reachable 2,604 (25.95%)
Runtime emulation failed
1,675 (16.69%)
Extraction/registration failed
498 (4.96%)
Architecture handling failed 0
2000
4000
Fig. 8. Final-state distribution for the full F IRM P ILOT run on LFwC.
Figure 8 shows that F IRM P ILOT’s failures are not dominated by one residual bottleneck. Among the 10,033 images, 5,256 reach a web service. The remaining images primarily fail because of runtime emulation failures (2,604 images), extraction or registration failures (1,675 images), and architecture-handling failures (498 images). This distribution shows that the remaining failures are spread across extraction, architecture handling, runtime execution, and late-stage service launch rather than concentrated in one probe or service check. Interpretation. The RQ1 result is strongest on the large LFwC corpus because LFwC contains vendor and version diversity that stresses fixed templates. F IRM P ILOT does not merely increase Ping success and then inherit Web success automatically. The absolute Web gain of 2,699 images is close to the network gain in scale but not identical in source, since many additional successes require persistent-state synthesis for service toggles and exposure recovery for bind-address mismatches after basic boot progress. The vendor-level table makes this visible. Large gains on Ubiquiti, Linksys, EnGenius, and D-Link indicate that environment reconstruction exposes service paths left dormant by static baselines. B. RQ2 Contribution of Specialized Agents Ablation setup. RQ2 evaluates specialized roles after holding the deterministic execution substrate fixed. The deterministic alternative is template-driven rehosting represented by FirmAE, which fixes image construction, boot inference, NVRAM defaults, and network setup through heuristic templates. We run leave-one-module ablations for SearchAgent, PlanAgent, FileAgent, NVRAMAgent, and NetworkAgent on the same 10,033-image LFwC set. Each ablation keeps QEMU execution, boot hooks, overlay infrastructure, renderers, and probes runnable while replacing one agent’s adaptive decisions with the closest deterministic fallback used by the substrate. State and exposure fidelity. The ablation results treat persistent state and network exposure as first-order recovery dimensions rather than implementation details. NVRAMAgent admits candidate keys from execution-consulted reads, missingkey traces, defaults and configuration files, static strings, and device/LAN/Web identity evidence. Its typed action interface preserves source and guard fields, while deterministic filters reject malformed fragments, web assets, symbol-like tokens, broad credential-series artifacts, and empty values before overlay materialization. NetworkAgent follows the same evidence discipline by deriving exposure deltas from interface/IP traces, bridge/VLAN/MAC events, service bind addresses, NVRAM LAN evidence, and the FirmAE-compatible seed,
then rendering only selected interface, address, bridge/VLAN, port, and QEMU-argument tuples. Full-success state/exposure distribution. We audit the same 5,256 Web-success images used throughout the downstream workflow analysis. NVRAM key files appear on 1,975 images, and nonempty NVRAM materialization covers 1,971 images with 479,440 key entries and 5,320 file-backed entries. The dominant key classes are system/default state (170,925), wireless/region state (119,536), LAN/interface state (105,064), WAN/DNS/gateway state (33,391), service exposure (30,995), identity/model state (13,068), and authentication/access state (11,781). Network exposure artifacts cover 5,191 images, materializing 5,370 exported IP entries, 2,703 QEMU networkargument artifacts, 2,607 bridge/VLAN/interface setup scripts, 161 multi-IP cases, and service-probe records on 2,696 images. Paired full-success effect. RQ2 measures each role on the 5,256 accepted Web-success executions from the full system. A loss is counted when one of these executions no longer reaches Web under the corresponding deterministic fallback. Replacing PlanAgent loses 2,472 images across 590 brand-device families; replacing NVRAMAgent loses 2,003 images across 462 families, with 1,671 also losing Ping; replacing SearchAgent, NetworkAgent, and FileAgent loses 1,568, 1,558, and 1,525 images, respectively. The NVRAM and Network dependency sets are coupled but distinct. In this split, 1,484 images fail under both removals, 519 only without NVRAMAgent, and 74 only without NetworkAgent. In an ASUS DSL-N10_C1 case, the full run reconstructs 198 execution-consulted NVRAM keys across LAN, WAN, UPnP, region, and interface state, starts /usr/sbin/httpd, and passes both probes; the no-NVRAM run loses Web reachability. In an ASUS DSL-AC88U case, the full run recovers br0 at 192.168.1.1, maps the management address through the observed bridge/interface relation, and exposes HTTP to the probe path; the no-Network run loses Web reachability. These cases link aggregate losses to concrete transitions and show corpus-scale state and exposure recovery. Agent contributions. Figure 9 reports how many full-system Web successes remain reachable after each leave-one-module replacement. The full system reaches 5,256 images. Replacing PlanAgent with a fixed schedule causes the largest drop, retaining 2,784 of them and losing 2,472. Replacing NVRAMAgent with deterministic defaults retains 3,253 and loses 2,003. Replacing FileAgent, NetworkAgent, and SearchAgent with their deterministic fallbacks retains 3,731, 3,698, and 3,688, respectively. These paired losses show that the agents are complementary. Planning, persistent-state synthesis, network exposure, retrieval, and filesystem/init recovery each contribute measurable gains beyond template-driven fallbacks. Mechanism interpretation. PlanAgent is the most influential component because environment transitions are order-sensitive. Network exposure before persistent-state recovery, or startup changes without runtime evidence, can create conflicting states. NVRAMAgent confirms the importance of executionconsulted state such as region, device mode, interface names, and authentication defaults, accounting for 2,003 lost Web
5,256
Full system No FileAgent
3,731 (-1,525)
No NetworkAgent
3,698 (-1,558)
No SearchAgent
3,688 (-1,568)
12.43M
CI fuzz payloads 1.85M
BOF fuzz payloads 63.6K
RouterSploit runs
11.7K
Candidate analysis dirs 10
3,253 (-2,003)
No NVRAMAgent
16.3K
Open-port observations
4
10
5
10
6
10
7
Fig. 10. RQ4 Workflow Scale on 5,256 Web-Success Images 2,784 (-2,472)
No PlanAgent
Code has 231 unique web successes, but F IRM P ILOT has 4,942 unique web successes. Measured against F IRM P ILOT’s Retained Web-success images web-success set, Claude Code recovers only 5.97% of the Fig. 9. Retained full-system Web successes under leave-one-module ablations; images that F IRM P ILOT can expose as web services. This gap labels on ablated bars show lost full-system successes. indicates that the task is not simply to generate plausible shell TABLE IV commands or code edits. The system must identify firmwareG ENERAL -AGENT BASELINE ON THE 10,033-I MAGE LF W C C ORPUS specific state assumptions, apply bounded artifact transitions through structured interfaces, and accept changes only after Metric Claude Code F IRM P ILOT Delta / ratio execution-grounded validation. Ping reachable 729 (7.27%) 7,217 (71.93%) +64.66 pp / 9.90× Web reachable 545 (5.43%) 5,256 (52.39%) +46.96 pp / 9.64× Resource profile. The token profile strengthens this conclusion. Using the same 10,033-image LFwC denominator, Claude Model tokens/image 3.61M 12.4K 291× fewer Code’s latest-result aggregation consumes 36.23B model tokens, Shared Web successes 314 314 1.00× Unique Web successes 231 4,942 21.4× more corresponding to 3.61M tokens per firmware. F IRM P ILOT’s Coverage of F IRM P ILOT Web set 5.97% 100% 16.8× main-run logs consume 12.4 thousand tokens per firmware. Claude Code therefore uses about 291 times more model successes when replaced by deterministic defaults. SearchAgent, context while recovering far fewer services. NetworkAgent, and FileAgent each account for more than Failure pattern. The baseline often spends its budget 1,500 lost Web successes, covering long-tail evidence retrieval, exploring a local workspace without converging on the coupled log-grounded interface and address exposure, and boot/init device environment assumptions that govern firmware startup. recovery. The different loss patterns show that F IRM P ILOT’s In contrast, F IRM P ILOT exposes these assumptions as explicit gain comes from coordinated agent roles validated through recovery surfaces covering filesystem/init selection, persistent execution feedback, rather than from a single permissive rule key–value state, and network exposure. The evidence is or isolated artifact edit. cumulative. RQ1 separates F IRM P ILOT from template-driven C. RQ3 General-Purpose Coding Agent Baseline deterministic rehosting, RQ2 identifies which specialized Baseline protocol. RQ3 tests whether general coding-agent decisions carry the gain after the shared substrate is held fixed, capability transfers to firmware rehosting without the domain- and RQ3 shows that a strong general coding agent does not specific recovery structure used by F IRM P ILOT. We evaluate recover the same service surface when it lacks the domain Claude Code as a single general-purpose coding-agent baseline action space and feedback loop. Scalable firmware rehosting on the same 10,033-image LFwC set and the same 2,400- therefore benefits from a domain-specific control plane that second per-image wall-clock budget as F IRM P ILOT. For each decides when model reasoning may enter the loop, what form image, the baseline receives the same firmware package, its output may take, and how execution results accept or reject extracted workspace, low-level rehosting utilities, run-script that output. context, serial-log evidence, and probe-output format available to the recovery workflow, but it is isolated from F IRM P ILOT’s D. RQ4 Support for Downstream Security Workflows Workflow scope. RQ4 evaluates whether web-reachable multi-agent orchestrator and typed action interfaces. The comparison uses the same Ping and Web-service predicates rehosts preserve enough service behavior for standard downdefined above, the same guest-IP and HTTP(S) probe logic, and stream security workflows. We analyze all 5,256 LFwC one latest-result record per firmware image for reachability and web-success images, covering 622 brand-device families resource accounting. Under this controlled protocol, Claude after collapsing firmware-version variants by brand and Code reaches 729 images by Ping and 545 images by Web, device_name. The audit uses service discovery, RouterSploit corresponding to 7.27% and 5.43%. In contrast, F IRM P ILOT known-check interaction, and protocol-aware fuzzing, crediting only per-firmware artifacts written by the workflow. reaches 7,217 images by Ping and 5,256 images by Web. Success-set overlap. The Web-success partition in Table IV Audit coverage. Figure 10 summarizes the downstream further clarifies the difference between generic exploration and artifact scale. Each audited image contributes service-discovery domain-specific environment recovery. Only 314 web successes records, web-fingerprint evidence, RouterSploit transcripts, are shared between Claude Code and F IRM P ILOT. Claude and protocol-input artifacts. These records bind the recovered 0
1000
2000
3000
4000
5000
TABLE V LF W C S ERVICE D ISCOVERY AND W EB -F INGERPRINT C OVERAGE Metric Web-success images analyzed Nmap-valid images Web-fingerprint evidence Detected service records Distinct service labels HTTP/HTTPS service images Telnet service images DNS/domain images UPnP service images
FirmAE
Full audit
Gain
2,557 2,557 2,557 8,654 46 2,332 561 911 226
5,256 5,256 5,256 16,290 57 4,865 1,542 1,607 358
+2,699 +2,699 +2,699 +7,636 +11 +2,533 +981 +696 +132
TABLE VII P ROTOCOL -AWARE F UZZING S IGNAL C LASSES Signal class CI crash candidates BOF crash candidates Any crash candidate HTTP/CGI handlers Control/config daemons UPnP/SOAP/SSDP parsers Runtime command markers
Images 1,331 1,522 1,565 554 280 64 31
Hits Representative evidence 4,012 4,591 8,603 1,263 1,708 134 69
actions/cookies; SOAP/HNAP; SSDP 256–12,000B buffers, cyclic markers deduplicated CI/BOF evidence httpd, mini_httpd, uhttpd, CGI rc, scfgmgr, procd, heartbeat upnp, miniupnpd, SOAP/HNAP fields command-dispatch payload markers
sites from recovered CGI parameters, forms, HNAP or SOAP fields, UPnP and SSDP identifiers, and cookie-bearing entries. TABLE VI CI mode exercises command-oriented inputs, while BOF mode ROUTER S PLOIT 1-DAY VALIDATION R ESULTS stresses request and protocol parsers with cyclic or fixed-size Validation layer Count Images Families Evidence buffers. Crash triage uses runtime log tails and deduplicates RouterSploit audit set 63,561 5,256 586 33 public 1-day modules signals by image, mode, phase, crash kind, marker, and excerpt. Netgear CVE-2017-5521 4,066 989 819 41 positives / 23 images D-Link CVE-2015-2051 1,715 646 451 HNAP RCE path Runtime feedback. Table VII reports 4,012 CI crash hits, Misfortune Cookie CVE-2014-9222 1,138 415 286 cookie path 4,591 BOF hits, and 1,565 images with at least one imageStrict known-check positives 65 35 35 3 known-check classes level crash candidate. The 1,288-image CI/BOF overlap shows Default-credential positives 343 302 275 FTP/SSH/Telnet auth All strict positives 408 336 309 module-confirmed outputs that recovered services expose input paths with both semantic reachability and memory-stress feedback, giving downstream endpoint to the tool interaction that exercised it, so RQ4 testing tools concrete endpoint, protocol, and process context. measures executable workflow evidence rather than raw openRQ4 takeaway. RQ4 shows that F IRM P ILOT turns reachaport counts. bility gains into downstream workflow support. The recovered Service discovery. Table V shows that the recovered envi- environments expose broader service surfaces than FirmAE, ronments expose a substantially broader service surface than sustain RouterSploit known-check interactions with manually the deterministic baseline. Across the same LFwC denominator, validated findings, and support protocol-aware input delivery F IRM P ILOT increases the average number of detected services with runtime feedback. These results show that evidence-guided per firmware from 0.86 to 1.62. The valid service records environment recovery improves not only Web reachability but include protocol labels, endpoints, and available banners or also the practical utility of rehosted firmware executions. fingerprints, giving downstream tools concrete targets rather than only reachability labels. E. Threats to Validity and Mitigations Known-check interaction. RouterSploit turns the recovered Dataset scope. LFwC provides metadata for 10,913 firmware service surface into executable known-check workflows. We run records; the executable set contains the 10,033 records whose RouterSploit over the full 5,256-image RQ4 audit set and obtain packages remained downloadable from recorded vendor or output and module records for every recovered web-success archival locations, and all main LFwC comparisons use this image. Across this full set, 63,561 module runs span 586 denominator. brand-device families and 33 public router modules covering Workflow fidelity. RQ4 counts service discovery, fuzzing, command execution, information disclosure, path traversal, and RouterSploit evidence only from per-image workflow authentication bypass, default credentials, and related checks. records, making the evaluation stricter than raw reachability. Because these modules require service dialogue, request paths, Baseline scope. FirmAE is the direct deterministic baseline authentication state, or protocol exchanges, positive transcripts because it supports the same full-image QEMU execution validate behavior beyond service discovery. and Ping/Web predicates. Claude Code evaluates the separate Positive transcripts. The CVE-linked checks cover Net- question of whether a general-purpose coding agent can gear password disclosure (CVE-2017-5521), D-Link HNAP substitute for the domain-specific action space under the same RCE (CVE-2015-2051), and Misfortune Cookie (CVE-2014- denominator, timeout, evidence bundle, and success predicates. 9222). The strict subset contains 408 module-confirmed LLM variability. F IRM P ILOT reports outcomes through positives across 336 images and 309 families, including 65 fixed probes and per-image execution records, while constrainknown-check positives from exploit-oriented modules and 343 ing model outputs to typed actions and deterministic artifact FTP/SSH/Telnet default-credential authentications across 302 renderers. Replications should keep probes and denominators images. Manual validation confirms 193 reproducible Router- fixed across models. Sploit findings, including 30 known-vulnerability findings and VI. R ELATED W ORK 163 authentication-exposure findings. These results show that the recovered service surface supports stateful validation of Scalable Firmware Rehosting. Prior systems improve known-vulnerability paths and authentication exposures. different parts of the firmware-analysis stack. FirmadyneProtocol-aware testing. The recovered service surfaces also style systems construct disk images, infer architecture and support protocol-aware input delivery. The fuzzer derives input kernels, synthesize run scripts, and derive network parameters
through deterministic templates [1]. FirmAE strengthens this full-image lineage with arbitrated emulation, NVRAM defaults, and network heuristics, making it the closest template-driven deterministic baseline for the Ping/Web reachability protocol used in this work [2]. FirmGuide and PANDaWAN focus on diagnosing or guiding rehosting failures [3], [4], while Greenhouse, FIRMWELL, and user-space dependency-aware rehosting reduce execution complexity by isolating services or user-space dependency environments [6], [7]. F IRM P ILOT targets a different point in this space because it keeps the full-system execution setting but replaces fixed environment templates with evidence-grounded recovery over boot, persistent state, and exposure artifacts. The distinction is important for scale. Identifying a missing init dependency, state key, or service condition is not enough unless the system can translate that evidence into a repeatable execution environment. F IRM P ILOT makes that path part of the system. Evidence is converted into typed deltas, deltas are applied through deterministic renderers, and the resulting firmware execution is re-probed under the same predicates. This design keeps the evaluation aligned with executable environment recovery rather than isolated failure explanation. Emulation Fidelity. Hybrid execution combines emulation with hardware for complex peripherals [8], [16], while modeling and abstraction improve portability and fuzzing effectiveness [9], [10], [26]. Invalidity-guided inference uses execution failures to discover missing assumptions [27]. These works enrich the execution substrate; F IRM P ILOT instead coordinates cross-layer service recovery on scalable whole-system rehosts. Firmware Security Analysis. Embedded security analyses require reachable services and stable execution for validation and triage. Prior work covers authentication-bypass detection [28], application-driven IoT fuzzing [29], keywordbased bug discovery [30], greybox fuzzing [31], firmware fuzzing [11], [32], driver/peripheral analysis [33], [34], and tracing/replay [35], [36]. F IRM P ILOT complements these tools by making more firmware executions reachable and analyzable. LLM Agents for Automated Analysis. Agentic LLM systems support penetration testing [21], fuzzing guidance [22], program/binary analysis [23], [24], multi-agent coordination [20], [37], and program-transformation pipelines with retrieval, planning, patch generation, and validation [38]–[40]. F IRM P ILOT adapts this paradigm to firmware rehosting by binding agent decisions to typed environment updates, where retrieval supplies evidence, planning orders recovery actions, specialized agents update state and exposure, and reruns with fixed probes validate the resulting artifacts. Firmware rehosting differs from many code-repair or penetration-testing settings because the target behavior is not specified by a test suite or a live remote service. The desired environment must be reconstructed from partial device assumptions, including init scripts, persistent-state reads, networkinterface conventions, and probe responses. F IRM P ILOT therefore uses LLM agents as bounded environment-recovery components rather than open-ended operators. The Claude Code baseline in RQ3 shows why this distinction matters because
broad coding ability does not replace domain-specific evidence routing, action schemas, and execution-grounded acceptance. VII. D ISCUSSION Rationale for a Multi-Agent Architecture. Rehosting failures are cross-layer and sequential, so progress requires coordinated evidence interpretation and constrained actuation. Role specialization separates what to infer from what to change by assigning Search and planning to symbolic evidence and File, NVRAM, and Network agents to concrete artifact surfaces. The ablations show that removing one role disables a specific recovery channel rather than a monolithic prompt. Security-Workflow Fidelity. Web reachability is only the entry criterion. RQ4 requires recovered services to sustain enumeration, manual validation of 1-day and credential-path findings, protocol-input delivery, and crash triage, turning reachability into an execution substrate for security analysis. Verifiability. The unit of evidence is a per-image execution trace rather than a model explanation. Each trace links observations, retrieved support, typed deltas, rendered artifacts, probes, final labels, and downstream tool outputs. This structure makes the reported counts auditable from the same records that drive the recovery loop and keeps failures inspectable at the transition where execution stopped. Limitations and Future Work. F IRM P ILOT does not fully address peripheral and device-specific I/O modeling, and customized stacks can still fail because of late-stage dependencies such as missing libraries, certificates, or vendor wrappers. Future work can combine environment recovery with hardware-in-the-loop execution [8], [16] and MMIO modeling [11]. Security and Ethical Considerations. The evaluation runs offline in isolated rehosting environments and does not scan live devices; generated deltas target analysis artifacts, and realdevice follow-up should follow responsible disclosure. VIII. C ONCLUSION F IRM P ILOT addresses firmware-rehosting brittleness through an evidence-guided multi-agent environment-recovery loop over boot, state, and network layers. Specialized agents interpret runtime evidence and apply typed, replayable artifact deltas. Across LFwC and the public FirmAE benchmark, F IRM P ILOT improves service reachability and supports service discovery, manual validation of RouterSploit findings, and protocol-aware fuzzing. The results show that agentic reasoning is effective for firmware rehosting when constrained by evidence, explicit transition interfaces, and repeated execution. R EFERENCES [1] D. D. Chen, M. Woo, D. Brumley, and M. Egele, “Towards automated dynamic analysis for linux-based embedded firmware.” in NDSS, 2016. [2] M. Kim, D. Kim, E. Kim, S. Kim, Y. Jang, and Y. Kim, “Firmae: Towards large-scale emulation of iot firmware for dynamic analysis,” in Proceedings of the 36th Annual Computer Security Applications Conference, 2020, pp. 733–745.
[3] Q. Liu, C. Zhang, L. Ma, M. Jiang, Y. Zhou, L. Wu, W. Shen, X. Luo, Y. Liu, and K. Ren, “Firmguide: Boosting the capability of rehosting embedded linux kernels through model-guided kernel execution,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2021, pp. 792–804. [4] I. Angelakopoulos, G. Stringhini, and M. Egele, “Pandawan: quantifying progress in linux-based firmware rehosting,” in 33rd USENIX Security Symposium (USENIX Security 24), 2024, pp. 5859–5876. [5] E. Johnson, M. Bland, Y. Zhu, J. Mason, S. Checkoway, S. Savage, and K. Levchenko, “Jetset: Targeted firmware rehosting for embedded systems,” in 30th USENIX Security Symposium, 2021, pp. 321–338. [6] H. J. Tay, K. Zeng, J. M. Vadayath, A. S. Raj, A. Dutcher, T. Reddy, W. Gibbs, Z. L. Basque, F. Dong, Z. Smith et al., “Greenhouse:{SingleService} rehosting of {Linux-Based} firmware binaries in {User-Space} emulation,” in 32nd USENIX Security Symposium, 2023, pp. 5791–5808. [7] C. Qin, C. Zhang, Y. Zheng, P. Liu, J. Zhang, Y. Li, W. Zhang, Y. Liu, and L. Sun, “User-space dependency-aware rehosting for linux-based firmware binaries,” in Proceedings of the Network and Distributed System Security Symposium, 2026. [8] M. Muench, D. Nisi, A. Francillon, and D. Balzarotti, “Avatar2 : A multitarget orchestration platform,” in Proceedings 2018 Workshop on Binary Analysis Research, 2018. [9] A. A. Clements, E. Gustafson, T. Scharnowski, P. Grosen, D. Fritz, C. Kruegel, G. Vigna, S. Bagchi, and M. Payer, “{HALucinator}: Firmware re-hosting through abstraction layer emulation,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 1201– 1218. [10] B. Feng, A. Mera, and L. Lu, “{P2IM}: Scalable and hardwareindependent firmware testing via automatic peripheral interface modeling,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 1237–1254. [11] T. Scharnowski, N. Bars, M. Schloegel, E. Gustafson, M. Muench, G. Vigna, C. Kruegel, T. Holz, and A. Abbasi, “Fuzzware: Using precise {MMIO} modeling for effective firmware fuzzing,” in 31st USENIX Security Symposium, 2022, pp. 1239–1256. [12] F. Bellard, “QEMU, a fast and portable dynamic translator,” in USENIX Annual Technical Conference, FREENIX Track, 2005, pp. 41–46. [13] A. Fasano, T. Ballo, M. Muench, T. Leek, A. Bulekov, B. Dolan-Gavitt, M. Egele, A. Francillon, L. Lu, N. Gregory et al., “Sok: Enabling security analyses of embedded systems via rehosting,” in Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, 2021, pp. 687–701. [14] E. Gustafson, M. Muench, C. Spensky, N. Redini, A. Machiry, Y. Fratantonio, D. Balzarotti, A. Francillon, Y. R. Choe, C. Kruegel et al., “Toward the analysis of embedded firmware through automated re-hosting,” in 22nd International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2019), 2019, pp. 135–150. [15] N. Redini, A. Machiry, R. Wang, C. Spensky, A. Continella, Y. Shoshitaishvili, C. Kruegel, and G. Vigna, “Karonte: Detecting insecure multibinary interactions in embedded firmware,” in 2020 IEEE Symposium on Security and Privacy (SP), 2020, pp. 1544–1561. [16] J. Zaddach, L. Bruno, A. Francillon, D. Balzarotti et al., “Avatar: A framework to support dynamic security analysis of embedded systems’ firmwares.” in NDSS, 2014, pp. 1–16. [17] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al., “Retrievalaugmented generation for knowledge-intensive nlp tasks,” Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020. [18] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022. [19] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems, vol. 36, pp. 8634–8652, 2023. [20] G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “CAMEL: Communicative agents for “mind” exploration of large language model society,” Advances in Neural Information Processing Systems, vol. 36, pp. 51 991–52 008, 2023. [21] G. Deng, Y. Liu, V. Mayoral-Vilches, P. Liu, Y. Li, Y. Xu, T. Zhang, Y. Liu, M. Pinzger, and S. Rass, “{PentestGPT}: Evaluating and harnessing large language models for automated penetration testing,” in 33rd USENIX Security Symposium, 2024, pp. 847–864.
[22] R. Meng, M. Mirchev, M. Böhme, and A. Roychoudhury, “Large language model guided protocol fuzzing,” in Proceedings 2024 Network and Distributed System Security Symposium, 2024. [23] H. Li, Y. Hao, Y. Zhai, and Z. Qian, “The hitchhiker’s guide to program analysis: A journey with large language models,” arXiv preprint arXiv:2308.00245, 2023. [24] P. Liu, C. Sun, Y. Zheng, X. Feng, C. Qin, Y. Wang, Z. Li, and L. Sun, “Harnessing the power of llm to support binary taint analysis,” arXiv preprint arXiv:2310.08275, 2023. [25] R. Helmke, E. Padilla, and N. Aschenbruck, “Mens sana in corpore sano: Sound firmware corpora for vulnerability research,” in Proceedings of the 2025 Network and Distributed System Security Symposium (NDSS), 2025. [26] C. Cao, L. Guan, J. Ming, and P. Liu, “Device-agnostic firmware execution is possible: A concolic execution approach for peripheral emulation,” in Proceedings of the 36th Annual Computer Security Applications Conference, 2020, pp. 746–759. [27] W. Zhou, L. Guan, P. Liu, and Y. Zhang, “Automatic firmware emulation through invalidity-guided knowledge inference,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 2007–2024. [28] Y. Shoshitaishvili, R. Wang, C. Hauser, C. Kruegel, and G. Vigna, “Firmalice-automatic detection of authentication bypass vulnerabilities in binary firmware.” in NDSS, 2015. [29] J. Chen, W. Diao, Q. Zhao, C. Zuo, Z. Lin, X. Wang, W. C. Lau, M. Sun, R. Yang, and K. Zhang, “Iotfuzzer: Discovering memory corruptions in iot through app-based fuzzing.” in NDSS, 2018, pp. 1–15. [30] L. Chen, Y. Wang, Q. Cai, Y. Zhan, H. Hu, J. Linghu, Q. Hou, C. Zhang, H. Duan, and Z. Xue, “Sharing more and checking less: Leveraging common input keywords to detect bugs in embedded systems,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 303–319. [31] A. Fioraldi, D. Maier, H. Eißfeldt, and M. Heuse, “AFL++: Combining incremental steps of fuzzing research,” in 14th USENIX Workshop on Offensive Technologies (WOOT 20), 2020. [32] N. Bars, M. Schloegel, T. Scharnowski, N. Schiller, and T. Holz, “Fuzztruction: Using fault injection-based fuzzing to leverage implicit domain knowledge,” in 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 1847–1864. [33] Z. Shen, R. Roongta, and B. Dolan-Gavitt, “Drifuzz: Harvesting bugs in device drivers from golden seeds,” in 31st USENIX Security Symposium (USENIX Security 22), 2022, pp. 1275–1290. [34] S. M. S. Talebi, H. Tavakoli, H. Zhang, Z. Zhang, A. A. Sani, and Z. Qian, “Charm: Facilitating dynamic analysis of device drivers of mobile systems,” in 27th USENIX Security Symposium (USENIX Security 18), 2018, pp. 291–307. [35] L. Craig, A. Fasano, T. Ballo, T. Leek, B. Dolan-Gavitt, and W. Robertson, “PyPANDA: Taming the PANDAmonium of whole system dynamic analysis,” in Proceedings 2021 Workshop on Binary Analysis Research, 2021. [36] Y. Shoshitaishvili, R. Wang, C. Salls, N. Stephens, M. Polino, A. Dutcher, J. Grosen, S. Feng, C. Hauser, C. Kruegel et al., “Sok:(state of) the art of war: Offensive techniques in binary analysis,” in 2016 IEEE Symposium on Security and Privacy (SP), 2016, pp. 138–157. [37] T. Abramovich, M. Udeshi, M. Shao, K. Lieret, H. Xi, K. Milner, S. Jancheska, J. Yang, C. E. Jimenez, F. Khorrami et al., “EnIGMA: Enhanced interactive generative model agent for ctf challenges,” arXiv preprint arXiv:2409.16165, 2024. [38] M. Seo, W. Choi, M. You, and S. Shin, “Autopatch: Multi-agent framework for patching real-world cve vulnerabilities,” arXiv preprint arXiv:2505.04195, 2025. [39] H. Li, Y. Tang, S. Wang, and W. Guo, “Patchpilot: A cost-efficient software engineering agent with early attempts on formal verification,” arXiv preprint arXiv:2502.02747, 2025. [40] Z. Yu, Z. Guo, Y. Wu, J. Yu, M. Xu, D. Mu, Y. Chen, and X. Xing, “PATCHAGENT: A practical program repair agent mimicking human expertise,” in 34th USENIX Security Symposium (USENIX Security 25), 2025, pp. 4381–4400.