Conceptio › Archive › arXiv CS
arXiv CSopen access

POZZER: A Power Side Channel-guided Fuzzer for Black-Box Embedded Systems

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

P OZZER: A Power Side Channel-guided Fuzzer for Black-Box Embedded Systems Pouya Narimani, Kseniia Rogova, Addison Crump, Martin Mohl, Meng Wang, Ulysse Planta, Pansilu Pitigalaarachchi and Ali Abbasi

arXiv:2609.23583v1 [cs.CR] 20 Sep 2026

CISPA Helmholtz Center for Information Security

Abstract

is an effective technique for discovering software and hardware vulnerabilities. Coverage-guided fuzzing, in particular, is effective because it continuously uses execution feedback to guide input generation toward previously unexplored execution paths. Existing firmware fuzzers obtain such feedback through instrumentation [41], hardware debug interfaces [23], or rehosting [52]. However, these approaches require some degree of access to the target firmware or its execution environment, which are often unavailable in practice. Modern embedded systems commonly integrate third-party modules supplied by different vendors. Many of these modules, such as wireless communication chips [24], global navigation satellite system (GNSS) receivers [58], secure elements [56], and sensor controllers [50], contain their own microcontroller (MCU) executing proprietary firmware. The firmware of these modules is often inaccessible to system integrators and end users, and debugging interfaces are often disabled to prevent firmware extraction [21, 63]. Consequently, many real-world embedded systems can only be tested through their exposed communication interfaces, entrenching a strict black-box setting in which the attacker has no access to the firmware source code or binary image. Under such conditions, existing coverage-guided fuzzers cannot obtain the execution feedback required to guide a fuzzer and therefore fall back to a blind fuzzer without execution feedback, which explores the input space significantly less efficiently. Obtaining execution feedback, therefore, becomes the fundamental challenge for black-box embedded firmware fuzzing. Physical side-channels offer an alternative source of execution feedback, without requiring firmware instrumentation, access to the firmware binary or source code [32], making physical side-channels well-suited for strict black-box settings. In particular, power side channels (PSCs) leak information about executed instructions [31] and processed data [30], making them a promising feedback source for black-box embedded fuzzing. However, transforming raw PSC traces into practical fuzzing feedback is challenging as measurement noise, temporal alignment, and execution-dependent signal variations

Firmware fuzzing is an effective technique for discovering vulnerabilities in embedded systems. However, existing coverage-guided firmware fuzzers typically obtain feedback through firmware instrumentation, hardware debug interfaces, or firmware rehosting, which requires access to the firmware source code or binary image. Such requirements are often infeasible for off-the-shelf embedded devices, where firmware binaries are inaccessible, unrehostable, immodifiable, or undebuggable, necessitating fuzzing under black-box conditions. In this paper, we present P OZZER, a power side-channelguided fuzzer for black-box embedded systems. P OZZER uses power traces as feedback to identify previously unseen behavior via an incrementally constructed graph-based representation of observed executions, guiding the fuzzer toward unexplored execution paths. Its non-profiling design requires neither prior firmware knowledge nor a clone device, extracting meaningful feedback from a single power trace per execution while remaining robust to measurement noise. We evaluate P OZZER on 15 firmware targets across two platforms and two real-world commercial embedded devices. Across the resulting target-platform combinations, P OZZER outperforms a blind fuzzer under the same time budget in 26 out of 30 targetplatform combinations. Furthermore, P OZZER discovers two previously unknown vulnerabilities in one of the commercial devices, both confirmed by the vendor, demonstrating its potential for identifying vulnerabilities in black-box embedded systems.

1

Introduction

Embedded systems have become an integral part of modern society, powering applications ranging from consumer internet of things (IoT) devices [22] to industrial automation [61], automotive systems [49], and medical devices [33]. As these systems process increasingly security-sensitive data and control safety-critical operations, vulnerabilities in their firmware can have severe security and safety consequences. Fuzzing 1

make it difficult to distinguish execution-related discrepancies from measurement artifacts. Therefore, practical PSCguided fuzzing requires a fast, online approach that extracts meaningful execution-related feedback from a single noisy power trace under black-box conditions. Although prior sidechannel-guided fuzzers explore similar ideas [16, 39, 54, 60], they rely on the availability of ground-truth data to train models [54], repeating executions [60], or unrealistic threat models [54]. In this paper, we first empirically analyze PSC traces from diverse embedded firmware targets to determine what execution-related information can be reliably extracted from raw PSC traces in a black-box setting, and identify the challenges posed by noisy environments and high-throughput fuzzing. We analyze the relationship between firmware execution and the corresponding PSC traces, characterizing execution-related features that consistently appear in the traces as well as artifacts introduced by the underlying hardware and firmware implementation. Furthermore, our analysis identifies several sources of variation in PSC traces, including redundant code segments, localized measurement noise, and temporal shifts, which can obscure the distinction between previously observed and genuinely new execution behavior and thereby affect the stability of a PSC-guided fuzzer. Guided by these observations, we design P OZZER, a PSCguided fuzzer for black-box settings. P OZZER addresses the challenge of applying the general idea of graph-based feedback to PSC traces collected from embedded systems to guide a black-box fuzzer. The key intuition behind P OZZER is that discrepancies between PSC traces indicate potential discrepancies in firmware control flow. Building upon this intuition, P OZZER incrementally constructs a pseudo execution-flow graph by identifying discrepancy points among the observed traces. The constructed graph serves as a global reference for determining whether newly captured traces exhibit previously unseen behavior and deserve more attention. To address the challenges of PSC-based feedback, P OZZER combines several PSC-specific techniques for robust PSC trace comparison, hardware noise handling, and efficient executionflow graph construction. The empirical findings from our analysis directly inform the design of P OZZER’s hardware and software-level noise-handling techniques, enabling reliable single power trace feedback without requiring repeated measurements, profiling, prior knowledge of the firmware, or firmware instrumentation. To evaluate P OZZER, we conduct two sets of experiments. First, we fuzz 15 firmware targets spanning two different hardware platforms. Although firmware source code is available for these targets, P OZZER operates on uninstrumented firmware, and it does not receive direct execution feedback during fuzzing; its guidance is derived solely from PSC traces. We compare P OZZER against a blind fuzzer as the baseline and an instrumented coverage-guided fuzzer as an upperbound fuzzer with direct execution feedback. Under the same

time budget, P OZZER outperforms the blind fuzzer in 26 out of 30 targets and, in some cases, achieves performance comparable to that of the instrumented coverage-guided fuzzer. Second, we conduct case studies on two real-world commercial off-the-shelf (COTS) devices under a black-box setting, where firmware instrumentation and binary extraction are unavailable. P OZZER identifies two previously unknown vulnerabilities in one of the commercial devices, both of which are confirmed by the vendors, whereas the blind fuzzer fails to discover them, demonstrating P OZZER’s applicability to real-world black-box embedded systems. In summary, this paper makes the following contributions: • We present P OZZER, a PSC-guided fuzzer for black-box embedded systems that outperforms a blind fuzzer without requiring firmware instrumentation, firmware binary extraction, or profiling. • We conduct an empirical analysis of PSC traces in the context of black-box embedded fuzzing, identifying hardware and software-level characteristics of embedded systems that affect the reliable extraction of executionrelated feedback from PSC traces. • We develop and integrate a set of PSC-specific techniques, including signature-based search, chunk-based comparison, first derivative comparison, and time-shiftaware neighbor search into P OZZER’s graph-based feedback pipeline, enabling meaningful feedback extraction from a single noisy PSC trace per execution. • We demonstrate both the effectiveness and practical applicability of P OZZER: across 15 firmware targets spanning two hardware platforms, P OZZER improves average coverage by 12.06% over a blind fuzzer under the same time budget, while case studies on two real-world COTS devices show that it can identify vulnerabilities under a black-box threat model without firmware instrumentation or binary access.

2

Background

In this section, we first define P OZZER’s threat model, detailing the information available to P OZZER about the target and the constraints imposed by it. We then introduce the concepts and terminology underlying coverage-guided fuzzing and physical side-channels, providing the background necessary to understand the design of P OZZER.

2.1

Threat model

P OZZER leverages PSC leakage to fuzz embedded systems, operating at the intersection of hardware and software. Therefore, we define the threat model from both the hardware and software perspectives. 2

Software. At the software level, we assume a strict blackbox threat model. P OZZER has no access to the target firmware’s source code or binary image. The attacker can interact with the target only through its exposed communication interfaces and observe its responses, if any. However, the attacker cannot obtain code-coverage feedback through source code or binary instrumentation. Thus, conventional instrumentation-based guidance is unavailable, making the setting equivalent to black-box blind fuzzing. Hardware. At the hardware level, we assume that the target’s debug interfaces are inaccessible and no instruction traces or debugging information can be obtained from the device. However, the attacker has physical access to the device and can passively monitor PSC traces by installing a shunt resistor on the target’s power supply. We further assume that the attacker knows the exposed interface pins and can use them to send test inputs to the target during fuzzing.

2.2

possesses a set of side-channel measurements with their corresponding inputs and/or outputs [48], and performs the attack directly on these measurements [19]. Accordingly, P OZZER requires no profiling stage, training data, or leakage models from similar devices. Instead, it derives feedback solely from PSC traces collected from the target during fuzzing, consistent with our black-box threat model.

2.4

Coverage-guided fuzzing has been proven to be a powerful tool to discover vulnerabilities in software and hardware systems. It relies on code-coverage feedback to identify inputs that reach previously unseen basic blocks, add them to a corpus, and iteratively mutate them to explore deeper program states. While highly effective, traditional coverage-guided fuzzers rely heavily on compile-time instrumentation or dynamic binary translation (e.g., QEMU [17]) to retrieve this coverage feedback, both of which are intrusive and incompatible with our black-box threat model. In this work, we replace traditional software-based coverage feedback with a custom PSC-based feedback mechanism. This enables P OZZER to perform coverage-guided fuzzing using only PSC traces as passive side-channel leakage, remaining completely oblivious to the target’s internal software constraints.

Power Side-Channels

PSCs were initially exploited to recover cryptographic keys from smart cards by leveraging data-dependent variations in power consumption [35]. Such attacks exploit implementation-induced information leakage rather than weaknesses in the underlying cryptographic algorithms. Subsequently, PSCs have also been used to identify executed instructions, as different instructions exhibit distinguishable PSC patterns [43, 47]. Crucially, these distinct patterns mean that different execution paths and control-flow transitions induce measurable differences in a device’s PSC traces. This makes PSCs a highly promising modality for passively inferring program state. However, using PSCs to provide online execution feedback differs fundamentally from traditional side-channel analysis. Unlike cryptographic attacks, which typically target short execution windows and tolerate extensive offline processing, fuzzing requires extracting meaningful execution information from program execution in real time. This introduces new challenges, including handling measurement noise [44], clock synchronization [27], and efficient processing of long time-series traces [20].

2.3

Coverage-guided fuzzing

3

Empirical analysis of PSC traces as feedback

This section presents an empirical analysis of PSC traces that motivates P OZZER’s design. As discussed in Section 2, PSC traces leak information about executed instructions and processed data [37, 43]. However, their suitability as an online feedback source for fuzzing remains largely unexplored. Unlike traditional PSC analysis, fuzzing requires extracting reliable execution-related information from a single PSC trace, without prior knowledge of the target firmware or repeated measurements. Therefore, we investigate which executionrelated characteristics are reliably observable in PSC traces and which software- and hardware-level effects affect their interpretation. Specifically, we ask: • Which firmware behaviors introduce variability or other artifacts into PSC traces, and how do these effects influence the interpretation of execution-related information?

Profiling vs. Non-Profiling Attacks

Side-channel attacks are commonly classified as profiling or non-profiling, depending on the attacker’s capabilities. Profiling attacks assume access to a clone device that is sufficiently similar to the target and controlled by the attacker. This enables the collection of ground-truth execution data to build a leakage model [55], which is then applied to the target device. However, profiling a clone device is not feasible in black-box scenarios because it needs access to the source code or binary image of the firmware to obtain ground-truth data. In contrast, a non-profiling attack assumes that the attacker has no access to a clone device. Instead, the attacker only

• Which hardware and measurement factors introduce temporal or voltage variations in PSC measurements, and how do these factors affect the reliability of PSC trace comparison? To answer these questions, we design our empirical analysis as follows. We select and execute ten firmware programs on two hardware platforms, ST STM32F3 and Microchip SAM4S. For each firmware, we manually construct ten test cases that exercise different execution paths. We execute each test case 3

Redundant Power Trace Segments

Power Traces with DC offset 0.3

Trace 1

Trace 2

Trace 3

4240

4245

4250

Trace 4

Trace 5

0.25

0.2

Amplitude (V)

−0.25

Amplitude (V)

0.00 Trace 1

0.25 0.00 −0.25

Trace 2

0.1 0.0 −0.1 −0.2

0.25

−0.3 4230

0.00 −0.25

4235

4255

Time (Sample index)

4260

4265

Trace 3

0

20

40

60

Time (Sample index)

80

100

Figure 2: Five PSC traces with identical code executions, with one execution exhibiting a noticeable local amplitude shift.

Figure 1: Three different power traces, all having the same code segment (inside the highlighted area) executed with different prior code execution paths (plotted in red).

firmware execution paths converge, causing identical code segments to appear at different temporal locations. As a result, direct time-domain comparison is insufficient for distinguishing previously observed execution behavior from genuinely new executions. Localized Measurement Noise. Repeated executions of the same test case reveal localized hardware-induced variations in the measured PSC traces. Figure 2 shows an example in which one execution exhibits a local amplitude shift despite executing the same code. These localized distortions originate from the measurement process, such as quantization errors, rather than the firmware itself. Since they affect only a portion of the PSC trace, they can obscure execution related features and make traces from identical executions appear different. Consequently, localized measurement noise can induce ambiguity when interpreting execution behavior directly from raw PSC traces. Time Shifts. Another behavior consistently observed during our measurements is the presence of small time shifts between repeated executions of the same test case. Although the executed instructions remain identical, small timing variations cause corresponding features in the PSC traces to become misaligned. As a result, traces generated from the same execution may no longer align sample by sample along the temporal axis, making it difficult to reliably compare execution behavior directly in the time domain.

1000 times while capturing the corresponding PSC traces, resulting in 200,000 traces in total. To characterize firmware-induced effects in PSC traces, we compare traces generated by different test cases that exercise different execution paths. This comparison allows us to identify execution dependent variations that can complicate feedback extraction. Motivated by observations reported in prior work [38], we examine recurring similarities and differences across the traces. Our analysis reveals that similar code segments can appear at different temporal locations across PSC traces, resulting in redundant code segments and temporal shifts that affect trace comparison. To characterize hardware and measurement-induced effects, we compare measurements for repeated executions of the same test case. Since repeated executions of the same test case are generally expected to follow the same firmware execution, differences between their traces should primarily arise from hardware and measurement effects. Comparing these repeated traces reveals localized measurement noise and temporal or voltage variations, demonstrating how these effects can affect the reliability of PSC trace comparison. The remainder of this section examines three key characteristics identified through the above empirical analysis: redundant code segments, localized measurement noise, and time shifts. These characteristics affect the reliability of PSCbased feedback extraction and motivate the design choices presented in Section 4. Redundant Code Segments. Prior work has shown that identical code segments may occur at different temporal positions in PSC traces [38], a behavior we also observe. We refer to these recurring segments as redundant code segments because they represent previously observed code execution rather than new execution behavior. Consequently, two traces can differ substantially while still containing segments corresponding to the same code execution. Figure 1 illustrates an example in which multiple execution paths converge to an identical code segment. Although the preceding execution differs, the latter parts of the traces correspond to the same code execution. This behavior naturally arises when distinct

4

PSC-based Graph Construction

In this section, we explain the design of P OZZER for using PSC traces as real-time fuzzing feedback. P OZZER must derive reliable execution-related feedback from noisy PSC measurements while operating under a strict black-box threat model. Our empirical analysis in Section 3 provides insights into several factors that affect PSC traces. These observations constitute the design of P OZZER and its approach to extracting, comparing, and interpreting execution-related feedback. In addition to these trace-specific challenges, P OZZER must satisfy two general requirements of online fuzzing. First, it must maintain high throughput. A fuzzing campaign executes many test cases, each producing a large PSC trace that must 4

Raw traces

be acquired and processed. Second, it requires a global reference against which each new observation can be compared to determine whether the corresponding test case exhibits previously unseen behavior. Therefore, P OZZER must derive execution-related feedback from PSC traces while limiting the overhead of trace acquisition and processing in addition to accounting for measurement variability and noise. These requirements guide the design of P OZZER and lead to the following three design goals: • Efficient feedback extraction: P OZZER should minimize the overhead of capturing and processing PSC traces so the fuzzer can achieve high throughput within a given time budget. • Reliable single-trace feedback: P OZZER should account for temporal and voltage variations in PSC traces without relying on repeated measurements. In particular, it should extract stable execution-related information from a single trace for each test case. • Global execution-flow modeling: P OZZER should maintain a global representation of execution flow behavior observed during fuzzing. This representation should allow each newly captured trace to be compared with previously observed behavior and support the identification of previously unseen execution-flow transitions. In addition, non-profiling operation is a constraint imposed by our threat model rather than an independent design goal. Therefore, the feedback mechanism of P OZZER must construct and update the global execution-flow representation during fuzzing using the PSC traces.

4.1

First-order derivatives

Trace 1 Trace 2

4.0

Trace 2: dy/dx

0.75

Amplitude (V)

3.5

Amplitude (V)

Trace 1: dy/dx

1.00

3.0 2.5 2.0 1.5

0.50 0.25 0.00 −0.25 −0.50 −0.75

1.0

−1.00 1.0

1.5

2.0

2.5

3.0

3.5

4.0

Time (Sample index)

4.5

(a) Raw PSC traces

5.0

1.0

1.5

2.0

2.5

3.0

3.5

4.0

Time (Sample index)

4.5

5.0

(b) First-order derivatives

Figure 3: PSC traces before and after first-order differentiation. Differentiation suppresses the local amplitude shift, allowing point-wise comparison to correctly detect similarity.

to mitigate hardware-induced noise in the next section. Similarity metric. PSC traces are floating-point time series that are heavily affected by noise. Therefore, exact equivalence is unsuitable for detecting discrepancies due to measurement variations. Hence, P OZZER requires a similarity metric to detect discrepancies between two PSC traces and consequently detect novelty. In this paper, we use L1 distance metric to compare the points in PSC traces. If the distance between two points is higher than a threshold, then the corresponding point is a discrepancy point. The L1 distance is inexpensive to compute, making it a fast and efficient metric for the repeated trace comparisons required by P OZZER’s online feedback.

4.2

Hardware Noise Handling

In this section, we explain the techniques used by P OZZER to mitigate hardware-induced noise of PSC traces. Our empirical analysis in Section 3 identified two types of hardware noise, localized amplitude shifts (Figure 2) and small time shifts. These variations can cause the discrepancy detection procedure to incorrectly identify measurement artifacts as execution differences. P OZZER therefore incorporates several essential noise-handling techniques at different stages of the feedback workflow to improve robustness while preserving execution-related information. Chunk-based Discrepancy Detection. Point-wise comparison is sensitive to localized amplitude variations because a disturbance affecting only a small portion of a trace can produce a sequence of apparent discrepancies. P OZZER therefore compares fixed-size chunks rather than individual samples. Each chunk represents a sequence of consecutive samples corresponding to a portion of the execution. Given two chunks X and Y , P OZZER computes their normalized L1 distance:

Graph-based Feedback

The concept of constructing a global reference for black-box fuzzers has been discussed in a prior paper [36]. This paper introduces the concept of Execution Divergence Graph (EDG) and uses it as a global reference for a black-box fuzzer. In this graph, nodes denote the execution traces, and edges denote transitions. For example, an edge A→B denotes that an execution trace A followed by an execution trace B was observed. Consequently, the graph stores observed prefixes of execution traces as paths from an initial node. If a program execution produces an execution trace that is not a prefix of any existing path, the graph is updated by splitting nodes or adding new edges to represent the newly observed program behavior. Although the EDG provides the foundations for graph construction in a noise-free environment, using it on PSC traces requires significant design improvements and the introduction of new techniques to address the challenges that arise when using PSC traces. These challenges include introducing a new similarity metric for PSC traces, designing techniques to mitigate hardware noise as discussed in Section 3, and controlling the graph size. Following, we first introduce the similarity metric that we use to detect discrepancies between PSC traces, then we explain the techniques used by P OZZER

DL1 (X,Y ) =

1 |CS|−1 ∑ |xi − yi |, |CS| i=0

where |CS| is the number of samples in a chunk. A chunk is considered discrepant when its distance exceeds a predefined threshold. Comparing chunks makes discrepancy detection 5

Neighbour Comparison Trace 1

Amplitude (V)

4

Graph trimming

2

Fuzzing engine

✅

0

Trace insertion

❌

Target interface

API

Target

Trace 2

4 2

Signature search 0

0

1

2

3

Time (Sample index)

4

Observer

Capture device

5

Feedback POZZER

Figure 4: Robustness of neighbor comparison to time shifts. Each point is compared with neighboring points within a radius of two samples, enabling correct matching despite small time shifts.

Figure 5: An overview of P OZZER’s architecture

ule, including graph construction, trace insertion, and graph trimming in a fuzzing loop. less sensitive to localized variations and enables the graph construction algorithm to operate at a lower temporal resolution, as it compares trace segments corresponding to instruction sequences rather than individual instruction profiles. Recall that Figure 2 already features a localized amplitude shift in PSC traces. A point-wise comparison would incorrectly identify this region as a discrepancy region and, consequently, as an interesting execution behavior. In contrast, P OZZER’s chunkbased discrepancy detection is considerably more robust to such disturbances. First-Order Derivatives. After identifying the discrepant chunk, P OZZER performs a point-wise comparison within that chunk to determine the exact discrepancy point. At this stage, point-wise comparison is more sensitive to local amplitude shifts because it operates at the sample level. Therefore, P OZZER performs this comparison in the first-order derivative domain rather than directly on the raw traces. Differentiation suppresses constant and slowly varying amplitude offsets while preserving changes in the trace signal. Figure 3 illustrates this effect, and Appendix A provides a mathematical analysis of the approach. Time Shifts. Small time shifts can also cause false discrepancy points because corresponding features may be displaced by one or more samples between traces. Since point-wise comparison ultimately determines the precise discrepancy location, such misalignment can introduce spurious graph nodes that do not represent new execution behavior. P OZZER therefore extends point-wise comparison with a neighbor search. For each sample in one trace, it compares the sample with the aligned sample and neighboring samples within a predefined window in the other trace. This allows corresponding features to match despite small temporal misalignments and reduces false discrepancy point detections. Figure 4 illustrates the approach using a neighborhood of two samples on either side.

5

5.1

Overview

In this subsection, we provide an overview of P OZZER and its online feedback workflow. Figure 5 shows the main components of P OZZER and their interactions. At each iteration, the fuzzing engine selects and mutates an input from the corpus to generate a new test case. The target interface sends the test case to the target through the target API, while the capture device collects the resulting PSC trace. The observer receives the trace and passes it to the feedback module, which compares the new trace with the execution flow observed so far and determines whether the test case exhibits previously unseen execution behavior. The feedback module maintains a graph-based representation of the observed execution behavior. It first uses a signature search to locate the corresponding portions of the new trace in the existing graph. If the trace contains previously unseen behavior, the trace insertion procedure incorporates the new execution into the graph, after which graph trimming removes redundant portions of the graph. The feedback module then reports whether new execution behavior was observed. If so, the fuzzing engine adds the corresponding test case to the corpus and continues with the next iteration; otherwise, it proceeds without adding the test case. The following subsections describe these mechanisms in detail. Subsection 5.2 presents P OZZER’s graph construction procedure and explains how it addresses challenges arising from firmware execution behavior and the interpretation of PSC traces.

5.2

Graph Construction

As described in Section 5.1, P OZZER represents observed execution behavior as a graph and uses this representation as the global reference for online feedback. This section describes how P OZZER constructs and maintains this graph from PSC traces. The construction process consists of three stages: signature-based search, trace insertion, and graph trimming.

Implementation

In this section, we first explain the high-level implementation of P OZZER, then we explain each stage of the feedback mod6

New trace

Graph

Graph

to be identified without comparing the full new trace against every branch in the graph. Once a matching signature is found, P OZZER resumes trace comparison from the corresponding graph location. Figure 6a illustrates this process. If a new trace transitions from discrepancy point D1 directly to a segment represented by D4, bypassing D2, signature matching identifies the transition and allows the trace to rejoin the existing execution path. Without this search, the transition would be integrated as a new branch from D1, resulting in unnecessary graph expansion.

New trace

D1 D2

New node

D3

D4

(a) Jump detection Graph

(b) New discrepancy point New trace

Graph

New trace

D1

New branch D2

5.2.2

D3

Trace Insertion

D4

(c) New branch

After capturing a new PSC trace and comparing it against the constructed graph, P OZZER determines whether the trace exhibits previously unseen behavior; if so, it incorporates that behavior into the graph. This step is necessary to maintain the global execution reference used by subsequent fuzzing iterations. The decision falls into one of the following cases: New discrepancy point. If the new PSC trace matches the signature of an existing node but exhibits a discrepancy at a location not represented by an edge, P OZZER creates the corresponding nodes and edge. Figure 6b shows this case. Since the trace reveals a previously unseen execution branch (red), the associated test case is considered interesting and is added to the corpus. New branch. If from a discrepancy point onward, the new PSC trace does not match any existing signature, P OZZER creates a new branch from that discrepancy point. Since all traces share the same entry point, they necessarily match the graph for at least an initial segment. This case introduces previously unseen execution behavior without introducing another point of discrepancy. Figure 6c illustrates this case, where a new branch (red) is added to an already existing discrepancy point. The corresponding test case is also marked as interesting and added to the corpus. New Path. A new PSC trace may also consist entirely of segments already represented in the graph while connecting them in a previously unseen order. Such a trace does not introduce a new segment, but it represents a previously unobserved execution path. P OZZER therefore adds the new edge to the graph and retains the corresponding test input. Figure 6d shows an example in which the trace connects discrepancy points D1 and D3 through a path not previously represented in the graph. Full match. If the entire trace is represented by a prefix on an existing path in the graph, the execution behavior has already been observed. P OZZER therefore discards the corresponding test input and does not modify the graph.

(d) New path

Figure 6: Four trace-insertion cases handled by P OZZER and the corresponding updates to the execution-flow graph.

Signature-based search addresses the non-sequential appearance of execution segments in PSC traces, trace insertion incorporates previously unseen execution behavior into the graph, and graph trimming removes redundant representations introduced during insertion. 5.2.1

Signature-based Search

A newly captured PSC trace cannot be compared only with the local descendants of its current position in the graph. As discussed in Section 3, firmware execution paths can converge, causing the same code segment to appear at different temporal locations in different traces. Consequently, a new trace may deviate from the current graph path and later rejoin an already observed path. Restricting the search to the current branch would interpret such a transition as new behavior and unnecessarily expand the graph. P OZZER addresses this issue using a signature-based search that searches for matching signatures across the entire graph. Since all PSC traces share the same entry point, P OZZER first compares a new trace with the graph from the root until the first discrepancy point. Rather than restricting subsequent comparison to the descendants of that discrepancy point, it searches a global pool of signatures collected from the entire graph. A signature is a fixed-length segment extracted from the beginning of each graph branch and represents the entry point of that branch. Searching across signatures throughout the entire graph enables P OZZER to identify execution-flow transitions that cannot be detected by considering only the local descendants of the current discrepancy point, as program execution does not necessarily follow a single sequential path. This global search also improves the efficiency of trace comparison. Signatures are substantially smaller than the full PSC traces stored in the graph, allowing candidate matches

5.2.3

Graph Trimming

Trace insertion can introduce redundant graph structure when the same execution segment appears at different temporal loca7

tions in PSC traces. As discussed in Section 3, such redundant segments arise naturally when different firmware execution paths converge. If the converged segment starts at an already existing discrepancy point, trace insertion identifies it as a new path, as described in the previous subsection. However, if the convergence point is not an already existing discrepancy point, then the corresponding PSC trace is added to the graph, leading to unnecessary graph expansion. Therefore, P OZZER applies graph trimming after inserting a new trace to identify redundant execution segments and merge them. This reduces the graph size while preserving the paths. Trimming is performed only when the graph is modified, avoiding additional processing for traces that do not introduce new behavior.

6

CW Husky platforms with CW313 target boards. As target MCUs, we use an ST ARM Cortex-M4 STM32F303 and a Microchip ARM Cortex-M4 SAM4S. We evaluate 15 firmware targets spanning a diverse set of embedded firmware components, including protocol and data parsers, binary file parsers, media decoders, and command interpreters. Our target set also includes benchmark programs commonly used by firmware rehosting frameworks [25, 40, 52], enabling a direct comparison with prior embedded firmware fuzzing work, despite P OZZER operating under a more restrictive black-box threat model. The fuzzer runs on an x86 host machine with 32 GB of memory and communicates with the CW platforms over USB. Further details of the experimental setup are provided in Appendix C. Second, to demonstrate P OZZER’s flexibility and independence from the CW platform, we evaluate P OZZER on two real-world black-box targets whose firmware source code and binaries are not publicly available. The first target is a U-Blox ZED-F9P GNSS receiver module. The second is a security chip that cannot be explicitly identified due to its sensitive nature and a restrictive non disclosure agreement (NDA); throughout this paper, we refer to it as the Anonymous Target. Unlike CW target boards, the real-world targets neither include an on-board shunt resistor for PSC measurements nor provide an internal trigger signal. Therefore, we install an external shunt resistor in the power supply path and measure the resulting PSC traces across the resistor using a Tektronix MSO66B oscilloscope equipped with a high speed interface (HSI). To generate the trigger signal, we run P OZZER on a Raspberry Pi 5 and use its GPIO pins. The Raspberry Pi serves as the host, sending test cases to the target while receiving the captured PSC traces from the oscilloscope and processing the feedback.

Evaluation

In this section, we evaluate P OZZER against a blind fuzzer and compare its performance with a coverage-guided fuzzer that provides a reference point for performance under instrumentation. We first describe the experimental setup and evaluation metrics. We then use a simple example program to showcase the internal operation of P OZZER and provide a detailed analysis of its discrepancy detection algorithm using a representative execution flow. Next, we evaluate P OZZER on 15 firmware targets using the ChipWhisperer (CW) platform to compare its performance with the blind fuzzer, followed by a comparison with the related work. Finally, we evaluate P OZZER on two real-world black-box embedded systems whose firmware is not publicly available and investigate its ability to discover vulnerabilities. Throughout these experiments, we address the following research questions: R1: Can P OZZER outperform a blind fuzzer under the same time budget, and is the performance difference statistically significant? R2: How does the overhead introduced by P OZZER’s feedback affect its fuzzing throughput compared with the blind fuzzer, and what is the resulting corpus size of P OZZER? R3: Can P OZZER fuzz real-world black-box embedded systems and discover vulnerabilities?

6.1

6.2

Evaluation metric

We evaluate P OZZER using three primary metrics: code coverage, fuzzing throughput, and corpus size. These metrics capture complementary aspects of fuzzing effectiveness and efficiency. Code Coverage. Since the source code of the firmware targets in the CW setup is available, we use the source code line coverage as the primary measure of fuzzing effectiveness. We use line coverage following established fuzzing evaluation practices [53] to enable consistent comparisons across targets running on MCUs from different vendors. We collect coverage using a unified replay procedure for all experiments. Specifically, after each fuzzing campaign, we replay the resulting corpus on a gcov [29]-instrumented version of the target and record the lines the corpus exercises. Because we use the same replay procedure and source code for all campaigns, the resulting coverage measurements are directly comparable across targets and fuzzers. Fuzzing Throughput. We measure fuzzing throughput as

Experimental setup

In this work, we develop a prototype of P OZZER on top of the LibAFL framework [28]. The PSC-trace acquisition, processing, graph construction, and feedback logic are implemented within P OZZER’s prototype. P OZZER requires a calibration stage to determine a set of setup-specific parameters. The calibration stage does not require any ground-truth data from the target; hence, it does not violate our black-box threat model. The details of this stage are provided in Appendix B. We use two experimental setups to evaluate P OZZER. First, we use the widely adopted CW [45] platform to validate the effectiveness of P OZZER’s approach. Specifically, we use four 8

Constructed Execution Flow Power trace

Groun-truth discrepancy points

firmware with a nested if structure (see Appendix D). The program checks the input against a fixed character sequence, with each correctly matched character reaching the next if condition. Although simple, this structure provides a basic execution flow that reveals how P OZZER constructs executionflow information from PSC traces. We use the ARM debug and watchpoint trace (DWT) unit on the STM32F3 MCU to obtain cycle-accurate instruction traces as ground truth. We select eight test cases that progressively satisfy the nested if conditions by providing increasingly longer correct input prefixes. We then execute these test cases, capture the corresponding PSC traces, and record the instruction traces using the DWT. P OZZER incrementally constructs the execution-flow graph. The first test case does not satisfy the first if. Starting with the second trace, which reaches and satisfies the first if, it identifies a discrepancy in the PSC trace that corresponds to the branch point in the instruction trace. When a trace extending to the second if is observed, P OZZER matches the previously identified discrepancy and detects the new discrepancy corresponding to the second if. It repeats this process for progressively longer executions, incrementally constructing the execution-flow graph and identifying the discrepancy associated with each branch. Figure 7 shows the resulting PSC trace and the eight discrepancy points identified by P OZZER. The red dashed lines indicate the ground-truth branch locations obtained from the DWT traces, while the blue stars indicate the discrepancy points detected by P OZZER. Their close correspondence illustrates how P OZZER recovers the execution-flow structure from PSC traces and identifies newly observed execution behavior. To illustrate the usefulness of this feedback, we run the blind fuzzer, P OZZER, and the greybox-guided fuzzer for five independent 12-hour campaigns on two target MCUs, STM32F3 and SAM4S, and compare how accurately they recover the eight discrepancy points. The execution-flow graph reconstructed from the blind fuzzer’s corpus contains, on average, three of the eight discrepancy points. P OZZER identifies all eight discrepancy points in two of the five campaigns on STM32F3 and three of the five campaigns on SAM4S; in the remaining campaigns on both platforms, it misses one discrepancy point. The greybox-guided fuzzer identifies all eight discrepancy points in every campaign. While this simple example is not intended to support broad performance conclusions, the results illustrate that P OZZER’s feedback can surpass the blind baseline and approach the coverage achieved by an instrumented feedback-driven fuzzer.

Discrepancy points detected by POZZER

0.3

Amplitude (V)

0.2 0.1 0.0 −0.1 −0.2 −0.3 −0.4 0

100

200

300

Time (Sample index)

400

500

Figure 7: Comparison of ground-truth execution discrepancies and the discrepancy points detected by P OZZER for the nestedif target.

the number of test cases executed per unit of time. This metric captures the execution overhead introduced by P OZZER’s PSC-based feedback mechanism and is used to assess its efficiency relative to the blind fuzzer. Because execution rates can vary across hardware platforms and firmware targets, we report throughput separately for each target and platform. Corpus Size. We also measure the number of test cases retained in the final corpus of each fuzzing campaign. Corpus size indicates how selectively the feedback mechanism identifies inputs as interesting and complements coverage when evaluating P OZZER’s feedback effectiveness. Fuzzing Configuration. For all experiments, we use the same mutation strategies and initial seeds across the compared fuzzers. We conduct five independent 12-hour campaigns for each target and report the average results across these repeated runs to account for the inherent randomness of fuzzing [53]. All fuzzers are given the same 12-hour time budget, enabling a time-based comparison despite differences in fuzzing throughput. For the comparisons, we use a blind fuzzer as the baseline because it operates under the same black-box threat model and attacker capabilities as P OZZER. The blind fuzzer uses the same fuzzing configuration as P OZZER, with only the observer and feedback modules disabled. Consequently, every generated test case is retained as interesting. We also use a reference coverage-guided fuzzer executed on a conventional computer as a reference point for performance achievable with direct execution feedback. This fuzzer uses GCC coverage instrumentation and therefore has access to information unavailable to P OZZER under the black-box threat model. Since this fuzzing cannot be performed directly on CW target hardware, we project its results onto the same 12-hour time axis based on the number of test cases executed by the blind fuzzer. This comparison favors the traditional coverage-guided fuzzer by ignoring instrumentation overhead and therefore acts as a conservative baseline for P OZZER’s performance.

6.3

6.4

Results on CW

To validate that P OZZER performs well on real targets, we conduct an extensive evaluation of P OZZER on 15 firmware targets across two hardware platforms. Table 1 reports the average coverage achieved by each fuzzer over five campaigns for each target and MCU. Coverage is computed only over

Example of graph construction

To analyze P OZZER’s graph construction algorithm in a controlled environment, we implement a contrived target 9

Lwjson firmware coverage Edge-guided

90

POZZER SAM4S

Betaflight firmware coverage

Blind

Edge-guided

40

POZZER SAM4S

Jpegdecoder firmware coverage

Blind

Edge-guided

30

50 40 30 20

50 40 30 20

30

Coverage (%)

60

POZZER SAM4S

Blind

20

10

20

10

10

10 0

2

4

6

8

Time (h)

10

0

12

0

Lwjson firmware coverage Edge-guided

90

POZZER STM32F3

2

4

6

Time (h)

8

10

0

12

0

Betaflight firmware coverage

Blind

Edge-guided

70

80

POZZER STM32F3

2

4

6

Time (h)

8

10

0

12

0

Inav firmware coverage

Blind

Edge-guided

40

POZZER STM32F3

2

4

6

Time (h)

8

10

12

Jpegdecoder firmware coverage

Blind

Edge-guided

30

POZZER STM32F3

Blind

6

10

60 50 40 30 20

50 40 30 20

30

Coverage (%)

70

Coverage (%)

60

Coverage (%)

Coverage (%)

Inav firmware coverage

Blind

Coverage (%)

70

20

10

20

10

10

10 0

POZZER SAM4S

60

Coverage (%)

Coverage (%)

80

0

Edge-guided

70

0

2

4

6

8

Time (h)

10

12

0

0

2

4

6

Time (h)

8

10

12

0

0

2

4

6

Time (h)

8

10

12

0

0

2

4

Time (h)

8

12

Figure 8: Coverage plots for evaluated targets. The first row shows targets running on SAM4S and the second row shows targets running on STM32F3. Table 1: Coverage (Cov.) achieved by P OZZER, the blind and the guided fuzzers on STM32F3 and SAM4S platforms. Also shown are the execution speed (executions per second) of the blind fuzzer and P OZZER, the corpus size (Corp. size) of P OZZER, and the Mann-Whitney U-test p-values and A12 effect sizes comparing P OZZER and the blind-fuzzer corpora. SAM4S Target lwjson [14]

Blind

STM32F3

P OZZER

p-value

A12

2760.60

0.0119

Cov.

Speed

Cov.

Speed

Corp. size

60.71%

70.20

71.39%

31.98

Blind

Greybox Guided

P OZZER

p-value

A12

1660.20

0.0116

1.0

74.04%

Cov.

Speed

Cov.

Speed

Corp. size

1.0

59.12%

69.54

67.43%

32.35

Cov.

cAT [6]

18.18%

69.41

20.98%

23.02

270.20

0.0361

0.92

17.95%

65.69

27.30%

22.10

1150.60

0.0079

1.0

43.48%

minmea [12]

28.22%

70.83

62.68%

32.19

3906.00

0.0094

1.0

29.30%

67.48

58.22%

31.57

1313.80

0.0111

1.0

71.46%

lwgps [13]

24.91%

70.00

47.98%

30.82

507.20

0.0211

0.96

23.33%

67.56

30.00%

28.45

152.80

0.2358

0.74

84.74%

jsonparser [5]

79.72%

69.82

90.09%

31.37

3265.00

0.0106

1.0

78.44%

65.94

87.16%

31.23

1369.20

0.0116

1.0

91.74%

nanomodbus [1]

25.43%

68.35

26.11%

28.40

2584.00

0.2873

0.72

26.08%

70.89

24.85%

26.10

3644.60

0.0231

0.06

30.95%

regex [7]

84.04%

69.03

87.17%

25.00

4412.60

0.0116

1.0

83.84%

66.13

81.82%

23.88

2104.40

0.1732

0.22

87.88%

libelf [9]

7.19%

66.86

33.35%

30.68

1771.00

0.0085

1.0

7.07%

69.19

6.83%

29.21

1624.80

0.2330

0.32

35.29%

jpegdecoder [4]

7.78%

64.11

21.56%

26.86

3306.80

0.0109

1.0

7.32%

67.66

16.70%

24.05

2614.80

0.0116

1.0

24.53%

microshell [8]

82.88%

65.82

88.00%

24.94

4046.80

0.0593

0.88

84.24%

66.60

77.76%

24.42

2247.20

0.0119

0.0

95.60%

betaflight [10]

31.69%

69.90

48.97%

24.65

2480.40

0.0119

1.0

32.19%

69.18

51.50%

24.14

1238.40

0.0119

1.0

62.21%

inav [11]

12.52%

67.49

28.61%

27.87

2112.60

0.0116

1.0

11.47%

67.45

27.71%

23.24

2297.40

0.0114

1.0

29.51%

drone [3]

9.14%

65.79

56.12%

24.91

2383.00

0.0101

1.0

8.79%

71.10

35.69%

24.40

709.40

0.0109

1.0

60.34%

cnc [3]

50.76%

68.95

60.48%

25.43

4359.40

0.0079

1.0

50.97%

70.12

54.82%

26.57

2378.00

0.0159

0.98

79.40%

stepper [2]

36.38%

67.75

43.28%

15.95

2353.00

0.0106

1.0

36.38%

68.17

43.20%

13.17

1466.20

0.0109

1.0

57.10%

the target source-code files and normalized by the number of lines considered instrumentable by gcov, rather than by the total number of lines across all firmware files. The Cov. columns under Blind report the average coverage achieved by the blind fuzzer on each target and platform, while the Cov. columns under P OZZER report the corresponding coverage achieved by P OZZER. Since the greybox-guided fuzzer runs on an x86 machine, we report a common reference coverage with instrumentation-based feedback for both platforms in the Greybox guided column. Overall, P OZZER outperforms the blind fuzzer on 26 out of 30 target-platform combinations. Figure 8 presents representative coverage plots; the remaining

plots are provided in Appendix E due to space constraints. Statistical Analysis. To further assess P OZZER’s performance, we statistically compare its coverage with that of the blind fuzzer across repeated campaigns. Since fuzzing is inherently randomized, a single campaign is insufficient to characterize a fuzzer’s performance. Unlike prior work [42], which reports results from a single campaign, we repeat each experiment five times to account for the effect of randomness. Following the statistical evaluation methodology recommended by Schloegel et. al. [53], we use the Mann-Whitney U-test [51] to assess statistical significance and the VarghaDelaney A12 effect size [59] to quantify the magnitude of the 10

Table 2: Comparison of edge coverage and fuzzing speed (executions per second) between FuzzEMup and P OZZER. GPS

Fuzzer

Soldering

CNC

Stepper

Edge Cov.

Speed

Edge Cov.

Speed

Edge Cov.

Speed

Edge Cov.

FuzzEMup [42]

909

0.04

1086

0.36

1164

0.03

1341

0.09

P OZZER (FuzzEMup duration)

1380

23.96

1076

4.40

1175

23.44

1346

15.16

POZZER (12-hours)

1509

23.26

1110

4.20

1176

23.00

1431

13.90

observed differences between P OZZER and the blind fuzzer. We use a significance level of 0.05. The A12 effect size represents the probability that a randomly selected run of P OZZER achieves higher coverage than a randomly selected run of the blind fuzzer. An A12 value of 0.5 indicates no difference between the two fuzzers, values greater than 0.5 favor P OZZER, and values less than 0.5 favor the blind fuzzer. Table 1 reports both the p-values obtained from the MannWhitney U-test and the corresponding A12 effect sizes for each target-platform combination. Reporting both metrics is important because statistical significance alone does not indicate the magnitude of the observed improvement. As shown in the table, 25 out of 30 target-platform combinations achieve a p-value below 0.05, indicating that the difference in coverage between P OZZER and the blind fuzzer is statistically significant. Furthermore, 26 out of 30 targets achieve an A12 value larger than 0.5, indicating that a randomly selected run of P OZZER is more likely to achieve higher coverage than a randomly selected run of the blind fuzzer.

Speed

R2: Despite the overhead introduced by its feedback mechanism, P OZZER achieves an average throughput of 26.30 executions per second and an average corpus size of 2216.35 test cases.

6.5

Comparison with related work

Several works have explored side channel-guided fuzzing for embedded systems [42, 60]. However, they either provide insufficient information about performance or contain inconsistencies in their threat models or evaluations. According to established fuzzing practices [18, 34, 53], a fair comparison of fuzzers should satisfy the following criteria: 1. Allocating an equal time budget to all fuzzers 2. Running the fuzzers for sufficiently long periods, ranging from several hours to a full day 3. Repeating fuzzing campaigns to account for randomness However, the related works do not satisfy these criteria. First, they evaluate fuzzers based on an equal input budget. Because capturing and processing PSC traces introduces additional overhead, this comparison ignores the difference in throughput between blind and PSC-guided fuzzers. Second, the comparisons use a limited number of test cases, typically around 1,000 inputs. However, fuzzing is inherently randomized and requires longer execution times to explore the program space effectively. Finally, the existing studies report results from only a single campaign, whereas a standard evaluation requires repeated campaigns to account for randomness [53]. In our evaluation, we address all of these limitations. We compare P OZZER with FuzzEMup [42]. To ensure a fair comparison with FuzzEMup, we contacted its authors and obtained their artifacts, including the seeds, corpus for each target, coverage collection script, and target source code. We then ported the targets to our CW setup and fuzzed them with P OZZER using the same seeds. Finally, we replayed the corpus generated by P OZZER using FuzzEMup’s coverage collection script and collected the corresponding edge coverage. Table 2 reports the execution speed and edge coverage of both fuzzers for each target. Because FuzzEMup’s execution speed is not reported directly in its paper, we obtained this information from its authors. In addition, FuzzEMup does not conduct standard 12-hour campaigns; hence, comparing its results directly with those from a 12-hour P OZZER campaign

R1: P OZZER outperforms the blind fuzzer under the same time budget on 26 of the 30 target-platform combinations, with the improvement being statistically significant on 25 of them. Speed. A guided fuzzer generally incurs overhead compared with a blind fuzzer because it must process feedback and determine whether each test case is interesting. For a PSCguided fuzzer, this overhead is particularly significant because PSC traces are large time-series signals whose processing can substantially reduce fuzzing throughput. To limit this cost, P OZZER uses the fast L1 distance for discrepancy detection. Its graph construction and trimming stage further removes redundancy from the constructed graph, thereby controlling the growth of the corpus and graph and reducing the cost of subsequent feedback decisions. While previous work [42] does not report speed or corpus size, the speed column in Table 1 reports the average fuzzing throughput and corpus size for each target over the five 12-hour campaigns. The Corp. size column also reports the corpus size as a percentage of the total number of executions performed by P OZZER. The results show that P OZZER’s throughput varies across targets, reflecting differences in target execution times. Nevertheless, the resulting corpus remains reasonably small in all campaigns. 11

7

would be unfair. Therefore, we report the results in two different rows. The P OZZER (FuzzEMup duration) row reports the results when P OZZER is allocated the same time budget as FuzzEMup, whereas the P OZZER (12-hour) row reports the results from a full 12-hour campaign. As shown in Table 2, under an equal time budget, P OZZER achieves a higher execution rate and covers more edges than FuzzEMup on three out of four targets. With a standard 12hour campaign, P OZZER outperforms FuzzEMup on all four targets. These experiments also highlight the importance of high-quality seeds. By measuring the unique edges contributed by each test case, we observe that most of the unique edge coverage is achieved by the initial high-quality seeds provided by FuzzEMup. However, such high-quality seeds might not be available in real-world black-box scenarios.

6.6

Discussion

In this work, we present P OZZER, the first PSC-guided blackbox fuzzer to outperform a blind fuzzer under the same time budget. Our evaluation highlights three aspects that warrant further discussion: seed selection, performance variation of P OZZER across hardware platforms, and scalability challenges arising from long executions and large graphs. Seed selection. Seed selection can strongly influence fuzzing performance, particularly for targets with structured input formats. In Section 6.4, we initialize all campaigns with a random seed, except for lwgps and minmea. These targets require inputs that conform to a strict grammar and begin with specific magic values. P OZZER, the blind fuzzer, and the traditional coverage-guided fuzzer generate inputs based on existing seeds, so both benefit equally from seeding with magic values. For the remaining targets, P OZZER starts from random seeds yet outperforms the blind fuzzer, suggesting that its feedback can guide input generation toward valid input formats and previously unexplored behavior. This result contrasts with the evaluation of FuzzEMup [42], which uses meaningful seeds for every target. As reported in Table 2, when we fuzz the FuzzEMup targets with P OZZER using the same seeds, P OZZER achieves higher edge coverage than FuzzEMup. Table 2 also shows that P OZZER achieves higher execution throughput. This higher throughput provides a plausible explanation for its improved coverage: by executing more test cases within the same time budget, P OZZER has more opportunities to explore the input space and discover inputs that exercise additional behavior. These results highlight the importance of feedback-processing efficiency and execution speed in PSC-guided fuzzing. Performance variation. A second discussion point is the performance variation of P OZZER across hardware platforms. Table 1 shows that P OZZER’s performance varies across platforms. In particular, it achieves lower average coverage on STM32F3 compared to SAM4S. To investigate this variation, we compare the compiled binary images for both platforms for the same targets. Although the binary images implement semantically similar firmware behavior, they differ in their instruction sequences and execution characteristics because the target MCUs uses different vendor-specific toolchains and implementations. These differences can affect the power characteristics of the observed executions, thereby influencing the quality of the PSC-based feedback and P OZZER’s resulting performance. Identifying the precise cause of this platform-dependent behavior would therefore require detailed knowledge of the target MCUs’s hardware implementations and power characteristics. Scalability. A third issue concerns scalability with execution length and graph size. Longer executions produce larger PSC traces and graphs, reducing P OZZER’s throughput. For the Anonymous Target, the execution length is 20 ms. At a sampling frequency of 25 MS/s, each test case there-

Real-world case study

After evaluating P OZZER extensively on the CW setup, we evaluate it on two real-world black-box embedded systems for which neither the firmware source code nor the binary image is publicly available. We fuzz a U-Blox ZED-F9P GNSS receiver module and an Anonymous Target. On the GNSS receiver, P OZZER achieves a throughput of 7.19 executions per second despite the substantially larger PSC traces compared with the CW setup. After three 4-hour fuzzing campaigns, P OZZER discovers two distinct crashes on the ZED-F9P. Both crashes were confirmed by the vendor, and one CVE1 assigned so far. Due to the sensitive nature of the Anonymous Target and the restrictions imposed by our NDA, we can not disclose the results for this target. Bug description. The ZED-F9P module communicates through u-blox’s proprietary UBX protocol, a modular binary protocol with a low-overhead checksum. UBX defines more than 100 message types organized into message classes for configuring the module and querying navigation information. We use the UBX interface as the input channel for fuzzing. According to the UBX protocol specification, floating-point fields are encoded according to the IEEE 754 standard. Our analysis shows that the module crashes when one of these fields contains an IEEE 754 encoding of NaN, a non-finite value that the firmware does not handle correctly. When the module receives a crafted input containing such a value, it crashes and reboots, indicating insufficient validation or handling of special floating-point values. R3: P OZZER operates independently of the CW platform, can fuzz real-world black-box embedded systems, and can discover vulnerabilities in such targets.

1 CVE id redacted for double blind review.

12

fore produces approximately 500,000 samples. Processing such a large PSC trace for every test case and constructing a graph from it further reduces P OZZER’s throughput. Moreover, when an oscilloscope is used, the trigger signal must also be transmitted to and processed by P OZZER’s observer, introducing additional time overhead. These limitations motivate future work on implementing P OZZER’s processing pipeline on a high-performance field-programmable gate array (FPGA). Such an implementation could capture and process each PSC trace, compare it with the execution-flow graph stored on the FPGA, and notify the host-side fuzzer when a trace exhibits previously unseen behavior. The host would then add the corresponding test case to the corpus. This design would eliminate the need to transfer every PSC trace to the host and allow the FPGA to maintain the graph and perform the feedback computation locally.

8

a source of fuzzing feedback. Sperl et al. [54] first proposed employing PSC together with a supervised machine-learning classifier to identify firmware branches. However, their approach relies on a profiling stage and assumes a white-box threat model. McClintick et al. [39] proposed an electromagnetic (EM) side-channel-guided fuzzer based on trace clustering and, to the best of our knowledge, presented the first timing-based comparison against blind fuzzing. Nevertheless, their evaluation is limited to a manually constructed parser and requires high sampling resolutions. DuskFuzz [60] integrates EM side-channel feedback into LibAFL [28] by exploiting a microarchitectural leakage originating from the instruction prefetch buffer. It uses this leakage to infer execution-flow discrepancies and guide the fuzzing process. However, the approach requires five executions for each test case, significantly reducing fuzzing throughput, and does not provide a timing-based comparison with blind fuzzing. Fuzz’EMup [42] applies a frequency-domain feature extractor on EM sidechannels and implements a divergence-based feedback mechanism for an EM-guided fuzzer. They evaluate their work on four firmware benchmarks and achieve better performance than a blind fuzzer, but they do not provide information about their fuzzer’s speed. Furthermore, they do not give the same time budget to the fuzzers and stop the fuzzers after at most 1000 executions. This makes their evaluation vague. Other side-channel-guided fuzzers have also been proposed [16,57]; however, they either assume stronger attacker models or do not demonstrate practical improvements over blind fuzzing under comparable time budgets. Unlike previous work, P OZZER is designed specifically for a strict black-box threat model. It requires neither firmware instrumentation, firmware binaries, nor profiling data, extracts meaningful feedback from a single power trace, and, to the best of our knowledge, is the first PSC-guided fuzzer to consistently outperform blind fuzzing under the same time budget.

Related work

This section reviews the work most closely related to P OZZER. We first discuss existing embedded firmware fuzzing techniques under different threat models. We then review blackbox fuzzing approaches that rely on software-generated feedback before discussing prior work on side-channel-guided fuzzing. Embedded firmware fuzzers differ primarily in the assumptions they make about the target system and the execution feedback available during fuzzing. When firmware source code is available, instrumentation-based approaches execute the instrumented firmware directly on the target device to obtain code coverage [41]. Alternatively, hardware debugging interfaces can be used to extract execution feedback without modifying the firmware [23]. Firmware rehosting provides another widely adopted solution by executing the firmware inside an emulator. Unlike instrumentation-based methods, rehosting does not require source code. For example, Fuzzware [52] performs coverage-guided fuzzing directly on firmware binaries through emulation. Although these approaches differ in their assumptions, they all require some level of access to the firmware, either through source code, firmware binaries, or hardware debugging interfaces. To relax these assumptions, recent work has explored fuzzing techniques that operate under black-box settings. Rather than relying on code coverage, these approaches leverage software-generated information exposed by the target, such as grammar specifications [62], protocol descriptions, system logs [15], or protocol snippets [26], to improve input generation. While effective in their respective scenarios, they still assume that execution-related information is available through the software interface. Consequently, they cannot exploit undocumented commands, hidden execution paths, or implementation-specific behaviors that are not reflected in the exposed interface. Several studies have investigated physical side channels as

9

Conclusion

In this paper, we presented P OZZER, a PSC-guided fuzzer for black-box embedded systems, and demonstrated that it achieves superior fuzzing performance compared with a baseline blind fuzzer under the same time budget. P OZZER uses PSC traces to construct a pseudo execution-flow graph of the target firmware and integrates hardware and software-level noise-handling techniques into its feedback mechanism. We extensively evaluated P OZZER on the CW platform using two different hardware platforms and 15 firmware targets. Our results show that, across more than 3,600 hours of fuzzing, P OZZER achieves, on average, 12.06% better performance than the baseline blind fuzzer. We further demonstrate that P OZZER is not limited to the CW platform by connecting it to an oscilloscope and evaluating it on two real-world black-box embedded systems. On one of these targets, P OZZER identified two crashes, which were both confirmed by the vendor. 13

Ethical Considerations We organize our ethical analysis around the stakeholders affected by the research, the potential benefits and harms to each, the measures taken to reduce those harms, and our decision to conduct and publish the work. Stakeholders and benefits. The principal stakeholders are users and operators of the evaluated software, project maintainers and contributors, the security-research community, and the research team. Users benefit when previously unknown defects are identified and corrected before independent discovery or exploitation. Maintainers receive limited input and technical evidence to support diagnosis and remediation, although processing reports requires engineering effort and may create reputational concerns. The research community benefits from a reproducible method for testing black-box embedded devices. At the same time, the resulting system is dual-use: it could help an adversary discover defects in software for which the source code or binary image is unavailable. We believe that the security improvements in hard-to-test devices resulting from publishing our tool outweigh the risks of potential adversarial use. We choose not to publish the harnesses for our real-world case studies. Research conduct. Our experiments were performed in a controlled lab environment. For the Chipwhisperer targets, we used publicly available projects and compiled them for the corresponding target platform. For the real-world case studies, we had two targets. The experiments for both targets were conducted in a controlled environment in our lab. For the anonymous target, we have shared the draft of our paper with the corresponding vendor, and they have confirmed that the current version is okay for publication. Vulnerability disclosure. For the Chipwhisperer experiment, we did not find any unknown vulnerability. For the real-world case study, we have found two vulnerabilities on the UBlox GNSS receiver, and we disclosed the findings to the vendor. They confirmed their reproducibility, and we have assigned one CVE so far. Reports included a minimal reproducer, affected versions, observed impact, and available root cause information, as described in the paper. We coordinated subsequent disclosure with the vendor. For the Anonymous Target, reports remain private. Therefore, we omit project identities and other information that would materially facilitate exploitation.

Open Science We will release our implementation’s source code and the results of the paper after its acceptance.

14

References

[13] GitHub - MaJerle/lwgps: Lightweight GPS NMEA parser for embedded systems — github.com. https: //github.com/{M}a{J}erle/lwgps, 2026.

[1] GitHub - debevv/nanoMODBUS: A compact MODBUS RTU/TCP C library for embedded/microcontrollers — github.com. https://github.com/debevv/ nano{M}{O}{D}{B}{U}{S}/tree/master. [Accessed 25-08-2026].

[14] GitHub - MaJerle/lwjson: Lightweight JSON parser for embedded systems — github.com. https://github. com/{M}a{J}erle/lwjson, 2026. [15] Yousra Aafer, Wei You, Yi Sun, Yu Shi, Xiangyu Zhang, and Heng Yin. Android {SmartTVs} vulnerability discovery via {log-guided} fuzzing. In 30th USENIX Security Symposium (USENIX Security 21), pages 2759– 2776, 2021.

[2] GitHub - RiS3-Lab/DICE-DMA-Emulation: DICE: Automatic Emulation of DMA Input Channels for Dynamic Firmware Analysis — github.com. https://github.com/{R}i{S}3-{L}ab/ {D}{I}{C}{E}-{D}{M}{A}-{E}mulation/tree/ master. [Accessed 25-08-2026].

[16] Jorge Barredo, Justyna Petke, David Clark, Daniel Blackwell, Maialen Eceiza, Jose Luis Flores, and Mikel Iturbe. Gaflerna ahoy! integrating em side-channel analysis into traditional fuzzing workflows. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pages 550–554, 2025.

[3] GitHub RiS3-Lab/p2im-real_firmware at d4c7456574ce2c2ed038e6f14fea8e3142b3c1f7 — github.com. https://github.com/ {R}i{S}3-{L}ab/p2im-real_firmware/tree/ d4c7456574ce2c2ed038e6f14fea8e3142b3c1f7. [Accessed 25-08-2026].

[17] Fabrice Bellard et al. Qemu, a fast and portable dynamic translator. In Usenix ATC, Freenix Track, pages 41–46, 2005.

[4] GitHub - cmumford/TJpgDec: Tiny JPEG Decompressor — github.com. https://github.com/cmumford/ {T}{J}pg{D}ec, 2021.

[18] Marcel Böhme, Cristian Cadar, and Abhik Roychoudhury. Fuzzing: Challenges and reflections. IEEE Software, 38(3):79–86, 2020.

[5] GitHub - rafagafe/tiny-json: The tiny-json is a versatile and easy to use json parser in C suitable for embedded systems. It is fast, robust and portable. — github.com. https://github.com/rafagafe/tiny-json, 2021.

[19] Eric Brier, Christophe Clavier, and Francis Olivier. Correlation power analysis with a leakage model. In Cryptographic Hardware and Embedded Systems-CHES 2004: 6th International Workshop Cambridge, MA, USA, August 11-13, 2004. Proceedings 6, pages 16–29. Springer, 2004.

[6] GitHub - marcinbor85/cAT: Plain C library for parsing AT commands for use in host devices. — github.com. https://github.com/marcinbor85/cAT, 2023. [7] GitHub - kokke/tiny-regex-c: Small portable regex in C — github.com. https://github.com/kokke/ tiny-regex-c, 2024.

[20] Elie Bursztein, Luca Invernizzi, Karel Král, Daniel Moghimi, Jean-Michel Picod, and Marina Zhang. Generalized power attacks against crypto hardware using longrange deep learning. arXiv preprint arXiv:2306.07249, 2023.

[8] GitHub - Nrusher/nr_micro_shell: shell for MCU. — github.com. https://github.com/{N}rusher/nr_ micro_shell/tree/master, 2025.

[21] Jiongyi Chen, Wenrui Diao, Qingchuan Zhao, Chaoshun Zuo, Zhiqiang Lin, XiaoFeng Wang, Wing Cheong Lau, Menghan Sun, Ronghai Yang, and Kehuan Zhang. Iotfuzzer: Discovering memory corruptions in iot through app-based fuzzing. In NDSS, pages 1–15, 2018.

[9] GitHub - 0intro/libelf: Libelf is a simple library to read ELF files. — github.com. https://github.com/ 0intro/libelf, 2026. [10] GitHub - betaflight/betaflight: Open Source Flight Controller Firmware — github.com. https://github. com/betaflight/betaflight, 2026.

[22] Jie Ding, Mahyar Nemati, Chathurika Ranaweera, and Jinho Choi. Iot connectivity technologies and applications: A survey. IEEE Access, 8:67646–67673, 2020.

[11] GitHub - iNavFlight/inav: INAV: Navigation-enabled flight control software — github.com. https:// github.com/i{N}av{F}light/inav.git, 2026.

[23] Max Eisele, Daniel Ebert, Christopher Huth, and Andreas Zeller. Fuzzing embedded systems using debug interfaces. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, pages 1031–1042, 2023.

[12] GitHub - kosma/minmea: a lightweight GPS NMEA 0183 parser library in pure C — github.com. https: //github.com/kosma/minmea, 2026. 15

[24] ESPRESSIF. ESP32 wireless module. https://www. espressif.com/en/products/modules.

[34] George Klees, Andrew Ruef, Benji Cooper, Shiyi Wei, and Michael Hicks. Evaluating fuzz testing. In Proceedings of the 2018 ACM SIGSAC conference on computer and communications security, pages 2123–2138, 2018.

[25] Bo Feng, Alejandro Mera, and Long Lu. {P2IM}: Scalable and hardware-independent firmware testing via automatic peripheral interface modeling. In 29th USENIX Security Symposium (USENIX Security 20), pages 1237– 1254, 2020.

[35] Paul Kocher, Joshua Jaffe, and Benjamin Jun. Differential power analysis. In Annual international cryptology conference, pages 388–397. Springer, 1999. [36] Yu-De Lin and Nils Ole Tippenhauer. Execution divergence graphs: Effective discovery of control-flows from execution traces as fuzzing feedback. arXiv preprint arXiv:2607.03396, 2026.

[26] Xiaotao Feng, Ruoxi Sun, Xiaogang Zhu, Minhui Xue, Sheng Wen, Dongxi Liu, Surya Nepal, and Yang Xiang. Snipuzz: Black-box fuzzing of iot firmware via message snippet inference. In Proceedings of the 2021 ACM SIGSAC conference on computer and communications security, pages 337–350, 2021.

[37] Chen Ling, Jinyuan Zhang, Ouchang Hai, Hangcheng Liu, and Xingshuo Han. Fusiondisassembler: A crossdevice approach for effective instruction disassembly in side-channel attacks. In 2025 IEEE 31th International Conference on Parallel and Distributed Systems (ICPADS), pages 1–8. IEEE, 2025.

[27] Eduardo Ferrufino, Luke Beckwith, Abubakr Abdulgadir, and Jens-Peter Kaps. Fobos 3: An open-source platform for side-channel analysis and benchmarking. In Proceedings of the 2023 Workshop on Attacks and Solutions in Hardware Security, pages 5–14, 2023.

[38] Yannan Liu, Lingxiao Wei, Zhe Zhou, Kehuan Zhang, Wenyuan Xu, and Qiang Xu. On code execution tracking via power side-channel. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pages 1019–1031, 2016.

[28] Andrea Fioraldi, Dominik Maier, Dongjia Zhang, and Davide Balzarotti. LibAFL: A Framework to Build Modular and Reusable Fuzzers. In Proceedings of the 29th ACM conference on Computer and communications security (CCS), CCS ’22. ACM, November 2022.

[39] Kyle W Mcclintick, Erez Binyamin, Brandon V John, Kyle W Ingols, Brendon R Chetwynd, and Emily K Shields. Side-channel assisted real-time fuzzing for embedded systems (scarfes): Fy23 cyber security linesupported program. 2024.

[29] Free Software Foundation. Gcov (Using the GNU Compiler Collection (GCC)). https://gcc.gnu.org/ onlinedocs/gcc/Gcov.html, 2021. [30] Davide Galli, Giuseppe Chiari, Davide Zoni, et al. Chameleon: A dataset for segmenting and attacking obfuscated power traces in side-channel analysis. IACR Transactions on Cryptographic Hardware and Embedded Systems, 2025(3):389–412, 2025.

[40] Alejandro Mera, Bo Feng, Long Lu, and Engin Kirda. Dice: Automatic emulation of dma input channels for dynamic firmware analysis. In 2021 IEEE Symposium on Security and Privacy (SP), pages 1938–1954. IEEE, 2021.

[31] Ognjen Glamočanin, Shashwat Shrivastava, Jinwei Yao, Nour Ardo, Mathias Payer, and Mirjana Stojilović. Instruction-level power side-channel leakage evaluation of soft-core cpus on shared fpgas. Journal of Hardware and Systems Security, 7(2):72–99, 2023.

[41] Alejandro Mera, Changming Liu, Ruimin Sun, Engin Kirda, and Long Lu. SHiFT: Semi-hosted fuzz testing for embedded applications. In 33rd USENIX Security Symposium (USENIX Security 24), pages 5323– 5340, Philadelphia, PA, August 2024. USENIX Association. URL: https://www.usenix.org/conference/ usenixsecurity24/presentation/mera.

[32] Yi Han, Matthew Chan, Zahra Aref, Nils Ole Tippenhauer, and Saman Zonouz. Hiding in plain sight? on the efficacy of power side Channel-Based control flow monitoring. In 31st USENIX Security Symposium (USENIX Security 22), pages 661– 678, Boston, MA, August 2022. USENIX Association. URL: https://www.usenix.org/conference/ usenixsecurity22/presentation/han.

[42] Fatemeh Moradihaghighi, Zihao Zhan, Yanan Guo, Ziming Zhao, Mashrur Chowdhury, and Zhenkai Zhang. Fuzz’emup: Leveraging em side-channel emanation to guide black-box embedded firmware fuzzing. In 2026 IEEE International Symposium on Hardware Oriented Security and Trust (HOST), pages 174–185. IEEE, 2026.

[33] Chenxi Huang, Jian Wang, Shuihua Wang, and Yudong Zhang. Internet of medical things: A systematic review. Neurocomputing, 557:126719, 2023.

[43] Pouya Narimani, Mohammad Ali Akhaee, and Seyed Amin Habibi. Side-channel based disassembler for avr micro-controllers using convolutional neural 16

networks. In 2021 18th International ISC Conference on Information Security and Cryptology (ISCISC), pages 75–80. IEEE, 2021.

[54] Philip Sperl and Konstantin Böttinger. Side-channel aware fuzzing. In European Symposium on Research in Computer Security, pages 259–278. Springer, 2019.

[44] Pouya Narimani, Meng Wang, Ulysse Planta, and Ali Abbasi. Exploring power side-channel challenges in embedded systems security. arXiv preprint arXiv:2410.11563, 2024.

[55] François-Xavier Standaert, François Koeune, and Werner Schindler. How to compare profiled sidechannel attacks? In International Conference on Applied Cryptography and Network Security, pages 485–498. Springer, 2009.

[45] Colin O’Flynn and Zhizhang (David) Chen. Chipwhisperer: An open-source platform for hardware embedded security research. In Emmanuel Prouff, editor, Constructive Side-Channel Analysis and Secure Design, pages 243–260, Cham, 2014. Springer International Publishing.

[56] STMicroelectronics. STSAFE. https://www.st.com/ en/secure-mcus/stsafe-a100.html. [57] Kai Su, Mark Giraud, Anne Borcherding, Jonas Krautter, Philipp Nenninger, and Mehdi Tahoori. Fuzz wars: The voltage awakens–voltage-guided blackbox fuzzing on fpgas. In 2024 IEEE 42nd VLSI Test Symposium (VTS), pages 1–7. IEEE, 2024.

[46] Kostas Papagiannopoulos, Ognjen Glamočanin, Melissa Azouaoui, Dorian Ros, Francesco Regazzoni, and Mirjana Stojilović. The side-channel metrics cheat sheet. ACM Computing Surveys, 55(10):1–38, 2023.

[58] U-Blox. U-Blox GNSS Receiver. https://www.u-blox.com/en/ positioning-chips-and-modules.

[47] Jungmin Park, Xiaolin Xu, Yier Jin, Domenic Forte, and Mark Tehranipoor. Power-based side-channel instruction-level disassembler. In Proceedings of the 55th annual design automation conference, pages 1–6, 2018.

[59] András Vargha and Harold D Delaney. A critique and improvement of the cl common language effect size statistics of mcgraw and wong. Journal of Educational and Behavioral Statistics, 25(2):101–132, 2000.

[48] Kangran Pu, Hua Dang, Fancong Kong, Jingqi Zhang, and Weijiang Wang. A quantitative analysis of nonprofiled side-channel attacks based on attention mechanism. Electronics, 12(15), 2023.

[60] Ulysse Vincenti, Thomas Hiscock, and David Hely. Duskfuzz: Encoding side-channel information to improve blackbox fuzzing.

[49] Md Abdur Rahim, Md Arafatur Rahman, Md Mustafizur Rahman, A Taufiq Asyhari, Md Zakirul Alam Bhuiyan, and Devarajan Ramasamy. Evolution of iot-enabled connectivity and applications in automotive industry: A review. Vehicular communications, 27:100285, 2021.

[61] Hansong Xu, Wei Yu, David Griffith, and Nada Golmie. A survey on industrial internet of things: A cyberphysical systems perspective. Ieee access, 6:78238– 78259, 2018.

[50] RENESAS. CAN Sensor Network Development Kit. https://www.renesas. com/en/design-resources/boards-kits/ tw007-indcanpocz.

[62] José Antonio Zamudio Amaya, Marius Smytzek, and Andreas Zeller. Fandango: evolving language-based testing. Proceedings of the ACM on Software Engineering, 2(ISSTA):894–916, 2025.

[51] Lothar Sachs. Applied statistics: a handbook of techniques. Springer Science & Business Media, 2012.

[63] Yaowen Zheng, Ali Davanian, Heng Yin, Chengyu Song, Hongsong Zhu, and Limin Sun. FIRM-AFL: HighThroughput greybox fuzzing of IoT firmware via augmented process emulation. In 28th USENIX Security Symposium (USENIX Security 19), pages 1099– 1114, Santa Clara, CA, August 2019. USENIX Association. URL: https://www.usenix.org/conference/ usenixsecurity19/presentation/zheng.

[52] Tobias Scharnowski, Nils Bars, Moritz Schloegel, Eric Gustafson, Marius Muench, Giovanni Vigna, Christopher Kruegel, Thorsten Holz, and Ali Abbasi. Fuzzware: Using precise {MMIO} modeling for effective firmware fuzzing. In 31st USENIX Security Symposium (USENIX Security 22), pages 1239–1256, 2022. [53] Moritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, Addison Crump, Arash Ale-Ebrahim, Nicolai Bissantz, Marius Muench, and Thorsten Holz. Sok: Prudent evaluation practices for fuzzing. In 2024 IEEE Symposium on Security and Privacy (SP), pages 1974–1993. IEEE, 2024.

A

First order derivatives

In this appendix, we mathematically show the effect of firstorder derivatives on low-frequency amplitude offsets. Consider a measured signal composed of the desired signal s[n] and an amplitude offset o[n]: 17

minimum temporal distance allowed between two consecutive discrepancy points in the constructed graph. Although the neighbor-search mechanism compensates for time shifts during the point-wise comparison stage, small time misalignments can still introduce spurious discrepancy points during chunk-based comparison. Enforcing a minimum node size prevents such closely spaced discrepancy points from being inserted into the graph, thereby reducing false positives and avoiding redundant graph nodes. The remaining of the sizerelated parameters are determined directly from the temporal resolution of the acquired PSC traces and therefore depend only on the sampling rate of the measurement device and the clock frequency of the target MCU. On the other hand, the threshold parameters depend on the electrical characteristics of the target and the measurement setup. To estimate these parameters, P OZZER performs a one-time calibration stage before each fuzzing campaign. During calibration, P OZZER automatically generates a set of test cases, repeatedly executes each test case on the target, and captures the corresponding PSC traces. Since each test case is repeated under the same conditions, we attribute the variation among the resulting traces to measurement noise. For calibration, P OZZER generates M = 100 random test cases and captures the corresponding PSC trace N = 100 times for each test case. Let Xm = {xm,n }Nn=1 denote the set of repeated PSC traces corresponding to the m-th test case. Following [46], we assume the measurement noise follows a Gaussian distribution for each sample. The sample covariance matrix of the repeated measurements for test case m is therefore estimated as

x[n] = s[n] + o[n]. The first-order discrete derivative is defined as: ∆x[n] = x[n] − x[n − 1]. Substituting x[n] gives: ∆x[n] = (s[n] + o[n]) − (s[n − 1] + o[n − 1]) = ∆s[n] + ∆o[n]. For a constant (DC) offset, o[n] = C, where C is a constant. Therefore, ∆o[n] = C −C = 0 Hence, ∆x[n] = ∆s[n] showing that the DC offset is eliminated by the first-order derivative. More generally, suppose the offset varies slowly with time, o[n] ≈ o[n − 1]. Then, ∆o[n] = o[n] − o[n − 1] ≈ 0,

Σm =

which implies

N 1 ∑ (xm,n − x̄m )(xm,n − x̄m )T , N − 1 n=1

where the mean of PSC trace for the m-th test case is ∆x[n] = ∆s[n] + ∆o[n] ≈ ∆s[n]. x̄m =

Therefore, the influence of low-frequency amplitude variations is significantly reduced after differentiation. Only rapidly varying components contribute appreciably to the derivative.

B

1 N ∑ xm,n N n=1

Repeating this procedure for all calibration test cases yields M covariance matrices. Since threshold estimation only requires the variance of each sample, we retain only the diagonal entries of Σm . For each test case, we then select the maximum variance as σ2m = max(diag(Σm )), which provides a conservative estimate of the worst-case measurement noise. Finally, the average worst-case noise variance over the entire calibration set is computed as

Calibration

As discussed earlier, P OZZER operates under a black-box threat model and therefore does not rely on any prior knowledge of the target firmware or profiling data. Nevertheless, a set of setup-specific parameters must be configured to account for differences in the measurement setup, such as the target MCU, printed circuit board (PCB), and the acquisition device. These parameters include chunk size, chunk threshold, signature size, signature threshold, time-shift window size, point-wise discrepancy threshold, and minimum node size. Among all parameters, the minimum node size plays a particularly important role in graph construction. It specifies the

σ̄2 =

1 M 2 ∑ σm . M m=1

The discrepancy threshold is then selected such that the probability of the measurement noise remaining below the threshold is upper-bounded to 0.9999. Under the zero-mean Gaussian assumption, this corresponds approximately to four standard deviations, 18

 P (xm,n,i < th) = Φ

th σ̄

 = 0.9999

=⇒

th ≈ 4σ̄

where Φ denotes the standard normal cumulative distribution function (CDF) and σ̄ is the average worst-case noise standard deviation over the entire calibration set. The calibration stage, therefore, produces the setup-specific parameter set used during fuzzing. Once calibration is complete, all parameter values remain fixed throughout the campaign. P OZZER then uses this fixed configuration to process each captured PSC trace and determine whether it represents previously unseen execution behavior.

C

Evaluation machine

In this appendix, we describe the machine used for the fuzzing experiments. The machine is equipped with an Intel Core Ultra 5 225 processor with 10 CPU cores and a maximum dynamic frequency of 4.4 GHz. It has 32 GB of memory and runs an Ubuntu 24.04.4 LTS.

D

A representative example: Nested-if structure

if (input[0] == 'P') if (input[1] == 'O') 3 if (input[2] == 'P') 4 if (input[3] == 'F') 5 if (input[4] == 'U') 6 if (input[5] == 'Z') 7 if (input[6] == 'Z') 8 if (input[7] == '!') 9 crash(); 1 2

Listing 1: Nested-if code structure

E

Coverage plots

In this appendix, we show the remaining coverage plots from Table 1. In total, we had 15 targets in Table 1 reporting results on two different hardware platforms. Therefore, here we present the remaining coverage plots for each platform separately. Figure 9 shows the coverage plots for the target firmwares running on SAM4S, while Figure 10 shows the coverage plots for the target firmware running on STM32F3.

19

Cat firmware coverage

Lwgps firmware coverage Edge-guided

100

POZZER SAM4S

Blind

Edge-guided

50

POZZER SAM4S

Minmea firmware coverage

Blind

50 40 30 20

30 20 10

0

2

4

6

Time (h)

8

10

0

12

0

2

Edge-guided

POZZER SAM4S

6

Time (h)

8

10

60 50 40 30 20

0

12

Blind

Edge-guided

40

POZZER SAM4S

2

4

6

Time (h)

8

10

Edge-guided

POZZER SAM4S

50 40 30 20

30

20

10

10 2

4

6

Time (h)

8

10

0

12

0

2

4

6

Time (h)

8

Cnc firmware coverage Edge-guided

90

POZZER SAM4S

80 70 60 50 40 30 20 0

12

0

2

Edge-guided

4

6

Time (h)

8

10

12

Edge-guided

POZZER SAM4S

Blind

50 40 30 20 10

0

2

4

6

Time (h)

POZZER SAM4S

8

10

0

12

0

2

4

6

Time (h)

8

10

12

Nanomodbus firmware coverage

Blind

Edge-guided

40

60 50 40 30 20

Coverage (%)

Coverage (%)

Coverage (%)

20

POZZER SAM4S

Blind

60

70

50 40 30 20

30

20

10

10

10 0

10

70

30

60

Stepper firmware coverage

Blind

80

40

70

10 0

50

Drone firmware coverage

Blind

Coverage (%)

60

60

Microshell firmware coverage 100

Coverage (%)

Coverage (%)

70

70

0

12

90

80

Blind

80

10

0

Blind

90

POZZER SAM4S

90

Elf firmware coverage

Regex firmware coverage 100

4

Edge-guided

100

10

10

Coverage (%)

Jsonparser firmware coverage

Blind

Coverage (%)

60

40

Coverage (%)

70

0

POZZER SAM4S

70

80

Coverage (%)

Coverage (%)

90

0

Edge-guided

80

0

2

4

6

Time (h)

8

10

0

12

0

2

4

6

8

Time (h)

10

0

12

0

2

4

6

8

Time (h)

10

12

Figure 9: This figure presents coverage plots for different targets running on SAM4S.

Cat firmware coverage

Lwgps firmware coverage Edge-guided

100

POZZER STM32F3

Blind

Edge-guided

50

POZZER STM32F3

Minmea firmware coverage

Blind

50 40 30 20

30 20 10

0

2

4

6

Time (h)

8

10

0

12

0

2

Edge-guided

POZZER STM32F3

6

Time (h)

8

10

60 50 40 30 20

0

12

Blind

Edge-guided

40

POZZER STM32F3

2

4

6

Time (h)

8

10

Edge-guided

POZZER STM32F3

50 40 30 20

30

20

10

10 6

Time (h)

8

10

0

12

0

2

4

6

Time (h)

Cnc firmware coverage Edge-guided

90

POZZER STM32F3

10

80 70 60 50 40 30 20 0

12

Edge-guided

70

20

0

2

4

Time (h)

Edge-guided

50 40 30 20

2

4

6

Time (h)

POZZER STM32F3

0

2

4

6

Time (h)

8

10

12

8

10

30 20

2

4

6

Time (h)

0

2

4

10

Time (h)

POZZER STM32F3

Blind

6

10

30

20

10

8

10

12

0

0

2

4

Time (h)

8

12

Figure 10: This figure presents coverage plots for different targets running on STM32F3.

20

6

20

0

12

Edge-guided

40

40

0

Blind

30

Nanomodbus firmware coverage

Blind

50

0

POZZER STM32F3

40

10

10

12

10 0

Coverage (%)

60

8

50

60

70

0

8

30

60

Stepper firmware coverage

Blind

80

Coverage (%)

4

Coverage (%)

2

40

70

10 0

50

Drone firmware coverage

Blind

Coverage (%)

60

10

60

Microshell firmware coverage 100

Coverage (%)

Coverage (%)

70

6

70

0

12

90

80

Blind

80

10

0

Blind

90

POZZER STM32F3

90

Elf firmware coverage

Regex firmware coverage 100

4

Edge-guided

100

10

10

Coverage (%)

Jsonparser firmware coverage

Blind

Coverage (%)

60

40

Coverage (%)

70

0

POZZER STM32F3

70

80

Coverage (%)

Coverage (%)

90

0

Edge-guided

80

8

12

Record · ID 1028612 · SHA-256 2ecd4afbb8082ff8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.