ClawGuard: Out-of-Band Detection of LLM Agent Workflow Hijacking via EM Side Channel Leo Linqian Gan, Jeffery Wu, Longyuan Ge, Lanqing Yang* , Yonghao Song, Jingkai Zhang, Haojia Jin, Weiyi Wang, and Guangtao Xue*
arXiv:2605.06205v1 [cs.CR] 7 May 2026
Shanghai Jiao Tong University, Shanghai, China Email: {leo-gan, jeffery2019, gly2000, yanglanqing, songyonghao, jinhaojia, weiyi_wang, gt_xue}@sjtu.edu.cn, [email protected]
Abstract—Autonomous Large Language Model (LLM) agents increasingly execute state-changing workflows by invoking files, databases, shells, network services, and external tools. This autonomy introduces a new integrity risk: workflow hijacking, in which an adversary inserts, omits, reorders, or substitutes skill invocations while preserving plausible high-level semantics. Existing defenses rely largely on host-internal telemetry such as audit logs, system calls, provenance graphs, or runtime monitors. Once the agent runtime or host OS is compromised, this telemetry shares the same trust boundary as the attacker and can be forged, suppressed, or blinded. We present ClawGuard, a passive out-of-band monitor that uses electromagnetic (EM) emanations as a software-independent physical channel for LLM-agent workflow auditing. The key observation is that agent skills are seconds-scale, compositional workloads: their mixtures of computation, DRAM traffic, storage bursts, network blocking, and idle intervals produce macroscopic EM envelopes that can be measured by external software-defined radios. ClawGuard converts continuous RF streams into skill-level and fine-window physical evidence through a drift-aware coarse–fine pipeline: dual-band SDR sensing, calibrated carrier selection, 320-dimensional spectral/temporal/cross-receiver features, cyclelocal normalization, temperature detrending, and record-level aggregation. We evaluate ClawGuard on a 7.82 TB main RF corpus containing 12,232 records across 16 benign skills and 22 attack skills, together with a separate new-bands replication corpus. On the production split, ClawGuard achieves AUC = 0.9945 and detects attacks at 100% true-positive rate with 1.16% falsepositive rate. On the survey-selected (80, 800) MHz carrier pair, the same pipeline reaches 88.3% record-vote accuracy and 90.3% attack recall on the surviving attack-class subset. These results show that passive EM sensing can serve as a practical physical consistency check for LLM-agent workflow hijacking under hostsoftware compromise. Index Terms—LLM Agents, Workflow Hijacking, Electromagnetic Side-Channel, Out-of-Band Security
I. I NTRODUCTION As global enterprise investments in Generative AI surge toward a projected $140 billion by 2030 [1], the underlying computing paradigm is undergoing a profound transformation. Large Language Models (LLMs) are evolving from passive conversational interfaces into autonomous, task-driven agents. Frameworks like LangChain [2] and research vehicles like OpenClaw [3] integrate contextual reasoning with the ability to invoke external tools, access local file systems, and execute dynamic code. * Corresponding authors.
Fig. 1: System overview. An adversary injects a malicious sub-skill into an OpenClaw execution on the target host, compromising host-internal telemetry. Consequently, the defender utilizes a co-located, passive SDR to observe far-field RF emanations as a secure out-of-band integrity channel.
However, this autonomy introduces a critical vulnerability: workflow hijacking. As enterprises deploy agents to handle increasingly sensitive financial and infrastructure tasks, the stakes of compromise have escalated. Recent security disclosures and research (e.g., PoisonedRAG [4], ToolHijacker [5], and ObliInjection [6]) demonstrate that adversaries can weaponize prompt injection or poisoned databases to subvert agent logic. Through these vectors, attackers can insert, reorder, or substitute tool invocations—such as injecting a covert data exfiltration branch into a routine database query— while preserving semantically plausible high-level interactions. The conventional approach to securing such workflows relies heavily on host-internal software telemetry. To detect Advanced Persistent Threats (APTs) and workflow deviations, the security community has extensively developed system-callbased provenance graphs (e.g., HOLMES [7], Unicorn [8]) and in-kernel eBPF observability frameworks (e.g., Kobra [9]). However, these software-layer defenses suffer from a fundamental architectural flaw: a symmetric threat model. The monitoring infrastructure shares the exact same physical and logical trust boundary as the execution environment it observes. Security audits, bootloader fuzzing campaigns [10], and controlled preemption attacks [11] consistently demonstrate that adversaries routinely achieve privilege escalation and kernellevel compromise. Once an attacker gains root access, they can effortlessly forge provenance records, blind eBPF sensors, or
suppress telemetry altogether. Software cannot reliably attest to its own integrity when the underlying operating system substrate is actively controlled by an adversary. To secure highstakes autonomous platforms, the defense mechanism must survive full host OS compromise. To break this symmetry, workflow monitoring must be established out-of-band, utilizing an observation channel that remains structurally isolated from the host OS. Electromagnetic (EM) side-channel signals offer such a hardware-rooted trust anchor. Arising unintentionally from the physical electrical activity of the CPU and memory bus, EM emanations reflect the ground-truth execution of the hardware. Crucially, they can be captured completely passively—without any invasive host modifications—and cannot be tampered with by softwarelevel adversaries. Prior EM side-channel literature has primarily focused on bit-level cryptanalysis (e.g., RSA/ECC key extraction [12], [13]) or coarse-grained application fingerprinting [14]. Monitoring an LLM agent requires a distinct, intermediate granularity: the skill level. Our key physical observation is that autonomous agent skills (e.g., database analytics, network streaming, or malicious shell executions) manifest as macroscopic, seconds-long compositional workloads. The structural choices of these skills—specifically their alternating bursts of computation, memory access, and I/O— generate highly separable, long-duration EM envelopes. Motivated by this, we design and implement ClawGuard, a fully out-of-band, physical-layer integrity monitor for LLM agent workflows. Deployed via software-defined radios (SDRs), ClawGuard observes the host non-intrusively. However, leveraging EM signals for semantic workflow monitoring presents a severe technical challenge: bridging the semantic gap. Raw EM signals are noisy, continuous analog streams heavily distorted by hardware non-stationarity and thermal drift. While prior side-channel literature successfully extracts bit-level cryptographic keys [12] or fingerprints coarsegrained, monolithic desktop applications [14], neither granularity suits LLM agents. Extracting high-level, discrete logical intent (e.g., distinguishing a benign file read from a malicious script execution) from continuous RF waves requires defining a completely new abstraction layer that maps macroscopic hardware physics to semantic workflows, all while remaining resilient to the temporal variance of long-running skills. To overcome these challenges, and bridge the semantic gap between continuous analog RF streams and discrete agent intent, we introduce an event-aware coarse–fine windowing pipeline. This architecture carefully handles the substantial temporal variance and thermal drift inherent in long-running workloads. ClawGuard extracts a robust feature representation to recover a physical-layer skill execution trace, which is subsequently validated against the planner’s intended logic using a confusability-weighted edit distance. Our contributions can be summarized as follows:
workflows, establishing a hardware-rooted trust anchor capable of surviving full OS-level compromise. • Robust Out-of-Band Architecture: We design an eventaware, drift-compensated physical-layer monitor (ClawGuard) that bridges the analog-semantic gap without host instrumentation, effectively translating continuous RF streams into discrete workflow integrity verdicts. • Methodological Insight on RF Artifacts: We expose a critical measurement pitfall in EM side-channel evaluations by isolating CPU-governor-driven frequency modulation from true hardware power signatures. Consequently, we provide an OS-agnostic band-selection methodology to prevent systemic evaluation artifacts in future RF research. • Large-Scale Feasibility Validation: We validate the system on a custom 7.82 TB RF corpus spanning 38 skills (including 22 ported attack primitives). The system achieves a production-split AUC of 0.9945 and a 100% true positive rate at 1.16% false positive rate with a median inference latency of 18 ms, demonstrating the practical viability of outof-band workflow auditing. Our contributions can be summarized as follows: • Formalizing Skill-Level Side Channels: We define a mid-
tier physical observation granularity for LLM agents: the skill. This granularity is coarse enough to produce stable macroscopic EM envelopes, yet fine enough to audit workflow integrity beyond whole-application fingerprinting. • Out-of-Band Workflow-Integrity Architecture: We design ClawGuard, a passive, dual-SDR physical-layer monitor that remains outside the host software trust boundary. ClawGuard converts continuous RF streams into driftcompensated skill-level and attack-state evidence through carrier calibration, coarse–fine windowing, and record-level aggregation. • Large-Scale LLM-Agent RF Corpus: We construct a 7.82 TB RF corpus for LLM-agent workflow monitoring, containing 12,232 records across 16 benign skills and 22 attack skills, together with a separate new-bands replication corpus. The corpus captures synchronized dual-band IQ streams, workflow/event metadata, and temperature traces, enabling evaluation of skill-level EM separability, workflowhijacking detection, drift, and carrier transfer. • Methodological Insight on RF Artifacts: We expose a measurement pitfall in EM side-channel evaluation by separating CPU-governor-driven frequency modulation from workload-dependent hardware power signatures. Based on a 1–3000 MHz band survey, we present a measured carrierselection methodology that avoids relying on nominal CPUclock or DRAM-frequency assumptions. • End-to-End Feasibility Validation: We evaluate ClawGuard on the main RF corpus and the new-bands replication corpus. On the production split, ClawGuard achieves AUC = 0.9945 and detects attacks at 100% true-positive rate with 1.16% false-positive rate; on the survey-selected (80, 800) MHz carrier pair, the same pipeline reaches 88.3% record-vote accuracy and 90.3% attack recall.
• Formalizing Skill-Level Side-Channels: We define a new
mid-tier granularity for physical observation that maps macroscopic hardware emanations to semantic LLM agent
2
TABLE I: Granularity of EM observation. ClawGuard targets the middle tier: coarse enough to produce stable physical envelopes, but fine enough to audit agent workflow integrity.
II. P RELIMINARY A. Physical Basis of Workload-Dependent EM Emanations Electromagnetic (EM) emanations are a physical by-product of digital computation. In CMOS systems, instruction execution, cache activity, DRAM transactions, I/O transfers, and clock-distribution logic all induce time-varying current flows. At a coarse abstraction, the dynamic power of a digital subsystem can be written as Pdyn (t) ≈ α(t)CeffV (t)2 f (t),
(1)
where α(t) is the switching-activity factor, Ceff is the effective switched capacitance, V (t) is the supply voltage, and f (t) is the operating frequency. Different workloads alter α(t) through their instruction mix, cache behavior, memory intensity, storage bursts, network blocking, and library calls. On modern platforms, dynamic voltage and frequency scaling (DVFS) may further modulate V (t) and f (t) in response to sustained activity. These current variations couple into PCB traces, powerdelivery networks, clock lines, memory buses, and attached cables, which act as unintended radiators. A passive receiver tuned to carrier band fc therefore observes a noisy projection of the host’s aggregate hardware activity: ! Xr (t; fc ) = Hr ( fc )
∑ E j (t; fc ) + Nr (t),
Level
Physical target
Why insufficient / useful
Bit / operation
Cryptographic rounds, key-dependent branches
Application
Whole programs or user activities
Skill
Seconds-scale tool invocations
Precise but too lowlevel for agent workflow semantics [12], [13] Robust but too coarse to separate tool-level agent behavior [14], [17] Matches the unit whose identity and order determine workflow integrity
the same runtime. Skill-level monitoring sits between these regimes. For a skill s, we view its physical manifestation as a distribution over windowed EM features: s 7→ Ps = Pr[φ (Xt:t+τ ) | s] ,
(3)
where φ (·) extracts spectral, temporal, and cross-receiver statistics from a window of duration τ. Two skills are physically separable when their induced distributions differ more than the nuisance variation introduced by temperature, cycle state, receiver noise, and ambient RF:
(2)
j∈C
where Xr is the IQ stream captured by receiver r, C denotes hardware components such as CPU cores, DRAM, regulators, buses, and I/O controllers, E j is the component-specific emission, Hr is the environment- and placement-dependent channel response, and Nr is receiver and ambient noise. Prior EM side-channel work demonstrates that such emissions can expose cryptographic secrets at bit or operation granularity [12], [13], [15], [16], and can also fingerprint coarse application behavior [14], [17], [18]. ClawGuard targets a different point in this abstraction hierarchy: the skill-level behavior of autonomous LLM agents.
d(Ps , Ps′ ) > ddrift + dnoise .
(4)
Equation 4 is not assumed to hold for all skill pairs. It is an empirical condition that motivates both our band survey and our staged detector design. C. Workflow Hijacking as Skill-Sequence Deviation The security object in an LLM-agent system is not an isolated process, but an ordered workflow. A benign workflow consists of a planner-authorized skill sequence, while a hijacked workflow deviates from that sequence. Recent attacks on RAG systems, tool documentation, and agent execution show that adversaries can induce such deviations while preserving high-level semantic plausibility: PoisonedRAG [4], ObliInjection [6], GRAGPOISON [19], and ToolHijacker [5] all demonstrate ways to redirect agent behavior through poisoned context, retrieval, or tool interfaces. We use workflow hijacking to denote unauthorized changes to the intended skill sequence. The main forms are: 1) Insertion: an unauthorized skill is added; 2) Omission: a required skill, such as validation or logging, is skipped; 3) Substitution: a benign skill is replaced by a different one; 4) Reordering: the same skills execute in a harmful order; 5) Branch injection: a covert malicious branch is added; 6) Parameter manipulation: the skill identity remains but its arguments change;
B. Why Agent Skills Form a Physical Granularity LLM agents do not execute user requests as a single monolithic program. Frameworks such as LangChain [2] and research rigs such as OpenClaw [3] decompose a task into named tool invocations or skills. A document-processing task, for example, may expand into OpenFile→Summarize→ ComposeEmail→SendEmail. Each skill is a bounded, seconds-scale workload that invokes host-side computation, files, databases, shells, network services, or external APIs. The skill abstraction is physically meaningful because it is long enough to produce a stable macroscopic EM envelope, yet semantically fine-grained enough to express workflow integrity. Bit-level side channels are too fine: they capture short cryptographic kernels rather than agent intent. Wholeapplication fingerprinting is too coarse: it cannot distinguish whether an agent performed a benign database query, a file backup, an email send, or a malicious shell utility within
3
64
(a) idle
64
(b) normal-light
64
(c) normal-heavy
64
produce distinct band-aggregated spectrograms. These visual profiles support the central premise of ClawGuard: the skill is not merely a software label, but a workload class with a measurable physical envelope.
(d) attack
48
48
48
32
32
32
32
16
16
16
16
2
1
0
0 0
1
2
3
4
5
0 0
1
2
time (s)
3
4
5
0 0
1
2
time (s)
3
4
5
0
1
2
time (s)
3
4
5
z-score
PSD band index
3
48
Second, the informative carrier bands are deploymentspecific and must be measured rather than assumed. Our original corpus used the (1.8 GHz, 105 MHz) pair, motivated by the nominal CPU-clock harmonic and a presumed DRAM-related band. The later 1–3000 MHz survey (shown in Fig. 3.) shows that the 1.8 GHz channel was partly a governor-driven activityfrequency-modulation artifact, while the (80 MHz, 800 MHz) pair gives a more defensible CPU/DRAM separation under fixed-frequency methodology. Figure 7 summarizes this carrier-selection result. This is why ClawGuard treats band selection as a calibration step for each rig.
time (s) 0
(a)
3
±σ over time mean PSD
z-score
2
(b)
3
±σ over time mean PSD
2
1
(c)
3
±σ over time mean PSD
2
1
1
0
0
0
−1
−1
−1
−2
±σ over time mean PSD
2
1
0 −1 −2
(d)
3
−2
−1
−2 −2
0
10
20
30
40
PSD band index
50
60
0
10
20
30
40
PSD band index
50
60
0
10
20
30
40
PSD band index
50
60
0
10
20
30
40
50
60
PSD band index
Solid red box: bands 30-34 (1.8 GHz CPU-clock harmonic, dominant attack signature). Dashed red box: bands 50-54 (secondary side-band / DRAM coupling). ‘(a) idle’ uses big48 PSD as baseline; only ‘(d) attack’ contains the injected attack pulse.
Fig. 2: Power spectral density (top) and spectrogram (bottom) of two representative skills (db_analytics vs. log_rotate_compress) on the CPU-correlated band. Skill identity manifests as both stationary spectral shape and time–frequency burst patterns.
Third, the limitations bound the theory. Our stress testing across the 16-class big48 experiment and the full 22class campaign (discussed in §VI-D) shows that skill-level EM recognition is not obtained by simply training a larger classifier. Cycle state, thermal drift, small per-class sample sizes, and cross-run distribution shift can dominate class information. These failures motivate the final architecture: ClawGuard uses a drift-aware coarse–fine front-end, performs skill recovery only where the class set is physically separable, and supplements sequence comparison with a focused finewindow attack detector.
E. Out-of-Band Trust Anchor
Fig. 3: Band-aggregated time–frequency fingerprints of the 15skill attack catalog.
Host-based workflow monitors rely on software telemetry: system calls, provenance graphs, audit logs, kernel hooks, or eBPF programs [7]–[9], [20]. These mechanisms are valuable when the OS is trusted, but they share the same trust boundary as the workload they observe. If an attacker escalates to root or kernel privilege, the host can forge logs, suppress events, modify binaries, or blind in-kernel monitors. This is the symmetric-monitoring failure that motivates ClawGuard.
7) Tool-result poisoning: a legitimate tool returns corrupted output to the planner. From ClawGuard’s perspective, these attacks matter when they alter the physical execution trace. The monitor does not infer whether a natural-language plan is morally benign. Instead, it checks whether the hardware activity emitted by the host is consistent with the intended skill-level workflow. This distinction keeps the problem measurable: semantic attacks become observable only through their downstream effect on host execution.
A passive EM receiver changes the trust boundary. It does not execute code on the device under test, does not require host-side instrumentation, and does not consume telemetry generated by the compromised OS. Instead, it measures an external physical projection of hardware activity and compares the recovered trace with an out-of-band intended workflow.
D. Empirical Evidence for Skill-Level EM Separability The above theory predicts that agent skills can be distinguishable when they induce different mixtures of CPU activity, memory traffic, storage I/O, network blocking, and idle intervals. We validate this prediction in three ways. First, empirical spectral analysis reveals that representative skills such as db_analytics and log_rotate_compress differ in both stationary spectral shape and time-frequency burst structure, as shown in Fig. 2. Extending this observation to the attack-skill catalog, we note that different shell-backed and bad-tool-result skills
This guarantee is deliberately narrow. EM monitoring is independent of host-side software, not immune to all physical attacks. A root attacker may change CPU-governor settings or shape workloads to reduce separability; a physical attacker may jam the RF channel or tamper with sensors. These adaptive cases are discussed in §VIII. The threat model in the next section therefore focuses on the setting where the host software may be fully compromised, but the external receiver and policy channel remain outside the attacker’s control.
4
III. P ROBLEM S TATEMENT
D. Scope of Detectability
obtained from a trusted out-of-band policy channel. The host may instead execute a physical workflow
ClawGuard detects workflow deviations that induce measurable changes in the host’s physical execution envelope. It is not a semantic oracle: if a malicious payload is deliberately shaped to be EM-indistinguishable from an authorized skill under the selected receiver bands, passive RF monitoring alone may not separate the two. The security value of ClawGuard is therefore an independent physical consistency check that remains outside the compromised host’s software trust boundary.
W ⋆ = ⟨s⋆1 , s⋆2 , . . . , s⋆m ⟩,
IV. T HREAT M ODEL
A. Workflow Integrity Target Let S denote the skill alphabet exposed by an agent runtime. An intended workflow is an ordered sequence W = ⟨s1 , s2 , . . . , sn ⟩,
si ∈ S ,
(5)
(6)
We consider an enterprise environment where an LLM agent autonomously executes complex tool workflows on a host computer. The fundamental motivation for our threat model is the inherent fragility of symmetric software defenses: once the host is compromised, the very telemetry these defenses rely upon can be fundamentally manipulated from the same vantage point. Attacker Goals and Entry Vectors. We assume a sophisticated adversary whose primary objective is workflow hijacking—the unauthorized insertion, reordering, substitution, or omission of agent skills to achieve malicious ends while maintaining high-level semantic plausibility. The adversary gains initial arbitrary code execution (ACE) on the host via application-layer vectors, such as indirect prompt injection, poisoned RAG databases, or supply-chain compromises within the agent’s toolchain. Post-Exploitation Capabilities. Upon initial compromise, we assume the attacker successfully escalates privileges to the administrative (root) or kernel level. Within this compromised boundary, the adversary exercises full control over the host operating system. Critically, the attacker can actively blind symmetric defenses by manipulating system binaries, intercepting system calls, bypassing in-kernel eBPF monitors, and forging or suppressing audit logs. This ensures the workflow hijacking remains completely invisible to host-internal telemetry. Defender Setup and Trust Anchor. To break this symmetric trust boundary, ClawGuard relocates the monitoring channel entirely out-of-band. The defender deploys a passive software-defined radio (SDR) in close, non-intrusive proximity (e.g., 2 cm) to the target chassis. This monitor relies solely on unintentional electromagnetic (EM) emanations from the CPU and memory bus, requiring absolutely no host-side software, OS instrumentation, or physical hardware modification. Out-of-Band Policy Knowledge. We assume the defender possesses a trusted, out-of-band record of the agent’s intended workflow prior to execution. This expectation, derived from the LLM planner’s original task description, is transmitted via a secure side-channel (e.g., a network-isolated policy store) and serves as the immutable ground-truth policy against which the physical EM execution trace is validated. Environmental Scope and Out-of-Scope Threats. The physical deployment must contend with realistic operational conditions, including ambient electromagnetic interference, dynamic voltage and frequency scaling (DVFS) policies, and natural thermal drift. While the adversary possesses absolute
where m may differ from n due to insertion, omission, branch injection, or other workflow-hijacking behavior. The defender does not trust host-side telemetry to reveal W ⋆ . Instead, the defender observes passive EM traces X = {Xr (t; fr )}Rr=1 ,
(7)
captured by external receivers. The goal is to decide whether the physical execution represented by X is consistent with the intended workflow W . B. Physical Trace Recovery The first task is to map noisy EM observations to skill-level evidence. Given a windowed feature sequence Z = ⟨z1 , z2 , . . . , zq ⟩,
z j = φ (Xt j :t j +τ ),
(8)
the monitor estimates either a skill label or an attack-related state for each window. In the ideal sequence-level setting, these predictions can be aggregated into an estimated physical trace Ŵ = ⟨ŝ1 , ŝ2 , . . . , ŝm̂ ⟩.
(9)
In the present system, we instantiate this abstraction in two complementary ways. Stage 1 performs skill recovery on physically separable skill subsets. Stage 2 performs finewindow attack detection for short malicious sub-skills embedded within longer task envelopes. This design reflects the empirical finding that large open-set multi-class skill recognition is unstable under cycle drift and limited per-class samples. C. Integrity Decision At the sequence level, workflow integrity can be expressed as a comparison between W and Ŵ : ( BENIGN, D(W, Ŵ ) ≤ δ , g(W, Ŵ ) = (10) HIJACKED, D(W, Ŵ ) > δ , where D may be instantiated as a weighted edit distance over skill insertions, deletions, substitutions, and reorderings. However, the quantitative evaluation in this paper focuses on the physical trace-recovery components that are implemented and measured: skill-level classification on separable subsets and fine-window attack detection. The full sequence-level editdistance comparator is part of the system design, but its largescale evaluation over a complete workflow attack catalog is left to future artifact expansion.
5
software control, physical-layer attacks are strictly out of scope. We assume the attacker cannot physically alter the host’s hardware, access the defender’s SDR sensors, or deploy external RF transmitters to actively jam the environment.
computation, DRAM traffic, storage bursts, network blocking, and idle gaps. Carrier choice is treated as a measured calibration problem rather than a hardware-specification assumption. For a candidate carrier f , we estimate the workload-induced deltas
V. ClawGuard D ESIGN
∆CPU ( f ) = Pcpu ( f )−Pidle ( f ),
∆RAM ( f ) = Pram ( f )−Pidle ( f ), (12) where Pidle , Pcpu , and Pram are median band powers under idle, CPU-intensive, and memory-intensive calibration workloads. A useful CPU channel should have high ∆CPU ; a useful memory channel should have high ∆RAM with limited CPU confounding. This calibration step explains the relationship between the original corpus and the new-bands replication. The original experiments used (1.8 GHz, 105 MHz), motivated by the CPUclock harmonic and a presumed memory-related band. The later 1–3000 MHz survey in §VI-B shows that the 1.8 GHz channel partly captures governor-driven activity-frequency modulation. The measured (80 MHz, 800 MHz) pair, summarized in Figure 7, provides a more defensible CPU/DRAM separation under fixed-frequency methodology. The rest of ClawGuard’s pipeline is independent of the specific carrier pair; the carrier pair is an instantiation parameter. In the evaluated prototype, both SDRs sample at 20 MS/s with 8-bit IQ. The original corpus uses the legacy (1.8 GHz, 105 MHz) pair, while the replication corpus uses the survey-selected (80 MHz, 800 MHz) pair. Temperature is sampled in parallel and used only for drift compensation, not as a security label.
A. Design Overview ClawGuard is a passive RF monitor that turns electromagnetic emanations into workflow-integrity evidence for LLMagent executions. Its input is a policy-approved workflow envelope and the raw IQ streams captured by external softwaredefined radios (SDRs). Its output is a record-level verdict indicating whether the observed physical execution is consistent with benign skill execution or contains attack-like activity. The design is intentionally staged. Raw EM streams are noisy, continuous, and strongly affected by temperature, cycle state, and carrier selection. Directly classifying an entire RF recording into a workflow verdict would mix four separate problems: sensing, temporal alignment, drift compensation, and security decision. ClawGuard therefore decomposes the task into the pipeline shown in Figure 4: X −→ {Fi, j } −→ {xi, j } −→ {pi, j , ai, j } −→ v.
(11)
Here, X is the multi-receiver EM trace, Fi, j is a fine window inside the i-th skill envelope, xi, j is its physical feature vector, pi, j is skill-level evidence, ai, j is attack-state evidence, and v ∈ {BENIGN, HIJACKED} is the final verdict. This pipeline corresponds directly to the claims evaluated in §VI. Stage 1 asks whether skill-level EM envelopes are separable on controlled skill subsets. Stage 2 asks whether short malicious activities embedded in longer task envelopes can be detected from fine-window RF evidence. The production operating curve in §VI-C is therefore a record-level attack-detection result, not a claim that ClawGuard has fully evaluated arbitrary sequence-level workflow verification. This distinction is important: the current system provides a physical consistency check for workflow hijacking, instantiated through skill recovery and fine-window attack detection. The design is guided by three constraints. First, the monitor must remain outside the host software trust boundary: it cannot rely on system calls, audit logs, eBPF events, or runtime self-reports. Second, the monitored unit must be the agent skill rather than a bit-level cryptographic operation or a whole application. Empirical profiling motivates this choice: different skills produce different macroscopic time-frequency envelopes. Third, the inference pipeline must be drift-aware, because §VI-D shows that cycle and thermal variation can dominate class information if left untreated.
C. Coarse–Fine Physical Evidence Extraction The central representation in ClawGuard is not a whole RF recording, but a set of fine-window physical evidence vectors aligned to a coarse skill envelope. A coarse window captures the expected wall-clock duration of one skill invocation. Inside it, ClawGuard extracts overlapping fine windows: Fi, j = [ai + jρ, ai + jρ + τ],
(13)
where [ai , bi ] is the coarse envelope of skill i, τ is the finewindow length, and ρ is the stride. This design is motivated by the attack structure. A malicious payload may occupy only a short interval inside a much longer task envelope. Whole-record features dilute such activity, while very short global sliding windows lose workflow context. Coarse–fine windowing preserves both: the coarse envelope keeps the skill-level structure, and fine windows expose localized attack bursts. Figure 5 illustrates this mechanism. Each fine window is transformed into a 320-dimensional V10 feature vector. The feature groups are deliberately statistical rather than instruction-level:
B. Passive RF Sensing and Carrier Calibration ClawGuard observes the host through two passive SDR receivers. The intended role of the two channels is complementary: one channel should emphasize CPU-correlated activity, while the other should emphasize memory-correlated activity. This dual view is necessary because agent skills differ not only in average compute load, but also in their mixture of
• Spectral
energy: log-PSD band energies that capture stationary workload power differences; • Spectral shape: centroid, bandwidth, flatness, rolloff, and spectral moments that capture carrier-local structure;
6
Fig. 4: Overview of ClawGuard. Passive dual-band SDR capture feeds a coarse–fine windowing front-end. Each fine window is transformed into a 320-d physical feature vector, compensated for drift, and classified into skill-level or attack-state evidence. Coarse-window aggregation produces the record-level workflow-integrity verdict evaluated in §VI.
and k = 80 for larger settings. These parameters match the evaluation protocol in §VI-A. For supervised training and offline evaluation, the OpenClaw rig provides event metadata that identifies the active WORK interval inside each coarse record. These events are used as labels and alignment metadata for the physical corpus. They are not treated as trusted runtime telemetry. This choice keeps the evaluated claim precise: the prototype measures whether properly localized skill and attack intervals are physically separable in RF, while the host-independent deployment path relies on the out-of-band workflow schedule rather than host logs.
F11. Small-window attack detection (iter-8, attack_pilot 30 records, LOCO-10) spec_seq @ 50 ms x 64 PSD band -> K small windows -> V10 stat / win -> RF -> pool (a) K x pool grid macro-F1 (b) smallwin variants vs baseline mean-pool dominates; max-pool fails (attack is distributed not bursty) * K=16 mean-pool f1 = 0.683 / acc = 0.733 (+0.013 / +0.033)
(c) ROC @ K=4 (best AUC config) mean-pool dominates; signal is connected not sparse 1.0
0.8 0.75
0.525
0.653
0.733
0.612
0.7
0.60 K8 (K = 8 small win)
0.433
0.631
0.554 0.55
0.50
0.45 K16 (K = 16 small win)
0.495
0.683
macro-F1
0.65
score (LOCO-10, 30 records)
0.70
0.700 0.670
0.700 0.653
0.700
0.682
0.631
0.8
0.633 0.612
0.6 0.533 0.525
0.5
0.4
0.3
0.2
top-k pool
0.4
K=4 mean (AUC = 0.694) K=4 topk (AUC = 0.684) K=4 max (AUC = 0.660) chance
baseline f1=0.670
0.1
random 0.5 macro-F1 acc
mean-pool
0.6
0.2
0.598 0.40
max-pool
True Positive Rate
K4 (K = 4 small win)
0.0 baselinesmallwin K=4 smallwin K=8 smallwin K=16 smallwin K=4 smallwin K=4 V10 single mean mean mean * topk max (320-d)
0.0 0.0
0.2
0.4
0.6
0.8
1.0
False Positive Rate
Fig. 5: Coarse–fine windowing. The coarse envelope preserves the skill-level task structure, while overlapping fine windows expose short attack payloads that would be diluted by wholerecord aggregation.
D. Inference and Record-Level Verdict ClawGuard uses the same drift-compensated feature representation for two measured inference tasks.
• Temporal envelope: amplitude moments, short-lag auto-
correlation, and burst statistics that capture the timing of computation and I/O; • Cross-receiver coupling: correlation and phase-derived statistics between the two SDR channels that capture CPU/DRAM co-variation. Let xi, j ∈ R320 be the raw feature vector extracted from Fi, j . Before inference, ClawGuard applies three leakage-controlled preprocessing steps. First, cycle-local normalization removes coarse amplitude offsets: x̃c,k =
xc,k − µc,k . σc,k + ε
Stage 1: skill evidence. For skill-recovery experiments, a classifier maps each fine-window feature vector to a posterior distribution over a skill set S1 : pi, j (s) = Pr(s | x′i, j ),
(16)
Fine-window posteriors are pooled inside the coarse window: p̄i (s) = Pool j pi, j (s),
ŝi = arg max p̄i (s). s
(17)
This stage corresponds to the focused3 and pairwise/big48 results in §VI-B and §VI-D. Its purpose is to test the physical premise that some agent skills have recoverable EM envelopes. It is not presented as a solved open-set 22-class recognition problem; our stress campaigns explicitly show the limits of that setting.
(14)
Second, temperature detrending removes low-order thermal bias: d
xk′ = x̃k − ∑ βℓ,k T ℓ .
s ∈ S1 .
(15)
ℓ=0
Stage 2: attack-state evidence. For workflow-hijacking detection, ClawGuard uses a fine-window state detector
Third, ANOVA feature selection keeps the top-k discriminative dimensions on the training fold only. The evaluated configurations use d = 1, k = 65 for the focused three-skill task,
ai, j = hψ (x′i, j ) ∈ {background, normal, attack}. (18)
7
A coarse record is flagged when the aggregation of its finewindow states indicates attack-like physical activity: ( HIJACKED, A({ai, j } j ) > η, (19) vi = BENIGN, A({ai, j } j ) ≤ η, where A(·) is the record-level aggregation rule and η is the operating threshold. In the evaluated prototype, A is implemented by majority voting or score aggregation over fine windows, yielding the record-level accuracy, recall, ROC, and PR results reported in §VI-C. Both stages are implemented with balanced random forests in the headline prototype. The default model uses 500 trees, maximum depth 15, and minimum leaf size 2. This choice is empirical rather than architectural: tree ensembles are stable under small physical datasets, heterogeneous feature groups, and non-Gaussian drift. Table II summarizes how the design components map to the evaluation sections. This alignment is intentional: every mechanism in the design is either evaluated directly or identified as an implementation parameter, avoiding unmeasured sequence-verifier claims.
Raspberry Pi target
Laptop target
Fig. 6: Prototype deployments. ClawGuard passively monitors real hosts using external SDRs. The Laptop deployment supports the headline and new-bands results, heavily emphasizing realistic enterprise scenarios; the Pi deployment is used as an additional stress setting.
as the prominent platform, and a Raspberry Pi agent host for an additional portability stress campaign. These deployments exercise the full sensing path: passive RF capture, temperature logging, feature extraction, model inference, and record-level verdict generation. b) Corpora.: Table IV summarizes the RF corpora. The main corpus contains 12,232 records and approximately 7.82 TB of raw IQ, covering 16 benign skills and 22 attack skills. The new-bands corpus is a separate re-collection using the measured (80, 800) MHz carrier pair. It runs for 98 minutes over 10 cycles and produces 550 IQ files (276 CPU-channel and 274 RAM-channel files), totaling 440 GB of raw IQ. The session also records 221 per-skill event files and 11,114 temperature samples. The Pi-side openclaw_attack_v1 corpus is used only as a stress campaign because it uses a third carrier pair and has small per-class support. c) Evaluation protocol.: All reported classification results use leave-one-cycle-out (LOCO) cross-validation unless stated otherwise. LOCO is stricter than random splitting because recordings in the same cycle share temperature, background RF, and collection-order effects. Cycle-local normalization, temperature detrending, ANOVA feature selection, and scaling are fit only on the training fold. This prevents collection-time leakage from inflating results.
VI. E VALUATION We evaluate ClawGuard as a complete physical monitoring system for LLM-agent workflow integrity. The evaluation is organized around four claims that match the system design in §V and the paper’s central claim in the introduction. • C1: Implementability. ClawGuard can be deployed as a passive dual-SDR monitor on real agent hosts and can collect synchronized RF, temperature, and workflow records at scale. • C2: Physical validity. Agent skills induce measurable macroscopic EM envelopes, and informative RF bands can be selected by measurement rather than assumed from hardware specifications. • C3: Detection efficacy. Fine-window physical evidence can detect workflow-hijacking activity at the record level with low false positives. • C4: Robustness and practicality. The pipeline remains useful under carrier re-selection, drift-aware evaluation, and realistic runtime constraints. Table III maps these claims to the corresponding datasets, figures, and measurements. This structure is deliberate: every result in the evaluation validates one part of the implemented system, and we avoid claiming an unevaluated arbitrary sequence-level verifier.
B. Physical Validity of EM Evidence This subsection validates the physical premise behind ClawGuard: agent skills must produce measurable EM structure, and the receiver bands must be selected from measurements rather than assumptions. a) Skill envelopes are visible.: Empirical spectral analysis confirms that representative skills such as db_analytics and log_rotate_compress differ in both stationary spectral shape and time-frequency burst structure. Extending the observation to the attack-skill catalog, different attack primitives produce distinct band-aggregated spectrograms. These profiles support the paper’s central
A. End-to-End Prototype and Corpus a) Prototype.: We implement ClawGuard with two HackRF One SDRs, each sampling at 20 MS/s with 8-bit IQ. The radios are placed near the target host without any electrical or software connection to it. A temperature sensor records thermal state for drift compensation. Figure 6 shows the two physical deployments used in the evaluation: a laptop target for the headline and new-bands corpora, establishing it
8
TABLE II: Alignment between ClawGuard’s design components and the evaluation. The method is written around the mechanisms that are quantitatively evaluated in the paper. Design component
Prototype instantiation
Evaluation support
Carrier calibration
Legacy (1.8 GHz, 105 MHz) and surveyed (80 MHz, 800 MHz) 20 s coarse envelope; 0.5 s fine windows with 0.25 s stride in the original pipeline 320-d dual-receiver spectral, temporal, and crosschannel statistics Cycle-local normalization, degree-1 temperature detrending, ANOVA top-k selection Random-forest posterior pooling over fine windows
Band survey and new-bands replication (§VI-B, §VI-C) Stage-2 record-level detection and case studies (§VI-C, §VI-D) Feature attribution and drift diagnosis (§VI-D)
Coarse–fine windowing V10 physical features Drift compensation Skill evidence Attack-state evidence
Fine-window {background, normal, attack} detector with record-level aggregation
LOCO protocol and ablations (§VI-A, §VI-D) focused3 and big48/pairwise skill recovery (§VI-B, §VI-D) Production ROC/PR and new-bands attack recall (§VI-C)
TABLE III: Evaluation roadmap. Each block validates a concrete system claim and corresponds to the abstractions introduced in §V. Claim
What is validated
Evidence
C1: Implementability C2: Physical validity
Passive RF prototype, synchronized corpus, real host deployment Carrier selection and skill-level EM separability
C3: Detection efficacy
Record-level workflow-hijacking detection
C4: Robustness/practicality
Cross-carrier replication, drift analysis, calibration, latency
Hardware stack, main 7.82 TB corpus, new-bands 440 GB session Band survey, PSD/spectrograms, pairwise skill separability Production ROC/PR, coarse–fine attack detection, case studies New-bands accuracy, cycle-leakage diagnosis, calibration, runtime breakdown
TABLE IV: Evaluation corpora. The main corpus supports the headline results. The new-bands corpus validates carrier reselection. The Raspberry Pi campaign is used as a boundary stress test, not as a headline accuracy source. Dataset
Host
Skills
Records / files
Trace
Carrier pair
focused3 big48 attack_skills_pilot Main corpus New-bands corpus openclaw_attack_v1
Laptop Laptop Laptop Laptop Laptop Pi
3 benign 16 benign 22 attacks 16 + 22 22 attacks + idle 22 attacks
123 records 1,529 records 299 records 12,232 records 550 IQ files 155 records
8s 8s 20 s mixed 20 s 20 s
1.8 GHz, 105 MHz 1.8 GHz, 105 MHz 1.8 GHz, 105 MHz 1.8 GHz, 105 MHz 80 MHz, 800 MHz 248 MHz, 800 MHz
physical claim that skill executions are not merely software labels; they induce macroscopic RF envelopes. b) Carrier selection must be measured.: We sweep 1– 3000 MHz under idle, CPU-heavy, and RAM-heavy conditions. For each candidate carrier f , we compute
Figure 8 shows steady progress across the 98-minute run; Figure 9 shows balanced per-skill record counts across channels; Figure 10 shows stable per-skill IQ amplitude. The CPU channel sits around 1.5–3.0 on the signed 8-bit scale, while the RAM channel sits around 3.0–4.5, with small within-skill variance for most skills. This validates the RF collection chain ∆CPU ( f ) = Pcpu ( f )−Pidle ( f ), ∆RAM ( f ) = Pram ( f )−Pidle ( f ). as a repeatable measurement instrument. The original corpus used (1.8 GHz, 105 MHz), motivated by d) Skill separability is structured.: On focused3, the nominal CPU harmonic and a presumed memory-related stage 1 reaches macro-F1 = 0.894 ± 0.009 across five seeds band. The survey shows that this choice is not robust un- under LOCO, with the worst fold above 0.878. This confirms der fixed-frequency methodology: the 1.8 GHz region partly that some skill groups are recoverable from RF evidence. The captures governor-driven activity-frequency modulation rather broader big48 catalog is more heterogeneous: a flat 16-class than a stable idle-vs-busy power signature. The measured classifier reaches only macro-F1 = 0.146. However, pairwise (80, 800) MHz pair is more defensible: 80 MHz is CPU- analysis reveals strong structure. Figure 11 shows that 12 of correlated, and 800 MHz aligns with LPDDR4 activity. Figure 120 pairwise classifiers exceed 0.80 F1 , and the most separable 7 summarizes the resulting CPU/RAM separation map. skill, build_release_pipeline, exceeds 0.94 against c) New-bands collection is stable.: After selecting several operational skills. This result justifies the design choice (80, 800) MHz, we re-collected a complete attack benchmark. in §V: use skill evidence where physical separability is strong,
9
Records per skill, by HackRF channel (full session)
Per-window CPU/RAM separation (v2 survey, 20 MHz windows, 5 MHz step) 8
attack_bad_tool_result_curl
RAM-only
attack_bad_tool_result_dig attack_bad_tool_result_mysql attack_bad_tool_result_nginx attack_bad_tool_result_redis
6
attack_bad_tool_result_ssh
2500
attack_bad_tool_result_wget_bash attack_bad_tool_result_wget_jpeg attack_bad_tool_result_wget_pdf attack_bad_tool_result_wget_py
4
attack_third_party_cat attack_third_party_chmod
ΔRAM [dB]
2
1500
0
−2 1000
attack_third_party_chown
Window center [MHz]
2000
attack_third_party_cp attack_third_party_date attack_third_party_ln attack_third_party_ls attack_third_party_mkdir attack_third_party_mv attack_third_party_rm attack_third_party_stat
0.0
−8 −8
−6
−4
−2
5.0
7.5
10.0
12.5
15.0
17.5
Fig. 9: Per-skill record counts in the new-bands corpus. Counts are balanced across SDR channels, with only minor one-record differences at shutdown.
500
Selected CPU 80 MHz Selected RAM 800 MHz Old CPU 1.8 GHz (rejected) Old RAM 105 MHz (rejected)
2.5
IQ records
−4
−6
hackrf1 (CPU 80 MHz) hackrf2 (RAM 800 MHz)
sim_idle_baseline
CPU-only 0
2
4
6
Per-skill IQ amplitude distribution (mean ± 1σ across cycles)
8
ΔCPU [dB]
attack_bad_tool_result_curl attack_bad_tool_result_dig attack_bad_tool_result_mysql attack_bad_tool_result_nginx
Fig. 7: Carrier-separation map from the 1–3000 MHz survey. The selected (80, 800) MHz pair provides a more defensible CPU/RAM split than the legacy (1.8 GHz, 105 MHz) pair.
attack_bad_tool_result_redis attack_bad_tool_result_ssh attack_bad_tool_result_wget_bash attack_bad_tool_result_wget_jpeg attack_bad_tool_result_wget_pdf attack_bad_tool_result_wget_py attack_third_party_cat attack_third_party_chmod attack_third_party_chown attack_third_party_cp attack_third_party_date
Benchmark collection progress: 10 cycles, 550 IQ files, 440 GB
attack_third_party_ln attack_third_party_ls attack_third_party_mkdir
500
attack_third_party_mv
Cumulative IQ files written
attack_third_party_rm attack_third_party_stat
CPU 80 MHz RAM 800 MHz
sim_idle_baseline
400
0.0
0.5
1.0
1.5
2.0
2.5
3.0
3.5
4.0
⟨|I, Q|⟩ over 2 KB sample (8-bit signed)
300
Fig. 10: Per-skill IQ amplitude in the new-bands corpus. The two receiver channels provide distinct and stable amplitude regimes across the session.
200
100
0
0
20
40
60
80
Wall time [min]
0.861. The result supports the coarse–fine design: short malicious payloads that are diluted in a whole-record representation become visible when the record is decomposed into fine physical windows. c) Why anomaly detection is insufficient.: One-class anomaly baselines do not solve the task. IsolationForest, OneClassSVM, and Mahalanobis-style scoring remain near chance in the combined setting. The reason is that “attack” is not equivalent to density anomaly: benign skills are diverse, and many attack windows lie near benign workload manifolds. Supervised fine-window attack evidence is therefore necessary. d) Cross-carrier detection.: The new-bands corpus tests whether the detection pipeline survives carrier re-selection. On the (80, 800) MHz corpus, the same V10 pipeline reaches 83.6% sub-window accuracy and 88.3% record-vote accuracy. Attack recall is 88.1% at the sub-window level and 90.3% at the record level (Table VI). This result is lower than the legacycorpus AUC, but it is more deployment-relevant because the carrier pair is selected by measurement rather than by the governor-sensitive 1.8 GHz channel. Figure 13 shows that most new-bands errors are background-as-attack, not attack-as-background. This is the conservative failure direction for an integrity monitor: the
Fig. 8: New-bands collection progress. The 10-cycle (80, 800) MHz session runs steadily for 98 minutes.
and use fine-window attack evidence for localized hijacking behavior. C. Workflow-Hijacking Detection This subsection evaluates the main security claim: ClawGuard detects workflow-hijacking activity from passive physical evidence. a) Production operating curve.: On a production split of 11,800 records (1,650 attack and 10,150 normal), ClawGuard achieves ROC-AUC = 0.9945 and PR-AUC = 0.9305. At the selected operating point, it reaches 100% true-positive rate at 1.16% false-positive rate. Even under the more conservative region FPR ≤ 1%, TPR remains approximately 0.99. Figure 12 shows the curve. b) Coarse–fine detection.: Table V evaluates the finewindow attack detector used by the record-level pipeline. Across four configurations, record-level accuracy is between 0.9252 and 0.9398, and attack recall is between 0.833 and
10
a
0.86 0.94 0.96 0.74 0.95 0.78 0.82 0.91 0.87 0.91 0.82 0.73 0.85 0.83 0.80 0.86
0.96 0.78 0.45
0.76 0.63 0.48 0.60 0.60 0.68 0.55 0.51 0.63 0.59
0.78 0.71 0.70 0.69 0.43 0.76
git_dev_workflow
0.66 0.65 0.60 0.49 0.55 0.65 0.53 0.53
0.91 0.77 0.51 0.52 0.76 0.48 0.69 0.66
simple_email
0.54 0.48 0.62 0.52 0.63 0.62
0.91 0.67 0.61 0.48 0.74 0.60 0.67 0.60 0.54 0.54
calculator_develop
0.62 0.54 0.49 0.50
0.73 0.73 0.71 0.59 0.52 0.55 0.52 0.55 0.59 0.62 0.57 0.62
clock_develop
0.5
20
15
10
0.56 0.50 0.49
0.85 0.72 0.55 0.57 0.69 0.51 0.58 0.65 0.45 0.52 0.51 0.54 0.56
0.61 0.53
0.83 0.68 0.68 0.67 0.44 0.63 0.56 0.53 0.57 0.63 0.50 0.49 0.50 0.61
local_wiki_search
0.6
0.57 0.57 0.51 0.50 0.54
0.82 0.62 0.74 0.67 0.49 0.68 0.54 0.49 0.60 0.48 0.57
package_install_update
0.7
0.51 0.54 0.60 0.59 0.45 0.57 0.61
0.87 0.61 0.66 0.63 0.60 0.60 0.60 0.65 0.51
file_sync_backup
0.8
0.53 0.69 0.60 0.67 0.54 0.52 0.58 0.56 0.55
0.82 0.66 0.68 0.65 0.63 0.63 0.53
browser_session_heavy
0.9
0.78 0.43 0.63 0.76 0.60 0.74 0.49 0.52 0.69 0.44 0.51
0.95 0.77 0.54 0.54 0.78
wiki_email
median = 0.615
0.81 0.54 0.69 0.65 0.52 0.63 0.48 0.67 0.59 0.57 0.67 0.69
0.74 0.73 0.78 0.81
db_analytics
31% pairs < chance chance ≈ 0.55
1.0
0.45 0.78 0.54 0.70 0.68 0.51 0.66 0.61 0.74 0.71 0.55 0.68 0.58
Pairwise binary macro-F1
video_streaming sensor_polling_iot
25
0.78 0.78 0.73 0.77 0.71 0.66 0.77 0.61 0.67 0.62 0.73 0.72 0.68 0.73
0.94 0.78
log_rotate_compress
Number of pairs
background system_maintenance
5
0.4
0.80 0.73 0.58 0.69 0.51 0.59 0.55 0.53 0.61 0.62 0.54 0.50 0.49 0.53 0.50
Sub-win acc.
Record acc.
bg recall
atk recall
v1 v2 v3 no-temp
0.816 0.787 0.809 0.809
0.9398 0.9252 0.9398 0.9398
0.250 0.139 0.222 0.222
0.833 0.844 0.861 0.861
TABLE VI: LOCO performance on the new-bands (80, 800) MHz corpus. The result is computed on the 14 attack classes that survived event-anchor filtering.
0 0.6
_r ele as e_ p ste b ipeli n m ac log _m kgro e a u _r ota inte nd te_ na vid com nce eo pr _ se e ns stre ss or _p amin oll g db ing_ _a iot na w lytic g br s ow it_d iki_ e e se r_ v_w mail se ss orkflo ion w _ h s file imple eavy _ _ ca sync em a lcu _ lat bac il or pa _d kup c ck e ag lock velo e_ _ ins dev p loc tall elo al_ _up p wik d i_s ate ea rch
ild
Config
0.50
0.8
Pairwise macro-F1
sy
bu
TABLE V: Coarse–fine record-level attack detection on big48_chunk1 + attack_skills_pilot_20260423b.
(bb) Distribution of 240 pairwise scores
(a) 16-class pairwise separability (120 pairs)
build_release_pipeline
Fig. 11: Pairwise skill separability on big48. Skill-level EM evidence is structured: some pairs are highly separable, while broad open-set multi-class recognition remains difficult.
Overall accuracy Background recall Attack recall
Sub-window
Record vote
83.6% 75.0% 88.1%
88.3% 70.0% 90.3%
Figure 23. Production operating curve — combined 1797 records (251 atk + 1546 normal) AUC=0.9945 AP=0.9305 → TPR=100% @ FPR=1.16%, threshold=0.20
(a) ROC (log-x for low-FPR detail) 1.0
0.8
0.6
Precision
True Positive Rate
0.8
0.4
0.6
0.4
AUC = 0.9945 chance Youden TPR=1.000 FPR=0.012 thr=0.201
0.2
0.2
AP = 0.9305 baseline (atk-rate 0.140)
0.0 10
b) Feature attribution explains the legacy result.: Figure 15 shows that the legacy model relies heavily on bands 30–34, corresponding to the 1.8 GHz harmonic region. This explains the strong original-corpus AUC, but also reveals why the new-bands replication is necessary: the legacy channel partly reflects governor-driven activity-frequency modulation. Spectral masking confirms that the model is using RF structure rather than random leakage; masking the dominant cluster degrades AUC substantially. c) Probability outputs are stable.: The raw randomforest probabilities are already well calibrated, with Brier score 0.011 and ECE 0.009. Post-hoc isotonic recalibration degrades the metrics, as shown in Figure 16. The production operating point therefore uses the raw model scores. d) Transfer stress test.: The Pi-side openclaw_attack_v1 campaign runs the 22-class attack catalog on a separate host and a third carrier pair (248, 800) MHz. We do not use this campaign as a headline accuracy result because it has only 155 records across 22 classes and mixes host, run, and carrier shifts. Instead, it is a boundary test. It shows that naive single-shot 22-class recognition is statistically underpowered at small per-class sample sizes, and that cross-run transfer can collapse under drift. This finding supports the staged design: ClawGuard should not be evaluated as a universal 22-class classifier, but as a physical consistency monitor built from separable skill evidence and fine-window attack evidence. e) Runtime.: ClawGuard’s post-feature inference latency is small relative to skill duration. Median per-record inference latency is 18 ms, and p99 is 29 ms (Figure 17). Batched prediction amortizes to approximately 0.15 ms per record. Since agent skills are seconds-scale and fine windows are hundreds of milliseconds long, sensing and windowing dominate wallclock delay; model inference is not the bottleneck. f) Evaluation takeaways.: The evaluation supports the following claims. First, ClawGuard is implemented as a real passive RF monitor, not a simulation. Second, measured carrier
(b) Precision-Recall
1.0
0.0 −4
10
−3
10
−2
10
−1
10
0
0.0
0.2
0.4
0.944 (thr=0.71)
1.000 (thr=0.20)
1.000 (thr=0.01)
0.8
1.0
(d) TPR/FPR vs threshold Production sweet-spot ≈ 0.20
(c) TPR @ FPR operating points 1.0
0.6
Recall
False Positive Rate (log scale)
1.0
1.000 (thr=0.00)
0.8
0.6 0.6
TPR FPR Youden thr=0.201
Rate
True-positive rate
0.8
0.446 (thr=0.95)
0.4
0.4
0.2
0.2 0.024 (thr=1.00)
0.0
0.0 ≤0.1%
≤0.5%
≤1.0%
≤2.0%
≤5.0%
≤10.0%
Target false-positive rate
0.0
0.2
0.4
0.6
0.8
1.0
Threshold
Fig. 12: Production operating curve on the 11,800-record split. ClawGuard achieves AUC = 0.9945 and detects attacks at 100% TPR with 1.16% FPR.
system is more likely to raise a false alert than to silently miss an attack. D. Robustness, Transfer, and Cost This final subsection explains why the reported results should be read as a measured physical system rather than a lucky split, and it states the known boundaries of the current prototype. a) Drift is measured and controlled.: Cycle and thermal drift are substantial. On big48, the average mutual information between V10 features and cycle index is roughly 18× larger than the mutual information with skill class. Figure 14 shows this effect. This is why all headline experiments use LOCO rather than random splits, and why the pipeline includes cycle-local normalization, temperature detrending, and training-fold-only feature selection.
11
Final LogReg 3-session sub-window confusion (LOCO-20)
395 (84%)
background
33 (7%)
44 (9%)
Final LogReg 3-session record-vote confusion (LOCO-20)
1.0
77 (88%)
background
0.8
8 (9%)
Figure 26. Attack-binary feature attribution — narrow bands + level statistics only 1.0
(a) RF feature importance heatmap (top-30 features) bright = important
3 (3%)
0.8
(b) Per-statistic aggregated importance band 30-34 (89%) band 51-54 (8%)
0.06 0.13
mean std
58 (95%)
0 (0%)
30 (10%)
attack
0 (0%)
0.4
0.2
273 (90%)
8 (9%)
attack
0 (0%)
max
0.06 0.08
min
0.08
p25
0.08
0.06
p75
0.09 0.10
0.04
Statistic
3 (5%)
normal
0.4
0.2
85 (91%)
normal Predicted
0.0
attack
background
normal Predicted
0.0
attack
0.000 0.175
max 0.08
slope
background
std
0.10
Recall
1 (0%)
True
229 (94%)
Recall
14 (6%)
normal
0.6
RF importance
True
0.6
0.251
mean
0.12
0.234
min 0.117
p25
0.223
p75 slope
0.000
peak_rate
0.000
0.02
peak_rate 0.00 0
4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
0.00
0.05
Frequency band index (0..63)
(a) Sub-window
(b) Record vote
0.47
Total RF importance
Fig. 13: New-bands confusion matrices. Most errors are conservative: background windows are sometimes flagged as attack, rather than attacks being missed.
b30 0.26
0.3 0.2
b32 0.12
0.1 0.0 4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
Frequency band index F9 Cycle 泄露诊断: feature 的信息归属 class vs cycle 每特征算 MI, 对比哪些特征是 'class-relevant' vs 'cycle-confounded'
cycle-MI > class-MI (87.5%)
50
class-MI > cycle-MI (12.5%) ★
0.8
gap = 0
feat #311 feat #318 feat #312 feat #297 feat #319 feat #296 feat #306 feat #251 feat #310 feat #300 feat #47 feat #295 feat #292 feat #145 feat #173 feat #171 feat #244 feat #197 feat #294 feat #147
mean gap = -0.234
对角线 y=x
0.6
# features
MI(feat ; class)
40
0.4
0.2
0.0 0.0
30
20
10
0.2
0.4
MI(feat ; cycle)
0.6
0.8
0
-0.8
-0.6
-0.4
-0.2
0.0
0.2
0.000
0.20
0.25
Fig. 15: Band-level feature attribution on the original corpus. The concentration near the 1.8 GHz harmonic motivates measured carrier re-selection and the new-bands replication.
(c) Top-20 'class-favoring' 特征 (按 gap 排序) 这些是 RF top-k 真正应该选的; ANOVA F-test 大致捕到一致子集
Figure 33. RF probability calibration — raw RF Brier=0.0107, ECE=0.0086 (already well-calibrated) Isotonic recalibration HURTS (ECE +0.0191); production uses raw RF directly MI(feat ; class) MI(feat ; cycle)
0.025
0.050
0.075
0.100
gap = MI(class) − MI(cycle)
0.125
0.150
0.175
(a) Reliability diagram — raw RF
0.200
1.0
MI
big48 1529 records, 16 classes, 48 cycles, V10 320-d. MI estimator: sklearn mutual_info_classif (k=5 NN). MI(class) mean=0.0138, max=0.2124; MI(cycle) mean=0.2475, max=0.9170.
Empirical fraction positive
Fig. 14: Cycle-leakage diagnosis. Many RF features are more predictive of collection cycle than of skill identity, motivating LOCO evaluation and drift compensation.
(b) Reliability diagram — isotonic-recalibrated (worse) 1.0
perfect raw RF (ECE=0.0086)
0.8
Empirical fraction positive
(b) Gap 直方图 整体 gap mean = −0.234 → 特征对 cycle 信息 18× 多于对 class → 16-class 天花板 0.14 是 intrinsic 特征瓶颈, 非 模型瓶颈
0.15
Top-15 features (RF importance) ──────────────────────────────── # band stat rf_imp 1 33 mean 0.1338 ★ 2 33 p75 0.0994 ★ 3 30 p75 0.0861 ★ 4 33 max 0.0793 ★ 5 33 min 0.0777 ★ 6 33 p25 0.0767 ★ 7 30 mean 0.0638 ★ 8 30 max 0.0568 ★ 9 30 min 0.0534 ★ 10 32 min 0.0301 11 32 p75 0.0288 12 32 mean 0.0269 13 32 p25 0.0224 14 31 max 0.0171 15 53 min 0.0165
0.4
0
(a) 每特征 MI 散点 (320 维 V10) 对角线下方 = 特征被 cycle 漂移主导; 仅 12.5% 真正分类有用
0.10
Total RF importance (top-80)
(c) Per-band aggregated importance — b33concentrated bands 30-34 & 51-54
0.6
0.4
0.2
0.0
0.6
0.4
0.2
0.0 0.0
selection is essential and exposes a methodological artifact in the original 1.8 GHz channel. Third, the main record-level detector achieves high AUC and low false-positive operation on the main corpus. Fourth, the same pipeline remains effective after re-collecting at the measured (80, 800) MHz carrier pair. Finally, the negative and stress results are consistent with the design: broad open-set skill recognition is fragile, while separable skill evidence and fine-window attack evidence provide the reliable basis for workflow-integrity monitoring.
perfect isotonic (ECE=0.0277)
0.8
0.2
0.4
0.6
0.8
1.0
0.0
0.6
0.8
1.0
(c) Calibration metrics — isotonic worsens all 3 0.2153
1522
10
3
180
−1
0.0453 0.0277
0.0249
Sample count
Metric value (log scale)
0.4
Mean predicted P(attack) (isotonic)
(d) Sample distribution per probability bin (most records cluster near 0 = idle/normal)
raw RF (lower=better) isotonic 10
0.2
Mean predicted P(attack)
10
2
48 24 14
10
1
6 2
0.0107
10
−2
0.0086
10
1
0
0 Brier
log-loss
ECE
0.0
0.2
0.4
0.6
0.8
1.0
RF P(attack) bin
Fig. 16: Probability calibration. Raw random-forest outputs are nearly identity-calibrated; isotonic recalibration is unnecessary.
VII. R ELATED W ORK To contextualize ClawGuard’s contributions, we systematically review the literature across host-based auditing, LLM agent security, and physical side-channel analysis. Table VII summarizes this landscape, highlighting the structural gaps that ClawGuard addresses.
utilizing eBPF (e.g., Kobra [9]) have become the standard for low-overhead tracing. Despite their sophistication, these software-layer defenses share a fundamental vulnerability: the symmetric threat model. They inherently assume the uncompromised integrity of the host OS and the hypervisor [10], [11]. As demonstrated by robust bootloader fuzzing campaigns [10] and controlled preemption attacks [11], once an adversary escalates privileges to the kernel or runtime environment, provenance records can be selectively forged, and eBPF sensors blinded. ClawGuard completely sidesteps this limitation by relocating the monitoring trust anchor outside the host’s physical and logical boundaries.
A. Host-Based Telemetry and Provenance Auditing Traditional host-based intrusion detection systems (HIDS) rely on OS-level telemetry to reconstruct causal execution flows. A massive body of literature has explored system-call-based provenance graphs for Advanced Persistent Threat (APT) detection, including seminal works like HOLMES [7], Sleuth [22], Unicorn [8], and ProvDetector [20]. Recent advancements have integrated graph representation learning (MAGIC [21], ThreaTrace [30]), causal inference (CausalIL [31]), and alert triage optimization (NoDoze [23], PrioTracker [25], Nodemerge [32]) to filter semantic workflow anomalies from massive enterprise logs [23], [24], [33]. Furthermore, in-kernel observability frameworks
B. LLM Agent Attacks and Defenses The transition from passive LLMs to autonomous, toolwielding agents has catalyzed a novel attack surface. Early
12
TABLE VII: Taxonomy of Workflow Integrity and Side-Channel Defenses vs. ClawGuard. (OS-Resilient: Survives kernel/root compromise; Granularity: Target abstraction level; OOB: Out-of-band isolated channel). Defense Paradigm
Representative Systems
Granularity
Host-based Provenance
HOLMES [7], Unicorn [8], ProvDetector [20], MAGIC [21], Sleuth [22] Kobra [9], NoDoze [23], Log2vec [24], PrioTracker [25] ACE [26], SAGA [27], StruQ [28], AgentDojo [29] RSA/ECC Extraction [12], [13], Screaming [15], BlueScream [16] EMMA [14], Callan et al. [17], Yilmaz et al. [18] This paper
In-Kernel eBPF / Logs LLM Agent Guardrails Bit-Level EM Crypto App-Level EM Profiling ClawGuard (Ours)
Figure 30. Inference latency — per-record p50=18.2ms, p99=28.6ms (real-time-able for 20s captures)
(a) Per-stage latency breakdown 30
17.01
Latency (ms)
Latency (ms)
28.6 24.8
25
20
15
10
22.6
20
19.3
18.2
15 10
5
5 0.73
1.12
0 18.7 feature
extract
0.18
0.28
0.35
0.21
standard scaler
anova select
0 rf predict 18.7 proba
p50
(c) Batch latency — RF predict scales sub-linearly
18.741
Per-record latency (ms, log scale)
18.75
Median ms per batch
18.50 18.25 18.00 17.75
17.1
17.0
17.50 17.25
10
p90
p95
p99
max
mean
(d) Amortised per-record latency drops 100× at B=128
1
2.134
10
0
0.533
0.146
17.00 1
8
Batch size
32
128
1
8
OOB?
System Call / Process Instruction / Event
×
N/A
×
×
N/A
×
Semantic / Prompt
×
N/A
×
Cryptographic Bit
✓
×
✓
Application / OS Loop
✓
×
✓
Agent Skill
✓
✓
✓
Physical side-channels have long been investigated for out-of-band monitoring. Classic electromagnetic (EM) security research heavily targets offensive cryptanalysis, achieving bit-level extraction of RSA/ECC keys via near-field probes [12], [13], [15], [16]. Moving up the abstraction hierarchy, instruction- and application-level fingerprinting frameworks (EMMA [14], Callan et al. [17]) successfully distinguish among small sets of desktop applications or detect smartphone camera activations [18]. Other physical modalities, such as power line monitoring (WattsUp [40], HardFails [41]) and acoustic snooping (RefleXnoop [42]), similarly exploit hardware physics for security. ClawGuard pioneers a distinct, mid-tier granularity: skilllevel workflow monitoring. Unlike bit-level cryptanalysis (which targets tightly unrolled mathematical loops) or applevel classification (which identifies monolithic programs), LLM agent skills are seconds-long, compositional, and highly dynamic macroscopic workloads. Translating these noisy, nonstationary RF streams into semantic intent requires overcoming severe thermal drift—a challenge largely unaddressed in shortwindow crypto-EM literature. By leveraging a drift-aware coarse–fine windowing pipeline on SDRs at 2 cm, ClawGuard provides a resilient hardware trust anchor that definitively bridges the gap between analog emanations and autonomous workflow integrity.
32.8
27.04
Drift-Aware?
C. Out-of-Band and EM Side-Channel Monitoring
(b) Total per-record latency
p50 p95 p99
25
OS-Resilient?
32
128
Batch size
Fig. 17: Inference latency. Median post-feature latency is 18 ms and p99 is 29 ms.
offensive research exposed indirect prompt injections [34]– [36], which have rapidly evolved into sophisticated workflow hijacking vectors. Adversaries can now weaponize poisoned Retrieval-Augmented Generation (RAG) architectures (e.g., PoisonedRAG [4], ObliInjection [6], RAGPoison [37]) and malicious tool documentation (e.g., ToolHijacker [5], AgentSmith [38]) to seamlessly subvert an agent’s planning logic without direct API access. To mitigate these threats, the defense community has proposed extensive application-layer safeguards. These range from prompt isolation and privilege separation architectures (ACE [26], SAGA [27], StruQ [28]) to comprehensive safety benchmarks (AgentDojo [29]). While these guardrails successfully restrict non-privileged logical flaws, they execute entirely within the host’s memory space. If the execution environment itself is compromised via a supply-chain vulnerability (e.g., malicious PyPI packages in the toolchain [39]), software-level constraints fail. ClawGuard addresses this by treating the agent execution as a black box, verifying its structural physical footprint instead of its self-reported logs.
VIII. D ISCUSSION AND L IMITATIONS The Physics of Skill-Level Observation and Feature Bottlenecks. Our evaluation demonstrates a sharp granularity gradient in side-channel analysis. Unlike bit-level cryptanalysis or monolithic app-level fingerprinting, LLM agent skills manifest as seconds-long, compositional workloads whose EM envelopes are dominated by structural resource choices (compute vs. I/O vs. memory). However, our mutualinformation diagnosis reveals a fundamental feature-space bottleneck: while architecturally distinct skills exhibit extreme pairwise separability (F1 > 0.94), scaling to a flat 16-class identification hits a plateau (macro-F1 ≈ 0.146). This is not
13
a generic detector failure but a physical reality—distinct skills with similar underlying resource constraints generate highly confusable power draws. This physical overlap explicitly necessitates ClawGuard’s two-stage design: relying on coarse-fine sequence voting and structural workflow editdistance comparisons rather than brittle, single-shot multi-class predictions. Future work may lift this ceiling by exploring contrastive pretraining objectives or phase-structure features that inherently subtract cycle-level nuisance variation. Adversarial Adaptation and Asymmetric Defense. A sophisticated adversary aware of ClawGuard might attempt several adaptive evasion strategies. The attacker could shape a malicious payload to mimic the EM profile of a benign skill, execute payloads in sub-window bursts (< 0.5 s), or exploit dynamic voltage and frequency scaling (DVFS) manipulation to erase the targeted harmonic footprint. Against EM-mimicry, ClawGuard’s confusability-weighted edit distance already penalizes unlikely structural sequences, though detecting perfectly crafted identical-EM payloads remains an open challenge. Against DVFS manipulation, our in-situ band selection methodology explicitly abandons highly volatile CPU clock harmonics in favor of stable, low-frequency PMIC and DRAM fundamentals. Furthermore, if an attacker resorts to deploying host-controlled GPIO/PWM emitters to actively jam the inband RF spectrum, a hardened ClawGuard deployment can trivially fallback to a spectrum-anomaly monitor that flags the environment as "signal-degraded" rather than emitting falseclean verdicts, maintaining the asymmetric defense advantage. Deployment Practicality and Generalization Limits. Operating via passive SDRs at a 2 cm standoff cleanly matches realistic deployment constraints in shared enterprise environments (e.g., rack-mounted industrial gateways), entirely preserving the out-of-band trust anchor. With a hardware Bill of Materials (BoM) around $1,500, batched inference latency of 0.15 ms, and an online feature extraction pipeline that shrinks raw IQ storage to ∼ 50 GB/day, the system is highly practical for real-time SOC auditing. However, generalizing physicallayer models remains a recognized challenge. While ClawGuard successfully mitigates within-session thermal drift via dynamic polynomial detrending, and proves remarkably robust across carrier frequencies (the 88.3% new-bands replication), cross-device (e.g., different ARM SoCs) and cross-session generalization (e.g., across distinct days with altered ambient interference) without re-calibration remain limited. Transitioning from the current technical-feasibility single-device study to a multi-DUT, multi-environment deployed system requires lightweight, automated deployment-time calibration protocols, which we earmark as the primary trajectory for future research.
audit logs, and eBPF monitors—can be fundamentally blinded or forged. To address this, we introduced ClawGuard, the first out-of-band, physical-layer integrity monitor for agent workflows. By capturing unintentional electromagnetic (EM) emanations via passive SDRs at 2 cm, ClawGuard grounds its security guarantees in unforgeable hardware physics rather than vulnerable host software. To bridge the semantic gap between continuous analog signals and discrete agent intent, we engineered a drift-resistant, event-aware coarse–fine windowing architecture that translates macroscopic hardware execution envelopes into deterministic security verdicts. Our extensive evaluation, underpinned by a 7.82 TB RF corpus spanning 38 benign and attack skills, demonstrates the profound efficacy of this approach. ClawGuard achieves a production-split ROC AUC of 0.9945 and a 100% true positive rate at just 1.16% false positive rate, delivering structural integrity validations with a median inference latency of a mere 18 ms. Beyond the system itself, we contribute critical methodological insights to the side-channel community by exposing the measurement pitfalls of OS-governor-driven frequency modulation, establishing a robust, OS-agnostic band-selection standard. By formalizing and successfully validating skilllevel EM monitoring, ClawGuard establishes a practical, hostindependent hardware trust anchor, opening a new frontier for securing the next generation of autonomous AI platforms. R EFERENCES [1] Statista, “Generative AI - worldwide | statista market forecast,” https://www.statista.com/outlook/tmo/artificial-intelligence/ generative-ai/worldwide, 2024, accessed: 2026-05-07. [2] LangChain, “LangChain: A framework for developing applications powered by language models,” https://www.langchain.com/, 2024. [3] OpenClaw, “Openclaw — personal ai assistant,” 5 2026, [Online; accessed 2026-05-07]. [Online]. Available: https://openclaw.ai/ [4] W. Zou, R. He, T. Bachmann, M. Salehi et al., “PoisonedRAG: Knowledge corruption attacks to retrieval-augmented generation of large language models,” in Proceedings of the 34th USENIX Security Symposium (USENIX Security 25). USENIX Association, 2025. [5] Anonymous, “Prompt injection attack to tool selection in LLM agents,” in Proceedings of the 2026 Network and Distributed System Security Symposium (NDSS ’26). Internet Society, 2026. [6] S. Xu et al., “ObliInjection: Order-oblivious prompt injection attack to LLM agents with multi-source data,” in Proceedings of the 2026 Network and Distributed System Security Symposium (NDSS ’26). Internet Society, 2026, arXiv:2512.09321. [7] S. M. Milajerdi, R. Geng, S. Khalighinejad, H. Agarwal, M. Egele, and N. Nikiforakis, “HOLMES: Real-time apt detection through correlation of suspicious information flows,” in 2019 IEEE Symposium on Security and Privacy (SP), 2019. [8] X. Han, T. Pasquier, A. Bates, J. Mickens, and M. Seltzer, “Unicorn: Runtime provenance-based detector for advanced persistent threats,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2020. [9] R. Farkhani et al., “Kobra: Targeted activity monitoring with ebpf,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2023. [10] Z. Zhong et al., “A comprehensive memory safety analysis of bootloaders,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2025. [11] Y. Zhu, B. Chen, Z. N. Zhao, and C. W. Fletcher, “Controlled preemption: Amplifying side-channel attacks from userspace,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 25). ACM, 2025.
IX. C ONCLUSION As LLM-driven agents evolve into autonomous, task-driven infrastructure, securing their execution against workflow hijacking requires transcending the symmetric trust boundaries of host-internal software. When adversaries can achieve arbitrary code execution and compromise the underlying operating system, traditional software telemetry—such as system calls,
14
[12] D. Genkin, I. Pipman, and E. Tromer, “Get your hands off my laptop: Physical side-channel key-extraction attacks on PCs,” in Cryptographic Hardware and Embedded Systems – CHES 2014, ser. Lecture Notes in Computer Science, vol. 8731. Springer, 2014, pp. 242–260. [13] D. Genkin, L. Pachmanov, I. Pipman, E. Tromer, and Y. Yarom, “ECDSA key extraction from mobile devices via nonintrusive physical side channels,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS ’16). ACM, 2016, pp. 1626–1638. [14] N. Sehatbakhsh, B. B. Yilmaz, A. Zajić, and M. Prvulovic, “EMMA: EM-based anomaly detection for embedded systems,” in Proceedings of the 29th USENIX Security Symposium (USENIX Security ’20). USENIX Association, 2020, pp. 1245–1262. [15] G. Camurati, S. Poeplau, M. Muench, T. Hayes, and A. Francillon, “Screaming channels: When electromagnetic side channels meet radio transceivers,” in Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security (CCS ’18). ACM, 2018, pp. 163–177. [16] P. Ayoub, R. Cayre, A. Francillon, and C. Maurice, “BlueScream: Screaming channels on bluetooth low energy,” in Proceedings of the 40th Annual Computer Security Applications Conference (ACSAC ’24). ACM, 2024. [17] R. Callan, A. Zajić, and M. Prvulovic, “A practical methodology for measuring the side-channel signal available to the attacker for instruction-level events,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’14). IEEE, 2014, pp. 242–254. [18] B. B. Yilmaz, E. E. Ugurlu, A. Zajić, and M. Prvulovic, “Detecting cellphone camera status at distance by exploiting electromagnetic emanations,” in Proceedings of the 2019 IEEE Military Communications Conference (MILCOM). IEEE, 2019, pp. 1–6. [19] J. Liang, Y. Wang, C. Li, and T. Wang, “GraphRAG under fire: Exposing vulnerabilities of GraphRAG to targeted poisoning attacks,” in Proceedings of the 2026 IEEE Symposium on Security and Privacy (S&P ’26). IEEE, 2026. [20] Q. Wang, W. U. Hassan, A. Bates et al., “ProvDetector: A provenancebased stealthy malware detection system,” in Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2020. [21] Z. Zheng et al., “MAGIC: Detecting advanced persistent threats via masked graph representation learning,” in 30th USENIX Security Symposium (USENIX Security 21), 2021. [22] M. N. Hossain et al., “Sleuth: Real-time attack scenario reconstruction from cots audit data,” in 26th USENIX Security Symposium (USENIX Security 17), 2017. [23] W. U. Hassan, S. Guo, D. Li, Z. Chen, K. Jee, Z. Li, and A. Bates, “Nodoze: Combatting threat alert fatigue with automated provenance triage,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2019. [24] F. Liu et al., “Log2vec: A heterogeneous graph embedding based approach for detecting cyber threats within enterprise,” in Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2019. [25] Y. Liu et al., “PrioTracker: Tuning ephemeral trace events for reliable threat detection,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2021. [26] Anonymous, “ACE: A security architecture for LLM-integrated app systems,” in Proceedings of the 2026 Network and Distributed System Security Symposium (NDSS ’26). Internet Society, 2026, arXiv:2504.20984. [27] ——, “SAGA: Governing AI agent security,” arXiv:2504.21034, 2025. [28] S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “StruQ: Defending against prompt injection with structured queries,” in Proceedings of the 34th USENIX Security Symposium (USENIX Security 25). USENIX Association, 2025. [29] E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents,” in Advances in Neural Information Processing Systems 37 (NeurIPS 2024) Datasets and Benchmarks Track, 2024. [30] S. Wang et al., “Threatrace: Detecting and tracing host-based threats in node level through graph convolutional networks,” in 31st USENIX Security Symposium (USENIX Security 22), 2022.
[31] Y. Chen et al., “CausalIL: Causal graph learning for host-based intrusion detection,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2023. [32] W. U. Hassan et al., “Tactical provenance analysis for endpoint detection and response systems,” in 2020 IEEE Symposium on Security and Privacy (SP), 2020. [33] Z. Zheng et al., “Poirot: Aligning attack behavior with threat intelligence,” in Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2019. [34] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world LLMintegrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23). ACM, 2023, pp. 79–90. [35] F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” arXiv preprint arXiv:2211.09527, 2022. [36] J. Liu et al., “Formalizing and detecting indirect prompt injection attacks,” in Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2024. [37] W. Zou et al., “Poisoning retrieval-augmented generation for large language models,” in 33rd USENIX Security Symposium (USENIX Security 24), 2024. [38] G. Chen et al., “Agent smith: A single image can hijack your autonomous agent,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. [39] Y. Xiao, D. Kirat, D. L. Schales, J. Jang, L. Xing, and X. Liao, “JBomAudit: Assessing the landscape, compliance, and security implications of Java SBOMs,” in Proceedings of the Network and Distributed System Security Symposium (NDSS 25). Internet Society, 2025. [40] S. S. Clark et al., “Wattsupdoc: Power side channels to nonintrusively discover untargeted malware on embedded medical devices,” in USENIX Workshop on Health Information Technologies (HealthTech), 2013. [41] G. Dessouky et al., “Hardfails: Insights into software-exploitable hardware bugs,” in 28th USENIX Security Symposium (USENIX Security 19), 2019. [42] P. Wang, J. Hu, C. Liu, and J. Luo, “RefleXnoop: Passwords snooping on NLoS laptops leveraging screen-induced sound reflection,” in Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS ’24). ACM, 2024.
A PPENDIX This appendix supports the robustness discussion in §VI-D. It records a deliberately difficult stress campaign and explains why ClawGuard is evaluated as a staged physical-consistency monitor rather than as a universal single-shot 22-class skill recognizer. a) Corpus and protocol.: The openclaw_attack_v1 campaign exercises the full 22-skill attack catalog on a separate host and a third carrier pair (248, 800) MHz. This corpus is distinct from both the headline laptop corpus and the newbands replication corpus in Table IV. It is therefore used only as a boundary stress test: it mixes host shift, run shift, carrier shift, and small per-class support. Twenty-two attack-related skills were ported: 11 skills with prefix attack_third_party_, 10 skills with prefix attack_bad_tool_result_, and one idle anchor. Each attack skill is a one-to-one reimplementation of a corresponding m4p1e/agent-sentinel attack entry and emits attack_begin, payload_start, payload_end, and attack_end events for offline alignment. The usable feature set contains 155 records across all 22 classes. b) Negative result.: A flat 22-class classifier does not recover reliable skill identity at this sample budget. Table VIII summarizes the evidence. The combined A+B point estimate, macro-F1 = 0.0818, is only modestly above random, and its
15
TABLE VIII: Key evidence openclaw_attack_v1 stress campaign.
from
the
v2 deltas with selected bands marked 8
Selected CPU 80 MHz Selected RAM 800 MHz
Finding
Flat 22-class
macro-F1 = 0.0818; bootstrap CI [0.043, 0.114] overlaps random baseline 3-class and 6-class settings remain at-random after baseline correction A-only training to B testing reaches 0 macroF1 in E9-2 26/605 features have |d| ≥ 0.8; max |d| = 6.7 Single-channel and cross-channel-only models are at or below random 14/22 classes have F1 = 0; the apparent stat bright spot vanishes in a binary test
Class collapse Cross-run transfer Feature drift Channel ablation Per-class behavior
4 2 0 −2
8 6
ΔRAM [dB]
Test
ΔCPU [dB]
6
4 2 0 −2
500
1000
1500
2000
2500
3000
Frequency [MHz]
Fig. 18: Band survey v2 workload deltas with the selected CPU/RAM bands highlighted.
bootstrap confidence interval overlaps the random baseline. Collapsing the label space to 3 or 6 classes does not solve the problem: the apparent numerical gain is explained by the corresponding rise in the random baseline. The most important failure mode is distribution shift: training on A cycles and testing on B’s cycle 5 collapses to 0 macro-F1 . c) Implication for the design.: This stress campaign supports two design choices made in §V. First, ClawGuard should not be claimed as a universal open-set 22-class skill recognizer under small per-class support. Second, drift-aware evaluation is necessary: LOCO splitting, cycle-local normalization, temperature detrending, and training-fold-only feature selection are safeguards against collection-cycle leakage. The robust operating setting is therefore the one evaluated in the body: use skill evidence where physical separability is strong, and use fine-window attack evidence for localized workflowhijacking activity. This appendix supports C1 and C2 in §VI: the system is a real passive RF prototype, and carrier selection must be measured rather than assumed from nominal hardware frequencies. d) Band-survey iterations.: The first survey iteration used the default ondemand governor and a single-threaded random-walk memory workload. Two pathologies appeared. First, the 1.8 GHz CPU-clock region was partly driven by governor activity. Second, the memory workload did not create a clear LPDDR4 signature. The second iteration pinned the host with the userspace governor and replaced the memory workload with a large inplace radix-sort workload. This produced visible CPU/RAM separation below 500 MHz and around 700–900 MHz. The memory-heavy signature increased from roughly +4.3 dB to +7.1 dB at the best window. The LPDDR4 zoom further showed that the RAM-heavy trace is elevated by roughly 3–4 dB across 700–900 MHz and peaks near 800 MHz. By contrast, the 1.8 GHz region remains tightly bundled under explicit governor pinning. This supports the methodological claim that the legacy 1.8 GHz channel partly reflects governorsensitive activity-frequency modulation rather than a stable idle-vs.-busy hardware-power signature.
Top windows ranked by activity-induced power excursion (v2 survey, 20 MHz window) Top 10 windows by ΔCPU
Top 10 windows by ΔRAM
66
681
81
711
76
696
276
66
311
56
91
276
306
211
281
216
271
221 ΔCPU ΔRAM stacked
251 0
2
4
6
8
10
12
14
ΔRAM ΔCPU stacked
251 0
dB delta Window center [MHz] (left) and stacked Δ
2
4
CPU + ΔRAM (right)
6
8
10
12
14
dB delta
Fig. 19: Top candidate windows ranked by CPU and RAM workload deltas. e) New-bands corpus.: After selecting the (80, 800) MHz carrier pair, we recollected a complete attack benchmark. A 2 s pre-flight on each HackRF unit confirmed clean tuning before the full run. The full 10-cycle session ran for 98 minutes and produced 550 IQ files, corresponding to 440 GB of raw IQ, together with 221 per-skill event files and 11,114 temperature samples. The session exited cleanly. Per-cycle progress was steady, per-skill counts remained balanced across the two channels, and the two receiver channels exhibited stable amplitude regimes over the run. The new-bands replication reaches 83.6% sub-window accuracy, 88.3% record-vote accuracy, and 90.3% record-level attack recall on the surviving attack-class subset. These results support cross-carrier feasibility while also bounding the claim: residual drift remains, especially in folds where background windows are conservatively classified as attack. This appendix supports C2–C4 in §VI. It preserves the additional analyses that substantiate the body claims: skill separability is structured but not universal; the coarse–fine detector improves localized attack detection; the new-bands result is useful but still affected by drift; and runtime overhead is small relative to skill duration.
16
log_rotate_compress
0.96 0.78 0.45
wiki_email git_dev_workflow browser_session_heavy simple_email file_sync_backup calculator_develop clock_develop package_install_update
0.74 0.73 0.78 0.81
0.78 0.43 0.63 0.76 0.60 0.74 0.49 0.52 0.69 0.44 0.51
0.95 0.77 0.54 0.54 0.78
0.53 0.69 0.60 0.67 0.54 0.52 0.58 0.56 0.55
ild
0.9
0.82 0.66 0.68 0.65 0.63 0.63 0.53
0.8
0.66 0.65 0.60 0.49 0.55 0.65 0.53 0.53
0.91 0.77 0.51 0.52 0.76 0.48 0.69 0.66
0.51 0.54 0.60 0.59 0.45 0.57 0.61
0.87 0.61 0.66 0.63 0.60 0.60 0.60 0.65 0.51
0.54 0.48 0.62 0.52 0.63 0.62
0.91 0.67 0.61 0.48 0.74 0.60 0.67 0.60 0.54 0.54
0.7 0.6
0.57 0.57 0.51 0.50 0.54
0.82 0.62 0.74 0.67 0.49 0.68 0.54 0.49 0.60 0.48 0.57
0.62 0.54 0.49 0.50
0.73 0.73 0.71 0.59 0.52 0.55 0.52 0.55 0.59 0.62 0.57 0.62
0.5
std
0.10
20
max
0.06 0.08
min
0.08
15
p25
0.08
0.06
p75
0.09 0.10
0.04
0.000 0.175
max 0.08
slope
10
0.251
mean
0.12
std
0.76 0.63 0.48 0.60 0.60 0.68 0.55 0.51 0.63 0.59
0.78 0.71 0.70 0.69 0.43 0.76
(b) Per-statistic aggregated importance band 30-34 (89%) band 51-54 (8%)
0.06 0.13
mean
0.234
min 0.117
p25
0.223
p75 slope
0.000
peak_rate
0.000
0.02
peak_rate 0.00 0
4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
0.00
0.05
Frequency band index (0..63)
0.10
0.15
0.20
0.25
Total RF importance (top-80)
0.56 0.50 0.49
0.85 0.72 0.55 0.57 0.69 0.51 0.58 0.65 0.45 0.52 0.51 0.54 0.56
0.61 0.53
0.83 0.68 0.68 0.67 0.44 0.63 0.56 0.53 0.57 0.63 0.50 0.49 0.50 0.61
0.4
5
(c) Per-band aggregated importance — b33concentrated bands 30-34 & 51-54 0.47
0.50
0.80 0.73 0.58 0.69 0.51 0.59 0.55 0.53 0.61 0.62 0.54 0.50 0.49 0.53 0.50
0 0.6
0.8
Pairwise macro-F1
0.4 b30 0.26
0.3 0.2
b32 0.12
0.1 0.0
sy
bu
median = 0.615
1.0
0.81 0.54 0.69 0.65 0.52 0.63 0.48 0.67 0.59 0.57 0.67 0.69
_r ele as e_ p ste b ipeli n m ac log _m kgro e a u _r ota inte nd te_ na vid com nce eo pr _ se e ns stre ss or _p amin oll g db ing_ _a iot na w lytic g br s ow it_d iki_ e e se r_ v_w mail se ss orkflo ion w _ h s file imple eavy _ _ ca sync em a lcu _ lat bac il or pa _d kup c ck e ag lock velo e_ _ ins dev p loc tall elo al_ _up p wik d i_s ate ea rch
local_wiki_search
(a) RF feature importance heatmap (top-30 features) bright = important
RF importance
db_analytics
0.45 0.78 0.54 0.70 0.68 0.51 0.66 0.61 0.74 0.71 0.55 0.68 0.58
Pairwise binary macro-F1
video_streaming sensor_polling_iot
25
0.78 0.78 0.73 0.77 0.71 0.66 0.77 0.61 0.67 0.62 0.73 0.72 0.68 0.73
0.94 0.78
31% pairs < chance chance ≈ 0.55
Statistic
0.86
Number of pairs
background
0.86 0.94 0.96 0.74 0.95 0.78 0.82 0.91 0.87 0.91 0.82 0.73 0.85 0.83 0.80
system_maintenance
Figure 26. Attack-binary feature attribution — narrow bands + level statistics only
(bb) Distribution of 240 pairwise scores
(a) 16-class pairwise separability (120 pairs)
Total RF importance
a build_release_pipeline
0
4
8
12
16
20
24
28
32
36
40
44
48
52
56
60
Frequency band index
Top-15 features (RF importance) ──────────────────────────────── # band stat rf_imp 1 33 mean 0.1338 ★ 2 33 p75 0.0994 ★ 3 30 p75 0.0861 ★ 4 33 max 0.0793 ★ 5 33 min 0.0777 ★ 6 33 p25 0.0767 ★ 7 30 mean 0.0638 ★ 8 30 max 0.0568 ★ 9 30 min 0.0534 ★ 10 32 min 0.0301 11 32 p75 0.0288 12 32 mean 0.0269 13 32 p25 0.0224 14 31 max 0.0171 15 53 min 0.0165
Fig. 20: Pairwise 2-class separability on big48.
Fig. 21: Random-forest band-level feature attribution on the headline corpus.
f) Skill separability.: The aggregate 16-class result on big48 is low (macro-F1 = 0.146), but the pairwise structure is informative. Among 120 pairwise 2-class tests, 12 pairs exceed 0.80 F1 , while 37 pairs fall below 0.55. The most separable skill is build_release_pipeline, which reaches 0.956 against log_rotate_compress, 0.947 against sensor_polling_iot, and 0.941 against system_maintenance. The least separable pair is db_analytics versus video_streaming at 0.435. The best operational meta-class triples, such as background/build/sensor, achieve macro-F1 ∈ [0.74, 0.79] under LOCO. This supports the staged design: skill-level EM evidence is useful where the physical classes are separable, but broad open-set recognition remains fragile. g) Ablations.: On focused3, the top-k feature sweep peaks at k = 65 with macro-F1 = 0.8977 and degrades beyond k = 80. On big48, k = 40 marginally outperforms k = 80, confirming that feature-count selection is task dependent and must be fit on the training fold only. Temperature detrending also depends on sample support. On focused3, a first-order detrend improves over no detrend by 0.018 macro-F1 and over a second-order detrend by 0.045, while smaller datasets can be harmed by detrending. Adding raw temperature as a feature changes macro-F1 by less than 0.005. Raw randomforest scores are already well calibrated (Brier score 0.011, ECE 0.009), and post-hoc isotonic recalibration degrades both metrics. Band attribution confirms that the headline legacy model uses RF structure but also motivates the new-bands replication. Bands 30–34, corresponding to the legacy 1.8 GHz region, carry most cumulative importance; masking this cluster drops AUC substantially. Cycle leakage is also substantial: many V10 features have higher mutual information with cycle index than with skill class, which motivates LOCO evaluation and leakage-controlled preprocessing. h) Coarse–fine attack detection.: Pilot A captures 30 tasks in which the planner is induced to call attack_third_party_rm. A flat 2-class within-session classifier on 20 s traces reaches macro-F1 = 0.670 with attack recall 0.55. The coarse–fine stage-2 detector improves the
TABLE IX: Coarse–fine 3-class attack-detection results on big48_chunk1+20260423b. Config
Sub-win acc.
Record acc.
bg recall
atk recall
v1 v2 v3 no-temp
0.816 0.787 0.809 0.809
0.9398 0.9252 0.9398 0.9398
0.250 0.139 0.222 0.222
0.833 0.844 0.861 0.861
record-level result to 0.9398 accuracy with attack recall 0.833. The larger attack pilot exercises the full 22-skill attack catalog over 10 cycles and 299 records. Across four stage-2 configurations, record-level accuracy remains at least 0.9252 and attack recall lies between 0.833 and 0.861, as shown in Table IX. This supports the body claim in §VI-C: fine-window evidence exposes short malicious payloads that are diluted in whole-record representations. i) New-bands stability.: Figures 22 and 23 expand the new-bands replication: cycles 3, 7, and 8 reach 100% recordvote accuracy, while cycles 1, 5, 6, and 10 are lower. In cycles 1, 6, and 10, every background sub-window is classified as attack, explaining much of the per-fold variance. This does not invalidate the detection claim, but it bounds it: the newbands corpus supports cross-carrier feasibility, while a larger or temperature-stratified recollection is needed to reduce false alerts under new-band drift. j) Baselines, failure modes, and runtime.: A naive crosscorpus binary attack classifier is not sufficient. Training a flat 2-class model with big48 normals as benign and the 22-class attack corpus as malicious yields apparent accuracy 0.993, but macro-F1 = 0.498 and catches zero of 11 attack records. Oneclass anomaly baselines also remain near chance, with AUC in [0.44, 0.58]. These failures support the design choice to use supervised fine-window attack evidence rather than density anomaly alone. Post-feature inference is not the bottleneck. The median perrecord inference latency is 18 ms, the p99 latency is 29 ms, and batched prediction amortizes to approximately 0.15 ms per record. Because agent skills last seconds and fine windows last
17
background
395 (84%)
33 (7%)
44 (9%)
Final LogReg 3-session record-vote confusion (LOCO-20)
1.0
0.8
background
77 (88%)
8 (9%)
3 (3%)
normal
3 (5%)
58 (95%)
0 (0%)
attack
8 (9%)
0 (0%)
85 (91%)
background
normal Predicted
attack
14 (6%)
229 (94%)
1 (0%)
attack
30 (10%)
0 (0%)
273 (90%)
background
normal Predicted
attack
0.8
0.6 True
normal
Recall
True
0.6
1.0
0.4
Recall
Final LogReg 3-session sub-window confusion (LOCO-20)
0.4
0.2
0.0
(a) Sub-window pooled.
0.2
0.0
(b) Record vote.
Sub-window Accuracy (%)
Fig. 22: Confusion matrices for the new-bands LOCO experiment. 100 90 80 70 60 50
Per-fold accuracy Mean = 84.7% 1
2
3
4
5
6
7
8
9
Figure 30. Inference latency — per-record p50=18.2ms, p99=28.6ms (real-time-able for 20s captures)
10
(a) Per-stage latency breakdown
(b) Total per-record latency 32.8
27.04
Hold-out Cycle (LOCO)
p50 p95 p99
25
30 25 17.01
Latency (ms)
Latency (ms)
20
Fig. 23: Per-fold accuracy across the 10 new-bands LOCO folds.
28.6 24.8
15
10
22.6
20
19.3
18.2
15 10
5
5 0.73
1.12
0
hundreds of milliseconds, sensing and windowing dominate wall-clock delay.
18.7 feature extract
0.18
0.28
0.35
0.21
standard scaler
anova select
0 rf predict 18.7 proba
p50
(c) Batch latency — RF predict scales sub-linearly
18.741
Per-record latency (ms, log scale)
18.75
Median ms per batch
18.50 18.25 18.00 17.75
17.1
17.0
17.50 17.25
10
p90
p95
p99
max
mean
(d) Amortised per-record latency drops 100× at B=128
1
2.134
10
0
0.533
0.146
17.00 1
8
Batch size
32
128
1
8
32
Batch size
Fig. 24: Per-record inference latency breakdown.
18
128