1
Detecting Logic Vulnerabilities Across the Contract and Device Layers of Blockchain-Enabled IoT With Multi-Agent Heterogeneous Graph Attention
arXiv:2609.18344v1 [cs.CR] 16 Sep 2026
Minfeng Qi, Jialin Li, Tianqing Zhu, Lefeng Zhang, and Zhe Sun
Abstract—Blockchain-enabled Internet of Things (IoT) systems integrate smart contracts with embedded devices to support decentralized device management and access control. Their security therefore depends jointly on the logic of on-chain contracts and off-chain device firmware. Logic flaws in either layer can violate the same system invariants, such as unauthorized access, improper state changes, or unguarded privileged operations. Existing approaches rely on contract analysis, firmware analysis, and graph-based vulnerability detection. However, these methods typically focus on a single layer or artifact and often depend on predefined vulnerability patterns, emulation fidelity, or homogeneous representations that obscure security-relevant component roles. They also lack a unified architecture that supports different security tasks while remaining deployable on resource-constrained gateways. To address these limitations, we extend MA-HGAT into a cross-layer multi-agent heterogeneous graph attention framework that models contracts, firmware artifacts, device fleets, and transaction streams with a unified four-role, nine-relation schema. Role-aligned agents exchange heterogeneous evidence through cross-attention, while graph-, link-, and node-level heads support multiple detection tasks and a role-based gateway–cloud partition enables lightweight edge inference. MA-HGAT thus provides a unified and deployable framework for detecting logic vulnerabilities across the contract and device layers of blockchain-enabled IoT systems. Index Terms—Blockchain-enabled IoT, smart contract security, firmware security, logic vulnerability detection, heterogeneous graph neural networks, multi-agent systems, edge–cloud collaboration.
I. I NTRODUCTION LOCKCHAIN-ENABLED Internet of Things systems replace a centralized trust anchor with a ledger on which smart contracts register devices, enforce access policies, and validate device data [2]–[4]. The devices themselves remain embedded systems running vendor firmware, while gateways relay requests, status information, and transactions between devices and the blockchain. This architecture has been adopted in smart grids, supply chains, and industrial monitoring because it reduces dependence on a single trusted server and
B
M. Qi, J. Li, T. Zhu, and L. Zhang are with the Faculty of Data Science, City University of Macau, Macau, China (e-mail: [email protected]; [email protected]; [email protected]; [email protected]). Z. Sun is with Guangzhou University, Guangzhou, China (e-mail: [email protected]). Corresponding author: Tianqing Zhu. A preliminary version of this article, covering the contract layer only, appeared in the International Symposium on Cyberspace Safety and Security (CSS) [1]. This version adds the system-level formulation for blockchainenabled IoT (Section III), the task heads (Section IV-F), the gateway–cloud partitioning (Section V-B), the per-graph implementation, and all device-layer and deployment experiments (Section VII).
provides a shared audit trail. However, system correctness no longer depends on one software boundary: it depends jointly on the contract logic that governs devices and the firmware logic that controls them. A design mistake at either side can therefore violate the same system policy, for example by admitting an unauthorized device or allowing an attacker to reconfigure an authorized one [5]–[8]. Existing security techniques address this problem from several directions. Smart contract analyzers use static analysis, symbolic execution, fuzzing, or semantic reasoning to identify vulnerable contract behavior [9]–[15]. Firmware analyzers recover authentication and input-validation weaknesses from binaries, emulate devices for testing, or search related firmware images for previously disclosed vulnerabilities [16]–[25]. Learning-based approaches further model program structure from labeled examples [26]–[28]. These approaches are effective within their intended settings, but they remain fragmented: rule-driven methods depend on predefined vulnerability patterns, dynamic methods depend on faithful execution environments, and learning-based methods often simplify security-relevant components into a uniform representation. More importantly, these methods usually analyze either contracts or firmware in isolation and are not designed around the gateway where the two sides of the system meet. The resulting gap is not simply the absence of another vulnerability detector, but the absence of a common basis for reasoning about security across the whole system. In practice, an operator must answer four related questions: whether a contract contains unsafe logic, what type of weakness is described by a device vulnerability report, which deployed devices are affected by a newly disclosed weakness, and whether the behavior of a device appears malicious. These questions are currently handled by separate tools even though they concern the same security policies. Fig. 1 illustrates why this separation is problematic. At the contract side, a device-registration operation may update the registry without checking ownership; at the firmware side, a configuration operation may update device settings without checking the current session. Although the concrete software components are different, both failures have the same underlying form: an operation changes protected state without the required guard. The central problem is therefore whether these cross-layer failures can be described and detected through a common security view, while still supporting the different decisions required by operators. Addressing this problem introduces three main challenges.
2
state, identity, and historical context can remain in the cloud. These observations suggest that cross-layer detection should be organized around shared security roles rather than software layers, and that deployment should be divided according to what each location naturally observes.
Smart contracts on the chain functions, storage, modifiers, events transactions, access decisions
Gateway relays every device transaction requests, status
Devices running vendor firmware handlers, parameters, auth gates, crashes, logs the same four roles carry the logic at both ends function runs the logic
state holds the data
modifier guards access
event leaves a trace
a logic flaw is a guard that is missing, at either layer on the chain
in the firmware
? no owner check
? no session check
writes
writes
register device
registry
set config
settings
Fig. 1. A blockchain-enabled Internet of Things system in which smart contracts and device firmware jointly enforce system logic. The same four security roles appear at both sides: components perform actions, hold state, guard access, and leave observable traces.
First, the evidence is structurally diverse: contracts, firmware requests, deployed device images, and transaction streams contain different kinds of components, yet the security meaning often lies in how those components interact rather than in any component alone. Treating all components as equivalent removes distinctions such as whether an operation is guarded, while analyzing each component type independently loses the relationships that expose the flaw. Second, the required decisions have different forms. Contract auditing and vulnerabilityreport analysis require judgments about an entire artifact, fleet analysis requires identifying which devices are related to a disclosed weakness, and online monitoring requires a decision for each observed action. A common approach must therefore preserve the same security reasoning while supporting these different decision forms. Third, deployment introduces a physical constraint: gateways observe device activity first but have limited computation, memory, and connectivity, whereas the broader context needed to interpret that activity is maintained in the cloud. Our design is motivated by two observations that connect these challenges. First, despite their different implementations, security-relevant components across the contract and device layers repeatedly play the same four roles: some perform actions, some hold state, some guard access, and some record what happened. A logic vulnerability can therefore be viewed as an incorrect or missing relationship among these roles rather than as a layer-specific code pattern. This view explains why the two examples in Fig. 1 are security-equivalent even though one occurs in a smart contract and the other in firmware. Second, among these roles, device actions are the information that a gateway observes directly, while broader
Based on this rationale, we build one framework that represents contracts, firmware evidence, device fleets, and device–contract activity using the same role-based security structure. The framework preserves the distinctions among actions, state, guards, and observable traces while allowing evidence from these roles to inform one another, addressing the first challenge. It then uses the resulting shared representation to support whole-artifact assessment, fleet-level association, and event-level monitoring, addressing the second challenge without requiring a separate reasoning pipeline for each task. Finally, only the part that processes locally observed device actions is placed on the gateway, while the broader system context remains in the cloud, addressing the deployment challenge while avoiding the transmission of raw device requests. We realize this design in our multi-agent heterogeneous graph attention framework. This article extends the conference version of our framework [1]. The main contributions are as follows:
Cross-layer problem formulation. We formulate logic vulnerability detection in blockchain-enabled Internet of Things systems as a unified problem spanning smart contracts, device firmware, deployed device fleets, and device–contract activity. We show that the same four security roles and their relationships provide a common basis for describing logic flaws across these settings (Section III, Table III). • One framework for four security decisions. We extend the original contract-oriented framework so that one shared representation supports contract auditing, vulnerability-report analysis, fleet-level vulnerability linking, and transaction monitoring rather than requiring a separate model for each task (Section IV). • Gateway–cloud deployment aligned with system roles. We divide inference according to the information naturally available at the gateway and in the cloud. Only the component that processes executed device actions runs at the gateway, reducing the local model footprint and replacing raw requests with fixed-size representations while preserving the overall decision process (Section V-B, Section VII-F). • Evaluation across contract and device layers. We evaluate the framework on contract vulnerability detection, device vulnerability-report analysis, and firmware vulnerability linking, together with a gateway–cloud deployment study. The evaluation further includes component ablations and an unseen-vendor setting to identify both the source of cross-layer transfer and the generalization boundary of the learned representation (Section VI, Section VII). •
3
II. R ELATED W ORK A. Security of Blockchain-Enabled IoT Blockchains entered IoT architectures as a decentralized substitute for the trusted server that registers devices, mediates access, and audits data [29], [30]. Novo [2] showed that access management for large device fleets scales when contracts enforce the policies and devices reach the chain through management hubs. Surveys catalog the resulting designs for supply chains, energy, healthcare, and industrial monitoring [3], [31]. Recent work evaluates such architectures on 5G edge deployments and releases simulated device–contract transaction logs for that purpose [4]. Security analyses of these systems concentrate on the ledger and the protocol, and they treat contract correctness and device integrity as assumptions. This article treats both as objects of analysis. The contract that enforces a policy and the firmware that a registered device runs are both places where a logic mistake breaks the policy. We detect logic vulnerabilities at both. B. Smart Contract Vulnerability Detection Research on smart contract vulnerability detection has followed three paths: static analysis, dynamic analysis, and formal verification [14], [32]. Static analyzers such as Mythril [10], Slither [9], and Oyente [11] combine symbolic execution, control- and data-flow analysis, and patternbased vulnerability templates. Securify [33], ZEUS [34], Sereum [35], and SmartShield [36] extend this line with compliance patterns, abstract interpretation, run-time monitoring, and automatic hardening. These tools are effective for well-defined semantic vulnerabilities, but they must predefine what to look for. This limits them on business-logic flaws that arise from implicit invariants or cross-function interactions [37]. Dynamic approaches include property-based testing with Echidna [38], hybrid symbolic-concrete exploration with Manticore [39], and fuzzers such as ContractFuzzer [12] and Smartian [40]. They uncover vulnerabilities that appear only under specific input sequences, but they remain bounded by path explosion and by the quality of the property specifications [41]. Formal verification with the K framework [13], F⋆ [42], or the Solidity SMTChecker [43] offers strong guarantees for high-assurance contracts at the price of substantial specification effort. It also assumes that the specification itself is sound, which is exactly what a logic vulnerability violates. LLM-based auditors such as GPTScan [14], SmartLLaMA-DPO [15], and VulnHunt-GPT [44] add semantic reasoning over code and documentation [45], [46]. However, they process contracts as sequential text and remain sensitive to prompt design and to hallucinated findings. C. IoT Firmware and Device Security Analysis Large-scale studies of embedded firmware established early that IoT devices ship with weak authentication, hard-coded credentials, and insecure management interfaces at scale [7], [47]. Surveys organize the resulting detection literature into static, dynamic, and hybrid categories [48]. Static approaches
recover authentication-bypass conditions from binaries (Firmalice [16]) or track tainted data across the multiple binaries that implement one service (Karonte [17]). Others use keywords shared between front-end and back-end code to reduce taint sources (SaTC [18]) or reason about protocol state in bare-metal images (FirmXRay [49]). Tools such as BinAbsInspector [50] package abstract interpretation for stripped binaries. Dynamic analysis depends on emulation. AVATAR [51] and Avatar2 [52] coordinate hardware-in-theloop execution, and P2IM [20] models peripheral interfaces automatically. FirmAE [21] scales full-system emulation, and FIRM-AFL [19] makes greybox fuzzing of emulated firmware practical. Fidelity and crash observability, however, remain fundamental obstacles [53]. A separate line targets recurring vulnerabilities. Vendors reuse code across product families, so a vulnerability found in one image usually affects many others. Genius [22] and Gemini [23] search for vulnerable functions through graph-based and neural binary similarity, and FirmUp [54] matches procedures across stripped firmware. FirmRec [24] combines vulnerability-specific signatures with symbolic reasoning. FirmVulLinker [25] profiles whole images along five dimensions to link homologous vulnerabilities, and it releases the labeled fleet that we use in Section VII. Most recently, LLM agents assisted by fuzzing (FirmAgent [55]) discover vulnerabilities in real firmware. Surveys of agentic AI for IoT cybersecurity [56] identify deployment topology (edge, fog, cloud) and latency at the edge as open problems. Our work complements all of these. It does not replace binary analysis or emulation. Instead, it learns over the heterogeneous relations that they expose and applies the same representation to the contracts that govern the devices. It can also be partitioned between gateways and the cloud.
D. Graph Neural Networks for Code and Security Analysis GNNs have become a standard tool for program analysis because they model control, data, and structural dependencies directly [57], [58]. Researchers have applied graph convolutional and attention networks [59], [60] to abstract syntax trees, control-flow graphs, and data-flow graphs for bug detection, code classification, and vulnerability identification. Examples include attention-based localization of fine-grained vulnerabilities [28] and explainable reentrancy localization in contracts [27]. In the IoT domain, graph embeddings form the basis of cross-architecture binary similarity for firmware bug search [22], [23]. Heterogeneous GNNs such as HAN [61] and R-GCN [62] distinguish node and relation types through hierarchical attention or relation-specific transformations. Recent contract detectors adopt heterogeneous graphs [26]. These models, however, are domain-agnostic. They neither prioritize state-transition semantics nor separate the reasoning of different component types. Their evaluations rarely test whether the typed representation is actually needed. Our schema is security-oriented by construction, and our evaluation includes homogeneous counterparts trained on the same graphs.
4
TABLE I ROLES AND THEIR MEANING IN THE CONTRACT GRAPH . Role
Symbol
Contract-layer meaning
Function
Vf
State variable Modifier
Vs Vm
Event
Ve
Encodes executable logic and interaction behavior Maintains persistent contract state Enforces constraints and access-control policies Exposes externally observable execution traces
TABLE II R ELATION TYPES AND THEIR MEANING IN THE CONTRACT GRAPH . Relation
Symbol
Contract-layer meaning
Calls Depends Modifies Triggers Affects
Ecalls Edepends Emodifies Etriggers Eaffects
Constrained by Returns to
Econstrained_by Ereturns_to
Invokes Uses
Einvokes Euses
Function f1 invokes function f2 Function f reads state variable s Function f writes state variable s Function f emits event e Event e is indexed by state variable s Function f is guarded by modifier m Modifier m returns control to function f Modifier m calls function f Modifier m accesses state variable s
E. Attention, Multi-Agent Reasoning, and Edge–Cloud Inference Attention mechanisms let a model weight informative context during representation learning [63], [64]. The KVQ formulation separates what is stored (values) from how relevance is computed (queries against keys). In program analysis, this separation helps to reveal the execution dependencies that uniform aggregation averages away. Multi-agent architectures distribute complex reasoning among specialized agents that capture complementary views of the input [65]. In graph learning, agent specialization has been shown to increase representational diversity. Separately, edge intelligence [66] and split computing [67], [68] study how to partition a neural network between a resource-constrained device and a server so that only intermediate representations cross the network. Existing partitioning work targets layer boundaries of homogeneous networks. MA-HGAT instead exposes a semantic boundary, the role whose entities a gateway observes, which we use in Section V-B. III. S YSTEM M ODEL , T HREAT M ODEL , AND U NIFIED S CHEMA A. System Model Fig. 1 shows the blockchain-enabled IoT system that this article targets. It follows the architecture that accessmanagement and data-integrity proposals for IoT have adopted [2], [4], [29], [30], [69]. Three layers interact. Contract layer. A set of smart contracts on a permissioned or public chain holds the system’s shared state and enforces its policies. A device-registry contract records which devices exist and who owns them. An access-control contract decides which principals may read data from, or send commands to, which
devices. A data-validation contract checks and time-stamps reported measurements before other parties consume them. Contract functions execute the policies, storage variables hold the registry and the permissions, modifiers guard the functions, and events expose what happened to off-chain observers. Gateway layer. Edge gateways connect devices to the chain. A gateway authenticates the devices behind it, relays their transactions, caches access decisions, and is the first component to observe what a device does. Gateways have limited resources: they are embedded Linux boards or microcontrollers with kilobytes to megabytes of memory. They often reach the cloud through intermittent or metered links. Device layer. The devices run vendor firmware. Beyond sensing and actuation, the firmware exposes management interfaces, typically HTTP, SOAP, or CGI handlers. Through these interfaces, gateways and administrators read status and change configuration. Request handlers execute the device’s logic, request parameters and configuration registers hold its state, authentication gates guard the handlers, and crashes, logs, and telemetry expose its behavior. Devices from a few vendors and product families are deployed in large numbers. A weakness in one firmware image therefore usually affects many registered devices. The layers jointly enforce the invariants that make the system trustworthy. Only registered devices act, only authorized principals change a device’s configuration, and every privileged action leaves a trace. Each invariant can be broken from either side. If the registry contract’s registration function lacks an ownership check, an attacker registers a rogue device whose data every party then trusts. If a device’s firmware exposes a configuration handler that does not validate the session cookie, an attacker reconfigures a legitimately registered device. The device then keeps acting on the chain with valid credentials. The first flaw is in contract logic and the second in firmware logic. Neither is a memory-safety bug: the code does what its author wrote, and the author’s design is wrong.
B. Threat Model The adversary can read the chain and the contract code and submit transactions to the contracts. It can send requests to the management interfaces of devices it can reach, and it fully controls any device it has compromised. The adversary cannot break the chain’s consensus or cryptography, cannot tamper with gateways, and cannot alter the model or its inputs. The defender is the operator of the system. The operator can audit contract code before deployment (task A) and triage vulnerability reports and proof-of-concept traces that concern the deployed device models (task B). The operator can also determine which registered firmware images share a newly disclosed weakness (task C) and monitor the transaction stream that its gateways relay (task D). Tasks A to C are offline analyses that run in the cloud. Task D is an online analysis whose first stage should run on the gateway.
5
TABLE III O NE SCHEMA , FOUR INSTANTIATIONS ACROSS THE LAYERS OF A BLOCKCHAIN - ENABLED I OT SYSTEM . E ACH COLUMN LISTS THE ENTITIES THAT POPULATE THE FOUR ROLES AND THE MAIN RELATIONS , THE PREDICTION TARGET, AND THE TASK HEAD USED . Contract layer Task A, DeFiHack Web3Bugs
/
Device layer: vulnerability reports Task B, IoTVulBench
Device layer: registered fleet Task C, FirmVulLinker
Gateway layer: transaction stream Task D, EdgeChainGuard blockchain transactions smart contracts (registry, access control, data validation) device identities
Function Vf State Vs
contract functions state variables
request dispatcher, endpoint handler HTTP query/body parameters
firmware images CVE entries
Modifier Vm
modifiers
vendors device families
time windows
k-NN profile similarity between images image ↔ CVE (known links)
consecutive transactions of one device
Event Ve
events
authentication gates (cookie, basic auth, SOAPAction, none) expected service behavior (crash)
calls
f 1 → f2
dispatcher → handler
depends/ modifies constrained_by triggers/ affects
reads / writes of s
GET / POST parameter of the handler
f guarded by m f emits e; e indexed by s
handler guarded by gate handler → crash; crash ↔ most anomalous parameter
image ↔ vendor image ↔ family
Target Head Role here
contract is vulnerable graph classification original results (Section VI)
vulnerability types (multi-label) graph classification accuracy evaluation
unknown image→CVE links link prediction accuracy evaluation
a
transaction reads contract; high-gas transaction writes it transaction guarded by device identity transaction → window; window → active contracts transaction is malicious node classification deployment workload onlya
EdgeChainGuard is a synthetic dataset [4]: its generator assigns attack subtypes at random, and its binary attack label coincides with transaction failure. We therefore use it only to measure the cost of gateway–cloud inference on a realistic graph shape, not to report detection accuracy.
C. Unified Heterogeneous Schema We represent every artifact that the four tasks examine as a heterogeneous graph G = (V, E, τ, ρ),
(1)
where V is the node set, E the edge set, τ : V → {f, s, m, e} assigns each node one of four roles, and ρ : E → R assigns each edge one of nine relation types. The roles are V = Vf ∪ Vs ∪ Vm ∪ Ve ,
(2)
with Vf (function) the components that execute logic and Vs (state) the components that hold or carry state. The set Vm (modifier) holds the components that guard execution, and Ve (event) holds the components that expose observable traces. The relation types are R = {calls, depends, modifies, triggers, affects, constrained_by, returns_to, invokes, uses},
(3)
which describe execution (calls), data access (depends, modifies, uses), guarding (constrained_by, returns_to, invokes), and observability (triggers, affects). Each node v carries a feature vector xv ∈ Rdτ (v) whose dimensionality depends on its role. Table I and Table II give the contractlayer meaning of the roles and relations as introduced in the conference version. Table III shows how the same schema is populated at the device and gateway layers. A proof-of-concept request against a device becomes a small graph. In it, the request dispatcher and the endpoint handler execute, the request parameters carry state, the authentication gates guard, and the expected service behavior (a crash) is the trace. A registered fleet becomes one graph. In it, firmware images execute, CVE entries are the state they may carry, and vendors guard, since they determine which code base an image inherits. Product families are the trace through which a weakness becomes visible across devices. A device–contract transaction stream becomes a graph in which
transactions execute, the contracts they touch hold state, device identities guard, and time windows are the observable trace. In every case a logic vulnerability corresponds to a wrong or missing relation between roles. One example is a handler that modifies configuration state without being constrained_by a gate. Another is an image that should depend on a CVE but is not yet known to. A third is a burst of transactions from one identity that triggers a window unlike its history. The graph constructions are given in Section V-D. D. Learning Objectives The encoder produces role-specific node embeddings hv ∈ Rd for all v ∈ V , and the four tasks attach three kinds of objectives to it. Graph classification (tasks A and B). For a contract or a proof-of-concept trace, the model maps a pooled representation hG to a label vector ŷ ∈ [0, 1]L . Here L = 1 for binary contract vulnerability. For the multi-label vulnerability types of IoTVulBench (command injection, buffer overflow, denial of service, remote code execution), L = 4, and a trace may carry several labels. Link prediction (task C). For a fleet graph, the score of a pair (u, c) of an image u ∈ Vf and a CVE c ∈ Vs is h⊤ u hc , (4) ∥hu ∥ ∥hc ∥ with temperature κ. For each CVE, the model ranks candidate images by this score to retrieve the registered images that share the weakness. Node classification (task D). For a transaction stream, each function node (transaction) receives a label ŷv from its embedding. In all cases the decision depends jointly on the local attributes of individual components and on the typed relations through which state, guards, and observable traces interact. The conference version established this intuition for contracts, and this article carries it to the device and gateway layers. score(u, c) = κ ·
6
one shared encoder, four streams what is analysed
its own typed attention projection twice
one graph, four roles
three heads, four answers
agents exchange
graph head
a gateway runs this first stage
A contract vulnerable?
modifier
A
contract code
B
vulnerability report
guards writes
function
agent
state
agent
modifier
agent
function C
firmware images
D
transaction stream
emits
state
B which weakness types? fuse link head C which devices share it? node head D transactions to flag?
agent
event
event
Fig. 2. The MA-HGAT framework. Each of the four artifacts becomes one graph of the schema of Section III-C, in which four roles are joined by named links. The four roles then stay apart as four streams: each has its own projection and its own agent, while the two attention layers are shared. Neighboring agents exchange evidence through cross-attention, the streams are fused, and three heads answer the four questions.
IV. MA-HGAT A RCHITECTURE Fig. 2 gives an overview. The encoder consists of a rolespecific feature projection layer and two heterogeneous graph attention layers with relation-specific parameters. These are followed by a state-centered KVQ attention module, a multiagent decision module with one agent per role, and a gated cross-type fusion layer. Task heads (Section IV-F) are attached to the encoder output. The description of the encoder follows the conference version. The task heads, the gating in the fusion layer, the link scorer, and the strictly per-graph computation of the attention and fusion stages are new. A. Feature Projection Layer Functions, state entities, modifiers, and events encode fundamentally different roles. A shared linear transformation would therefore hide type-level inductive biases that matter for security reasoning. Each node type t ∈ {f, s, m, e} therefore has its own projection, (0) ht = ReLU LN(Wt xt + bt ) , (5) where Wt ∈ Rd×dt and bt ∈ Rd are type-specific. Layer normalization (LN) stabilizes scales across types, and the projection maps all types into a common hidden dimensionality d. Dropout follows the activation. At the device and gateway layers, this layer also absorbs the domain differences. A 37-dimensional byte-level firmware profile, a 10-dimensional HTTP-parameter descriptor, and a one-hot vendor indicator all enter the same encoder through their own Wt . B. Heterogeneous Graph Attention Layers For each relation r ∈ R we compute relation-aware attention coefficients exp(erij ) r , k∈N r exp(eik )
r αij =P
(6)
i
erij = LeakyReLU a⊤ r [Wr hi ⊕ Wr hj ] ,
where Nir is the set of neighbors of node i under relation r, Wr is a relation-specific projection, ar the corresponding
attention vector, and ⊕ concatenation. We aggregate messages across relations, X X (l+1) (l) r hi =σ αij Wr hj , (7) r∈R j∈Nir
with multi-head attention inside each relation. We stack two such layers and sum their outputs, so that both one-hop and two-hop relational context are retained, (1)
hi = hi
(2)
+ hi .
(8)
Keeping a separate Wr for each of the nine relations lets the model weight edges by their type. For instance, it can weight a modifies edge from an unauthenticated handler to a configuration parameter differently from a depends edge that only reads it. C. State Attention Module (Key–Value–Query) Logic vulnerabilities often arise from abnormal or unintended state transitions [35]: inconsistent updates to privileged variables, violations of invariants, or improper sequencing of persistent operations. To emphasize such transitions the encoder applies a KVQ attention module over the state nodes, QK ⊤ Attention(Q, K, V ) = softmax √ V, (9) dk Q = WQ H̃s , K = WK H̃s , V = WV H̃s , H̃s = Hs + P , (10) where Hs stacks the state-node embeddings and P is a learned positional encoding along the ordering of the state nodes in the graph. This ordering is declaration order for contract variables and parameter order within a request. Queries express what the model is searching for, keys encode candidate transition points, and values carry the semantic context of each state entity. The output is added residually to Hs with layer normalization. For contracts, this module concentrates on the storage operations that underlie price manipulation, privilege escalation, and atomicity failures. At the device layer, its usefulness depends on whether the state nodes carry rich attributes. We quantify this effect in Section VII-D. The attention is computed inside each graph. When several graphs are batched, keys and values
7
of other graphs are masked out, and the positional encoding indexes state nodes within their own graph.
positive links and uniformly sampled negative images (three per positive). Inference ranks all candidate images per CVE.
D. Multi-Agent Decision Module
Node classification. A two-layer perceptron maps hfinal of v each function node to class logits.
To separate reasoning perspectives, MA-HGAT uses four agents aligned with the node types: the Function agent (F), the State agent (S), the Modifier agent (M), and the Event agent (E). Each agent first refines the embeddings of its own type through a residual feed-forward block, Aa = LN Ha + FFNa (Ha ) , a ∈ {F, S, M, E}. (11) The agents then exchange information through cross-attention over their per-graph summary tokens āa = mean(Aa ) (the mean over the nodes of role a in the graph at hand), [cF , cS , cM , cE ] = MultiHeadAttention([āF , āS , āM , āE ]), (12) and each agent’s nodes receive the exchanged context through a learned gate, ga = σ Wg [Aa ⊕ ca ] , A′a = LN Aa + ga ⊙ ca . (13) This design encourages complementary reasoning instead of collapsing all evidence into one homogeneous embedding. For example, the Function agent may see an authorization flaw as an anomalous call pattern. The Modifier agent may see it as an unguarded write, and the Event agent as a missing audit trace. The gate lets each agent decide how much of the others’ view to absorb. E. Gated Cross-Type Fusion Layer The fusion layer integrates the four agents once more at the graph level. A second multi-head attention block maps the type summaries ā′a to fused context vectors ua . A residual gate then injects these vectors back into every node of the corresponding type, hfinal = LN h + σ W [h ⊕ u ] ⊙ u (14) v u v τ (v) τ (v) . v The conference version used plain concatenation and projection. The gate instead lets global context modulate nodespecific embeddings without overwriting them. This property is necessary for the node-level and link-level heads introduced next. F. Task Heads Graph classification. All nodes of all types are pooled with a learned scalar gate γv = σ(wγ⊤ hfinal ), v P γv hfinal v , ŷ = W2 ReLU(W1 hG +b1 )+b2 , hG = Pv∈V v∈V γv + ϵ (15) so the evidence for a graph-level decision may come from any node type. The gate is trained end-to-end to select the informative nodes. Link prediction. The scorer of (4) is applied to the final embeddings of a function node (firmware image) and a state node (CVE). Training uses binary cross-entropy over known
G. Loss Function and Training For the binary contract task the conference version combines binary cross-entropy with focal loss [70], L = αLBCE +(1−α)LFocal ,
LFocal = −
X
(1−pi )γ log pi ,
i
(16) to emphasize rare but high-impact positives. The multilabel IoT typing task uses per-label binary cross-entropy, the link task uses binary cross-entropy over sampled pairs, and the node task uses class-weighted cross-entropy. We optimize all models with AdamW [71], with dropout on the projection, attention, and fusion layers. We use early stopping or a fixed epoch budget as stated per experiment. The implementation is built on PyTorch and the Deep Graph Library [72]. A single MAHGAT class exposes the three heads through a mode argument and three ablation switches (use_state_attn, use_multi_agent, use_cross_fusion). For efficiency, the graphs of a fold are processed as one batched graph. The state attention, the agent summaries, the cross-attention, and the fusion context are computed with segment masks, so every graph is processed exactly as if it were alone. We verified that the output for a graph is identical up to single-precision rounding (below 2 × 10−7 ) whether or not it is batched with other graphs.
H. Contract Graph Construction For completeness we recall how contract graphs are built in the conference version. Solidity sources are compiled and parsed to obtain abstract syntax trees and control-flow graphs. Functions, state variables, modifiers, and events become nodes. Their attributes include visibility, parameter and return signatures, storage type and initialization, guarded conditions, and emission locations. Call edges are derived from invocation sites, and depends and modifies edges from state reads and writes. The constrained_by, returns_to, and invokes edges come from modifier applications and their control transfers. The triggers and affects edges come from event emissions and indexed parameters [27], [40], [73].
V. D ETECTION W ORKFLOW AND G ATEWAY–C LOUD D EPLOYMENT This section describes the five-tier workflow in which MAHGAT serves the four tasks of Section III. The first four tiers generalize the pipeline of the conference version. The fifth tier is new and covers deployment on gateways, and Fig. 2 marks the split it introduces.
8
A. Tiers 1–4: Ingestion, Graph Construction, Detection, and Triage Ingestion and normalization. We normalize raw artifacts into a common record format. These artifacts are Solidity projects from DeFiHack [74] and Web3Bugs [75], PoC traces and advisories from IoTVulBench, firmware images and ground-truth tables from FirmVulLinker, and transaction logs from EdgeChainGuard. For contracts, we resolve compiler versions and dependencies [76] and filter out incomplete sources [77]. For PoC traces, we parse the HTTP request line, headers, query string, and body. For firmware, we profile each image at the byte level. For transaction logs, we sort records by time and accumulate per-device histories causally. Heterogeneous graph construction. We convert each subject into the schema of Section III-C. The contract construction is summarized above. The device- and gateway-layer constructions are described in Section V-D. Every construction guarantees that each canonical relation exists, possibly as a placeholder edge, so that graphs of different subjects can be batched. Parallel multimodal detection. Three engines can run in parallel. The first is the MA-HGAT detector. The second is a rule-based detector that encodes interpretable structural rules. Examples at the device layer are shell metacharacters in parameters, oversized single-character runs, or unauthenticated access to privileged endpoints. The third is an optional LLMbased semantic analyzer that reasons over documentation and decompiled code [14], [44]. In the conference experiments, the three engines were combined through score normalization, confidence calibration, and cross-engine consistency filtering. In the device-layer experiments of this article, we evaluate MA-HGAT on its own and report the rule-based detector as a baseline. This keeps the accuracy attributable to the graph model from being confounded with ensemble effects. Dataset-aware triage. This tier turns the outputs into the artifacts an analyst needs. These are per-type probabilities for PoC traces, a ranked list of candidate homologous images per CVE, and per-transaction alerts with the device and contract they involve. Dataset-specific post-processing is also applied at this tier: proxy and flash-loan patterns for DeFiHack, multi-file dependencies for Web3Bugs, and vendor and family metadata for firmware fleets. B. Tier 5: Gateway–Cloud Partitioning Stream monitoring (task D) differs from the offline tasks in where the computation may run. Gateways see the transactions and requests of their own devices immediately, but they have little memory and intermittent uplinks. The context that gives a transaction its meaning lives in the cloud: which contracts and devices exist, and what happened in other time windows. The modules of MA-HGAT are aligned with roles, and the Function agent is the only module whose inputs a gateway observes. We therefore split the encoder as follows. Gateway tier. Each gateway runs the function-role projection Wf and the calls-relation attention of the first layer. That is, it passes messages along the intra-device temporal chain
of transactions it observes. Its output is a d-dimensional embedding per event, which is uploaded instead of the raw request. Cloud tier. The cloud receives the gateway embeddings as the function-node inputs. It runs the projections of the other three roles and completes the first layer over all relations. It then runs the second layer, the state attention, the multi-agent module, the fusion layer, and the task head. The gateway embedding already summarizes the intra-device temporal chain of each device, so the raw request features never leave the gateway. The model is trained end to end with the split in place, so the additional local pass is accounted for rather than approximated. This split has three consequences, which we quantify in Section VII-F. First, the gateway model is small: the function projection plus one relation-specific attention block. Second, the upstream payload is a fixed-size embedding rather than a variable-size request whose content may be sensitive. Third, the cloud latency is independent of the number of gateways, because the gateway work runs in parallel with no coordination. Unlike layer-wise split computing [67], [68], the boundary here is semantic. It follows the role that a gateway observes rather than a depth in the network. C. What the Split Protects and What It Does Not Moving one stage of the model to the gateway changes what an adversary can reach, so we state what the boundary buys and what it does not. Evasion does not become easier. The gateway computes the function projection and the calls attention over the transactions of one device. An adversary who controls that device and knows the boundary can shape that chain. The chain is one input to the decision among several. The cloud still projects the other three roles, completes the first layer over all nine relations, and runs the second layer, the state attention, the agents, and the head. Those inputs are the contracts, the device identities registered on the chain, and the time windows. A single compromised device does not control them. The model is also trained end to end with the split in place. The gateway stage is part of one function, not a filter in front of it. Section VII-F confirms that the decisions match cloud-only inference. Content minimization is not a privacy guarantee. The uplink carries a d-dimensional embedding instead of the request. That embedding is trained for the detection objective and carries no bound on what it reveals. An adversary who holds the gateway model and observes its outputs may recover properties of the inputs, as the work on collaborative-inference inversion [78] and on leakage from embeddings [79] shows. Our claim is therefore the narrow one: request contents stay on the device side of the uplink, and the cloud never stores them. An operator who needs a stated bound has to add a mechanism such as calibrated noise, and we did not measure what that would cost in accuracy. A compromised gateway is not contained. The threat model assumes that gateways are not tampered with, and that
9
assumption carries weight here. A compromised gateway can upload any embedding for its own devices. Those embeddings enter the shared graph, so the second layer and the fusion carry them to nodes of the other roles and then to other devices. The boundary does not bound this influence. An operator who cannot trust gateways needs per-gateway rate limits or calibration on top of the model. Extraction of the gateway model gives one stage, not the decision. The 20.5 KB that a gateway holds are the function projection and one relation-specific attention block, 2.5% of the parameters. An adversary who extracts them learns neither the head nor the modules for the other three roles, so the decision function does not follow from the gateway alone. Extraction does give white-box access to the stage that the adversary’s own traffic passes through, which is the usual starting point for building evasive inputs. D. Device- and Gateway-Layer Graph Constructions Vulnerability reports (IoTVulBench, task B). Each CVE in IoTVulBench [80] ships a curated advisory and a raw HTTP PoC payload. The advisory (detail.yml) gives the name, description, CVSS score, severity, and tags. The payload reproduces the vulnerability against an emulated router. These routers are the kind of device whose management interface a gateway or administrator talks to in the system of Section III. We parse the payload into a graph as follows. There are two function nodes: the request dispatcher, which carries the method, and the endpoint handler. The handler carries path descriptors such as depth, length, and the presence of goform, HNAP, userRpm, or cgi, together with requestlevel statistics. There is one state node per query or body parameter. Its features are length, Shannon entropy, density of shell metacharacters, longest single-character run, URLencoding density, path-traversal markers, and credential-like names. There is one modifier node per authentication gate found in the headers: cookie, basic authorization, SOAPAction, or an explicit no authentication node. There is one event node for the expected service behavior (a crash, attributed with the request length). GET parameters attach to the handler through depends, and POST parameters attach through modifies. Gates attach through constrained_by, invokes, and returns_to. A cookie gate attaches through uses to the parameter whose name it references, or to the first parameter when none matches. The event node is triggered by the handler and affects the most anomalous parameter. Labels are taken only from the curated advisory text, and features only from the payload. The CVSS score of the advisory is not used as a feature. Registered fleet (FirmVulLinker, task C). FirmVulLinker [25] releases 54 known-defective firmware images (TP-Link and D-Link routers, 24 device families) and a ground-truth table of 74 CVEs. Each CVE entry lists a baseline image and the other images it affects (11.3 on average, from 1 to 34). We build one fleet graph. Firmware images are function nodes, each described by a 37-dimensional byte-level profile computed from the raw image bytes. When the image is an archive, we profile
its largest member. The profile holds the log size, overall and 4-KB block entropy statistics, a nibble histogram, and printable-string density and length. It also holds the normalized frequency, over the first 8 MB, of 14 tokens such as httpd, goform, login, and password. CVEs are state nodes (disclosure year), vendors are modifier nodes, and device families are event nodes. Images connect to their vendor (constrained_by), to their family (triggers), and to their three most similar images by cosine similarity of the profiles (calls). For training links only, images also connect to the CVEs known to affect them (depends). Device–contract transaction stream (EdgeChainGuard, task D). EdgeChainGuard [4] provides 500 blockchainmediated IoT transactions from 50 devices against three contracts (access control, device registry, and data validation). This is exactly the traffic that the gateway layer of Section III relays. Transactions are function nodes with causal features only. These features are the normalized gas fee, time of day, contract indicator, log inter-arrival time, burst count, and the prior failure rate and history length of the device. Contracts are state nodes, devices are modifier nodes, and twenty time windows are event nodes. Consecutive transactions of one device are chained by calls. A transaction depends on its contract and modifies it when its gas fee is in the top quartile. Devices invoke and constrain their transactions and use the contracts they touch. Windows are triggered by transactions and affect the contracts active in them. As noted in Table III, this graph is used only as a deployment workload. VI. C ONTRACT-L AYER R ESULTS This section summarizes the contract-layer evaluation (task A) of the conference version [1]. All numbers in this section are those reported there. The datasets are DeFi contracts rather than device-registry contracts, because no labeled corpus of IoT-specific contracts exists. The defect classes they cover are access-control violations, state-update ordering, and arithmetic logic. These are the classes that registry, access-control, and validation contracts are exposed to. A. Setup The conference version used two datasets with complementary characteristics. DeFiHack [74] contains 663 contracts with manually verified logic-vulnerability annotations. These cover DeFi attack scenarios such as price manipulation, flashloan attacks, reentrancy exploits, and state-inconsistency vulnerabilities. Web3Bugs [75] is a corpus of exploitable smart contract bugs. We use the 72-project subset released with the GPTScan evaluation [81], which spans access-control weaknesses, arithmetic errors, and logic design flaws. The contracts were compiled with solc 0.8.19 and converted into heterogeneous graphs. Each dataset was split 70/15/15 into training, validation, and test sets, with no contract shared across splits. The model used hidden size 256, four attention heads, dropout 0.2, AdamW with learning rate 10−3 , weight decay 0.01, batch size 32, and early stopping with patience 10. The reported metrics were accuracy, precision, recall, and F1-score.
10
TABLE IV M AIN DETECTION RESULTS ON D E F I H ACK AS REPORTED IN THE CONFERENCE VERSION [1] (%).
TABLE V A BLATION ON D E F I H ACK AS REPORTED IN THE CONFERENCE VERSION (%).
Method
Acc.
Prec.
Rec.
F1
Configuration
Acc.
Prec.
Rec.
F1
Slither [9] Mythril [10] Smartian [40] ContractFuzzer [12]
72.34 68.92 76.89 74.56
75.21 71.34 78.23 76.12
68.45 65.78 74.56 72.34
71.67 68.45 76.35 74.19
GCN+VulDetector ReGNN
79.23 81.56
80.67 82.34
77.45 80.23
79.03 81.27
MA-HGAT+DeepSeek-R1
87.04
86.55
83.51
85.35
Full system Without GNN layers Without MA-HGAT detector Without LLM analyzer Without rule detector Without CFG generator LLM-only analysis Rule-only detection
94.95 78.34 82.45 88.23 79.56 91.78 65.42 58.93
89.46 72.89 78.32 84.67 75.12 87.23 61.78 55.67
85.13 69.45 76.58 82.15 72.34 83.67 58.34 52.89
87.24 71.13 77.44 83.39 73.71 85.42 60.01 54.25
(a) DeFiHack (a) DeFiHack (n = 663)
(b) Web3Bugs (n = 72)
(b) Web3Bugs
Fig. 4. Effect of removing individual components on the contract datasets.
Fig. 3. Comparison with LLM-based detectors on the two contract datasets.
B. Main Results and Baseline Comparison Table IV summarizes the main results. MA-HGAT combined with the DeepSeek-R1 semantic analyzer achieved 87.04% accuracy, 86.55% precision, 83.51% recall, and 85.35% F1 on DeFiHack. On Web3Bugs it achieved 87.45% accuracy, 88.97% precision, 85.54% recall, and 88.42% F1. Compared with symbolic tools (Slither [9], Mythril [10], Smartian [40]) and fuzzing (ContractFuzzer [12]), the framework yielded consistently higher recall and F1. Relative to the two learning-based baselines of Table IV, it improved F1 by 4.08 points over the stronger one. C. Comparison with LLM-Based Detectors Fig. 3 compares MA-HGAT with LLM-based detectors. On DeFiHack, MA-HGAT (87.04%) outperformed Smart-LLaMA-DPO [15] (82.35%), GPTScan [14] (79.68%), VulnHunt-GPT [44] (76.42%), and general-purpose models such as DeepSeek-R1 [82] (85.23%) and GPT-4 [83] (80.36%). On Web3Bugs the margin over Smart-LLaMA-DPO grew to 6.22 points. Five-fold cross-validation gave mean accuracies of 87.04% (±0.42) on DeFiHack and 87.45% (±0.89) on Web3Bugs. Bootstrap confidence intervals supported the differences to the LLM baselines. D. Component Integration and Ablation Table V reports the ablation of the conference version on DeFiHack. Fig. 4 shows the effect of removing individual components on both datasets. The full ensemble system combined the graph detector, the rule-based detector, and the LLM analyzer by weighted voting, and it reached 94.95% accuracy. Within this system, removing the GNN layers caused the largest drop (16.61 points), followed by the rule detector (15.39) and the MA-HGAT detector (12.50). Within the graph model, removing the state attention module reduced
accuracy by 3.22 points, and removing the multi-agent module reduced it by 4.89. Removing cross-type fusion reduced accuracy by 4.33 points, and collapsing all node types into a homogeneous graph reduced it by 7.11. Per-type analysis showed the strongest results on logic-bypass (86.63% F1) and mathematical-logic (84.43% F1) vulnerabilities on DeFiHack. On Web3Bugs, the strongest results were on access-control (90.12%) and reentrancy (89.45%) vulnerabilities (Fig. 5). These results were obtained on contract graphs with tens to hundreds of typed nodes and richly populated relations. They established that heterogeneous relational modeling, statecentric attention, and multi-agent fusion each contribute measurably on such graphs. The next section asks whether the same holds at the device layer and what the gateway deployment costs. VII. D EVICE -L AYER AND D EPLOYMENT R ESULTS We now evaluate tasks B and C at the device layer and the deployment of task D at the gateway layer. The released code (per-graph implementation, payload-only features) produces all numbers in this section. Per-fold and per-seed values, persample predictions, and the paired tests are released with it. Four questions guide the evaluation. RQ1: Does MA-HGAT triage vulnerability reports and link registered firmware images to disclosed weaknesses at least as well as strong homogeneous, tabular, and rule-based alternatives trained on the same information? RQ2: Which of its components matter at the device layer, and why? RQ3: How much of that accuracy survives when the tested devices come from a vendor absent from training? RQ4: What does the role-aligned gateway– cloud split cost, and does it change decisions? A. Datasets, Tasks, and Protocol We evaluate the four tasks on four datasets rather than on one deployment, because no public dataset couples the
11
(a) DeFiHack
(b) Web3Bugs
Fig. 5. Per-type detection performance on the contract datasets.
contracts, the firmware, and the transaction stream of a single blockchain-enabled IoT system. Each dataset is the largest public one we found for its layer. IoTVulBench (task B, vulnerability-report triage). IoTVulBench [80] provides reproducible vulnerability environments for consumer routers (TP-Link, D-Link, Tenda, and others), each with a curated advisory and a raw PoC request. We use the 95 CVEs that have both. We derive the multi-label target from the advisory tags and text only: command injection (43 CVEs), buffer overflow (43), denial of service (60), and remote code execution (25). Every CVE carries at least one label, and 60 carry exactly two. We build the graphs as in Section V-D from the payload only. Each graph has two function nodes, one to fifteen state nodes (2.3 on average), one or two modifier nodes (only five requests are unauthenticated), and one event node. Two thirds of the graphs have a single state node; for 19 requests that node is a placeholder because no parameter could be parsed. We use five-fold cross-validation stratified on the label combination. Within each fold, we average the predicted probabilities of three random seeds before thresholding at 0.5. We report sample-wise accuracy (fraction of correct label decisions), macro-F1 over the four labels, exact-match rate, and per-label F1. Standard deviations are over folds. As a secondary check, we use a second target on the same graphs: whether the advisory rates the CVE critical (CVSS ≥ 9.0; 47 of 95). FirmVulLinker (task C, fleet linking). The fleet graph of Section V-D covers 54 images (46 TP-Link, 8 D-Link), 74 CVE groups, and 24 device families. Each CVE affects 11.3 images on average beyond its baseline (1 to 34). For each CVE, we keep its baseline image and a random 70% of the affected images as known training links. We hold out the remaining 30% (at least one image) as test links. At inference, (4) ranks all images except the baseline and the known positives. We report mean reciprocal rank (MRR), Hits@K for K ∈ {1, 3, 5, 10}, and the area under the ROC curve (AUC) per CVE, averaged over CVEs and over three random splits and seeds. EdgeChainGuard (task D, deployment workload). We use the 500-transaction graph (500 function, 3 state, 50 modifier, and 20 event nodes) to measure the gateway–cloud split in Section VII-F. Its generator assigns attack subtypes stochastically, and its binary attack label coincides with transaction failure [4]. Every detector, including simple rules and random
forests, therefore stays within a few points of the 53% majority rate on it. We do not report it as a detection result. Hyperparameters. For IoTVulBench: hidden size 64, four heads, dropout 0.25, AdamW with learning rate 10−3 and weight decay 10−3 , 150 full-batch epochs, per-label binary cross-entropy. For FirmVulLinker: hidden size 128, four heads, dropout 0.2, AdamW with learning rate 2 × 10−3 and weight decay 10−4 , 300 epochs, temperature κ = 10, three negatives per positive. We tuned no hyperparameter on test data. Baselines. The graph baselines receive exactly the same graphs as MA-HGAT. The tabular and rule baselines receive payload statistics only. Homogeneous GraphSAGE (labeled Homogeneous GCN in the released code) and Homogeneous GAT collapse the heterogeneous graph into a single node and edge type. Features are zero-padded or truncated to a common dimensionality. They use two layers of mean aggregation [59] or graph attention [60] with the same hidden size, optimizer, and epochs, followed by mean pooling for graph-level tasks. In the linking task they use the same scorer and negative sampling. The random forest operates on hand-crafted graphlevel statistics. For PoC traces, these are twelve request-level statistics such as payload length, maximum parameter length and run, entropy, metacharacter count, and endpoint indicators. For linking, they are the concatenated and differenced profiles of the candidate and the baseline image, together with vendor and family agreement. The rule-based detector flags injection when shell metacharacters or command tokens appear. It flags overflow when a parameter reaches 200 characters or contains a run of 150 identical characters. It flags denial of service when such a run appears or the request mentions reboot or ping. It flags code execution when injection indicators or such a run are present. For linking we add two more baselines. A cosine profile baseline ranks images by profile similarity to the CVE’s baseline image. A metadata rule ranks images of the same family first and of the same vendor second. Ablations disable the state attention, the multi-agent module, or the cross-type fusion. Statistical tests. For task B we compare methods with a paired bootstrap over the 95 CVEs (10,000 resamples of the macro-F1 difference) and McNemar’s test on exact-match correctness. For task C we use a paired bootstrap and a Wilcoxon signed-rank test over the 74 per-CVE reciprocal ranks (averaged over the three splits). We call a difference significant when the 95% bootstrap interval excludes zero. For the rank-based linking comparisons, whose per-CVE differences are heavy-tailed, we also call it significant when the Wilcoxon test gives p < 0.05. We report both statistics wherever they disagree. B. RQ1, Task B: Vulnerability-Report Triage on IoTVulBench Table VI and Fig. 6(b) report the results. MA-HGAT reaches 81.6% sample accuracy, 78.1% macro-F1, and 55.8% exact match. It is clearly better than the rule-based detector. The macro-F1 difference of 9.8 points has a 95% interval of [3.4, 16.1]. McNemar’s test on exact matches counts 52 CVEs that MA-HGAT labels completely right and the rules do not, against 8 in the other direction (p < 0.001). The rules fire often
12
TABLE VI TASK B: MULTI - LABEL VULNERABILITY- REPORT TRIAGE ON I OTV UL B ENCH (95 CVE S , FIVE - FOLD CV; GNN S AVERAGE THREE SEEDS PER FOLD ; %). B EST VALUE PER COLUMN IN BOLD , SECOND BEST UNDERLINED . ACC . IS SAMPLE - WISE ACCURACY OVER THE FOUR LABEL DECISIONS . E XACT IS THE FRACTION OF CVE S WHOSE ENTIRE LABEL VECTOR IS CORRECT. CI = COMMAND INJECTION , BO = BUFFER OVERFLOW, D O S = DENIAL OF SERVICE , RCE = REMOTE CODE EXECUTION . Method
Acc.↑
Macro-F1↑
Exact↑
F1 CI
F1 BO
F1 DoS
F1 RCE
MA-HGAT (full) w/o state attention w/o multi-agent module w/o cross-type fusion
81.6±3.5 81.1±2.6 80.8±3.5 81.3±3.3
78.1±4.4 77.7±3.6 78.0±4.7 77.9±3.8
55.8 55.8 52.6 54.7
82.5 77.9 85.1 82.5
90.7 90.7 89.4 90.7
79.7 79.7 76.6 78.9
59.6 62.5 60.9 59.6
Homogeneous GraphSAGE Homogeneous GAT Random forest (payload statistics) Rule-based detector
82.6±1.5 80.0±0.5 83.9±2.9 63.2
77.6±1.7 76.1±2.6 78.7±5.8 68.3
55.8 46.3 52.6 9.5
84.7 83.0 87.4 83.7
89.4 90.7 95.3 74.3
80.6 76.9 78.4 66.2
55.8 53.7 53.7 49.0
(89.8% recall at 57.5% precision) and label only 9.5% of the CVEs completely right. Against the learned baselines, MAHGAT has the highest macro-F1 among the graph models and the highest exact-match rate (tied with GraphSAGE). None of these differences is significant. The macro-F1 difference is +0.5 points against GraphSAGE (interval [−3.0, 4.4]) and −0.6 points against the random forest ([−5.8, 5.1]). The exactmatch comparison against GraphSAGE is a 7-to-7 tie. The only learned baseline that MA-HGAT beats on the paired test is the homogeneous GAT, on exact match (11 CVEs to 2, p = 0.027). The random forest is the strongest model for buffer overflow (95.3% F1). There, a single scalar, the longest parameter run, probably decides the label. Tied with the homogeneous GAT, it is also the weakest learned model for remote code execution (53.7%). MA-HGAT and its variants are the strongest there (59.6–62.5%). For this label, we would expect that evidence from the parameter content, the authentication gate, and the endpoint must be combined. On the secondary severity target, all learned models separate critical from non-critical advisories well above the 50.5% majority rate from payload-only graphs. MA-HGAT reaches 71.6% (±7.1) accuracy and 76.1% F1, GraphSAGE 75.8% and 78.9%, GAT 73.7% and 77.9%, and the random forest 72.6% and 72.3%. The fold-level standard deviations (6–8 points on 19 CVEs per fold) exceed every pairwise difference. Our reading is that typed attention has little to attend over on request graphs with five to nineteen nodes. One third of the graphs have exactly five nodes, and two thirds have a single state node. The information that separates the four classes appears to be concentrated in a few parameter statistics that every method receives. A mean aggregator recovers it as well as relation-specific attention does. MA-HGAT’s higher fold-to-fold variance (±4.4 versus ±1.7 for GraphSAGE) is consistent with a higher-capacity model trained on 76 graphs per fold. The graphs are this small because we build them from the request alone. Recovering the handler that parses the request, with its control and data flow, as Karonte [17] and SaTC [18] do, would populate the schema with far more nodes and relations. The contract results indicate that this is the density at which typed attention starts to pay.
C. RQ1, Task C: Fleet Linking on FirmVulLinker Table VII and Fig. 6(a) report the linking results. Learning over the fleet graph is clearly better than the non-graph heuristics on which vendor advisories implicitly rely. MA-HGAT reaches 0.586 MRR, 91.6% Hits@10, and 0.960 AUC. Raw profile similarity reaches 0.541 MRR and 72.0% Hits@10, and the vendor/family rule 0.537 MRR and 70.2% Hits@10. The per-CVE MRR differences of +0.045 and +0.049 are significant by the Wilcoxon test (p < 0.002). A few CVEs on which MA-HGAT ranks poorly, however, widen the bootstrap intervals to include zero. When an operator checks the top ten candidates per CVE, 92% of the held-out affected images appear in that list on average with MA-HGAT. With the heuristics, 70–72% do. The pairwise random forest is between them (0.566 MRR, 88.8% Hits@10; the difference to MAHGAT is not significant). The homogeneous GAT is clearly worse (0.529 MRR, p < 0.001). The homogeneous GraphSAGE baseline, however, is the best model on this task on every metric (0.628 MRR, 96.5% Hits@10, 0.976 AUC). Its advantage over MA-HGAT of 0.042 MRR is close to the significance threshold (bootstrap interval [−0.088, −0.002], p = 0.035; Wilcoxon p = 0.073). MA-HGAT is also less stable across splits (±0.027 MRR versus ±0.011). The fleet graph has only two modifier entities (vendors) and 24 event entities (families), and CVE nodes carry a single attribute. All discriminative content is in the 37-dimensional firmware profile of the function nodes and in the k-NN and CVE links between them. A mean aggregator over the collapsed graph therefore loses little. Attention that must be learned from 54 images and 74 CVEs probably adds variance rather than signal. We attribute the gap over the heuristics to the graph structure itself, which both graph models exploit. The depends links to known affected images can propagate CVE membership across the k-NN calls edges and the shared vendor and family nodes. This is the code-reuse structure that makes vulnerabilities recur [24], [25]. D. RQ2: Ablation Across Tasks The ablation rows of Table VI and Table VII, together with the paired tests, give a consistent picture of which components matter at the device layer. Only the multi-agent module helps on both tasks. On IoTVulBench, removing it lowers the exact-match rate from
13
100 MA-HGAT (full)
90
w/o state attn.
Hits@K (%)
80
w/o multi-agent
70
w/o cross-type fusion
60
Homo. GraphSAGE Homo. GAT
50 MA-HGAT Homo. GraphSAGE Homo. GAT
40 30
1
3
5
Random Forest Cosine profile Metadata rule
K
Random Forest
macro-F1 sample accuracy
Rule-based
10
40
(a) Task C, FirmVulLinker: ranking registered images per CVE
50
60
70
score (%)
80
90
(b) Task B, IoTVulBench: multi-label vulnerability-report triage
Fig. 6. Device-layer accuracy results. (a) Hits@K on FirmVulLinker for the six main methods; error bars are standard deviations over three splits. (b) Macro-F1 (filled) and sample accuracy (hollow) on IoTVulBench; error bars are standard deviations over five folds. Blue bars are MA-HGAT variants, orange bars are baselines; the hatched bar is the full model. TABLE VII TASK C: LINKING REGISTERED FIRMWARE IMAGES TO DISCLOSED CVE S ON F IRM V UL L INKER (54 IMAGES , 74 CVE GROUPS ; MEAN±STD OVER THREE RANDOM SPLITS AND SEEDS ). H ITS @K IN %. B EST VALUE PER COLUMN IN BOLD , SECOND BEST UNDERLINED . Method
MRR↑
Hits@1↑
Hits@3↑
Hits@5↑
Hits@10↑
AUC↑
MA-HGAT (full) w/o state attention w/o multi-agent module w/o cross-type fusion
0.586±0.027 0.598±0.015 0.574±0.018 0.602±0.012
44.4±3.5 45.7±2.2 43.7±2.5 46.1±2.4
65.2±2.7 66.6±1.3 62.8±2.0 68.0±0.5
76.1±3.0 77.2±1.0 74.7±1.5 77.1±2.0
91.6±2.7 93.7±1.4 91.8±1.6 93.1±1.8
0.960±0.012 0.966±0.005 0.957±0.008 0.967±0.006
Homogeneous GraphSAGE Homogeneous GAT Random forest (pairwise profiles) Cosine similarity of profiles Vendor/family metadata rule
0.628±0.011 0.529±0.009 0.566±0.021 0.541±0.010 0.537±0.012
47.3±1.5 37.5±0.8 40.9±4.8 45.2±1.2 44.2±2.2
73.0±1.0 60.5±1.2 65.4±1.4 56.9±1.4 58.4±0.7
83.3±0.5 71.2±0.9 74.2±0.5 61.4±1.2 63.1±0.1
96.5±0.3 84.9±1.7 88.8±1.2 72.0±0.6 70.2±0.2
0.976±0.003 0.916±0.004 0.946±0.008 0.797±0.003 0.836±0.003
55.8% to 52.6% (macro-F1 changes by 0.1 points). On FirmVulLinker, removing it lowers MRR from 0.586 to 0.574 and Hits@3 from 65.2% to 62.8%. The per-CVE reciprocal ranks are higher with the module on 22 CVEs and lower on 7, with 45 ties (Wilcoxon p = 0.005). The bootstrap interval is [−0.001, 0.023], so the effect does not meet our interval criterion. The effect is small, but it is the one that transfers. We think the per-role refinement and the gated exchange between agents let the sparse vendor and family nodes influence the image embeddings without being averaged away. Cross-type fusion is neutral or slightly harmful. On IoTVulBench, removing it changes macro-F1 by 0.2 points and exact match by 1.1 points (neither significant). On FirmVulLinker, removing it raises MRR from 0.586 to 0.602 (not significant, p = 0.19). The fusion layer re-injects graph-level context. In a single fleet graph that context is identical for every pair and only adds parameters. State attention helps only when states carry structure. On IoTVulBench, state nodes are HTTP parameters with ten attributes each, and their order matters. There, removing the KVQ module costs 0.4 macro-F1 points and 4.6 F1 points on command injection, the label most tied to parameter content.
Exact match is unchanged. On FirmVulLinker, a state node is a CVE with a single year attribute. There, removing the module slightly improves MRR (0.598 versus 0.586, p = 0.10) and Hits@10 (93.7% versus 91.6%). We attribute this to the module’s positional encoding and softmax over 74 nearidentical CVE embeddings, which have no structure to exploit. This mirrors, in reverse, the conference finding that state attention was worth 3.22 accuracy points on contracts, where state variables are the richest role. Takeaway. Typed heterogeneity pays off when (i) several roles carry informative attributes and (ii) the graphs are large enough for relation-specific attention to be estimated. Contract graphs satisfy both. The two device-layer benchmarks available today satisfy neither, and there a mean aggregator over the same graph is at least as good. At the device layer, the value of the schema is not higher accuracy. It lies in representing both layers of the system with one model and in the deployment property of Section VII-F. E. RQ3: Generalization to an Unseen Vendor The protocols above draw training and test cases from the same pool of devices, so we reran both tasks under a vendor
14
TABLE VIII V ENDOR HOLDOUT. TASK B REPORTS MACRO -F1, TASK C MEAN RECIPROCAL RANK , BOTH IN %. “I N DIST.” REPEATS TABLE VI AND TABLE VII. T HE HEURISTIC ROW IS THE RULE - BASED DETECTOR FOR TASK B AND THE VENDOR / FAMILY RULE FOR TASK C; THE RANDOM FOREST USES PAYLOAD STATISTICS FOR TASK B AND PAIRWISE PROFILES FOR TASK C. Task B, F1
Task C, MRR
Method
In dist. Unseen In dist. TP→D D→TP
MA-HGAT (full) Homogeneous GraphSAGE Homogeneous GAT Random forest Heuristic (no training) Cosine profiles
78.1 77.6 76.1 78.7 68.3 –
59.4 55.3 60.5 53.6 68.3 –
58.6 62.8 52.9 56.6 53.7 54.1
51.6 56.2 55.6 3.9 63.4 56.4
16.5 16.7 13.4 8.5 14.0 20.2
Random ranking
–
–
–
8.7
8.7
TABLE IX G ATEWAY– CLOUD PARTITIONING OF MA-HGAT ON THE 500- TRANSACTION E DGE C HAIN G UARD WINDOW (CPU INFERENCE , ONE THREAD ). Quantity Parameters at the gateway (of 204,802) Parameter footprint, gateway (KB) Parameter footprint, cloud (KB) Activation bytes, gateway (KB)a Activation bytes, cloud (KB)a Embedding buffer per window, gateway (KB) Latency, gateway tier (ms) Latency, cloud tier (ms) Latency, end to end (ms) Uplink per event (B)b Uplink per 500-event window (KB) Accuracy (%)c Macro-F1 (%)c
Cloud-only
Split
0 5,120 (2.5%) – 20.5 819.2 798.7 – 1,038 8,983 8,983 – 128 – 2.04 31.22 30.67 31.22 31.66 2,100 256 1,050 128 57.5±0.7 56.5±1.1 57.0±1.1 55.8±1.1
a Sum of the sizes of all tensors emitted by the modules during one forward
holdout. For task B we resolve each CVE’s vendor from the advisory text alone (54 D-Link, 24 Tenda, 17 TP-Link), leave one vendor out at a time, and pool the three folds. For task C the CVE groups are vendor-pure, so we train on one vendor’s groups and hold out every link of the other’s except the query image that defines each CVE. Everything else stays as in Section VII-A. Neither task transfers across vendors. On task B every learned model loses 19 to 25 macro-F1 points, and exact match falls from 55.8% to 21.1% for MA-HGAT. No difference between MA-HGAT and any baseline is significant on the paired bootstrap, and the widest interval is [−10.3, 1.8] points against the random forest. The rule-based detector does not train, so its row is unchanged, and under the shift it has the highest macro-F1 of all. On task C the two directions disagree because of the corpus composition. Only 8 of the 54 images are D-Link, so testing on D-Link rewards anything that ranks the query’s vendor first, and the metadata rule beats MAHGAT there (0.634 against 0.516, interval [−0.21, −0.04]). The reverse direction removes that shortcut. Every method then falls near profile similarity and far below the 0.586 of the in-distribution split. Reading. Both benchmarks describe a device through vendorspecific surfaces. These are HTTP endpoints and parameter names for task B, and firmware layout for task C. Under a vendor shift the evidence a model relies on changes, and the learned mapping does not follow it. Inside the device population it was trained on, the framework is competitive with strong baselines. Outside it, an operator should retrain rather than transfer. Carrying attributes that describe what a handler does, rather than where it sits, is the concrete next step for the schema. F. RQ4, Task D: Gateway–Cloud Deployment We measure the split of Section V-B on the EdgeChainGuard stream (500 transactions, 50 devices) with the IoTVulBench PoC requests as the reference for raw request sizes. All measurements are CPU inference in PyTorch with intraop threads pinned to one, 10 warm-up and 50 timed repetitions (20 for the per-gateway timings of Fig. 7(a)). The timed cloudonly forward pass corresponds to the full model of Section IV
pass over the window, an upper bound on the live activation working set. The cloud figure is for the full model. b Cloud-only uploads the raw request; the value is the mean of the 95 IoTVulBench PoC requests (median 1,084 B, maximum 47,064 B). The split uploads one 64-dimensional embedding. A pre-extracted 10dimensional feature vector would be 40 B. c Mean±std over three seeds on a held-out 200-transaction split (majority rate 50.5%; 300 training transactions, unweighted cross-entropy, 120 epochs). The split variant is trained end to end with the full split pipeline, including the fusion layer. Both values are near chance because the workload’s labels are synthetic. The row is a parity check, not a detection result.
on the whole window. The gateway-tier figure in Table IX includes the construction of the gateway’s local graph, whereas the per-gateway figures in Fig. 7(a) time the forward pass only. Table IX and Fig. 7 summarize the results. The gateway model is tiny. The Function agent’s gateway share (the function projection and the calls attention block) has 5,120 parameters, 2.5% of the 204,802-parameter model. That is 20.5 KB in single precision versus 819 KB for the full model. Its activations for a 500-event window sum to at most 1.04 MB (2.1 KB per event), one ninth of the 8.98 MB of the full model. Its output buffer is 128 KB (500 embeddings of 64 floats). A footprint of this size fits the gateway-class devices that relay device transactions. The split is almost free in latency. Cloud-only inference over the window takes 31.22 ms. The split takes 2.04 ms at the gateway (including the construction of the local graph) and 30.67 ms in the cloud. Measured end to end, it takes 31.66 ms, within 1.5% of the cloud-only figure. The conservative sum of the gateway and cloud tiers is 32.7 ms, within 5%. As Fig. 7(a) shows, the cloud time does not change with the number of gateways. The time of the slowest gateway stays between 1.2 and 1.7 ms from 1 to 64 gateways. The end-to-end latency of a window therefore stays between 31.9 and 32.4 ms for any partition, and a single gateway sustains about 3.3 × 105 events per second. These times come from a general-purpose CPU pinned to one thread, not from gateway hardware, so they transfer as ratios and not as absolute values. The parameter and activation footprints above transfer directly. Raw requests stay at the gateway. A gateway uploads a 256-byte embedding per event instead of the raw request.
cloud tier (remaining agents, fusion, head) gateway tier (function agent, slowest gateway)
40 30
1.69
1.51
1.60
1.58
1.59
1.22
1.65
20 10 0
1
2
4
8
16
number of edge gateways
32
64
(a) latency of the 500-transaction window (cloud + slowest gateway)
105
bytes per event (log)
end-to-end latency (ms)
15
47,064 B
104 2,100 B
10
3
1,084 B 256 B
102 101
40 B
raw trace (max)
raw trace (mean)
raw trace (median)
gateway engineered embedding features
(b) upstream payload per event (raw sizes from 95 real PoC traces)
Fig. 7. Cost of the gateway–cloud split. (a) End-to-end latency for the 500-transaction window as the window is partitioned over 1–64 gateways. The cloud tier (hollow, 30.67 ms) is independent of the number of gateways. The gateway tier (hatched, numbers above bars) is the latency of the slowest gateway. (b) Upstream payload per event: the 256-byte gateway embedding versus the raw PoC requests of IoTVulBench (mean, median, maximum) and a pre-extracted 40-byte feature vector.
In IoTVulBench, a raw request averages 2,100 B (median 1,084 B, maximum 47,064 B), so the reduction is 8.2× on average and 184× in the worst case. The embedding is larger than a pre-extracted 40-byte feature vector, so the benefit over feature-level offloading is not bandwidth. Instead, the gateway does not need to run and update a domain-specific feature extractor. Request contents, which may embed credentials or configuration values, also never leave the device side. Decisions do not change. Over three seeds, the model trained end to end with the split in place, including the fusion layer, reaches 56.5% (±1.1) accuracy and 55.8% (±1.1) macroF1. Cloud-only training reaches 57.5% (±0.7) and 57.0% (±1.1). The difference of about one point is comparable to the seed-to-seed spread. Both values are near chance because this workload’s labels are synthetic (Section VII-A). The row is a parity check, not a detection result. VIII. C ONCLUSION Blockchain-enabled Internet of Things systems rely on security-critical logic in both smart contracts and device firmware, yet existing approaches typically analyze these layers separately. This work addressed this gap by extending MA-HGAT into a unified cross-layer framework that represents contracts, firmware evidence, device fleets, and device– contract transactions through the same role-based structure. The framework supports contract auditing, vulnerability-report triage, fleet-level vulnerability linking, and transaction monitoring within one model, while allowing the device-action component to run on resource-constrained gateways. Our evaluation shows that MA-HGAT is effective across both contract and device layers and that the role-aligned multiagent design is the main source of cross-layer transfer. The gateway–cloud partition further reduces local computation and avoids transmitting raw requests with little impact on decision latency. At the same time, the unseen-vendor evaluation reveals a clear generalization limitation. Future work will focus
on richer firmware-derived representations, broader labeled datasets, and end-to-end evaluation on real deployments. R EFERENCES [1] J. Li, M. Qi, L. Zhang, Y. Xiong, Z. Sun, and T. Zhu, “MA-HGAT: Multi-agent heterogeneous graph attention network for smart contract logical vulnerability detection,” in Proceedings of the 16th International Symposium on Cyberspace Safety and Security (CSS), ser. Lecture Notes in Computer Science. Springer, 2026, to appear; conference version of this article. [2] O. Novo, “Blockchain meets IoT: An architecture for scalable access management in IoT,” IEEE Internet of Things Journal, vol. 5, no. 2, pp. 1184–1195, 2018. [3] A. Reyna, C. Martín, J. Chen, E. Soler, and M. Díaz, “On blockchain and its integration with IoT. Challenges and opportunities,” Future Generation Computer Systems, vol. 88, pp. 173–190, 2018. [4] M. J. C. S. Reis, “Blockchain-enhanced security for 5G edge computing in IoT,” Computation, vol. 13, no. 4, p. 98, 2025. [5] Hacken, “The Hacken 2025 half-year Web3 security report,” https:// hacken.io/insights/h1-2025-security-report/, 2025. [6] ZeroTime Technology, “Analysis report on Web3 on-chain security in the first half of 2025,” https://www.panewslab.com/zh/articles/q13ibu8al081, 2025, in Chinese; published on PANews. [7] M. Antonakakis, T. April, M. Bailey, M. Bernhard, E. Bursztein, J. Cochran, Z. Durumeric, J. A. Halderman, L. Invernizzi, M. Kallitsis, D. Kumar, C. Lever, Z. Ma, J. Mason, D. Menscher, C. Seaman, N. Sullivan, K. Thomas, and Y. Zhou, “Understanding the Mirai botnet,” in 26th USENIX Security Symposium (USENIX Security 17), 2017, pp. 1093–1110. [8] European Union Agency for Cybersecurity (ENISA), “ENISA threat landscape 2024,” https://www.enisa.europa.eu/publications/ enisa-threat-landscape-2024, 2024. [9] J. Feist, G. Grieco, and A. Groce, “Slither: A static analysis framework for smart contracts,” in 2019 IEEE/ACM 2nd International Workshop on Emerging Trends in Software Engineering for Blockchain (WETSEB). IEEE, 2019, pp. 8–15. [10] ConsenSys Diligence, “Mythril: Security analysis tool for EVM bytecode,” https://github.com/ConsenSysDiligence/mythril, 2023. [11] L. Luu, D. H. Chu, H. Olickel, P. Saxena, and A. Hobor, “Making smart contracts smarter,” in Proceedings of the ACM Conference on Computer and Communications Security (CCS). ACM, 2016. [12] B. Jiang, Y. Liu, and W. K. Chan, “Contractfuzzer: Fuzzing smart contracts for vulnerability detection,” in 33rd ACM/IEEE International Conference on Automated Software Engineering (ASE). ACM, 2018. [13] Runtime Verification Inc., “K: A rewrite-based executable semantic framework,” Runtime Verification Inc. official documentation, 2020, accessed: 2025-01-02. [Online]. Available: https://kframework.org/
16
[14] Y. Sun, D. Wu, Y. Xue, H. Liu, H. Wang, Z. Xu, X. Xie, and Y. Liu, “Gptscan: Detecting logic vulnerabilities in smart contracts by combining gpt with program analysis,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE). ACM, 2024. [15] L. Yu, Z. Huang, H. Yuan, S. Cheng, L. Yang, F. Zhang, C. Shen, J. Ma, J. Zhang, J. Lu, and C. Zuo, “Smart-LLaMA-DPO: Reinforced large language model for explainable smart contract vulnerability detection,” Proceedings of the ACM on Software Engineering, vol. 2, no. ISSTA, 2025. [16] Y. Shoshitaishvili, R. Wang, C. Hauser, C. Kruegel, and G. Vigna, “Firmalice: Automatic detection of authentication bypass vulnerabilities in binary firmware,” in Network and Distributed System Security Symposium (NDSS), 2015. [17] N. Redini, A. Machiry, R. Wang, C. Spensky, A. Continella, Y. Shoshitaishvili, C. Kruegel, and G. Vigna, “Karonte: Detecting insecure multibinary interactions in embedded firmware,” in 2020 IEEE Symposium on Security and Privacy (S&P), 2020, pp. 1544–1561. [18] L. Chen, Y. Wang, Q. Cai, Y. Zhan, H. Hu, J. Linghu, Q. Hou, C. Zhang, H. Duan, and Z. Xue, “Sharing more and checking less: Leveraging common input keywords to detect bugs in embedded systems,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 303–319. [19] Y. Zheng, A. Davanian, H. Yin, C. Song, H. Zhu, and L. Sun, “FIRMAFL: High-throughput greybox fuzzing of IoT firmware via augmented process emulation,” in 28th USENIX Security Symposium (USENIX Security 19), 2019, pp. 1099–1114. [20] B. Feng, A. Mera, and L. Lu, “P2IM: Scalable and hardwareindependent firmware testing via automatic peripheral interface modeling,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 1237–1254. [21] M. Kim, D. Kim, E. Kim, S. Kim, Y. Jang, and Y. Kim, “FirmAE: Towards large-scale emulation of IoT firmware for dynamic analysis,” in Annual Computer Security Applications Conference (ACSAC), 2020, pp. 733–745. [22] Q. Feng, R. Zhou, C. Xu, Y. Cheng, B. Testa, and H. Yin, “Scalable graph-based bug search for firmware images,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2016, pp. 480–491. [23] X. Xu, C. Liu, Q. Feng, H. Yin, L. Song, and D. Song, “Neural networkbased graph embedding for cross-platform binary code similarity detection,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2017, pp. 363–376. [24] H. Xiao, Y. Zhang, M. Shen, C. Lin, C. Zhang, S. Liu, and M. Yang, “Accurate and efficient recurring vulnerability detection for IoT firmware,” in Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS). ACM, 2024. [25] Y. Cheng, F. Xu, L. Xu, Y. Ge, J. Yang, W. Fan, W. Huang, and W. Liu, “FirmVulLinker: Leveraging multi-dimensional firmware profiling for identifying homologous vulnerabilities in Internet of Things devices,” Electronics, vol. 14, no. 17, p. 3438, 2025. [26] X. Sun and F. Komaki, “Bhgnn-rt: Capturing bidirectionality and network heterogeneity in graphs,” PloS One, vol. 20, no. 7, p. e0326756, 2025. [Online]. Available: https://doi.org/10.1371/journal.pone.0326756 [27] H. Liu, Y. Tong, S. Ji, and P. Zhang, “Smart contract reentrancy vulnerability localization using explainable graph neural networks,” in 49th IEEE Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 2025, pp. 1126–1135. [28] X. Duan, J. Wu, S. Ji, Z. Rui, T. Luo, M. Yang, and Y. Wu, “VulSniper: Focus your attention to shoot fine-grained vulnerabilities,” in Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), 2019, pp. 4665–4671. [29] K. Christidis and M. Devetsikiotis, “Blockchains and smart contracts for the Internet of Things,” IEEE Access, vol. 4, pp. 2292–2303, 2016. [30] A. Dorri, S. S. Kanhere, R. Jurdak, and P. Gauravaram, “Blockchain for IoT security and privacy: The case study of a smart home,” in IEEE International Conference on Pervasive Computing and Communications Workshops (PerCom Workshops), 2017, pp. 618–623. [31] M. Qi, Q. Wang, Z. Wang, M. Schneider, T. Zhu, S. Chen, W. Knottenbelt, and T. Hardjono, “Sok: Bitcoin layer two (l2),” ACM Computing Surveys, vol. 58, no. 3, pp. 1–37, 2025. [32] W. Deng, H. Wei, T. Huang, C. Cao, Y. Peng, and X. Hu, “Smart contract vulnerability detection based on deep learning and multimodal decision fusion,” Sensors, vol. 23, no. 16, p. 7246, 2023. [Online]. Available: https://doi.org/10.3390/s23167246 [33] P. Tsankov, A. Dan, D. Drachsler-Cohen, A. Gervais, F. Buenzli, and M. Vechev, “Securify: Practical security analysis of smart contracts,”
in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS). ACM, 2018, pp. 67–82. [34] S. Kalra, S. Goel, M. Dhawan, and S. Sharma, “Zeus: Analyzing safety of smart contracts,” in Network and Distributed System Security Symposium (NDSS). Internet Society, 2018. [35] M. Rodler, W. Li, G. O. Karame, and L. Davi, “Sereum: Protecting existing smart contracts against re-entrancy attacks,” in Network and Distributed System Security Symposium (NDSS). Internet Society, 2019. [36] Y. Zhang, S. Ma, J. Li et al., “Smartshield: Automatic smart contract protection made easy,” in 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 2020, pp. 23–34. [37] K. Li, Y. Xue, S. Chen, H. Liu, K. Sun, M. Hu, H. Wang, Y. Liu, and Y. Chen, “Static application security testing (SAST) tools for smart contracts: How far are we?” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, p. 65, 2024. [Online]. Available: https://doi.org/10.1145/3660772 [38] G. Grieco, W. Song, A. Cygan, J. Feist, and A. Groce, “Echidna: Effective, usable, and fast fuzzing for smart contracts,” in Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA). ACM, 2020, pp. 557–560. [39] Trail of Bits, “Welcome to manticore’s documentation!” 2019, manticore 0.3.7 documentation. [Online]. Available: https://manticore.readthedocs. io/ [40] J. Choi, D. Kim, S. Kim, G. Grieco, A. Groce, and S. K. Cha, “Smartian: Enhancing smart contract fuzzing with static and dynamic data-flow analyses,” in 36th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2021. [41] Z. Kong, C. Zhang, M. Xie et al., “Smart contract fuzzing towards profitable vulnerabilities,” Proceedings of the ACM on Software Engineering, vol. 2, no. FSE, 2025. [42] F* Development Team, “F*: A proof-oriented programming language,” F* official documentation and website, 2021, accessed: 2025-01-02. [Online]. Available: https://www.fstar-lang.org/ [43] Solidity Team, “Solidity smtchecker documentation,” 2023. [Online]. Available: https://docs.soliditylang.org/en/latest/smtchecker.html [44] B. Boi, C. Esposito, and S. Lee, “Vulnhunt-gpt: A smart contract vulnerabilities detector based on openai chatgpt,” in Proceedings of the 39th ACM/SIGAPP Symposium on Applied Computing (SAC), 2024. [45] N. Li, M. Qi, Z. Xu, X. Zhu, W. Zhou, S. Wen, and Y. Xiang, “Blockchain cross-chain bridge security: Challenges, solutions, and future outlook,” Distributed Ledger Technologies: Research and Practice, vol. 4, no. 1, pp. 1–34, 2025. [46] T. Jiao, Z. Xu, M. Qi, S. Wen, Y. Xiang, and G. Nan, “A survey of ethereum smart contract security: Attacks and detection,” Distributed Ledger Technologies: Research and Practice, vol. 3, no. 3, pp. 1–28, 2024. [47] A. Costin, J. Zaddach, A. Francillon, and D. Balzarotti, “A large-scale analysis of the security of embedded firmwares,” in 23rd USENIX Security Symposium (USENIX Security 14), 2014, pp. 95–110. [48] A. Qasem, P. Shirani, M. Debbabi, L. Wang, B. Lebel, and B. L. Agba, “Automatic vulnerability detection in embedded devices and firmware: Survey and layered taxonomies,” ACM Computing Surveys, vol. 54, no. 2, pp. 25:1–25:42, 2021. [49] H. Wen, Z. Lin, and Y. Zhang, “FirmXRay: Detecting Bluetooth link layer vulnerabilities from bare-metal firmware,” in Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2020, pp. 167–180. [50] Keen Security Lab, Tencent, “BinAbsInspector: Vulnerability scanner for binaries,” https://github.com/KeenSecurityLab/BinAbsInspector, 2022. [51] J. Zaddach, L. Bruno, A. Francillon, and D. Balzarotti, “AVATAR: A framework to support dynamic security analysis of embedded systems’ firmwares,” in Network and Distributed System Security Symposium (NDSS), 2014. [52] M. Muench, D. Nisi, A. Francillon, and D. Balzarotti, “Avatar2 : A multitarget orchestration platform,” in Workshop on Binary Analysis Research (BAR), co-located with NDSS, 2018. [53] M. Muench, J. Stijohann, F. Kargl, A. Francillon, and D. Balzarotti, “What you corrupt is not what you crash: Challenges in fuzzing embedded devices,” in Network and Distributed System Security Symposium (NDSS), 2018. [54] Y. David, N. Partush, and E. Yahav, “FirmUp: Precise static detection of common vulnerabilities in firmware,” in Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2018, pp. 392–404.
17
[55] J. Ji, C. Zhang, S. Gan, L. Jian, H. Liu, T. Liu, L. Zheng, and Z. Jia, “FirmAgent: Leveraging fuzzing to assist LLM agents with IoT firmware vulnerability discovery,” in Network and Distributed System Security Symposium (NDSS), 2026. [56] V. Nageshwaran and S. Ezekiel, “Agentic AI and large language models for autonomous IoT cybersecurity: A systematic survey, taxonomy, and research roadmap,” Electronics, vol. 15, no. 12, p. 2740, 2026. [57] M. Allamanis, M. Brockschmidt, and M. Khademi, “Learning to represent programs with graphs,” International Conference on Learning Representations, 2018, preprint on OpenReview. [Online]. Available: https://openreview.net/forum?id=BJOFETxR[58] B. Raju and G. Devi K, “An elegant intellectual engine towards automation of blockchain smart contract vulnerability detection,” Scientific Reports, vol. 15, no. 1, p. 26104, 2025. [59] W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 1024–1034. [60] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph attention networks,” in International Conference on Learning Representations (ICLR), 2018. [61] X. Wang, H. Ji, C. Shi, B. Wang, Y. Ye, P. Cui, and P. S. Yu, “Heterogeneous graph attention network,” in The World Wide Web Conference (WWW), 2019, pp. 2022–2032. [62] M. Schlichtkrull, T. N. Kipf, P. Bloem, R. van den Berg, I. Titov, and M. Welling, “Modeling relational data with graph convolutional networks,” in The Semantic Web (ESWC), 2018, pp. 593–607. [63] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems, 2017, arXiv preprint arXiv:1706.03762. [Online]. Available: https://arxiv.org/abs/1706.03762 [64] J. Alammar, “The illustrated transformer,” Personal blog, June 2018, visual explanation of Transformer architecture. [Online]. Available: https://jalammar.github.io/illustrated-transformer/ [65] P. K. Adjei, Z. Qin, I. A. Obiri, A. Badjie, C. N. A. Cobblah, A. Alqahtani, Y. H. Gu, and M. A. Al-antari, “A graph attention networkbased multi-agent reinforcement learning framework for robust detection of smart contract vulnerabilities,” Scientific Reports, vol. 15, no. 1, p. 29810, 2025. [66] Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proceedings of the IEEE, vol. 107, no. 8, pp. 1738–1762, 2019. [67] Y. Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neurosurgeon: Collaborative intelligence between the cloud and mobile edge,” in Proceedings of the 22nd International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2017, pp. 615–629. [68] Y. Matsubara, M. Levorato, and F. Restuccia, “Split computing and early exiting for deep learning applications: Survey and research challenges,” ACM Computing Surveys, vol. 55, no. 5, pp. 90:1–90:30, 2022. [69] M. S. Ali, M. Vecchio, M. Pincheira, K. Dolui, F. Antonelli, and M. H. Rehmani, “Applications of blockchains in the Internet of Things: A comprehensive survey,” IEEE Communications Surveys & Tutorials, vol. 21, no. 2, pp. 1676–1717, 2019. [70] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980–2988. [71] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations (ICLR), 2019. [72] M. Wang, D. Zheng, Z. Ye, Q. Gan, M. Li, X. Song, J. Zhou, C. Ma, L. Yu, Y. Gai, T. Xiao, T. He, G. Karypis, J. Li, and Z. Zhang, “Deep graph library: A graph-centric, highly-performant package for graph neural networks,” arXiv preprint arXiv:1909.01315, 2019. [73] T. Górski, “Smart contract design pattern for processing logically coherent transaction types,” Applied Sciences, vol. 14, no. 6, p. 2224, 2024. [74] SunWeb3Sec, “DeFiHackLabs: Reproduce DeFi hacked incidents using foundry,” https://github.com/SunWeb3Sec/DeFiHackLabs, 2022. [75] Z. Zhang, B. Zhang, W. Xu, and Z. Lin, “Demystifying exploitable bugs in smart contracts,” in Proceedings of the 45th IEEE/ACM International Conference on Software Engineering (ICSE). IEEE, 2023, pp. 615–627, dataset at https://github.com/ZhangZhuoSJTU/Web3Bugs. [76] Crytic, “Crytic compile,” 2023. [Online]. Available: https://github.com/ crytic/crytic-compile [77] A. Pinheiro, E. D. Canedo, R. d. O. Albuquerque, and R. T. de Sousa Júnior, “Validation of architecture effectiveness for the continuous monitoring of file integrity stored in the cloud using
blockchain and smart contracts,” Sensors (Basel, Switzerland), vol. 21, no. 13, 2021. [Online]. Available: https://doi.org/10.3390/s21134440 [78] Z. He, T. Zhang, and R. B. Lee, “Model inversion attacks against collaborative inference,” in Proceedings of the 35th Annual Computer Security Applications Conference (ACSAC). ACM, 2019, pp. 148–162. [79] C. Song and A. Raghunathan, “Information leakage in embedding models,” in Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (CCS). ACM, 2020, pp. 377– 390. [80] a101e-lab, “IoTVulBench: An open-source benchmark dataset for IoT security research,” https://github.com/a101e-lab/IoTVulBench, 2024. [81] MetaTrust Labs, “Scanning results of the GPTScan engine on 72 Web3Bugs projects,” https://github.com/MetaTrustLabs/ GPTScan-Web3Bugs, 2023. [82] DeepSeek-AI, “DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning,” 2025. [Online]. Available: https: //arxiv.org/abs/2501.12948 [83] OpenAI, “Chatgpt,” 2023. [Online]. Available: https://openai.com/ chatgpt