Conceptio › Archive › arXiv CS
arXiv CSopen access

SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent

Zeqin Liao et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

1

SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent

arXiv:2609.34983v1 [cs.SE] 28 Sep 2026

Zeqin Liao, Yuhong Nan, Member, IEEE, Henglong Liang, Zixu Gao, Lianyu Hu∗ , Yuqiang Sun, Zhijie Zhong, Xiaoyu Ma, Zibin Zheng, Fellow, IEEE, and Yang Liu, Senior Member, IEEE

and fiat-backed stablecoins [4]. However, security incidents targeting the on-chain contracts of these applications have become increasingly frequent [2], [3]. Our investigation reveals more than 73 incidents in the past five years, the majority of which stem from contract vulnerabilities [5]. Given its importance, the security of OFC contracts has not been systematically reviewed. More specifically, existing public reports and studies typically classify these incidents into seemingly-unrelated categories. For example, the vulnerabilities in PolyNetwork [6], Synapse [7], and THORChain [8] are attributed to an access-control vulnerability, accounting errors, and event forgery, respectively. This fragmentation also appears at the analysis-tool level, where existing tools often address separate vulnerability categories within OFC applications rather than providing a unified view, and thus often miss vulnerabilities outside their patterns. For example, SmartAxe [2], BridgeGuard [9], VCScope [10], and ChainSniper [11] focus on access-control and semantic vulnerabilities, interaction vulnerabilities, signature-verification vulnerabilities, and reentrancy or overflow vulnerabilities, respectively. Despite the diverse categories and consequences, our research identifies OFC inconsistency (OFCI) as the common underlying cause of OFC vulnerabilities. Specifically, OFC applications are expected to maintain value equivalence between the onchain and off-chain representations of the same asset [12]. We characterize this requirement as an equivalence invariant. While Index Terms—Vulnerability detection, Agent, On-chain-off- prior studies mainly target symptoms of a broken invariant, such as unauthorized state modification or accounting errors, chain Communication we focus on where and how equivalence invariant is structurally broken. Accordingly, we define OFC inconsistency as an I. I NTRODUCTION inconsistency that arises from incomplete asset-exchange logic Smart contracts are programs deployed on blockchains, between on-chain contracts and off-chain entities and ultimately supporting a wide range of decentralized applications (DApps) breaks the equivalence invariant. While OFCIs have been widely exploited in practice, research such as decentralized finance (DeFi) [1]. In the DeFi ecosystem, on OFC contract security remains limited. First, vulnerabilitythe demand for on-chain and off-chain communication (OFC) detection tools [2], [11], [10], [9] are dedicated to a specific is rapidly growing, with typical applications such as crossOFC application domain (i.e., cross-chain bridges), and thus chain bridges [2], real-world asset (RWA) tokenization [3], cannot cover general OFC exchange logic or generalize to ∗ Corresponding author: Lianyu Hu. other OFC applications such as RWA tokenization and fiatZeqin Liao, Lianyu Hu, Yuqiang Sun, Xiaoyu Ma, and Yang Liu are with backed stablecoins. Second, other related efforts [13], [14], [15] Nanyang Technological University, Singapore (e-mail: [email protected]. sg; [email protected]; [email protected]; [email protected]; address adjacent problems rather than vulnerability detection. [email protected]). DArcher [13] tests data synchronization, while XGuard [14] Yuhong Nan, Henglong Liang, Zixu Gao, Zhijie Zhong, and Zibin Zheng and HighGuard [15] target behavior-anomaly analysis. are with Sun Yat-sen University, China (e-mail: [email protected]; Furthermore, existing approaches have limited generality [email protected]; [email protected]; gaozx9@mail2. sysu.edu.cn; [email protected]). in detecting OFCIs due to two key limitations. First, most

Abstract—Smart contracts underpin decentralized finance, where growing demand for on-chain/off-chain communication (OFC) has driven diverse applications such as cross-chain bridges, real-world asset tokenization, and fiat-backed stablecoins. The OFC-related security incidents in these applications are increasingly frequent, but prior studies address separate vulnerability categories within OFC applications rather than providing a unified view, causing vulnerabilities outside known patterns to be missed. In this paper, we identify OFC inconsistency (OFCI) as a root cause of OFC vulnerabilities, which arises from business-logic flaw and ultimately breaks the equivalence between the onchain and off-chain asset representations to induce inconsistency. Automatically detecting OFCIs faces two challenges including (1) locating heterogeneous business logic, and (2) transferring existing vulnerability knowledge to identify unseen OFCI instances. To this end, we propose SmartMemory, the first framework to leverage a memory-based agent for OFCI detection. To address heterogeneity, SmartMemory maps diverse implementations of OFC contracts into a canonical business-semantic representation to locate the business logic for OFCI inspection. For knowledge reuse, SmartMemory integrates a memory-based agent to distill vulnerability knowledge from features into patterns and detection rules, enabling knowledge transfer across cases to identify unseen OFCIs. Lastly, SmartMemory performs taint analysis to verify the reachability, type, and impact of each candidate OFCI. We construct the first real-world OFCI dataset comprising 48 DApps with 81 OFCIs for evaluation, on which SmartMemory achieves 80.68% precision and 87.65% recall. In addition, through an analysis of 325 real-world OFC applications, SmartMemory detects 36 previously unknown OFCIs, all of which have been confirmed and fixed by corresponding parties.

2

approaches [2], [10], [11], [9] rely on application-agnostic analysis or predefined rules, whereas OFCIs are tightly coupled with application-specific business semantics. The same vulnerability may appear in different implementations of business logic, making it difficult for such analysis to capture. Second, these approaches [2], [11], [9] are predominantly pattern-based and rely on human heuristics. However, OFC applications involve n-to-n combinations of chains, assets, and implementations, which causes their business logic and vulnerability patterns to grow combinatorially. As a result, any finite pattern set is insufficient in practice, exposing the limitation of pattern-based methods. While the method [10] based on a standalone LLM can reason over business logic for vulnerability discovery, it lacks a systematic mechanism to abstract, store, and transfer vulnerability knowledge across cases. This limits its ability to detect previously unseen vulnerabilities. Our insight is that a memory mechanism for LLM agent can address this limitation by enabling the agent to distill knowledge from detected cases and transfer it to new detection cases. In this way, the agent can accumulate and reuse knowledge across cases, helping it handle unseen vulnerabilities. In this paper, we propose SmartMemory, an LLM-agentbased framework for detecting OFCIs in smart contracts. SmartMemory operates through a three-stage pipeline of locating, identifying, and verifying OFCIs. In the first stage, SmartMemory uses a general conceptual model to map heterogeneous implementations into a canonical business-semantic representation, thereby locating the OFC business logic for inspection (Section IV-B). In the second stage, SmartMemory uses a hierarchical-memory agent to check whether the target logic contains OFCI. The memory progressively abstracts vulnerability knowledge from features into patterns and rules, enabling knowledge accumulation and transfer across cases (Section IV-C). In the third stage, SmartMemory constructs a cross on-chain/off-chain data-flow graph (CDFG) and performs taint analysis on it to verify the reachability, type, and impact of the detected OFCI candidate (Section IV-D). Evaluation. To evaluate the effectiveness of SmartMemory, we first construct a manually labeled dataset from public security reports of real-world OFC attacks, covering 48 OFC applications with 81 OFCIs. Experimental results show that SmartMemory achieves 80.68% precision and 87.65% recall on the dataset. Compared with state-of-the-art (SOTA) tools within their respective detection scopes, SmartMemory significantly outperforms SmartAxe and VCScope. Furthermore, we construct a comprehensive OFC contract dataset consisting of 325 real-world OFC applications and deploy SmartMemory in collaboration with a professional auditing firm for large-scale security auditing. SmartMemory detects 36 previously unseen zero-day OFCIs, all of which have been confirmed and fixed. Contributions. This paper makes the following contributions: • We highlight OFCI as a root cause of OFC vulnerabilities,

which arises from business-logic flaw and ultimately breaks the equivalence invariants to induce inconsistency. • We propose SmartMemory, the first framework that leverages a memory-based agent to detect OFCIs for OFC contracts, to the best of our knowledge.

Fig. 1. Core workflows of on-chain and off-chain communications

• We perform extensive experiments to demonstrate the

effectiveness of SmartMemory. Through an analysis of 325 real-world applications, SmartMemory identifies 36 new OFCIs, all of which have been confirmed and fixed. • We build the first manually-labeled dataset of OFC vulnerabilities, as well as the largest OFC contract dataset. II. BACKGROUND A. On-chain and Off-chain Communication in DeFi As a kind of smart-contract-powered peer-to-peer application, Decentralized Finance has attracted a surge in popularity [16]. On-chain and off-chain communication (OFC). In the DeFi ecosystem, on-chain and off-chain communication (OFC) refers to interactions between on-chain smart contracts and off-chain entities or external settlements [3], [10]. Fig. 1(a) shows an OFC architecture with on-chain and off-chain components. On-chain contracts implement token-exchange logic, while offchain components relay information between contracts and entities. This allows users to deposit assets on one blockchain and redeem them off chain or on another blockchain, and vice versa. The OFC workflows are as follows: • Asset deposit. The contract executes a deposit to transfer

assets into the contract. 1 the contract • Token burning. Upon redemption (see ⃝), burns the corresponding on-chain tokens, and emits an event 2 as proof of asset outflow (see ⃝). • Off-chain verification. Off-chain witnesses or relayers verify 3 the on-chain proof and determine the asset value (see ⃝). • Off-chain audit. Off-chain data are packaged as proof (see 4 and submitted to the contract. ⃝) • Asset registration. Upon receiving the proof, the contract verifies the submitted asset information for admission. • Issuance and minting. Upon a minting request, the contract 5 under the obtains the valuation, and mints tokens (see ⃝) issuance rules for the user. Equivalence invariant. OFC applications are required to maintain value equivalence between the on-chain and off-chain representations of the same asset [17]. We characterize this requirement as equivalence invariant, i.e., the value represented on chain must remain equivalent to its off-chain counterpart across each relevant state transition of OFC applications. Since the two representations interact only through a small set of state transition points, the equivalence is constrained at these seams [18]. As shown in Fig. 1, OFC applications possess three equivalence invariants:

3

Fig. 2. Examples of OFCIs caused by different types of inconsistencies, namely, (a) asset-exchange inconsistency and (b) proof-semantic inconsistency.

B1 (Outflow equivalence): the asset transferred by the user must be value-equivalent to the on-chain proof emitted by 1 ⃝). 2 the contract (⃝≡ • B2 (Verification equivalence): the on-chain proof emitted by the contract must be value-equivalent to the asset value 2 ⃝). 3 parsed by the off-chain verifier (⃝≡ • B3 (Issuance equivalence): the off-chain proof information accepted by the contract must be value-equivalent to the 4 ⃝). 5 on-chain tokens minted from it (⃝≡ •

B. LLM Agent and Memory LLM-based agent is a rapidly developing research direction, which can continually learn and adapt to evolving task requirements [19], [20]. Memory is a key mechanism supporting this capability [21], [20], [22]. Analogous to human memory, it enables agents to learn from experience, reflect, and accumulate knowledge over time. It also distinguishes an agent from a standalone LLM [20]. Memory involves three core processes: (1) encoding, which transforms information into a storable form; (2) updating, which integrates encoded information into memory; and (3) retrieval, which extracts taskrelevant information to prompt the model when needed [23]. With memory, LLM agents can derive high-level insights from trajectories and experience. Produced through model introspection, these insights can be retrieved to improve current performance and support continual adaptation to dynamic task demands [20], [22], [23]. III. M OTIVATION A. Definition and Problem Statement We define an OFC inconsistency as a flaw in the contract’s business logic that violates the equivalence invariants (i.e., B1 , B2 , and B3 defined in Section II-A), thereby causing the value of an asset’s on-chain representation to diverge from its offchain counterpart. We classify OFCIs into two structural types

based on which equivalence invariant is violated and where the resulting inconsistency is observed. We use two motivating examples in Fig. 2 to show two types of OFCIs, which are drawn from real-world OFC applications exploited by attackers. We reorganize the original contract code for better clarity. Type-1: Asset-exchange inconsistency (B1 and B3 ). This type of inconsistency captures violations of asset-quantity equivalence in contract business logic, where the contract computes an output token quantity that is inconsistent with the corresponding input-side asset value. This type corresponds to violations of B1 and B3 . Fig. 2(a) illustrates such a case caused by flawed business logic. In contract Pledge, function mint supports minting tokens according to either on-chain assets (line 2) or off-chain collateral (lines 3–4). However, both paths incorrectly use chainAmount to determine the mint amount, although this parameter actually denotes the amount derived from on-chain assets. Consequently, an adversary can submit a minting request with minimal off-chain collateral (see Sa1 ), then increase the on-chain asset amount (see Sa2 ), and obtain an excessive number of minted on-chain tokens (see Sa3 ). This creates an inconsistency between the off-chain proof (see Sa4 ) and the minted on-chain tokens (see Sa3 ). The adversary then withdraws the originally deposited on-chain assets (see Sa6 ), gaining profit with only minimal off-chain collateral. Type-2: Proof-semantic inconsistency (B2 ). This type of inconsistency is caused by the flawed contract business logic for generating on-chain asset proofs, where the structure of an on-chain proof allows an off-chain verifier to parse a meaning inconsistent with the intended asset-flow semantics. This type corresponds to violations of B2 . Fig. 2(b) illustrates such a case. In contract Token, function deposit locks on-chain assets (lines 3–4) and emits a Deposit event, Edeposit , as the proof of on-chain assets (line 5). However, another contract Vault in the same DApp provides function returnAssets to batchreturn assets from the vault to users. After returning the assets (lines 9–10), this function also emits a Transfer event, Ereturn

4

Merkle root as message status (Line 14), and emits Process and Receive (Lines 16–17). Despite the implementation differences, OFC contracts are organized around core workflows (Section II-A), which endow their business logic with a set of shared underlying semantics. We invite domain experts to distill these semantics into primitive functions, intra-operations, and operated variables, forming a general business-logic conceptual model for heterogeneous OFC contracts (Fig. 5). We further define four key variables around the equivalence breakpoints defined in Section III-A, which capture OFC-specific on-chain/off-chain asset-exchange semantics that existing contract analysis techniques do not support. With this model, SmartMemory prompts the LLM to assign corresponding business-logic-related statements with the Fig. 3. An example of diversity of off-chain asset registration. triple labels of primitive function, intra-operation, and operated variable. Then, SmartMemory produces semantically-annotated (line 11), with the same structure as Edeposit . Moreover, source code to locate the OFC business logic for inspection each parameter of Ereturn can be arbitrarily crafted through (Section IV-B). externally controlled inputs (line 11). This design introduces an exploitable OFCI. An adversary can invoke returnAssets C2: Constructing memory mechanism for inconsistency to forge an on-chain asset proof (see Sp1 ), Ereturn . Because the detection. This challenge lies in the following aspects: off-chain relayer cannot distinguish Ereturn from Edeposit , it (1) knowledge representation, where memory must encode may treat both as valid proofs (see Sp2 ), causing inconsistency information that can support OFCI detection rather than store between the on-chain proof and the verification result. As a case details indiscriminately, (2) knowledge update and learning, result, the adversary may obtain a valid-looking asset proof where memory must distill reusable and generalizable OFCI (see Sp3 ) without actually transferring assets to the DApp, and knowledge from successive detection cases rather than merely accumulate isolated observations, and (3) knowledge retrieval, then withdraw assets from the OFC application for profit. where the relevant knowledge must be retrieved from a growing Scope. We target vulnerabilities where controllable inputs memory accurately for a given new OFCI detection. cause the violation of B1 , B2 , or B3 and induce on-chain/offTo address these challenges, we design an LLM agent chain value inconsistency, which consists of asset-exchange with a three-layer memory that organizes knowledge into inconsistency and proof-semantic inconsistency. These two feature, pattern, and rule graphs. For representation, the types cover the main OFC vulnerabilities, while fewer cases feature graph encodes OFCI-specific information for individual caused by non-contract issues (e.g., runtime/off-chain infrasvulnerabilities, while the pattern and rule graphs encode tructure) are outside the scope of our static analysis. higher-level knowledge abstracted across them. For update, operators with explicit triggering conditions progressively abstract features into patterns and rules, merge redundant B. Challenges and Solutions knowledge, remove outdated or falsified knowledge, and OFCI originates from logic incompleteness (e.g., conflicting decompose overfitted knowledge, enabling the agent to abstract, events or erroneous computations) and ultimately breaks store, and transfer vulnerability knowledge across cases. For equivalence invariant to induce inconsistency. Hence, a straight- retrieval, the agent augments similarity-based retrieval with forward idea to identify OFCIs has two key steps: (1) identify graph-structural expansion, considering both semantic and functions that implement incorrect or incomplete OFC-specific structural relevance to reduce noisy matches (Section IV-C). business logic and their relevant variables that an adversary can thereby manipulate; and (2) check whether critical variables IV. D ESIGN OF S MART M EMORY involved in the equivalence invariant are data-flow dependent on manipulable variables in (1), so as to determine whether A. Overview an adversary can indirectly alter them to induce profitable SmartMemory analyzes OFC smart contract source code and inconsistencies. However, implementing these steps is non- reports detected OFCIs along with their vulnerability traces. trivial. We present the challenges and corresponding solutions. As shown in Fig. 4, the workflow forms a three-stage pipeline. C1: Diversity in business-logic implementations. The first The first locates the OFC business logic for inspection, the challenge is the diversity of OFC business-logic implementa- second identifies business-logic incompleteness, and the third tions, which stems from heterogeneous implementations and verifies the reachability, type, and impact of detected OFCIs. non-standardized designs [24]. Fig. 3 shows two different Contract-logic semantic extraction. Guided by the OFC implementations of off-chain asset registration in OFC contracts. conceptual model in Fig. 5, SmartMemory identifies businessWormhole [25] verifies off-chain proofs through a guardian logic operations in the input contracts and annotates relevant multisignature quorum (Line 3), records message status with statements with triple labels of primitive functions, intraa boolean flag (Lines 5–6), and emits Redeem (Line 8). In operations, and operated variables. With these annotations, contrast, Nomad [26] uses Merkle proofs (Line 12), stores the SmartMemory generates semantically annotated source code

Wormhole { register(bytes memory encodedVm) { 3 VerifyVM(encodedVm); ... //P1: multi-sig verification 4 // P2: record message status (boolean flag) 5 setTransferCompleted(vm.hash); 6 completedTransfers[hash] = true; ... 7 // P3: emit event 8 emit redeem(ChainId, Token, Recipient, Amount);}} 9contract Nomad { 10function prove(bytes32 leaf, bytes32 proof, uint256 index) { 11 // P1: verify message via Merkle proof + time window 12 bytes32 root = MerkleLib.branch(leaf, proof, index);... 13 // P2: record message status (Merkle root) 14 messages[_hash] = root; ... 15 // P3: emit event 16 emit Process(Hash, success, Data); 17 emit Receive(nonce, token, recipient, amount);}} 1contract 2function

5

Fig. 4. Overview of SmartMemory.

to locate the OFC business logic for inspection and delimits the detection context with the constructed call graph. Agent-based inconsistency identification. SmartMemory employs an LLM agent with three-layer memory to examine whether the annotated target functions contain OFCIs. This memory learns from an external vulnerability knowledge base and the agent’s historical detection experience, abstracting lowlevel features into recurring patterns and generalized detection rules. For each target function, the agent retrieves relevant memory knowledge and combines it with the semanticallyannotated source code to prompt the LLM in determining whether the function is vulnerable. Vulnerability discovery via taint analysis. SmartMemory performs taint analysis on a cross on-chain/off-chain data-flow graph (CDFG). The analysis verifies whether an identified vulnerable function is externally reachable, whether taint can propagate to variables of equivalence invariant associated with possible inconsistencies, and determines the OFCI type. B. Contract Logical Semantic Extraction Conceptual model construction for OFC contracts. To build a generic business-logic model for heterogeneous OFC smart contracts, we first collect all the OFC applications from DeFiLlama [27]. We then invite domain experts to review the contracts in these DApps and identify relevant functions. The experts first categorize the functions by OFC workflow, then abstract the primitive intra-operations in each category, and summarize the semantics of the key variables in each operation. Fig. 5 presents the conceptual model of OFC contracts. The model is hierarchical: contracts contain functions, functions comprise intra-operations, and intra-operations operate over variables. The conceptual model consists of a set of primitive functions O = {Deposit, Burn, Register , Mint, Price, Transfer , Yield , Govern, Liquidate}. These nine primitive functions are systematically derived, rather than selected ad hoc, from the generic OFC workflows in Section II. The four OFC-specific primitives (i.e., Deposit, Burn, Register , Mint) abstract the core asset-exchange workflows, while the five general primitives (i.e., Price, Transfer , Yield , Govern, Liquidate) are included because they can become carriers that break the equivalence invariant. For instance, the transfer operation returnAssets induces an inconsistency in Fig. 2 (b). Further, each function is composed of a certain subset of primitive intra-operations S = {Check , Update, Emit, Verify, Assign} and each intra operation operates on a certain subset of variables V = {OnchainAsset, OffchainCollateral , Proof , Debt, Supply, Valuation, Permission, Config, Event}. Further, we define the variables with Vkey ∈ {OnchainAsset, OffchainCollateral , AssetProof , CollateralProof } as key variables, which correspond to the critical OFC variables at the

⟨OFCModel⟩ ::= C ; f ; τ ⟨Contract⟩ ::= C = {f1 , . . . , fn } ⟨Function⟩ ::= f = {s1 , . . . , sm } ⟨Operation⟩ ::= τ : S → ⟨o, κ, δ⟩ ⟨FunLabel O⟩ ::= Deposit | Burn | Register | Mint | Price | Transfer | Yield | Govern | Liquidate ⟨OpLabel K⟩ ::= Check | Update | Emit | Verify | Assign ⟨VarLabel V⟩ ::= on-chain asset | off-chain collateral | asset proof | collateral proof | balance | supply | debt | valuation | permission | config | event ⟨Composition⟩ ::= R(o) = {(κ, δ) | ∃s ∈ S, τ (s) = ⟨o, κ, δ⟩} ⟨Deposit⟩ ::= ⟨Check, on-chain asset⟩, ⟨Update, balance⟩, ⟨Update, supply⟩, ⟨Emit, proof⟩ ⟨Burn⟩ ::= ⟨Check, balance⟩, ⟨Update, balance⟩, ⟨Update, debt⟩, ⟨Update, supply⟩, ⟨Emit, proof⟩ ⟨Register ⟩ ::= ⟨Check, on-chain asset⟩, ⟨Verify, collateral proof⟩, ⟨Update, on-chain asset⟩, ⟨Emit, event⟩ ⟨Mint⟩ ::= ⟨Check, permission⟩, ⟨Check, on-chain asset⟩, ⟨Check, off-chain collateral⟩, ⟨Update, debt⟩, ⟨Assign, on-chain asset⟩, ⟨Emit, event⟩ ⟨Price⟩ ::= ⟨Check, permission⟩, ⟨Update, valuation⟩, ⟨Emit, event⟩ ⟨Transfer ⟩ ::= ⟨Check, permission⟩, ⟨Check, balance⟩, ⟨Update, balance⟩, ⟨Emit, event⟩ ⟨Yield⟩ ::= ⟨Update, balance⟩, ⟨Emit, event⟩ ⟨Govern⟩ ::= ⟨Update, valuation⟩, ⟨Update, config⟩ ⟨Liquidate⟩ ::= ⟨Check, debt⟩, ⟨Update, debt⟩, ⟨Emit, event⟩

Fig. 5. Conceptual model of OFC smart contracts

two ends of the equivalence breakpoints. Next, we introduce the composition of each primitive function in conceptual model. Deposit validates asset admission with ⟨Check, onchainasset⟩, updates the user balance and system supply with ⟨Update, balance⟩ and ⟨Update, supply⟩, and emits an on-chain inflow proof with ⟨Emit, proof⟩. Burn checks balance sufficiency with ⟨Check, balance⟩, deducts balance, debt, and total supply with ⟨Update, balance⟩, ⟨Update, debt⟩, and ⟨Update, supply⟩, and emits an assetoutflow proof with ⟨Emit, proof⟩. Register confirms that the asset is registered with ⟨Check, onchain-asset⟩, verifies the off-chain audit proof with ⟨Verify, collateralproof⟩, sets admission with ⟨Update, onchain-asset⟩, emits registration event with ⟨Emit, event⟩. Mint checks call compliance with ⟨Check, permission⟩ and validates the minting basis with ⟨Check, onchain-asset⟩ or ⟨Check, offchain-collateral⟩, allocates the minted amount with ⟨Assign, onchain-asset⟩, updates debt with ⟨Update, debt⟩, and emits a minting event with ⟨Emit, event⟩. Price restricts the caller to the oracle with ⟨Check, permission⟩, updates valuation with ⟨Update, valuation⟩, and emits a price-update event with ⟨Emit, event⟩. Transfer checks permission and balance with ⟨Check, permission⟩ and ⟨Check, balance⟩, updates both balances with ⟨Update, balance⟩, and emits transfer event ⟨Emit, event⟩. Yield credits computed yield to the user balance with ⟨Update, balance⟩ and emits a distribution event with ⟨Emit, event⟩. Govern updates risk parameters and system configuration, including contract upgrades, with ⟨Update, valuation⟩ and ⟨Update, config⟩, without a fixed order. Liquidate checks collateral insufficiency with ⟨Check, debt⟩, adjusts total supply and clears debt with ⟨Update, supply⟩ and ⟨Update, debt⟩, and emits liquidation event ⟨Emit, event⟩. Note that the conceptual model is used to provide prototypes to construct structured prompts that guides the LLM in identifying contract logic for inspection. Moreover, the OFC specificity

6

of the model lies in the key variables, i.e., {OnchainAsset, OffchainCollateral, AssetProof , CollateralProof }, and their corresponding equivalence breakpoints. These equivalences are meaningful only in on-chain/off-chain asset-exchange scenarios and are not addressed by general contract analysis. Semantic annotation. SmartMemory annotates the corresponding statement with a triple label ⟨o, κ, δ⟩ to generate semantically-annotated code, where o ∈ O, κ ∈ S, and δ ∈ V. SmartMemory uses an LLM for semantic identification and annotates source code with triple labels. Given a function’s source code and the conceptual-model prompt, the LLM identifies the ⟨o, κ, δ⟩ labels corresponding to source statement. The labels are attached to statements, and multiple labels on the same statement are treated as a union. Since annotation operates on statement semantics rather than parameter signatures, heterogeneous implementations of the same primitive operation can be mapped to the same label. Taking Fig. 3 as an example, Wormhole and Nomad differ in parameter names, parameter counts, and implementation mechanisms, but can be annotated with the same triple ⟨ Register, Update, onchain-asset⟩. This is the core mechanism by which SmartMemory represents business-logic variants, i.e., implementation diversity is absorbed by semantic roles. SmartMemory does not replace real functions with templates from the conceptual model. The source code remains intact, and the model only guides the LLM in identifying semantics. SmartMemory further constructs call graph, and then selects context functions along the call chain and those sharing the same triple label ⟨o, κ, δ⟩, and serializes them into semanticallyannotated code context. Here, the call graph provides the control-flow context (i.e., the call chain), while the triple label ⟨o, κ, δ⟩ delimits the scope of the same business-logic operation. C. Agent-based Inconsistency Identification The memory design of SmartMemory is inspired by manual auditing practice, where experienced auditors extract casespecific features, abstract them into reusable patterns and rules, and update or retrieve such knowledge as new evidence accumulates. Accordingly, SmartMemory structures its memory as three interconnected knowledge graphs including feature graph Gfeature , pattern graph Gpattern , and rule graph Grule . Fig. 7 illustrates the two-phase workflow of SmartMemory’s memory-based agent. Phase-1. Memory retrieval for vulnerability detection. Given the semantically annotated code context of a contract function, SmartMemory performs similarity-based top-down memory traversal to retrieve vulnerability features, patterns, and detection rules most relevant to the current detection task. SmartMemory then combines the retrieved knowledge with the code context to prompt the LLM to identify logic incompleteness. Phase-2. Memory update for continual vulnerability learning. After detection, SmartMemory extracts vulnerability features from the agent’s historical detection experience and external vulnerability knowledge, and abstracts them into higher-level patterns and detection rules. Three-layer memory structure. SmartMemory organizes its memory as a three-layer knowledge graph ordered by increasing

Fig. 6. Architecture of three-layer memory integrated by SmartMemory.

levels of abstraction (Fig. 6). Nodes within each layer encode vulnerability knowledge at the corresponding level, while interlayer edges capture abstraction and derivation relations. The feature layer captures structured knowledge of individual vulnerabilities. Each vulnerability is represented by six entitynode types: OFC Business Logic, Vulnerability, PoC, Patch, Subcontract, and Attack. Specifically, OFC Business Logic describes the involved business logic using the conceptualmodel triple ⟨o, κ, δ⟩. Vulnerability records the vulnerable code and a description of the inconsistency type and cause. PoC records the proof-of-concept code and attack steps. Patch records the patch code and fixing method. Subcontract denotes involved subcontracts. And Attack denotes the corresponding real-world attack event. Within each vulnerability, these entities are linked by edges to form its feature subgraph. Specifically, a vulnerability occurs in business logic, forms a PoC, and locate at a subcontract. An attack exploits the vulnerability and manipulates the business logic. A patch fixes the vulnerability and blocks the attack. Feature subgraphs of different vulnerabilities are connected by association edges based on shared business logic, reused subcontracts, or shared attack events. The pattern and rule layers contain natural-language nodes for vulnerability patterns and detection rules. Within these two layers, nodes are connected by the three edge types inherited from the feature layer (i.e., shared business logic, reused subcontract, and shared attack event), as well as inter-entity association edges introduced during memory update. Adjacent layers are linked by cross-layer edges that encode abstraction and derivation relations between upper and lower layer entities. These edges support top-down cross-layer projection during retrieval. Memory retrieval. SmartMemory traverses memory top-down, from the rule layer to the feature layer, as shown in Fig. 7. Similarity-based retrieval. At each layer, SmartMemory retrieves the top-k nodes most similar to a query q, where similarity is measured by the cosine similarity between the embeddings of the query and of each node’s content. S = TopKni ∈N

v(q) · v(ni ) . ∥v(q)∥ ∥v(ni )∥

(1)

Since embedding similarity can be noisy, SmartMemory further performs a one-hop expansion over the graph, adding the outgoing and incoming neighbors of each node in S to obtain

7

Fig. 7. Details of inconsistency identification in SmartMemory.

S̃, which captures both semantic and structural relevance. Top-down memory traversal. SmartMemory retrieves layer by layer, from the rule layer down to the feature layer. The query at each layer accumulates the knowledge retrieved at the layers above it. Moreover, except for the top layer, each layer first projects the results of the layer above onto itself along the cross-layer edges to form seed nodes, and then merges these seeds with its own retrieval results as the final results. (1) Rule layer. Using the targeted contract functions C as the query, SmartMemory retrieves and expands over Grule via Eq. (1) to obtain the rule-layer result R̃S . (2) Pattern layer. SmartMemory first projects R̃S along the rule-to-pattern cross-layer edges to the pattern nodes connected to them, forming a seed set Pproj . Here, projection follows the cross-layer edges to select adjacent nodes. SmartMemory then uses [C; R̃S ] as the query to retrieve and expand over Gpattern , and merges the result with Pproj to obtain P̃S . (3) Feature layer. Likewise, SmartMemory projects P̃S along the pattern-to-feature cross-layer edges to form the feature seeds Fproj . SmartMemory then uses [C; R̃S ; P̃S ] as the query to retrieve and expand over Gfeature via Eq. (1), and merges the result with Fproj to obtain F̃S . Memory learning. SmartMemory updates its memory using five operators, whose conditions and actions are as follows. • Abstract is triggered when a layer contains two or more

entities connected by association edges, such as shared business logic, shared attack events, or shared subcontracts. SmartMemory utilizes an LLM to abstract these entities into a higher-level entity in the adjacent upper layer.

• Merge merges two entities whose textual contents are

semantically equivalent or highly similar according to their embedding similarity. • Associate establishes an association edge between two entities whose neighborhood subgraphs are structurally similar, as measured by normalized graph edit distance [28]. • Remove deletes an entity if it is persistently not retrieved in subsequent detections or if new evidence falsifies it or renders it outdated. • Decompose utilizes an LLM to split a pattern-layer or rulelayer entity into finer-grained subentities when multiple detections show that it is overly case-specific, such as matching one case or failing to generalize to other variants. Due to page limit, the details of memory initialization, implementations of memory retrieval and update operators are available in our repository. SmartMemory adapts memory-based agent to OFCI detection through three domain-specific designs. (i) Content. The memory stores OFCI-specific entities (e.g., OFC business-logic entities). (ii) Semantics. The memory forms an audit-inspired abstraction hierarchy from features to patterns and rules. (iii) Operation. The memory-update operators and top-down retrieval operator are designed to refine and transfer OFCI knowledge. D. Vulnerability Discovery via Taint Analysis SmartMemory models the vulnerable functions identified above as OFCI indicators and analyzes whether they can be triggered by an external adversary and whether they can affect 1 ⃝ 5 in equivalence invariant) to critical OFC variables (e.g., ⃝–

8

induce inconsistency. To this end, SmartMemory constructs a cross on-chain/off-chain data-flow graph (CDFG) and performs taint analysis to verify each indicator’s reachability and impact, thereby determining the OFCI type. CDFG construction. For an OFC contract, an asset proof must be confirmed off chain by a relayer or witness through verification before being consumed by subsequent logic. This off-chain confirmation introduces a data-flow break between proof production and proof consumption, which conventional data-flow analysis cannot continuously capture. To address this, SmartMemory first uses Slither [29] to construct control-flow graphs (CFGs) for OFC contracts and then performs data-flow analysis on the CFGs to obtain local data-flow graphs (DFGs). 2 ⃝, 3 and ⃝ 4 in Fig. 1 serve Among the critical variables, ⃝, 3 represents as alignment points across local DFGs. Since ⃝ the asset value parsed off chain by the relayer or witness, SmartMemory does not analyze their implementation code 3 as a single node. SmartMemory then adds and abstracts ⃝ 2 and ⃝, 3 and the other two alignment edges, one between ⃝ 3 and ⃝. 4 These edges bridge proof production, offbetween ⃝ chain confirmation, and proof submission, thereby connecting local DFGs across contracts into a unified CDFG. Taint analysis. SmartMemory defines that sources are externally controllable public-function inputs while sinks are OFCI 1 ⃝. 5 SmartMemory indicators and OFC critical variables ⃝– adopts the taint-propagation method from a prior study [30] and confirms OFCIs using two conditions. Condition 1: External-call reachability. On CDFG, SmartMemory performs forward taint propagation from the sources. If taint reaches an OFCI indicator, the indicator is considered reachable, meaning that an external attacker can trigger it through controllable inputs. Condition 2: Critical-variable impact. Given a reachable indicator, SmartMemory checks whether taint can propagate to OFC critical variables involved in inconsistency checking. OFCI type depends on which critical-variable group the 1 and ⃝, 2 or ⃝ 4 and indicator affects. If the affected pair is ⃝ 5 SmartMemory classifies the case as Type-1 asset-exchange ⃝, 2 and ⃝, 3 SmartMemory inconsistency. If the affected pair is ⃝ classifies it as Type-2 proof-semantic inconsistency. We use the two cases in Fig. 2 as running examples to illustrate how SmartMemory identifies the two OFCI types. In Fig. 2(a), mint contains two minting paths. During contractlogic semantic extraction, SmartMemory annotates the two paths as ⟨Mint, Assign, onchain-asset⟩ at Line 2 and ⟨Mint, Assign, offchain-collateral⟩ at Lines 3–4. During agent-based inconsistency identification, SmartMemory inspects Lines 2–4 and finds that both paths incorrectly reuse chainAmount as the source amount. SmartMemory identifies and marks mint as an OFCI indicator. Taint analysis further verifies that mint can be triggered through a public entry and that taint can 4 and the minted asset ⃝. 5 affect the off-chain collateral proof ⃝ SmartMemory classifies this case as asset-exchange inconsistency. In Fig. 2(b), the event emitted by Token.deposit at Line 6 is annotated as ⟨Deposit, Emit, proof⟩, while the event emitted by Vault.returnAssets at Line 13 is annotated as ⟨Transfer, Emit, event⟩. SmartMemory inspects the two sites

and finds that their emitted events are structurally isomorphic. SmartMemory therefore marks returnAssets as an OFCI indicator. This proof isomorphism prevents the off-chain relayer or witness from distinguishing the forged proof from the genuine one. Taint analysis verifies that returnAssets is 2 and externally reachable and affects the produced proof ⃝ 3 the parsed value ⃝. And SmartMemory classifies it as proofsemantic inconsistency. V. E VALUATION A. Implementation and Evaluation Setup Research questions. The research questions are as follows: • RQ1. How effective is SmartMemory in detecting OFCIs? • RQ2. How do the individual components of SmartMemory contribute to OFCI detection? • RQ3. What are the runtime and monetary costs of SmartMemory? • RQ4. How effective is SmartMemory at finding zero-day vulnerabilities? We use GPT-5-mini as the backbone LLM for SmartMemory during evaluation. All results are averaged over 5 runs on an Ubuntu 20.04 server with an Intel i9-10980XE CPU (3.0 GHz), an RTX 3090 GPU, and 250 GB RAM. Dataset and groundtruth establishment. We prepare two datasets and collection protocol (sources, search, exclusion, deduplication, and annotation schema) is available in repository. Manually labeled OFCI dataset (Dlabeled ). To build ground truth for evaluating SmartMemory, we collect real-world OFCI incidents from public security reports and disclosures, covering 48 distinct OFC applications (i.e., DApps). By reviewing the reports and contract code, we manually annotate 81 OFCIs. Large-scale OFC contract dataset (Dlarge ). To evaluate the real-world performance of SmartMemory on previously unseen OFCIs, we exhaustively search the community and the Internet for OFC applications and collect 325 applications, comprising 5,158 code documents. We prevent information leakage through following isolation measures. First, the conceptual model contains only schemalevel labels and encodes no application-specific instances. For Dlabeled , under leave-one-out reconstruction, we remove any information relevant to target DApp from vulnerability base and historical trajectories. The memory is rebuilt from the remaining N-1 applications. Third, for Dlarge , we remove its overlap with Dlabeled before evaluation, so the agent trained on Dlabeled is tested only on contracts unseen in the labeled set. Together, these steps prevent target applications from entering memory construction or reappearing through labeled-set overlap. B. RQ1: Overall Effectiveness To answer RQ1, we evaluate SmartMemory on the manually labeled OFCI dataset (Dlabeled ) using precision and recall. Following the leave-one-out reconstruction, we apply SmartMemory to each OFC application and manually compare its results with the ground truth (81 OFCIs in 48 applications). Here, we evaluate SmartMemory’s overall effectiveness in RQ1

9

TABLE I OVERALL EFFECTIVENESS OF S MART M EMORY ON Dlabeled . OFCI AEI PI Total

TP 60 11 71

Precision FP Rate 14 81.08% 3 78.57% 17 80.68%

TP 60 11 71

Recall FN Rate 8 88.24% 2 84.62% 10 87.65%

Note: AEI: Asset-exchange inconsistency; PI: Proof-semantic Inconsistency.

TABLE II C OMPARISON RESULTS BETWEEN S MART M EMORY AND THE ABLATION BASELINES OVER Dlabeled . Component

Approach

TP Logic extraction w/o logic extraction 57 37 Inconsistency Baseline (A) A + memory (B) 51 detection B + retrieval ops (C) 67 SmartMemory Full version 71 (C + update ops)

Precision FP Rate 33 63.33% 25 59.68% 20 71.83% 24 73.63%

TP 57 37 51 67

Recall FN Rate 24 70.37% 44 45.68% 30 62.96% 14 82.72%

17 80.68% 71 10 87.65%

and conduct fair comparisons with SOTA tools within their respective detection scopes in RQ2 below. As shown in Table I, SmartMemory achieves 80.68% OFC contracts, and its impact is reflected in both precision and precision and 87.65% recall on Dlabeled , with consistently recall. We compare the full SmartMemory with an ablation strong performance on both subtypes. The results indicate that version without this module on Dlabeled . Table II reports the results. Without this module, precision SmartMemory identifies OFCIs for OFC contracts effectively. drops to 63.33% and recall to 70.37%, confirming that contractFalse positives and false negatives. We inspect all 17 false logic semantic extraction substantially improves both metrics. positives and find that most of them originate from incomplete Moreover, 14 of 24 false negatives introduced by the ablation contract availability. Some DApps do not open-source all version can be resolved by our proposed method. We revisit the deployed contracts, causing unresolved external calls during motivating example in Fig. 2(a). Without semantic extraction, static analysis. Due to control-flow breaks, SmartMemory it cannot distinguish the two minting paths and therefore inevitably performs analysis with incomplete logic context, misses the conflicting use of chainAmount, causing a false which increases false positives. We also analyze the 10 false negative. In contrast, SmartMemory recognizes that the two negatives and find that they mainly depend on non-contract paths correspond to different primitive operations with different components. Among them, 4 rely on cryptographic verification expected parameters and eliminates the false negative. logic in third-party libraries outside the OFC applications, and Impact of agent-driven inconsistency identification. As 6 depend on real-time data sources such as oracles. Temporal-holdout analysis. Since Dlabeled is from public discussed earlier, three-layer memory, top-down retrieval and security reports, some cases may have appeared in the backbone OFCI-oriented update are three key designs of memory model’s pre-training corpus and potentially introduce pre- mechanism to support agent-driven inconsistency identification. training leakage. To evaluate this risk, we split Dlabeled into We evaluate this component with three ablation versions. We first implement a baseline version of SmartMemory that pre-cutoff and post-cutoff subsets based on each OFCI’s removes all memory designs and relies on standalone LLM readisclosure date and the model’s knowledge cutoff date. Note soning. We then implement multiple versions of SmartMemory that vulnerabilities disclosed after a model’s knowledge cutoff by incrementally adding memory content, retrieval and update are unlikely to have been included in its pre-training corpus. operators. We run all versions of SmartMemory on Dlabeled , We run SmartMemory on the full Dlabeled using GPT-4o as and further measure their precision and recall. the holdout backbone. We use GPT-4o instead of GPT-5-mini As shown in Table II, each design of memory mechanism because its knowledge cutoff (i.e., October 2023) is earlier and contributes positively to the better performance of Smartleaves enough post-cutoff OFCIs for analysis. We compare the Memory, i.e., its precision and recall increase monotonically. precision and recall of SmartMemory on the two subsets. Overall, the results indicate that memory mechanism helps Among the 81 OFCIs, 56 were disclosed before the cutoff SmartMemory enhance both precision and recall. date and 25 after it. With GPT-4o as the backbone, SmartMemory achieves 81.48% (22/27) precision and 88.00% (22/25) Impact of taint-analysis-based vulnerability discovery. Taintrecall on the post-cutoff subset, comparable to the 80.00% analysis-based vulnerability discovery is another key compo(48/60) precision and 85.71% (48/56) recall on the pre-cutoff nent of SmartMemory to reveal the hidden OFCIs, and improves subset. The results suggest that SmartMemory remains effective the recall of OFCI detection. ChainSniper and BridgeGuard are on OFCIs without appearing in model pre-training, and its not open-sourced, XGuard, Darcher and HGuard do not support effectiveness is unlikely to be driven by pre-training leakage. OFCI detection. We conduct scope-matched comparisons with Notably, SmartMemory attains comparable precision and recall applicable SOTA tools. We compare SmartMemory with under both GPT-5-mini and GPT-4o, indicating that the method SmartAxe [2] on access control and semantic vulnerabilities using 24 DApps from Dlabeled , covering 41 OFCI instances. For is model-agnostic and remains effective for different LLMs. VCScope [10], which targets signature-verification vulnerability, we select 9 DApps from Dlabeled , covering 16 OFCI instances. C. RQ2: Ablation Study We then compare recall on each evaluation set. Impact of contract-logic semantic extraction. To answer As shown in Table III, SmartMemory achieves higher recall RQ2, we evaluate contract-logic semantic extraction (Sec- than both baselines within their respective scopes. Against the tion IV-B), a core component of SmartMemory. It maps 68.29% recall of SmartAxe, SmartMemory achieves 87.80% diverse implementations to a common conceptual model, recall. Against VCScope, SmartMemory reaches 87.50% recall, enabling business-logic identification across heterogeneous compared with 43.75%. Further, manual investigation shows

10

TABLE III S COPE - MATCHED COMPARISON BETWEEN S MART M EMORY AND SOTA TOOLS WITHIN THEIR RESPECTIVE SCOPES ON Dlabeled . Scope

Approach

Cross-chain bridge (24 DApps, 41 OFCIs) Signature verification (9 DApps, 16 OFCIs)

SmartAxe [2] SmartMemory VCScope [10] SmartMemory

TP 28 36 7 14

Recall FN Rate 13 68.29% 5 87.80% 9 43.75% 2 87.50%

that our method can eliminate 8 of 13 false negatives from SmartAxe and 7 of 9 from VCScope. The results indicate that taint-analysis-based vulnerability discovery helps SmartMemory outperform SOTA approaches. D. RQ3: Performance Overhead To answer RQ3, we evaluate the efficiency of SmartMemory on Dlabeled in terms of runtime and monetary cost. Table IV reports the average runtime and token usage per DApp for each stage. Contract-logic semantic extraction takes 110.66 seconds and 20,658 tokens (costing $0.0072); agent-driven inconsistency identification takes 57.70 seconds and 25,492 tokens (costing $0.0141); and vulnerability discovery takes 36.48 seconds. Overall, SmartMemory uses 204.84 seconds and 46,150 tokens per DApp on average, with an estimated cost of $0.0213. The results indicate that SmartMemory is practical in terms of runtime and cost. It can integrate locally deployed open-source LLMs to further reduce monetary cost. TABLE IV T HE AVERAGE TIME AND TOKENS FOR ANALYZING EACH DAPP IN Dlabeled . Efficiency of SmartMemory Contract logical semantic extraction Agent-based inconsistency identification Vulnerability discovery Total

Avg. time (s) Avg. tokens 110.66 20658 ($0.0072) 57.70 25492 ($0.0141) 36.48 – 204.84 46150 ($0.0213)

invited six domain experts and organized them into two annotator groups and one referee group. Each annotator had at least two years of experience, and each referee had at least four years of domain research experience. The annotator groups labeled independently and then cross-validated their results until consensus was reached. We used Cohen’s Kappa to measure agreement. Cases below the threshold were escalated to the referee group for final adjudication. The average Cohen’s Kappa across all pairs is 0.764, indicating substantial agreement. VI. R ELATED W ORK Smart contract vulnerability detection. Research on the topic has evolved through several stages. Early work relies on program analysis, including symbolic execution [31], pattern matching [32], [33], [34], and formal verification [35]. Later studies use deep learning, such as attention-based BiLSTM [36] and graph neural networks [37]. More recently, LLM-based methods have emerged. While prior study [38] combines LLMs with program analysis, other studies [39], [40] fine-tune smaller LLMs or propose a two-stage framework of detection and explanation for auditing. However, these methods either rely on predefined patterns or lack continual learning from detection experience, limiting their adaptability to emerging OFC-specific vulnerability variants. Program analysis for OFC contract. Several approaches have been proposed for OFC contract security. SmartAxe [2] constructs access control models and cross-chain controlflow/data-flow graphs to detect cross-chain vulnerabilities. VCScope [10] employs a two-stage LLM-based summarization method for vulnerability identification. ChainSniper [11] trains machine learning models on a limited dataset for cross-chain vulnerability detection. BridgeGuard [9] proposes a static analysis framework for detecting external interaction vulnerabilities. Other notable efforts focus on anomaly detection for contract behavior. XGuard [14] applies predefined rules to detect inconsistent behaviors in cross-chain contracts. Hguard [15] uses manually constructed dynamic response graphs as formal specifications to identify logic deviations behavior.

E. RQ4: Finding Zero-days in the Real World To evaluate SmartMemory’s ability to discover previously unseen vulnerabilities, we apply it to Dlarge , which includes 325 real-world OFC DApps. This large-scale study is conducted VII. C ONCLUSION with a professional security auditing firm that integrated SmartMemory into its auditing workflow. SmartMemory reports In this paper, we propose SmartMemory, an agent-driven 69 warnings (including 55 TPs and 14 FPs confirmed manually) static-analysis framework to detect OFCIs for OFC smart across the 325 applications. After further manual inspection, contracts. On a manually labeled dataset of 81 OFCIs from our domain experts confirm 36 of the 55 TPs are exploitable 48 OFC DApps, SmartMemory achieves 80.68% precision and zero-day OFCIs. The detailed inspection criteria are described 87.65% recall. We further apply SmartMemory to 325 realin Section V-F. For each confirmed vulnerability, we construct world OFC DApps in collaboration with a professional security a proof-of-concept (PoC) exploit to verify exploitability. After auditing firm, identifying 36 previously unseen zero-day OFCIs, responsible disclosure, we report all confirmed cases to the all of which have been acknowledged and patched. corresponding developers. As of submission, all of the 36 zeroday vulnerabilities have been acknowledged and fixed. The VIII. DATA AVAILABILITY disclosure materials are available in our repository. F. Threat to Validity One internal threat is subjective bias in manual analysis during dataset collection and evaluation. To mitigate it, we

The repository (https://doi.org/10.5281/zenodo.21006162) provides the artifact, datasets, vulnerability statistics, complete prompts, parameter settings, embeddings, dataset collection protocol, memory implementations and disclosure materials.

11

R EFERENCES [1] J. Chen, X. Xia, D. Lo, J. Grundy, X. Luo, and T. Chen, “Defining smart contract defects on ethereum,” IEEE Transactions on Software Engineering, vol. 48, no. 1, pp. 327–345, 2020. [2] Z. Liao, Y. Nan, H. Liang, S. Hao, J. Zhai, J. Wu, and Z. Zheng, “Smartaxe: Detecting cross-chain vulnerabilities in bridge smart contracts via fine-grained static analysis,” Proceedings of the ACM on Software Engineering, vol. 1, no. FSE, pp. 249–270, 2024. [3] S. Chen, M. Jiang, and X. Luo, “Exploring the security issues of real world assets (rwa),” in Proceedings of the Workshop on Decentralized Finance and Security, 2024, pp. 31–40. [4] M. Y. Guan, Y. Yu, T. Sharma, M. Z. Huang, K. Qin, Y. Wang, and K. Y. Wang, “Security perceptions of users in stablecoins: Advantages and risks within the cryptocurrency ecosystem,” in 2025 IEEE Symposium on Security and Privacy (SP). IEEE, 2025, pp. 2753–2771. [5] “The statistic source for ofc vulnerability,” https://doi.org/10.5281/zenodo. 21006162, 2026, [Accessed 27-June-2026]. [6] “Polynetwork exploit,” https://en.wikipedia.org/wiki/Poly Network exploit, 2022, [Accessed 27-June-2026]. [7] “Synapse exploit,” https://medium.com/synapse-protocol/11-06-2021post-mortem-of-synapse-metapool-exploit-3003b4df4ef4, [Accessed 27June-2026]. [8] Sebastian Sinclair, “Blockchain protocol thorchain suffers $8m hack,” https://www.coindesk.com/markets/2021/07/23/blockchain-protocolthorchain-suffers-8m-hack/, 2021, [Accessed 20-Sep-2023]. [9] Z. Zhou, X. Luo, X. Ji, J. Mao, T. He, J. Wang, and Q. Wu, “Bridgeguard: Checking external interaction vulnerabilities in cross-chain bridge router contracts based on symbolic dataflow analysis,” IEEE Transactions on Dependable and Secure Computing, 2025. [10] H. Wang, K. Yan, and W. Diao, “From patterns to precision: Llmguided detection of signature verification flaws in smart contracts,” IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), pp. 240–251, 2026. [11] T.-D. Tran, K. A. Vo, D. T. Phan, C. N. Tan, and V.-H. Pham, “Chainsniper: A machine learning approach for auditing cross-chain smart contracts,” in Proceedings of the 2024 9th International Conference on Intelligent Information Technology, 2024, pp. 223–230. [12] W. Ou, S. Huang, J. Zheng, Q. Zhang, G. Zeng, and W. Han, “An overview on cross-chain: Mechanism, platforms, challenges and advances,” Computer Networks, vol. 218, p. 109378, 2022. [13] W. Zhang, L. Wei, S. Li, Y. Liu, and S.-C. Cheung, “Darcher: Detecting on-chain-off-chain synchronization bugs in decentralized applications,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2021, pp. 553–565. [14] K. Wang, Y. Li, C. Wang, J. Gao, Z. Guan, and Z. Chen, “Xguard: Detecting inconsistency behaviors of crosschain bridges,” in Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering, 2024, pp. 612–616. [15] M. Eshghie, C. Artho, H. Stammler, W. Ahrendt, T. Hildebrandt, and G. Schneider, “Highguard: Cross-chain business logic monitoring of smart contracts,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, 2024, pp. 2378–2381. [16] L. Zhou, X. Xiong, J. Ernstberger, S. Chaliasos, Z. Wang, Y. Wang, K. Qin, R. Wattenhofer, D. Song, and A. Gervais, “Sok: Decentralized finance (defi) attacks,” in 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2023, pp. 2444–2461. [17] P. Han, Z. Yan, W. Ding, S. Fei, and Z. Wan, “A survey on cross-chain technologies,” Distributed ledger technologies: research and practice, vol. 2, no. 2, pp. 1–30, 2023. [18] H. Mao, T. Nie, H. Sun, D. Shen, and G. Yu, “A survey on crosschain technology: Challenges, development, and prospect,” IEEE Access, vol. 11, pp. 45 527–45 546, 2022. [19] P. Zhao, Z. Jin, and N. Cheng, “An in-depth survey of large language model-based artificial intelligence agents,” arXiv preprint arXiv:2309.14365, 2023. [20] Y. Hu, S. Liu, Y. Yue, G. Zhang, B. Liu, F. Zhu, J. Lin, H. Guo, S. Dou, Z. Xi et al., “Memory in the age of ai agents,” arXiv preprint arXiv:2512.13564, 2025. [21] Z. Zhang, Q. Dai, X. Bo, C. Ma, R. Li, X. Chen, J. Zhu, Z. Dong, and J.-R. Wen, “A survey on the memory mechanism of large language modelbased agents,” ACM Transactions on Information Systems, vol. 43, no. 6, pp. 1–47, 2025. [22] W.-C. Huang, W. Zhang, Y. Liang, Y. Bei, Y. Chen, T. Feng, X. Pan, Z. Tan, Y. Wang, T. Wei et al., “Rethinking memory mechanisms

of foundation agents in the second half: A survey,” arXiv preprint arXiv:2602.06052, 2026. [23] J. Xie, Z. Chen, R. Zhang, X. Wan, and G. Li, “Large multimodal agents: A survey,” arXiv preprint arXiv:2402.15116, 2024. [24] W. Li, Z. Liu, J. Chen, Z. Liu, and Q. He, “Towards blockchain interoperability: A comprehensive survey on cross-chain solutions,” Blockchain: Research and Applications, p. 100286, 2025. [25] “Wormhole bridge exploit incident analysis,” https://www.certik. com/blog/wormhole-bridge-exploit-incident-analysis, [Accessed 26-Mar2026]. [26] “Nomad bridge hack: Root cause analysis,” https://medium.com/nomadxyz-blog/nomad-bridge-hack-root-cause-analysis-875ad2e5aacd, [Accessed 27-June-2026]. [27] “Defillama,” https://defillama.com/, 2026, [Accessed 27-June-2026]. [28] H. Bunke, “On a relation between graph edit distance and maximum common subgraph,” Pattern recognition letters, vol. 18, no. 8, pp. 689– 694, 1997. [29] J. Feist, G. Grieco, and A. Groce, “Slither: a static analysis framework for smart contracts,” in 2019 IEEE/ACM 2nd international workshop on emerging trends in software engineering for blockchain (WETSEB). IEEE, 2019, pp. 8–15. [30] Z. Liao, S. Hao, Y. Nan, and Z. Zheng, “Smartstate: Detecting state-reverting vulnerabilities in smart contracts via fine-grained statedependency analysis,” in Proceedings of the 32nd ACM SIGSOFT International symposium on software testing and analysis, 2023, pp. 980– 991. [31] L. Luu, D.-H. Chu, H. Olickel, P. Saxena, and A. Hobor, “Making smart contracts smarter,” in Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, 2016, pp. 254–269. [32] S. Tikhomirov, E. Voskresenskaya, I. Ivanitskiy, R. Takhaviev, E. Marchenko, and Y. Alexandrov, “Smartcheck: Static analysis of ethereum smart contracts,” in Proceedings of the 1st international workshop on emerging trends in software engineering for blockchain, 2018, pp. 9–16. [33] Z. Liao, Z. Zheng, X. Chen, and Y. Nan, “Smartdagger: a bytecode-based static analysis approach for detecting cross-contract vulnerability,” in Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, 2022, pp. 752–764. [34] Z. Liao, Y. Nan, Z. Gao, H. Liang, S. Hao, J. Wu, and Z. Zheng, “Satellite: Detecting and analyzing smart contract vulnerabilities caused by subcontract misuse,” IEEE Transactions on Software Engineering, 2025. [35] P. Tsankov, A. Dan, D. Drachsler-Cohen, A. Gervais, F. Buenzli, and M. Vechev, “Securify: Practical security analysis of smart contracts,” in Proceedings of the 2018 ACM SIGSAC conference on computer and communications security, 2018, pp. 67–82. [36] P. Qian, Z. Liu, Q. He, R. Zimmermann, and X. Wang, “Towards automated reentrancy detection for smart contracts based on sequential models,” IEEE access, vol. 8, pp. 19 685–19 695, 2020. [37] Y. Zhuang, Z. Liu, P. Qian, Q. Liu, X. Wang, and Q. He, “Smart contract vulnerability detection using graph neural networks,” in Proceedings of the twenty-ninth international conference on international joint conferences on artificial intelligence, 2021, pp. 3283–3290. [38] Y. Sun, D. Wu, Y. Xue, H. Liu, H. Wang, Z. Xu, X. Xie, and Y. Liu, “Gptscan: Detecting logic vulnerabilities in smart contracts by combining gpt with program analysis,” in Proceedings of the IEEE/ACM 46th international conference on software engineering, 2024, pp. 1–13. [39] Z. Wei, J. Sun, Z. Zhang, X. Zhang, and Z. Hou, “Ftsmartaudit: A knowledge distillation-enhanced framework for automated smart contract auditing using fine-tuned llms,” arXiv preprint arXiv:2410.13918, 2024. [40] W. Ma, D. Wu, Y. Sun, T. Wang, S. Liu, J. Zhang, Y. Xue, and Y. Liu, “Combining fine-tuning and llm-based agents for intuitive smart contract auditing with justifications,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 2025, pp. 1742– 1754.

Record · ID 1108626 · SHA-256 4f14076765a56126
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.