Conceptio › Archive › arXiv CS
arXiv CSopen access

SkelOT: Reusing AOT Compilation Across EVM Contract Families

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

SkelOT: Reusing AOT Compilation Across EVM Contract Families Sipeng Xie1 , Qianhong Wu1 , Minghang Li1 , Qin Wang4 , Zhipeng Wang3 , Bo Qin2 1 Beihang University | 2 Renmin University of China 3 The University of Manchester | 4 Independent

arXiv:2609.24404v1 [cs.CR] 21 Sep 2026

Abstract

families: sets of contracts that share instruction structure and differ only in a small set of embedded constants. Factorydriven mass deployment, proxy architectures, and templatebased reuse have been documented across the ecosystem [20, 43, 47]. Naive per-hash AOT compiles each family member separately. The cost is redundant compile work, a nativeartifact cache that scales with deployment count rather than with the number of families, and reduced workload coverage under a finite compile budget. Prior smart-contract research has studied clone prevalence and code reuse [20, 24, 43], factory and proxy deployment patterns [47, 52], and contract similarity detection [27, 28]. These works treat contracts as deployment or similarity objects, not as units of compilation work. The compilationgranularity problem has not been raised.

Ahead-of-time (AOT) compilers (e.g., revmc, evmone, and DTVM) for the Ethereum Virtual Machine (EVM) reuse compilation artifacts at contract-code-hash granularity. This granularity is poorly matched to real EVM workloads dominated by contract families: factory-, proxy-, and templatedriven deployments that share instruction structure but differ in a small set of embedded constants. Across four EVM chains (Base, Ethereum, BSC, and Arbitrum), we find that 23.1–47.6% of unique compilable bytecodes map to shared family skeletons within 10K-block windows. Per-hash AOT therefore redundantly recompiles structurally equivalent code, inflating compile time and artifact footprint while reducing workload coverage under finite compile budgets. We present SkelOT, an AOT framework that lifts the unit of compilation reuse from code hash to family skeleton. SkelOT compiles one native artifact per family, bakes invariant constants into the artifact, and reads variant constants from a per-contract runtime table. Built on revmc/LLVM and evaluated on a 10K-block Base mainnet corpus (3.52M transactions), SkelOT reduces compilation units by 47.5%, artifact footprint by 57.4%, and compile time by 2.19×, while preserving byte-identical execution outcomes versus per-hash AOT. At runtime, SkelOT delivers a 1.31× median per-contract speedup across family members. Under a compile budget targeting 75% execution-time coverage, SkelOT needs far fewer artifacts than per-hash AOT, and the advantage holds at every coverage target.

Figure 1. SkelOT overview: Naive AOT (left) compiles per code hash; our SkelOT (right) compiles once per family and stores only differing constants in per-contract tables.

Keywords: Ethereum Virtual Machine, ahead-of-time compilation, code reuse, contract families, smart contracts

1

In this paper, we reframe the problem. Exact code hash is a good identity key for a deployed contract but a poor reuse unit for AOT compilation. The right reuse unit is the contractfamily skeleton. We define it by what an AOT compiler can reuse across family members, not by what a clone detector identifies as similar. We present SkelOT, a skeleton-aware AOT framework. SkelOT extracts the shared skeleton from each contract family and applies an invariant/variant constant split. It makes this split once per family, by comparing the bytecode of every member. Invariants hold the same value across every family member and are baked into the compiled skeleton as compile-time constants. Variants differ across members and are read at runtime from a per-contract table. One compile

Introduction

Production ahead-of-time (AOT) compilers for the Ethereum Virtual Machine (EVM) index compilation artifacts by percontract bytecode hash. Systems such as revmc [37], evmone [21], and DTVM [53] all use this granularity. Each deployed contract is an independent compilation target. Two contracts that differ only in their embedded constants are compiled and cached as separate artifacts. Per-hash AOT, therefore, does not share compiled code across contracts. Modern EVM workloads, however, are not flat collections of independent contracts. They are dominated by contract ⊲ Accepted by European Conference on Computer Systems (EuroSys’27). ★ Most of Z. Wang’s work was made before The University of Manchester.

1

therefore covers every member of the family. Under a finite compile budget, this either saves compile time or covers more of the workload. We implement SkelOT on revmc [37] and evaluate it on four representative EVM-compatible chains, one 10K-block window per chain. Across compile time, native-artifact footprint, runtime dispatch, and budgeted coverage, SkelOT improves over Naive AOT while preserving byte-identical execution outcomes. Our contributions are summarized as follows:

a stack machine over 256-bit words. A contract’s bytecode is a sequence of opcodes. A small subset of those opcodes, the PUSH opcodes PUSH1 through PUSH32, carry an inline immediate operand of 1 to 32 bytes loaded onto the stack at execution time. Every other opcode acts on values already on the stack. Control flow uses JUMP and JUMPI, whose target must coincide with a JUMPDEST opcode in the same contract, and targets that fail this check trap. A few opcodes also read a contract’s bytecode as data, such as CODESIZE and CODECOPY. Each deployed contract is identified on chain by a code hash, the keccak digest of its bytecode. Solidity also appends a metadata trailer to the bytecode, a short fingerprint of the source and compiler settings. The trailer is data and never executes, yet the code hash covers it like every other byte. Existing EVM AOT engines lower each contract’s bytecode into native code without changing this execution model. The host compiler we build on (revmc, sitting on LLVM [25]) translates the opcode stream into LLVM IR. It materializes the 256-bit stack as IR-level memory that LLVM can promote to SSA values [13], and lowers arithmetic and stack manipulation directly into IR operations the optimizer can fold and specialize. Memory and storage opcodes, by contrast, become opaque calls to host runtime routines, so the IR pipeline treats them as black-box functions and does not optimize across them. Control flow is reconstructed from the bytecode and lowered into native branches, with one analysis pass folding constant-target jumps into direct branches before the IR is handed to LLVM. A developer compiles source to bytecode once with solc [42] before contract deployment. Every node that wants native execution must compile deployed bytecode, and it keeps compiling as new contracts arrive, on the same machine resources that serve execution. Fees are charged by the opcodes of the original bytecode, so native code only changes how fast an opcode runs.

• We identify a structural mismatch between Naive AOT and modern EVM workloads: per-hash compilation follows bytecode identity, whereas family-driven deployment creates many contracts with the same instruction skeleton but different embedded constants. Across four public EVM-compatible networks (Base, Ethereum, BSC, and Arbitrum), 23.1–47.6% of unique compilable bytecodes collapse into shared family skeletons. We frame the issue as a compilation-granularity problem (§2, §3). • We propose skeleton extraction with an invariant/variant constant split as a reuse abstraction. It preserves compile-time optimization for invariant constants while externalizing only the variant subset that differs across family members (§4). • We deliver an end-to-end implementation covering skeleton extraction, variance analysis, family-level compilation, runtime constant-table execution, and family-aware caching and dispatch (§5). • We conduct a large-scale empirical validation on a 10K-block Base mainnet corpus with 3.52M transactions. SkelOT reduces compilation units by 47.5%, native-artifact footprint by 57.4%, and compile time by 2.19×, while preserving byte-identical outcomes against per-hash AOT across 9,997 non-empty blocks. It also achieves a 1.31× median runtime speedup. SkelOT also needs a smaller compile budget to cover a given share of execution time than per-hash AOT, with the advantage holding at every coverage target (§6).

2

2.2

We call the conventional design of existing EVM AOT compilers Naive per-hash AOT. This baseline compiles each distinct code hash separately. In other words, if two deployed contracts do not have exactly the same bytecode, they are treated as different compilation targets, even when most of their code structure is shared. This design is simple and safe. A code hash is a clear onchain identifier, so it gives the compiler an easy cache key and a clean correctness boundary. Each native artifact maps to one deployed contract, and the compiler only needs to iterate over the unique hashes in the workload. The limitation is that code hash is used not only as a deployment identity, but also as the compilation-reuse boundary. As a result, existing AOT systems can reuse native code only when two contracts are byte-for-byte identical, missing reuse opportunities among contracts that share the same structure but differ only in embedded constants.

Why Is Native AOT Insufficient?

Existing EVM AOT engines compile at code-hash granularity, so deployment identity also becomes the reuse boundary. This boundary is simple but too fine-grained for modern workloads. We show why in this section. 2.1

Naive Per-Hash AOT

EVM Execution And AOT Compilation

The chain keeps a global state that holds account balances, contract storage, and the bytecode of every deployed contract. The state advances one block at a time, as each block’s transactions execute the bytecode of the contracts they call, reading and writing the rest of the state. Contracts also call other contracts, so a single transaction can expand into a stack of nested calls. Ethereum Virtual Machine [49, 51] is 2

2.3

Table 1. Cross-chain skeleton deduplication.

Why Identity-Based Reuse Fails In Practice

The assumption fails in practice because three forces shape modern EVM deployment and decouple structural novelty from deployment count. First, contracts are typically built from shared libraries, audited base implementations, and source-level templates, so composition-level reuse propagates almost unchanged into bytecode-level near-duplicates [20, 43]. Second, factory contracts massinstantiate near-identical logic from a single source, and members of the resulting family differ only in constructorbaked constants such as owner addresses, token identifiers, pool parameters, or fee tiers [47]. Third, proxy-heavy architectures separate deployment identity from implementation, so a single implementation serves many proxy instances, and both the proxy layer and the implementation layer produce clusters of structurally identical contracts in the on-chain workload [47, 52]. Each force produces families of contracts that share instruction structure and differ only in a bounded set of embedded constants. Every member of such a family still carries a distinct code hash, so Naive AOT compiles each member separately.

2.4

Chain (blocks) BASE (38,004,930–38,014,929) ETH (23,000,000–23,009,999) BSC (95,897,884–95,907,883) ARB (458,863,320–458,873,319)

#codes #skels Redu. (%) Window range (%) 19,130 35,215 14,444 2,895

10,025 24,406 9,996 2,225

47.60 30.69 30.79 23.14

[40.91, 42.42] [26.83, 28.43] [24.67, 27.13] [16.34, 21.03]

- Reduction (Redu.) is computed over each 10K-block corpus. Window range reports

ten within-chain 1K-block sub-windows.

3

Contract Families in EVM Workloads

We use family to denote a compiler-relevant reuse boundary, formally defined in §3.1. This section shows that the boundary is not only common in the studied EVM workloads, but also appears where compilation cost is high and remains stable over time. 3.1

Defining A Contract Family

Our contract-family definition is driven by compilation reuse. Given an EVM bytecode sequence, its skeleton removes the immediate-value bytes of each PUSH instruction, strips Solidity metadata and trailing zero bytes, and keeps all opcodes in their original positions. A family is then the set of contracts whose skeleton byte sequences are bytewise identical [14]. §4.2 explains why this byte-level representation is suitable for compilation. Family membership is therefore a deterministic byte-equality check, rather than a similarity score. • For example, PUSH20 0xaaa... PUSH1 0x05 ADD and PUSH20 0xbbb... PUSH1 0x05 ADD share the skeleton PUSH20 * PUSH1 * ADD; replacing ADD with another opcode would place the contract in a different family. This bytewise boundary is necessary because SkelOT reuses compiled code, not merely analysis results. Ethereum clone-detection techniques usually cluster contracts under a similarity threshold [48]. Such clusters are useful for software analysis, but they do not identify a safe compiler reuse boundary. The compiler must share one instruction stream while specializing only contract-specific constants. Skeleton equality provides this boundary. Contracts in the same family have the same opcode layout and control-flow structure, and may differ only in immediate values that SkelOT later classifies as invariant or variant (§4.4). Contracts that are only similar, but not skeleton-identical, therefore fall outside SkelOT’s reuse model.

The Threefold Cost Of The Mismatch

The cost of the identity-equals-reuse assumption is threefold. First, the compiler performs redundant translation and optimization work across near-duplicate family members. It lowers, optimizes, and register-allocates the same instruction sequence once per member rather than once per family. Dynamic-compilation studies have flagged this redundancy as a dominant source of wasted warmup and end-to-end latency across short-lived deployments [38]. Second, the native code cache stores structurally duplicated artifacts. Its footprint grows with the deployment count of each family rather than with the number of distinct families in the workload, so memory and disk store the same compiled logic repeatedly. Production code-cache designs, by contrast, key on identity so that a cache hit substitutes a stored artifact for a full recompile [9, 44]. Third, and most consequential in production, redundant compile work reduces how much of the workload the AOT pipeline can cover under a finite budget. Every real deployment compiles under such a budget, whether a wall-clock window before a block is due, a memory ceiling on the compile cache, or a background-compilation cap shared with execution. The JIT-policy literature shows that such budgets force selective compilation, threshold tuning, and explicit tradeoffs between how aggressively methods are compiled and how quickly the hot working set is covered [22]. Compile budgets make the first two costs operationally visible. Wasted translation work and duplicated artifacts become uncovered contracts, which fall back to the interpreter and dominate tail latency.

3.2

Prevalence On The Corpus

We apply the same filter (§6) to one 10K-block window on each of four EVM-compatible networks: Base [12], Ethereum [49], BSC (BNB Smart Chain) [6, 26], and Arbitrum [7]. Table 1 lists the inspected block ranges and the per-chain corpus and dedup statistics. On every chain, the corpus of unique compilable bytecodes collapses onto a substantially smaller set of skeletons, yielding 10K-window reductions of 3

Table 2. Size-stratified skeleton deduplication on the 10K-block Base corpus. Bucket

#bytecode

#skeletons

Reduction (%)

Byte share (%)

tiny small medium large huge

570 1,204 4,052 6,114 7,190

166 513 2,609 4,256 2,484

70.9 (±10.1) 57.4 (±8.0) 35.6 (±5.6) 30.4 (±3.5) 65.5 (±5.8)

0.01 0.27 5.21 26.09 68.41

Total

19,130

10,025

47.6

100.00

- Reduction is reported per bucket with population std. dev. over ten 1K-block

windows. Byte share is the share of raw bytecode bytes. - Buckets follow member bytecode length, so a skeleton whose members straddle a boundary counts once in each bucket it spans; the skeleton rows sum to 10,028, three more than the 10,025 distinct skeletons.

47.60% on Base, 30.69% on Ethereum, 30.79% on BSC, and 23.14% on Arbitrum. The reduction also stays bounded across each chain’s 𝑛=10 non-overlapping 1K-block sub-windows, with ranges of [40.91%, 42.42%] on Base, [26.83%, 28.43%] on Ethereum, [24.67%, 27.13%] on BSC, and [16.34%, 21.03%] on Arbitrum. These bounded ranges show that the reduction is uniform across sub-windows rather than driven by one deployment burst. The corpus-level rate sits above every subwindow because a larger window merges more members into each family. Family structure is therefore a property of the EVM workload, not of any one chain.

Figure 2. Four-chain family-size distribution

A corpus-wide reduction would be less relevant to compilation if it were driven mainly by tiny proxies or stubs. On the Base corpus, it is not. We bucket bytecode by length into tiny (<100 B), small (100–999 B), medium (1,000–4,999 B), large (5,000–14,999 B), and huge (≥15,000 B) contracts. Skeleton reduction is strongest in the largest contracts, where any eventual compile-side reuse is most valuable. The huge bucket accounts for 68.41% of raw input bytes and reduces by 65.5%, while the tiny bucket accounts for only 0.01% of raw input bytes. Table 2 reports the full per-bucket breakdown with across-window variability and input-byte share. §6 then quantifies compile-side savings on this corpus.

on that chain, and the heavy-tailed shape recurs on every corpus. At the exact-code level, 1,672 code hashes recur on at least two chains and 145 recur on all four; at the skeleton level, the counts rise to 2,970 and 302. Thus the recurring pattern is not only byte-for-byte redeployment. It is shared contract structure across chains. Every cross-chain code hash maps to the same skeleton, which serves as a consistency check on the extraction. The singleton tail marks the boundary of family-level reuse. For instance, on the Base corpus, 9,200 of the 10,025 distinct skeletons appear only once. These singleton skeletons fall back to Naive AOT because there is no cross-member disagreement from which to derive variant positions. The reuse opportunity comes from the 825 multi-member families, which cover 9,930 unique compilable bytecodes, or 51.9% of the corpus.

3.4

4

3.3

Where Family Structure Concentrates

Family-Size Distribution

SkelOT Design

SkelOT raises the reuse unit of EVM AOT compilation from one deployed contract’s code hash to the skeleton shared by a contract family. The design therefore shares one compiled artifact across family members while keeping invariant constants visible to the optimizer, preserving per-contract execution behavior, and leaving most of the host AOT pipeline unchanged. Figure 3 shows the SkelOT compilation pipeline applied to a contract family.

Across Base, Ethereum, BSC, and Arbitrum, a small number of large families produce the reduction, not many accidental near-duplicates. The top families correspond to recognizable, long-lived protocol and deployment templates. They include Uniswap V3 [1] and PancakeSwap V3 [36] pool skeletons, PancakeSwap V3 liquidity-mining pool variants [36], EIP1167 [33] minimal proxies, TransparentUpgradeableProxy families [35], ClankerToken ERC-20 [11, 46] deployments, and Ethereum ERC1155Creator proxy/NFT-collection families [30, 40]. These templates appear repeatedly across the top-family lists and recur in the deployment stream rather than appearing as isolated duplicates. Figure 2 shows the four chains as matched log–log ranksize panels; each panel uses arrow callouts for the top families

4.1

Design Principles

SkelOT is guided by one reuse objective. The objective is to maximize reuse of compiled structure across contracts in the same family, so that compilation effort for one member benefits every member sharing its skeleton. 4

Two contracts belong to the same family exactly when their skeleton byte sequences match. The prototype identifies each skeleton by a fast hash of its byte sequence [2]. The hash only locates candidates, and the prototype then confirms membership by comparing the skeleton bytes, so the family check stays exact at an 𝑂 (1) lookup cost, and the hash is the family identifier used by the caching and dispatch machinery of §5. Byte equality also handles compiler-version skew. If two Solidity versions generate different code, the skeletons differ and the contracts never share an artifact. The EVM draws no boundary between code and data, so stripping has to establish one. The trailer’s last two bytes give its length. We read that length, check that the covered bytes have the CBOR shape of Solidity metadata, and only then strip them. A jump could still target a JUMPDEST byte inside the trailer, so we mark stripped positions as invalid jump targets in both the native code and our interpreter, and such a jump fails the same way on both sides. However, inside the code region there is no such boundary. If reachable code remains after a skeleton’s last terminator, execution could run off the end into the stripped bytes, so we conservatively reject that skeleton (§6.2 reports the counts). EOF-style formats [5] would separate code from data at the format level and remove the need for both checks. Two consequences follow from this minimal definition. First, two contracts that both push the value zero but use different opcodes to do so (a dedicated zero-PUSH opcode versus a one-byte PUSH opcode with a zero immediate) end up in different families, because the skeleton preserves the opcode choice rather than normalizing across alternative encodings of the same value. This is the one point at which the design trades a little reuse for a locally decidable notion of family equality that needs no cross-contract normalization. Second, the constant classification below requires at least two members. A singleton family compiles identically to percontract compilation, treating every PUSH immediate as a fixed compile-time constant, and gains no reuse benefit until a second member arrives and supplies the cross-member comparison that exposes variants.

Figure 3. SkelOT compilation pipeline.

There are three design constraints. The first constraint is semantic preservation. Shared execution must match percontract compilation in success or failure, gas, output, and post-state for every transaction. The second is runtime efficiency. Reuse should add no measurable overhead beyond noise. The third is deployability. The system should integrate with existing EVM AOT pipelines by reusing their standard bytecode translation, optimization, and native-code emission path rather than introducing a specialized execution engine. 4.2

Choosing A Byte-Level Skeleton Representation

Recall from §3.1 that a skeleton is a contract’s bytecode with PUSH immediates, the appended Solidity metadata section, and any trailing zero bytes removed, with every remaining opcode preserved in place; two contracts share a family when their skeleton byte sequences are bytewise identical. The representation is deliberately a lean byte sequence rather than a richer structural form such as an abstract syntax tree or a normalized control-flow graph. The host compiler’s existing bytecode analysis already reconstructs the control-flow graph and instruction boundaries from the opcode sequence during translation, and sharing that reconstruction across family members is exactly the reuse SkelOT exploits. A richer skeleton would not expand what can be shared; it would only add transformation stages between the raw bytecode and the work the backend already performs. 4.3

4.4

The Invariant/Variant Constant Split

We classify once, when a family forms, reading the bytecode of every member. A PUSH position is invariant if all members hold the same value there, and variant if any member differs. The family descriptor records this split. The compiler bakes the invariant values into one shared block of native code and leaves a table slot for each variant position. We number the slots in the order the variant positions appear in the skeleton, so every member fills a table with the same layout. We keep the descriptor fixed once written, so a later contract must pass admission (§4.7) to join the family; §7.4 discusses more details. Algorithm 1 lists the steps. Our central design insight is invariant/variant constant split. In real contract families, most PUSH immediates are

Skeleton Extraction And Family Membership

The skeleton extractor performs a single linear scan of the bytecode. Each opcode byte is emitted as-is, immediate bytes following a PUSH opcode are skipped, and the appended Solidity metadata section and any trailing zero bytes are removed, so deployments of the same logical contract that differ only in those trailing bytes share a single skeleton. 5

identical across members; we call these positions invariants. Only a small subset differs across members; we call these positions variants. SkelOT handles them differently. Invariants remain compile-time constants in the generated code, preserving the backend’s ability to fold constants, specialize surrounding code, and use them as static branch targets or comparison pivots. This folding is distinct from the folding that source-level compilers such as solc [42] perform while generating EVM bytecode, whose results are already frozen inside the PUSH immediates. The folding we mean happens one stage later, when the backend lowers bytecode to native code and can simplify around a literal immediate but not around a table load. SkelOT excludes variants from the compiled skeleton and loads them at runtime from a compact per-member table in a fixed skeleton-defined order. Thus, unlike per-contract compilation, which keeps every immediate constant but forgoes reuse, SkelOT retains the stable constants needed for optimization while externalizing only the values required for cross-member code sharing. The opposite design point is an all-variant scheme, which classifies every eligible immediate as a variant. This maximizes skeleton sharing but replaces compile-time constants with runtime table loads. We exclude immediates folded into control-flow targets, since changing them would invalidate the shared branch structure rather than simply update a table entry. For all remaining immediates, once the code loads a value from the per-member table, the backend can no longer fold arithmetic, specialize comparisons, or resolve branches around that value. These optimizations are a key reason why per-contract compilation outperforms interpretation on EVM workloads. An all-variant design therefore sacrifices compile-time specialization even for values that never vary across the family, leading to larger native artifacts, longer compile time, and extra runtime loads. We quantify these compile-side costs in §6.2. The observed structure of contract families motivates the split. Real families exhibit a stable skeleton with shallow variance, so preserving invariant positions as compile-time constants lets the backend keep the optimizations that matter for per-contract compilation. §6.1 quantifies this structure, showing that the multi-member population is overwhelmingly invariant, that the variant share shrinks further with contract size, and that the invariant classification is stable across the deployment horizon, not only within the measurement window. The ablation in §6.2 confirms that preserving invariants as compile-time constants accounts for much of the retained optimization benefit. 4.5

contract’s table and pushes the 256-bit value, exactly as a plain PUSH of an invariant constant would. The deployed bytecode is never rewritten. We choose between an inline immediate and a table load only when lowering to native code, so offsets, jump targets, and gas stay untouched. The shared artifact therefore carries the skeleton, the invariant constants, and the optimizations built around them, while the per-member table carries only what the family allows to differ. Each variant PUSH instruction is compiled with a fixed slot index that locates its entry in the per-member table, and the same slot assignment applies to every member, so a single artifact serves all members without per-member code generation. 4.6

Correctness Boundary

Let 𝐵𝑖 be an admitted member of a family, let 𝑠 = 𝑆 (𝐵𝑖 ) be its skeleton, and let 𝑥 be a transaction input. For each PUSH position 𝑝 in 𝑠, SkelOT classifies 𝑝 as invariant when every family member carries the same immediate at that position, and as variant otherwise. The descriptor assigns each variant position 𝑝 a slot 𝑗 (𝑝) in appearance order, and member 𝐵𝑖 supplies the table entry 𝑇𝑖 [ 𝑗 (𝑝)] containing exactly the immediate value at 𝑝 in its original bytecode. The relative preservation claim is Obs(SkelOT(𝑠,𝑇𝑖 , 𝐵𝑖 , 𝑥)) = Obs(NaiveAOT(𝐵𝑖 , 𝑥)), where Obs records success or failure, gas, output, and poststate. The claim is relative. Assuming the underlying percontract AOT compiler preserves the EVM semantics formalized by prior work [10], SkelOT preserves the behavior of that compiler while changing the reuse unit. Proof (sketch). The proof sketch is a case analysis over the executed instruction stream. Skeleton equality fixes the opcode layout, instruction boundaries, and PUSH widths for every family member. For any non-PUSH instruction, SkelOT and Naive AOT therefore execute the same EVM opcode at the same logical position. For an invariant PUSH position, SkelOT embeds the same 256-bit value that Naive AOT would have embedded for 𝐵𝑖 . For a variant PUSH position 𝑝, SkelOT loads 𝑇𝑖 [ 𝑗 (𝑝)], defined to be the same 256-bit value the original bytecode of 𝐵𝑖 would have pushed. Thus each semantic step sees the same opcode and the same stack operands in both executions. An induction over the trace gives identical observable behavior, provided the host compiler does not bake a value into native control flow and later allow that value to vary. □ Gas preservation. Gas equality follows from the same stepby-step argument. At step 𝑘, SkelOT and Naive AOT execute the same EVM opcode with the same stack operands. Static gas is therefore the same, since it is charged by opcode in the original instruction stream. Dynamic gas is also the same, since memory growth, copying, storage, and calls are all computed from the same operands in the same EVM context. The variant-table read does not change this argument. It

Runtime Model

At runtime, every member of a family executes the same compiled artifact. Each member supplies separately a small per-member table holding its variant values, one entry per variant position. When execution reaches a variant position, the compiled code reads the matching entry from the current 6

Algorithm 1 Preparing shared artifacts and member tables

happens inside the LLVM artifact before the EVM PUSH value reaches the stack, so it is not itself an EVM instruction and creates no separate gas event. Once the value reaches the stack, it is exactly the original immediate from 𝐵𝑖 , so any later gas formula that depends on that value receives the same input under SkelOT and Naive AOT.

Require: Contracts 𝐶 and host ahead-of-time compiler 𝐻 Ensure: Family registry 𝑅 and member map 𝑀 1: 𝑅 ← ∅; 𝑀 ← ∅ 2: G ← the partition of 𝐶 by skeleton hash. 3: for all family group 𝐺 ∈ G do 4: if |𝐺 | = 1 then 5: Let 𝑐 be the only contract in 𝐺. 6: 𝑎𝑐 ← 𝐻 .compileContract(𝑐 ) 7: 𝑀 [codeHash(𝑐 ) ] ← Direct(𝑎𝑐 ) ⊲ singleton path 8: continue 9: end if 10: 𝑠 ← skeletonHash(𝐺 ) 11: 𝐷 ← classifyPushes(𝐺 ) ⊲ split invariant and variant PUSHes 12: Choose a representative contract 𝑟 ∈ 𝐺. 13: 𝑎𝑠 ← 𝐻 .compileFamily(𝑟, 𝐷 ) ⊲ attempt shared compilation 14: if 𝑎𝑠 = ⊥ then 15: for all contract 𝑐 ∈ 𝐺 do 16: 𝑎𝑐 ← 𝐻 .compileContract(𝑐 ) 17: 𝑀 [codeHash(𝑐 ) ] ← Direct(𝑎𝑐 ) 18: end for 19: continue ⊲ fall back to direct artifacts 20: end if 21: 𝑅 [𝑠 ] ← (𝑎𝑠 , 𝐷 ) ⊲ register the shared artifact 22: for all contract 𝑐 ∈ 𝐺 do 23: 𝑇𝑐 ← ∅ 24: for all (𝑝, 𝑗 ) ∈ Variants(𝐷 ) do 25: 𝑇𝑐 [ 𝑗 ] ← pushImmediate(𝑐, 𝑝 ) 26: end for 27: 𝑀 [codeHash(𝑐 ) ] ← Shared(𝑠,𝑇𝑐 ) ⊲ member table 28: end for 29: end for 30: return (𝑅, 𝑀 )

Code environment. Code observation can also be stated in terms of the environment each opcode sees. The native artifact is only the execution vehicle. The EVM context presented to the compiled code still contains 𝐵𝑖 as the currentframe bytecode, along with the same memory, gas meter, return-data buffer, and host state that Naive AOT uses. For current-frame code queries, CODESIZE and CODECOPY read the executing member’s bytecode length and bytes from the per-call EVM context rather than from any compile-time constant, so the same artifact returns the correct length and bytes for every family member regardless of metadata or trailing-zero differences. Queries about another account, including EXTCODESIZE, EXTCODECOPY, and EXTCODEHASH, go through the same host-state lookup as Naive AOT and observe the code, length, or hash stored for the addressed account. SkelOT therefore changes the native artifact selected for execution, but not the EVM-visible code environment. Folded control flow. The final boundary concerns constants that the compiler has already turned into a native branch structure. If a pushed value is recognized as a static jump target, the compiler may bake that target directly into the native branch. That value can no longer vary safely across family members. A later member could supply a different jump target in its bytecode, but the shared native code would still branch to the baked target from the earlier member. SkelOT therefore keeps such positions outside the variant set. If a family member needs one of those values to differ, that member lies outside the current shared-artifact boundary rather than being admitted with an unsound table entry. Concretely, if a pushed jump target is folded into a native branch for one family member, reusing that branch for another member whose target differs would jump to the first member’s target. SkelOT fences that PUSH position from the variant set, so the family either shares an artifact with the target fixed or does not share an artifact at all. 4.7

Admission is intentionally conservative. If a candidate disagrees with an invariant position, SkelOT rejects that candidate from the existing family rather than rewriting the descriptor, changing the variant set, or migrating the compiled artifact. This keeps the online rule aligned with the batch correctness boundary. Versioned families and online reclassification could recover more reuse, but they require an artifact invalidation policy, which we leave to §7.4. An alternative all-variant configuration of SkelOT, evaluated in §6.2 and discussed in §7.4, trades a larger artifact and longer compile time for a more permissive admission rule over contracts that share the skeleton and do not require folded control-flow operands to vary.

Admission

5

The admission mechanism determines whether a newly observed contract can join an existing family and avoid a full recompile. Given the contract bytecode and the family’s invariant/variant classification, the admission descriptor checks two conditions: the contract skeleton must match the family skeleton, and every invariant position must contain the family’s expected value. If both checks pass, the contract is admitted, and its per-member table is filled with the constants from its variant positions.

Implementation

Compiler integration. The SkelOT prototype adds familylevel reuse to an existing EVM ahead-of-time compiler at four points. It extracts a skeleton key, classifies invariant and variant PUSH positions, compiles one shared artifact per reusable skeleton, and lowers only variant PUSH positions to indexed table loads. The host compiler’s bytecode analysis, control-flow reconstruction, optimization passes, and native object emission remain unchanged. 7

Artifact and member state. Compiled artifacts and member data live in two tiers. A family registry, keyed by the skeleton hash, stores each successfully shared native object with its descriptor. A member map, keyed by contract code hash, stores either a direct per-contract artifact or a sharedfamily key with the member’s variant table. Dispatch looks up the map entry. For a shared entry, it passes the member’s table pointer in the call context. Each call carries its own context, so nested calls into the same family each read their own table. The shared artifact stays stateless. Executing a family member needs no per-member code generation and no per-PUSH dispatch. The shared artifact carries the skeleton, the invariant constants, and the optimizations built around them, while the per-member table carries only what the family allows to differ.

95.89% (n=1833)

Ethereum

96.91% (n=632)

BSC

93.72% (n=173)

Arbitrum 90

92 94 96 98 Invariant PUSH-byte share (%) n = multi-member families on that chain

Figure 4. Cross-chain invariant share of embedded constants. Across 4 10K-block corpora, 93.72%–96.91% of pushimmediate bytes in multi-member families are invariant.

Preparation path. Algorithm 1 summarizes this preparation path. The procedure groups contracts by skeleton hash and treats single-member groups as ordinary per-contract compilation (lines 1–9). For a reusable group, it builds a family descriptor that records the invariant PUSH immediates and the table slots for variant PUSH positions (lines 10– 13). If the host compiler cannot build a shared artifact for that descriptor, the group falls back to per-contract artifacts (lines 14–20). When shared compilation succeeds, the registry stores the shared artifact and its descriptor once under the skeleton key (line 21). The member map then stores one entry per contract (lines 22–27). Each entry points to the shared skeleton and carries that contract’s concrete variant table, so dispatch can reuse the shared code while supplying the member-specific PUSH values. SkelOT populates the registry only when shared compilation succeeds, and populates the member map for every compiled contract. Direct entries point to ordinary per-contract artifacts. Shared entries point through the family key and carry the concrete table used by that member.

6

94.92% (n=825)

Base

Direct wall-clock comparisons with other EVM execution engines would mix this variable with implementation language, intermediate representation design, compiler backend, cache organization, and JIT or AOT policy. We treat external engines such as evmone and DTVM as contextual systems in §8 rather than as causal baselines. We use Base mainnet as the primary corpus because it is a high-volume EVM-compatible chain with an active contractdeployment ecosystem, which stresses both compile-side and runtime-side metrics under realistic load. We arbitrarily sample a window of 10,000 blocks, blocks 38,004,930 through 38,014,929 inclusive, containing 3.52M transactions. Three blocks are empty, so the experiments materialize 9,997 block state pre-images. EOF-format bytecode is deferred in the EVM roadmap, so we exclude empty bytecode and bytecode beginning with 0xEF, leaving 19,130 unique compilable bytecodes. To verify that this arbitrary choice is representative, we check the window in two ways. Skeleton reduction rates stay within narrow ranges across ten 1K-block sub-windows on every chain (§3), and the top families span months to years of deployment (Table 3), so the window captures persistent workload structure. Experiments run on an Intel Xeon Platinum 8275CL with sixteen hardware threads and 30 GiB DRAM under Ubuntu 24.04. Both modes are built with rustc 1.91 and LLVM 21.1 and run with sixteen worker threads. Object footprint is the cumulative byte length of the LLVM-emitted .o files per mode. The runtime then links and loads these objects as .so files; we measure the .o stage, identically for both modes. We report per-member runtime variant tables separately as runtime metadata. The variant-table payload totals 5.4 MiB across 9,910 members of multi-member families, with a median table size of 352 bytes per member. The full registry file additionally stores a per-member hash identifier and per-family bookkeeping fields, totaling 9.6 MiB on disk.

Evaluation

The evaluation addresses four research questions. • RQ1: How common are reusable contract families in real EVM workloads? • RQ2: How much does SkelOT reduce compilation cost and artifact size, and where do the savings come from? • RQ3: Does SkelOT preserve correctness and maintain competitive dispatch time? • RQ4: How much native code is needed to cover different shares of execution time? Our primary baseline is Naive (per-hash) AOT, implemented in the same LLVM-based AOT pipeline as SkelOT. Both build on revmc [37], and the baseline is revmc’s own per-hash compilation path, run at the same compiler version, optimization level, and thread count. This choice isolates the design variable, the unit at which compiled code is reused. 8

Table 3. Deployment-time spans for top skeleton families across four chains.

6.1

Family

#members

Span (days)

Clanker ERC-20 [Base] EIP-1167 proxy [ARB] EIP-1167 proxy [Base] EIP-1167 proxy [BSC] EIP-1167 proxy [ETH] ERC-1155 creator [ETH] ERC-20 template [BSC] PancakeSwap LmPool [ARB] PancakeSwap LmPool [Base] PancakeSwap LmPool [BSC] PancakeSwap V3 [ARB] PancakeSwap V3 [Base] PancakeSwap V3 [BSC] Swap helper [ETH] Transparent proxy [ARB] Transparent proxy [ETH] Uniswap V3 Pool [ARB] Uniswap V3 Pool [Base] Uniswap V3 Pool [BSC] Uniswap V3 Pool [ETH]

465 33 342 160 629 167 164 28 212 243 37 402 1,123 127 45 164 116 3,323 405 2,722

99.4 1,690.1 805.8 1,912.7 1,732.7 887.4 57.8 835.4 763.7 973.2 971.4 789.0 1,090.9 670.2 267.1 534.5 1,762.5 752.9 1,142.3 1,355.1

EIP-1167 proxies, transparent proxies, and token or collection templates. This temporal evidence supports the interpretation that the invariant-heavy families observed in the corpora are persistent deployment templates, not contracts that merely appear together in short measurement windows. The same spans also support the classification. Members deployed years apart hold the same values at every invariant position, so the invariant sets have stayed stable. 6.2

RQ2: Compile-Side Savings And Attribution

For the compile-side benchmark, we compile the filtered corpus once under Naive AOT and once under SkelOT from an empty object cache. We measure compile units, wall-clock compile time, object footprint, and fallback count. SkelOT reduces (Figure 5) the number of compile units by 47.5%, shrinks the native-object footprint by 57.4%, and cuts compile time by 2.19×. Fallback is small. A conservative check on where the stripped code region ends rejects 7 skeletons covering 17 members, and one family of 3 members degrades because its members jump to different static targets, so these 20 contracts compile per-hash instead. To separate skeleton sharing from the invariant/variant split, we run a three-mode compile-side ablation over all 817 shared families. Naive AOT compiles each member on its own, and we sum the per-member cost from the full perhash build. SkelOT shares one skeleton artifact and keeps invariant immediates as compile-time constants. All-variant is SkelOT with the classification step skipped, so it still shares the skeleton but routes every eligible PUSH immediate through the per-contract table. Figure 6 shows both parts of the benefit. The Naive bars measure family sharing relative to SkelOT, with compiletime ratios from 5.7× to 29.6× and object-size ratios from 5.4× to 34.7× across buckets. Compiling each family once takes 10,385 s instead of 248,239 s and produces 141.4 MB instead of 3,040 MB, a 23.9× and 21.5× difference over the 9,910 members. The all-variant bars hold skeleton sharing fixed and isolate the constant-classification policy, with compiletime ratios from 1.11× to 1.58× and object-size ratios up to

RQ1: Workload Structure

RQ1 asks whether the invariant/variant split reflects a recurring workload property rather than a Base-specific artifact. Across all four studied chains, multi-member families have the same qualitative shape. Only a small fraction of PUSHimmediate bytes vary across family members. The byteweighted variant share is 5.08% on Base, 4.11% on Ethereum, 3.09% on BSC, and 6.28% on Arbitrum, computed over 825, 1,833, 632, and 173 multi-member families respectively. Correspondingly, 93.72%–96.91% of embedded-constant bytes are invariant and can be baked into the shared skeleton as compile-time constants. The split is decided per position, and 99.52% of PUSH positions on Base are invariant, so the byte-weighted share above is a conservative figure. Larger contracts also vary less. In an average family from the tiny bucket, 12.6% of the PUSH bytes are variant. In the huge bucket, the average is 0.4%. Figure 4 reports the four-chain breakdown. These results show that the invariant/variant split matches the common case in deployed contracts. Most constants stay the same across members, so the per-contract runtime table carries only a small residual. The same family structure is durable over deployment time, so it is not merely a short-window co-occurrence effect. Table 3 reports the five largest multi-member families on each chain and traces sampled members through full chain history using archive RPC. The top Base families span 99–806 days, Ethereum reaches 1,733 days, BSC reaches 1,913 days, and Arbitrum reaches 1,762 days. The long-span families include Uniswap V3 and PancakeSwap V3 pool skeletons,

Cost normalized to Naive AOT

Naive AOT

Units

SkelOT

Compile time

Object footprint

1.25 1.00

19,130

27,758 s

5.05 GB

0.75 10,037

12,701 s

0.50

2.15 GB

0.25 0.00

Naive

SkelOT

Naive

SkelOT

Naive

SkelOT

Figure 5. Compile-side cost. SkelOT cuts units by 47.5%, footprint by 57.4%, and compile time by 2.19×. 9

Compile time (s)

Naive AOT

SkelOT

Table 4. End-to-end replay of all 10K blocks, three modes.

All-variant

100k

Mode

Exec (s)

Family (s)

Other (s)

Compile (s)

Fam. build (s)

10k

Interpreter Naive AOT SkelOT

534.4 459.1 450.6

— 93.8 84.8

— 365.2 365.8

— 27,758.3 12,701.5

— — 1.1

1k 100

Object size (MB)

10

1,842 invoked family members after 2 warmup and 5 measured rounds per mode. For each member we run a paired 𝑡-test on the per-round differences and correct for testing all 1,842 hypotheses at once with the Benjamini-Hochberg procedure at 𝑞 = 0.05 [4]. For each measured family member, speedup is the mean Naive AOT dispatch time divided by the mean SkelOT dispatch time over the measured rounds, so the metric compares per-member dispatch time during block replay. To confirm execution-level correctness, we replay every non-empty block of the corpus once under Naive AOT and once under SkelOT dispatch and compare each transaction’s success/failure status, gas consumption, and output between the two modes. Across all 9,997 non-empty blocks (3.52M transactions), SkelOT dispatch produces byte-identical execution results to Naive AOT, with zero mismatches. The replay also covers the hardest case for table binding, where a call into one family member runs inside a call into another member of the same family and the two frames must read different tables. The corpus contains 3.4M such nested family frames, and all replay with zero mismatches. Runtime timing shows no aggregate regression against Naive AOT. In the paired per-contract comparison, SkelOT has a 1.31× median per-member speedup, with a short slowdown tail. The median dispatch time falls from 23.08 𝜇s under Naive AOT to 18.12 𝜇s. After the Benjamini-Hochberg correction, 1,171 contracts are faster, 644 equivalent, and 27 slower, with a bootstrap 95% confidence interval on the median speedup of [1.286, 1.345]. Figure 7 shows the full per-pair distribution. Most of the distribution lies near or above the no-change boundary, with larger gains in the right tail, and every sampled block has a block-internal median above 1.0×. We attribute this 1.31× median speedup to instruction-fetch locality under cross-instance code-region reuse. The counter evidence in §7.1 supports this instruction-cache explanation. Table 4 gives the broader picture, the wall-clock cost of replaying all 10,000 blocks under each mode, including compilation. Compilation accounts for most of the difference between Naive AOT and SkelOT. Execution time also favors SkelOT, and the gain is isolated. Family-member frames drop from 93.8 s to 84.8 s, while state access, transaction assembly, and precompiled-contract execution stay flat, since compilation strategy touches none of them. Family-table construction adds 1.1 s, negligible next to either compile cost.

1,000 100 10 1

small 102 fam.

medium 298 fam.

large 256 fam.

huge 161 fam.

Figure 6. Compile-side ablation over all 817 shared families. Naive AOT, SkelOT, and All-variant compile the same families under identical settings. Compile time is the sum of per-unit wall clock.

1.21×. The invariant/variant split contributes a measurable compile-side benefit beyond skeleton sharing alone. Both gains grow with contract size. The huge bucket alone holds about half of the members. Naive AOT would spend most of its 3,040 MB on this bucket, while SkelOT serves it with 73 MB. Larger contracts have more instructions per skeleton, so skipping one recompilation saves more work. They also hold more invariant constants, which keeps more values folded into the shared code. The all-variant regression has a concrete per-PUSH mechanism. Families with more PUSH positions forced to be variants produce larger native artifacts. Removing compile-time constants creates extra native code, rather than only reducing the optimizer’s ability to simplify existing code. On Arbitrum, our least favorable corpus, the same effect holds at smaller scale, cutting compilation units by 23.1%, footprint by 26.2%, and compile time by 1.29×. New contracts keep arriving after the initial corpus is compiled. Over the later windows, every 1,000 blocks bring 764 new unique bytecodes on average. Under per-hash reuse, every one of them is a fresh compilation candidate. Under SkelOT, at least 44.6% join an existing family and reuse its artifact at once, so the candidate pool shrinks by nearly half. 6.3

RQ3: Correctness And Runtime Behavior

Because paired timing is expensive, we evaluate runtime behavior by replaying every hundredth block of the corpus under both Naive AOT and SkelOT dispatch. Across 99 sampled blocks, we record paired per-member timings for 10

p95 2.60x

10k

200 150

Naive, matching SkelOT

1k 300 100 10

3,000

0.5×

1×

1.5× 2×

3×

Footprint (MB)

50 5×

Per-pair speedup ratio (n = 1,842)

1,000 300 100 30 10 3

Figure 7. Per-pair runtime speedup. Median speedup over 1,842 family members is 1.31×, with 1,171 faster pairs, 644 equivalent pairs, and 27 slower pairs.

6.4

Naive, own hot set

30

100

0

SkelOT

3k

Artifacts

p75 1.68x

250

p50 1.31x

p25 1.07x

Family-member pairs

300

50%

75% 90% 95% 99% Execution-time coverage target

Figure 8. Compile budget under three policies. Each policy compiles hottest first until it reaches the coverage target. Naive, own hot set ranks individual contracts by their own execution time. Naive, matching SkelOT compiles every member of the units SkelOT selects, one artifact per contract.

RQ4: Budget Efficiency

Budget efficiency asks how much native code an AOT system must materialize and store to cover a target share of execution time. For each target, we take the smallest executionweighted SkelOT skeleton prefix that reaches the target. That prefix defines both the shared artifacts SkelOT compiles and the concrete code hashes those artifacts serve. To compare against Naive AOT on the same deployment scope, we count how many native artifacts Naive AOT would need under its one-artifact-per-code-hash policy, and compare the cumulative artifact footprint. We also report a more conservative reading, in which Naive AOT ranks contracts by their own execution time and compiles only its own hot set. Figure 8 shows that SkelOT serves the same deployment scope with substantially fewer native artifacts and lower footprint across the measured coverage range. The hottest part of the workload shows the largest gap. At 75% executiontime coverage, 35 SkelOT artifacts cover 3,933 distinct code hashes, while Naive AOT needs one artifact per hash. This is a 112× reduction in artifact count and a 437× reduction in footprint. The advantage remains large deeper in the ranking, with count and footprint compression of 33.9× and 168× at 90% coverage, and 6.2× and 62× at 99%. Under the conservative reading, Naive AOT needs 1.9× the artifacts and 5.9× the footprint at 75%, and 2.5× and 13.9× at 95%. Compression decreases as the coverage target moves into the long tail, where more singleton families enter the prefix and each additional SkelOT artifact serves fewer additional deployments. Even there, footprint compression stays above artifact-count compression, which indicates that the hot shareable families near the head are also expensive for Naive AOT to compile separately. SkelOT thus spends compilation budget where it matters first, on highly executed

templates that would otherwise produce many large perhash artifacts. The 99% point in Figure 8 is slightly below full coverage of the compile corpus because a small set of compile units did not execute in the measured trace. The conservative reading instead widens, from 1.6× at 50% to 2.5× at 99% in artifacts, because contracts without a family dominate the head of the workload while families sit in the mid and long tail.

7

Discussion

Family-level reuse changes the operating point of EVM AOT systems, not only their compile budget. This section interprets the measured gains, the workloads that support them, and the boundary that still limits transfer. 7.1

Implications For EVM AOT Systems

A code hash remains the right way to name a deployed contract, but it is too fine-grained as the sole unit of AOT compilation. On the Base corpus, compiling at family granularity cuts compilation units, native-artifact footprint, and compile time while preserving per-contract dispatch identity. The budget result sharpens the same point. A finite AOT budget should buy coverage over deployed behavior, not redundant native objects for members of the same template family. Runtime measurements on the same corpus are also favorable, with no measured penalty and a 1.31× median speedup over Naive AOT. 11

Table 5. Top-level TMA slot attribution for SkelOT versus Naive AOT. Values are SkelOT minus Naive; negative values are savings.

When Naive AOT compiles each contract on its own, a block that runs many similar contracts keeps jumping to different native binaries. Each jump can make the CPU bring in another piece of code before execution can continue. Those waits are the main source of the runtime gap. SkelOT instead gives all members of the same family one shared binary. Once that binary is hot, later calls to other family members reuse the same code in the instruction cache. The hardware counters match this explanation. We use the CPU’s top-level slot accounting because its four buckets add up cleanly. Almost the entire reduction comes from slots where the CPU is waiting for code to arrive at the front of the pipeline. The other buckets do not explain the speedup. Backend waiting increases slightly, wasted speculative work falls a little, and useful retired work is nearly unchanged (Table 5). More detailed counters point to the same cause. L1 instruction-cache stalls fall by about 1.0K and 0.25K cycles per call, and L2 code-read misses also fall. Counters for frontend decode bandwidth move by only a few cycles, and offcore code-read cycles are zero. Our machine does not expose the CPU events needed to split the frontend bucket cleanly into smaller additive pieces, so we use these detailed counters only as supporting evidence for an instruction-cache explanation. We read counters per call over the same fixed block order in both modes, repeat each measurement several times after warmup, and report averaged paired deltas in Table 5. We do not pin CPU frequency or ASLR; to remove that noise, we use repeated paired measurements. The same mechanism predicts where SkelOT will not help, and the data agrees. A shared binary only stays in cache when it is hit often enough; a family with only a handful of members generates too few hits to keep its code warm. SkelOT also pays a small fixed cost per call. It loads variant PUSH constants from a per-member data table rather than inlining them as immediate operands. When the family is large, this cost disappears against the cache savings; when the family is small, the cost is real and the savings never materialize. In our measurements on the Base corpus, the runtime ratio rises sharply with family size. Families with fewer than one hundred members stay essentially flat near 1.07×, and only families of at least one hundred members produce the 1.58× median that drives the headline. The contracts on which SkelOT actually slows things down sit almost entirely in the small-family stratum, exactly where the mechanism predicts no benefit. A practical AOT system should therefore enable family compilation only when a family is large enough to carry its weight. Below that threshold, plain per-contract compilation is the safer choice; above it, the locality win compounds with every additional member. The exact cutoff depends on the calling profile and on how much per-call work the host EVM does outside the JIT’d code, but our bucket-level evidence puts a sensible default in the range of roughly one hundred members. With this admission rule in place, SkelOT

TMA bucket

Delta (slots/call)

Interpretation

Frontend bound Backend bound Bad speculation Retiring

−8,149.7 +165.0 −141.4 −10.3

fewer frontend undersupply slots offset: more backend-bound slots fewer wasted speculative slots nearly unchanged useful-work slots

Total slots

−8,136.5

net slot reduction

keeps its wins on the long-tail high-reuse families that drive the headline number and avoids the regressions on small families. 7.2

When SkelOT Helps

SkelOT helps most when deployments reuse the same bytecode structure and differ only in a small set of immediates. That is the common case in our measurements, especially for large contracts. When a workload has less repeated structure, SkelOT exposes less sharing and falls back toward fine-grained compilation. The current boundary is intentionally precise. Differences in structure, control flow, or opcode choice separate families. Future abstractions can admit more cases while keeping the same rule. Compiled code should be shared only when member-specific behavior is explicitly represented. Similarly, SkelOT does not normalize PUSH widths. Such normalization can raise reuse but is not semantics-free in the EVM. PUSH width affects instruction boundaries and jump targets. Metadata and byte length are observable through code-observation instructions, so SkelOT strips metadata and trailing zero bytes only while preserving the original code view at execution. A safe normalization layer for opcode encodings would need the same separation between the reuse key and the EVM-visible bytecode. SkelOT leaves that larger design point to follow-on work and keeps the present reuse predicate byte-level and auditable. Families benefit in different ways. In some, a few hot members [1, 36] carry most of the execution time. Later members with the same structure and different constants get native code at once, with no new compilation. In others, execution time is spread over many members [15, 16, 24, 47]. One compilation and one artifact then serve the whole family. In every case, the members differ only in constructor parameters that deployment writes into the code. 7.3

External Validity

The evidence has two scopes. Workload-structure measurements use one 10K-block window from each of Base, Ethereum, BSC, and Arbitrum, and show the same invariantheavy family structure across chains. Compile-side savings, all-block correctness replay, timing, and hardware-counter 12

attribution use Base as the primary corpus on a revmcbased LLVM AOT prototype under a fixed version of the EVM rules. A smaller Arbitrum compile run confirms the compile-side direction (§6.2). Within the Base window, family structure is stable across sub-windows and over deployment time (§6.1). Prior studies of factory deployment, proxy contracts, composition, on-chain dependencies, and bytecode clones report that family-structured deployment extends well beyond this window [20, 23, 43, 47, 52]. Broader replication should quantify how compile-side ratios, runtime attribution, and budget compression transfer across fork rules, protocol mixes, chains, and compiler backends. The design requirement is narrower. A backend must keep invariant PUSH values as compile-time constants, lower variant PUSH values as table loads, and bind the current member’s table during dispatch. 7.4

Current admission

All-variant admission

shared artifact: token0 := load table[0] token1 := load table[1] fee := 500 baked

shared artifact: token0 := load table[0] token1 := load table[1] fee := load table[2]

contract A table: slot0= 0x10 ; slot1= 0xaa slot2= 500 ✓

contract A table: slot0= 0x10 ; slot1= 0xaa slot2= 500 ✓

contract C table: slot0= 0x30 ; slot1= 0xcc slot2= 3000 ✗

contract C table: slot0= 0x30 ; slot1= 0xcc slot2= 3000 ✓

Figure 9. Admission trade-off. Split rejects a candidate changing a baked invariant (left); all-variant admission accepts it by moving the position into the member table (right).

value. Over the shared family artifacts, where the two designs differ, all-variant takes 22.2% longer to compile, produces 19.4% larger artifacts, runs family member frames 18.2% slower, and grows the registry file from 9.6 MiB to 526 MiB. The current split therefore favors deployments that can amortize a stricter admission boundary across many same-skeleton members, whereas unconditional admission favors deployments that prioritize join-on-arrival simplicity over per-family efficiency. Two possible improvements look promising. One is to recompute families periodically, regrouping the contracts rejected since the last pass. This needs a policy for retiring live artifacts. The other is to declare which constants vary at the source level. The classification is then known when the first member arrives. Deduplicating the artifact store is another promising optimization. Content-defined chunking [34, 50] splits binaries at boundaries derived from the content, so similar binaries share most chunks on disk. Our prototype does not use it, but it could shrink the on-disk store further. Versioned families are a possible middle ground beyond the current design. They could recover some currently rejected members without pushing every position into the runtime table, for example by compiling a second artifact when a once-invariant position begins to vary. That extra flexibility would need a policy for when to fork a family, how to route old and new members, and whether to retire or migrate existing artifacts. An arrival-order replay could compare conservative admission, unconditional admission, and versioned families in terms of acceptance rate, artifact churn, compile cost, and runtime effect.

Admission Policy Trade-Offs

We run classification and table construction in a batch phase, which matches how AOT works. Family classification only adds a comparison over the compile batch. We checked whether our way of choosing families is stable over time. We split the window in half, and families chosen from the first half alone cover 80.2% of execution time in the second half, while the gold-standard answer is 80.8%. The classification also goes stale slowly, since the deployment spans in §6.1 show top families staying on chain for months to years. This paper evaluates batch family reuse, but the same design choice also appears at deployment time. A deployed system needs an online rule for newly arriving contracts: either preserve the compiled artifact exactly and admit only members that match its descriptor, or widen the descriptor so that more same-skeleton contracts can join. SkelOT’s current rule is deliberately conservative. A new contract joins an existing compiled family only when its skeleton matches and every invariant position preserves the family’s expected value. When a candidate changes a current invariant, SkelOT rejects that contract from the family rather than rewriting the shared artifact online. This boundary keeps admission monotone, keeps the cache key stable after compilation, and avoids artifact invalidation or migration. The cost is that some same-skeleton contracts remain outside the already compiled family. A more permissive endpoint is unconditional admission through an all-variant classification. Under this rule, every eligible PUSH value is lowered as a runtime table load, so any contract with the same skeleton can join an existing family by binding its own table at deployment; only diverging static-jump targets still block admission (§4.6). Figure 9 sketches the difference in a listing-style example aligned with Figure 1. Moving an otherwise stable immediate from the native artifact into the table removes a compile-time constant, which can prevent constant folding, comparison specialization, and other backend simplifications around that

8

Related Work

We summarize SkelOT and relevant studies in Table 6. EVM execution and AOT compilation. Ahead-of-time compilation of EVM bytecode to native code is an established engineering direction. The revmc project lowers EVM bytecode to LLVM IR and produces one native artifact per code hash [37]; evmone offers an optional AOT path [21]; 13

Table 6. Summary and comparison of related works. Works

AOT/native engines [21, 37, 53] Workload studies [14, 20, 43, 47, 52] Clone detection [19, 27, 28, 32, 48] Reusable compilation [17, 18, 31, 38, 39] Below-translation reuse [3, 8, 29, 41, 45] SkelOT Legend:

supplied;

Capability supplied

Refs.

partial;

Role for SkelOT

Native exec.

Family signal

Reuse

(AOT/native) (orthogonal) (not EVM AOT) (native blocks)

(per hash) (factories/proxies) (clusters) (no families) (post-translation)

(artifact reuse) (no compiler) (not units) (context reuse) (dedup/cache)

execution substrate opportunity evidence boundary warning conceptual ancestry contrast point

(EVM AOT)

(skeleton)

(family reuse)

family-aware AOT reuse

not supplied.

and DTVM explores a hybrid lazy-JIT architecture on a Wasm/dMIR substrate [53]. SkelOT inherits the LLVM-based AOT execution model but changes the compilation-reuse unit from the full bytecode or code hash to a contract-family skeleton, as detailed in §4. DTVM’s design axis is orthogonal to the reuse granularity we study.

reusing compiled functions, optimized IR, or specialization results across runs and virtual-machine instances [31, 38, 39]. These systems target dynamic runtimes, where reused code can be guarded, replay-validated, or deoptimized when the context changes. SkelOT works in a stricter EVM AOT setting. Shared native code must preserve per-contract behavior and gas accounting before execution. We therefore place reuse at the contract-family skeleton and share code only through an explicit invariant/variant split.

Empirical structure of EVM workloads. Production EVM contracts exhibit substantial structural reuse. Factory-driven deployment accounts for more than 90 % of contracts created since 2020 [47]; subcontracts often cluster around recurring imported patterns [43]; and code-sharing proxies and clones are widely used [20, 52]. These studies show that many deployed contracts come from recurring templates, motivating family-level optimization, but they stop at workload characterization rather than compilation reuse. DiAngelo et al. [14] are closest to our abstraction. They use bytecode skeletons to avoid repeated large-scale weakness analysis of bytecodes that differ only outside the chosen abstraction. SkelOT builds on the same observation but changes the boundary’s role. For AOT compilation, skeleton equality alone is insufficient. The compiler must decide which constants can remain baked into shared native code and which must be externalized into per-contract tables. SkelOT turns skeletons from a static-analysis boundary into an AOT reuse boundary through an invariant/variant constant split and an EVM-specific correctness condition.

9

Conclusion

We show that code hash is too fine-grained as the reuse unit for EVM AOT compilation. Modern EVM workloads contain contract families that share instruction skeletons while differing only in a few embedded constants. We design SkelOT to exploit this structure by compiling one artifact per family, baking invariant constants into shared code, and loading only variants from per-contract tables. Across four EVM-compatible chains, we find that 23.1– 47.6% of unique compilable bytecodes collapse into shared skeletons. On Base, SkelOT cuts compilation units by 47.5%, artifact footprint by 57.4%, and compile time by 2.19×, while preserving Naive-AOT behavior and achieving a 1.31× median runtime speedup.

References

Smart-contract clone detection. Ethereum clone detection is mature but orthogonal to compilation. The representative three approaches (i.e., birthmark-based, semantic sketch, structural embedding) identify similar contracts or clone clusters, but they do not define executable reuse units [19, 27, 28, 32, 48]. For AOT reuse under EVM gas equivalence, similarity is not enough. The system must know which constants can be shared safely and which must remain contract-specific. SkelOT supplies this missing correctness layer through its invariant/variant constant split.

[1] Hayden Adams, Noah Zinsmeister, Moody Salem, River Keefer, and Dan Robinson. 2021. Uniswap v3 Core. Whitepaper. https://app. uniswap.org/whitepaper-v3.pdf; Accessed 2026-04-28. [2] Jean-Philippe Aumasson and Daniel J Bernstein. 2012. SipHash: a fast short-input PRF. In International Conference on Cryptology in India (INDOCRYPT). Springer, 489–508. [3] Fabrice Bellard. 2005. QEMU, a Fast and Portable Dynamic Translator. In Proceedings of the USENIX Annual Technical Conference (ATC), FREENIX Track. USENIX Association, 41–46. [4] Yoav Benjamini and Yosef Hochberg. 1995. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 57, 1 (1995), 289–300. [5] Alex Beregszaszi, Paweł Bylica, and Andrei Maiboroda. 2021. EIP3540: EOF - EVM Object Format v1. Ethereum Improvement Proposal. https://eips.ethereum.org/EIPS/eip-3540; Accessed 2026-09-01. [6] BNB Chain. 2026. BNB Smart Chain: High Performance DeFi Hub. BNB Chain Documentation. https://docs.bnbchain.org/bnb-smart-

Reusable compilation and JIT-code reuse. Reusable compilation is conceptually related to SkelOT, but it usually assumes a different correctness model. Partial evaluation specializes a program with respect to known inputs [17, 18]. Recent JIT-reuse systems reduce repeated compilation by 14

chain/overview/; Accessed 2026-05-06. [7] Lee Bousfield, Rachel Bousfield, Chris Buckland, Ben Burgess, Joshua Colvin, Edward W. Felten, Steven Goldfeder, Daniel Goldman, Braden Huddleston, Harry Kalodner, Frederico Arnaud Lacs, Harry Ng, Aman Sanghi, Tristan Wilson, Valeria Yermakova, and Tsahi Zidenberg. 2022. Arbitrum Nitro: A Second-Generation Optimistic Rollup. Arbitrum Documentation. https://docs.arbitrum.io/nitro-whitepaper.pdf; Accessed 2026-05-06. [8] Derek Bruening, Timothy Garnett, and Saman Amarasinghe. 2003. An Infrastructure for Adaptive Dynamic Optimization. In Proceedings of the International Symposium on Code Generation and Optimization (CGO). 265–275. [9] Bytecode Alliance. 2026. Cache Configuration of wasmtime. https: //docs.wasmtime.dev/cli-cache.html. Accessed 2026-04-24. [10] Franck Cassez, Joanne Fuller, Milad K Ghale, David J Pearce, and Horacio MA Quiles. 2023. Formal and executable semantics of the ethereum virtual machine in dafny. In International Symposium on Formal Methods (FM). Springer, 571–583. [11] Clanker. 2025. ClankerToken v3.1.0 and v4.0.0. Clanker Documentation. https://clanker.gitbook.io/clanker-documentation/references/ core-contracts/clankertoken-v3.1.0-and-v4.0.0; Accessed 2026-05-06. [12] Coinbase. 2023. Base: An Ethereum L2 Built on the OP Stack. https: //docs.base.org/. Mainnet launched 2023-08-09; Accessed 2026-04-28. [13] Ron Cytron, Jeanne Ferrante, Barry K. Rosen, Mark N. Wegman, and F. Kenneth Zadeck. 1991. Efficiently Computing Static Single Assignment Form and the Control Dependence Graph. ACM Transactions on Programming Languages and Systems (TOPLAS) 13, 4 (1991), 451–490. [14] Monika di Angelo, Thomas Durieux, João F. Ferreira, and Gernot Salzer. 2024. Evolution of Automated Weakness Detection in Ethereum Bytecode: A Comprehensive Study. Empirical Software Engineering (ESE) 29, 2 (2024), 41. [15] Monika di Angelo and Gernot Salzer. 2019. Mayflies, Breeders, and Busy Bees in Ethereum: Smart Contracts Over Time. In Proceedings of the 3rd ACM Workshop on Blockchains, Cryptocurrencies and Contracts (BCC). 1–10. [16] Monika di Angelo and Gernot Salzer. 2020. Characteristics of Wallet Contracts on Ethereum. In 2nd Conference on Blockchain Research & Applications for Innovative Networks and Services (BRAINS). 232–239. [17] Chris Fallin and Maxwell Bernstein. 2025. Partial Evaluation, WholeProgram Compilation. In Proceedings of the ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI). [18] Yoshihiko Futamura. 1999. Partial Evaluation of Computation Process— An Approach to a Compiler-Compiler. Higher-Order and Symbolic Computation 12, 4 (1999), 381–391. Reprint of the original 1971 paper in Systems, Computers, Controls 2(5):45–50.. [19] Zhipeng Gao, Vinoj Jayasundara, Lingxiao Jiang, Xin Xia, David Lo, and John Grundy. 2019. SmartEmbed: A Tool for Clone and Bug Detection in Smart Contracts through Structural Code Embedding. In IEEE International Conference on Software Maintenance and Evolution (ICSME). [20] Ningyu He, Lei Wu, Haoyu Wang, Yao Guo, and Xuxian Jiang. 2020. Characterizing Code Clones in the Ethereum Smart Contract Ecosystem. In Financial Cryptography and Data Security (FC). 654–675. [21] Ipsilon. 2024. evmone: Fast Ethereum Virtual Machine Implementation. https://github.com/ipsilon/evmone. Apache-2.0 licensed open-source software; successor to the Ewasm/ethereum/evmone repository; accessed 2026-04-22. [22] Michael R. Jantz and Prasad A. Kulkarni. 2013. Exploring Single and Multilevel JIT Compilation Policy for Modern Machines. ACM Transactions on Architecture and Code Optimization (TACO) 10, 4 (2013), 40:1–40:29. [23] Xiangfu Jin, Zihao Liu, and Martin Monperrus. 2025. OnChain Analysis of Smart Contract Dependency Risks on Ethereum. arXiv:2503.19548 [cs.SE] arXiv:2503.19548.

[24] Faizan Khan, Istvan David, Daniel Varro, and Shane McIntosh. 2022. Code Cloning in Smart Contracts on the Ethereum Platform: An Extended Replication Study. IEEE Transactions on Software Engineering (TSE) 49, 4 (2022), 2006–2019. [25] Chris Lattner and Vikram Adve. 2004. LLVM: A Compilation Framework for Lifelong Program Analysis & Transformation. In Proceedings of the International Symposium on Code Generation and Optimization (CGO). 75–86. [26] Rujia Li, Jingyuan Ding, Qin Wang, Keting Jia, Haibin Zhang, and Sisi Duan. 2025. Does finality gadget finalize your block? A case study of Binance consensus. In 34th USENIX Security Symposium (USENIX Sec). 4109–4125. [27] Han Liu, Zhiqiang Yang, Yu Jiang, Wenqi Zhao, and Jiaguang Sun. 2019. Enabling Clone Detection For Ethereum Via Smart Contract Birthmarks. In IEEE/ACM International Conference on Program Comprehension (ICPC). 105–115. [28] Han Liu, Zhiqiang Yang, Chao Liu, Yu Jiang, Wenqi Zhao, and Jiaguang Sun. 2018. EClone: Detect Semantic Clones in Ethereum via Symbolic Transaction Sketch. In Proceedings of the ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE) — Demo/Tool Track. [29] LLVM Project. 2026. lld: The LLVM Linker. https://lld.llvm.org/. Linker documentation including identical code folding (ICF); Accessed 202604-28. [30] Manifold. 2025. Manifold Creator. Manifold Documentation. https://docs.manifold.xyz/v/manifold-for-developers/manifoldcreator-architecture/overview; Accessed 2026-05-06. [31] Meetesh Kalpesh Mehta, Sebastián Kryński, Hugo Musso Gualandi, Manas Thakur, and Jan Vitek. 2023. Reusing Just-in-Time Compiled Code. Proceedings of the ACM on Programming Languages (PACMPL) 7, OOPSLA2 (2023), 1176–1197. [32] Ran Mo, Haopeng Song, Wei Ding, and Chaochao Wu. 2025. Code Cloning in Solidity Smart Contracts: Prevalence, Evolution, and Impact on Development. In Proceedings of the IEEE/ACM International Conference on Software Engineering (ICSE). 3060–3071. [33] Peter Murray and Nate Welch. 2018. EIP-1167: Minimal Proxy Contract. Ethereum Improvement Proposal. https://eips.ethereum.org/EIPS/eip1167; Accessed 2026-04-28. [34] Athicha Muthitacharoen, Benjie Chen, and David Mazières. 2001. A Low-Bandwidth Network File System. In Proceedings of the ACM Symposium on Operating Systems Principles (SOSP). [35] OpenZeppelin. 2026. Proxy. OpenZeppelin Contracts Documentation. https://docs.openzeppelin.com/contracts/5.x/api/proxy; Accessed 2026-05-06. [36] PancakeSwap. 2023. PancakeSwap V3 Contracts. GitHub repository. https://github.com/pancakeswap/pancake-v3-contracts; Accessed 2026-05-06. [37] Paradigm. 2024. revmc: JIT and AOT Compiler for the Ethereum Virtual Machine. https://github.com/paradigmxyz/revmc. [38] Andrej Pečimúth, David Leopoldseder, and Petr Tůma. 2024. An Analysis of Compiled Code Reusability in Dynamic Compilation. In Proceedings of the ACM SIGPLAN International Workshop on Virtual Machines and Intermediate Languages (VMIL). [39] Andrej Pečimúth, David Leopoldseder, and Petr Tůma. 2025. Reusing Highly Optimized IR in Dynamic Compilation. In Proceedings of the 39th European Conference on Object-Oriented Programming (ECOOP). [40] Witek Radomski, Andrew Cooke, Philippe Castonguay, James Therien, Eric Binet, and Ron Evans. 2018. EIP-1155: Multi Token Standard. Ethereum Improvement Proposal. https://eips.ethereum.org/EIPS/eip1155; Accessed 2026-05-06. [41] Vijay Janapa Reddi, Dan Connors, Robert Cohn, and Michael D. Smith. 2007. Persistent Code Caching: Exploiting Code Reuse Across Executions and Applications. In Proceedings of the International Symposium on Code Generation and Optimization (CGO). 74–88.

15

How Far Are We? Proceedings of the ACM on Software Engineering (ASE/FSE) 2 (2025), 1249–1269. [49] Gavin Wood. 2014. Ethereum: A Secure Decentralised Generalised Transaction Ledger. Ethereum Project Yellow Paper. https://ethereum. github.io/yellowpaper/paper.pdf; Accessed 2026-04-28. [50] Wen Xia, Yukun Zhou, Hong Jiang, Dan Feng, Yu Hua, Yuchong Hu, Qing Liu, and Yucheng Zhang. 2016. FastCDC: A Fast and Efficient Content-Defined Chunking Approach for Data Deduplication. In Proceedings of the USENIX Annual Technical Conference (ATC). [51] Sipeng Xie, Qianhong Wu, Minghang Li, Qiyuan Gao, Bo Qin, and Qin Wang. 2026. MHOT: Height-Optimized Authenticated Data Structure for Blockchain State Commitment. arXiv preprint arXiv:2606.11736; USENIX Security 2026, to appear (2026). [52] Mengya Zhang, Preksha Shukla, Wuqi Zhang, Zhuo Zhang, Pranav Agrawal, Zhiqiang Lin, Xiangyu Zhang, and Xiaokuan Zhang. 2025. An Empirical Study of Proxy Contracts at the Ethereum Ecosystem Scale. In Proceedings of the IEEE/ACM International Conference on Software Engineering (ICSE). [53] Wei Zhou, Xiong Xu, Changzheng Wei, Ying Yan, Wei Tang, Zhihao Chen, Xuebing Huang, et al. 2025. DTVM: Revolutionizing Smart Contract Execution with Determinism and Compatibility. arXiv:2504.16552 [cs.DC] http://arxiv.org/abs/2504.16552v2.

[42] Solidity Team. 2026. Solidity Documentation. Online documentation. https://docs.soliditylang.org; Accessed 2026-08-24. [43] Kairan Sun, Zhengzi Xu, Chengwei Liu, Kaixuan Li, and Yang Liu. 2023. Demystifying the Composition and Code Reuse in Solidity Smart Contracts. In Proceedings of the ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). [44] Leszek Swirski. 2019. Code Caching for JavaScript Developers. https: //v8.dev/blog/code-caching-for-devs. Accessed 2026-04-24. [45] Sriraman Tallam, Cary Coutant, Ian Lance Taylor, Xinliang David Li, and Chris Demetriou. 2010. Safe ICF: Pointer Safe and Unwinding Aware Identical Code Folding in the Gold Linker. GCC Developers’ Summit 2010. https://research.google/pubs/safe-icf-pointer-safe-andunwinding-aware-identical-code-folding-in-gold/; Accessed 2026-0428. [46] Fabian Vogelsteller and Vitalik Buterin. 2015. EIP-20: Token Standard. Ethereum Improvement Proposal. https://eips.ethereum.org/EIPS/eip20; Accessed 2026-04-28. [47] Ziyue Wang, Zongwen Shen, Lei Chen, Wei Song, Jidong Ge, LiGuo Huang, and Bin Luo. 2026. Empirical Analysis of Smart Contract Factories on EVM-compatible Chains. ACM Transactions on Software Engineering and Methodology (TOSEM) (2026), 3787216. [48] Zuobin Wang, Zhiyuan Wan, Yujing Chen, Yun Zhang, David Lo, Difan Xie, and Xiaohu Yang. 2025. Clone Detection for Smart Contracts:

16

Record · ID 1028589 · SHA-256 18379504437fc346
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.