ConceptioArchivearXiv CS
arXiv CSopen access

The Audit Gap in Blockchain Security: A Four-Year Empirical Study of Public Audit Findings and Real-World Exploit Incidents

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

The Audit Gap in Blockchain Security: A Four-Year Empirical Study of Public Audit Findings and Real-World Exploit Incidents

arXiv:2606.15465v1 [cs.CR] 13 Jun 2026

Stefan Beyer1 1 Oak Security

[email protected]

April 2026 Abstract This paper presents an empirical analysis of the Web3 security landscape over the fouryear and three-month period from 1 January 2022 to 27 March 2026. The dataset combines 23,818 public audit findings produced by 22 independent security firms with 218 real-world exploit incidents documented by rekt.news, representing aggregate losses of approximately US$7.76 billion. We report three central findings. First, the distribution of audit findings—by severity, category, and technology stack—is substantially stable across the observation window, with the Critical-plus-High share remaining within a 15– 17% band in every complete year. Second, the categorical distribution of realised exploit losses does not correspond to the categorical distribution of audit findings: private-key compromise, phishing, and social-engineering vectors account for approximately 49.6% of cumulative losses yet represent a negligible share of published audit findings. Third, realised losses exhibit extreme concentration: the eight largest incidents account for 50.6% of cumulative dollar losses and the twenty largest for 71.4%, a distributional shape inconsistent with Gaussian assumptions. Throughout, we adopt the analytical convention that audit outputs and exploit outputs describe different populations and present the two datasets in parallel rather than as directly comparable samples.

Keywords: Web3 security; smart-contract audits; exploit incidents; empirical software engineering; operational security; heavy-tailed loss distributions.

1 Introduction Empirical study of the Web3 security domain is hindered by the distributed and heterogeneous nature of its evidentiary base. Audit reports are produced by dozens of firms, in varying formats, published through both repository-based and web-native channels, and they differ substantively in how findings are classified. Exploit data is similarly fragmented, tracked primarily by a small number of journalistic and research publications, and subject to disagreement about what should count as a distinct incident, how to value non-liquid assets, and how to attribute root cause. This paper attempts a consolidated view across both sides of that evidentiary base over a four-year and three-month observation window. The motivating question is simple: across

The Audit Gap in Blockchain Security

2

all public audit output and all documented real-world incidents, do the patterns that auditors surface correspond to the patterns that result in realised losses, and if not, in what ways do they differ? The analysis is conducted at the aggregate level; no single audit report or exploit is examined in isolation. Related work. Prior systematisations of Web3 security have largely been taxonomic or mechanism-focused. Atzei et al. [1] provide an early survey of attacks on Ethereum smart contracts, organising vulnerabilities by language- and blockchain-level cause. Werner et al. [5] survey the DeFi protocol stack and its security assumptions, and Zhou et al. [4] catalogue DeFi attack techniques and incidents, reporting—consistent with the present study—that price-oracle manipulation and permissionless-interaction attacks are among the most frequent incident types yet receive disproportionately little academic attention. A second line of work bears directly on the gap this paper measures: the weak coupling between vulnerabilities that are detectable and those that are exploited. Perez and Livshits [2] show that of 23,327 contracts flagged as vulnerable by six analysis tools, fewer than 2% were ever exploited, and that the funds genuinely at risk were concentrated in a small number of contracts. Chaliasos et al. [3] evaluate five widely used security tools against 127 high-impact real-world attacks and survey practitioners, finding that the tools would have prevented only a minority of the attacks and that the most damaging incidents fall outside their detection scope. Domain-specific systematisations of cross-chain bridge hacks [6, 7] likewise attribute the largest bridge losses to key- and signature-level compromise rather than to contract-logic defects. These works either evaluate tools or enumerate mechanisms within a single vulnerability class or protocol type. The present study is complementary and deliberately aggregate: it measures the joint empirical distribution of what auditors report and what attackers realise, across the full public record of both audit findings and incidents over a multi-year window, and quantifies the divergence between the two. Contributions. (i) We assemble and classify the largest cross-firm corpus of public audit findings analysed to date (23,818 findings, 22 firms) alongside a quality-filtered incident dataset (218 events, US$7.76 B). (ii) We document that the audit-finding distribution is stable across the window while the realised-loss distribution is not, and that the two distributions are categorically misaligned. (iii) We characterise the realised-loss distribution as heavy-tailed and concentrated by chain and protocol type, with implications for risk provisioning. The paper is organised as follows. Section 2 documents the data sources and methodology. Section 3 reports the overall shape of the audit-findings dataset. Section 4 reports the temporal evolution of audit output. Section 5 presents the incident dataset. Section 6 contrasts the two datasets. Section 7 examines human-vector attacks. Section 8 reports supplementary findings. Section 9 concludes.

2 Data and methodology 2.1 Audit-findings dataset

The audit-findings dataset draws on two categories of source. The first is repository-based publication by firms that maintain public GitHub archives of PDF audit reports. The second is

The Audit Gap in Blockchain Security

3

web-native publication through structured online channels, including vendor APIs, contentmanagement REST endpoints, and certificate or report portals. In total, the public output of 22 independent security firms is included, spanning both publication modes. Individual firms are not identified in this paper: the analysis is conducted strictly at the aggregate level, and no firm-level results are reported. For repository-based sources, the date on which a report was first committed to the repository is used as the report’s publication date, obtained via git log --diff-filter=A. This avoids the bias introduced when the existence of a report in a repository is conflated with its first publication; reports committed prior to the observation window are excluded even if they remain in the repository. 2.2 Classification

Findings are assigned to a flat taxonomy of approximately twenty categories through a threestage procedure. The first stage applies title-level pattern matching to all findings. The second stage applies extended-context reclassification to findings that did not match in stage one, using body text extracted from the surrounding report where available. The third stage applies targeted large-language-model reclassification to remaining Critical- and High-severity findings and to the residual ‘Other’ bucket. The ‘Other’ rate on the full 23,818-finding dataset is 7.9%. Findings originating from web-native sources classify more reliably than PDF-extracted findings because their presentation is already machine-readable; PDF-extracted findings are noisier where titles are truncated or line-broken during extraction. 2.3 Incident dataset

Incident-level data is sourced from rekt.news with the publisher’s explicit written permission. Each incident is tagged with a date, an approximate loss amount in US dollars, a chain, a protocol type, and a primary root cause. Loss amounts are used only where the scraped value reasonably corresponds to the dollar value of assets actually stolen; entries where the scraped amount clearly reflects total value locked or token market capitalisation rather than an actual theft are excluded, or the loss is normalised with reference to the published post-mortem. Root-cause labels are retained in our own taxonomy rather than the rekt.news free-text field, because the two use different terminological conventions. The quality filter (a present, theftconsistent loss amount, excluding rug-pulls and editorial entries) reduces an initial archive of 242 entries to the 218 incidents analysed here, with aggregate losses of US$7.764 B. 2.4 Observation window

The window covers 1 January 2022 through 27 March 2026. The final three months represent a partial year and are labelled 2026* throughout. Year-over-year comparisons that include 2026 are restricted to those where the partial-year nature does not materially bias interpretation. Severity backfill lag—the empirical tendency of informational-severity findings to appear in public archives later than higher-severity ones—is documented in Section 4 and treated as a known source of apparent early-year skew.

The Audit Gap in Blockchain Security

4

% of all findings

40

37.5%

30 20 10 0

22.9%

22.4% 11.2% 6.0%

Critical

High

Medium

Low

Informational

Figure 1. Distribution of severity classifications across all 23,818 findings in the observation window.

2.5 Analytical caveat

Public audit findings and realised exploit incidents describe different populations. The audit dataset enumerates vulnerabilities identified in code reviewed within a defined scope. The incident dataset enumerates vulnerabilities actually exploited in deployed systems, which may or may not have been audited. The two are presented in parallel to enable comparison of categorical and temporal patterns, not to imply that either is a statistical sample of the other. Dataset construction, filtering, and classification were performed with AI-assisted tooling under human editorial oversight; methodology, data-quality thresholds, and interpretation were defined and validated by the author. Analytical code and intermediate datasets are retained for reproducibility.

3 Audit findings: overall distribution 3.1 Distribution by severity

Across the full dataset, 17.2% of findings are classified as Critical or High, comprising 1,439 Critical and 2,659 High findings. Medium-severity findings constitute 22.4%; Low-severity 37.5%; Informational 22.9% (Figure 1). The large share of Low and Informational findings is consistent with industry practice, in which audit reports function as holistic code-quality assessments rather than narrowly focused vulnerability enumerations. 3.2 Distribution by vulnerability category

Table 1 reports the frequency of vulnerability categories. The five most frequent—logic/businesslogic, code quality, input validation, access control, and initialisation/upgradeability—together account for 54.7% of all findings. The distribution has a long tail: the top sixteen categories account for 92.1% of the dataset, with the remaining 7.9% in the ‘Other’ residual. Three observations are worth noting. First, reentrancy—historically emblematic of smartcontract vulnerability—accounts for only 2.1% of findings, indicating that auditors now surface this class substantially less frequently than in the mid-2010s. Second, logic and business-logic errors constitute the single largest category, reflecting that many findings con-

The Audit Gap in Blockchain Security

5

Table 1. Category frequencies across all 23,818 findings, 2022–Q1 2026. Category

Count

Percent

Logic Error / Business Logic Code Quality Input Validation Access Control / Authorization Other Initialization / Upgradeability Integer Overflow / Arithmetic Oracle / Price Manipulation Gas / Efficiency Cross-chain / Bridge Signature / Replay Attack Denial of Service Precision / Rounding Errors Token Standard Compliance Unchecked Return Values Reentrancy

3479 3098 2372 2339 1882 1749 1278 1016 992 881 797 781 562 544 496 491

14.6 13.0 10.0 9.8 7.9 7.3 5.4 4.3 4.2 3.7 3.3 3.3 2.4 2.3 2.1 2.1

cern protocol-specific invariant violations rather than named CWE patterns. Third, oracle and price-manipulation findings account for only 4.3% of audit output but, as Section 6 reports, are associated with a substantially larger share of realised losses. 3.3 Distribution by technology stack

Solidity and EVM-compatible chains generate between 79% and 84% of findings in every year (Figure 2). The residual 16–21% is distributed across a changing composition of non-EVM stacks. Rust/Solana rises from approximately 3% in 2023 to a local peak of 8% in 2025 before retreating to 4% in Q1 2026. TON/FunC does not appear as a distinct category before 2023 and stabilises thereafter at 4–5%. CosmWasm and Cosmos SDK combined range between 5% and 8%, with Cosmos SDK becoming more visible from 2024 as chain-module audits entered public circulation. Move-based ecosystems (Aptos, Sui) constitute a smaller but fastest-growing segment in proportional terms, reaching 4% of Q1 2026 findings.

4 Temporal evolution, 2022–Q1 2026 4.1 Volume and severity

Published audit volume more than doubled between 2022 and 2024, rising from 2,526 findings to 7,412 (Figure 3). Volume retreated to 6,504 in 2025 and is running at 1,755 findings for Q1 2026, a rate consistent with the 2023 pace on an annualised basis. Growth from 2022 through 2024 reflects, in substantial part, the entry of new firms into public-report publication (twelve firms in 2022, seventeen in 2024) rather than increased identification rates per report. The 2025 contraction is consistent with a softer audit market. The Critical-plus-High share is remarkably stable across complete years: 22.7% in 2022, 16.8% in 2023, 16.0% in 2024, and 15.0% in 2025. The elevated 22.2% reading for Q1 2026 is partially attributable to the empirical pattern that Informational and Low-severity findings

% of annual findings

The Audit Gap in Blockchain Security

90.0 87.5 85.0 82.5 80.0 77.5 75.0 72.5 70.0

6

EVM share

Non-EVM tail

Rust / Solana TON / FunC CosmWasm Cosmos SDK (Go) Move / Aptos / Sui Other / multi-chain

8 6 4 2 2022

2023

2024

2025

0

2026*

2022

2023

2024

2025

2026*

Published findings

8000

Critical + High Medium / Low / Info

6000

7,412 6,504

5,621

4000 2000 0

2,526

2022

1,755

2023

2024

2025

Critical + High (% of findings)

Figure 2. Distribution of audit findings across technology stacks, as percentage of annual findings. Left: dominant EVM share. Right: composition of the non-EVM tail.

25.0 22.5

22.7

22.2

20.0 16.8

17.5

16.0

15.0

15.0

2022 25 mean 17.6%

12.5 10.0

2022

2023

2024

2025

2026*

2026*

Figure 3. Left: annual count of published findings, decomposed into Critical+High versus all other severities (2023 total inferred as the residual of the annual totals). Right: share of Critical+High findings per year, with the 2022–2025 mean shown for reference.

are backfilled to public archives with greater lag than higher-severity findings; early-year readings therefore over-represent the serious categories. Over the four complete years, the Critical-plus-High share is essentially flat. 4.2 Category evolution

The categorical composition exhibits several distinct movements within an otherwise stable overall distribution (Figure 4). The share of logic and business-logic findings declined from approximately 19% in 2022 to approximately 11% in 2025–2026; a substantial portion of this decline is reclassification-driven, as the pipeline routes findings that earlier datasets captured under a broad ‘logic’ label into more specific categories. The share of initialisation and upgradeability findings rose from approximately 4% to approximately 11%, consistent with the industry-wide shift from monolithic contract architectures to proxy-upgradeable patterns and the vulnerability classes associated with initialisation, admin-role management, and storage-slot collisions. The share of oracle and price-manipulation findings more than doubled, from approximately 2% to 6–7%. The share of reentrancy findings declined further, from approximately 6% to approximately 3%; Section 5 reports that the residual reentrancy bugs reaching production remain associated with disproportionate realised losses. Cross-chain and bridge findings peaked in 2023 and have trended downward, consistent with reduced

% of annual findings

The Audit Gap in Blockchain Security

7

20.0 Logic / business logic 17.5 15.0 12.5 10.0 7.5 Reentrancy 5.0 Initialisation / upgradeability 2.5 Oracle / price manip. 0.0 2022

11% 6.5% 3%

2025

Figure 4. Movement in the share of audit findings for four categories between 2022 and 2025 (stated endpoints). Table 2. Technology-stack distribution of audit findings, percentage of annual findings. Stack

2022

2023

2024

2025

2026*

Solidity / EVM Rust / Solana TON / FunC CosmWasm Cosmos SDK (Go) Move / Aptos / Sui Other / multi-chain

80% 4% — 4% 0% 0% 9%

82% 3% 4% 3% 1% 1% 6%

79% 5% 5% 2% 3% 1% 8%

80% 8% 5% 2% 1% 1% 3%

84% 4% 4% 1% 1% 4% 2%

novel-bridge launch rates and increased audit coverage of the existing bridge generation. 4.3 Stability alongside movement

Against the categorical drift documented above, three features of the dataset are essentially unchanged over the complete four-year period: the Critical-plus-High share (within a 15– 17% band); the identity of the top five categories (access control, logic, initialisation, input validation, and arithmetic, with some reordering); and the dominance of Solidity/EVM (79– 84% of annual findings). The audit picture does not reconfigure itself year-over-year; it is stable at a level consistent with an industry that has developed its review practices to address a known set of recurring failure modes. 4.4 Technology-stack shifts

Table 2 presents the year-over-year evolution of the technology-stack distribution. The EVM share is stable; the composition of the non-EVM tail is not. TON/FunC is absent from 2022 public audit output because reporting conventions for that stack crystallised in mid-2023. The 2025 Rust/Solana peak coincides with a renewed phase of Solana DeFi activity. The Move ecosystem exhibits the fastest proportional growth, though from a small absolute base.

3.0

8

70

2.91

60

2.5

2.36

50

Incident count

Aggregate losses (US$B)

The Audit Gap in Blockchain Security

2.0

40

1.5

1.28

1.0

30

1.08

20

0.5 0.0

0.13

2022

2023

2024

2025

2026*

10 0

Figure 5. Documented exploit incidents (line, right axis) and aggregate annual losses (bars, left axis), 2022–Q1 2026. Table 3. Annual decomposition of exploit activity. The ‘Audited’ column reports the count of incidents where the affected protocol had received at least one public audit. Year

Incidents

Total losses

Audited

Audited losses

2022 2023 2024 2025 2026*

46 53 45 59 15

US$2.91 B US$1.28 B US$1.08 B US$2.36 B US$0.13 B

23 37 23 17 5

≈ US$1.42 B ≈ US$820 M ≈ US$350 M ≈ US$1.67 B ≈ US$40 M

Full window

218

US$7.76 B

105

≈ US$4.30 B

5 Real-world exploit incidents Over the window, the rekt.news archive documents 218 incidents meeting the quality filters of Section 2, with aggregate losses of approximately US$7.76 B. Table 3 summarises the annual decomposition. Losses are volatile year-over-year because single large events dominate annual totals: 2022 is influenced heavily by the Ronin Network exploit (US$624 M); 2025 by the Bybit multi-signature phishing incident (US$1.44 B); the 2025 DeFi-only loss ranking is headed by the Cetus arithmetic-overflow exploit on Sui (US$223 M). 5.1 Root-cause distribution

Table 4 reports the distribution of aggregate losses by root cause, ordered by total dollar impact. Three categories dominate. Private-key compromise alone accounts for US$1,894 M (24.4% of cumulative losses). Phishing and social engineering account for US$1,511 M (19.5%). Accesscontrol failures account for US$994 M (12.8%). Together these three represent 56.7% of aggregate realised losses. Private-key compromise combines a relatively high incidence count (45) with a high mean loss (≈US$42 M) to produce the largest aggregate. Phishing and socialengineering incidents are few (12) but exhibit the second-highest mean loss (≈US$126 M), disproportionately influenced by the single Bybit event. Bridge exploits and signature/replay attacks are rare but exhibit the highest mean per-incident losses of any category: the four bridge events average US$148 M each, and the two signature-replay events—Wormhole (February

The Audit Gap in Blockchain Security

9

Table 4. Distribution of aggregate realised losses by root-cause category, full observation window. Root cause Private Key Compromise Phishing / Social Engineering Access Control Oracle / Price Manipulation Bridge Exploit Signature / Replay Attack Logic Error / Business Logic Integer Overflow / Arithmetic Flash Loan Reentrancy Supply Chain / Dependency Governance Attack

Incidents

Total losses

% losses

Mean / incident

45 12 30 43 4 2 37 7 5 14 7 5

US$1,894 M US$1,511 M US$994 M US$666 M US$593 M US$516 M US$298 M US$288 M US$266 M US$256 M US$243 M US$206 M

24.4 19.5 12.8 8.6 7.6 6.6 3.8 3.7 3.4 3.3 3.1 2.7

US$42.1 M US$125.9 M US$33.1 M US$15.5 M US$148.3 M US$258.0 M US$8.1 M US$41.1 M US$53.1 M US$18.3 M US$34.7 M US$41.2 M

2022) and Nomad (August 2022)—average US$258 M.

6 Divergence between audit output and exploit activity The central empirical finding is that the categorical distribution of published audit findings does not correspond to the categorical distribution of realised exploit losses. Figure 6 presents the two distributions side by side. 6.1 Observed misalignment

Of the twelve most frequent audit categories and the twelve largest exploit-loss root causes, access control is the only category appearing in the top four on both sides (Table 5). The remaining top-four audit categories—logic errors, code quality, and input validation—together account for 37.6% of audit output but only 3.8% of realised dollar losses (predominantly through the Logic Error category, in position seven on the loss side). Conversely, private-key compromise, phishing and social engineering, and supply-chain dependency attacks together account for approximately 47% of realised losses but constitute a negligible share of audit findings. 6.2 Interpretation

The misalignment admits a specific interpretation. A conventional smart-contract audit reviews the source code of a defined commit against a defined scope. It does not review the deployment environment, signer key-management practices, the continuous-integration pipeline, front-end hosting infrastructure, or the supply chain of third-party dependencies. The categories that audits surface are therefore, by construction, those visible in source code under review. The categories that drive realised losses include several classes of attack—principally operationalsecurity failures and human-facing manipulation—that are not visible in source code, whether because they do not exist there (key theft, phishing) or because they arise in artefacts outside the audited scope (supply-chain compromise, governance-vote manipulation after a discrete contract has been audited).

The Audit Gap in Blockchain Security

10

Audit findings (% of count) Logic / Business Logic Code Quality Input Validation Access Control Initialisation / Upgrade. Arithmetic Oracle / Price Manip. Cross-chain / Bridge Signature / Replay Denial of Service Reentrancy

Exploit losses (% of US$)

14.6 Private Key Compromise 13.0 / Soc. Eng. Phishing 10.0 Access Control 9.8Oracle / Price Manip. 7.3 Bridge Exploit 5.4 Signature / Replay 4.3 Logic / Business Logic 3.7 Arithmetic 3.3 Flash Loan 3.3 Reentrancy 2.1 Supply Chain / Dependency

0.0

2.5

5.0

7.5

10.0 12.5 15.0

24.4 19.5 12.8 8.6 7.6 6.6 3.8 3.7 3.4 3.3 3.1

0

5

Human-vector root cause

10

15

20

25

Figure 6. Left: audit findings by category (% of count, 2022–Q1 2026). Right: realised exploit losses by root cause (% of US$). Human-vector root causes on the right panel are outlined in black. Table 5. Rank comparison of audit-finding categories against exploit-loss root causes. Rank by audit frequency

Rank by exploit losses

1. Logic Error / Business Logic (14.6%) 2. Code Quality (13.0%) 3. Input Validation (10.0%) 4. Access Control / Authorization (9.8%) 5. Initialization / Upgradeability (7.3%) 6. Integer Overflow / Arithmetic (5.4%) 7. Oracle / Price Manipulation (4.3%) 8. Reentrancy (2.1%) 9. Denial of Service (3.3%) 10. Cross-chain / Bridge (3.7%)

1. Private Key Compromise (US$1.89 B) 2. Phishing / Social Engineering (US$1.51 B) 3. Access Control (US$994 M) 4. Oracle / Price Manipulation (US$666 M) 5. Bridge Exploit (US$593 M) 6. Signature / Replay Attack (US$516 M) 7. Logic Error / Business Logic (US$298 M) 8. Integer Overflow / Arithmetic (US$288 M) 9. Reentrancy (US$256 M) 10. Supply Chain / Dependency (US$243 M)

This should not be read as a claim that audits are ineffective. The relevant conclusion is more specific: the effective scope of a conventional audit is narrower than the risk surface that the word ‘security’ is implicitly taken to cover in the Web3 context. Access control is the exception that illustrates the rule, because it sits at the interface between source code and operational reality, and is therefore the one category where a static code review and a deployed-system attack can meaningfully meet. This aggregate divergence is consistent with prior findings at the level of individual contracts: that flagged vulnerabilities are rarely exploited in practice [2], and that contemporary analysis tools would have prevented only a minority of high-impact real-world attacks [3].

7 The prominence of human-vector attacks We classify four root-cause categories as human-vector: private-key compromise, phishing and social engineering, supply-chain or dependency compromise, and governance attack. The

The Audit Gap in Blockchain Security

11

Table 6. Annual incidence and loss share of human-vector attacks. Incidents

Human-vector

% of incidents

H-V losses

% of losses

2022 2023 2024 2025 2026*

46 53 45 59 15

11 19 17 19 3

23.9 35.8 37.8 32.2 20.0

US$502 M US$843 M US$803 M US$1,667 M US$34 M

17.3 65.6 74.6 70.8 25.3

Full window

218

69

31.7

US$3,849 M

49.6

Human-vector share (%)

Year

80 70 60 50 40 30 20 10 0

% of annual losses % of incidents

2022

2023

2024

2025

2026*

Figure 7. Share of annual losses and of incident counts attributable to human-vector root causes.

defining feature common to this grouping is that the proximate cause of the realised loss sits outside the static source code of the affected protocol—whether at the operational-security layer (keys, signer workflows), the distribution layer (front-ends, build artefacts, third-party packages), or the governance layer (vote mechanisms operating after deployment). 7.1 Temporal pattern

Human-vector attacks accounted for the majority of annual losses in each of 2023, 2024, and 2025, with loss shares of 65.6%, 74.6%, and 70.8% respectively (Table 6, Figure 7). The 2022 pattern is the exception: that year’s losses concentrated in a small number of large bridge exploits (Ronin, Wormhole, Nomad, Harmony, BNB Bridge) whose root causes were either contract-level or signature-level rather than operational. 7.2 Illustrative incidents

A small number of large incidents define the human-vector pattern. Bybit (February 2025, US$1.44 B): a multi-signature approval workflow was compromised through targeted phishing; signers approved a disguised upgrade of the cold-wallet contract. The underlying contract had been audited; the approval interface had not. This is the largest single cryptocurrency theft to date and accounts for 18.4% of cumulative losses in the window. DMM Bitcoin (May 2024, US$304 M): private keys stolen, no audit-visible code defect. WazirX (July

The Audit Gap in Blockchain Security

12

Table 7. Largest exploit incidents affecting protocols with prior public audit coverage. Audit attribution is drawn from the rekt.news archive. Protocol Bybit BNB Bridge Mixin Network Euler Finance Nomad Bridge Beanstalk Poloniex Harmony Bridge Heco / HTX Orbit Bridge FEI / Rari Qubit Finance

Date

Loss

2025-02-22 2022-10-07 2023-09-25 2023-03-14 2022-08-02 2022-04-18 2023-11-10 2022-06-24 2023-11-22 2024-01-03 2022-05-01 2022-01-28

US$1,430 M US$586 M US$200 M US$197 M US$190 M US$181 M US$126 M US$100 M US$99 M US$82 M US$80 M US$80 M

Root cause Phishing of multi-sig signers Signature replay on IAVL Merkle proofs Supply-chain (cloud DB credentials) Flash-loan donation attack Initialisation bug in signature check Governance attack (flash-loaned votes) Private-key compromise Private-key compromise (2 of 5 signers) Private-key compromise Private-key compromise Reentrancy Logic error in bridge deposit

2024, US$235 M): multi-signature workflow compromise. Mixin Network (September 2023, US$200 M): supply-chain compromise of cloud-hosted database credentials. Ronin Network (March 2022, US$624 M): validator-key compromise of five of nine validators, not a contractlayer event. Poloniex (November 2023, US$126 M): hot-wallet private-key compromise at an exchange. 7.3 The audited-but-exploited pattern

Of the 218 incidents, 105 (48%) affected protocols that had received at least one public audit prior to the event. These account for approximately US$4.3 B of the US$7.76 B aggregate, or roughly 55% of cumulative losses. Table 7 lists the twelve largest such incidents. Of these twelve, nine fall into root-cause categories outside the conventional smart-contract audit scope (phishing, private-key compromise, signature replay against infrastructure rather than the audited contract, supply-chain compromise, or governance-mechanism manipulation). The remaining three—Nomad (initialisation), Euler (donation attack), Qubit (logic error)—are cases where the audit-scope root cause was real but the affected version was either post-audit, outside the explicit scope, or arose in a path not fully enumerated during review.

8 Supplementary findings 8.1 Concentration of losses (Pareto)

Aggregate losses exhibit a steeply concentrated distribution (Figure 8). The single largest incident (Bybit) alone accounts for 18.4% of cumulative losses. The eight largest account for 50.6%; the twenty largest for 71.4%. The remaining 198 incidents—91% of the count— together account for less than 29% of cumulative losses. The distributional shape is inconsistent with an assumption of approximately Gaussian losses; risk models assuming Gaussianity will systematically underestimate annual worst-case outcomes.

The Audit Gap in Blockchain Security

13

Cumulative % of losses

100 80 20 incidents 71.4%

60

8 incidents 50.6%

40 20 0

0

25

50

75

100

125

150

Incident rank (by loss, descending)

175

200

Figure 8. Cumulative share of aggregate losses as a function of incident rank.

Incidents

120 100

101

5

98

4.61

4

80

3

60

2.74

2

40 19

20 0

Aggregate losses (US$B)

Ethereum

BNB Chain Other chains

1 0

0.42 Ethereum

BNB Chain Other chains

Figure 9. Incident counts and aggregate losses by chain (BNB Chain aggregates BSC and BNB Beacon Chain).

8.2 Chain concentration

Incident activity is concentrated by chain to an even greater degree than by root cause (Figure 9). Ethereum and BNB Chain together host 89% of all incidents (199 of 218) and 94% of all losses (US$7.35 B of US$7.76 B). Every other chain—Solana, Sui, Arbitrum, Stacks, Stellar, Cosmos, Base, Sonic, and several smaller chains—contributes a single-digit incident count across the full window. The pattern admits two interpretations: either chains outside the top two have held up well under adversarial pressure, or most do not yet host sufficient total value locked to justify large-scale attacker investment, and incident concentration will broaden as they grow. The 2025 Cetus exploit on Sui (US$223 M) is the first substantive indication that the second interpretation is in part correct: a single large loss can shift the risk profile of a chain materially. 8.3 Distributional shape by root cause

Because the loss distribution is heavy-tailed, mean per-incident values are unreliable summary statistics for categories that include one or more catastrophic events (Figure 10). The mean

US$M per incident

The Audit Gap in Blockchain Security

160 140 120 100 80 60 40 20 0

14

Mean Median

148 126

42 8.4

6.5

Private Key Compromise

Phishing / Soc. Eng.

3.5

Bridge Exploit

Figure 10. Mean versus median loss per incident, for three root causes with n ≥ 3 incidents (median for private-key compromise inferred from the reported mean-tomedian ratio). Table 8. Distribution of aggregate losses by protocol type. Protocol type CEX (centralised exchange) Bridge Lending DEX Wallet / custody Derivatives Yield / staking DAO

Incidents

Total losses

Mean / incident

Typical failure mode

24 22 48 31 7 11 8 3

US$2,284 M US$2,401 M US$736 M US$603 M US$400 M US$71 M US$40 M US$6 M

US$95.2 M US$109.1 M US$15.3 M US$19.5 M US$57.2 M US$6.5 M US$5.0 M US$2.0 M

Key compromise, phishing Signature/replay, key compromise Oracle manipulation, logic Arithmetic, oracle, reentrancy Key compromise Oracle, logic Logic, reentrancy Governance

phishing incident is US$126 M; the median is US$6.5 M, a ratio of approximately 19. The mean bridge exploit is US$148 M; the median is US$3.5 M, a ratio of approximately 42, reflecting the outsized 2022 events. Even private-key compromise, the most uniformly large human-vector category, exhibits a mean-to-median ratio of approximately 5. The practical implication is that planning against a category of attack should be conducted against the tail of the distribution rather than its mean. For centralised-exchange and bridge risk in particular, the mean per-incident loss substantially understates the distribution of plausible outcomes. 8.4 Protocol-type concentration of losses

Centralised exchanges and bridges together account for approximately 60% of cumulative losses across 21% of incidents (Table 8). Lending, the most numerous protocol type by incident count, accounts for less than 10% of cumulative losses. Protocol type thus functions as a stronger predictor of tail-loss exposure than the specific vulnerability class most commonly associated with that type. 8.5 The audit-coverage paradox, reconsidered

A naïve reading of Table 3—48% of incidents affected audited protocols, accounting for 55% of losses—might support the conclusion that audits do not reduce realised losses. The

The Audit Gap in Blockchain Security

15

data does not support that conclusion. The majority of audited-but-exploited incidents have root causes outside the audited scope, as Section 7 documents. Where a contract-level bug was the direct cause (Nomad initialisation, Euler donation flash-loan, and others), the audit and the deployed vulnerable code diverged in identifiable ways: subsequent code changes, explicitly out-of-scope paths, or undeployed audited versions. A more accurate reading is that audits are narrowly effective but broadly incomplete relative to the risk surface they are often implicitly relied upon to defend. The category of protocol whose failure mode is primarily operational—custodians, bridges, exchanges—is the same category for which a static contract review provides the least defensive coverage and for which complementary operational-security engagements are most impactful on the loss distribution.

9 Conclusion The analysis supports five empirical claims across the observation window. (1) The distribution of public audit findings is substantively stable: the Critical-plus-High share has remained within a narrow band (15–17% in complete years), and the identity of the five most frequent categories has not changed; drift is present at the individual-category level—most notably the rise of initialisation and oracle findings, and the decline of reentrancy—but the overall distribution does not exhibit regime change. (2) The distribution of realised exploit losses has shifted substantially: from 2023 onwards, human-vector attacks account for the majority of annual losses, in contrast to the comparatively stable code-review picture described by claim (1). (3) Audit output and realised exploit output describe different populations whose categorical distributions do not correspond; access control is the only category in the top four of both, and the misalignment substantially reflects the narrower scope of a conventional audit (source code of a specific commit) relative to the broader risk surface on which losses are realised (operational security, deployment infrastructure, and human-facing approval workflows). (4) The distribution of realised losses is heavy-tailed: eight incidents account for 50.6% of cumulative losses and twenty for 71.4%; chain-level and protocol-type concentration (89% of incidents on two chains; 60% of losses in two protocol types) reinforces the tail dependence, and security programmes designed against mean outcomes systematically underprovision against tail outcomes. (5) Solidity and EVM remain the dominant stack throughout, but the long tail of non-EVM stacks is present and growing, and teams designing security programmes for multi-stack products on the assumption of single-stack audit availability are accumulating coverage gaps the public-report record does not yet fully illuminate. The overall picture does not support a claim that the ecosystem has become more or less secure in any simple sense. The code-review problem, as surfaced by public audit output, is approximately where it was in 2022. The operational-security problem, as surfaced by the incident record, grew substantially through 2023–2025. Both components require continued attention, from disciplines that have historically operated with relatively little overlap. The data supports a portfolio view of Web3 security practice in which code review and operationalsecurity engagements are treated as complementary rather than substitutable: audits for the bugs that exist in the code as written, and parallel assessments—of key management, signer workflows, build-pipeline hardening, and dependency supply chain—for the categories that drive the empirically larger share of realised losses.

The Audit Gap in Blockchain Security

16

Data and reproducibility Incident-level metadata is licensed from rekt.news under explicit written permission for the factual fields used here (dates, loss amounts, chain, protocol, and root-cause classification); no article text is reproduced. All rekt.news content remains the copyright of Rekt News (EU trademark registration 018857408). Analytical code and intermediate datasets are retained and available on request.

Acknowledgements The author gratefully acknowledges rekt.news for granting explicit written permission to use their incident archive as a primary data source; the real-world exploit analysis in this paper would not have been possible without that collaboration. The author also thanks the independent security firms whose public audit output forms the basis of the findings dataset. Public disclosure of audit results—whether through repository-based publication or webnative report portals—is what makes empirical work of this kind possible. Any errors in the aggregation, classification, or interpretation of the data are the author’s alone.

References [1] N. Atzei, M. Bartoletti, and T. Cimoli. A Survey of Attacks on Ethereum Smart Contracts (SoK). In Principles of Security and Trust (POST), LNCS 10204, pp. 164–186, Springer, 2017. [2] D. Perez and B. Livshits. Smart Contract Vulnerabilities: Vulnerable Does Not Imply Exploited. In 30th USENIX Security Symposium, pp. 1325–1341, 2021. arXiv:1902.06710. [3] S. Chaliasos, M. A. Charalambous, L. Zhou, R. Galanopoulou, A. Gervais, D. Mitropoulos, and B. Livshits. Smart Contract and DeFi Security Tools: Do They Meet the Needs of Practitioners? In IEEE/ACM 46th International Conference on Software Engineering (ICSE), Article 60, 2024. arXiv:2304.02981. [4] L. Zhou, X. Xiong, J. Ernstberger, S. Chaliasos, Z. Wang, Y. Wang, K. Qin, R. Wattenhofer, D. Song, and A. Gervais. SoK: Decentralized Finance (DeFi) Attacks. In IEEE Symposium on Security and Privacy (S&P), pp. 2444–2461, 2023. [5] S. M. Werner, D. Perez, L. Gudgeon, A. Klages-Mundt, D. Harz, and W. J. Knottenbelt. SoK: Decentralized Finance (DeFi). In ACM Conference on Advances in Financial Technologies (AFT), pp. 30–46, 2022. [6] S.-S. Lee, A. Murashkin, M. Derka, and J. Gorzny. SoK: Not Quite Water Under the Bridge: Review of Cross-Chain Bridge Hacks. In IEEE International Conference on Blockchain and Cryptocurrency (ICBC), 2023. [7] N. Belenkov, V. Callens, A. Murashkin, K. Bak, M. Derka, J. Gorzny, and S.-S. Lee. SoK: A Review of Cross-Chain Bridge Hacks in 2023. arXiv:2501.03423, 2025. [8] Rekt News. The Rekt Leaderboard. https://rekt.news/leaderboard/. Accessed 27 March 2026.

Record · ID 280141 · SHA-256 66155645f5260a05
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.