A dataset of early blockchain-registered AI agents on Ethereum Yulin Liu, Quantum Economics AI Lab, Zurich, Switzerland [email protected]
Abstract This study presents a structured dataset of blockchain-registered artificial intelligence agents under the ERC-8004 standard on Ethereum. The dataset integrates on-chain identity records, minting transactions, transfer events, reputation summaries, and individual feedback records, together with resolved off-chain metadata where available. Data were collected from Ethereum mainnet using Web3 RPC queries and processed into tabular form to enable reproducible analysis. The dataset covers 10,000 agents within a defined block range and includes both event-level records and aggregated summaries. It enables empirical research on agent identity formation, reputation systems, service exposure, and early-stage decentralized AI ecosystems. This resource supports studies in blockchain analytics, decentralized trust infrastructure, and the emerging agentic economy.
Background & Summary Blockchains[1] have become an important source of large-scale, machine-readable records for studying digital economic activity[2-4]. Prior work has shown that leveraging blockchain’s immutable provenance and decentralized governance can establish a trustworthy foundation for data exchange, thereby enhancing the auditability and integrity of large-scale scientific datasets[5,6]. In the Bitcoin[7-11] context, cohort-based processing has been used to reorganize the full transaction history into economically meaningful datasets and indicators, making a rapidly growing ledger more tractable for research[12,13]. In Ethereum[14-16], data-intensive empirical
work has likewise demonstrated the value of combining on-chain transaction records with protocol-level context to study transaction fee mechanisms and user behavior[17]. Building on this line of research, the present work extends blockchain data curation from transactions and fees to a new emerging object of study: on-chain registered artificial intelligence agents. Recent advances in large language models, autonomous software agents[18-20], and agent-to-agent interaction frameworks have accelerated interest in an “agentic economy” in which software agents discover one another, exchange services, coordinate tasks, and engage in financial activity. In practice, once an AI agent begins to hold assets, receive payments, trigger transactions, or participate in decentralized applications, it requires a blockchain-compatible identity and wallet layer. However, communication protocols alone do not solve the full coordination problem. Protocols such as Model Context Protocol[21] and agent-to-agent[22] messaging standards help define how agents expose capabilities or exchange messages, but they do not by themselves provide a persistent answer to who an agent is, how it can be addressed across systems, or whether it has accumulated credible evidence of trustworthy behavior. ERC-8004 addresses this gap by proposing a trust infrastructure for autonomous agents built around three linked components: identity, reputation, and validation[23]. The Identity Registry establishes unique on-chain agent identities as ERC-721 tokens[24], associating each agent with a registry namespace, token identifier, ownership record, and a metadata pointer that can resolve to off-chain descriptive files. The Reputation Registry records historical feedback submitted by counterparties, allowing agents to accumulate performance signals over time through quantitative values and flexible user-defined tags. The broader architectural intent also includes a Validation Registry for higher-assurance verification of agent outputs, although this component remains immature in current deployments and does not yet provide comparable observable data. Taken together, these
registries suggest an infrastructure for machine-readable identity, accountability, and trust in multi-agent environments. Despite the conceptual importance of this standard, ERC-8004 data remain difficult to analyze in practice. Information is distributed across smart contract storage, emitted events, token metadata URIs, and heterogeneous off-chain JSON documents. The protocol is intentionally minimalist on-chain: only the minimum data needed for interoperability and verification are recorded at the contract layer, while richer descriptive content, service endpoints, and operational details are delegated to off-chain metadata. This design improves flexibility, but it also fragments the observable record. As a result, researchers and builders currently lack a structured public dataset that documents the early formation of ERC-8004 agent identity and reputation at scale. Here we present a dataset of the first 10,000 ERC-8004 AI agents registered on Ethereum mainnet. Observation begins with agent 0, minted at block 24,339,925 on 2026-01-29 10:31:11 UTC, and extends through block 24,839,925, spanning 500,000 Ethereum blocks. Within this window, we document all agent identifiers from 0 through 9,999. For each registered agent, the core dataset records mint block, mint timestamp, mint transaction hash, owner wallet, agent wallet where available, token URI, metadata hosting type, and lifecycle status. Across the 10,000 agents, 7,856 are observed in a minted_only state and 2,144 in a metadata_linked state. To complement these on-chain identity records, we resolve and normalize off-chain metadata for the subset of agents whose token-linked metadata could be retrieved and parsed successfully. In the current release, structured metadata were resolved for 72 agents, including agent names, descriptions, image URLs, active flags, x402 support indicators, and counts of registered services, trust entries, and registrations. We additionally extracted 112 service records from metadata, covering service types such as MCP, A2A, OASF, x402, web, docs, email, and HTTP-JSON-RPC endpoints, along with
optional version, skills, and domains fields. These metadata reveal how ERC-8004 identities can serve as a bridge between NFT-style on-chain registration and richer agent capability descriptions hosted off-chain. The dataset also captures the early reputation layer of ERC-8004. 980 underlying feedback records were observed from 197 unique client addresses. These feedback records store value and precision fields together with flexible tag dimensions, enabling heterogeneous assessment criteria rather than a single fixed rating scheme. In our observation window, 930 feedback entries contain a first tag and 899 contain a second tag; common examples include liveness, trust, quality, helpfulness, and reachability, with paired operational tags such as liveness-check and oracle-screening. No feedback revocations were observed in this release. This structure is important because it reflects the protocol’s design philosophy: trust signals are intended to be composable, application-specific, and extensible rather than reduced to one universal score. This dataset is designed to support both scientific and applied reuse. For researchers, it provides an early empirical record of how agent identity and reputation are instantiated on-chain, enabling work on trust formation, service discovery, reputation systems, agent ecosystems, and the financial behavior of autonomous software entities. For builders, the dataset provides a practical index of registered ERC-8004 agents, metadata endpoints, and service declarations that can support dashboards, registries, search tools, analytics pipelines, and protocol monitoring. More broadly, the dataset offers a reproducible baseline for studying the earliest large-scale wave of on-chain AI agent registration on Ethereum, analogous to how earlier blockchain datasets enabled systematic study of transactions, fees, and user activity in prior stages of crypto-economic development[25-28].
Methods
The overall data collection and processing pipeline is illustrated in Fig. 1.
Figure 1: Data collection and processing pipeline for ERC 8004 agents. Data are collected from the Ethereum blockchain by querying ERC 8004 identity and reputation contracts. Agent identities are discovered from event logs and contract calls, and then processed through three parallel pipelines: identity extraction, metadata resolution, and reputation extraction. These pipelines generate structured tables including agents_core, agent_metadata, agent_services, agent_reputation_summary, and agent_feedback_records. All outputs are integrated into a unified dataset anchored to a fixed observation block and exported for downstream analysis. Data sources and study scope
Data were collected from Ethereum mainnet, focusing on the ERC-8004 identity registry at 0x8004A169FB4a3325136EB29fA0ceB6D2e539a432 and the reputation registry at 0x8004BAa17C55a88189AE136b182e5fdA19dE9b63. The dataset corresponds to the namespace eip155:1:0x8004A169FB4a3325136EB29fA0ceB6D2e539a432. Collection was performed using Python, Web3.py, Requests, and a PostgreSQL-compatible database. The study covers agents with identifiers 0–9999. The observation window spans blocks 24,339,925 to 24,839,925, anchored at a fixed observation block 24,839,925 for state queries. Transfer events were reconstructed from on-chain logs emitted by the ERC-8004 identity registry contract using an event-based scanning pipeline. Due to the absence of a dedicated indexed transfer history in the contract state, events were extracted by iterating over a bounded block interval and decoding log topics corresponding to agent ownership changes. To balance computational feasibility with empirical coverage, we restricted extraction to a predefined observation window spanning blocks 24,399,925 to 24,499,925. All detected transfer events within this interval were normalized into a structured table, including source and destination addresses, agent identifiers, and associated transaction metadata. Because minting operations emit transfer-like events in the ERC-8004 standard, these were programmatically distinguished and cross-referenced with the mint economics table to avoid duplication in downstream analyses. Snapshot design
A fixed observation block was used to ensure that all extracted data represent a consistent state of each agent at a specific point in time. Immutable attributes such as agent identifiers and minting information are independent of the observation block, whereas dynamic attributes—including ownership, reputation metrics, feedback counts, service declarations, and lifecycle status—depend on the cumulative history of on-chain events and therefore vary across blocks. Querying contract state at a single predefined block
ensures consistency, reproducibility, and comparability across agents in downstream use cases. Agent discovery
Agents were identified by scanning Transfer(address,address,uint256) events emitted by the identity registry. Mint events were detected by filtering logs where the sender topic equals the zero address. Discovered token IDs were filtered to the range 0–9999 and deduplicated. Identity extraction
For each agent, the identity state was queried at the observation block using ownerOf(tokenId) and tokenURI(tokenId). Mint block and transaction hash were obtained from event logs, and timestamps were derived from block headers. Records were stored in agents_core with normalized addresses and transaction hashes. Metadata hosting type was inferred from URI prefixes. Lifecycle status was assigned as metadata_linked if a token URI was present, and minted_only otherwise. Mint transaction data
Mint transactions were enriched by querying transaction objects, receipts, and block headers. Extracted fields include gas used, gas price, transaction value, nonce, block gas limit, and block gas used. Mint cost was computed as gas_used × gas_price. Results were stored in mint_economics, keyed by transaction hash. Metadata resolution
Token URIs were resolved to HTTP-accessible endpoints. IPFS URIs were mapped to gateway URLs, and base64-encoded data URIs were decoded when present. HTTP responses were required to be non-empty and valid JSON. Parsed fields include name, description, image URL, active status, x402 support, and counts of services, trust entries, and registrations. Service entries were normalized into agent_services, preserving order, endpoint,
version, and optional skills and domains. Cross-chain registrations were parsed and stored separately. Agents without resolvable metadata were retained in the dataset. Reputation extraction
Reputation data were collected at the observation block using the ERC-8004 reputation registry. For each agent, client addresses were obtained via getClients. Aggregate statistics were computed using getSummary. Individual feedback records were retrieved by iterating over indices returned by getLastIndex and reading entries via readFeedback. Feedback records include value, decimals, tags, and revocation status, and were stored in agent_feedback_records. Aggregate results were stored in agent_reputation_summary. Database construction
The dataset was organized as a relational schema with tables for identity, metadata, services, cross-chain registrations, reputation summaries, feedback records, and mint transaction data. Records were inserted using upsert operations. Child tables were refreshed by deleting and rewriting rows per agent to maintain consistency. Concurrency and fault tolerance
Data collection used bounded thread pools with staged execution. Remote Procedure Call (RPC) requests were rate-limited using semaphores and global throttling, with exponential backoff and retry for transient errors. A second-pass retry stage reprocessed failed records serially. Metadata requests were executed with explicit timeouts and strict JSON validation. Inclusion and normalization
All agents in the target ID range were retained regardless of metadata or reputation availability. Addresses and transaction hashes were normalized to lowercase. Timestamps were stored in UTC. Numeric reputation values were
stored as raw integers with explicit decimal precision. Revoked feedback entries were preserved.
Data Record The structure of the dataset is illustrated in Fig. 2.
Figure 2: Relational schema of the ERC 8004 agent dataset. Relational schema of the ERC 8004 agent dataset. The agents_core table serves as the primary index of agent identities. Metadata-related information is stored in agent_metadata and extended by agent_services. Reputation data are represented by agent_reputation_summary and detailed feedback records in agent_feedback_records. Additional tables include mint_economics for transaction-level data, crosschain_registrations for cross-chain mappings, and transfer_history for ownership transfers. Relationships between tables are
defined through agent_id or transaction hashkeys, enabling structured integration of on-chain and off-chain data.
Technical Validation Completeness of agent coverage
The dataset was designed to cover all ERC-8004 agents with identifiers in the range 0–9999. Completeness was verified by comparing the observed agent identifiers against the full expected range. No missing identifiers were detected, indicating that all agents within the target interval were successfully indexed and included in the agents_core table. Transfer activity for ERC-8004 agents was found to be sparse over the observed period, with relatively few ownership changes compared to the total number of registered agents. Additionally, reconstructing transfer histories requires full log traversal rather than direct state queries, resulting in substantial computational overhead. Given these characteristics, we provide transfer data as an auxiliary table derived from a bounded observation window rather than a full historical reconstruction. This design preserves representative behavioral signals while avoiding excessive infrastructure requirements that may hinder dataset reproducibility. Users should therefore interpret the transfer_history table as a partial but systematically extracted record of agent ownership dynamics, rather than an exhaustive ledger of all transfers since contract deployment. Consistency of identity data
Identity records were derived directly from on-chain events and contract state at a fixed observation block. Minting information, including block number, transaction hash, and timestamp, was cross-referenced between event logs and block headers. Ownership data and token URIs were retrieved via contract calls at the observation block, ensuring that all identity fields correspond to a consistent on-chain state.
Metadata validation and coverage
Metadata resolution was performed by retrieving and parsing token-linked URIs. A subset of agents (72 out of 10,000) yielded valid JSON metadata. The primary cause of metadata retrieval failure was that many token URIs resolved to HTTP endpoints returning non-JSON content, preventing structured parsing. This outcome is consistent with the ERC-8004 design, where metadata is optional and externally hosted, and therefore not guaranteed to be uniformly available or machine-readable. Reputation data integrity
Reputation data were validated at both the aggregate and record levels. Summary statistics were successfully computed, while individual feedback records were observed for 628 agents, totaling 980 entries. This discrepancy reflects the design of the reputation registry, where summary values may be defined even in the absence of extensive feedback activity. All feedback entries include raw values and decimal precision, preserving exact on-chain representations. No anomalous values, such as negative feedback scores or invalid decimal ranges, were observed. Additionally, no revoked feedback entries were present in the dataset. Reproducibility
The dataset was constructed using a deterministic pipeline based on a fixed observation block. All contract state queries were executed using historical eth_call, enabling exact reconstruction of the dataset given the same block height and RPC source. The full data collection logic is provided as open-source code, allowing independent verification and regeneration of the dataset.
Usage Notes The dataset is designed for both scientific analysis and practical applications involving ERC-8004 AI agents. The primary table, agents_core, provides a complete index of the first 10,000 registered agents and can be used as an entry point for all downstream analyses. Researchers may use the dataset to study the emergence of on-chain AI agent identity systems, including the distribution of ownership, the temporal dynamics of agent creation, and the early-stage formation of reputation signals. The availability of both aggregated (agent_reputation_summary) and record-level (agent_feedback_records) data enables analysis of reputation mechanisms, including feedback sparsity, client participation, and the use of multi-dimensional tags such as liveness, trust, and quality. The metadata and service tables (agent_metadata and agent_services) provide a partial but structured view of how agents expose capabilities and endpoints. These tables can be used to analyze interoperability patterns across emerging protocols such as MCP, A2A, OASF, and x402, as well as to build registries or discovery systems for agent services. Because metadata is externally hosted and optional, users should account for incomplete coverage when performing analyses that depend on descriptive fields or service declarations. Similarly, feedback data is sparse relative to the full agent population, reflecting the early-stage adoption of the reputation registry. All numerical values in the reputation layer are stored in raw integer form together with decimal precision. Users should convert these values appropriately depending on their analytical requirements. The dataset is anchored to a fixed observation block. Users interested in longitudinal dynamics or agent lifecycle evolution should reconstruct additional snapshots at different block heights using the provided codebase.
Data Availability The dataset described in this study is publicly available via the Harvard Dataverse repository. It includes all tables described in the Data Records section, including agents_core, agent_metadata, agent_services, agent_reputation_summary, agent_feedback_records, mint_economics, and a data dictionary file. The dataset is provided in machine-readable format CSV to facilitate reuse and integration into downstream analysis pipelines. The version reported in this study corresponds to a snapshot of Ethereum mainnet at block 24,839,925. A persistent identifier (DOI) is assigned to the dataset via Harvard Dataverse to ensure long-term accessibility and reproducibility. The dataset is available at: https://doi.org/10.7910/DVN/HJZW8Q
Code Availability All code used to collect, process, and validate the dataset is publicly available on GitHub. The repository includes scripts for data collection from Ethereum using Web3.py, metadata resolution, reputation extraction, and database construction. The codebase is organized as a reproducible pipeline with configurable parameters, including RPC endpoint, block range, and database connection settings. It supports re-execution to reconstruct the dataset from the same observation block. The code is available at: https://github.com/YulinLiu20/ERC8004
References 1. Zheng, Z., Xie, S., Dai, H., Chen, X. & Wang, H. An overview of blockchain technology: Architecture, consensus, and future trends. IEEE Int. Congr. Big Data 557–564 (2017). https://doi.org/10.1109/BigDataCongress.2017.85 2. Harvey, C. R., Ramachandran, A. & Santoro, J. Defi And The Future Of Finance. (John Wiley, 2021). 3. Liu, Y. & Zhang, L. Cryptocurrency valuation: An explainable ai approach. SSRN Electron. J. https://doi.org/10.2139/ssrn.3657986 (2021). 4. Zhang, Y., Chen, Z., Sun, Y., Liu, Y., & Zhang, L. (2023, July). Blockchain network analysis: A comparative study of decentralized banks. In Science and information conference (pp. 1022-1042). Cham: Springer Nature Switzerland. 5. Wang, R., Ye, F., Tang, S., Zhang, H., He, J., Zhang, X., & Xu, C. (2025). Blockchain Technology for Big-data Sharing in Material Genome Engineering. Scientific Data, 12(1), 1813. 6. Nguyen, T. L., Nguyen, L., Hoang, T., Bandara, D., Wang, Q., Lu, Q., ... & Chen, S. (2025). Blockchain-empowered trustworthy data sharing: Fundamentals, applications, and challenges. ACM Computing Surveys, 57(8), 1-36. 7. Nakamoto, S. Bitcoin: A peer-to-peer electronic cash system. Decentralized Bus. Rev. 21260, https://www.debr.io/article/21260-bitcoin-a-peer-to-peer-electronic-cas h-system (2008). 8. Böhme, R., Christin, N., Edelman, B. & Moore, T. Bitcoin: Economics, technology, and governance. J. Econ. Perspectives 29, 213–238, https://doi.org/10.1257/jep.29.2.213 (2015). 9. Subacchi, P. From gold to bitcoin and beyond. Nat. 597, 626–627, https://doi.org/10.1038/d41586-021-02615-2 (2021). 10. Yu, H., Sun, Y., Liu, Y., & Zhang, L. (2024, March). Bitcoin gold, litecoin silver: An introduction to cryptocurrency valuation and trading
strategy. In Future of Information and Communication Conference (pp. 573-586). Cham: Springer Nature Switzerland. 11. Pagnotta, E. S. Decentralizing money: Bitcoin prices and blockchain security. The Rev. Financial Stud, https://doi.org/10.1093/rfs/hhaa149 (2021). 12. Liu, Zhang, Zhao. Deciphering bitcoin blockchain data by cohort analysis. Sci. Data https://doi.org/10.1038/s41597-022-01254-0 (2022). 13. Liu, Y., Zhang, L. & Zhao, Y. Replication data for: “deciphering bitcoin blockchain data by cohort analysis”. Harv. Dataverse. https://doi.org/10.7910/DVN/XSZQWP (2021). 14. Kushwaha, S. S., Joshi, S., Singh, D., Kaur, M., & Lee, H. N. (2022). Ethereum smart contract analysis tools: A systematic review. Ieee Access, 10, 57037-57062. 15. John, K., Monnot, B., Mueller, P., Saleh, F. & Schwarz-Schilling, C. Economics of ethereum. J. Corp. Finance 91, 102718, https://doi.org/10.1016/j.jcorpfin.2024.102718 (2025). 16. Somin, S., Altshuler, Y. & Pentland, A. Crypto-asset trading on top of Ethereum Blockchain comprehensive dataset. Sci Data 12, 1407 (2025). https://doi.org/10.1038/s41597-025-05662-w 17. Liu et al. Empirical analysis of EIP-1559: Transaction fees, waiting times, and consensus security. CCS (2022). 18. Kim, J., & Im, I. (2023). Anthropomorphic response: Understanding interactions between humans and artificial intelligence agents. Computers in Human Behavior, 139, 107512. 19. Zhao, P., Jin, Z., & Cheng, N. (2023). An in-depth survey of large language model-based artificial intelligence agents. arXiv preprint arXiv:2309.14365. 20. Kühl, N., Schemmer, M., Goutier, M., & Satzger, G. (2022). Artificial intelligence and machine learning. Electronic Markets, 32(4), 2235-2244. 21. Hou, X., Zhao, Y., Wang, S., & Wang, H. (2025). Model context protocol (mcp): Landscape, security threats, and future research
directions. ACM Transactions on Software Engineering and Methodology. 22. Ray, P. P. (2025). A review on agent-to-agent protocol: Concept, state-of-the-art, challenges and future directions. Authorea Preprints. 23. De Rossi, M., Crapis, D., Ellis, J., & Reppel, E. (2025, August 13). EIP-8004: Trustless agents. Ethereum Improvement Proposals. https://eips.ethereum.org/EIPS/eip-8004 24. Entriken, W., Shirley, D., Evans, J., & Sachs, N. (2018, January 24). EIP-721: Non-Fungible Token Standard. Ethereum Improvement Proposals. https://eips.ethereum.org/EIPS/eip-721 25. John, K., Monnot, B., Mueller, P., Saleh, F., & Schwarz-Schilling, C. (2025). Economics of ethereum. Journal of Corporate Finance, 91, 102718. 26. Biais, B., Capponi, A., Cong, L. W., Gaur, V., & Giesecke, K. (2023). Advances in blockchain and crypto economics. Management Science, 69(11), 6417-6426. 27. Schilling, L., & Uhlig, H. (2019). Some simple bitcoin economics. Journal of Monetary Economics, 106, 16-26. 28. Liu, Z., Luong, N. C., Wang, W., Niyato, D., Wang, P., Liang, Y. C., & Kim, D. I. (2019). A survey on blockchain: A game theoretical perspective. IEEE Access, 7, 47615-47643.
Acknowledgement The author would like to thank all the attendees of the Swiss QuantEcon AI Workshop for their valuable feedback and comments.
Competing interests The author declares no competing interests.