ConceptioArchivearXiv CS
arXiv CSopen access

TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems Refat Ishrak Hemel, Ehsan Hallaji, and Roozbeh Razavi-Far

arXiv:2607.09528v1 [cs.LG] 10 Jul 2026

Trustworthy and Secure AI Lab (TSAI Lab), Faculty of Computer Science, University of New Brunswick, Canada {refatishrak.hemel, e.hallaji, roozbeh.razavi-far}@unb.ca Abstract—The emergence of metaverse platforms has created virtual economies that introduce new challenges related to fraud, bot activity, and illicit financial behavior. Despite growing interest in trustworthy metaverse analytics, existing datasets typically focus on user behavior, authentication, or financial transactions in isolation, limiting the development and reproducible evaluation of multimodal fraud detection methods. To address this gap, we present TSAI-MetaFraud, a multimodal, multi-task benchmark dataset for fraud analytics in virtual economies. TSAI-MetaFraud integrates behavioral, transactional, and graph-structured information while incorporating realistic fraud and automated bot scenarios. We define benchmark tasks including transaction fraud detection, cross-modal node classification, temporal link prediction, and weakly supervised fraud detection, and provide baseline evaluations using machine learning models and graph neural networks. By jointly capturing behavioral activity, financial interactions, and relational structure within a unified virtual economy, TSAI-MetaFraud provides a benchmark for advancing multimodal learning, graph mining, fraud analytics, and trustworthy AI in emerging metaverse ecosystems. Index Terms—Metaverse, fraud detection, graph mining, behavioral analytics, virtual economies.

I. I NTRODUCTION The rapid development of metaverse technologies is transforming virtual environments into persistent digital ecosystems where users socialize, collaborate, trade virtual assets, and participate in increasingly sophisticated economic activities. Platforms supporting virtual commerce, digital ownership, and avatar-based interactions are expected to play a significant role in future online economies [1]. As these environments evolve, concerns regarding trust, security, and financial integrity have become increasingly important [2], [3]. Similar to traditional financial systems, metaverse economies are vulnerable to fraudulent transactions, automated bot activity, account manipulation, identity abuse, and coordinated malicious behavior [3], [4]. Recent advances in machine learning, graph mining, and fraud analytics have demonstrated considerable success in detecting illicit activities in domains such as banking, ecommerce, and cryptocurrency networks [5], [6]. However, applying these techniques to metaverse environments remains challenging due to the scarcity of publicly available datasets. Existing fraud datasets primarily focus on transaction networks, while metaverse datasets often emphasize user interDataset available at: https://github.com/tsai-unb/MetaFraud.

actions, virtual environments, or visual content. Consequently, researchers lack realistic benchmarks that jointly capture user behavior, financial activity, and interaction networks within a unified virtual-world setting. The metaverse introduces unique characteristics that further distinguish it from conventional fraud detection domains. Financial transactions are closely intertwined with avatar behavior, social interactions, and movement patterns [3]. Malicious actors may exploit both behavioral and financial channels simultaneously, making it necessary to analyze heterogeneous information sources rather than relying on transactions alone [7]. Furthermore, the dynamic and evolving nature of virtual environments creates complex network structures and temporal dependencies that are not adequately represented in existing benchmarks. To address these limitations, we introduce TSAI-MetaFraud, a multimodal benchmark dataset collected from a comprehensive metaverse environment built using OpenSimulator. TSAIMetaFraud captures both behavioral and financial aspects of virtual-world activity through a combination of avatar interactions, transaction records, biometric attributes, and graphbased relationships. The dataset incorporates realistic benign and malicious behaviors, including automated bots, fraudulent transactions, and hybrid attack scenarios, providing a comprehensive testbed for fraud analytics and trustworthy AI research. Beyond releasing the dataset, we establish a suite of benchmark tasks designed to facilitate reproducible evaluation across multiple research communities. These tasks include heterogeneous fraud detection, cross-modal node classification, and temporal link prediction. We further provide baseline results using representative machine learning and Graph Neural Network (GNN) approaches to characterize the challenges posed by the dataset and establish reference performance levels. Hence, the primary contributions of this work are as follows: • We introduce TSAI-MetaFraud, a multimodal benchmark that jointly captures behavioral activity, financial transactions, graph-structured interactions, and fraud-related behaviors within a metaverse environment. • We define a suite of benchmark tasks spanning fraud detection, bot detection, temporal link prediction, and weakly supervised learning, enabling systematic evaluation of multimodal and graph-based methods.

TABLE I C OMPARISON OF TSAI-M ETA F RAUD WITH EXISTING METAVERSE

Dataset

Behavior

Transactions

Fraud Label

Bot Label

Metaverse

DATASETS AND RELATED BENCHMARKS FROM NEIGHBORING DOMAINS .

256-MetaverseRecords [8] MetaPed [9] MGAD [10] Deep Motion Masking [11] Metaverse Financial Transactions [12] Cresci-15 [13] Cresci-17 [14] TwiBot-22 [15] Game Bot Detection [16] Game-Contagion [17] TSAI-MetaFraud (Ours)

✓ ✓ ✓ ✓ – ✓ ✓ ✓ ✓ ✓ ✓

– – – – ✓ – – – – – ✓

– – – – ✓ – – – – – ✓

– – – – – ✓ ✓ ✓ ✓ ✓ ✓

✓ ✓ ✓ ✓ ✓ – – – – – ✓

from head and hand motion data. MetaData [4] further showed that adversarially designed virtual environments can infer sensitive user characteristics through behavioral responses during gameplay. These datasets have significantly advanced research on privacy, authentication, and behavioral analytics in virtual environments. However, they primarily focus on motion telemetry and user profiling, and do not contain financial transactions, fraud labels, or interaction networks required for studying security threats in virtual economies. Financially oriented metaverse datasets remain relatively scarce. The Metaverse Financial Transaction Dataset [12] provides transaction records designed for anomaly detection and fraud analysis in virtual economies. Nevertheless, it lacks peer-reviewed benchmark tasks and does not incorporate user behavior, avatar interactions, or relational information beyond transactions. B. Related Domains

We provide baseline results using machine learning and graph neural network models and publicly release the dataset and evaluation framework to support reproducible research. II. R ELATED W ORK

To position TSAI-MetaFraud within the current research landscape, this section reviews existing metaverse datasets, discusses related benchmarks from neighboring domains, and highlights the limitations of current resources that motivate the proposed benchmark. A. Metaverse and Virtual-World Datasets The growing interest in metaverse technologies has stimulated the creation of datasets spanning virtual economies, user behavior, immersive interactions, security, and privacy. However, existing resources typically focus on a single aspect of virtual environments and do not provide a unified benchmark for studying financial fraud, behavioral anomalies, and relational interactions simultaneously. One line of research focuses on virtual-world content and behavioral data. The 256-MetaverseRecords dataset [8] provides annotated recordings collected from multiple metaverse environments and supports tasks such as event detection, interaction recognition, and scene understanding. Similarly, MetaPed [9] employs metaverse simulations to generate synthetic pedestrian behaviors for autonomous driving applications. While these datasets demonstrate the potential of virtual environments for large-scale data generation, they primarily target computer vision tasks and do not contain transactional or security-related information. A second group of datasets investigates authentication, privacy, and user profiling in immersive environments. The Metaverse Gait Authentication Dataset [10] focuses on gait-based biometric authentication using synthetic avatar movements. Deep Motion Masking [11] introduced a dataset for studying privacy-preserving anonymization of VR motion telemetry, while Nair et al. released datasets demonstrating large-scale user identification [18] and personal attribute inference [19]

Although metaverse-specific datasets remain limited, several neighboring domains have contributed benchmarks that capture aspects of user behavior, bot activity, and fraud detection. In social networks, benchmark datasets such as Cresci15 [13] and Cresci-17 [14] have played a central role in advancing social bot detection by providing labeled human and automated accounts exhibiting varying levels of sophistication. More recently, TwiBot-22 [15] introduced a largescale graph-based benchmark that incorporates diverse user attributes, relationships, and interaction structures, facilitating the development of GNN approaches for bot detection. Related efforts have also emerged in online gaming environments. The Game Bot Detection Dataset [16] provides longterm behavioral logs collected from a large-scale MMORPG and includes both legitimate players and confirmed bots. Similarly, Game-Contagion [17] captures the diffusion of cheating behavior through social interaction networks, enabling the study of behavioral influence and fraud propagation in virtual communities. These datasets demonstrate the importance of combining behavioral signals with relational information when analyzing malicious activity in large-scale digital ecosystems. Beyond fraud analytics, behavioral datasets have been widely studied in the human activity recognition community [20]. These benchmarks typically rely on wearable or mobile sensors to classify user activities and have contributed substantially to behavior modeling and representation learning. However, unlike metaverse environments, they do not incorporate virtual interactions, economic transactions, or adversarial behaviors. Collectively, these neighboring domains provide valuable insights into behavioral analytics, bot detection, and graphbased learning. Nevertheless, none simultaneously capture avatar behavior, financial transactions, interaction networks, fraud labels, and bot activity within a metaverse economy. TSAI-MetaFraud addresses this gap by unifying these components into a single multimodal benchmark specifically designed for fraud analytics and trustworthy AI in virtual environments.

C. Research Gap Table I summarizes the characteristics of existing metaverse, social-network, and online-gaming datasets relevant to fraud analytics and behavioral modeling. While prior datasets provide valuable resources for studying individual aspects of virtual environments, such as user behavior, authentication, bot detection, or financial transactions, none simultaneously combine behavioral activity, financial transactions, fraud annotations, and bot labels within a metaverse setting. This limitation restricts the development and reproducible evaluation of multimodal methods that jointly leverage behavioral and transactional information for fraud analytics. TSAI-MetaFraud addresses this gap by providing a unified benchmark that integrates these components and supports multiple learning tasks, including fraud detection, bot detection, temporal link prediction, and weakly supervised inference within a realistic virtual economy. The use of a simulated environment is motivated by the limited availability of real-world metaverse fraud datasets, which are often restricted by privacy concerns, platform policies, and the scarcity of verified fraud labels. Simulation enables controlled generation of both benign and malicious behaviors while providing complete ground-truth annotations for reproducible evaluation. Similar simulation-based approaches have been widely adopted in virtual-world, cybersecurity, and autonomous-system research when real-world data are inaccessible or difficult to label. III. DATASET G ENERATION AND C OLLECTION This section describes the methodology used to construct TSAI-MetaFraud, including the virtual-world environment, data generation pipeline, participant modeling strategy, and post-processing procedures used to produce the final benchmark. A. Metaverse Environment TSAI-MetaFraud was generated within a virtual-world environment built using OpenSimulator (OpenSim), an opensource platform compatible with the Second Life protocol. OpenSim provides a persistent multi-user environment supporting avatar interactions, virtual asset ownership, in-world currency transfers, and user authentication, making it suitable for simulating realistic virtual economies. The virtual world was designed as a multi-region environment consisting of commercial, trading, and residential areas. Avatars could navigate the environment, interact with virtual objects, purchase items, and perform peer-to-peer financial transactions using the native currency system. These activities generated both behavioral and transactional traces that were recorded throughout the simulation. The simulated population consisted of 936 active avatars representing different behavioral and financial profiles. Benign users performed normal movements and legitimate financial transactions. Behavioral fraud accounts represented automated bots exhibiting synthetic interaction patterns and abnormal keystroke characteristics. Financial fraud accounts followed

realistic movement behaviors while participating in illicit financial activities such as transaction layering and structuring. Hybrid fraud accounts combined automated behavioral patterns with fraudulent financial activity. In addition, a subset of accounts was assigned to an unknown category, where labels were intentionally masked to emulate the uncertainty commonly encountered in real-world investigations. The final dataset contains 45 behavioral fraud accounts, 16 financial fraud accounts, 10 hybrid fraud accounts, and 400 unknown accounts, with the remaining avatars representing benign users. B. Data Generation Pipeline Behavioral and transactional data were collected through a multi-layer acquisition pipeline integrated directly into the OpenSimulator environment, as shown in Fig. 1. Avatar activities were monitored using in-world Linden Scripting Language scripts that recorded movement trajectories, object interactions, and user actions in real time. Simultaneously, server-side transaction logs captured all financial events, including timestamps, sender and receiver identifiers, transaction amounts, transaction types, and associated fees. To model realistic user behavior, avatar movement events were linked with keystroke timing characteristics. Human behavioral distributions were derived from the public KMT keystroke dynamics dataset, from which hold-time and flighttime statistics were extracted. Gaussian Kernel Density Estimators (KDEs) were then fitted to these timing distributions and used to generate realistic behavioral traces for benign avatars. In contrast, automated accounts were generated using specialized behavioral models designed to emulate varying levels of sophistication. Two categories of automated accounts were introduced. Naive bots followed rigid timing distributions exhibiting highly predictable interaction patterns, whereas stealth bots sampled from the human-derived KDE distributions while maintaining substantially reduced behavioral variance. This design creates meaningful overlap between benign and malicious behavior and prevents trivial separation of classes. Consequently, successful detection requires models to learn subtle behavioral characteristics rather than relying on simple timing thresholds. Server-side APIs connected the OpenSimulator database to the data processing framework, enabling automated extraction and synchronization of behavioral events, biometric characteristics, and transaction records. C. Ethical Considerations TSAI-MetaFraud is generated within a fully simulated environment and does not contain personally identifiable information, financial records, or sensitive user data. The behavioral timing distributions used to model human activity were derived from the publicly available KMT keystroke dynamics dataset, which was originally collected under established consent procedures. All financial activities within TSAI-MetaFraud utilize virtual currency and do not correspond to real-world monetary transactions.

Fig. 1. Architecture of the TSAI-MetaFraud data generation pipeline. Behavioral telemetry and financial transactions are collected from the OpenSimulator environment, processed through behavioral modeling and fraud synthesis modules, and transformed into graph and tabular benchmark representations.

D. Data Processing and Quality Assurance Following data generation, several preprocessing and quality-control procedures were applied. Inactive accounts that did not participate in behavioral interactions or financial transactions were removed to eliminate isolated nodes from the final graph representation. Missing behavioral information associated with transaction events was represented using zeropadded session-level features in the tabular representation and handled natively within graph-based learning frameworks. To support temporal analysis, the continuous data stream was partitioned into a sequence of uniform, non-overlapping chronological intervals. This enables capturing a complete activity cycle, while the resulting temporal segmentation enables dynamic graph construction and temporal prediction tasks. Finally, all features were validated to ensure consistency and realism. Timing variables were constrained to physically meaningful ranges, transaction records were verified for integrity, and graph connectivity was examined to prevent disconnected components and information leakage between temporal partitions. IV. DATASET D ESCRIPTION This section presents the structure, statistical properties, and key characteristics of the TSAI-MetaFraud benchmark. The dataset is designed to support research on fraud analytics, behavioral modeling, and graph-based learning within virtual economies. A. Data Representation To serve diverse research communities, TSAI-MetaFraud is distributed in a dual-format representation: a modular relational graph format for GNNs, and a unified flat tabular format for standard Machine Learning models. 1) Graph Dataset: The modular graph format represents avatars and sessions as entities, and transactions as directed relationships. This schema comprises four tables summarized in Table II: avatar attributes (node_avatars.csv), session biometrics (node_sessions.csv), transaction

edges (edge_transactions.csv), and avatar labels (node_avatars_classes.csv). 2) Tabular Dataset: To support conventional machine learning models without requiring graph-based processing, we additionally provide a unified tabular representation of the dataset. For every transaction edge, we construct a 24dimensional feature vector by merging the following features. 1) Transaction Features: Attributes including transaction value, fees, timestamp, type, and the target Fraud_Label (8 features). 2) Sender Avatar Features: Network centralities and metrics of the sending account (6 features). 3) Receiver Avatar Features: Network centralities and metrics of the receiving account (6 features). 4) Sender Session Biometrics: Keystroke and movement biometrics of the active login session under which the sender executed the payment (4 features). B. Statistical Overview Table III summarizes the main characteristics of the dataset. TSAI-MetaFraud contains 936 active avatars interacting within a simulated metaverse economy during the data collection period. During this interval, the environment generated 74,671 financial transactions and 230,490 behavioral interaction events, producing a rich multimodal benchmark that combines behavioral, transactional, and relational information. C. Distribution Analysis The dataset exhibits several distributions commonly observed in real-world fraud analytics settings. Fig. 2 presents the class distribution of the benchmark. Similar to many financial crime datasets, TSAI-MetaFraud is highly imbalanced, with benign users representing the majority of the population and fraudulent entities comprising only a small fraction of the network. Fig. 3 illustrates transaction flows between account categories. While benign and unknown entities account for the majority of transaction activity, fraudulent accounts participate in transactions involving multiple classes rather than

TABLE II ATTRIBUTE SCHEMA FOR AVATARS , SESSIONS , AND TRANSACTIONS IN THE BENCHMARK DATASET Description

Field Name

Data Type

Attribute Description

Avatar Nodes

Avatar_ID Account_Age_Days pagerank hits_hub hits_authority in_degree out_degree

String Float Float Float Float Float Float

Unique UUID identifying the avatar account. Lifetime of the avatar account (days). PageRank centrality representing overall transaction activity. HITS Hub score capturing outgoing transaction volume. HITS Authority score capturing incoming transaction volume. Normalized in-degree centrality. Normalized out-degree centrality.

Session Nodes

Session_ID Avatar_ID Session_Duration_Min Activity_Intensity Path_Regularity Action_Mean

String String Float Float Float Float

Unique identifier for the chronological session window. The avatar performing the login session. Life span of the session (minutes). Number of micro-actions executed per minute. Inverse variance of action intervals (path path regularity). Average action hold duration based on human KDE (seconds).

Transaction Edges

Transaction_ID Sender_Avatar_ID Receiver_Avatar_ID Timestamp Amount_Sent Transaction_Type Transaction_Fee Multi_Wallet_Indicator Fraud_Label Session_ID

String String String Float Float String Float Integer String String

Unique transaction edge transaction. ID of the sending account. ID of the receiving account. Raw epoch timestamp. Financial value of the transaction. Category of transaction (e.g. transfer, payment). Fees paid. Binary indicator (0/1) for multi-wallet aggregation. Ground truth transaction label. Active session under which the payment occurred.

Avatar_ID Class_Label

String String

Unique UUID identifying the avatar account. Ground truth avatar label.

Avatar Labels

TABLE III S UMMARY STATISTICS OF THE TSAI-M ETA F RAUD DATASET Statistic

Value

Active avatars Behavioral sessions Behavioral interactions Financial transactions Unique transacting nodes

936 936 230,490 74,671 936

forming isolated communities. This cross-class interaction pattern creates a more realistic fraud detection setting in which illicit behavior is embedded within normal economic activity, making detection more challenging than simple communitybased identification. Fig. 4 visualizes avatar movement patterns within the virtual environment. Activity is concentrated around commercial and trading regions, while peripheral areas remain comparatively sparse. These localized hotspots create realistic spatial heterogeneity and reflect the tendency of economic activity to cluster around specific areas of the virtual world. D. Network Structure Visualization To illustrate the relational structure of TSAI-MetaFraud, we provide visualizations of the transaction and interaction networks. Fig. 5(a) presents the transaction graph connecting sender and receiver accounts through financial transfers. Fig. 5(b) illustrates behavioral interactions among avatars, while Fig. 5(c) depicts the heterogeneous graph structure integrating avatar entities, behavioral sessions, and transaction relationships. These visualizations highlight the complex topology,

Fig. 2. Class distribution of nodes and edge labels, highlighting severe class imbalance.

sparsity, and heterogeneity of the benchmark and motivate the use of graph-based learning approaches. E. Data Characteristics TSAI-MetaFraud possesses several characteristics that make it a challenging benchmark for machine learning and data mining research. 1) Non-Stationarity: Behavioral patterns and transaction activity evolve throughout the simulation period, producing temporal distribution shifts that require models to operate under changing conditions. 2) User Heterogeneity: The virtual economy contains users with diverse behavioral profiles, transaction frequencies, and network positions. This results in a heterogeneous graph containing both highly connected hub accounts and sparsely connected participants.

tasks spanning fraud analytics, behavioral modeling, graph learning, and semi-supervised inference. Collectively, these tasks evaluate a model’s ability to detect illicit financial activities, identify stealthy automated accounts, forecast future transactions, and operate under realistic conditions of limited supervision. Table IV summarizes the inputs, prediction targets, and evaluation metrics associated with each task. A. Transaction Fraud Detection This task evaluates multiclass fraud detection under a strict inductive setting. Models are required to classify transaction edges involving previously unseen accounts using both transactional attributes and graph-derived structural information. The objective is to classify each transaction into one of four categories: normal transactions (Benign), Behavioral Fraud, Financial Fraud, or hybrid (Both), while generalizing to new participants that were not observed during training. Fig. 3. Transaction flow between account categories. Ribbon widths are proportional to the number of transactions exchanged between classes, illustrating interactions among benign, fraudulent, hybrid, and unlabeled entities.

Fig. 4. Spatial density heatmap showing avatar coordinates and active hotspots across the virtual world grid.

3) Graph Sparsity: Despite the large number of transactions, only a small subset of all possible avatar pairs interact directly. Consequently, the resulting transaction network exhibits sparsity characteristics similar to those observed in real financial systems. 4) Multimodality: The dataset combines behavioral session information with financial transaction data, enabling the study of methods that jointly leverage behavioral and relational evidence for fraud detection. 5) Relational Structure: Financial transactions are naturally represented as directed interactions between entities, allowing the dataset to support graph-based tasks such as node classification, link prediction, and fraud propagation analysis. V. B ENCHMARK TASKS To facilitate reproducible evaluation and comparison of future methods, TSAI-MetaFraud defines a suite of benchmark

B. Cross-Modal Node Classification This task focuses on identifying avatar profiles using multimodal behavioral and structural information. Session-level behavioral features (i.e., from login sessions) are combined with graph-derived account characteristics (centralities) to classify avatar nodes into one of the four categories: real humans (Benign), automated scripts (Behavioral Fraud), illicit financial accounts (Financial Fraud), and hybrid automated fraud accounts (Both). By evaluating models on stealthy bot profiles and illicit financial flows concurrently, this task measures a model’s ability to generalize across orthogonal data modalities. C. Temporal Link Prediction This task evaluates a model’s ability to forecast future transactions in an evolving financial network. Given a sequence of historical transaction graph snapshots, the objective is to predict whether a directed transaction will occur between two accounts in a future time interval. The task captures the temporal and non-stationary nature of activity within virtual economies. D. Weakly Supervised Fraud Detection This task simulates realistic fraud investigation scenarios in which only a small subset of entities is labeled. Models must leverage graph structure and behavioral information to infer the multiclass label (i.e., similar to the four defined labels of previous tasks) of a largely unlabeled population. By masking the majority of available labels (90% masked), the benchmark evaluates the ability of learning algorithms to operate effectively under severe label scarcity. VI. E VALUATION AND A NALYSIS To characterize the difficulty of the proposed benchmark and establish reference performance levels for future research, we evaluate a diverse set of baseline models spanning classical machine learning, graph representation learning, and graphbased label propagation. These baselines were selected to reflect common approaches used in fraud analytics, bot detection, and network mining.

Fig. 5. Network representations of TSAI-MetaFraud, including the transaction-to-transaction money flow graph, avatar interaction graph, and avatar–transaction graph. Together, these views illustrate the relational and heterogeneous structures underlying the benchmark. TABLE IV B ENCHMARK TASKS SUPPORTED BY TSAI-M ETA F RAUD . #

Task

Input Data

Prediction Target

1 2 3 4

Transaction fraud detection Cross-modal node classification Temporal link prediction Weakly supervised fraud detection

Transaction attributes and graph-based account features Features of behavioral session and account-level graph Sequence of temporal transaction graphs Partially labeled transaction graph

Benign, Behavioral Fraud, Financial Fraud, Both Benign, Behavioral Fraud, Financial Fraud, Both Future transaction formation probability Multiclass fraud labels of unlabeled accounts

TABLE V BASELINE MODELS USED FOR EVALUATING TSAI-M ETA F RAUD . Model

Category

Data

Logistic Regression Random Forest XGBoost GraphSAGE (Node Classifier) GraphSAGE (Link Predictor) Label Propagation

Linear Classification Tree-Based Learning Gradient Boosting Graph Neural Networks Temporal Graph Learning Graph Diffusion

Tabular Tabular Tabular Graph Graph Graph

A. Baseline Models The evaluated baselines are summarized in Table V. Logistic Regression, Random Forest [21], and XGBoost [22] are trained on the tabular representation of the dataset and serve as representative classical baselines. These methods provide insight into the predictive value of handcrafted behavioral and transactional features without explicitly modeling graph structure. To evaluate graph-based representation learning, we employ a heterogeneous GraphSAGE model operating on the multirelational transaction network [23]. The model propagates information across avatar and session entities through transactional and behavioral relationships, enabling node-level and edge-level prediction tasks. For temporal link prediction, a GraphSAGE encoder coupled with a link prediction head is applied to sequences of transaction graph snapshots. Finally, a standard Label Propagation algorithm is included as a nonparametric graph diffusion baseline for weakly supervised fraud detection.

B. Evaluation Protocol To prevent information leakage and better reflect realistic deployment scenarios, task-specific evaluation protocols are employed. For transaction fraud detection and behavioral bot detection, a strict inductive split is adopted in which unique avatars are partitioned into training (80%) and testing (20%) sets. Consequently, all test entities remain unseen during training. For temporal link prediction, transaction records are divided into chronological one-hour snapshots. Models are trained on the historical sequence of transaction snapshots preceding the final observation window, and evaluated on their ability to forecast links that emerge in the subsequent, unseen period. Negative examples are generated through random sampling of non-existent transaction pairs. For weakly supervised fraud detection, the inductive split is retained while additionally masking 90% of the available training labels, simulating realistic conditions in which only a small subset of suspicious accounts can be manually investigated. C. Experimental Setup To prevent information leakage and evaluate realistic coldstart performance, we adopt a strict inductive split based on unique Avatar_IDs. Specifically, 80% of avatars are assigned to the training set, while the remaining 20% are reserved for testing. Transaction edges are partitioned accordingly: the training set contains 34,253 transactions for which both sender and receiver belong to the training avatar set, whereas the inductive test set contains 8,593 transactions involving at least one previously unseen avatar. This protocol

ensures that models are evaluated on their ability to generalize to new participants rather than memorizing previously observed accounts. For classical machine learning baselines, hyperparameter selection is performed using five-fold crossvalidation on the training split. GNN models are trained for 100 epochs using cross-entropy loss. All experiments were conducted using Python. Graph-based models were implemented using PyTorch and PyTorch Geometric, while classical baselines were implemented using Scikit-Learn and XGBoost. Experiments were executed on a workstation equipped with an Apple Silicon processor and 16GB of unified memory. Given the highly imbalanced nature of fraud detection tasks, we report class-specific Precision, Recall, and F1score, together with overall Macro F1-score. Macro-averaged metrics are particularly important as they provide a balanced evaluation across both majority and minority classes.

TABLE VI C OMPARATIVE BASELINE PERFORMANCE ON THE INDUCTIVE TEST SET FOR TRANSACTION CLASSIFICATION (TASK 1) Model (Data)

Target Class

Precision

Recall

F1-Score

Logistic Regression (Tabular)

Behavioral Fraud Both (Hybrid) Financial Fraud Real (Benign)

0.81 0.00 0.02 0.96

0.69 0.00 0.20 0.62

0.75 0.00 0.03 0.76

Random Forest (Tabular)

Behavioral Fraud Both (Hybrid) Financial Fraud Real (Benign)

0.86 0.00 0.00 0.97

1.00 0.00 0.00 1.00

0.93 0.00 0.00 0.98

XGBoost (Tabular)

Behavioral Fraud Both (Hybrid) Financial Fraud Real (Benign)

0.81 0.00 0.00 0.97

0.69 0.00 0.00 0.99

0.75 0.00 0.00 0.98

GraphSAGE (Graph)

Behavioral Fraud Both (Hybrid) Financial Fraud Real (Benign)

0.34 0.55 0.37 0.93

0.49 0.62 0.41 0.88

0.40 0.58 0.39 0.90

D. Main Results The objective of the baseline evaluation is not to establish state-of-the-art performance, but rather to characterize the difficulty of the proposed benchmark and provide reference results for future research. The evaluated methods span classical tabular learning, graph representation learning, temporal link prediction, and semi-supervised graph inference. Table VI presents the results for transaction fraud detection (Task 1) under the strict inductive setting, where transactions involving previously unseen avatars are reserved for testing. The results reveal a clear distinction between behavioral and financial fraud detection. While Random Forest and XGBoost achieve strong performance on the Behavioral Fraud class, they completely fail to identify Financial Fraud and hybrid fraud transactions. In contrast, the heterogeneous GraphSAGE model achieves meaningful performance across all fraud categories, suggesting that graph structure provides critical information for detecting complex illicit transaction patterns. Table VII reports the results for cross-modal node classification (Task 2). Tree-based methods achieve the strongest overall performance on benign accounts (Real) and automated scripts (Behavioral Fraud), with Random Forest obtaining F1-scores of 0.98 and 0.90 respectively. This indicates that session-level behavioral features (e.g., average action hold durations, path regularity) provide highly discriminative signals for identifying basic automated anomalies. However, all tabular baselines struggle to identify Financial Fraud and hybrid nodes, as their malicious patterns are expressed through multi-hop transactional relationships which are unavailable to standard models. Table VIII summarizes the results for temporal link prediction (Task 3). Traditional topological heuristics such as Common Neighbors, Jaccard Coefficient, and Preferential Attachment perform poorly in forecasting future transactions. GraphSAGE substantially outperforms these methods, demonstrating the importance of learned structural representations for modeling dynamic transaction networks. Table IX presents the macro-averaged baseline performance on weakly supervised multiclass fraud propagation (Task 4)

TABLE VII C OMPARATIVE BASELINE PERFORMANCE ON THE INDUCTIVE TEST SET FOR AVATAR NODE CLASSIFICATION (TASK 2) Data

Model

Logistic Tabular Regression

Target Class

Precision

Recall

F1-Score

Real (Benign) Behavioral Fraud Financial Fraud Both (Hybrid)

0.95 0.89 0.00 0.50

0.66 0.89 0.00 0.50

0.78 0.89 0.00 0.50

Tabular

Random Forest

Real (Benign) Behavioral Fraud Financial Fraud Both (Hybrid)

0.97 0.82 0.00 0.00

1.00 1.00 0.00 0.00

0.98 0.90 0.00 0.00

Tabular

XGBoost

Real (Benign) Behavioral Fraud Financial Fraud Both (Hybrid)

0.96 0.80 0.00 0.00

1.00 0.89 0.00 0.00

0.98 0.84 0.00 0.00

GraphSAGE

Real (Benign) Behavioral Fraud Financial Fraud Both (Hybrid)

0.89 0.11 0.00 0.00

0.78 0.22 0.00 0.00

0.83 0.14 0.00 0.00

Graph

under extreme label scarcity, where 90% of training labels are masked. All methods experience a substantial performance degradation compared to the fully supervised setting. Interestingly, Label Propagation and GraphSAGE show limited effectiveness compared to Logistic Regression under severe label scarcity. This behavior is largely driven by the difficulty of identifying the highly imbalanced and extremely sparse minority classes (e.g., Financial Fraud and Both), for which only a handful of labeled examples are available during training. These results highlight the challenges of fraud detection under realistic low-label conditions and motivate the development of more robust label-efficient graph learning methods for metaverse security. E. Discussion and Insights The baseline results reveal several characteristics that make TSAI-MetaFraud a challenging benchmark for fraud analyt-

TABLE VIII C OMPARATIVE BASELINE PERFORMANCE ON DYNAMIC TEMPORAL LINK PREDICTION (TASK 3) Data

Model

AUROC

Average Precision

Graph Graph Graph Graph

Common Neighbors Jaccard Coefficient Preferential Attachment GraphSAGE

0.2358 0.1008 0.3470 0.6089

0.4051 0.3288 0.4611 0.6485

TABLE IX M ACRO - AVERAGED BASELINE PERFORMANCE ON WEAKLY- SUPERVISED MULTICLASS FRAUD PROPAGATION (TASK 4) Data

Model

Precision

Recall

F1-Score

Tabular Tabular Tabular Graph Graph

Logistic Regression Random Forest XGBoost Label Propagation GraphSAGE

0.4911 0.4423 0.4398 0.2176 0.2359

0.5071 0.4722 0.4722 0.2500 0.2405

0.4873 0.4566 0.4553 0.2327 0.2377

ics and graph learning. First, the results highlight a clear distinction between behavioral and financial fraud detection. Tree-based models such as Random Forest and XGBoost achieve strong performance on classifying benign users (Real) and automated scripts (Behavioral Fraud). This suggests that session-level behavioral features, including activity intensity and timing statistics, provide highly discriminative signals for identifying basic automated behavior. However, these same models fail to detect Financial Fraud and hybrid (Both) node categories, achieving zero F1-scores on both classes. This observation indicates that financial fraud patterns are primarily expressed through transaction structures and multihop interaction flows rather than local account attributes. Additionally, GraphSAGE node classifier struggles on Behavioral Fraud (0.14 F1-score) compared to tree baselines. This highlights an important cross-modal difficulty: while graph neural networks are well-suited for topological propagation, the standard message-passing mechanism tends to aggregate feature representations over the neighborhood, which smooths out the local, high-frequency keystroke timing anomalies that define automated behavior. The temporal link prediction results demonstrate the limitations of traditional topological heuristics in evolving virtual economies. Common Neighbors, Jaccard Coefficient, and Preferential Attachment achieve poor predictive performance, while GraphSAGE attains substantially higher AUROC and Average Precision scores. These results suggest that future transaction formation depends on complex structural and temporal dependencies that cannot be captured using static proximity measures alone. Finally, the weakly-supervised setting (Task 4) exposes the combined limits of graph diffusion and label scarcity. Under this condition, the rare classes are represented by only a single anchor node. Standard semi-supervised GNNs and Label Propagation struggle to propagate signals effectively from these single points without introducing massive noise,

resulting in macro F1-scores of approximately 0.23–0.24. This highlights the severe challenges posed by the long-tailed class distribution when supervision is sparse, establishing TSAIMetaFraud as a rigorous testbed for few-shot and weaklysupervised graph learning. F. Modality Analysis TSAI-MetaFraud combines two complementary sources of information: behavioral session characteristics and financial transaction networks. The baseline results provide insight into the relative contribution of each modality. Behavioral features prove highly effective for identifying automated scripts. Models trained on session-level characteristics are able to identify Behavioral Fraud and Real accounts with high accuracy, indicating that behavioral signals contain strong discriminative information. However, these same models struggle to identify structural financial fraud, suggesting that illicit transaction routing patterns are not easily observable from local behavioral features alone. Conversely, graph-based methods are better suited for capturing financial fraud and hybrid fraud activities. By leveraging transaction relationships and network structure, GraphSAGE achieves meaningful performance on fraud classes that are largely missed by tabular baselines. This demonstrates the importance of relational information when modeling illicit financial behavior. Taken together, these results suggest that behavioral and transactional modalities provide complementary information. Behavioral features facilitate bot detection, whereas graph structure is essential for identifying complex money-laundering patterns. Consequently, TSAIMetaFraud provides a valuable benchmark for evaluating multimodal learning approaches that jointly exploit behavioral and relational information. VII. D ISCUSSION Beyond the benchmark tasks presented in this work, TSAIMetaFraud enables several research directions in fraud analytics, graph learning, and trustworthy AI, while also presenting limitations that motivate future extensions of the dataset. A. Research Opportunities TSAI-MetaFraud enables several research directions at the intersection of graph mining, fraud analytics, and virtual-world security. 1) Multimodal Fraud Detection: The benchmark combines behavioral telemetry, transactional records, and relational network information within a unified framework. This creates opportunities for developing multimodal learning methods capable of jointly modeling user behavior and financial activity to identify sophisticated fraud patterns. 2) Dynamic and Evolving Virtual Economies: Unlike many existing fraud datasets that provide static snapshots, TSAIMetaFraud contains temporally evolving interactions and transaction networks. This supports research on temporal graph learning, online fraud detection, and adaptive learning methods capable of operating under non-stationary environments.

3) Learning Under Limited Supervision: The severe class imbalance and label scarcity exhibited by the benchmark reflect practical challenges encountered in fraud investigation. Consequently, TSAI-MetaFraud provides a testbed for semisupervised learning, graph-based label propagation, active learning, and label-efficient fraud detection techniques. 4) Benchmarking Multimodal Graph Learning: TSAIMetaFraud provides a unified benchmark for evaluating models that combine behavioral, transactional, and relational information. Such multimodal graph learning approaches have attracted growing interest in fraud detection, financial analytics, and trustworthy AI, yet suitable benchmark datasets remain limited. TSAI-MetaFraud enables systematic evaluation of feature fusion, representation learning, and GNN architectures under realistic conditions involving class imbalance, temporal evolution, and limited supervision. B. Limitations TSAI-MetaFraud represents a specific virtual-world ecosystem implemented within OpenSimulator and may not capture all behaviors and economic dynamics observed across different metaverse platforms. In addition, the benchmark focuses on behavioral bot activity, transaction anomalies inspired by classic money laundering typologies and cryptocurrency financial crimes, and therefore does not encompass every possible form of malicious behavior. Despite these limitations, TSAIMetaFraud provides a unique benchmark that combines behavioral telemetry, financial transactions, and graph structure within a unified framework, enabling the evaluation of fraud detection, graph learning, and multimodal analytics methods in emerging virtual economies. VIII. C ONCLUSION This paper introduced TSAI-MetaFraud, a multimodal benchmark dataset for fraud analytics in virtual economies. The dataset combines behavioral telemetry, financial transactions, and heterogeneous graph structures within a unified framework and is released in both graph and tabular formats. To support reproducible evaluation, we defined four benchmark tasks spanning transaction fraud detection, cross-modal node classification, temporal link prediction, and weakly supervised fraud detection. This dataset fills an important benchmarking gap by providing a unified platform for evaluating multimodal learning, graph mining, and fraud analytics methods in virtual economies, thereby supporting reproducible research in an emerging and rapidly evolving application domain. Baseline experiments demonstrate that the benchmark presents significant challenges arising from multimodality, graph structure, temporal dynamics, class imbalance, and limited supervision. We hope TSAI-MetaFraud will facilitate future research on graph learning, multimodal analytics, and security in emerging virtual-world ecosystems. R EFERENCES [1] H. Huawei, Z. Qinnan, L. Taotao, Y. Qinglin, Y. Zhaokang, W. Junhao, Z. Xiong, Z. Jianming, J. Wu, and Z. Zheng, “Economic systems in the metaverse: Basics, state of the art, and challenges,” ACM Computing Surveys, vol. 56, no. 4, 2023.

[2] Y. Huang, Y. J. Li, and Z. Cai, “Security and privacy in metaverse: A comprehensive survey,” Big Data Mining and Analytics, vol. 6, no. 2, pp. 234–247, 2023. [3] R. McAmis, B. Durak, M. Chase, K. Laine, F. Roesner, and T. Kohno, “Handling identity and fraud in the metaverse,” IEEE Security & Privacy, vol. 23, no. 1, pp. 27–37, 2025. [4] V. Nair, G. Munilla Garrido, D. Song, and J. O’Brien, “Exploring the privacy risks of adversarial VR game design,” Proceedings on Privacy Enhancing Technologies, vol. 2023, no. 4, p. 238–256, 2023. [5] S. Motie and B. Raahemi, “Financial fraud detection using graph neural networks: A systematic review,” Expert Systems with Applications, vol. 240, p. 122156, 2024. [6] D. Cheng, Y. Zou, S. Xiang, and C. Jiang, “Graph neural networks for financial fraud detection: a review,” Frontiers of Computer Science, vol. 19, no. 9, 2025. [7] A. Karapatakis, “Metaverse crimes in virtual (un)reality: Fraud and sexual offences under english law,” Journal of Economic Criminology, vol. 7, p. 100118, 2025. [8] P. Steinert, S. Wagenpfeil, I. Frommholz, and M. L. Hemmje, “256 metaverse records dataset,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, p. 4256–4263. [9] N. H. H. Aung, P. Sangwongngam, R. Jintamethasawat, and L. Wuttisittikulkij, “Fully synthetic pedestrian anomaly behavior dataset generation in metaverse for enhancing autonomous driving object detection,” IEEE Access, vol. 12, pp. 166 630–166 642, 2024. [10] S. Ravikanti, “Metaverse gait authentication dataset (MGAD),” 2025. [Online]. Available: https://dx.doi.org/10.21227/rvh5-8842 [11] V. Nair, W. Guo, J. F. O’Brien, L. Rosenberg, and D. Song, “Deep motion masking for secure, usable, and scalable real-time anonymization of ecological virtual reality motion data,” in 2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), 2024, pp. 493–500. [12] J. Ward, “Metaverse financial transaction dataset,” 2023. [Online]. Available: https://www.kaggle.com/datasets/faizaniftikharjanjua/ metaverse-financial-transactions-dataset [13] S. Cresci, R. Di Pietro, M. Petrocchi, A. Spognardi, and M. Tesconi, “Fame for sale: Efficient detection of fake twitter followers,” Decision Support Systems, vol. 80, pp. 56–71, 2015. [14] S. Cresci, R. D. Pietro, M. Petrocchi, A. Spognardi, and M. Tesconi, “Social fingerprinting: Detection of spambot groups through dna-inspired behavioral modeling,” IEEE Transactions on Dependable and Secure Computing, vol. 15, no. 4, pp. 561–576, 2018. [15] S. Feng, Z. Tan, H. Wan, N. Wang, Z. Chen, B. Zhang, Q. Zheng, W. Zhang, Z. Lei, S. Yang, X. Feng, Q. Zhang, H. Wang, Y. Liu, Y. Bai, H. Wang, Z. Cai, Y. Wang, L. Zheng, Z. Ma, J. Li, and M. Luo, “Twibot-22: towards graph-based twitter bot detection,” in Proceedings of the 36th International Conference on Neural Information Processing Systems, 2022. [16] A. R. Kang, S. H. Jeong, A. Mohaisen, and H. K. Kim, “Multimodal game bot detection using user behavioral characteristics,” SpringerPlus, vol. 5, no. 1, p. 523, 2016. [17] J. Woo, S. W. Kang, H. K. Kim, and J. Park, “Contagion of cheating behaviors in online social networks,” IEEE Access, vol. 6, pp. 29 098– 29 108, 2018. [18] V. Nair, W. Guo, J. Mattern, R. Wang, J. F. O’Brien, L. Rosenberg, and D. Song, “Unique identification of 50,000+ virtual reality users from head & hand motion data,” in 32nd USENIX Security Symposium (USENIX Security 23), 2023, pp. 895–910. [19] V. Nair, C. Rack, W. Guo, R. Wang, S. Li, B. Huang, A. Cull, J. F. O’Brien, M. Latoschik, L. Rosenberg, and D. Song, “Inferring private personal attributes of virtual reality users from ecologically valid head and hand motion data,” in IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), 2024, pp. 477–484. [20] Y. Woo, S. Hwang, S. Oh, M. Kang, S. Lee, J. Kim, J. Cha, and K. K. Kim, “Human activity recognition dataset for pedestrians with mobility disabilities,” Scientific Data, 2026. [21] L. Breiman, “Random forests,” Machine learning, vol. 45, no. 1, pp. 5–32, 2001. [22] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, p. 785–794. [23] W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, p. 1025–1035.

Record · ID 361404 · SHA-256 010ab59a59ac6b6f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.