MGT EVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
arXiv:2604.25152v1 [cs.CR] 28 Apr 2026
Yuanfan Li1, † , Qi Zhou1, † , Chengzhengxu Li1, † , Zhaohan Zhang2, † , Chenxu Zhao1, † , Zepu Ruan1, † , Chao Shen1 , Xiaoming Liu1, ∗ 1 Faculty of Electronic and Information Engineering, Xi’an Jiaotong University 2 Queen Mary University of London † Equal contribution, ∗ Corresponding author: [email protected] Project Website: http://uncoverai.cn Project Codebase: https://github.com/Liyuuuu111/MGT-Eval Demo Video: https://www.youtube.com/watch?v=1CVoGQFW4KU To address this gap, we develop MGT EVAL, Abstract a unified platform for systematically evaluating the capabilities of existing MGT detectors with a user-friendly graphical interface. MGT EVAL not only supports document detection using selected detectors, similar to commercial products such as GPTZero1 (see Figure 1 for demonstration), but also provides a fully customizable, full-cycle workflow for constructing and evaluating MGT detectors. Specifically, it consists of four components: Dataset Building, Dataset Attack, Detector Training, and Performance Evaluation. The users are allowed to customize their dataset for training and testing by uploading human-written text corpora and configuring the generator for generating corresponding MGT. To support robustness evaluation, MGT EVAL includes a Dataset Attack module which contains 12 common attack methods such as humanization (Wang et al., 2024a) and deletion (Kukich, 1992), for perturbing the test set and providing a comprehensive testbed for assessing detector performance. Detector Training module provides a configurable training pipeline for training the detectors. To simplify the training process, we train 26 existing detectors using the standardized datasets from the Dataset Building module and release as off-the-shelf models. In the Performance Evaluation component, we automatic the evaluation process and report the common metrics such as Accuracy, F1 score, AUROC, and TPR at FPR = 0.01 and 0.001. Users can customize detector evaluation through either the command-line interface or a web-based interface without rewriting any code. MGT EVAL adopts a modular design, in which all detector and attack modules are managed through a registry mechanism. This design greatly improves extensibility and maintainability, enabling users to seamlessly add new methods,
We present MGT EVAL, an extensible platform for systematic evaluation of MachineGenerated Text (MGT) detectors. Despite rapid progress in MGT detection, existing evaluations are often fragmented across datasets, preprocessing, attacks, and metrics, making results hard to compare and reproduce. MGT EVAL organizes the workflow into four components: Dataset Building, Dataset Attack, Detector Training, and Performance Evaluation. It supports constructing custom benchmarks by generating MGT with configurable LLMs, applying 12 text attacks to test sets, training detectors via a unified interface, and reporting effectiveness, robustness, and efficiency. The platform provides both command-line and Webbased interfaces for user-friendly experimentation without code rewriting.
1
Introduction
The rapid advancement of large language models (LLMs), such as ChatGPT (Achiam et al., 2023), LLaMA (Dubey et al., 2024), and DeepSeek (Liu et al., 2025a), has made it increasingly easy to generate fluent, human-like text at scale. Although this capability enables a wide range of beneficial applications, it also lowers the barrier to producing misleading or malicious content, including fabricated news and sophisticated phishing messages. To address these risks, researchers have developed various methods to detect machine-generated text (MGT) (Su et al., 2023; Hans et al., 2024; Bao et al., 2024; Liu et al., 2024; Li et al., 2025), aiming to distinguish LLM-generated content from human-written text and inform the users about the provenance of the textual contents. However, the emerging MGT detectors lack a consistent platform for fair comparison under the same setting, leading to difficulty in evaluating the performance and robustness of the detector.
1
1
https://gptzero.me/
1
4
2
3
Figure 1: Demo section of our MGT EVAL. Before detecting, users can select the detectors (subplot 1), and then select the models (subplot 2) and the parameters (subplot 3) used for detection. Then the users can input the text and run the detection (subplot 4), our detector will output the detection result (human-written or machine-generated) and the confidence of the result. Try our MGT EVAL in http://uncoverai.cn/.
reuse existing components, and configure evaluation pipelines flexibly without modifying the core framework. Our contributions are summarized as follows:
attack modules improves extensibility and maintainability, making it easy to integrate new methods and support future research.
2
• Unified Evaluation Pipeline. We present MGT EVAL, a unified framework for MGT detector evaluation that standardizes the full workflow, including dataset construction, dataset attack, detector training, and performance evaluation. By unifying the data preparation, training, and evaluation process, MGT EVAL provides a reproducible, configurable, and fair platform for comparison across detectors.
Related Work
Machine-Generated Text (MGT) Detectors. As LLMs continue to advance, a growing body of work has focused on detecting machine-generated text (MGT). Existing methods can be broadly categorized into metric-based and model-based detectors. Metric-based detectors (Su et al., 2023; Bao et al., 2024; Hans et al., 2024; Xu et al., 2024; Zhu et al., 2025) require little or no supervised training and identify MGT using statistical cues, but they are often sensitive to adversarial perturbations and may yield poorly calibrated scores, making threshold selection difficult in real-world deployment. By contrast, model-based detectors (Liu et al., 2022; Hu et al., 2023; Liu et al., 2024; Koike et al., 2024; Li et al., 2025) fine-tune text classifiers (e.g., BERT/RoBERTa) and generally achieve stronger detection performance, but they typically depend on substantial labeled data to reach high accuracy. MGT EVAL integrates a wide range of advanced metric-based and model-based detectors under a unified framework for standardized evaluation and comparison. Machine-Generated Text (MGT) Benchmarks. With the rapid development of MGT detection, there have recently been several efforts toward
• Comprehensive Benchmarking Suite. MGT EVAL supports systematic evaluation of 26 existing MGT detectors on both clean and attacked datasets, with comprehensive metrics such as Accuracy, F1, AUROC, and TPR at low FPR operating points. This design facilitates multi-dimensional analysis of detector performance, including effectiveness and robustness under text perturbations. • User-Friendly Extensible Framework. MGT EVAL provides both command-line and web-based interfaces, allowing users to customize evaluation pipelines without rewriting code. In addition, its modular, registry-based design for detector and 2
Section 3.3 Detector Training
Section 3.1 Dataset Building Human Corpora
Train/Val Dataset Insertion
Attack Methods
Paraphrasing
... ... Transposition
Section 3.2 Dataset Attack
Test Dataset
MGT Detector Section 3.4 Detector Evaluation
Evaluation Metrics ACC
F1 Score
AUROC
AUPR
ASR
Latency
TPR@FPR
GPU Peak
... ...
Figure 2: Pipeline of our MGT EVAL. Users can use human corpora and configurable LLMs to build dataset (Section 3.1), and choose different attacks to generate attacked test dataset (Section 3.2). Then users can use train/val dataset to train a detector (Section 3.3), test the detector in the test dataset, and obtain output metrics (Section 3.4).
Feature
configuration-driven design: all modules exchange a canonical record schema and save run manifests (configuration, random seed, and output artifacts), which reduces integration cost across new datasets and detectors.
MGTBench MGTBench 2.0 Stumbling Blocks MGT EVAL
# Detectors # Attacks # Built-in Metrics
13 3 –
12 0 –
8 12 3
26 12 ≥5
Custom Dataset Custom Attacks
% %
% %
% %
% %
Front-end UI CLI Support
% !
% !
% !
! !
3.1 Table 1: Comparison of features in MGTEval with existing MGT detectors evaluation framework.
The Dataset Building module standardizes heterogeneous corpora into a unified binary detection format and provides a reproducible pipeline for dataset loading and custom dataset construction.
building systematic evaluation frameworks for MGT detectors. Representative examples include MGTBench (He et al., 2024), MGTBench 2.0 (Liu et al., 2025b), and Stumbling Blocks (Wang et al., 2024a), which respectively advance benchmark standardization, academicdomain evaluation, and robustness stress testing under realistic attacks. These works provide valuable resources and evaluation protocols for analyzing detector performance from different perspectives. However, as shown in Table 1, existing frameworks still have limited support for unified end-to-end evaluation functionalities, such as integrated dataset construction, detector training, user-friendly front-end interaction, and comprehensive multi-metric reporting within a single system. In contrast, our MGT EVAL is designed as an integrated and extensible framework that unifies dataset construction, attack generation, detector training, and detector evaluation for systematic MGT detector benchmarking.
3
Dataset Building
Dataset Loading. MGT EVAL supports loading datasets from local files (.jsonl, .json, .csv), directory-level recursive discovery, and iterable in-memory objects. In addition to flat binary records, it natively handles common MGT benchmark schemas such as HC3-style QA records and paired generation records, enabling users to seamlessly ingest both benchmark corpora and custom collections. Supported dataset structures are summarized in Table 2, and MGT EVAL covers widely used benchmarks including HC3 (Guo et al., 2023), SemEval 2024 (Wang et al., 2024b), and M4 (Wang et al., 2024c). Custom Dataset Construction. Beyond directly loading existing datasets, MGT EVAL enables userdefined dataset construction by pairing humanlabeled corpora with machine-generated counterparts produced by configurable LLM backends. Given a corpus consisting of human-written texts, users can specify one or multiple LLM generators and fully configure generation parameters, including prompt templates, decoding temperature, topk/top-p sampling, maximum length, and random seeds, to synthesize machine-written samples. The resulting dataset is automatically labeled (label=0 for human, label=1 for machine) and annotated with generation metadata (e.g., model, prompt ID,
MGT EVAL Framework
In this section, we present the MGT EVAL framework. MGT EVAL consists of four components: Build Dataset, Attack Dataset, Train Detector, and Evaluate Detector. Our pipeline can be found in Figure 2. To make the pipeline extensible and reproducible, MGT EVAL adopts a registry-based, 3
Format Type
where α, β are learned by minimizing binary crossentropy:
Example Structure
Flat binary data
[{"text": "...", "label": 0/1, "id": "...", "source": "...", "lang": "...", "model": "..."}] HC3-style QA data [{"human_answers": ["..."], "chatgpt_answers": ["..."], "source": "..."}] Paired generation data {"original": [{"text": "..."}], "sample": [{"text": "..."}], "sampled": [{"text": "..."}], "rewritten": [{"text": "..."}]} Attack-aligned paired {"sample": [{"text": "...", "attack": "..."}], "meta": {"base_id": "...", "acdata tive_attack": "..."}} Standardized data {"id": "...", "text": "...", "label": 0/1, "source": "...", "lang": "...", "model": "...", "attack": "..."}
" ∗
(α , β ) = arg min − α,β
Dataset Attack
After dataset construction, Dataset Attack is used to build adversarial evaluation sets for robustness analysis. The key objective is to stress-test detectors under realistic evasion perturbations while preserving the human reference distribution. Given either (i) datasets generated by Dataset Building or (ii) any external dataset supported by the unified loader, MGT EVAL first aligns all records to the standard binary format. As is common practice in prior work (Wang et al., 2024a), it then attacks only machine-generated samples (label=1), while keeping human samples unchanged. This setting avoids unnecessary humantext distortion and ensures a clean robustness comparison between clean and attacked machine outputs. For each eligible machine sample, the module generates one or multiple attacked variants and stores them with provenance metadata (e.g., sample identity and attack type), so that paired and traceable robustness evaluation can be performed downstream.
(2) .
Model-based Detectors. For model-based detectors, MGT EVAL performs end-to-end supervised optimization on the unified binary dataset. These detectors directly output decision-ready logits/probabilities, so no separate score-to-probability fitting is required. The framework supports configurable optimization and data-scale settings, including backbone/model selection, learning rate, batch size, number of epochs, maximum sequence length, weight decay, warmup ratio, gradient accumulation, precision/device options, train/validation/test split ratios, and explicit subsampling for controlled databudget experiments. Detectors that are inferenceonly (training-free) are skipped in this stage and evaluated directly in the next stage. 3.4
Performance Evaluation
After training (or direct loading for training-free detectors), MGT EVAL evaluates detectors on heldout test sets, and optionally on attacked test sets generated by Dataset Attack. We organize evaluation into three aspects: effectiveness, robustness, and efficiency. For protocol consistency, threshold-dependent metrics use a fixed threshold determined during the training/calibration stage, and the same threshold is reused for both clean and attacked evaluation of the same detector. This design ensures fair comparison across settings and avoids threshold re-tuning on attacked data.
Detector Training
Using the aligned training split from Build Dataset, MGT EVAL supports both metric-based and modelbased detector training under a unified interface. Metric-based Detectors. For metric-based detectors, the model output is a scalar score rather than a calibrated probability. Therefore, MGT E VAL treats calibration as the training step for this detector family. Given a training set {(xi , yi )}N i=1 , where yi ∈ {0, 1}, we first compute detector scores si = g(xi ). We then fit a one-dimensional logistic regression to map scores into probabilities: pi = σ(αsi + β) =
#
This mapping enables unified thresholding and fair comparison across heterogeneous metric detectors. MGT EVAL also exposes detector-specific controls, such as scoring model choice, perturbation settings (e.g., perturbation number/strength), calibration dataset path, threshold policy, and training sample size (e.g., sample_k, random seed).
decoding settings), facilitating controlled benchmarking across generators and decoding regimes.
3.3
N 1 X yi log pi N i=1
+ (1 − yi ) log(1 − pi )
Table 2: Supported dataset structures in MGT EVAL.
3.2
∗
Effectiveness Metrics To measure detection quality on clean test sets (and optionally attacked test sets), MGT EVAL reports a set of standard classification and ranking metrics. Specifically, we report Accuracy (ACC) for overall correctness, Precision, Recall, and F1 for positive-class detection quality, and threshold-free ranking metrics
1 , (1) 1 + exp(−(αsi + β)) 4
including AUROC and AUPR. To characterize performance under strict deployment constraints, we also report low-false-positive operating-point metrics such as TPR@FPR=α (e.g., α = 0.01 and 0.001), which are particularly important in realworld moderation and integrity scenarios.
put path controls, and synchronizes with configuration files used by the CLI. During execution, users can monitor real-time logs, job status, and resourcerelated metadata, and directly access exported artifacts such as JSON/CSV reports and plots. In addition to the core pipeline pages, the WebUI includes an extra Demo section (Figure 1) for interactive single- or small-batch testing, where users can quickly inspect detector outputs and compare clean versus attacked text behavior before launching fullscale experiments. More screenshots of interface can be found in Appendix A.
Robustness Metrics To evaluate robustness against evasive modifications, MGT EVAL supports attacked test sets generated by the Attack Dataset module and reports robustness-oriented metrics on these sets. In particular, we use Attack Success Rate (ASR) to quantify how often attacks can flip detector decisions on samples that were originally classified correctly in the clean setting. This paired evaluation protocol enables traceable robustness analysis across attack types and detector families.
5
Experiment Setting. We evaluate MGT EVAL on a clean held-out test set constructed by the Build Dataset module. Specifically, human-written texts are sampled from the SemEval 2024 dataset (Wang et al., 2024b), while machine-generated texts are produced by Qwen3-4B-2507-Instruct (Yang et al., 2025). The final dataset contains 2,000 samples in total, with a balanced human-to-machine ratio of 1:1. We split the dataset into training/validation/test sets with a ratio of 8:1:1 and report results on the test split. For detector implementation consistency in MGT EVAL, modelbased detectors use RoBERTa-Base (Liu et al., 2019) as the backbone, while metric-based detectors use GPT-Neo-125M (Gao et al., 2020) as the scoring model and t5-small (Raffel et al., 2020) as the perturbation model (for methods requiring perturbation-based scoring). All detectors are evaluated under the same pipeline in MGT EVAL, and we report effectiveness, robustness, and efficiency metrics in Table 3. Experiment Results. We report the performance of 26 detectors in Table 3 and draw the following observations. (i) Strong overall performance under a unified protocol. Longformer achieves the best results on all major effectiveness and low-FPR robustness metrics, reaching 99.50% Accuracy, 99.50% F1, 99.99% AUROC, 100.00% TPR@FPR=0.01, and 99.00% TPR@FPR=0.001. Compared with the second-best detector in Accuracy/F1 (RAiDAr), it improves by +1.50 pp Accuracy and +1.52 pp F1, indicating a clear margin under the same data and evaluation settings. (ii) Clear efficiency–effectiveness trade-offs across detector families. While Longformer provides the strongest overall detection quality, it is not the most efficient method. The fastest detector is
Efficiency Metrics To measure practical usability, MGT EVAL reports runtime-oriented metrics during detector evaluation. These include evaluation time (end-to-end runtime for processing a test set), throughput (samples processed per second), and latency per sample (average inference cost per sample). Together, these metrics help users compare not only detection quality but also deployment efficiency under different detector architectures and configurations. Output and Reporting Beyond scalar summary metrics, MGT EVAL can export structured evaluation results for downstream analysis, including detector-level summaries and (when available) sample-level predictions with metadata such as source, language, generator, and attack type. This supports grouped/sliced evaluation and facilitates reproducible comparison across datasets, attacks, and detectors.
4
Experiments and Results
Interface Design
The MGT EVAL WebUI is designed as a taskoriented control panel that mirrors the end-to-end workflow of the framework while minimizing configuration friction for non-expert users, an overview can be found in Figure 2. The interface organizes core operations into four pages consistent with the backend pipeline: Build Dataset (Figure 2 Part 1), Attack Dataset (Figure 2 Part 2), Train Detector (Figure 2 Part 3), and Evaluate Detector (Figure 2 Part 4). Each page provides structured parameter forms with default presets, inline help, and validation to prevent invalid runs. To support reproducibility, the UI exposes seed, data path, and out5
Detector Binoculars (Hans et al., 2024) CoCo (Liu et al., 2022) DetectGPT (Mitchell et al., 2023) DeTeCtive (Guo et al., 2024) DNA-DetectLLM (Zhu et al., 2025) DNA-GPT (Yang et al., 2023) Entropy (Gehrmann et al., 2019) Fast-DetectGPT (Bao et al., 2024) GLTR (Gehrmann et al., 2019) GREATER (Li et al., 2025) RoBERTa-Base (Liu et al., 2019) LASTDE (Xu et al., 2024) LASTDE++ (Xu et al., 2024) Likelihood (Gehrmann et al., 2019) LogRank (Gehrmann et al., 2019) Longformer (Beltagy et al., 2020) LRR (Su et al., 2023) MPU (Tian et al., 2023) NPR (Su et al., 2023) OpenAI Detector PECOLA (Liu et al., 2024) RADAR (Hu et al., 2023) RAiDAr (Mao et al., 2024) Rank (Gehrmann et al., 2019) SimpleAI Detector (Guo et al., 2023) TOCSIN (Ma and Wang, 2024)
Accuracy (%, ↑) F1 (%, ↑) AUROC (%, ↑) TPR@FPR=0.01 (%, ↑) TPR@FPR=0.001 (%, ↑) Avg Latency (ms/sample) (↓) GPU Peak (GiB) (↓) 92.00 97.50 50.00 70.50 89.50 79.00 72.50 90.50 90.00 97.50 93.50 88.50 91.50 87.00 88.00 99.50 85.50 90.50 50.00 73.50 94.50 26.50 98.00 77.50 56.00 59.00
91.75 97.56 66.44 67.76 89.01 80.00 69.61 90.26 90.20 97.46 93.90 88.21 91.10 86.87 87.88 99.50 84.82 91.32 66.67 65.81 94.74 29.67 97.98 78.47 27.87 70.50
95.84 99.95 62.99 80.69 95.42 88.83 18.02 95.61 94.01 99.92 99.72 94.94 96.24 94.06 94.73 99.99 95.13 99.88 80.04 81.25 99.69 15.65 99.63 85.69 60.99 95.03
83.00 99.00 1.00 23.00 80.00 23.00 0.00 85.00 20.00 96.00 87.00 32.00 84.00 30.00 36.00 100.00 47.00 98.00 3.00 41.00 99.00 0.00 97.00 33.00 12.00 79.00
81.00 98.00 1.00 23.00 80.00 19.00 0.00 80.00 1.00 96.00 87.00 3.00 84.00 0.00 0.00 99.00 43.00 98.00 0.00 39.00 60.00 0.00 86.00 5.00 9.00 0.00
21.85 73.13 585.28 42.67 31.03 295.89 16.80 13.64 26.69 7.01 6.91 14.99 63.10 14.59 18.03 20.29 18.86 6.87 1132.42 959.82 6.86 2067.45 427.40 14.87 611.88 110.73
3.64 0.63 2.92 1.82 3.26 5.59 0.50 1.04 1.01 0.14 0.32 0.41 1.19 0.41 1.03 0.83 1.01 0.32 0.96 0.32 0.32 0.76 14.99 1.03 0.32 3.73
Table 3: Detector performance summary. Avg Latency denotes the average detection time per sample during evaluation, measured in milliseconds (ms/sample), and GPU Peak denotes the peak GPU memory consumption during evaluation (GiB). Best values are in bold and second-best values are underlined.
6
PECoLA (6.86 ms/sample), followed closely by MPU (6.87 ms/sample), RoBERTa-Base (6.91 ms/sample), and GREATER (7.01 ms/sample). Among these efficient detectors, GREATER is particularly attractive because it combines near-SOTA performance (97.50% Accuracy, 99.92% AUROC, 96.00%/96.00% TPR at FPR=0.01/0.001) with the lowest GPU peak memory usage (0.14 GiB), substantially lower than the second-best memory tier (0.32 GiB). In contrast, some detectors (e.g., RAiDAr) achieve strong accuracy but incur much higher computational cost (427.40 ms/sample and 14.99 GiB GPU peak), highlighting the importance of reporting efficiency metrics together with detection quality. (iii) Multi-metric evaluation is necessary to reveal calibration and threshold sensitivity. Several detectors exhibit large discrepancies between threshold-free and threshold-dependent metrics. For example, TOCSIN obtains a strong AUROC (95.03%) but much lower Accuracy (59.00%) and a TPR@FPR=0.001 of 0.00%, suggesting sensitivity to thresholding and operating-point selection. Similarly, DetectGPT and NPR are substantially slower (585.28 and 1132.42 ms/sample, respectively) while delivering limited Accuracy gains (both 50.00% Accuracy in this setting). We also observe severe degradation for some methods (e.g., RADAR and Entropy), which further demonstrates the value of MGT EVAL as a unified benchmark for systematically comparing detector effectiveness, robustness, and efficiency under identical protocols.
Conclusion
We introduced MGT EVAL, a unified framework for systematic evaluation of machine-generated text detectors. Unlike prior benchmark efforts that focus on specific datasets, tasks, or robustness settings, MGT EVAL provides an end-to-end evaluation pipeline that unifies dataset construction, attack generation, detector training/calibration, and detector evaluation under a consistent protocol. The framework integrates 26 detectors and 12 attacks, supports comprehensive metrics covering effectiveness, robustness, and efficiency, and offers both CLI and WebUI interfaces to improve usability and reproducibility. MGT EVAL is designed as a modular, registry-based infrastructure to facilitate future extensions. In future work, we plan to expand support for multilingual and cross-domain benchmarks, richer attack compositions, attribution-oriented tasks, and broader detector families (including newer LLM- and API-based detectors), further advancing standardized and reproducible MGT detection evaluation.
Limitation Despite the practical utility and comprehensive design of MGT EVAL for systematic MGT detector evaluation, it still exhibits several limitations: i) Although MGT EVAL unifies dataset construction, attack generation, detector training, and evaluation under a single framework, running the full pipeline (especially for multiple detectors and attacks) can be computationally expensive and time-consuming, 6
which may limit usability in resource-constrained environments. ii) While MGT EVAL integrates a broad set of detectors and attacks through a registrybased modular design, some detector-specific implementation details or original hyperparameter settings may not be fully preserved under a unified protocol, which can lead to discrepancies from results reported in the original papers. iii) The current version of MGT EVAL mainly focuses on binary MGT detection (human vs. machine) and a fixed set of text-level attacks. As a result, it does not yet fully cover more complex settings, such as fine-grained source attribution, multilingual large-scale benchmarking, or adaptive attack–defense co-evaluation, which may constrain its general applicability in broader real-world scenarios.
ing clear documentation, controlled benchmarking settings, and avoiding claims that attacks guarantee bypass in real systems. Risk of false positives and deployment harm. MGT detectors can make incorrect predictions, including false positives on human-written text and false negatives on machine-generated text. These errors may cause harm if detector outputs are used as sole evidence in high-stakes decisions (e.g., academic penalties, moderation sanctions, or fraud accusations). MGT EVAL is an evaluation framework, not a decision policy. We recommend that detector outputs be used as one signal among others, with human oversight in sensitive applications. Data and privacy considerations. MGT EVAL is designed to operate on user-provided or publicly available corpora. Users remain responsible for ensuring that the datasets they load, generate, or export comply with applicable licenses, privacy requirements, and institutional policies. When evaluating on sensitive data, users should adopt appropriate anonymization and access-control practices.
Acknowledgment This work is supported by National Natural Science Foundation of China (62272371, 62103323) and Fundamental Research Funds for the Central Universities under grant (xzy012024144, xzy012025043). The author Xiaoming Liu gratefully acknowledges the support of K. C. Wong Education Foundation.
Bias and generalization. Detector performance may vary across languages, domains, and writing styles, and benchmarking results may reflect biases in the chosen datasets and generators. We encourage users to evaluate detectors on diverse datasets and report subgroup performance whenever possible. Overall, we view MGT EVAL as infrastructure for more transparent and reproducible MGT detector evaluation, and we encourage its use in ways that prioritize safety, fairness, and responsible deployment.
Ethics Statement This work presents MGT EVAL, a benchmarking and evaluation framework for machine-generated text (MGT) detectors. Our goal is to improve the rigor, reproducibility, and transparency of detector evaluation, which may support safer deployment of MGT detection systems in applications such as content moderation, platform integrity, and academic integrity analysis. Potential benefits. By standardizing dataset construction, attack generation, training/calibration, and evaluation, MGT EVAL reduces protocol inconsistency and helps the community identify detector strengths and failure modes more reliably. We believe this can improve scientific comparability and encourage more robust detector development.
References Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2024. Fast-detectgpt: Efficient zero-shot detection of machine-generated text via conditional probability curvature. In ICLR.
Dual-use risks. The framework includes attackgeneration functionality (Attack Dataset) for robustness evaluation. Such functionality could potentially be misused to study detector weaknesses for evasion purposes. We include attacks for defensive benchmarking and stress testing, not to facilitate harmful misuse. In practice, we recommend responsible release and use of attack modules, includ-
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman,
7
Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, and 1 others. 2025a. Deepseek-v3. 2: Pushing the frontier of open large language models. arXiv preprint arXiv:2512.02556.
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, and 1 others. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
Shengchao Liu, Xiaoming Liu, Yichen Wang, Zehua Cheng, Chengzhengxu Li, Zhaohan Zhang, Yu Lan, and Chao Shen. 2024. Does detectgpt fully utilize perturbation? bridging selective perturbation to finetuned contrastive learning detector would be better. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1874–1889.
Sebastian Gehrmann, SEAS Harvard, Hendrik Strobelt, and Alexander M Rush. 2019. Gltr: Statistical detection and visualization of generated text. ACL 2019, page 111.
Xiaoming Liu, Zhaohan Zhang, Yichen Wang, Hang Pu, Yu Lan, and Chao Shen. 2022. Coco: Coherenceenhanced machine-generated text detection under data limitation with contrastive learning. arXiv preprint arXiv:2212.10341.
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv preprint arXiv:2301.07597.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
Xun Guo, Yongxin He, Shan Zhang, Ting Zhang, Wanquan Feng, Haibin Huang, and Chongyang Ma. 2024. Detective: Detecting ai-generated text via multi-level contrastive learning. Advances in Neural Information Processing Systems, 37:88320–88347.
Yule Liu, Zhiyuan Zhong, Yifan Liao, Zhen Sun, Jingyi Zheng, Jiaheng Wei, Qingyuan Gong, Fenghua Tong, Yang Chen, Yang Zhang, and 1 others. 2025b. On the generalization and adaptation ability of machinegenerated text detectors in academic writing. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 5674–5685.
Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Spotting llms with binoculars: zero-shot detection of machine-generated text. In Proceedings of the 41st International Conference on Machine Learning, pages 17519–17537.
Shixuan Ma and Quan Wang. 2024. Zero-shot detection of llm-generated text using token cohesiveness. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17538–17553.
Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2024. Mgtbench: Benchmarking machine-generated text detection. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 2251–2265.
Chengzhi Mao, Carl Vondrick, Hao Wang, and Junfeng Yang. 2024. Raidar: generative ai detection via rewriting. arXiv preprint arXiv:2401.12970.
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2023. Radar: Robust ai-text detection via adversarial learning. Advances in neural information processing systems, 36:15077–15095.
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023. Detectgpt: Zero-shot machine-generated text detection using probability curvature. In International Conference on Machine Learning, pages 24950–24962. PMLR.
Ryuto Koike, Masahiro Kaneko, and Naoaki Okazaki. 2024. Outfox: Llm-generated essay detection through in-context learning with adversarially generated examples. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 21258–21266. Karen Kukich. 1992. Techniques for automatically correcting words in text. ACM computing surveys (CSUR), 24(4):377–439.
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
Yuanfan Li, Zhaohan Zhang, Chengzhengxu Li, Chao Shen, and Xiaoming Liu. 2025. Iron sharpens iron: Defending against attacks in machine-generated text detection with adversarial training. arXiv preprint arXiv:2502.12734.
Jinyan Su, Terry Zhuo, Di Wang, and Preslav Nakov. 2023. Detectllm: Leveraging log rank information for zero-shot detection of machine-generated text. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12395–12412.
8
Yuchuan Tian, Hanting Chen, Xutao Wang, Zheyuan Bai, Qinghua Zhang, Ruifeng Li, Chao Xu, and Yunhe Wang. 2023. Multiscale positive-unlabeled detection of ai-generated texts. arXiv preprint arXiv:2305.18149. Yichen Wang, Shangbin Feng, Abe Bohan Hou, Xiao Pu, Chao Shen, Xiaoming Liu, Yulia Tsvetkov, and Tianxing He. 2024a. Stumbling blocks: Stress testing the robustness of machine-generated text detectors under attacks. arXiv preprint arXiv:2402.11638. Yuxia Wang, Jonibek Mansurov, Petar Ivanov, jinyan su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold, Chenxi Whitehouse, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, and Preslav Nakov. 2024b. Semeval-2024 task 8: Multidomain, multimodel and multilingual machine-generated text detection. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), pages 2041–2063, Mexico City, Mexico. Association for Computational Linguistics. Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, and Preslav Nakov. 2024c. M4: Multi-generator, multi-domain, and multi-lingual black-box machine-generated text detection. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1369–1407, St. Julian’s, Malta. Association for Computational Linguistics. Yihuai Xu, Yongwei Wang, Yifei Bi, Huangsen Cao, Zhouhan Lin, Yu Zhao, and Fei Wu. 2024. Training-free llm-generated text detection by mining token probability sequences. arXiv preprint arXiv:2410.06072. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Xianjun Yang, Wei Cheng, Yue Wu, Linda Petzold, William Yang Wang, and Haifeng Chen. 2023. Dnagpt: Divergent n-gram analysis for training-free detection of gpt-generated text. arXiv preprint arXiv:2305.17359. Xiaowei Zhu, Yubing Ren, Fang Fang, Qingfeng Tan, Shi Wang, and Yanan Cao. 2025. Dna-detectllm: Unveiling ai-generated text via a dna-inspired mutationrepair paradigm. arXiv preprint arXiv:2509.15550.
9
A
More Screenshots of MGT EVAL
applying attacks to create an attacked test set. In the training page (Figure 5), users can choose datasets and supported MGT detectors to train a model, while the evaluation page (Figure 6) reports holistic metrics for assessing performance.
Figure 4–Figure 6 provides additional screenshots of MGT EVAL. Users can build custom datasets by uploading human-written texts and selecting an LLM to generate MGT (Figure 4), optionally
1
3
2
Figure 3: The Dataset Building Page. The users are allowed to specify the input path for human-written texts and the output directory where the constructed dataset will be saved (subplot 1). Users can also select the LLM used to generate machine-generated texts (MGTs) that mimic the uploaded human samples (subplot 2). This page also provides additional configurable options, including the LLM temperature, maximum output tokens, and sampling hyperparameters such as Top-k and Top-p. (subplot 3)
1
4
2
3
Figure 4: The Dataset Attack Page. Users are allowed to select the dataset to be attacked as the Input Data and specify the output directory for the attacked dataset (subplot 1). Users apply multiple attack methods to the dataset by selecting the supported attacks (subplot 2). For each selected attack, users adjust method-specific parameters, for example, controlling the proportion of tokens to be modified (subplot 3). To facilitate informed configuration, the page also presents descriptions and usage examples for all available attack methods (subplot 4).
10
1
4
2
3
Figure 5: The Detector Training Page. Users are allowed to select the detector to train from the available options (subplot 1) and access a concise summary of its metadata, including a high-level description, the corresponding paper, and its publication venue (subplot 2). The interface further allows users to configure training-related settings, such as the choice of training dataset and the number of samples to be used (subplot 3). After training, the system presents the training results, including evaluation accuracy on the validation set and, when available, on the test set, as well as the loss trajectory recorded throughout the training process (subplot 4).
1
4
2
3
Figure 6: The Performance Evaluation Page. Users are allowed to select a detector to evaluate from the available options (subplot 1) and choose the existing evaluation dataset and checkpoint to be used (subplot 2). The interface also allows configuration of evaluation parameters, such as batch size and random seed (subplot 3). Once the evaluation is completed, the system presents a comprehensive set of results, including Accuracy, F1 score, AUROC, AUPR, ASR, the confusion matrix, and additional metrics (subplot 4).
11