JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
1
Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning
Abstract—With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic face detection methods have achieved significant progress, they suffer from the inherent limitation of overconfidence due to their reliance on the Softmax activation function. Thus, these methods often lead to unreliable predictions when encountering unknown Out-ofDistribution (OOD) images, and cannot ascertain the model’s uncertainty in its prediction. Meanwhile, most existing methods require massive high-quality annotated data, which greatly limits their practicability across diverse scenarios. To address these limitations, we propose EMSFD (Evidence-based decision Modeling for Synthetic Face Detection with uncertainty-driven active learning), an approach designed to enhance detection reliability and generalizability. Specifically, EMSFD models class evidence using the Dirichlet distribution and explicitly incorporates model uncertainty into the prediction process. Furthermore, during training, the estimated uncertainty is exploited to prioritize more informative samples from the unlabeled pool for annotation, thereby reducing labeling cost and improving model generalization. Extensive experimental evaluations demonstrate that our method enhances the interpretability of synthetic face detection. Meanwhile, our method yields a 15% increase in accuracy compared to existing state-of-the-art (SOTA) baselines, which demonstrates the superior detection performance and generalizability of our approach. Our code is available at: https://github.com/hzx111621/EMSFD. Index Terms—Synthetic Face Detection, Evidential Deep Learning, Active Learning.
I. I NTRODUCTION Synthetic face detection is a critical task [1] aimed at identifying the authenticity of images found on social media. The rapid proliferation of generative models [2] has significantly blurred the boundary between real and synthetic imagery, This work was supported by the National Natural Science Foundation of China Under Grant 62402182, 62322309, 62572125. Natural Science Foundation of Shanghai under Grant 25ZR1401019. The Shanghai Explorer Program under Grant 24TS1411700, and the Open Research Fund of The State Key Laboratory of Blockchain and Data Security, Zhejiang University. Qingchao Jiang is with the Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, East China University of Science and Technology, Shanghai 200237, China (e-mail: [email protected]) Zhenxuan Hou and Zhiying Zhu are with the School of Information Science and Engineering, East China University of Science and Technology, Shanghai 200237, China (e-mail:[email protected]; [email protected]) Zhenxing Qian and Xinpeng Zhang are with College of Computer Science and Artificial Intelligence, Fudan University, Shanghai 200433, China(e-mail: [email protected]; [email protected]). Zaiwang Gu is with the Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore.(e-mail: [email protected]; ). Corresponding author: Zhiying Zhu.
Need Huge Labeled Samples
1
Probability
Softmax Classifier
arXiv:2605.09935v1 [cs.CV] 11 May 2026
Qingchao Jiang Senior Member, IEEE, Zhenxuan Hou, Zhiying Zhu, Zhenxing Qian, Senior Member, IEEE, Xinpeng Zhang, Member, IEEE, Zaiwang Gu
Really? 𝑝𝐺𝐴𝑁 𝑝𝐷𝑀 𝑝𝑅𝐸𝐴𝐿
(a) Existing Works Training Only Limited Labeled Samples
1
Inference
Probability
𝑝𝐺𝐴𝑁 𝑝𝐷𝑀 𝑝𝑅𝐸𝐴𝐿
It’s real, and the uncertainty value is 0.0012
Uncertainty
(b) Ours
Fig. 1. Comparison between existing works and our proposed method. Existing methods only provide a single prediction without model confidence. Our proposed method not only outputs the prediction but also yields the uncertainty value. Furthermore, the estimated uncertainty is seamlessly integrated into an active learning loop, enabling the model to iteratively select the most informative samples, thereby drastically reducing the dependency on largescale annotated datasets.
making distinguishing images increasingly challenging. Images generated by techniques such as Generative Adversarial Networks (GANs) [3] and Diffusion Models (DMs) [4] exhibit high fidelity, rendering them nearly indistinguishable to the human eye. In particular, the emergence of diffusion-based models has ushered image generation into a transformative new era. Non-expert users can now easily obtain high-quality AI-generated images and perform content editing through commercial APIs, such as Midjourney [5] and Flux [6]. These technologies have been widely utilized in fields such as virtual reality and digital entertainment. However, technology is a double-edged sword. These advancements are also being exploited for malicious purposes, including identity fraud and the dissemination of disinformation. Given that the misuse of synthetic images seriously jeopardizes social stability [7], the development of detection methods capable of effectively discriminating between real and synthetic faces is imperative. Existing research primarily focuses on the intrinsic fingerprints left by different generative mechanisms [8]–[13]. The core idea behind these approaches is to exploit model-specific statistical discrepancies or structural cues embedded in the data distribution to enhance detection performance. Notably, [13] utilized model reconstruction discrepancy to achieve effective ternary discrimination by performing inverse reconstruction on images using multiple generative models. Despite their promising results, as illustrated in Figure 1, previous works
0000–0000/00$00.00 © 2021 IEEE
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
only output a single prediction. It is impossible to ascertain the model’s confidence in its own judgment. This is primarily because they almost exclusively employ Softmax as the activation function, providing a point estimate for the sample’s class probability without offering the associated uncertainty. Meanwhile, since Softmax converts logits into probabilities via exponential operations, even minor differences in logits are amplified. Consequently, even when encountering unknown OOD images (e.g., images generated by unseen architectures), the model tends to yield erroneous predictions with high confidence scores, failing to express ’I’m not sure’ regarding these unknown OOD images. Given the critical importance of accurately detecting OOD images for synthetic face detection, addressing this fatal issue is paramount for ensuring detection model reliability. Furthermore, the majority of existing methods rely on a fully supervised training paradigm, necessitating extensive high-quality annotated data. In real-world scenarios, however, acquiring labeled samples that encompass a wide range of generative models is often prohibitively expensive. This significantly constrains the scalability and practical deployment value of these methods. To mitigate the lack of model confidence information, we draw inspiration from evidential deep learning (EDL) [14], [15]. EDL models the distribution of class probabilities using the Dirichlet distribution, thereby explicitly quantifying the uncertainty inherent in model predictions. Specifically, the model outputs are interpreted as ’evidence’ reflecting the confidence for each class, which subsequently serves as the parameters for the Dirichlet distribution to model uncertainty. During the training phase, EDL learns to generate evidence directly from input data, empowering the model to evaluate uncertainty and integrate it into the decision-making process. We observe that integrating the EDL strategy into the synthetic face detection task allows for the effective quantification of the model’s uncertainty when dealing with images. Crucially, model uncertainty is not only valuable during the inference phase but also serves as an effective criterion for sample selection throughout training. The primary objective of Active Learning (AL) is to prioritize the selection of the most informative samples from vast unlabeled pools for manual annotation, thereby maximizing model performance while minimizing labeling overhead under constrained budgets. In the context of synthetic face detection, AL offers inherent advantages. By leveraging the uncertainty quantified by EDL to assess the informational value of each sample, we can effectively transform ’model ignorance’ into ’annotation priority.’ This strategy significantly improves data efficiency and enhances the model’s detection performance and generalization capability. Motivated by the above observations, we propose EMSFD, a synthetic face detection framework that integrates evidencebased decision modeling with uncertainty-driven active learning. Specifically, the proposed framework decomposes synthetic face detection into two stages. In the first stage, a binary classification task is performed in the spatial domain to determine whether the input image is real or synthetic. In the second stage, for images identified as synthetic, source attribution is further conducted in the frequency domain to distinguish whether they are generated by GANs or diffusion
2
models. To support these two stages, we design a Face Evidence Extraction (FEE) module, which learns discriminative representations and constructs evidence-based outputs under a Dirichlet distribution, enabling the model to provide both class predictions and corresponding uncertainty estimates. Building upon this design, we further incorporate model uncertainty into the active learning process. During training, the uncertainty estimated by EDL is used to evaluate unlabeled samples, and those with higher uncertainty, which are considered more informative, are preferentially selected for manual annotation and iterative model updating. With this mechanism, EMSFD not only alleviates the overconfidence issue commonly encountered by conventional methods under OOD scenarios, but also improves detection performance and cross-generator generalization under limited labeling budgets. Extensive experimental results demonstrate that the proposed method consistently outperforms existing mainstream approaches under multiple benchmark settings, validating the effectiveness of jointly leveraging evidence modeling and active learning for synthetic face detection. Our contributions can be summarized as follows: 1) We pioneer the application of Evidential Deep Learning to synthetic face detection, which not only improves the detection accuracy but also effectively quantifies predictive uncertainty. 2) We further propose an uncertainty-driven active learning strategy that reduces the labeling cost of model training while improving model performance and generalization. 3) Comprehensive experiments validate that our method achieves superior performance compared to existing SOTA methods. The paper is structured as follows. Section II presents an overview of related works. Section III details the proposed methodology. Section IV presents experimental results and comprehensive analysis. Finally, Section V concludes the paper. II. R ELATED W ORK This section provides a brief overview of image generation techniques, synthetic image detection methods, evidential deep learning, and active learning. A. Image Generation The dominant paradigms in modern image generation are GANs [3] and DMs [4]. Driven by a minimax game mechanism, GANs have long dominated the landscape of highfidelity image synthesis. To address the instability associated with training on high-resolution images, Reference [16] proposed ProGAN, which employs a progressive growing strategy that incrementally increases resolution layer by layer, effectively enhancing both training stability and generation quality. Building on this foundation, StyleGAN [17] introduced a stylebased generator architecture that leverages Adaptive Instance Normalization (AdaIN) to achieve disentangled control over high-level semantic features. Subsequently, StyleGAN2 [18] redesigned the normalization operations to eliminate dropletlike artifacts and further optimize image quality. Furthermore,
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
approaching the problem from a signal processing perspective, StyleGAN3 [19] resolved the issue of texture sticking during generation, thereby achieving strict translation and rotation equivariance. In addition, to synergize the local feature perception of Convolutional Neural Networks (CNNs) with the long-range dependency modeling capabilities of Transformers, Reference [20] proposed VQGAN. By learning a discrete visual codebook, this method effectively models image structures while maintaining high perceptual quality. Despite the advantage of GANs in terms of inference speed, the issue of Mode Collapse when dealing with complex distributions remains a primary bottleneck. In recent years, diffusion models have emerged as a new paradigm in image generation, leveraging their superior capability in covering data distributions and achieving stable training. Reference [21] proposed Ablated Diffusion Model (ADM), which demonstrates for the first time on ImageNet that diffusion models surpass GANs in terms of image fidelity. Subsequently, Reference [22] introduced Improved DDPM (IDDPM), which further refines the noise scheduling strategy. By learning the variance schedule, IDDPM achieves a reduction in sampling steps while improving log-likelihood performance. However, performing iterative denoising directly in the pixel space incurs substantial computational overhead. To address this limitation, Reference [23] proposed Latent Diffusion Model (LDM), which leverages perceptual compression techniques to shift the diffusion process to a low-dimensional latent space. This approach significantly reduces computational complexity while preserving semantic details of images. Building upon the LDM architecture, Stable Diffusion further validates the effectiveness of this methodology on large-scale datasets, demonstrating superior capabilities in text-to-image generation and strong generalization performance, establishing itself as the mainstream generative foundation model in current practice. Since GANs and DMs are currently the most mainstream face generation models, we aim to construct a method capable of detecting both of them simultaneously. B. Synthetic Image Detection Synthetic image detection aims to distinguish whether an image is real (originating from the real world) or synthetic (generated by a generative model). Existing methods primarily fall into two categories: one leverages the artifacts generated during the image synthesis process. For instance, reference [11] proposed a hierarchical multi-level approach that achieves source attribution for synthetic methods through three distinct levels. Reference [9] demonstrates that each GAN model possesses a unique intrinsic fingerprint, which can be exploited for image attribution and forensic analysis. Building on this, [24] identified consistent anomalous patterns in synthetic images through frequency domain analysis, leading to the development of a frequency-based detector for GAN-generated imagery. Reference [25] effectively leverages Large Vision-Language Models (LVLMs) to detect subtle yet discernible artifacts in synthetic images, achieving excellent performance. The other exploits reconstruction discrepancies among different generative models. DIRE [10] detects images
3
generated by DMs by measuring the residual between the input image and the image reconstructed after inversion using a pretrained Diffusion Model. SPAI [26] detects AI-generated images as OOD samples by self-supervised learning the spectral distribution of real images and subsequently utilizing spectral reconstruction similarity. Regardless of the specific method, they predominantly rely on the Softmax activation function, which inevitably leads to overconfident predictions for certain classes, while severely lacking interpretability. C. Evidential Deep Learning EDL estimates the uncertainty in classification predictions by modeling the Dirichlet distribution of class probabilities. In contrast to Bayesian Neural Networks (BNNs) [27], which indirectly infer predictive uncertainty via weight distributions, EDL explicitly models uncertainty based on Subjective Logic theory and the Dempster-Shafer Theory (DST) [28] of evidence. The specific mechanism of the EDL framework is described as follows. For a given input xi and K classes, the neural network outputs an evidence vector ei = [ei1 , ei2 , . . . , eiK ]T , where each element represents the model’s support for the corresponding class. To ensure the non-negativity of evidence, an activation function such as ReLU, Softplus, or an exponential function is applied to the network output. The evidence vector ei is subsequently utilized to define the parameters αi of the Dirichlet distribution, formulated as: αik = eik + 1,
(1)
where k denotes the index of each class. The resulting Dirichlet distribution pi ∼ Dir(αi ) (where pi represents the class probabilities) allows for the modeling of predictive uncertainty. According to Subjective Logic theory, the uncertainty ui associated with the input xi is defined as: K ui = PK
k=1 αik
,
(2)
PK where k=1 αik represents the sum of Dirichlet parameters across all classes. Here, ui is inversely correlated with the accumulated evidence. A scenario with zero total evidence yields maximum uncertainty (ui = 1), signifying complete ignorance of the model regarding the input. D. Active Learning AL aims to select the most informative samples from largescale unlabeled data for manual annotation, thereby training high-performance models at minimal labeling costs. By prioritizing valuable samples and reducing annotation redundancy, AL provides an effective solution for improving data efficiency under limited labeling budgets. Traditional AL approaches typically rely on heuristic query strategies, such as uncertainty sampling, query-by-committee, and expected model change. In recent years, with the rapid advancement of deep learning, AL methods tailored for deep neural networks have emerged, including loss-prediction-based LL4AL [29], core-set-based selection [30], and uncertainty estimation methods grounded
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
4
Which is real ?
STEP 1
STEP 2 GANs or DMs?
REAL or Synthetic ? 1
FEE
2D FFT
augmentation
FEE
augmentation
1
𝑝𝐺𝐴𝑁 𝑝𝐷𝑀 𝑢𝑛𝑐𝑒𝑟𝑡𝑎𝑖𝑛𝑡𝑦
e2 exponential
evidence
1
+
2
Class probabilities
1 1 + 2 2 p2 = 1 + 2
D( p | )
p1 =
alpha
1
Dirichlet distribution
1
Backbone Input
Linear layer
Evidence Learning
u= Output
2
1 + 2
REAL
Uncertainty
or
GANs
Architecture of FEE
𝑝𝐴𝐼 𝑝𝑅𝐸𝐴𝐿 𝑢𝑛𝑐𝑒𝑟𝑡𝑎𝑖𝑛𝑡𝑦
e1
DMs
REAL GANs DMs
Fig. 2. The overall workflow of EMSFD. The given image is initially screened to determine its reality. If the image is classified as synthetic, it is then passed to the second stage to discriminate among the generative models. The FEE module is responsible for providing both the prediction and the uncertainty value. The augmented images are fed into the FEE module and first processed by the backbone network for evidence extraction. The extracted evidence is then used to construct a Dirichlet distribution, from which the prediction probabilities and uncertainty are derived.
in Bayesian deep learning. These methods have substantially improved the practicality of AL in complex learning scenarios and significantly promoted its application in a broad range of computer vision tasks, such as image classification, object detection, and semantic segmentation. Despite its extensive exploration in general visual recognition tasks, the application of AL to Synthetic Face Detection remains relatively limited. Most existing Synthetic Face Detection methods follow a fully supervised training paradigm, which relies heavily on large-scale, high-quality annotated datasets, thus facing prohibitive labeling costs and strong data dependency. To date, only a handful of studies have attempted to introduce AL mechanisms into this field to mitigate the scarcity of labeled resources. III. P ROPOSED M ETHOD This section presents the technical details of the proposed EMSFD. Section III-A introduces the motivation of our method. Section III-B describes the overall framework of EMSFD. Section III-C details the Face Evidence Extraction (FEE) module. Section III-D introduces the uncertainty-driven active learning strategy. Section III-E further presents the training objective. A. Motivation Currently, most SOTA synthetic face detection methods rely on the Softmax function to convert the network output logits into class probabilities. Although this mechanism performs well in distinguishing known categories, it essentially produces
point estimates of probabilities through exponential normalization. Formally, the probability pij is computed as:
exp(zij ) pij = PK , k=1 exp(zik )
(3)
where i indexes the input samples, j indexes the classes, zij denotes the logit output for class j of sample i, and K is the total number of classes. Because Softmax utilizes exponential normalization, it tends to inflate the probability of the winning class, leading to a phenomenon known as pathological overconfidence. Although this mechanism is beneficial for separability in standard tasks, it renders the model incapable of providing reliable uncertainty estimates. Even when encountering OOD samples or adversarial attacks that differ vastly from the training manifold, the model often predicts with high confidence rather than expressing uncertainty. A classic example is a digit classifier processing a cat image: instead of signaling an ”unknown” status, it forces a confident classification into one of the digit classes. This inability to express ”I don’t know” poses severe risks in safetycritical applications, particularly when defending against OOD spoofing attacks where misleading high confidence predictions can be fatal. Given that EDL has achieved significant success in the field of fake audio detection [31], to mitigate the pathological overconfidence in OOD predictions for synthetic face detection, we propose the EMSFD, inspired by EDL.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
5
Generative Adversarial Networks
Diffusion Models
Fig. 3. Comparison of 2D FFT spectra across diverse generative models. The figure illustrates the frequency-domain characteristics of images synthesized by StyleGAN, StyleGAN2, VQGAN, ADM, IDDPM, and LDM. These spectra reveal distinct frequency artifacts inherent to each specific architecture during the image synthesis process.
B. EMSFD Framework The top panel of Figure 2 illustrates the pipeline of the image detection process. Drawing inspiration from the multilevel deepfake detection and recognition framework presented in [11], the proposed EMSFD implements a two-stage progressive classification scheme. The first stage performs a coarsegrained binary classification (Real vs. Synthetic), treating outputs from GANs and DMs as a unified ’synthetic’ class. The second stage follows with a fine-grained classification (GANs vs. DMs) to identify the specific source of the synthetic images. Furthermore, we observed that images generated by distinct mechanisms exhibit more significant discrepancies in the frequency domain. Consequently, we employ a dual stream learning mechanism. The spatial stream is utilized to differentiate between real and fake faces, while the frequency stream is dedicated to distinguishing between generative architectures. The inference process operates in two stages. In STEP 1, the FEE module takes RGB images as input for Real vs. Synthetic binary classification, outputting class evidence parameters and uncertainty scores. Upon identifying a sample as ’Synthetic’, the workflow transitions to STEP 2, where the image is converted to the frequency domain via a two-dimensional fast Fourier transform (2D FFT). The corresponding frequency-domain representation is presented in Figure 3. The transform is computed independently for each RGB channel across the spatial dimensions, resulting in a complex-valued representation that preserves the original spatial resolution. We maintain standard FFT output configurations to ensure end-to-end consistency within the framework. This spectral representation is then processed by the FEE module to further classify the source as either GANs or DMs. An active learning mechanism is further introduced during the training phase. Specifically, the model is initially trained on a small labeled seed set. Subsequently, the most informative samples are selected from the unlabeled pool based on their uncertainty scores and added to the labeled set for iterative model updates. C. Face Evidence Extraction To effectively learn discriminative features, we design a FEE module as shown at the bottom of Figure 2. This module is structured with three integral components: a frozen backbone
for feature extraction, a Linear layer for feature projection and dimensionality reduction, and an Evidence layer to quantify uncertainty based on the Dirichlet distribution. To adapt the complex-valued spectra for ViT in STEP 2, we decompose the FFT output into log-magnitude (log(|F | + 10−8 )) and phase (∠F ) components. The 3-channel log-magnitude and 3-channel phase maps are concatenated along the channel dimension, resulting in a 6-channel frequency-domain representation. To accommodate this, the ViT patch embedding layer is modified for 6-channel input. Unlike traditional Softmax classifiers that directly predict class probabilities, the FEE module explicitly models the uncertainty of predictions through the Dirichlet distribution, learning evidence rather than probabilities. Let α = [α1 , α2 , . . . , αK ] be the evidence parameters. The Dirichlet distribution is defined as: D(p1 , . . . , pK |α) =
1 Y αk −1 pk B(α)
(4)
k
where B(α) is the multivariate Beta function: Q Γ(αk ) B(α) = kP Γ ( k αk )
(5)
and Γ is the Gamma function. The corresponding class probabilities and uncertainty estimates are: pk =
αk K ,u = , S S
S=
X
αk
(6)
k
The output of the FEE module comprises the corresponding class probabilities and uncertainty estimates. The uncertainty u reflects the model’s confidence in its predictions. A larger value of uncertainty indicates higher uncertainty in the prediction. It is generally accepted that OOD samples should exhibit high uncertainty, whereas In-Distribution (ID) samples are expected to maintain low uncertainty. D. Uncertainty-Driven Active Learning We further introduce an uncertainty-driven active learning strategy into the training process of EMSFD. Let the training set consist of a labeled subset DL and an unlabeled pool DU . At the beginning of training, we construct an initial labeled set by sampling a small number of instances from each class in a class-balanced manner. The model is first trained on DL for
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
6
several epochs. After each active learning round, the current model is used to estimate the uncertainty of samples in DU , and the most informative samples are selected for annotation and added to the labeled set for the next round of training. A larger uncertainty value indicates that the model has less confidence in the current sample and therefore considers it more informative for subsequent training. Based on the uncertainty scores, we rank all samples in the unlabeled pool and select the top-B samples with the highest uncertainty as the query set: Q = TopBx∈DU u(x),
(7)
where B is the query budget for each active learning round. After manual annotation, the queried samples are moved from the unlabeled pool to the labeled set: DL ← DL ∪ Q,
DU ← DU \ Q.
(8)
The model is then further optimized on the updated labeled set, and this procedure is repeated until the annotation budget is exhausted or the maximum number of active learning rounds is reached. E. Training Objective To simultaneously guarantee both high classification accuracy and robust inter-class separation, we carefully design a multi-task learning framework that integrates the EDL loss and a contrastive learning loss. LEDL =
B X i=1
wi
K X
yik (ψ(Si ) − ψ(αik ))
(9)
k=1
where yik is the one-hot encoded ground truth label, ψ is the digamma function,P wi is the sample weight for handling class imbalance, Si = k αik is the sum of evidence, B is the batch size, and K is the number of classes. Lcontrastive =
X −1 X exp(sim(zi , zp )/τ ) log P |P (i)| a∈A(i) exp(sim(zi , za )/τ ) i∈I
p∈P (i)
(10) where zi , zp , za are the L2 normalized feature vectors for the anchor, positive, and contrastive samples, respectively; sim(zi , zj ) = ziT zj denotes the cosine similarity (dot product) between normalized embeddings; I ≡ {1, . . . , B} is the set of indices within a mini-batch of size B; A(i) ≡ I \ {i} is the set of all indices in the batch excluding the anchor i; P (i) ≡ {p ∈ A(i) : yp = yi } is the set of indices of all positive samples sharing the same class label yi as the anchor i, with |P (i)| denoting its cardinality; and τ = 0.1 is the temperature parameter that scales the similarity scores. The two losses are combined with equal weights to form the total loss function: Ltotal = LEDL + Lcontrastive
(11)
This design enables the model to simultaneously optimize classification accuracy and feature space structure, ensuring both strong discriminative performance and features with good clustering and separation properties.
IV. E XPERIMENTS A. Experimental Setup Backbone and Data Augmentation. In this paper, we employ the Vision Transformer [32] (ViT-B/16) as the backbone network, incorporating the Exponential function as the activation mechanism. To mitigate overfitting and enhance model robustness, we implement a comprehensive data augmentation strategy as shown in Figure 4. Specifically, input images undergo Random Resized Crop and Random Horizontal Flip. Geometric and photometric diversities are further enriched through Random Rotation, Random Affine transformations, and Color Jitter. Training Details. The model is optimized using the AdamW optimizer with an initial learning rate of α = 10−5 and a weight decay of λ = 5 × 10−4 . A Cosine Annealing scheduler is utilized with a period of 100 steps, coupled with gradient norm clipping (max norm set to 1.0) to stabilize convergence. The batch size is fixed at 16. Under the active learning setting, half of the training set is first selected as the initial labeled subset, while the remaining half forms the unlabeled pool. The model is then trained on the current labeled subset for 10 epochs, after which the uncertainty scores produced by EDL are used to evaluate the unlabeled samples. In each active learning round, the top 20% most uncertain samples from the unlabeled pool are selected for annotation and added to the labeled subset for the next round of training. In total, five rounds of active learning are conducted in our experiments. The same active learning protocol is adopted for both Step 1 and Step 2, while the two stages are optimized independently according to their different input forms. Specifically, Step 1 takes spatial-domain RGB images as input, whereas Step 2 uses frequency-domain representations obtained via a two-dimensional fast Fourier transform. All experiments are conducted on a single NVIDIA RTX 5060Ti GPU. All active learning comparisons are conducted under an identical annotation budget. Datasets and Metrics. Real face images were sampled from the Asian faces within the FFHQ1 dataset, while synthetic face samples were derived from the ASFD2 [13] dataset. We selected StyleGAN and ADM to construct the training set, while the remaining generators served as the OOD test set to evaluate the model’s generalization capability. In Step 1, the task is formulated as a binary classification between real and synthetic faces. The training set comprises a total of 6,000 images, consisting of 2,000 real images, 2,000 GAN-generated images, and 2,000 DM-generated images. During training, GAN and DM images are collectively categorized into the ’synthetic’ class; therefore, Step 1 involves a distribution of 2,000 real and 4,000 synthetic samples. In Step 2, the objective shifts to source attribution for samples previously identified as synthetic. The training set for this stage contains 4,000 synthetic images, comprising 2,000 GAN-generated and 2,000 DM-generated images. Under the AL setting, the training set for Step 1 is further partitioned into 3,000 initial labeled samples and 3,000 unlabeled samples. Similarly, for Step 2, 1 https://github.com/NVlabs/ffhq-dataset 2 https://github.com/Hurrice-star/ASFD
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Random Resized Crop
RAW
7
Random Horizontal Flip
Random Rotation
(b)
(c)
(a)
Color Jitter
Random Affine
(d)
(e)
Fig. 4. Visual comparison of images before and after data augmentation. (a) Random Resized Crop bolsters scale invariance and local feature recognition. (b) Random Horizontal Flip expands spatial orientation via mirror symmetry. (c) Random Rotation simulates real-world pose variations and axial tilts. (d) Color Jitter (perturbing brightness, contrast, saturation, and hue) mitigates illumination and device disparities. (e) Random Affine introduces complex geometric deformations through translation and shearing. TABLE I P ERFORMANCE COMPARISON WITH EXISTING DETECTION METHODS (VALUES EXCEEDING 90 ARE HIGHLIGHTED IN GREEN , WHILE VALUES BELOW 30 ARE MARKED IN RED .) Method
ID REAL
OOD
StyleGAN
ADM
StyleGAN2
Average
VQGAN
IDDPM
LDM
Acc(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
CNNSpot [9] LGrad [33] DIRE [10] DeepFeatureX [12] Cutting-Edge [11] MDL [13]
99.5 0.2 96.8 95.3 99.6 99.9
100 99.5 99.4 99.8 100 99.8
100 98.3 99.9 100 100 100
99.8 98.2 98.3 74.9 99.4 99.3
99.9 47.4 99.8 94.4 99.9 99.9
36.4 10.2 76.9 5.5 35.9 91.4
92.6 44.9 97.1 49.4 81.8 99.9
1.4 9.8 3.6 78.1 4.4 25.5
45.7 45.5 57.6 95.2 63.8 64.1
68.5 89.1 91.1 98.6 76.7 97.7
86.4 26.6 97.8 99.6 96.1 99.5
18.1 89.1 17.1 1.6 3.1 56.1
80.3 26.6 56.7 55.6 47.2 73.7
60.5 56.5 69.1 66.9 59.9 81.6
84.2 48.2 84.8 82.4 81.5 89.5
Ours
99.8
97.4
99.8
99.0
99.8
93.5
99.9
86.1
79.9
99.6
96.4
99.3
97.9
96.4
96.2
TABLE II P ERFORMANCE COMPARISON WITH DIFFERENT ACTIVATION FUNCTIONS (VALUES EXCEEDING 90 ARE HIGHLIGHTED IN GREEN , WHILE VALUES BELOW 30 ARE MARKED IN RED .) Activation
ID REAL
Softplus ReLU Exponential
OOD
StyleGAN
ADM
StyleGAN2
VQGAN
Average IDDPM
LDM
Acc(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
98.2 95.9 99.8
71.9 94.3 97.4
97.3 99.1 99.8
99.9 100 99.0
98.1 98.1 99.8
71.9 94.3 93.5
98.2 99.2 99.9
71.9 94.1 86.1
93.1 98.1 79.9
99.9 100 99.6
99.7 99.7 96.4
98.5 51.3 99.3
89.7 86.9 97.9
87.5 89.9 96.4
97.0 96.8 96.2
the training set is subdivided into 2,000 initial labeled samples and 2,000 unlabeled samples. Performance is evaluated using Accuracy (Acc) and Average Precision (AP). B. Comparison with Existing Detectors We conducted a comparative evaluation against representative SOTA methods, including CNNSpot3 [9], LGrad4 [33], DIRE5 [10], DeepFeatureX6 [12], Cutting-Edge7 [11], and Model Discrepancy Learning (MDL) [13]. To ensure a fair comparison, all selected baseline models were retrained on the ASFD dataset using their publicly available source code. Specifically, for the models proposed in [9], [33], and [10], the classifier heads were modified from a binary to a ternary classification configuration. We tested the generalization capabilities of these models across a diverse set of generators. Table I presents the detailed performance comparison. 3 https://peterwang512.github.io/CNNDetection/ 4 https://github.com/chuangchuangtan/LGrad 5 https://github.com/ZhendongWang6/DIRE 6 https://github.com/opontorno/block-based deepfake-detection 7 https://iplab.dmi.unict.it/mfs/Deepfakes/MasteringDeepfake2023/
Overall, existing methods generally exhibit significant performance degradation on specific generative models, particularly VQGAN, IDDPM, and ADM, revealing a severe deficiency in cross-distribution generalization. For instance, the Acc of CNNSpot on VQGAN is a mere 1.4%, while LGrad plummets to 0.2% on StyleGAN. Similarly, DIRE and DeepFeatureX demonstrate substantial performance fluctuations. In contrast, our method maintains stable and superior performance across all tested generators. Notably, it achieves an AP of over 97% on StyleGAN, LDM, IDDPM and ADM. Even on the most challenging VQGAN, our method also sustains high performance. With an average performance of 96.4% Acc and 96.2% AP, our approach significantly outperforms existing methods. These results indicate that our method effectively captures the underlying statistical regularities of images, thereby maintaining robust generalization capabilities even when confronting diverse generative mechanisms, demonstrating high potential for practical applications. Meanwhile, we perform an Expected Calibration Error (ECE) [34] analysis to quantitatively assess the reliability of our method. As summarized in Table III, our method
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
TABLE III E XPECTED C ALIBRATION E RROR ACROSS DIFFERENT METHODS
Method
OOD Average ECE ↓
CNNSpot [9] DIRE [10] Ours
0.6179 0.5803 0.0516
8
STEP 1 t-SNE
STEP 1 t-SNE (by Uncertainty)
REAL Synthetic
achieves superior calibration performance in OOD scenarios. This substantial reduction in ECE indicates that our approach effectively suppresses the overconfidence inherent in standard Softmax-based classifiers. Furthermore, we evaluate the generalization performance of our model on an unseen generative method, FLUX.2. Given that FLUX.2 is built upon a novel Flow Matching architecture, its underlying generative logic differs fundamentally from conventional GANs and DMs. Consequently, this experiment focuses exclusively on firststage binary detection, omitting the second-stage source model attribution. For the evaluation, we constructed a test set of 1,000 images generated using the prompt ’An Asian human’. As summarized in Table IV, while the detection performance of all models declines when encountering this novel architecture compared to their performance on GANs and DMs, our method consistently outperforms existing baselines across all metrics. Specifically, our approach achieves an Acc of 64.9% and an AP of 67.4% with the lowest ECE (0.346), demonstrating its superior cross-paradigm generalization capability. TABLE IV D ETECTION PERFORMANCE ON THE UNSEEN GENERATIVE PARADIGM (FLUX.2).
Method
Acc (%)
AP (%)
ECE
CNNSpot [9] DIRE [10] Ours
54.1 0.0 64.9
18.9 0.0 67.4
0.459 1.000 0.346
C. Visualizations and Uncertainty Estimation We employ the t-SNE visualization method [35] to visualize the feature vectors extracted from the model’s final layer. As illustrated in Figure 5, the EMSFD achieves excellent interclass separation in both the STEP 1 and STEP 2 binary classification tasks. Specifically, the feature clusters for each class are compact, exhibiting minimal overlap. Furthermore, we superimposed the model’s uncertainty estimates onto the t-SNE projection. Observations reveal that high uncertainty is predominantly concentrated at the decision boundaries between the two classes, whereas low uncertainty is distributed within the class interiors. This demonstrates that our approach not only effectively performs synthetic face classification but also accurately quantifies prediction confidence, thereby providing interpretability for the model’s decision-making process. D. Ablation Study The selection of an efficient backbone network and an appropriate activation function is paramount for enhancing the
(a) STEP 2 t-SNE
(b) STEP 2 t-SNE (by Uncertainty)
GANs DMs
(c)
(d)
Fig. 5. Visualization of the spatial representation of our model. (a) and (c) present the t-SNE visualizations, while (b) and (d) display the t-SNE plots overlaid with uncertainty values. The uncertainty is low at the cluster centers but high at the cluster boundaries. The average uncertainty in the first step is lower than that in the second step, suggesting that identifying model attribution is relatively more challenging.
overall performance of our approach. To this end, we conducted comprehensive ablation studies to identify the optimal configuration. In this paper, we evaluated the performance of EfficientNet-B3 [36], ResNet-101 [37], DenseNet-121 [38], and ViT-B/16 under identical experimental conditions, measuring both Acc and AP. As presented in Table V, ViTB/16 exhibited superior feature extraction capabilities in both Steps, significantly outperforming other candidate backbones. Consequently, we adopted ViT-B/16 as the optimal solution for all stages of our approach. We must acknowledge that cascaded architectures are inherently susceptible to error propagation. However, in our experiments, Step 1 achieves nearsaturated performance across all evaluated generators, with an accuracy exceeding 99.8%. At this level of precision, the vast majority of synthetic samples are correctly routed to Step 2. Consequently, the overall statistical impact of cascaded errors is negligible. Meanwhile, we further investigated the impact of different activation functions on the stability of evidence output and overall detection performance. Specifically, we constructed the evidence branch using Softplus, Exponential, and ReLU activations, respectively, to validate the critical role of activation functions in deep evidential learning. Table II presents the detection results for these three activation types across different generators. The experimental results indicate that the exponential activation delivers the most stable and superior overall performance. It maintains high Acc and AP across all generators, significantly outperforming Softplus while avoiding the performance collapse observed with ReLU on the ADM generator. These findings suggest that exponential activation effectively enhances the dynamic range of evidence values, enabling the model to capture robust forgery traces more readily
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
9
TABLE V P ERFORMANCE COMPARISON ACROSS DIFFERENT BACKBONES . Method
ID REAL
OOD
StyleGAN
ADM
StyleGAN2
VQGAN
IDDPM
LDM
Average
Acc(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
ECE
99.8 86.3 92.1 91.5
97.4 88.8 82.7 92.5
99.8 96.0 94.5 97.8
99.0 95.9 95.9 99.2
99.8 90.7 89.2 95.2
93.5 89.5 83.6 92.9
99.9 98.4 97.8 99.3
86.1 60.1 82.2 65.4
79.9 77.4 92.9 83.8
99.6 98.1 97.9 99.6
96.4 98.8 97.9 99.5
99.3 90.7 92.9 97.2
97.9 75.7 79.1 78.1
96.4 87.1 89.6 91.2
96.2 89.0 92.9 92.9
0.052 0.297 0.154 0.226
ViT-B/16 ResNet-101 EfficientNet-B3 DenseNet-121
TABLE VI P ERFORMANCE COMPARISON BETWEEN S OFTMAX AND E XPONENTIAL Method
ID REAL
Softmax Exponential
OOD
StyleGAN
ADM
StyleGAN2
ECE
VQGAN
IDDPM
LDM
Average
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
89.4 99.8
95.6 100.0
95.7 97.4
99.7 99.8
96.7 99.0
92.0 99.8
73.2 93.5
94.7 99.9
72.8 86.1
85.7 79.9
99.9 99.6
99.8 96.4
96.0 99.3
84.8 97.9
89.1 96.4
93.2 96.2
0.203 0.052
TABLE VII P ERFORMANCE COMPARISON BETWEEN O URS AND CNNS POT AFTER ADDING ACTIVE LEARNING . Method
ID REAL
OOD
StyleGAN
ADM
StyleGAN2
VQGAN
Average IDDPM
LDM
Acc(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
Acc(%)
AP(%)
CNNSpot [9]
99.5
94.9
99.4
98.4
99.5
94.9
99.4
55.9
70.2
98.4
99.5
73.8
73.2
88.0
91.6
Ours
99.8
97.4
99.8
99.0
99.8
93.5
99.9
86.1
79.9
99.6
96.4
99.3
97.9
96.4
96.2
when confronting diverse generative mechanisms. Thus, the exponential activation stands out as the optimal choice for the EMSFD.
100% 90% 80% 70% 60%
To investigate the individual contributions of our architectural modules, we compare the ViT + EDL framework against a ViT + Softmax baseline (Table VI). The Softmaxbased variant exhibits obviously diminished accuracy and increased ECE relative to our complete model. These results demonstrate that the incorporation of EDL yields significant performance gains for our proposed framework. Nevertheless, it maintains performance on par with MDL. While a powerful backbone can inherently yield performance gains, we provide a comparative analysis in Figure 6 to demonstrate that the superior performance of our method is not solely attributable to the strength of the backbone network. Specifically, substituting the original CNNSpot backbone with the same ViT utilized in our framework boosts performance from 60.5% to 70.6%; however, these results are still substantially eclipsed by the 96.4% accuracy of our approach. Ultimately, these results confirm that the proposed modules are essential for achieving optimal performance. Table VII presents a detailed performance comparison between our method and CNNSpot integrated with active learning. The results demonstrate that our approach consistently outperforms CNNSpot across most evaluation metrics. Specifically, our method achieves nearperfect detection accuracy on both the IDDPM and LDM datasets, substantially surpassing CNNSpot even when the latter is augmented with active learning. These results underscore the inherent superiority of our proposed approach.
50% 40% 30% 20% 10% 0%
CNN Spot
CNN Spot(VIT) Avg Acc
Ours
Avg AP
Fig. 6. Performance comparison between CNNSpot and our EMSFD under an identical backbone.
V. C ONCLUSION In this paper, we have proposed EMSFD, a synthetic face detection framework that integrates evidence-based decision modeling with uncertainty-driven active learning. By modeling class evidence with the Dirichlet distribution, the proposed method explicitly incorporates model uncertainty into the synthetic face detection task, thereby effectively alleviating the overconfidence issue commonly exhibited by conventional Softmax classifiers under OOD scenarios and improving the reliability and interpretability of model predictions. Building upon this design, we further exploit uncertainty information in the active learning process by prioritizing high-uncertainty and thus more informative samples from the unlabeled pool for annotation, which effectively improves model performance
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
and cross-generator generalization. In addition, we construct a two-stage detection framework, where real/synthetic discrimination is conducted in the spatial domain and source attribution between GANs and diffusion models is performed in the frequency domain. Experimental results demonstrate that EMSFD consistently outperforms existing mainstream methods under multiple benchmark settings, validating the effectiveness of the collaborative design of evidence modeling and active learning for synthetic face detection. In future work, we will further explore more efficient uncertainty estimation and sample querying strategies, and investigate how the proposed framework can be extended to handle increasingly complex forged samples generated by emerging generative architectures, thereby further improving the robustness and practical value of synthetic face detection systems in open environments.
R EFERENCES [1] F. Boutros, V. Struc, J. Fierrez, and N. Damer, “Synthetic data for face recognition: Current state and future prospects,” Image and Vision Computing, vol. 135, p. 104688, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0262885623000628 [2] A. Zulfiqar, S. Muhammad Daudpota, A. Shariq Imran, Z. Kastrati, M. Ullah, and S. Sadhwani, “Synthetic image generation using deep learning: A systematic literature review,” Computational Intelligence, vol. 40, no. 5, p. e70002, 2024. [Online]. Available: https: //onlinelibrary.wiley.com/doi/abs/10.1111/coin.70002 [3] V. L. Trevisan de Souza, B. A. D. Marques, H. C. Batagelo, and J. P. Gois, “A review on generative adversarial networks for image generation,” Computers & Graphics, vol. 114, pp. 13–25, 2023. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S009784932300064X [4] F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10 850–10 869, 2023. [5] Midjourney, “Midjourney,” 2024, accessed: Mar. 27, 2026. [Online]. Available: https://www.midjourney.com/ [6] Black Forest Labs, “FLUX.1: Text-to-Image Generation Model,” 2024, accessed: Mar. 27, 2026. [Online]. Available: https://blackforestlabs.ai/ [7] S. J. Nightingale and H. Farid, “Ai-synthesized faces are indistinguishable from real faces and more trustworthy,” Proceedings of the National Academy of Sciences, vol. 119, no. 8, p. e2120481119, 2022. [Online]. Available: https://www.pnas.org/doi/abs/10.1073/pnas.2120481119 [8] M. Barni, K. Kallas, E. Nowroozi, and B. Tondi, “Cnn detection of gangenerated face images based on cross-band co-occurrences analysis,” in 2020 IEEE International Workshop on Information Forensics and Security (WIFS), 2020, pp. 1–6. [9] S.-Y. Wang, O. Wang, R. Zhang, A. Owens, and A. A. Efros, “Cnngenerated images are surprisingly easy to spot... for now,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8695–8704. [10] Z. Wang, J. Bao, W. Zhou, W. Wang, H. Hu, H. Chen, and H. Li, “Dire for diffusion-generated image detection,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 22 388– 22 398. [11] L. Guarnera, O. Giudice, and S. Battiato, “Mastering deepfake detection: A cutting-edge approach to distinguish gan and diffusion-model images,” ACM Trans. Multimedia Comput. Commun. Appl., vol. 20, no. 11, Sep. 2024. [Online]. Available: https://doi.org/10.1145/3652027 [12] O. Pontorno, L. Guarnera, and S. Battiato, “Deepfeaturex net: Deep features extractors based network for discriminating synthetic from real images,” in Pattern Recognition, A. Antonacopoulos, S. Chaudhuri, R. Chellappa, C.-L. Liu, S. Bhattacharya, and U. Pal, Eds. Cham: Springer Nature Switzerland, 2025, pp. 177–193. [13] Q. Jiang, Z. Xu, Z. Zhu, N. Chen, H. Wang, and Z. Ba, “Model discrepancy learning: Synthetic faces detection based on multi-reconstruction,” in 2025 IEEE International Conference on Multimedia and Expo (ICME), 2025, pp. 1–6.
10
[14] M. Sensoy, L. Kaplan, and M. Kandemir, “Evidential deep learning to quantify classification uncertainty,” in Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., vol. 31. Curran Associates, Inc., 2018. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2018/file/a981f2b708044d6fb4a71a1463242520-Paper.pdf [15] A. Amini, W. Schwarting, A. Soleimany, and D. Rus, “Deep evidential regression,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 14 927– 14 937. [Online]. Available: https://proceedings.neurips.cc/paper files/ paper/2020/file/aab085461de182608ee9f607f3f7d18f-Paper.pdf [16] T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” 2018. [Online]. Available: https://arxiv.org/abs/1710.10196 [17] T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4401– 4410. [18] T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8110–8119. [19] T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks,” in Proc. NeurIPS, 2021. [20] P. Esser, R. Rombach, and B. Ommer, “Taming transformers for highresolution image synthesis,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 12 873–12 883. [21] P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 8780– 8794. [Online]. Available: https://proceedings.neurips.cc/paper files/ paper/2021/file/49ad23d1ec9fa4bd8d77d02681df5cfa-Paper.pdf [22] A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 8162–8171. [Online]. Available: https://proceedings.mlr.press/v139/nichol21a.html [23] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “ High-Resolution Image Synthesis with Latent Diffusion Models ,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2022, pp. 10 674–10 685. [Online]. Available: https: //doi.ieeecomputersociety.org/10.1109/CVPR52688.2022.01042 [24] J. Frank, T. Eisenhofer, L. Schönherr, A. Fischer, D. Kolossa, and T. Holz, “Leveraging frequency analysis for deep fake image recognition,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 3247–3258. [Online]. Available: https://proceedings.mlr.press/v119/ frank20a.html [25] P. Sharma, S. Garg, and D. Toshniwal, “Mirage: Unveiling hidden artifacts in synthetic images with large vision-language models,” 2025. [Online]. Available: https://arxiv.org/abs/2510.03840 [26] D. Karageorgiou, S. Papadopoulos, I. Kompatsiaris, and E. Gavves, “Any-resolution ai-generated image detection by spectral learning,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 18 706–18 717. [27] D. J. C. MacKay, “A practical bayesian framework for backpropagation networks,” Neural Computation, vol. 4, no. 3, pp. 448–472, 05 1992. [Online]. Available: https://doi.org/10.1162/neco.1992.4.3.448 [28] A. P. Dempster, “A generalization of bayesian inference,” Journal of the Royal Statistical Society: Series B (Methodological), vol. 30, no. 2, pp. 205–232, 1968. [Online]. Available: https://rss.onlinelibrary.wiley. com/doi/abs/10.1111/j.2517-6161.1968.tb00722.x [29] D. Yoo and I. S. Kweon, “Learning loss for active learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 93–102. [30] O. Sener and S. Savarese, “Active learning for convolutional neural networks: A core-set approach,” arXiv preprint arXiv:1708.00489, 2017. [31] J. Y. Kang, J. Won Yoon, S. Kim, M. H. Han, and N. Soo Kim, “Fadel: Uncertainty-aware fake audio detection with evidential deep learning,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–5.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
[32] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021. [Online]. Available: https://arxiv.org/abs/2010.11929 [33] C. Tan, Y. Zhao, S. Wei, G. Gu, and Y. Wei, “ Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection ,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2023, pp. 12 105–12 114. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/CVPR52729.2023.01165 [34] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html [35] L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of machine learning research, vol. 9, no. Nov, pp. 2579–2605, 2008. [36] M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 6105–6114. [Online]. Available: https://proceedings.mlr.press/v97/tan19a.html [37] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. [38] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4700–4708.
Qingchao Jiang (Senior Member, IEEE) received the B.E. and Ph.D. degrees from the Department of Automation, East China University of Science and Technology, Shanghai, China, in 2010 and 2015, respectively. From March 2015 to September 2015, he was a Post-Doctoral Fellow with the Department of Chemical and Materials Engineering, University of Alberta, Edmonton, AB, Canada. From September 2015 to September 2016, he was a Humboldt Research Fellow with the Institute for Automatic Control and Complex Systems (AKS), University of Duisburg-Essen, Duisburg, Germany. From October 2016 to May 2017, he was a Visiting Research Fellow with the Department of Chemical and Biomolecular Engineering, The Hong Kong University of Science and Technology, Hong Kong. He is currently a Professor with the East China University of Science and Technology. His research interests include data mining and analysis, data-driven soft sensing, multivariate statistical process monitoring, and deep learning-based process modeling. Prof. Jiang is an Early Career Advisory Board Member of the IFAC Journal Control Engineering Practice. He was a recipient of the 2021 World Artificial Intelligence Conference Youth Outstanding Paper Nomination Award. He is an Associate Editor of IEEE ACCESS.
Zhenxuan Hou received the B.Eng. degree in Engineering from Tiangong University in 2024. Currently, he is pursuing a Master’s degree in Control Engineering at East China University of Science and Technology, Shanghai, China. His research interests encompass Industrial Control System (ICS) security, focusing on adversarial examples and backdoor attacks against industrial models, as well as AI-Generated Content (AIGC) forensics, with a specific emphasis on synthetic face detection.
11
Zhiying Zhu received the B.S. degree from Hohai University (HHU) in 2016, the M.S. degree from Shanghai University (SHU) in 2019, and the Ph.D. degree from Fudan University (FDU) in 2023. He is currently an Associate Professor with the college of computer science and software engineering, Hohai University, China. His research interests include trustworthy machine learning, information hiding, and multimedia security.
Zhenxing Qian (Senior Member, IEEE) received the B.S. and Ph.D. degrees from the University of Science and Technology of China (USTC), in 2003 and 2008, respectively. He is a Professor with the School of Computer Science, Fudan University, where he is also the Vice Dean with the Key Laboratory of China Culture and Tourism Ministry. He has published over 230 peer-reviewed papers, many of them were published in IEEE Transactions and conferences like NerurIPS, CVPR, ICCV, AAAI, etc. His research interests include multimedia security and neural network security. He is an Associate Editor of IEEE TCYB, TCSVT, TASLP, etc. He also served as the Area Chair of IEEE ICME 2023 and ACM MM 2025.
Xinpeng Zhang (Member, IEEE) received the B.S. degree in computational mathematics from Jilin University, China,in 1995, and the M.E. and Ph.D. degrees in communication and information systems from Shanghai University, China, in 2001 and 2004, respectively. Since 2004, he has been with the faculty of the School of Communication and Information Engineering, Shanghai University, where he is currently a Professor. He is also with the faculty of the School of Computer Science, Fudan University. He was a Visiting Scholar with The State University of New York at Binghamton from 2010 to 2011, and also an experienced Researcher with Konstanz University, sponsored by the Alexander von Humboldt Foundation, from 2011 to 2012. His research interests include multimedia security, image processing, and digital forensics. He has published over 200 papers in these areas. He is an Associate Editor of the IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY.
Zaiwang Gu received the B.S. and MS degrees from the Hohai University and Shanghai University, in 2016 and 2019, respectively. He is a senior research staff with Agency for Science, Technology and Research (A*STAR), Singapore. His research interests include trustworthy AI and medical image analysis.