ConceptioArchivearXiv CS
arXiv CSopen access

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang School of Computer Science and Technology, Harbin Institute of Technology, Weihai {2023210984,2023211640,shaoyi.fang}@stu.hit.edu.cn {beiczhang,qizb,xuyuanrong,xinyliu,wgzhang}@hit.edu.cn

arXiv:2604.16207v1 [cs.CV] 17 Apr 2026

Abstract As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND.

CCS Concepts • Security and privacy → Human and societal aspects of security and privacy.

Keywords Face Forgery Detection, Incremental Learning, Semantic Anchors, Fine-Grained Visual-Text Alignment ACM Reference Format: Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang. 2026. AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection. In Proceedings of the 2026 International Conference on Multimedia Retrieval (ICMR ’26). ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/ 3805622.3810877 ∗ Corresponding author.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. ICMR ’26, Amsterdam, Netherlands © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/10.1145/3805622.3810877

(a) Existing IFFD methods Replay Set

Task t Real?

Classifier Fake?

(b) Ours

Task t

Dynamic Matching

Image Encoder

Binary label Head

Artifact-Probe Attention Inconsistent Lighting Blurry Eyes Unnatural Jawline

Text Encoder

Real? Fake? Blur eyes 0.95

Multi label Head

......

Smooth lip 0.08

(c) Results

Figure 1: Comparison between AIFIND and other methods. (a) Conventional methods depend on replay sets to mitigate forgetting. (b) Our method realizes a data-replay-free paradigm via stable semantic anchors. Visual features are continuously aligned with stable semantic priors, ensuring consistent decision boundaries across tasks.

1

Introduction

With the rapid development of generative models, face forgery has advanced and achieved unprecedented realism, posing serious public safety risks. Thus, developing forgery detection becomes crucial for maintaining security in the digital ecosystem. Existing mainstream face forgery detection methods [5, 20, 31, 36, 49] attempt to train a generalizable detector with limited and static datasets. However, new forgery techniques emerge endlessly, which quickly render existing detection methods ineffective. In view of this, Incremental Face Forgery Detection (IFFD) is proposed to continuously train the model with the latest forged samples. Currently, existing methods [4, 24, 37] for IFFD mainly rely on data replay, which preserves knowledge by storing a small subset of past samples. Representative approaches such as DFIL [24] and SURLID [4] follow this paradigm. However, merely replaying discrete samples fails to explicitly constrain the topology of the feature space. Regardless of data replay or regularization, as shown in Fig. 1(a), current IFFD methods rely on coarse binary supervision, failing to leverage fine-grained artifact cues. Critically, in incremental learning settings, without stable anchors to stabilize the latent space, the learned feature distribution is prone to unconstrained drift. Consequently, when accommodating new forgery types, the model inadvertently overwrites previously learned representations, resulting in catastrophic forgetting.

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

Furthermore, while recent works [5, 20, 36] have sought to integrate Vision-Language Models (VLMs) to enhance deepfake detection, they primarily focus on static scenarios and fail to apply them in incremental learning. Existing VLM-based methods often treat semantic information as a coarse global label, neglecting the fine-grained semantic discrepancies between authentic and forged facial regions. Motivated by this, we rethink the IFFD paradigm by exploiting the intrinsic semantic sensitivity of pre-trained VLMs toward specific artifact-prone regions. We posit that linguistic concepts possess a natural invariance: while the visual manifestations of local anomalies vary across datasets, the semantic distinction between authentic and manipulated facial components remains constant. Leveraging this property, we propose to utilize these region-aware authenticity priors as invariant semantic anchors. By anchoring visual features to these stable semantic coordinates, we can explicitly stabilize the evolving visual feature space. As shown in Fig. 1(b), we interpret artifacts as invariant semantic anchors. This formulation establishes a stable, high-dimensional reference for fine-grained supervision, enabling semantically guided artifact inference to resist feature drift. Based on this, we propose Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection (AIFIND). AIFIND consists of the following cooperative components: (1) Artifact-Driven Semantic Prior Generator (ASPG) constructs the initial semantic anchors by interpreting volatile low-level artifact cues into stable sparse textual labels; (2) Artifact-Probe Attention (APA) module injects selected textual artifact cues into the image encoder for fine-grained visual–text alignment; (3) Semantic-Guided Incremental Detector (SGID) governs the learning process within this anchored space, leveraging APA and dual supervision to simultaneously discriminate authenticity and identify specific artifact types; and (4) Adaptive Decision Harmonizer (ADH) aligns binary and multi-label heads by strictly preserving the angular relationships relative to the semantic anchors, guaranteeing consistency in decision boundaries. During training, AIFIND executes a dynamic matching strategy to achieve stable incremental learning. First, ASPG instantiates semantic anchors to form a fixed coordinate system. The model then transitions to a dynamic matching mechanism, autonomously recalibrating targets based on similarity. Then, APA injects these matched semantic anchors to enforce fine-grained anchoring. This process continuously rectifies volatile visual features via stable semantic definitions. Finally, ADH harmonizes the classifiers by preserving their angular relationships relative to the semantic anchors, maintaining geometric consistency across tasks. As a result, AIFIND effectively mitigates catastrophic forgetting, enabling semantically coherent knowledge evolution without replay buffers. Experiments under multiple incremental protocols [4] validate the superiority and strong generalization capability of our method. Our main contributions are summarized as follows:

• We propose a data-replay-free framework for IFFD, which reinterprets artifacts as semantic anchors to explicitly constrain the feature space without storing past data. • We formulate an artifact-probe attention that effectively anchors evolving visual forgery cues to immutable semantic anchors, preventing feature drift.

• We conduct comprehensive experiments that empirically validate the superiority of our framework, proving the effectiveness of semantic anchors in incremental learning.

2 Related Work 2.1 Face Forgery Detection Face forgery detection has become a prominent research topic in computer vision. Early studies mainly relied on detecting anomalies in biometric cues such as eye blinking [18], head pose [46], and pupil morphology [7]. With the rapid advancement of deep learning, attention gradually shifted toward learning forgery traces from multimodal clues, such as frequency domain [13, 49] and temporal consistency [8, 45]. Recently, the emergence of vision–language models such as CLIP [25] has opened up new possibilities for semanticlevel deepfake detection. RepDFD [20] reprograms CLIP by injecting universal perturbations into input images, while ForAda [5] leverages a Forensics Adapter to learn hybrid boundary artifacts specific to facial forgeries. VLFFD [36] attempts fine-grained visual–text alignment by incorporating detailed textual semantics, but its dependence on paired real–fake data limits its applicability to incremental learning. These approaches [5, 31, 42, 45, 49] aim to extract universal forgery representations from limited observed data and achieve effective transfer to unseen manipulation types.

2.2

Incremental Face Forgery Detection

As forgery techniques evolve rapidly, constructing a general detector from limited training datasets has become increasingly impractical. This motivates the exploration of Incremental Face Forgery Detection (IFFD), which enables models to continuously learn emerging forgery patterns while preserving previous knowledge. Incremental learning methods are commonly categorized into three paradigms: parameter isolation [38, 39], parameter regularization [1, 29], and data replay [2, 22, 34]. In the context of IFFD, most existing studies emphasize knowledge distillation and replay mechanisms. CoReD [15] preserves old knowledge through taskaware distillation while adapting to new domains. DFIL [24] replays representative and challenging samples from previous datasets, and DMP [37] dynamically expands prototypes to accommodate newly emerging forgery types. HDP [35] achieves replay through UAPs [23], while SUR-LID [4] aligns latent features across tasks to mitigate mutual interference. However, these methods still treat all forgery instances as a single Fake category, which limits their ability to capture fine-grained artifact semantics. Meanwhile, Vision–Language Models (VLMs) such as CLIP [47] have inspired new paradigms for incremental learning. Promptbased approaches [16, 28, 33, 40] design task-specific prompts to guide CLIP models, while adapter-based techniques [9, 17, 48] insert lightweight trainable modules at intermediate layers to encapsulate new knowledge. Despite their success in general incremental learning, these methods are not specifically designed for the unique challenges of deepfake detection, where subtle visual cues play a critical role, leading to limited effectiveness when applied directly.

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

Artifact-Probe Attention

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Artifact-driven Semantic Prior Generator top-k

normalized

raw scores

mask

mask

Query

Blurry Eyes Unnatural Jawline

Layer Norm Self-Attention

Inconsistent Lighting

Adaptive Decision Harmonizer Binary Consensus Boundary Semantic Boundary Align Align

Task t

push

Semantic-Guided Incremental Detector Task t

Key&Value

Multi label Head

Backbone

ASPG Multi-Head Cross-Attention

Binary label Head

Incrementing Task t Boundary Align

Incrementing gate

Task t+1

Backbone

ASPG

Image Encoder

Attention Map

Layer Norm

ArtifactProbe Attention

Inconsistent Lighting Blurry Eyes

Feed-Forward Network

Features

ℒce

push

Binary label Head

Task t+1

Dynamic Matching

Top-k

ℒtag

Text Encoder

Unnatural Jawline

Semantic Align

ℒdis

Multi label Head

push

Embedding Similarity Matrix

Figure 2: Overall framework of AIFIND. Artifact-Driven Semantic Prior Generator (ASPG) instantiates semantic anchors to build a fixed coordinate system. Semantic-Guided Incremental Detector (SGID) uses the Artifact-Probe Attention (APA) to perform the anchoring, which constrains volatile visual features to these stable semantic anchors. Adaptive Decision Harmonizer (ADH) maintains geometric semantic consistency, preserving semantic angular relationships across tasks. Table 1: Correspondence between Indicators and Facial Regions. ✓ indicates that the specific operator is applied. Facial Region

Blur

Color

Structure

Texture

Boundary

Eyes Nose Cheeks Mouth Jawline Boundary

✓ ✓ ✓ ✓ ✗ ✗

✓ ✓ ✓ ✓ ✗ ✗

✓ ✓ ✓ ✓ ✗ ✗

✓ ✓ ✓ ✓ ✗ ✗

✗ ✗ ✗ ✗ ✓ ✓

3 Method 3.1 Framework overview As illustrated in Fig. 2, we introduce AIFIND to enable stable incremental learning via a semantically anchored feature space. Specifically, ASPG transforms low-level cues into semantic anchors. These anchors are injected into image encoder by SGID using APA to ensure fine-grained visual–text alignment. Finally, ADH aligns classifier weights to preserve geometric consistency across tasks.

3.2

Artifact-Driven Semantic Prior Generator

According to [36] and our analysis, we define 5 representative forgery dimensions I = {Iblur, Icolor, Istructure, Itext, Iboundary }. To spatially ground these dimensions, we utilize MediaPipe Face Mesh [21] to locate 6 facial regions R = {𝑅eyes, 𝑅nose, 𝑅cheeks, 𝑅mouth, 𝑅jawline, 𝑅boundary }. Additionally, a global skin reference region 𝑆 is extracted to serve as the baseline for computing relative inconsistencies. The correspondence between facial regions and indicators is in Tab. 1 and definitions of indicators are as follows:

Blur Indicator. We measure local sharpness using the variance of the Laplacian operator:  Iblur = Var ∇2 𝐼 M , (1) where 𝐼 M denotes the grayscale intensities within the region mask. Color Indicator. To capture lighting inconsistencies, we calculate the luminance deviation between the target region 𝑅 and the skin reference 𝑆 in CIELAB space: Icolor = |𝐿𝑅 − 𝐿𝑆 | ,

(2)

where 𝐿𝑅 and 𝐿𝑆 represent the mean 𝐿-channel values, respectively. Structural Indicator. We assess the structural compatibility between the target region and the surrounding skin using the Structural Similarity Index (SSIM): Istruct = SSIM(𝑃𝑅 , 𝑃𝑆 ) ,

(3)

where 𝑃𝑅 and 𝑃𝑆 are normalized grayscale patches from 𝑅 and 𝑆. Texture Indicator. To expose statistical anomalies in skin texture, we extract local contrast using the Gray-Level Co-occurrence Matrix (GLCM): Itext = Contrast(GLCM(𝐼 M )) ,

(4)

where 𝐼 M is quantized to 64 levels to retain details. Boundary Indicator. We identify potential blending artifacts along edges by computing the average gradient magnitude: √︃  Ibound = E (𝑥,𝑦) ∈ M (5) 𝐺𝑥2 + 𝐺 𝑦2 , where 𝐺𝑥 and 𝐺 𝑦 denote Sobel derivatives. To bridge pixel-level statistics with high-level semantics, for each artifact dimension 𝑖 ∈ I and region 𝑔 ∈ R, we leverage an

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

LLM to generate a candidate set of 𝐾 contrastive text pairs, denoted 𝐾 , and align them with a support set Ω . The as T𝑖,𝑔 = {(𝑡𝑘r , 𝑡𝑘f )}𝑘=1 𝑖,𝑔 optimal anchor 𝐴𝑖,𝑔 is selected by calculating CLIP similarity: ∑︁   𝐴𝑖,𝑟 = argmax Sim(𝑡 f, 𝑥 f ) + Sim(𝑡 r, 𝑥 r ) , (6) (𝑡 r ,𝑡 f ) ∈ T𝑖,𝑔 (𝑥 r ,𝑥 f ) ∈Ω

𝑖,𝑔

where 𝑟 refers to real and 𝑓 refers to fake. By iterating this process across all indicators and facial regions, we construct a Semantic Anchor Library A = {𝐴𝑖,𝑔 | 𝑖 ∈ I, 𝑔 ∈ R}. For a forged image 𝑥 f , we target the most severe anomalies by selecting the top-𝑁 highest scores, and retrieve their corresponding forgery descriptions 𝑡 f . Conversely, for a real image 𝑥 r , we focus on the most pristine facial details by selecting the bottom-𝑁 lowest scores, assigning the linked authentic descriptions 𝑡 r .

3.3

Semantic-Guided Incremental Detector

Building upon the instantiated semantic anchors, we propose the Semantic-Guided Incremental Detector (SGID) to explicitly inject these semantic anchors into the visual learning process. The framework is underpinned by two components: Artifact-Probe Attention (APA) and Dual Supervision. Artifact-Probe Attention. To effectively incorporate the semantic priors, we introduce the APA module within the vision transformer. Let 𝑋 ∈ R𝑃 ×𝐷 denote the intermediate visual embeddings, where 𝑃 is the number of patches. Simultaneously, let 𝑆 = Φtext (A𝑚𝑎𝑡𝑐ℎ𝑒𝑑 ) ∈ R𝑁 ×𝐷 represent the textual embeddings of the matched semantic anchors, where 𝑁 is the number of selected anchors. The APA module functions as a cross-modal bridge, employing a Multi-Head Attention mechanism where visual tokens serve as queries to probe the semantic details in the embeddings: 𝑋˜ = MHA(𝑄 = 𝑋, 𝐾 = 𝑆, 𝑉 = 𝑆). (7) To ensure adaptive integration, we employ a residual connection with a learnable gating parameter 𝑔 ∈ R𝐷 : ˜ 𝑋 fused = 𝑋 + 𝑔 ⊙ 𝑋, (8) where ⊙ denotes element-wise multiplication. The gating coefficient 𝑔 dynamically modulates the injection of semantic priors, preventing the overriding of intrinsic visual cues. The fused representation 𝑋 fused then proceeds to the subsequent layers. In practice, APA is injected into the top-𝑀 transformer layers. Dual Supervision. To enforce the alignment between visual features and the selected semantic anchors, we employ a dual supervision strategy. Given an input image 𝑥, let 𝐹 = Φimg (𝑥, 𝑆) denote the visual features extracted by the APA-enhanced image encoder. First, a binary classifier 𝐶𝑡 is employed to predict global authenticity, formulated as: Lcls = CE(𝐶𝑡 (𝐹 ), 𝑌bin ),

(9)

where 𝑌bin ∈ {0, 1} is the ground-truth label (Real/Fake), and CE(·) denotes the standard cross-entropy loss. Simultaneously, to ensure the model comprehends the specific forgery patterns described by the semantic anchors, a multi-label head 𝐻𝑡 is utilized to predict the presence of the defined artifact dimensions (Iblur, Icolor, Istructure, Itext, Iboundary ). The artifact dimensions prediction loss is defined as: Lind = BCE(𝐻𝑡 (𝐹 ), 𝑌ind ),

(10)

where 𝑌ind ∈ {0, 1} | I | denotes the binary artifact indicator vector corresponding to the dimensions in I, and BCE(·) is the binary cross-entropy loss applied independently to each attribute. By jointly optimizing Lcls and Lind , the model is encouraged to align its visual features with the stable semantic anchors. The stationary semantic space acts as a persistent reference throughout the incremental learning process, effectively preventing feature drift and mitigating catastrophic forgetting.

3.4

Adaptive Decision Harmonizer

In incremental learning, the decision boundaries of classifiers tend to drift as new tasks are introduced. To mitigate this, we propose ADH to align classifier weights to preserve geometric consistency across tasks. Let 𝑊˜ 𝑗(𝑏 ) and 𝑊˜ 𝑗(𝑚) respectively denote the normalized classifier weights of the binary-label and multi-label heads from previous tasks. To quantify the semantic affinity between the current task and historical knowledge, ADH computes adaptive similarity weights for each head ℎ ∈ {𝑏, 𝑚}:  exp cos(𝑊˜ 𝑖 (ℎ) , 𝑊˜ 𝑗(ℎ) )/𝜏 (ℎ) 𝜔𝑗 = Í (11) , ˜ (ℎ) , 𝑊˜ (ℎ) )/𝜏 𝑙≠𝑖 exp cos(𝑊𝑖 𝑙 where 𝜏 is a temperature parameter and 𝑗 ≠ 𝑖. High similarity implies that the current artifacts share underlying semantic traits with task 𝑗, warranting stronger alignment. (ℎ) We then construct a Global Semantic Reference 𝑊˜ ref by aggregating previous classifiers according to their semantic affinity: ! ∑︁ (ℎ) (ℎ) (ℎ) 𝑊˜ = norm 𝜔 𝑊˜ . (12) ref

𝑗

𝑗

𝑗≠𝑖

This reference vector captures the stable historical decision trend, serving as a robust alignment anchor to prevent the model from forgetting previously learned forgery patterns. To incorporate the historical knowledge without distorting the feature space, we perform a spherical semantic alignment. Unlike Euclidean interpolation, this process respects the geometric structure of the hypersphere, rotating the current decision boundary toward the global reference along the geodesic path: sin((1 − 𝑡 (ℎ) )𝜃 ) ˜ (ℎ) sin(𝑡 (ℎ) 𝜃 ) ˜ (ℎ) (ℎ) 𝑊˜ new = 𝑊𝑖 + 𝑊ref , (13) sin 𝜃 sin 𝜃 where 𝜃 represents the angular distance between the current and reference weights. Crucially, the alignment coefficient 𝑡 (ℎ) ∈ [0, 1] is adaptively determined to balance plasticity and stability: (ℎ) 𝑡 (ℎ) = cos(𝑊˜ 𝑖 (ℎ) , 𝑊˜ ref ),

(14)

where a high cosine similarity indicates that the current task is semantically consistent with history, triggering a gentle update to preserve existing anchors. Finally, to decouple the semantic directional alignment from the magnitude-dependent feature strength, we rescale the aligned weights to recover the norm of 𝑊𝑖 (ℎ) : (ℎ) (ℎ) 𝑊new = 𝑊˜ new · ∥𝑊𝑖 (ℎ) ∥ 2,

(15)

which ensures that the alignment modifies only the semantic direction without degrading the detector’s discriminative sensitivity.

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

By constraining the update trajectory to the spherical manifold, ADH prevents the decision boundaries from drifting into semantically ambiguous regions, guaranting that the evolving detector remains geometrically consistent and semantically coherent with the accumulated global anchor library.

3.5

Training strategy and overall loss

Training strategy. During the initial 𝑛 epochs, the semantic anchors 𝑆 are fixed to the descriptions generated by ASPG, providing stable and consistent guidance to the backbone. Starting from the (𝑛 + 1)-th epoch, we activate the dynamic matching mechanism. In this phase, 𝑆 is adaptively retrieved by calculating the cosine similarity between visual features and textual embeddings, enabling the model to capture fine-grained, instance-specific forgery patterns. Overall loss. Following [24], we also maintain the previous-task learned information via knowledge distillation loss, which is: Ldis = ∥Φimg (𝑥, 𝑆; 𝜃 𝑡 ) − Φimg (𝑥, 𝑆; 𝜃 𝑡 −1 )∥ 22,

(16)

where Φimg (·; 𝜃 𝑡 −1 ) is the frozen backbone extractor trained on the previous (𝑡 − 1)-th task, which serves as a reference to regularize the feature space on the current data. The total loss function is defined as: Loverall = Lcls + 𝜇 1 Lind + 𝜇 2 Ldis,

(17)

where 𝜇1 and 𝜇2 are trade-off coefficients.

4 Experiment 4.1 Experimental Settings Datasets. Our datasets follow the protocol in [4] to ensure a comprehensive evaluation. We use several forgery face datasets, including DeepFake Detection Challenge Preview (DFDCP) [6], Celeb-DF-v2 (CDF) [19], and FaceForensics++ [27]. Furthermore, we incorporate recently released datasets that feature more diverse and sophisticated forgery methods, namely MCNet [10], BlendFace [32], StyleGAN3 [12] from DF40 [43] and SDv21 [26] from DiffusionFace [3]. Evaluation Protocol. To systematically evaluate the performance of our model, we adopt the standard evaluation protocols in [4]. • Protocol 1 (P1): Datasets Incremental with {SDv21, FF++, DFDCP, CDF}. This protocol simulates scenarios in which the model is required to adapt to entirely novel data environments across different incremental steps. • Protocol 2 (P2): Forgery Categories Incremental with {Hybrid (FF++), Face-Reenactment (MCNet), Face-Swapping (BlendFace), Entire Face Synthesis (StyleGAN3)}. The model continuously learns to defend against new forgery types, while the distribution of genuine data remains constant. Implementation Details. We use CLIP-ViT-L/14 [25] as backbone, fine-tuned via LN-tuning [47]. The Adam optimizer is employed with a learning rate of 8 × 10−5 , 20 epochs, and batch size of 32. For replay-based baselines, the replay buffer size is set to 500 for each task. We set 𝜇 1 =0.1, 𝜇2 =1 and set the number of selected semantic anchors to 𝑁 = 3. To ensure training stability, we set the warm-up period to 𝑛 = 5. The APA modules are injected into the last 𝑀 = 4 layers of the image encoder. We employ Frame-level Area Under the Curve (AUC) as the evaluation metric.

4.2

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Comparison with Other Methods

To comprehensively evaluate the effectiveness of our proposed framework, we conduct extensive comparisons under Protocol 1 (Cross-Dataset) and Protocol 2 (Cross-Manipulation). As reported in Tab.2, our method consistently outperforms all baselines across both protocols, achieving a superior trade-off between retaining past knowledge and adapting to new forgeries. Comparison with Replay-based IFFD Methods. Unlike replaybased approaches such as DFIL [24] and SUR-LID [4], which rely on storing historical samples to mitigate forgetting, our framework achieves better performance in a strict data-replay-free setting. This effectively highlights the efficiency of semantic anchors. By substituting raw data storage with invariant semantic priors, we effectively circumvent the storage constraints and privacy concerns inherent in replay-based paradigms. Comparison with General Replay-free ViT-based Methods. Directly transferring general Incremental Learning (IL) methods, including prompt-based (Coda-Prompt [33]) and adapter-based (CL-LoRA [9]) techniques, to the IFFD task leads to significant degradation. This reveals their inherent limitation: designed for object recognition, they primarily focus on high-level semantic content rather than subtle, low-level forgery artifacts. Consequently, they fail to decouple forensic traces from content, lacking the finegrained discriminative cues required for this specific task. Fairness Validation. To ensure a rigorous comparison and rule out the influence of backbone capacity, we replace the backbones of DFIL [24] and SUR-LID [4] with the ViT-L/14 used in our method. Even under this aligned configuration, our method still demonstrates clear superiority in performance. These results conclusively verify that our performance gains stem from the proposed methods rather than the raw power of the backbone. It underscores that the core challenge of IFFD lies in effective feature alignment, which cannot be solved solely by scaling up model parameters.

4.3

Ablation Study

Overall Ablation. As reported in Tab.3, the ablation study validates the indispensability of each component. The absence of ADH results in significant decision boundary drift. Removing Lind impairs feature disentanglement, degrading the model’s ability to explicitly encode specific artifacts. Furthermore, omitting APA weakens the cross-modal alignment, depriving the visual encoder of fine-grained semantic guidance. These results demonstrate that optimizing both decision harmonization and semantic-visual interaction is essential for IFFD, enabling the unified framework to mitigate catastrophic forgetting while adapting to novel forgery patterns. Effect of Alignment Method. We evaluate the impact of different classification-head alignment methods within ADH by comparing Linear (LERP), Exponential Moving Average (EMA), and Weighted Mean (WM) against our method. Unlike Euclidean-based methods that distort weight magnitudes, as shown in Tab.5, our method yields superior performance, outperforming baselines that suffer from high-dimensional directional drift (LERP) or delayed adaptation (EMA). This confirms that strictly constraining the update trajectory to the spherical space preserves semantic angular consistency, ensuring geometrically coherent decision boundaries that robustly mitigate catastrophic forgetting across incremental tasks.

Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Table 2: Performance comparisons (AUC) with Protocol 1 (Datasets Incremental) and Protocol 2 (Forgery Categories Incremental). Task 1 (T1) to Task 4 (T4) represent current incremented tasks in {SDv21, FF++, DFDCP, CDF} or {Hybrid, FR, FS, EFS}. The bold denotes the best ones. † indicates results copied from [4]. ‡ indicates that the backbones of DFIL [24] and SUR-LID [4] are replaced with ViT-L/14 in the experiments for a fair comparison. Method

Replays

Protocol 1

Task SDv21

FF++

DFDCP

Protocol 2 CDF

Avg.

Hybrid

FR

FS

EFS

Avg.

Methods with CNN Backbone CoReD (MM’21)† [15]

500

T1 T2 T3 T4

0.9998 0.7459 0.8555 0.8718

0.9433 0.9096 0.8376

0.8154 0.7987

0.9341

0.9998 0.8446 0.8602 0.8606

0.9665 0.9355 0.8907 0.8454

0.7988 0.7929 0.6429

0.8605 0.8417

0.9263

0.9665 0.8671 0.8480 0.8141

DFIL (MM’23)† [24]

500

T1 T2 T3 T4

0.9998 0.7400 0.9692 0.9326

0.9466 0.8164 0.7397

0.9088 0.7908

0.9881

0.9998 0.8433 0.8981 0.8628

0.9646 0.5574 0.6071 0.5083

0.9975 0.6649 0.9556

0.9903 0.7081

0.9996

0.9646 0.7775 0.7541 0.7929

HDP (IJCV’24)† [35]

500

T1 T2 T3 T4

0.9998 0.8373 0.9341 0.9055

0.9507 0.8532 0.8039

0.8737 0.8412

0.9501

0.9998 0.8940 0.8870 0.8752

0.9671 0.6741 0.6300 0.5989

0.9545 0.7135 0.7006

0.9509 0.8934

0.9373

0.9671 0.8143 0.7648 0.7826

SUR-LID (CVPR’25)† [4]

500

T1 T2 T3 T4

0.9999 0.9937 0.9986 0.9971

0.9485 0.8844 0.8479

0.9161 0.9067

0.9744

0.9999 0.9711 0.9330 0.9315

0.9685 0.8291 0.9050 0.8790

0.9242 0.9626 0.9679

0.9794 0.9356

0.9907

0.9685 0.8766 0.9490 0.9433

DFIL (ViT-L/14)‡

500

T1 T2 T3 T4

0.9999 0.8342 0.9897 0.9413

0.9581 0.8372 0.7523

0.9192 0.8241

0.9923

0.9999 0.8962 0.9154 0.8775

0.9682 0.5831 0.6598 0.5748

0.9981 0.7324 0.9627

0.9941 0.7597

0.9997

0.9682 0.7906 0.7954 0.8242

SUR-LID (ViT-L/14)‡

500

T1 T2 T3 T4

0.9999 0.9999 0.9997 0.9997

0.9090 0.9012 0.8838

0.9238 0.9394

0.9714

0.9999 0.9545 0.9416 0.9486

0.9406 0.9235 0.9043 0.9035

0.9902 0.9886 0.9898

0.9756 0.9773

0.9994

0.9406 0.9568 0.9562 0.9675

Coda-Prompt (CVPR’23) [33]

0

T1 T2 T3 T4

0.9999 0.7958 0.8564 0.8049

0.9323 0.8315 0.7492

0.9026 0.8131

0.9413

0.9999 0.8645 0.8635 0.8271

0.9631 0.6132 0.5816 0.5245

0.8231 0.7346 0.5938

0.8831 0.6264

0.9461

0.9631 0.7181 0.7331 0.6727

CL-LoRA (CVPR’25) [9]

0

T1 T2 T3 T4

0.9997 0.7892 0.8015 0.7742

0.9421 0.8896 0.8256

0.8991 0.8549

0.9626

0.9997 0.8656 0.8634 0.8543

0.9531 0.6216 0.5831 0.5367

0.9846 0.7216 0.5966

0.9733 0.7591

0.9988

0.9531 0.8031 0.7593 0.7228

0

T1 T2 T3 T4

0.9999 0.9999 0.9999 0.9988

0.9484 0.9397 0.9327

0.9999 0.9742 0.9554 0.9686

0.9719 0.9605 0.9447 0.9286

0.9948 0.9909 0.9897

0.9817 0.9892

0.9999

0.9719 0.9776 0.9724 0.9769

Methods with Vision Transformer

Traditional replay-free ViT-based Incremental Learning Methods

Our Method AIFIND(Ours)

0.9267 0.9551

Effect of APA Injection Layers. We investigate the impact of APA injection depth by evaluating five configurations: Low-level (1–4), Medium-level (11–14), High (21–24), Multi-level (5, 10, 15, 20) and All Layers (0–24). As shown in Tab. 4, the High-level setting yields superior performance, outperforming other configurations. This validates that semantic anchors align most effectively with high-level visual features that share similar semantic granularity, whereas other settings introduce semantic noise that tends to interfere with the extraction of high-level visual features. Hyperparameter Analysis. We further investigate the impact of key hyperparameters: the number of selected semantic anchors 𝑁 , the warm-up period 𝑛, the number of APA injection layers 𝑀 and the trade-off coefficients 𝜇 1 and 𝜇 2 . As shown in Tab.6, 𝑁 = 3 achieves the optimal trade-off between guidance and noise, 𝑛 = 5

0.9879

prevents premature alignment with immature features, and 𝑀 = 4 ensures sufficient interaction with high-level representations, and 𝜇1 = 0.1, 𝜇2 = 1 prove optimal for balancing the training objectives. Table 3: Ablation study (AUC) for each proposed component. Bold indicates the best performance. Variant

SDv21

FF++

DFDCP

CDF

Avg.

w/o All w/o ADH w/o APA w/o L𝑖𝑛𝑑

0.9671 0.9897 0.9896 0.9897

0.8746 0.9161 0.9156 0.9173

0.9035 0.9280 0.9283 0.9316

0.9513 0.9546 0.9597 0.9706

0.9241 0.9471 0.9483 0.9523

Ours

0.9988

0.9327

0.9551

0.9879

0.9686

AUC (%)

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

100 90 80 70 60 50

Color Saturation

0

2

Severity

100 90 80 70 60 50

4

Color Contrast

0

2

Severity

4

100 90 80 70 60 50

Block Wise

0

2

4

Severity HDP

DFIL

100 90 80 70 60 50

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Gaussian Blur

0

2

Severity SUR-LID

100 JPEG Compression 100 Gaussian Noise 90 90 80 80 70 70 60 60 50 50 0 2 4 0 2 4

4

Severity

Ours

Severity

Figure 3: Robustness under unseen perturbations (following Protocol 1, average AUC is used as the evaluation metric). Table 4: Performance comparisons (AUC) of APA injection layers. Bold indicates the best performance.

Table 6: Performance (Avg. AUC on Protocol 1) under different parameter settings. Bold indicates the best result.

Depth

SDv21

FF++

DFDCP

CDF

Avg.

Parameter

Low Medium Multi All

0.9108 0.9833 0.9466 0.9597

0.7003 0.9092 0.8144 0.9073

0.8751 0.9258 0.9011 0.9236

0.9695 0.9893 0.9639 0.9726

0.8639 0.9519 0.9065 0.9408

High (Ours)

0.9988

0.9327

0.9551

0.9879

0.9686

Effect of APA Gating Strategy. We evaluate the impact of the semantic injection gate within the APA module by comparing fixed scales {0.01, 0.1, 1} against a learnable parameter. Unlike static settings that impose a rigid injection intensity, as shown in Fig. 4, the learnable strategy yields superior performance, outperforming fixed scalars that suffer from either insufficient guidance (gate = 0.01) or excessive feature perturbation (gate = 1). This confirms that adaptively modulating the injection ratio enables the model to dynamically balance semantic anchor integration with visual preservation, ensuring optimal fine-grained alignment without disrupting the underlying feature topology. trainable

gate=0.01

gate=0.1

gate=1

1.00

0.98

0.98

0.96

0.96

AUC

AUC

1.00

0.94 0.92 0.90

trainable

gate=0.01

gate=0.1

gate=1

FF++

MCNet

BlendFace

Style-GAN3

0.94 0.92

SDv21

FF++

DFDCP

CDF

0.90

(a) Protocol 1

(b) Protocol 1

Figure 4: Ablation study on the gating strategy within the Artifact-Probe Attention (APA) module.

𝑛

𝑁

𝑀

𝜇1

𝜇2

4.4

Settings & Performance 0

5

10

15

20

0.9582

0.9686

0.9613

0.9601

0.9593

1

2

3

5

10

0.9588

0.9621

0.9686

0.9614

0.9522

1

2

3

4

5

0.9603

0.9611

0.9632

0.9686

0.9651

0.01

0.05

0.1

0.2

0.5

0.9592

0.9621

0.9686

0.9576

0.9531

0.1

0.5

1

1.5

2

0.9498

0.9572

0.9686

0.9581

0.9486

Generalization and Robustness Evaluations

To ensure a fair comparison, we standardize the backbone for all baseline methods to ViT-L/14, eliminating performance disparities caused by different backbones. Cross-Dataset Generalization. To rigorously evaluate the generalization capability of AIFIND against unseen domains, we conduct cross-dataset experiments. The model, fully trained under Protocol 1, is directly evaluated on four unseen benchmarks: DeepFakeDetection (DFD) [44], UniFace [41] (from DF40 [43]), SDv15 [26] (from DiffusionFace [3]), and FakeAVCeleb (FAVC) [14]. These datasets encompass a wide spectrum of generative mechanisms, ranging from conventional face swapping and GAN-based synthesis to diffusionbased text-to-image generation, which imposes significant challenges, requiring the model to overcome substantial domain shifts and rely on intrinsic, transferable forensic features rather than dataset-specific artifacts.

Table 5: Performance (AUC) under different classificationhead alignment method. Bold indicates the best result.

Table 7: Cross-dataset generalization results (AUC) on unseen datasets. Bold indicates the best performance.

Method

SDv21

FF++

DFDCP

CDF

Avg.

Method

DFD

UniFace

SDv15

FAVC

Avg.

LERP EMA WM

0.9905 0.9902 0.9935

0.9314 0.9286 0.9291

0.9421 0.9403 0.9462

0.9792 0.9807 0.9774

0.9608 0.9599 0.9615

DFIL HDP SUR-LID

0.8212 0.8342 0.8825

0.6549 0.6977 0.8279

0.8236 0.8231 0.8522

0.7157 0.7495 0.8143

0.7564 0.7761 0.8442

Ours

0.9988

0.9327

0.9879

0.9686

0.9686

Ours

0.9332

0.8987

0.8847

0.8755

0.8980

Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Original Original Original Original

Ours Ours Ours

SUR-LID SUR-LID SUR-LID

(a) SDv2.1 (a)SDv2.1 SDv2.1 (a) SDv2.1 (a)

Original Original Original

Ours Ours Ours

SUR-LID SUR-LID

Original Original Original

(b)FF++ FF++ (b) (b) FF++ (b)

Ours Ours Ours

SUR-LID SUR-LID SUR-LID

Ours Ours Ours

Original Original Original

(c) DFDCP (c) (c)DFDCP DFDCP (c) DFDCP

SUR-LID SUR-LID SUR-LID

(d) (d) CDF (d) CDF (d)CDF CDF

Figure 5: Visualization of Grad-CAM heatmaps across different datasets. Fake Artifact Probability Distribution asymmetry in eye shape or size 0.081 eye color inconsistent with face 0.054 eyes are blurry unnatural artifacts or seams on the jawline 0.025 mouth structure inconsistent 0.020 mouth area is blurry 0.016 abrupt gradient discontinuity at jawline 0.010 nose area is blurry 0.002 nose structure deviates 0.001 unnatural texture on lips or teeth 0.001 inconsistent lighting on the nose 0.000 unnatural texture on the cheeks 0.000 inconsistent lighting on the mouth 0.000 seam or splice detected along face edge 0.000 unnatural texture on the nose 0.000

Fake Artifact Probability Distribution 0.790

mouth area is blurry mouth structure inconsistent seam or splice detected along face edge unnatural artifacts or seams on the jawline inconsistent lighting on the mouth unnatural texture on lips or teeth abrupt gradient discontinuity at jawline asymmetry in eye shape or size eyes are blurry eye color inconsistent with face nose structure deviates nose area is blurry unnatural texture on the cheeks inconsistent lighting on the nose unnatural texture on the nose

Top 3 Artifacts Others

0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Softmax Probability

(a) Grad-CAM Heatmap

(b) Probability Distribution

0.026 0.012 0.010 0.002 0.001 0.001 0.001 0.000 0.000 0.000 0.000 0.000

0.275

0.340 0.330

Top 3 Artifacts Others

0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Softmax Probability

(c) Grad-CAM Heatmap

(d) Probability Distribution

Figure 6: Consistency Between Attention Heatmaps and Fake Artifact Probability Distributions. As shown in Tab.7, our framework achieves superior crossdataset generalization across all unseen benchmarks. Unlike baselines that may overfit to dataset-specific patterns, AIFIND learns intrinsic and transferable artifact representations, enabling robust detection even in unknown domains. Robustness Evaluations For robustness evaluation, we adopt the rigorous perturbation protocols from [11], which include five severity levels across six different perturbation types. As shown in Fig.3, our model consistently achieves higher AUC scores across all perturbation levels compared to other baselines. These results suggest that by aligning visual features with stable semantic anchors, AIFIND preserves discriminative ability under significant image distortions, ensuring reliable performance in practical scenarios.

4.5

Visualizations

We employ Grad-CAM [30] to visualize the attention maps of the model trained under Protocol 1. As illustrated in Fig. 5, the baseline SUR-LID [4] exhibits relatively dispersed attention, often appearing distracted by the background or irrelevant facial areas. In contrast, our method tends to concentrate more on semantically sensitive facial components, such as the eyes and mouth, which are notoriously prone to manipulation artifacts. This improved focus suggests that the semantic anchors effectively guide the model to attend to critical forensic regions rather than low-level noise, validating that our linguistic supervision successfully directs visual attention to physically meaningful areas. Furthermore, the results in Fig. 6 demonstrate a notable **spatialsemantic consistency** between the attention heatmaps and the inference outcomes. Specifically, the regions receiving high visual activation generally align with the artifact categories that yield

elevated predicted probabilities. For instance, when the heatmap highlights the eye region, the probability score for eye-related artifact classes rises distinctively. These observations indicate that our model learns to associate discriminative forensic traces with their corresponding semantic priors, thereby mitigating the risk of overfitting to spurious cues and enhancing the interpretability and reliability of the decision-making process.

5

Conclusion

In this paper, we propose AIFIND, an Artifact-Aware Interpreting Fine-Grained Alignment framework. Unlike traditional methods that rely on sample replay, AIFIND leverages a semantic anchor library to guide the model in learning invariant forgery representations. Our method ensures that visual features are aligned with stable semantic anchors, effectively mitigating catastrophic forgetting. Extensive experiments demonstrate that our method achieves stateof-the-art performance and exhibits spatial-semantic consistency in visualization. In the future, we plan to extend our framework to broader multimodal scenarios, exploring more adaptive semantic anchors via Large Vision-Language Models to tackle increasingly diverse forgery patterns in open-world settings.

Acknowledgments This work is partially supported by the National Natural Science Foundation of China under Grants 62441232, 62476068, 62306092, 62502115, and projects ZR2025ZD01, ZR2024QF066, ZR2025QC1516 supported by Shandong Provincial Natural Science Foundation, and projects 2024DXZD0004 supported by Inner Mongolia Department of Science and Technology.

AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection

References [1] Hunar Batra and Ronald Clark. 2024. Evcl: Elastic variational continual learning with weight consolidation. arXiv preprint arXiv:2406.15972 (2024). [2] Pietro Buzzega, Matteo Boschini, Angelo Porrello, and Simone Calderara. 2021. Rethinking experience replay: a bag of tricks for continual learning. In 2020 25th International Conference on Pattern Recognition. 2180–2187. [3] Zhongxi Chen, Ke Sun, Ziyin Zhou, Xianming Lin, Xiaoshuai Sun, Liujuan Cao, and Rongrong Ji. 2024. Diffusionface: Towards a comprehensive dataset for diffusion-based face forgery analysis. arXiv:2403.18471 (2024). [4] Jikang Cheng, Zhiyuan Yan, Ying Zhang, Li Hao, Jiaxin Ai, Qin Zou, Chen Li, and Zhongyuan Wang. 2025. Stacking brick by brick: Aligned feature isolation for incremental face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13927–13936. [5] Xinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou, and Junyu Dong. 2025. Forensics adapter: Adapting clip for generalizable face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19207– 19217. [6] Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, and Cristian Canton Ferrer. 2019. The Deepfake Detection Challenge (DFDC) Preview Dataset. [7] Hui Guo, Shu Hu, Xin Wang, Ming-Ching Chang, and Siwei Lyu. 2022. Eyes tell all: Irregular pupil shapes reveal gan-generated faces. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing. 2904–2908. [8] Zonghui Guo, Yingjie Liu, Jie Zhang, Haiyong Zheng, and Shiguang Shan. 2025. Face Forgery Video Detection via Temporal Forgery Cue Unraveling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7396–7405. [9] Jiangpeng He, Zhihao Duan, and Fengqing Zhu. 2025. CL-LoRA: Continual LowRank Adaptation for Rehearsal-Free Class-Incremental Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 30534– 30544. [10] Fa-Ting Hong and Dan Xu. 2023. Implicit identity representation conditioned memory compensation network for talking head video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision. [11] Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. 2020. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2889–2898. [12] Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks. Advances in Neural Information Processing Systems 34 (2021), 852–863. [13] Hossein Kashiani, Niloufar Alipour Talemi, and Fatemeh Afghah. 2025. Freqdebias: Towards generalizable deepfake detection via consistency-driven frequency debiasing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8775–8785. [14] Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. 2021. FakeAVCeleb: A novel audio-video multimodal deepfake dataset. arXiv preprint arXiv:2108.05080 (2021). [15] Minha Kim and Shahroz Tariq. 2021. Cored: Generalizing fake media detection with continual representation using distillation. In Proceedings of the 29th ACM International Conference on Multimedia. 337–346. [16] Youngeun Kim, Yuhang Li, and Priyadarshini Panda. 2024. One-stage promptbased continual learning. In European Conference on Computer Vision. 163–179. [17] Jiashuo Li, Shaokun Wang, Bo Qian, Yuhang He, Xing Wei, Qiang Wang, and Yihong Gong. 2025. Dynamic integration of task-specific adapters for class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 30545–30555. [18] Yuezun Li, Ming-Ching Chang, and Siwei Lyu. 2018. In Ictu Oculi: Exposing AI Generated Fake Face Videos by Detecting Eye Blinking. In IEEE International Workshop on Information Forensics and Security. [19] Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. 2020. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [20] Kaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao, and Bin Li. 2025. Standing on the shoulders of giants: Reprogramming visual-language model for general deepfake detection. In Proceedings of the AAAI Conference on Artificial Intelligence. 5262–5270. [21] Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, WanTeh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. 2019. MediaPipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172 (2019). [22] Sriram Mandalika, Harsha Vardhan, and Athira Nambiar. 2025. Replay to Remember (R2R): An Efficient Uncertainty-driven Unsupervised Continual Learning Framework Using Generative Replay. arXiv preprint arXiv:2505.04787 (2025). [23] Seyed-Mohsen Moosavi-Dezfooli and Alhussein Fawzi. 2017. Universal adversarial perturbations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1765–1773.

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

[24] Kun Pan, Yifang Yin, Yao Wei, Feng Lin, Zhongjie Ba, Zhenguang Liu, Zhibo Wang, Lorenzo Cavallaro, and Kui Ren. 2023. Dfil: Deepfake incremental learning by exploiting domain-invariant forgery clues. In Proceedings of the 31st ACM International Conference on Multimedia. 8035–8046. [25] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PmLR, 8748–8763. [26] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10684–10695. [27] Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 1–11. [28] Anurag Roy, Riddhiman Moulick, Vinay K Verma, Saptarshi Ghosh, and Abir Das. 2024. Convolutional prompting meets language models for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23616–23626. [29] Krisanu Sarkar. 2025. Adaptive Variance-Penalized Continual Learning with Fisher Regularization. arXiv preprint arXiv:2508.16632 (2025). [30] Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, and Devi Parikh. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 618–626. [31] Kaede Shiohara and Toshihiko Yamasaki. 2022. Detecting deepfakes with selfblended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18720–18729. [32] Kaede Shiohara, Xingchao Yang, and Takafumi Taketomi. 2023. Blendface: Redesigning identity encoders for face-swapping. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7634–7644. [33] James Seale Smith, Leonid Karlinsky, Vyshnavi Gutta, Paola Cascante-Bonilla, Donghyun Kim, Assaf Arbelle, Rameswar Panda, Rogerio Feris, and Zsolt Kira. 2023. Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11909–11919. [34] James Seale Smith, Lazar Valkov, Shaunak Halbe, Vyshnavi Gutta, Rogerio Feris, Zsolt Kira, and Leonid Karlinsky. 2024. Adaptive memory replay for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3605–3615. [35] Ke Sun, Shen Chen, Taiping Yao, Xiaoshuai Sun, Shouhong Ding, and Rongrong Ji. 2025. Continual face forgery detection via historical distribution preserving. International Journal of Computer Vision 133, 3 (2025), 1067–1084. [36] Ke Sun, Shen Chen, Taiping Yao, Ziyin Zhou, Jiayi Ji, Xiaoshuai Sun, ChiaWen Lin, and Rongrong Ji. 2025. Towards general visual-linguistic face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19576–19586. [37] Jiahe Tian, Cai Yu, Xi Wang, Peng Chen, Zihao Xiao, Jizhong Han, and Yesheng Chai. 2024. Dynamic mixed-prototype model for incremental deepfake detection. In Proceedings of the 32nd ACM International Conference on Multimedia. 8129– 8138. [38] Qiang Wang, Xiang Song, Yuhang He, Jizhou Han, Chenhao Ding, Xinyuan Gao, and Yihong Gong. 2025. Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4839–4849. [39] Zhicheng Wang, Yufang Liu, Tao Ji, Xiaoling Wang, Yuanbin Wu, Congcong Jiang, Ye Chao, Zhencong Han, Ling Wang, Xu Shao, et al. 2023. Rehearsal-free continual language learning via efficient parameter isolation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10933–10946. [40] Zifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang, Ruoxi Sun, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer Dy, and Tomas Pfister. 2022. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [41] Chao Xu, Jiangning Zhang, Yue Han, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang, and Yong Liu. 2022. Designing one unified framework for high-fidelity face reenactment and swapping. In European Conference on Computer Vision. Springer, 54–71. [42] Zhiyuan Yan, Yuhao Luo, Siwei Lyu, Qingshan Liu, and Baoyuan Wu. 2024. Transcending forgery specificity with latent space augmentation for generalizable deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8984–8994. [43] Zhiyuan Yan, Taiping Yao, Shen Chen, Yandan Zhao, Xinghe Fu, Junwei Zhu, Donghao Luo, Chengjie Wang, Shouhong Ding, Yunsheng Wu, et al. 2024. Df40: Toward next-generation deepfake detection. Advances in Neural Information Processing Systems 37 (2024), 29387–29434.

ICMR ’26, June 16-19, 2026, Amsterdam, Netherlands

Hao Wang, Beichen Zhang∗ , Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu, Weigang Zhang

[44] Zhiyuan Yan, Yong Zhang, Xinhang Yuan, Siwei Lyu, and Baoyuan Wu. 2023. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. In Advances in Neural Information Processing Systems, Vol. 36. 4534–4565. [45] Zhiyuan Yan, Yandan Zhao, Shen Chen, Mingyi Guo, Xinghe Fu, Taiping Yao, Shouhong Ding, Yunsheng Wu, and Li Yuan. 2025. Generalizing deepfake video detection with plug-and-play: Video-level blending and spatiotemporal adapter tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12615–12625. [46] Xin Yang, Yuezun Li, and Siwei Lyu. 2019. Exposing deep fakes using inconsistent head poses. In ICASSP 2019-2019 IEEE international conference on acoustics, speech

and signal processing. 8261–8265. [47] Andrii Yermakov, Jan Cech, and Jiri Matas. 2025. Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection. arXiv (2025). [48] Jiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu, Dong Wang, Huchuan Lu, and You He. 2024. Boosting continual learning of vision-language models via mixture-ofexperts adapters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 23219–23230. [49] Jiaran Zhou, Yuezun Li, Baoyuan Wu, Bin Li, Junyu Dong, et al. 2024. Freqblender: Enhancing deepfake detection by blending frequency knowledge. Advances in Neural Information Processing Systems 37 (2024), 44965–44988.

Record · ID 31309 · SHA-256 8ab1ab8f67dd34bc
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.