Research Article
YUV20K: A Complexity-Driven Benchmark and TrajectoryAware Alignment Model for Video Camouflaged Object Detection
arXiv:2604.09985v1 [cs.CV] 11 Apr 2026
Yiyu Liu † · Shuo Ye † · Chao Hao · Zitong Yu ∗
·
Received: date / Accepted: date
Abstract Video Camouflaged Object Detection (VCOD) is currently constrained by the scarcity of challenging benchmarks and the limited robustness of models against erratic motion dynamics. Existing methods often struggle with Motion-Induced Appearance Instability and Temporal Feature Misalignment caused by complex motion scenarios. To address the data bottleneck, we present YUV20K, a pixel-level annoated complexitydriven VCOD benchmark. Comprising 24,295 annotated frames across 91 scenes and 47 kinds of species, it specifically targets challenging scenarios like largedisplacement motion, camera motion and other 4 types scenarios. On the methodological front, we propose a novel framework featuring two key modules: Motion Feature Stabilization (MFS) and Trajectory-Aware Alignment (TAA). The MFS module utilizes frame-agnostic Semantic Basis Primitives to stablize features, while the TAA module leverages trajectory-guided deformable sampling to ensure precise temporal alignment. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art competitors on existing datasets and establishes a new baseline on the challenging YUV20K. Notably, our framework exhibits superior cross-domain generalization and robustness when confronting complex spatiotemporal scenarios. Our code and dataset will be available at https: //github.com/K1NSA/YUV20K Keywords Camouflaged Object detection · Video Datasets · Deformable Alignment †
The authors contributed equally to this work. ∗ Correspondence: [email protected] Yiyu Liu, Shuo Ye, Chao Hao and Zitong Yu are with the Department of Computer Science, Great Bay University, Dongguan, Guangdong, 523000, China (Email: [email protected], [email protected], [email protected] and [email protected]).
1 Introduction Camouflaged Object Detection (COD) aims to identify and segment objects that exhibit high visual similarity to their surroundings [1]. Unlike generic object detection that focuses on salient targets, COD tackles scenarios where the foreground and background share homogeneous textures, luminance, and patterns. Consequently, this task requires models to effectively decouple these confusing visual cues and accurately localize subtle structural boundaries. Consequently, COD not only facilitates practical applications in medical image segmentation [2,3], agricultural pest monitoring [4] but also drives the development of more robust and perceptive computer vision systems. Image COD (ICOD) focuses on detecting camouflaged targets within static images. Since these targets exhibit extreme spatial consistency with their surroundings, ICOD models are restricted to relying solely on static appearance cues, making the detection highly challenging. Consequently, motion cues for video camouflaged object detection effectively break the targetbackground homogeneity [5, 6], transforming the illposed static problem into a tractable spatiotemporal inference task. However, while temporal motion cues are crucial for breaking camouflage, they simultaneously introduce a new layer of complexity that has been overlooked. Existing methods typically operate under the assumption of stable motion, characterized by linear trajectories and rigid poses [5]. In contrast, real-world camouflaged creatures exhibit highly complex motion dynamics, such as an octopus squeezing through a narrow crevice (nonrigid deformation) or a predator launching a lethal strike (large-displacement motion). Overlooking these scenarios introduces two fundamental impediments to effective
2
Yiyu Liu et al
rigid sampling on non-rigid target
T T+1 deformable sampling on non-rigid target
T
T+1
misalignment
{Framet }T=5 t=1 correct alignment
{Framet }T=5 t=1
Fig. 1: Top: Misalignment of rigid sampling convolutions. When handling non-rigid moving objects, standard sampling convolutions with fixed sampling grids fail to adapt to the target’s continuous deformation. This mismatch causes the network to inevitably sample irrelevant background regions, indicated by the blue points, rather than the target features highlighted in yellow. This leads to severe background noise accumulation over time. Bottom: Alignment via deformable sampling. By dynamically adjusting the sampling grid, our deformable strategy faithfully tracks the target’s features across frames, ensuring all sampled points remain yellow. This actively avoids background interference, mitigating noise accumulation and ensuring accurate spatiotemporal feature alignment. segmentation: Motion-Induced Appearance Instability and Temporal Feature Misalignment. First, in the spatial domain, complex motion induces appearance instability. Rapid motion results in blur that obliterates texture details [7], while slow motion renders motion cues imperceptible, hindering effective figure-ground separation [8]. Second, in the temporal domain, irregular motion leads to severe feature misalignment. Camouflaged targets frequently undergo extreme non-rigid deformations or erratic displacements, which fundamentally violate the rigid grid assumptions of standard aggregation modules (e.g., 3D convolutions) [9]. As visually demonstrated in Fig. 1 (Top), when handling such non-rigid targets, standard rigid sampling convolutions fail to adapt to continuous deformations. Consequently, the fixed sampling grid inevitably captures irrelevant background regions, causing severe noise accumulation and spatial misalignment across frames.
feature space. Functioning as robust anchors during global query operations, the SBPs enable the model to effectively lock onto the camouflaged target within the global feature space. To resolve Temporal Feature Misalignment, driven by the dynamic alignment depicted in Fig. 1 (Bottom), we propose the Trajectory-Aware Alignment (TAA) module. Drawing inspiration from deformable paradigms [10, 11], TAA moves beyond rigid aggregation. Specifically, it leverages temporal features to explicitly predict pixel-wise motion offsets, effectively modeling the target’s trajectory. These offsets are then applied to warp the spatial sampling grid, ensuring that features are aggregated strictly along the object’s actual deformation path. Unlike standard sampling which performs indiscriminate averaging on a fixed grid, our trajectory-guided deformable sampling precisely aligns distinct features, significantly optimizing spatiotemporal feature integrity.
To address Motion-Induced Feature Instability, we propose the Motion Feature Stabilization module (MFS). Within this module, we introduce a set of Semantic Basis Primitives (SBPs) initialized via a Gaussian distribution. Designed as frame-agnostic learnable parameters, these primitives remain invariant across dynamic input frames. By injecting them into the feature stream, we augment the semantic representation, expand the
Beyond algorithmic design, the field is constrained by the scarcity of benchmarks that fully capture the complexity of VCOD [14]. While pioneering datasets have established a solid baseline, capturing the full spectrum of erratic dynamics inherent to wild biological behaviors remains an open challenge. This leaves a critical gap in representing challenging scenarios, such as large-displacement dashes, severe non-rigid deformations
YUV20K
3
Table 1: Comparison of YUV20K with existing VCOD datasets. YUV20K provides the most comprehensive annotations. Dataset
Venue
videos/frames
species
Class
B.Box
Mask
CAD [12] MOCA [5] MOCA-Mask [13]
ECCV’16 ACCV’20 CVPR’22
9/839 141/37,250 87/22,939
6 67 44
✓ × ✓
× ✓ ×
✓ × ✓
-
91/24,295
47
✓
✓
✓
YUV20K
and multi-object interactionss gap, we present YUV20K. Comprising 91 clips, YUV20K is positioned as a challenging and scenario diverse benchmark. By including a higher proportion of complex scenarios (e.g., large displacement motion, camera motion), YUV20K aims to expose the limitations of current models and drive the development of more robust spatiotemporal reasoning systems. Our main contributions are summarized as follows: – We propose the Motion Feature Stabilization (MFS) module to address motion-induced appearance instability. By introducing frame-agnostic Semantic Basis Primitives as anchors, MFS effectively stabilizes semantic features against blur and rapid changes, enabling robust target locking even under erratic motion dynamics. – We design the Trajectory-Aware Alignment (TAA) module to resolve temporal feature misalignment. Moving beyond rigid grid assumptions, TAA utilizes trajectory-guided deformable sampling to predict pixel-wise motion offsets, ensuring precise feature aggregation along the object’s complex deformation path. – We present YUV20K, a complexity-driven benchmark comprising 91 clips that captures the full spectrum of wild biological behaviors. Extensive experiments demonstrate that our method achieves stateof-the-art performance on this challenging dataset and existing benchmarks, setting a new baseline for robust video camouflaged object detection.
2 Related Work 2.1 Camouflaged Object Detection Image-based COD aims to identify concealed objects from a single static RGB image. This task is inherently challenging because camouflaged targets share high intrinsic similarities with their backgrounds. Biological studies [15] reveal that predatory animals typically employ a two-step strategy to break camouflage: first scanning the environment to locate potential prey, followed by precise identification. Inspired by this mechanism,
Fan et al. [1, 16] proposed a coarse-to-fine framework that initially generates a localization map to discover the target, which is subsequently refined into a precise pixel-level segmentation mask. Another line of research attempts to tackle the high visual similarity by exploring alternative feature spaces and relational modeling. For instance, recent works [17, 18] utilize frequency cues, fusing complementary information from the frequency domain with spatial representations to discern subtle differences. Meanwhile, methods like MGL [19] introduce mutual graph learning to explicitly decouple the underlying topological relationship between the camouflaged object and its surroundings. With the development of deep learning architectures, recent methods rely heavily on multi-scale feature integration. Models like ZoomNet [20] and FSPNet [21] effectively model context at both global and local scales, capturing scale-specific semantics. Despite these advancements, image-based methods rely solely on static appearance cues, making them vulnerable to complex scenes where boundaries remain completely ambiguous without temporal information.
2.2 Video Camouflaged Object Detection Compared to static images, motion cues between video frames provide critical information for breaking camouflage. Existing VCOD strategies primarily focus on how to effectively utilize spatial-temporal information, which can be broadly categorized into explicit and implicit motion modeling. Explicit motion methods typically employ off-theshelf optical flow estimators (e.g., RAFT [22]) to separate moving foregrounds from backgrounds. Lamdouar et al. [5] take optical flow and difference images as inputs, designing a differentiable registration module for background alignment and a motion segmentation module to discover moving objects. Similarly, recent two-stream architectures [23] simultaneously conduct camouflaged object segmentation and explicit optical flow estimation. However, explicitly extracting optical flow introduces several challenges. It struggles when the camouflaged object remains stationary or under severe camera motion. Furthermore, due to the repetitive textures of camouflaged objects, pixel-level explicit motion estimation is often noisy and incurs considerable computational expense [24]. To overcome the limitations of explicit optical flow, recent works have shifted toward learning implicit interframe feature correspondences for temporal alignment. SLT-Net [13] introduced a model that implicitly captures motion over both short and long temporal intervals,
4
ensuring temporal consistency while directly predicting pixel-level masks. To further reduce architectural redundancy, recent trends aim to unify static and dynamic processing. For example, ZoomNeXt [25] uses a difference-aware routing mechanism to adaptively propagate inter-frame temporal cues. Additionally, models like SAM-PM [26] explore spatio-temporal cross-attention mechanisms to enforce temporal consistency across consecutive frames. Despite these improvements, existing methods still struggle with motion-induced appearance instability and temporal feature misalignment under complex trajectories. Furthermore, while pioneering datasets like MoCA [5], CAD [12], and MoCA-Mask [13] have significantly advanced VCOD, they primarily serve as foundational benchmarks. As algorithms evolve, evaluating spatial-temporal reasoning under highly complex dynamic scenes (e.g., severe environmental degradation and sustained complex trajectories) demands benchmarks with larger scales and higher annotation densities.
Yiyu Liu et al
overall scene diversity and motion complexity remain limited. Similarly, another available dataset, CAD [12], contains only 9 short video sequences, which is far from sufficient for comprehensive model evaluation. Compared to the abundance of ICOD data, the limited scale and lack of challenging scenarios in current VCOD datasets severely hinder the rigorous evaluation of spatiotemporal models [14]. To bridge this critical gap and provide a robust testing ground for the community, we introduce the YUV20K dataset, featuring highly diverse distributed environments, complex scenarios, and high-quality frame-level annotations.
3.1 Dataset Construction
Dataset Collection and Annotation. To ensure biological diversity and scene complexity, we curated raw videos from open-source platforms (e.g., YouTube, Google) using a diverse set of keywords covering 47 animal families. We strictly screened the videos based on three criteria: (1) High Camouflage Degree, excluding instances with obvious salient features; (2) Motion Richness, prioritizing clips containing erratic movements or deformations; and (3) High Resolution, ensuring sufficient detail for pixel-level annotation. This process yielded 91 video clips, comprising a total of 24,295 frames with an average length of 266 frames per CAD MoCA-mask YUV20K clip. Annotation Pipeline. Annotating camouflaged obFig. 2: Spatial distribution density maps. Comjects frame-by-frame is labor-intensive and error-prone. pared to CAD and MoCA-Mask, where targets tend To balance efficiency and quality, we employed a semito appear in the central area, our YUV20K provides a automatic pipeline. First, we utilized X-AnyLabeling broader and more uniform spatial distribution. This di[30] integrated with state-of-the-art foundation models YUV20Kfor robust verse placement introduces greater challenges MoCA (SAM 2 [31] and Grounding SAM [32]) to generate inisegmentation. tial coarse masks. Second, these proposals underwent a rigorous manual refinement phase to correct boundary errors, with special attention given to motion-blurred regions and intricate biological structures (e.g., feathers, limbs). Finally, a cross-validation round was conducted 3 YUV20K Dataset to resolve any remaining ambiguities. Diversity Explanation Category Diversity. As illusLarge-scale, high-quality datasets serve as the cornertrated in Tab. 1, our dataset comprises 47 distinct anistone for advancing and rigorously evaluating detection mal species, providing a rich taxonomic diversity that algorithms. The rapid progress in Image COD (ICOD) is spans a broad spectrum of ecological niches. This comlargely attributed to the availability of well-established prehensive coverage exposes models to a vast array of benchmarks, such as CAMO [27], CHAMELEON [28], biological camouflage strategies, effectively preventing COD10K [1], and NC4K [29]. algorithms from overfitting to specific appearance priConversely, the evolution of Video COD (VCOD) ors and ensuring robust generalization across unseen has been heavily bottlenecked by data scarcity. For species. instance, the pioneering VCOD dataset, MoCA [5], primarily provides bounding box annotations rather than Habitat Diversity. As illustrated in Fig. 3 (a) and dense pixel-level masks. Cheng et al. [13] subsequently (b), we construct a wide range distribution of ecologiintroduced MoCA-Mask to supply pixel-level labels, the cal habitats to accurately reflect real-world scenarios.
YUV20K
(a)
5
Aquatic Animals
Flying Animals
(b)
Terrestrial Animals
(c) M-Obj
Hunt 20
20
18
18
18
16
16
16
14
14
14
12
12
12
10
10
10
8
8
8
6
6
6
4
4
4
2
2
2
0
0
CAD
octopus
rabbit
colugo
Terrestrial
T-Obj
20
MoCA YUV20K
Ldm
CM
45
0
CAD
MoCA YUV20K Occ
45
40
40
40
35
35
35
30
30
30
25
25
25
20
20
20
15
15
15
10
10
10
5
5
5
0
0
CAD
MoCA YUV20K
CAD
MoCA YUV20K
CAD
MoCA YUV20K
45
0
CAD
MoCA YUV20K
Fig. 3: Overview of the proposed YUV20K dataset. (a) Representative frames showcasing diverse camouflage strategies across aquatic, aerial, and terrestrial environments. (b) A hierarchical sunburst chart illustrating the rich taxonomic diversity of the annotated animal species. (c) Attribute-based comparison with existing datasets (CAD and MoCA-Mask) across six complex scenarios: Large Displacement Motion (Ldm), Camera Motion (CM), Occlusion (Occ), Multiple Objects (M-Obj), Hunting (Hunt), and Tiny Object (T-Obj). YUV20K provides a significantly larger number of challenging videos across all attributes, notably establishing a new benchmark in extreme cases like M-Obj and Hunt, where previous datasets are severely lacking.
6
The dataset systematically covers terrestrial, aquatic, and aerial domains, thereby ensuring a comprehensive representation of diverse background clutter, varying complex environmental structures. Scenario Diversity. To faithfully reflect real-world complexities, YUV20K incorporates a wide spectrum of challenging spatiotemporal scenarios. As detailed in Fig. 3 (c), the dataset encompasses not only complex motion dynamics—such as camera motion, large displacements, and non-rigid deformations—but also diverse target attributes and behaviors, including hunting events, tiny-scale targets, and multi-object instances. This comprehensive profile provides a rigorous testbed for evaluating the robustness of VCOD algorithms in the wild. Spatial Distribution Diversity A common limitation in video datasets is a noticeable central tendency, where subjects frequently appear near the center of the frame due to human filming habits. As illustrated in Fig 2, the spatial density maps reveal that YUV20K provides a substantially broader and more uniform spatial distribution. This extensive spatial coverage effectively mitigates center bias, requiring models to develop robust spatiotemporal segmentation capabilities.
4 Method 4.1 Overall Architecture As illustrated in Fig. 4, our framework is designed to process video sequences and accurately segment camouflaged objects by overcoming motion-induced instability and temporal feature misalignment. Given a video clip T consisting of 5 consecutive frames, denoted as I = {Its } , s = 0.5, 1, 1.5. We first feed them into a Triplet Feature Encoder [33] to extract hierarchical spatial features. While these multi-scale features capture rich spatial details, they are fundamentally vulnerable to the complex motion dynamics inherent in real-world camouflaged scenarios. To tackle this, the extracted features are first processed by the self-attention with Motion Feature Stabilization (MFS) module. The MFS introduces robust semantic anchors to alleviate the appearance instability caused by motion or severe deformations, ensuring feature consistency within the spatial domain. The output feature is defined as the spatial features. To further capture inter-frame motion dynamics, the stabilized features are fed into the 3DCDCST module. The output of this module effectively encapsulates continuous motion cues and is defined as the temporal features.Finally, to resolve the issue of temporal misalignment, the Trajectory-Aware Alignment (TAA)
Yiyu Liu et al
module takes both features as input. It leverages the extracted temporal features to explicitly predict motion offsets, which are then used to perform deformable sampling on the spatial features, achieving precise spatiotemporal alignment. The aligned, multi-scale features are ultimately aggregated and fed into a decoder to generate the final camouflaged object masks. In the following subsections, we detail the design of these core components.
4.2 Motion Feature Stabilization module 4.2.1 Semantic Basis Primitives Complex motion inevitably induces appearance instability (e.g., blur), causing significant disturbances in feature representations. Consequently, a stable reference or is requisite to rectify these features. To address this, we propose the Motion Feature Stabilization (MFS) module. Drawing inspiration from recent advances in fine-grained visual classification and confusing image recognition [34–36], we shift our attention to the intrinsic semantic attributes of camouflaged targets (e.g., biological textures like skin, feather), termed as Semantic Basis Primitives (SBPs). Unlike dynamic video features that fluctuate over time, these primitives function as anchors. They are injected into the feature stream to implicitly discover and aggregate generalized camouflage concepts. By utilizing these anchors to perform global query operations, we can leverage their stability to probe the motion-perturbed feature space and efficiently localize the camouflaged targets. Technically, we instantiate the SBPs as a set of global learnable parameters, independent of the input video frames. To enhance the discriminative power specifically to distinguish subtle camouflaged patterns from complex backgrounds, we structure them into paired − N components. Let P = {(p+ i , pi )}i=1 denote the set of + N primitive pairs, where pi ∈ RC represents the posiC tive capturing target-related semantics, and p− i ∈ R represents the negative responsible for suppressing background semantics. We initialize these embeddings via a standard Gaussian distribution to ensure a diverse semantic starting point. This paired structure allows the module to explicitly model both the presence of camouflaged attributes and the absence of background distractors, forming a comprehensive concept basis. 4.2.2 Stabilization Injection For each semantic primitive i, a learnable logit li governs the mixing probability αi = σ(li ), where σ represents for sigmoid operation. The final concept token ci is
YUV20K
7
{I𝑡1.0 }𝑇𝑡=1
{I𝑡1.5 }𝑇𝑡=1
{I𝑡0.5 }𝑇𝑡=1 P+
Probability Gate
Semantic Basis Primitives query Q
Semantic Injection
P-
Triplet Feature Encoder
Self-Attention
Motion Feature Stabilization
3D-CDCST
Spatial Feature
Temporal Feature
Trajectory Aware Alignment
query Q’
Feature Map (T=5)
Prediction Mask
Fig. 4: Overall architecture of the proposed framework. Given a video sequence, we construct a multi-scale input pyramid (0.5×, 1.0×, 1.5×) and extract hierarchical features via a shared backbone. Then features are refined by self-attention with Motion Feature Stabilization module to capture global context. The next step decouples the representations into two pathways. In the temporal pathway, the 3D Central DIfference Spatial-Temporal Convolution modules extract robust temporal cues resilient to severe motion dynamics. In spatial way, the output of self-attention directly push into Trajectory-Aware Alignment module(TAA). Guided by these temporal features, the TAA module dynamically samples and aligns the spatial features. Finally, a decoder aggregates these aligned representations for accurate camouflage mask prediction. synthesized via convex combination. This mechanism allows the module to adaptively emphasize or suppress specific camouflage patterns.
C=
N X − [αi p+ i + (1 − αi )pi ]
(1)
i
Q′ = XWQ′ ,
K = XWK ,
Fs = Softmax
Q′ K⊤ √ dk
V = XWV ,
(3)
V.
(4)
. Upon obtaining the stable concept tokens, we integrate them into the feature stream via a self-attention mechanism. Specifically, we inject these vectors into the feature map Fs generated query Q to formulate an augmented query set, denoted as Q′ , which can be seen in top-right of Fig. 4. This augmented representation is subsequently linearly projected to generate the corresponding Keys (K) and Values (V). By equipping the query with these stable concept vectors, the attention mechanism gains the capability to globally screen motion-induced interference and actively probe for generalized camouflage attributes (e.g., texture patterns). Consequently, this interaction yields a rectified feature map, where the target representation is stabilized against dynamic perturbations. X = Concat[Fs , Linear(C)],
(2)
4.3 Trajectory Aware Alignment 4.3.1 3D Central Difference Spatial-Temporal Convolution In VCOD scenarios, the efficacy of trajectory alignment heavily relies on the quality of motion guidance. However, due to the high visual homogeneity between the target and the background, standard feature extractors (e.g., vanilla 3D convolutions) that primarily aggregate intensity information often fail to capture reliable motion boundaries. Consequently, the subsequent offset prediction module (TAA) lacks reliable guidance, ultimately leading to severe feature misalignment. To effectively guide trajectory alignment, it is crucial to extract fine-grained motion cues while suppressing static background textures—a task where standard
8
Yiyu Liu et al
intensity-based 3D convolutions often fail. To address this, we introduce the 3D Central Difference Spatial Temporal Convolution (3D-CDCST) [37, 38], which integrates gradient-level differential information into spatiotemporal aggregation. Specifically, we formulate 3DCDCST as a hyperparameter-controlled fusion of a vanilla intensity term and a central difference gradient term. For efficient implementation, we derive it as a unified operator that shares weights without introducing extra parameters:
y(p0 ) = θ ·
X
w(pn )(x(p0 + pn ) − x(p0 ))
pn
|
{z
}
Gradient Term
+ (1 − θ) ·
X
∆P = ϕof f (Ft ) ,
| =
{z
Vanilla Term
(5)
w(pn ) · x(p0 + pn ) {z
}
Standard 3D Conv
− θ · x(p0 ) ·
X
w(pn ),
pn ∈R
|
{z
M = ϕmod (Ft ) ∈ RT ×H×W .
(6)
}
pn ∈R
|
Standard sampling operations rely on fixed grids, which are suboptimal for capturing the complex, non-rigid motion patterns of camouflaged targets. To address this, we extend deformable convolution [10,11] into a TrajectoryAware Alignment (TAA) module, as illustrated in Fig. 5. The primary goal is to resolve the spatial misalignment between the static appearance features Fs and the instantaneous motion features Ft . Trajectory Prediction. We utilize the motion-enhanced feature Ft , which explicitly encodes motion boundaries as a guidance proxy. A lightweight 3D convolution layer ϕ is applied to Ft to simultaneously regress the trajectory offsets ∆P and a modulation mask M:
w(pn )x(p0 + pn )
pn
X
4.3.2 Deformable Grid Sampling
}
spatio-temporal term
where p0 is the center position, R is the local neighborhood, and θ ∈ [0, 1] balances the contribution of gradient versus intensity. This design highlights motion boundaries, producing a high-fidelity motion proxy Ft to accurately guide the subsequent deformable grid sampling.
Temporal Feature
Offset Predictor
Offset
Standard Grid
Modulator
Mask
Deformable Grid
Here, ∆P models the non-rigid deformation field, while the mask M serves as a confidence map to assign lower weights to unreliable sampling points (e.g., background noise). Alignment and Fusion. Guided by the predicted offsets ∆P and the sampling grid G, we perform bilinear sampling S to warp the appearance features Fs to align with the target’s current position. The aligned features are then modulated and fused back into the motion stream via a residual connection: Fout = Ft + Linear S Fs , G + ∆P ⊙ M ,
(7)
where S(·) denotes the bilinear sampling operator. By employing the TAA module, the stable appearance attributes are effectively synchronized with the motion features, yielding a rectified representation that is robust to complex dynamic scenarios.
Spatial Feature
5 Experiments Grid Sample
Out
5.1 Experimental Setup 5.1.1 Datasets
Fig. 5: Detailed architecture of the Trajectory-Aware Alignment (TAA) module. The motion-enhanced temporal feature Ft is utilized to predict trajectory offsets ∆P and a modulation mask M. The offsets dynamically deform the standard sampling grid to accurately sample the static spatial feature Fs . The aligned spatial representations are then modulated by M and add back into the temporal stream.
We comprehensively evaluate our proposed framework on three pixel-level annotated VCOD datasets: CAD [12], MoCA-Mask [13], and our YUV20K. The Camouflaged Animal Dataset (CAD) comprises 9 brief video sequences (839 frames) and serves exclusively as a test set. Built upon the original MoCA dataset [5], the MoCA-Mask [13] benchmark provides precise pixel-level annotations. Following the standard split in [13], it is partitioned into a training set of 71
YUV20K
sequences (19,313 frames) and a test set of 16 sequences (3,626 frames). Furthermore, our proposed YUV20K contains 91 video clips (24,295 frames) covering 47 animal species, which is carefully divided into a training set of X sequences (Y frames) and a test set of Z sequences (W frames). To rigorously assess both in-domain performance and cross-domain generalization, we establish two distinct training paradigms. First, when optimized on the MoCA-Mask training set, the model is evaluated indomain on the MoCA-Mask test set, and cross-domain on CAD and the entirety of the YUV20K dataset (combining its training and testing splits). Second, conversely, when trained on the YUV20K training set, the model is evaluated in-domain on the YUV20K test set, and cross-domain on CAD and the entirety of the MoCAMask dataset (all training and testing splits).
5.1.2 Evaluation Metrics To quantitatively evaluate the segmentation performance, we adopt seven widely used metrics implemented via the PySODMetrics library [39]. These include Structuremeasure (Sm ) [40], maximum F-measure (Fm ) [41], weighted F-measure (Fβw ) [42], Enhanced-alignment measure (Em ) [43], and Mean Absolute Error (M ) [44]. Additionally, we report the mean Dice coefficient (mDice) and mean Intersection over Union (mIoU) to provide a comprehensive assessment of the region-level segmentation accuracy.
5.1.3 Implementation Details Our proposed framework is implemented in PyTorch. The spatial encoder is initialized with the parameters of PVTv2 [33] pretrained on ImageNet. The network is optimized using the AdamW optimizer (β1 = 0.9, β2 = 0.999, weight decay=1 × 10−6 ) with an initial learning rate of 1×10−4 , which is regulated by a cosine annealing learning rate scheduler. To ensure a fair comparison and facilitate our crossdomain evaluation, we follow the widely adopted twostage training protocol [13]. Specifically, the model is first pretrained on the static image camouflage dataset (COD10K-TR) to acquire fundamental appearance-based camouflage representations. Subsequently, to model complex spatiotemporal dynamics, we conduct two independent fine-tuning branches: one optimized exclusively on the MoCA-Mask training set, and the other solely on our YUV20K training set. Both models are fine-tuned for 10 epochs.
9
5.2 Comparison with State-of-the-Art Methods To rigorously evaluate the effectiveness of our proposed framework, we conduct a comprehensive comparison with 7 recent state-of-the-art (SOTA) methods. These encompass both prominent image-based models (e.g., ZoomNet [20]), and recent strong video-based baselines (e.g.,ZoomNext [25], EMIP [23]). To ensure absolute fairness, all comparison methods are evaluated using their official codebases and provided weights, or retrained strictly under our identical experimental protocols. Following our dual-source training paradigm, the quantitative results are reported in two distinct settings: First, Tab. 2 summarizes the performance of models optimized on the MoCA-Mask training set. In this setting, alongside Video COD models, we purposefully include top-tier Image COD methods as static baselines. It is worth noting that these image-based models are evaluated directly using their officially released static weights without any video-specific fine-tuning, serving as a reference for pure appearance-based detection capabilities. Second, Tab. 3 presents the results of models trained on our proposed YUV20K dataset. Since YUV20K is specifically designed to benchmark highly complex scenarios, this table focuses exclusively on comparing dedicated Video COD architectures. This allows for a rigorous assessment of each model’s true spatiotemporal learning capacity when confronted with extreme nonrigid deformations and large displacements. Overall, our method consistently achieves superior performance across both training paradigms. Finally, qualitative visual comparisons against competitive baselines are presented in Fig. 6 and Fig. 7, further demonstrating our model’s robustness in effectively suppressing motion feature instability and accurately aligning moving camouflaged targets. 5.2.1 Quantitative Evaluation Tab. 2 and Tab. 3 report the quantitative performance of all competing methods across three benchmarks: CAD, MoCA-Mask, and our proposed YUV20K. As observed, our framework consistently sets new state-of-the-art records across both training paradigms, though the performance dynamics vary to reflect the intrinsic complexity of the source domains. Specifically, when optimized on the MoCA-Mask dataset (Tab. 2), our method significantly surpasses recent strong competitors, such as ZoomNext [25] and EMIP [23], by a notable margin across most metrics. Conversely, training on our proposed YUV20K benchmark (Tab. 3) presents a formidable challenge due to its highly complex scenarios. This unprecedented difficulty
10
inherently narrows the performance margins among toptier methods. Furthermore, it is noteworthy that most models achieve relatively high absolute metric scores on the YUV20K test set. We primarily attribute this favorable performance to the high-definition (1080p) nature of our dataset. Unlike earlier benchmarks plagued by low-resolution compression artifacts, the 1080p sequences in YUV20K inherently preserve fine-grained texture cues and delicate boundary details. This high spatial fidelity equips all models with richer foundational signals, thereby raising the overall performance baseline. Beyond raising the evaluation baseline, this high spatial fidelity also makes YUV20K an exceptionally superior training data source. A compelling testament to this is the remarkable cross-domain performance of EMIP on MoCA-Mask (e.g., achieving an Sα of 0.879) despite being trained exclusively on YUV20K. The rich, noise-free spatiotemporal details in our 1080p sequences empower optical-flow-based methods to extract highly accurate motion boundaries and learn explicit motion priors to their maximum potential. This demonstrates that YUV20K serves as a powerful foundational resource that significantly enhances the representation capabilities of existing VCOD architectures. However, while our high-quality data maximizes the potential of these baselines, it simultaneously exposes their inherent architectural limitations. Even with strong priors acquired from YUV20K, EMIP exhibits a fundamental trade-off. By architecturally enforcing a strict reliance on explicit motion, EMIP perfectly aligns with the distinct, large-scale movements in MoCA-Mask but suffers significant degradation in subtle or static-like camouflage scenarios, as evidenced by its sharp performance drop on the cross-domain CAD dataset. Conversely, our framework avoids these aggressive motion shortcuts. Our explicit trajectory-aware alignment (TAA) and motion feature stabilization (MFS) mechanisms prioritize global semantic stability. Consequently, we trade negligible in-domain structural gains for significantly superior cross-domain robustness and exceptionally clean background suppression, achieving a leading Mean Absolute Error (M of 0.015) on YUV20K. This confirms that our model is optimized for balanced, real-world generalization rather than overfitting to specific motion distributions. 5.2.2 Qualitative Evaluation Visual comparison of with recent SOTA methods are shown in Fig. 6 and Fig. 7, which demonstrate the robustness of our model across various complex scenarios. Specifically, Fig. 6 presents a comprehensive visual comparison across the three benchmarks, with
Yiyu Liu et al
each row highlighting a distinct and demanding challenge in video camouflaged object detection. The top three rows are sampled from our complex YUV20K dataset: the first row (frog) illustrates sudden and rapid movements; the second row (birds in snow ) presents the extreme difficulty of identifying multiple, tiny targets; and the third row (flounder ) showcases severe non-rigid deformable motion. The subsequent two rows are selected from the MoCA-Mask benchmark: the fourth row (flower spider ) features subtle partial body motion where only specific part of spider are active; and the fifth row (hedgehog) depicts extreme camouflage where the foreground heavily blends into the background. Finally, the last row showcases an example from the CAD dataset (scorpion), introducing targets with highly complex topological shapes.
All methods suffer from severe motion-induced feature instability and temporal feature mismatch which often resulting in missing parts, blurred boundaries, or false positives under these demanding conditions. Our approach consistently maintains sharp contours and object integrity. This superior performance demonstrates the effectiveness of our proposed MFS and TAA modules in stabilizing motion features and achieving precise spatiotemporal alignment.
Furthermore, Fig. 7 illustrates the temporal consistency and tracking stability of our approach across sampled continuous video sequences. To comprehensively validate our model, we visualize five distinct scenarios. The first sequence (jumping frog), sampled from the MoCA benchmark, represents typical camouflage and motion. The subsequent four sequences are drawn from our proposed YUV20K dataset, carefully selected to highlight extreme environmental challenges: an underwater worm (complex white background), a nightjar (intricate texture blending), birds in snow (multiple object and color assimilation), and a safari scene in the dark (low-light degradation). Note that while each example displays five continuous frames, the temporal strides vary adaptively to capture the specific motion speeds of different targets. As visually evident, when handling continuous dynamic changes, baseline methods frequently suffer from severe temporal inconsistency, exhibiting target flickering, boundary collapse. In stark contrast, our framework maintains highly accurate, crisp, and temporally coherent segmentation masks throughout the entire temporal span, powerfully demonstrating the effectiveness of our spatiotemporal alignment and feature stabilization mechanisms.
YUV20K
11
Table 2: Quantitative comparison of models trained on MoCA-Mask. All models in this table are trained exclusively on the MoCA-Mask training set and evaluated on the test sets of CAD (cross-domain), MoCA-Mask (in-domain), and our proposed YUV20K (cross-domain). Bold and red indicate the best performance. ↑ denotes higher is better, and ↓ denotes lower is better. The symbol “-” indicates the results are not available. CAD [12]
Method
MoCA-Mask [16]
YUV20K (Ours)
Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ Image-based Methods SINet [16] ZoomNet [20]
.621 .405 .380 .045 .720 .652 .430 .410 .040 .750
.310 .340
.401 .440
.580 .250 .230 .040 .650 .600 .280 .250 .035 .670
.210 .230
.280 .310
.775 .645 .599 .039 .791
.535
.615
SLT-Net-ST [13] SLT-Net-LT [13] IMEX [45] EMIP [23] ZoomNext [25]
.719 .720 .684 .710 .759
.832 .849 .813 .835 .861
.410 .424 .370 .415 .485
.509 .521 .469 .528 .565
.631 .621 .661 .669 .699
.701 .722 .778 .789 .696
.258 .259 .409 .326 .367
.344 .346 .319 .424 .434
.842 .839 .824 .823
.888 .898 .890 .851
.652 .664 .621 .628
.738 .748 .715 .698
Ours
.771 .613 .584 .019 .871
.507
.589
.700 .456 .432 .008 .733
.377
.449
.858 .812 .780 .035 .920
.697
.772
Video-based Methods .517 .530 .534 .613
.479 .498 .452 .504 .562
.032 .030 .033 .029 .020
.311 .311 .400 .445
.288 .293 .371 .374 .421
.029 .029 .020 .017 .008
.740 .750 .725 .720
.712 .724 .690 .691
.031 .029 .033 .036
Table 3: Quantitative comparison of models trained on our proposed YUV20K. All models in this table are trained exclusively on the YUV20K training set and evaluated on the test sets of CAD (cross-domain), MoCA-Mask (cross-domain), and YUV20K (in-domain). This demonstrates the strong generalization ability acquired from our dataset. Bold and red indicate the best performance. ↑ denotes higher is better, and ↓ denotes lower is better. CAD [12]
Method
MoCA-Mask [16]
YUV20K (Ours)
Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ Sα ↑ Fβ ↑ Fβω ↑ M↓ Em ↑ mIoU↑ mDice↑ SLT-Net-ST [13] .720 .492 .445 .039 .806 .725 .533 .483 .030 .839 EMIP [23] ZoomNext [25] .775 .569 .565 .018 .815
.398 .405 .503
.499 .518 .577
.798 .642 .597 .029 .848 .879 .789 .775 .012 .925 .827 .719 .685 .018 .853
.559 .722 .624
.659 .789 .707
.848 .684 .648 .025 .908 .857 .727 .702 .020 .924 .856 .704 .685 .017 .895
.607 .643 .635
.697 ..717 .690
Ours
.505
.578
.831 .726 .693 .017 .857
.632
.714
.858 .706 .688 .015 .897
.639
.692
.775 .582 .569 .017 .830
Table 4: Attribute-based performance on the proposed YUV20K dataset. We evaluate the video-based methods across six challenging motion scenarios: Ldm (Large displacement motion), CM (Camera Motion), Occ (Occlusion), M-Obj (Multiple Objects), Hunt (Hunting), and T-Obj (Tiny Object). We report the structuremeasure (Sα ↑), weighted F-measure (Fβω ↑), and mean Intersection-over-Union (mIoU ↑) to demonstrate the robustness of different methods. Bold and red indicate the best performance. Ldm
Method
CM
Occ
M-Obj
Hunt
T-Obj
Sα ↑
Fβω ↑
mIoU ↑
Sα ↑
Fβω ↑
mIoU ↑
Sα ↑
Fβω ↑
mIoU ↑
Sα ↑
Fβω ↑
mIoU ↑
Sα ↑
Fβω ↑
mIoU ↑
Sα ↑
Fβω ↑
mIoU ↑
SLT-Net-ST [13] SLT-Net-LT [13] ZoomNet [20] ZoomNext [25] EMIP [23]
.855 .851 .799 .841 .832
.745 .754 .644 .738 .709
.689 .698 .583 .675 .647
.885 .880 .826 .876 .859
.802 .809 .695 .793 .758
.747 .754 .636 .734 .695
.877 .875 .791 .901 .863
.784 .801 .648 .842 .768
.716 .731 .573 .768 .689
.733 .728 .678 .773 .721
.542 .546 .428 .611 .520
.474 .477 .358 .539 .447
.782 .779 .726 .824 .736
.624 .638 .525 .699 .552
.547 .559 .440 .624 .468
.639 .638 .637 .691 .629
.364 .364 .281 .426 .321
.307 .306 .235 .366 .271
Ours
.856
.765
.700
.886
.816
.755
.901
.846
.769
.780
.631
.553
.825
.706
.627
.695
.446
.372
5.3 Ablation Study In this section, we perform a comprehensive ablation analysis to investigate the contribution of each proposed component. To rigorously evaluate the cross-domain generalization capability of these modules, all experiments in this section are conducted using models trained exclusively on the MoCA-Mask training set. The quantitative results on CAD, MoCA-Mask, and our YUV20K benchmarks are summarized in Tab. 5. Synergy of MFS and TAA. As shown in Tab. 5, introducing either the Motion Feature Stabilization (MFS) or the Trajectory-Aware Alignment (TAA) module individually brings noticeable improvements to the in-domain
MoCA-mask dataset compared to the baseline. However, when deployed to the complex and unseen YUV20K dataset, deploying a single module occasionally leads to ω performance fluctuations (e.g., a drop in Fm ). This suggests that while individual modules excel in MoCA bias motion, single-axis optimization is insufficient to bridge the vast domain gap toward wild, erratic deformations. Crucially, when both modules are integrated into our Full Model (last row), the cross-domain performance on YUV20K experiences a massive surge (Sm jumps ω to .858 and Fm to .780). This compelling evidence proves that MFS and TAA are highly complementary. The appearance stabilization provided by MFS acts as a prerequisite for TAA to accurately predict trajectory
12
Yiyu Liu et al
Fig. 6: Qualitative visual comparison of our proposed approach against state-of-the-art methods. Representative sequences are selected from the YUV20K MoCA, and CAD datasets (top to bottom). From left to right: the original input frames, ground truth (GT) masks, predictions of our model, ZoomNeXT, EMIP, and SLT-Net (Short-term and Long-term variants). Table 5: Ablation study on three datasets. ↓ indicates lower is better, others are higher is better. The best results are highlighted in bold.
CAD
Method
MoCA-mask
Sm ↑
Em ↑ M ↓ Sm ↑
Baseline + MFS + TAA w/ vanilla 3D + TAA
.759 .757 .762 .763
.562 .543 .573 .575
.861 .020 .699 .421 .852 .021 .715 .438 .877 .019 .707 .443 .866 .019 .715 .459
.696 .722 .715 .728
.851 .863 .860 .867
.036 .032 .029 .030
Ours
.771 .584
.871
.733 .008 .858 .780 .920
.035
.019
offsets, while TAA perfectly aligns the stabilized features. Together, they form an indispensable synergy that successfully conquers severe spatiotemporal complex scenarios. Impact of 3D-CDC in TAA. To justify the design choice of using 3D-CDC within the TAA module, we conduct an internal ablation by substituting it with a standard vanilla 3D convolution (denoted as ‘+ TAA w/o 3D-CDC’ in Tab. 5). Comparing the 3rd and 4th rows, we observe an intriguing phenomenon. On the source domain (MoCAMask), the variant without 3D-CDC performs competitively. However, when generalizing to the YUV20K dataset, removing the 3D-CDC operator leads to a no-
.700
ω Fm ↑
YUV20K (Ours)
ω Fm ↑
.432
ω Em ↑ M ↓ Sm ↑ Fm ↑ Em ↑ M ↓
.008 .008 .008 .007
.823 .807 .808 .805
.691 .467 .664 .657
ticeable performance drop. This degradation suggests that standard intensity-based 3D convolutions struggle to distinguish subtle, non-rigid motion patterns from static backgrounds in unseen environments. In contrast, by explicitly leveraging gradient-level differential cues, our full TAA module (equipped with 3D-CDC) captures fine-grained motion dynamics and guides the trajectory alignment significantly more effectively under extreme conditions.
5.4 Hyperparameter Analysis To further investigate the properties of our proposed framework, we conduct detailed analyses on two critical
YUV20K
13
Fig. 7: Qualitative comparison of video camouflaged object detection. Visual sequences comparing our proposed method against state-of-the-art models (e.g., EMIP, SLT-LT) on challenging scenarios from the YUV20K dataset (Part 1). By actively avoiding background interference through our deformable strategy, our method consistently maintains temporal coherence, predicts sharper boundaries, and effectively suppresses background noise compared to other approaches.
14
Yiyu Liu et al
Fig. 7: Qualitative comparison of video camouflaged object detection (Continued). Additional sequence results under extreme challenges such as heavy occlusion and tiny objects. Our model demonstrates superior robustness in faithfully tracking non-rigid targets across frames. hyperparameters: the trade-off coefficient θ in the 3DCDCST module and the dimension of Semantic Basis Primitives (SBP) in the MFS module. Impact of θ in 3D-CDCST. In our 3D-CDCST module, the parameter θ regulates the contribution ratio between the standard 3D convolution and the central difference convolution. Interestingly, empirical results reveal that the optimal θ value is highly dependent on the intrinsic motion characteristics of the training domain. Specifically, when optimized on the MoCA dataset, which features relatively simple scenarios, setting θ = 0.458 yields the best performance. However, on our highly complex YUV20K dataset with erratic
and severe deformations, the optimal value shifts to θ = 0.158. This discrepancy indicates that different motion distributions require distinct levels of gradient-level differential cues. Motivated by this insight, expanding θ into a dynamically learnable parameter to further enhance scene-adaptive generalization remains a promising direction for our future work. Dimension of Semantic Basis Primitives. The dimension of the Semantic Basis Primitives, denoted as K, dictates the representational capacity for appearance stabilization in the MFS module. A smaller dimension fails to provide sufficient semantic anchors to comprehensively encompass the complex diversity of camouflaged
YUV20K
objects and highly textured backgrounds. Conversely, an excessively large dimension not only introduces burdensome computational overhead but also provokes feature redundancy and over-smoothing, which dilutes critical discriminative cues. Extensive experiments demonstrate that setting the dimension to K = 384 strikes the perfect balance. This configuration grants the network ample capacity to robustly anchor features against severe motion dynamics without suffering from over-parameterization.
6 Conclusion In this paper, we addressed the critical challenges in Video Camouflaged Object Detection (VCOD) stemming from data scarcity and complex motion scenarios. We introduced YUV20K, a comprehensive and diverse benchmark that fills the gap in representing wild biological behaviors, offering a higher proportion of complex scenarios such as occlusion and hunting behavior compared to prior datasets. To effectively process these challenging inputs, we proposed a unified spatiotemporal framework designed to rectify motion-induced feature instability and temporal misalignment. Our approach integrates the Motion Feature Stabilization (MFS) module, which leverages Semantic Basis Primitives to anchor semantic features, and the Trajectory-Aware Alignment (TAA) module, which utilizes trajectory-guided deformable sampling for precise feature alignment and aggregation. Experimental results confirm that our framework sets a new baseline for the field. Notably, it consistently achieves state-of-the-art performance not only under standard in-domain evaluations but also in rigorous cross-domain settings, exhibiting exceptional generalization capabilities on our challenging YUV20K dataset. We believe this work significantly advances VCOD research from both data and algorithmic perspectives, paving the way for the development of more resilient spatiotemporal reasoning systems in wild, complex scenarios.
Abbreviations SBP, semantic basis primitives; MFS, motion feature stabilization; 3D-CDCST, 3 dimension center-difference convolution spatial temporal; TAA, trajectory aware alignment.
Author Contributions Y.L. and Y.S. contributed equally to this work. Y.L. conceived the original idea and implemented the main
15
methodology. Y.S. performed partial experiments and contributed to drafting the manuscript. C.H. provided guidance and critically revised the manuscript. Z.Y. supervised the project, organized the research discussions, and finalized the writing. All authors read and approved the final manuscript.
Funding This work was supported by the National Natural Science Foundation of China (Grant No. 62576076). The computational resources are supported by SongShan Lake HPC Center (SSL-HPC) in Great Bay University.
Data Availability The datasets supporting this study’s findings are publicly accessible and include two camouflaged video datasets: CAD [12] (Google Drive), MoCA-Mask [13] (Google Drive). Our YUV20K will be available at https:// github.com/K1NSA/YUV20K.
Declarations Competing interests The authors declare no competing interests.
References 1. Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling Shao. Concealed object detection. IEEE transactions on pattern analysis and machine intelligence, 44(10):6024– 6042, 2021. 2. Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. Pranet: Parallel reverse attention network for polyp segmentation. In International conference on medical image computing and computer-assisted intervention, pages 263–273. Springer, 2020. 3. Deng-Ping Fan, Tao Zhou, Ge-Peng Ji, Yi Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. Infnet: Automatic covid-19 lung infection segmentation from ct images. IEEE transactions on medical imaging, 39(8):2626–2637, 2020. 4. Dan Jeric Arcega Rustia, Chien Erh Lin, Jui-Yung Chung, Yi-Ji Zhuang, Ju-Chun Hsu, and Ta-Te Lin. Application of an image and environmental sensor network for automated greenhouse insect pest monitoring. Journal of Asia-Pacific Entomology, 23(1):17–28, 2020. 5. Hala Lamdouar, Charig Yang, Weidi Xie, and Andrew Zisserman. Betrayed by motion: Camouflaged object discovery via motion segmentation. In Proceedings of the Asian conference on computer vision, 2020.
16 6. Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, and Weidi Xie. Self-supervised video object segmentation by motion grouping. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7177–7188, 2021. 7. Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander SorkineHornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 724–732, 2016. 8. Fan Yang, Qiang Zhai, Xin Li, Rui Huang, Ao Luo, Hong Cheng, and Deng-Ping Fan. Uncertainty-guided transformer reasoning for camouflaged object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4146–4155, 2021. 9. Xintao Wang, Kelvin CK Chan, Ke Yu, Chao Dong, and Chen Change Loy. Edvr: Video restoration with enhanced deformable convolutional networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 0–0, 2019. 10. Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 11. Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al. Internimage: Exploring large-scale vision foundation models with deformable convolutions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14408–14419, 2023. 12. Pia Bideau and Erik Learned-Miller. It’s moving! a probabilistic model for causal motion segmentation in moving camera videos. In European Conference on Computer Vision, pages 433–449. Springer, 2016. 13. Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong, Mehrtash Harandi, Tom Drummond, and Zongyuan Ge. Implicit motion handling for video camouflaged object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13864– 13873, 2022. 14. Fengyang Xiao, Sujie Hu, Yuqi Shen, Chengyu Fang, Jinfa Huang, Chunming He, Longxiang Tang, Ziyun Yang, and Xiu Li. A survey of camouflaged object detection and beyond. arXiv preprint arXiv:2408.14562, 2024. 15. Joanna R Hall, Innes C Cuthill, Roland Baddeley, Adam J Shohet, and Nicholas E Scott-Samuel. Camouflage, detection and identification of moving targets. Proceedings of the Royal Society B: Biological Sciences, 280(1758):20130064, 2013. 16. Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng, Jianbing Shen, and Ling Shao. Camouflaged object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2777–2787, 2020. 17. Runmin Cong, Mengyao Sun, Sanyi Zhang, Xiaofei Zhou, Wei Zhang, and Yao Zhao. Frequency perception network for camouflaged object detection. In Proceedings of the 31st ACM international conference on multimedia, pages 1179–1189, 2023. 18. Yanguang Sun, Chunyan Xu, Jian Yang, Hanyu Xuan, and Lei Luo. Frequency-spatial entanglement learning for camouflaged object detection. In European Conference on Computer Vision, pages 343–360. Springer, 2024. 19. Qiang Zhai, Xin Li, Fan Yang, Chenglizhao Chen, Hong Cheng, and Deng-Ping Fan. Mutual graph learning for camouflaged object detection. In Proceedings of the
Yiyu Liu et al IEEE/CVF conference on computer vision and pattern recognition, pages 12997–13007, 2021. 20. Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang, and Huchuan Lu. Zoom in and out: A mixed-scale triplet network for camouflaged object detection. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 2160–2170, 2022. 21. Zhou Huang, Hang Dai, Tian-Zhu Xiang, Shuo Wang, Huai-Xin Chen, Jie Qin, and Huan Xiong. Feature shrinkage pyramid for camouflaged object detection with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5557–5566, 2023. 22. Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In European conference on computer vision, pages 402–419. Springer, 2020. 23. Xin Zhang, Tao Xiao, Ge-Peng Ji, Xuan Wu, Keren Fu, and Qijun Zhao. Explicit motion handling and interactive prompting for video camouflaged object detection. IEEE Transactions on Image Processing, 2025. 24. Wenjun Hui, Zhenfeng Zhu, Shuai Zheng, and Yao Zhao. Endow sam with keen eyes: Temporal-spatial prompt learning for video camouflaged object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19058–19067, 2024. 25. Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang, and Huchuan Lu. Zoomnext: A unified collaborative pyramid network for camouflaged object detection. IEEE transactions on pattern analysis and machine intelligence, 46(12):9205–9220, 2024. 26. Muhammad Nawfal Meeran, Bhanu Pratyush Mantha, et al. Sam-pm: Enhancing video camouflaged object detection using spatio-temporal attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1857–1866, 2024. 27. Trung-Nghia Le, Tam V Nguyen, Zhongliang Nie, MinhTriet Tran, and Akihiro Sugimoto. Anabranch network for camouflaged object segmentation. Computer vision and image understanding, 184:45–56, 2019. 28. Przemyslaw Skurowski, Hassan Abdulameer, Jakub Blaszczyk, Tomasz Depta, Adam Kornacki, and Przemyslaw Koziel. Animal camouflage analysis: Chameleon database. Unpublished manuscript, 2(6):7, 2018. 29. Yunqiu Lyu, Jing Zhang, Yuchao Dai, Aixuan Li, Bowen Liu, Nick Barnes, and Deng-Ping Fan. Simultaneously localize, segment and rank the camouflaged objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 30. Wei Wang. Advanced auto labeling solution with added features. https://github.com/CVHub520/X-AnyLabeling, 2023. 31. Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 32. Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. Grounded sam: Assembling open-world models for diverse visual tasks. arXiv preprint arXiv:2401.14159, 2024. 33. Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvt v2: Improved baselines with pyramid vision transformer. Computational visual media, 8(3):415–424, 2022. 34. Mateo Espinosa Zarlenga, Pietro Barbiero, Gabriele Ciravegna, Giuseppe Marra, Francesco Giannini,
YUV20K Michelangelo Diligenti, Zohreh Shams, Frederic Precioso, Stefano Melacci, Adrian Weller, et al. Concept embedding models: Beyond the accuracy-explainability trade-off. Advances in neural information processing systems, 35:21400–21413, 2022. 35. Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. Concept bottleneck models. In International conference on machine learning, pages 5338–5348. PMLR, 2020. 36. Shuo Ye, Lixin Chen, Qiaoqi Li, Jiayu Zhang, Chaomeng Chen, and Shutao Xia. IKA2: Internal Knowledge Adaptive Activation for Robust Recognition in Complex Scenarios. Machine Intelligence Research, 23(2):429–443, April 2026. 37. Zitong Yu, Chenxu Zhao, Zezheng Wang, Yunxiao Qin, Zhuo Su, Xiaobai Li, Feng Zhou, and Guoying Zhao. Searching central difference convolutional networks for face anti-spoofing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5295–5305, 2020. 38. Zitong Yu, Yunxiao Qin, Hengshuang Zhao, Xiaobai Li, and Guoying Zhao. Dual-cross central difference network for face anti-spoofing. arXiv preprint arXiv:2105.01290, 2021. 39. Youwei Pang. Pysodmetrics: A simple and efficient implementation of sod metrcis, 2020. 40. Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji. Structure-measure: A new way to evaluate foreground maps. In ICCV, pages 4548–4557, 2017. 41. Radhakrishna Achanta, Sheila Hemami, Francisco Estrada, and Sabine Susstrunk. Frequency-tuned salient region detection. In 2009 IEEE conference on computer vision and pattern recognition, pages 1597–1604. IEEE, 2009. 42. Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal. How to evaluate foreground maps? In CVPR, pages 248–255, 2014. 43. Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, MingMing Cheng, and Ali Borji. Enhanced-alignment measure for binary foreground map evaluation. In IJCAI, pages 698–704, 2018. 44. Federico Perazzi, Philipp Krähenbühl, Yael Pritch, and Alexander Hornung. Saliency filters: Contrast based filtering for salient region detection. In CVPR, pages 733–740, 2012. 45. Wenjun Hui, Zhenfeng Zhu, Guanghua Gu, Meiqin Liu, and Yao Zhao. Implicit-explicit motion learning for video camouflaged object detection. IEEE Transactions on Multimedia, 26:7188–7196, 2024.
17