1
JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering
arXiv:2606.31745v1 [cs.CV] 30 Jun 2026
Ziyuan Liu, Ruifei Zhu, Ouqiao Ma, and Yuantao Gu
Abstract—Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but neither what nor why. To bridge this semantic gap, we introduce JL1-CC&QA, a multi-task benchmark that extends the JL1-CD dataset with two complementary annotation layers: change captioning (CC) and change question answering (QA). Built upon 5,000 bi-temporal image pairs acquired by the Jilin-1 satellite at 0.5–0.75 m ground sample distance, the benchmark comprises: (i) JL1-CC, providing 17,021 quality-verified captions that describe diverse land-cover transformations; and (ii) JL1-QA, offering 20,060 question–answer pairs across eight question types, enabling fine-grained, interactive interrogation of surface changes. All annotations are produced via a threestage pipeline consisting of multi-modal large language model (LLM) generation, vision-grounded LLM judging, and human expert verification. We hope that JL1-CC&QA, as a benchmark unifying binary change masks, change captions, and changeoriented QA over the same image set, will serve as a valuable resource for the community to advance multi-task change understanding in remote sensing. The dataset is available at https://github.com/circleLZY/JL1-CD. Index Terms—Benchmark, change captioning, change question answering, change detection, remote sensing.
I. I NTRODUCTION
R
EMOTE sensing change detection (CD) aims to identify surface changes from multi-temporal imagery acquired over the same geographic area, serving as an important technology for urban planning, environmental monitoring, disaster response, and resource management [1], [2]. Over the past decade, the rapid growth of both Earth observation data and deep learning methods has propelled CD into one of the most active research frontiers in remote sensing. The predominant formulation of CD is binary change detection (BCD), which classifies each pixel as either changed or unchanged. A large number of BCD benchmarks have been established, spanning diverse geographic contexts, spatial resolutions, and change categories (see Table I for a comprehensive summary) [3]–[19]. While the vast majority of these benchmarks adopt the bi-temporal setting and have driven extensive algorithmic development [10], [11], [20]– [24], the community has also explored single-temporal CD that exploits pretrained semantic representations to infer changes from unpaired imagery [25]–[27], as well as multi-temporal Ziyuan Liu and Yuantao Gu are with the Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China (e-mail: [email protected]; [email protected]). Ruifei Zhu is with Chang Guang Satellite Technology Co., Ltd. (CGSTL) Changchun 130102, China (e-mail: [email protected]). Ouqiao Ma is with the College of Communications Engineering, Army Engineering University of PLA, Nanjing 210007, China (e-mail: [email protected]). (Corresponding author: Yuantao Gu.)
monitoring that tracks continuous change trajectories from dense satellite time series [28]–[31]. Despite these differences, all of the above approaches produce pixel-level masks that encode where change occurred but remain silent on what changed or why. In parallel, the range of input modalities has expanded considerably. While optical RGB imagery remains the dominant data source for CD research [7], [10], [13], [15], [19], multispectral and hyperspectral sensors offer richer spectral discrimination for fine-grained change analysis [5], [8], [9], [39], [54]–[56], and cross-modal fusion of optical and SAR data enables all-weather, day-and-night monitoring [14], [17], [45], [57]–[60]. Although these modalities provide richer visual information, the output of BCD remains a pixel-level binary mask devoid of semantic content. Semantic change detection (SCD) and building damage assessment (BDA) partially bridge this gap by assigning categorical labels to changed pixels. SCD introduces perpixel land-cover annotations at both time phases, enabling identification of change types (e.g., farmland→buildings) [30], [34], [37], [41]–[43], [47], [48], [50], [51], [61]. BDA further introduces ordinal damage scales (e.g., minor/major/destroyed) for disaster response [32], [35], [38], [46], [49], [52], [53]. Nevertheless, the semantic content in these datasets is still encoded as discrete numerical class indices drawn from a closed taxonomy. This gap calls for a shift toward naturallanguage-grounded change understanding. The advent of vision–language models (VLMs) has begun to close this gap. In the natural image domain, change captioning benchmarks first demonstrated the feasibility of describing visual differences in natural language [62], [63]. The remote sensing community quickly adopted this paradigm, producing a growing family of change captioning (CC) datasets that range from low-resolution multitemporal pairs to very-highresolution urban and disaster scenes [64]–[70]. In parallel, change detection visual question answering (CDVQA) has emerged as a complementary task, enabling users to query specific aspects of surface changes through natural-language questions [71], [72], while instruction-tuning datasets have further extended the interaction to multi-turn conversational analysis [73]–[75]. Despite this rapid progress, rare benchmarks jointly provide binary change masks, change captions, and change-oriented question–answer pairs on the same set of image pairs, which limits multi-task learning and cross-task evaluation. The evolution of datasets has been accompanied by corresponding advances in model architectures (see Fig. 1). Methods built on CNN backbones with change decoders [10],
2
TABLE I C OMPREHENSIVE S UMMARY OF R EPRESENTATIVE R EMOTE S ENSING C HANGE D ETECTION DATASETS Dataset
Year Task Modal. Source
GSD
Size (px)
Scale
Temp. Change Elements
SZTAKI [3] AICD [4] ABCD [32] MtS-WH [33] OSCD [5] CDD [6] WHU-CD [7] HRSCD [34]
2009 2011 2017 2017 2018 2018 2018 2019
BCD BCD BDA BCD BCD BCD BCD SCD
Opt. Opt. Opt. Opt. MS Opt. Opt. Opt.
1.5 m — 0.4 m 1m 10 m VHR 0.2 m 0.5 m
952×640 800×600 128×128 7200×6000 600×600 256×256 15354×32507 10000×10000
13 1000 8506 1 24 16000 1 291
Bi Bi Bi Bi Bi Bi Bi Bi
River [8] Hyper. CDD [9] xBD [35] LEVIR-CD [10] LEVIR-CD+ [10] DSIFN [11] CD Data GZ [12]
2019 2019 2019 2020 2020 2020 2020
BCD BCD BDA BCD BCD BCD BCD
HS HS Opt. Opt. Opt. Opt. Opt.
Optical-SAR-CD [36] SYSU-CD [15]
2021 BCD Cross 2021 BCD Opt.
S2Looking [13] HTCD [14] SECOND [37]
2021 BCD Opt. 2021 BCD Opt. 2021 SCD Opt.
ChangeOS [38] ACDA [39] SpaceNet 7 [29] WHU-OPT-SAR [40]
2021 2021 2021 2022
CLCD [16] Hi-UCD [41]
2022 BCD Opt. 2022 SCD Opt.
BDA BCD BCD BCD
Opt. HS Opt. Cross
Landsat-SCD [42] 2022 SCD MS MSBC [17] 2022 BCD Multi DynamicEarthNet [28] 2022 SCD MS BANDON [43] 3DCD [44] CAU-Flood [45] WUSU [30]
2023 2023 2023 2023
SCD BCD BCD SCD
Opt. Opt. Cross Opt.
RescueNet [46] EGY-BCD [18] CropSCD [47]
2023 BDA Opt. 2023 BCD Opt. 2024 SCD Opt.
ClearSCD [48]
2024 SCD Opt.
QQB [49] TSCD [31] JL1-CD [19]
2024 BDA Cross 2024 BCD Opt. 2025 BCD Opt.
PRO-HRSCD [50]
2025 SCD Opt.
TripleS [51]
2025 SCD Opt.
BRIGHT [52] CDFNet [53]
2025 BDA Cross 2025 BDA Opt.
Aerial Synthetic Aerial IKONOS Sentinel-2 GE Aerial Aerial (IGN)
Buildings, forest, foundations Synthetic (buildings, vehicles) Building damage Parking, water, bldg., farm, industrial Buildings, roads Seasonal, vehicles, bldg., trees Buildings Artificial, agriculture, forest, wetland, water Airborne — 463×241 1 Bi Water Airborne — Varies 3 Bi Crop transition Satellite VHR 1024×1024 22068 Bi Building damage GE 0.5 m 1024×1024 637 Bi Buildings GE 0.5 m 1024×1024 985 Bi Buildings GE VHR 512×512 3940 Bi Roads, bldg., farmland, water GE VHR 256×256 4000+ Bi Water, road, farmland, bare soil, forest, bldg., ships Opt.+SAR 10 m 600×600 3 Bi Mixed land cover Aerial 0.5 m 256×256 20000 Bi Bldg., road, veg., water, marine construction Satellite 0.5–0.8 m 1024×1024 5000 Bi Buildings Sat.+UAV VHR Varies 1293 Bi Buildings, roads, urban features Aerial 0.5–3 m 512×512 4662 Bi Ground, tree, low veg., water, bldg., playground Satellite VHR Varies 2568 Bi No damage / minor / major / destroyed Air./Sat. — Varies 4 Bi Anomalous spectral changes Satellite 4m 1024×1024 101 AOIs Multi Buildings GF-2+S1 1–10 m Varies 2695 Bi Farmland, urban, rural, water, forest, road GF-2 0.5–2 m 512×512 600 Bi Cropland Aerial 0.1 m 1024×1024 40800+ Multi Water, grass, forest, bare soil, bldg., greenhouse, road, bridge Landsat 30 m 256×256 8468 Bi Desert, farmland, bldg., water GF-2+S1+S2 0.8–10 m Varies 4838 Bi Buildings PlanetScope 3 m 1024×1024 75 AOIs Multi Impervious, agriculture, forest, wetland, soil, water, snow Aerial VHR 2048×2048 2283 Bi Buildings Aerial VHR 512×512 472 Bi Buildings, roads, bridges Opt.+SAR 10 m 256×256 2016 Bi Flood inundation GF-2 0.8 m 512×512 10000+ Multi Bldg., road, cropland, forest, grass, river, lake, excavation, bare soil UAV VHR 3000×4000 2396 Single Bldg., water, road, vehicle, tree, pool Satellite 0.25 m 256×256 6091 Bi Buildings GF-2 VHR 256×256 4202 Bi Water, forest, plantation, grassland, impervious, greenhouse, road, bare soil Various 0.5–3 m 512×512 4662 Bi Ground, tree, low veg., water, bldg., playground SAR+Opt. VHR 256×256 3200 Bi Buildings WorldView-2 0.5 m 256×256 5140 Multi Buildings, greenhouses JL-1 0.5–0.75 m 512×512 5000 Bi Bldg., road, PV, hardened, forest, grass, cropland, water Aerial (IGN) 0.5 m 10000×10000 291 Bi Artificial, agriculture, forest, wetland, water Various 0.5–3 m 512×512 4662+ Bi Bare soil, water, bldg., farmland, veg., road SAR+Opt. 0.3–1 m 1024×1024 4538 Bi Buildings Satellite VHR Varies 1500+ Single Multi-level building damage
Abbreviations: Opt. = Optical RGB; MS = Multispectral; HS = Hyperspectral; Cross = Cross-modal; Multi = Multi-modal fusion; GE = Google Earth; GF-2 = Gaofen-2; JL-1 = Jilin-1; S1/S2 = Sentinel-1/2; PV = Photovoltaic; Bi = Bi-temporal; Bldg. = Building(s); Veg. = Vegetation; VHR = Very High Resolution (<1 m). Dataset names are hyperlinked to download pages where available.
3
Fig. 1. Timeline of the development of mainstream deep learning-based CD methods.
[11], [20], [21], [76]–[97] established the dominant Siamese paradigm, while Transformer [22], [23], [54], [98]–[108] introduced global cross-temporal attention, and Mamba-based methods [24], [60], [109]–[115] achieved comparable perception at linear computational cost. Foundation model adaptations have enabled zero-shot and few-shot CD by transferring pretrained CLIP and SAM representations [116]–[124]. Generative approaches based on GANs [57], [125]–[127] and diffusion models [128]–[134] have addressed data scarcity through synthetic bi-temporal image generation. Agent-based systems have integrated LLMs as reasoning engines for interactive change analysis [135]–[138]. A critical observation is that the text modality has been a key enabler at each stage: CLIPbased models use text prompts to guide change semantics, captioning models translate visual change into language, VQA
models support interactive querying, and agents conduct multistep reasoning in natural language. This trajectory underscores the need for language-grounded benchmarks that can support the training and evaluation of these increasingly capable architectures. Motivated by the above analysis, we present JL1-CC&QA, extending the JL1-CD benchmark [19], which consists of 5,000 bi-temporal image pairs from the Jilin-1 satellite with a resolution of 0.5–0.75 m, with two new annotation layers: (i) JL1-CC, providing 17,021 quality-verified change captions spanning both anthropogenic and natural changes; and (ii) JL1-QA, offering 20,060 question–answer pairs across eight question types (existence, description, location, magnitude, temporal comparison, causation, relative comparison, and visual detail). All annotations are produced via a three-stage
TABLE II S PECIFICATIONS OF THE JL1-CD S OURCE DATASET Attribute
Value
Satellite Spatial resolution Image size Spectral bands Total image pairs Training / test split Annotation Temporal coverage Geographic coverage
Jilin-1 (JL1) Constellation 0.5–0.75 m GSD 512 × 512 pixels RGB (3 channels) 5,000 4,000 / 1,000 Pixel-level binary mask 2022–2023 Multiple provinces in China
pipeline: multi-modal LLM generation, vision-grounded LLM judging, and human expert verification. Our main contributions are summarized as follows: 1) We construct a change understanding benchmark that provides binary change masks, natural-language change descriptions, and diverse QA pairs over 5,000 bi-temporal satellite image pairs, facilitating joint training and evaluation across CD, CC, and QA tasks. 2) We design a scalable, reproducible annotation pipeline that couples multi-modal LLM generation with visiongrounded LLM judging and human verification. 3) We provide comprehensive dataset statistics, and release all data and code to support community research toward unified remote sensing change understanding. II. DATASET C ONSTRUCTION This section describes the construction of JL1-CC&QA. We first introduce the source dataset JL1-CD (Section II-A), then detail the change captioning pipeline (Section II-B) and the question answering pipeline (Section II-C). A. JL1-CD JL1-CC&QA is built upon JL1-CD [19], which comprises 5,000 bi-temporal image pairs captured by the Jilin-1 highresolution optical satellite. The imagery was acquired across multiple provinces in China, including Shandong, Ningxia, Anhui, Hebei, and Hunan, between early 2022 and late 2023, and was carefully curated to exclude blur, cloud occlusion, and extreme illumination conditions. Each image pair consists of a pre-event image, a post-event image, and a pixel-level binary change mask annotated by professional interpreters. Key specifications are summarized in Table II. Two properties of JL1-CD make it suitable for change captioning and question answering. First, the dataset is allinclusive in terms of change types: it encompasses both anthropogenic changes (buildings, roads, hardened surfaces, photovoltaic panels, etc.) and natural changes (woodlands, grasslands, croplands, water bodies, etc.). This diversity ensures that the derived CC and QA annotations cover a broad spectrum of real-world surface dynamics. Second, the change area ratio (CAR)—defined as the proportion of changed pixels in each image pair—spans the full range from near-zero to 100%, with a mean of 9.7% and a median of 2.9% (Fig. 2).
Number of Image Pairs
4
Median = 2.9% Mean = 9.7%
3000 2500 2000 1500 1000 500 0
0
20
40
60
Change Area Ratio (%)
80
100
Fig. 2. Distribution of change area ratio (CAR) across 5,000 image pairs in JL1-CD.
This long-tailed distribution mirrors the natural imbalance of real-world change scenarios and poses a meaningful challenge for language-grounded change understanding at varying scales. B. JL1-CC 1) Task Definition: Given a bi-temporal image pair (IA , IB ) and its corresponding binary change mask M , the change captioning task requires generating a set of natural-language sentences {c1 , c2 , . . . , ck } that accurately describe the semantic content of the observed surface changes, including the type, location, and extent of change. 2) Annotation Pipeline: As illustrated in Fig. 3, the JL1-CC annotation pipeline consists of three stages. Stage 1: Multi-modal LLM Generation. For each image pair, we prompt a multi-modal LLM (Kimi-K2.6) with three visual inputs: the pre-event image IA , the post-event image IB , and the binary change mask M , together with spatial metadata including the CAR and a textual description of the primary change region (e.g., “upper-left area of the image”). The model is instructed to generate five diverse captions, each describing the same change from a different perspective: change type, spatial location, visual appearance, scale, or implication. Stage 2: Vision-Grounded LLM Judging. A second LLM call evaluates each caption by examining the original image pair alongside the generated text. Each caption is scored on a 1–10 integer scale across five criteria: (1) accuracy: whether the description matches the visible changes, with heavy penalties for hallucination; (2) specificity: whether concrete landcover terms are used; (3) spatial correctness: whether location references are accurate; (4) naturalness: whether the English is fluent; and (5) informativeness: whether the caption conveys meaningful detail. The top-3 scoring captions are retained, and ties at the cutoff are preserved. Stage 3: Human Expert Verification. A subset of the generated captions is reviewed by domain experts to verify factual accuracy and identify systematic errors, ensuring that the automated pipeline produces reliable annotations at scale. 3) Statistics: Table III summarizes the JL1-CC dataset statistics. The pipeline generates 25,000 candidate captions (5 per pair) and retains 17,021 after quality filtering, yielding a pass rate of 68.1%. The judge score distribution (Fig. 4) shows clear discrimination: 39.8% of captions score 9–10, 42.3%
5
Fig. 3. Overview of the JL1-CC annotation pipeline. Stage 1: a multi-modal LLM generates five candidate captions from the image pair and spatial metadata. Stage 2: a vision-grounded LLM judge scores each caption and retains the top three. Stage 3: human experts verify factual accuracy.
TABLE III JL1-CC DATASET S TATISTICS
Number of Captions
8000 6000
Statistic
Rejected
4000
Accepted
2000 0
1
2
3
4
5
6
Judge Score
7
8
9
10
Fig. 4. Judge score distribution for JL1-CC. Captions scoring below 7 (red) are rejected; those scoring 7–8 (green) and 9–10 (blue) are retained.
Fig. 5. Word cloud of the 17,021 selected captions in JL1-CC.
score 7–8, and 17.9% score below 7 and are rejected. The selected captions have a mean length of 26.2 words, with a vocabulary of 7,458 unique tokens. Fig. 5 presents a word cloud of the selected captions, where the most prominent terms— upper, lower, bare soil, agricultural, building, road, water— reflect both the spatial referencing style and the diversity of land-cover types in JL1-CD.
Image pairs Generated captions Selected captions Avg. selected / pair Avg. caption length (words) Vocabulary size Judge score (mean / median)
Train
Test
Total
4,000 20,000 13,616 3.40 26.2 6,837 7.8 / 8.0
1,000 5,000 3,405 3.40 26.4 4,064 7.8 / 8.0
5,000 25,000 17,021 3.40 26.2 7,458 —
C. JL1-QA 1) Task Definition: Given a bi-temporal image pair (IA , IB ) and a natural-language question Q, the change question answering task requires generating a textual answer A that accurately responds to the question based on the visible surface changes. 2) Question Taxonomy: We define eight question types adapted for open-ended answer generation, as summarized in Table IV. Each type targets a distinct aspect of change understanding, ranging from binary existence judgments to causal reasoning, ensuring comprehensive coverage of user information needs. 3) Annotation Pipeline: As illustrated in Fig. 6, the JL1-QA annotation pipeline shares the same three-stage architecture as JL1-CC, with two key enhancements in the generation stage. Stage 1: Context-Enriched QA Generation. Beyond the three visual inputs (IA , IB , M ), the LLM additionally receives contextual metadata from JL1-CC: the change area ratio, the spatial change region description, and up to three selected change captions. This context enrichment grounds the QA generation in verified change descriptions, yielding more accurate and diverse question–answer pairs. The model generates five QA pairs per image, each covering a different question type randomly sampled from the taxonomy. Crucially, the prompt explicitly prohibits copying exact numerical metadata (e.g.,
6
Fig. 6. Overview of the JL1-QA annotation pipeline. The generation stage receives the image pair, change mask, and JL1-CC metadata (captions, CAR, change region) as context. The judge evaluates each QA pair for accuracy, quality, completeness, and redundancy.
10000
Type
Example Question
YES/NO
Has the vegetation in the lower half been removed? What happened to the farmland in the center? Where did the most significant change occur? Is the change large-scale or localized? What was present in the upper-left before the change? What type of development likely caused the changes? What do the new structures appear to be? Which area shows the most dramatic change?
WHAT WHERE HOW MUCH BEFORE/AFTER CAUSE DETAIL COMPARE
change percentages) into answers, as such precision cannot be derived from visual inspection alone. Stage 2: Multi-Criteria Judging. Each QA pair is scored on a 1–10 integer scale across four criteria: (1) answer accuracy: whether the answer is factually correct given the images; (2) question quality: whether the question is clear, natural, and non-trivial; (3) answer completeness: whether the answer adequately addresses the question; and (4) redundancy: whether the QA pair is substantially different from the others for the same image. QA pairs scoring below 7 are discarded. The most common rejection reasons are hallucinated precise percentages (score 1–4), redundancy with other QA pairs (score 5–6), and vague answers (score 5–6). Stage 3: Human Expert Verification. As with JL1-CC, domain experts review a subset of selected QA pairs to verify factual accuracy and check for systematic errors in question formulation or answer content. 4) Statistics: Table V summarizes the JL1-QA dataset statistics. From 24,995 generated QA pairs, 20,060 pass the quality threshold (score ≥ 7), yielding a pass rate of 80.3%.
Number of QA Pairs
TABLE IV Q UESTION T YPES IN JL1-QA WITH E XAMPLES
8000 6000
Rejected
Accepted
4000 2000 0
1
2
3
4
5
6
Judge Score
7
8
9
10
Fig. 7. Judge score distribution for JL1-QA. QA pairs scoring below 7 are rejected.
The higher pass rate compared to JL1-CC (68.1%) is attributed to the context enrichment from change captions, which reduces hallucination in the generation stage. The judge score distribution (Fig. 7) confirms effective quality discrimination: 51.4% of QA pairs score 9–10, 28.4% score 7–8, and 20.2% are rejected. The question type distribution (Fig. 8) shows broad coverage across all eight categories, with YES/NO (21.2%), WHERE (18.3%), and WHAT (16.9%) being the most frequent. Questions average 11.1 words and answers 19.3 words in length. III. C ONCLUSION In this paper, we presented JL1-CC&QA, a multi-task benchmark that extends the JL1-CD binary change detection dataset with two complementary annotation layers: change captioning (JL1-CC) and change question answering (JL1-QA). Both layers are produced through a three-stage pipeline—multi-modal LLM generation, visiongrounded LLM judging, and human expert verification— balancing scalability with factual reliability. We hope this
7
DETAIL YES/NO
4.4%
21.2%
BEFORE/AFTER
10.9%
HOW MUCH
11.6%
WHERE
18.3%
16.7%
OTHER
16.9%
WHAT Fig. 8. Distribution of question types in the 20,060 selected QA pairs of JL1-QA. TABLE V JL1-QA DATASET S TATISTICS Statistic Image pairs Generated QA pairs Selected QA pairs Avg. selected / pair Avg. question length (words) Avg. answer length (words) Judge score (mean / median)
Train
Test
Total
3,999 19,995 16,055 4.01 11.1 19.3 7.9 / 9.0
1,000 5,000 4,005 4.01 11.1 19.3 7.9 / 9.0
4,999 24,995 20,060 4.01 11.1 19.3 —
benchmark serves as a valuable resource for the community to advance multi-task change understanding in remote sensing. R EFERENCES [1] Z. Lv, T. Liu, J. A. Benediktsson, and N. Falco, “Land cover change detection techniques: Very-high-resolution optical images: A review,” IEEE Geoscience and Remote Sensing Magazine, vol. 10, no. 1, pp. 44–63, 2022. [2] T. Bai, L. Wang, D. Yin, K. Sun, Y. Chen, and W. Li, “Deep learning for change detection in remote sensing: A review,” Geo-spatial Information Science, vol. 26, no. 3, pp. 262–288, 2023. [3] C. Benedek and T. Szirányi, “Change detection in optical aerial images by a multilayer conditional mixed Markov model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 47, no. 10, pp. 3416–3430, 2009. [4] N. Bourdis, D. Marraud, and H. Sahbi, “Constrained optical flow for aerial image change detection,” in 2011 IEEE International Geoscience and Remote Sensing Symposium. Vancouver, BC, Canada: IEEE, 2011, pp. 4176–4179. [5] R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Urban change detection for multispectral earth observation using convolutional neural networks,” in IGARSS 2018 – 2018 IEEE International Geoscience and Remote Sensing Symposium. Valencia, Spain: IEEE, 2018, pp. 2115– 2118. [6] M. A. Lebedev, Y. V. Vizilter, O. V. Vygolov, V. A. Knyaz, and A. Y. Rubis, “Change detection in remote sensing images using conditional adversarial networks,” The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. XLII-2, pp. 565–571, 2018. [7] S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 1, pp. 574–586, 2019. [8] Q. Wang, Z. Yuan, Q. Du, and X. Li, “GETNET: A general end-to-end 2-D CNN framework for hyperspectral image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 57, no. 1, pp. 3–13, 2019.
[9] J. López-Fandiño, D. B. Heras, F. Argüello, and M. Dalla Mura, “GPU framework for change detection in multitemporal hyperspectral images,” International Journal of Parallel Programming, vol. 47, no. 2, pp. 272–292, 2019. [10] H. Chen and Z. Shi, “A spatial-temporal attention-based method and a new dataset for remote sensing image change detection,” Remote Sensing, vol. 12, no. 10, p. 1662, 2020. [11] C. Zhang, P. Yue, D. Tapete, L. Jiang, B. Shangguan, L. Huang, and G. Liu, “A deeply supervised image fusion network for change detection in high resolution bi-temporal remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 166, pp. 183– 200, 2020. [12] D. Peng, L. Bruzzone, Y. Zhang, H. Guan, H. Ding, and X. Huang, “SemiCDNet: A semisupervised convolutional neural network for change detection in high resolution remote-sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 59, no. 7, pp. 5891–5906, 2021. [13] L. Shen, Y. Lu, H. Chen, H. Wei, D. Xie, J. Yue, R. Chen, S. Lv, and B. Jiang, “S2Looking: A satellite side-looking dataset for building change detection,” Remote Sensing, vol. 13, no. 24, p. 5094, 2021. [14] R. Shao, C. Du, H. Chen, and J. Li, “SUNet: Change detection for heterogeneous remote sensing images from satellite and UAV using a dual-channel fully convolution network,” Remote Sensing, vol. 13, no. 18, p. 3750, 2021. [15] Q. Shi, M. Liu, S. Li, X. Liu, F. Wang, and L. Zhang, “A deeply supervised attention metric-based network and an open aerial image dataset for remote sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–16, 2022. [16] M. Liu, Z. Chai, H. Deng, and R. Liu, “A CNN-Transformer network with multiscale context aggregation for fine-grained cropland change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 4297–4306, 2022. [17] H. Li, F. Zhu, X. Zheng, M. Liu, and G. Chen, “MSCDUNet: A deep learning framework for built-up area change detection integrating multispectral, SAR, and VHR data,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 5163– 5176, 2022. [18] S. Holail, T. Saleh, X. Xiao, and D. Li, “AFDE-Net: Building change detection using attention-based feature differential enhancement for satellite imagery,” IEEE Geoscience and Remote Sensing Letters, vol. 20, pp. 1–5, 2023. [19] Z. Liu, R. Zhu, L. Gao, Y. Zhou, J. Ma, and Y. Gu, “JL1-CD: A new benchmark for remote sensing change detection and a robust multi-teacher knowledge distillation framework,” arXiv preprint arXiv:2502.13407, 2025. [20] R. C. Daudt, B. Le Saux, and A. Boulch, “Fully convolutional siamese networks for change detection,” in 2018 25th IEEE International Conference on Image Processing (ICIP). Athens, Greece: IEEE, 2018, pp. 4063–4067. [21] S. Fang, K. Li, J. Shao, and Z. Li, “SNUNet-CD: A densely connected siamese network for change detection of VHR images,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022. [22] H. Chen, Z. Qi, and Z. Shi, “Remote sensing image change detection with transformers,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022. [23] W. G. C. Bandara and V. M. Patel, “A transformer-based siamese network for change detection,” in IGARSS 2022 – 2022 IEEE International Geoscience and Remote Sensing Symposium. Fort Worth, TX, USA: IEEE, 2022, pp. 207–210. [24] H. Chen, J. Song, C. Han, J. Xia, and N. Yokoya, “ChangeMamba: Remote sensing change detection with spatiotemporal state space model,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1– 20, 2024. [25] Z. Zheng, A. Ma, L. Zhang, and Y. Zhong, “Change is everywhere: Single-temporal supervised object change detection in remote sensing imagery,” in IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE, 2021, pp. 15 193–15 202. [26] H. Chen, J. Song, C. Wu, B. Du, and N. Yokoya, “Exchange means change: An unsupervised single-temporal change detection framework based on intra- and inter-image patch exchange,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 206, pp. 87–105, 2023. [27] T. Zhou, F. Luo, C. Fu, T. Guo, X. Wang, and B. Du, “STMNet: Singletemporal mask-based network for self-supervised hyperspectral change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–12, 2024. [28] A. Toker, L. Kondmann, M. Weber, M. Eisenberger, A. Camero, and J. Hu, “DynamicEarthNet: Daily multi-spectral satellite dataset
8
for semantic change segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 21 126–21 135. [29] A. Van Etten, D. Hogan, J. M. Manso, J. Shermeyer, N. Weir, and R. Lewis, “The multi-temporal urban development SpaceNet dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 2021, pp. 6394– 6403. [30] S. Shi, Y. Zhong, Y. Liu, J. Wang, Y. Wan, J. Zhao, P. Lv, L. Zhang, and D. Li, “Multi-temporal urban semantic understanding based on GF-2 remote sensing imagery: From tri-temporal datasets to multi-task mapping,” International Journal of Digital Earth, vol. 16, no. 1, pp. 3321–3347, 2023. [31] Y. Zhao, H.-C. Li, S. Lei, N. Liu, J. Pan, and T. Celik, “COUD: Continual urbanization detector for time series building change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 19 601–19 615, 2024. [32] A. Fujita, K. Sakurada, T. Imaizumi, R. Ito, S. Hikosaka, and R. Nakamura, “Damage detection from aerial images via convolutional neural networks,” in 2017 Fifteenth IAPR International Conference on Machine Vision Applications (MVA). Nagoya, Japan: IEEE, 2017, pp. 5–8. [33] C. Wu, L. Zhang, and B. Du, “Kernel slow feature analysis for scene change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 55, no. 4, pp. 2367–2384, 2017. [34] R. C. Daudt, B. Le Saux, A. Boulch, and Y. Gousseau, “Multitask learning for large-scale semantic change detection,” Computer Vision and Image Understanding, vol. 187, p. 102783, 2019. [35] R. Gupta, B. Goodman, N. Patel, R. Hosfelt, S. Sajeev, E. Heim, J. Doshi, K. Lucas, H. Choset, and M. Gaston, “Creating xBD: A dataset for assessing building damage from satellite imagery,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2019, pp. 10–17. [36] S. Saha, P. Ebel, and X. X. Zhu, “Self-supervised multisensor change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–10, 2022. [37] K. Yang, G.-S. Xia, Z. Liu, B. Du, Y. Wen, and M. Pelillo, “Asymmetric siamese networks for semantic change detection in aerial images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–18, 2022. [38] Z. Zheng, Y. Zhong, J. Wang, A. Ma, and L. Zhang, “Building damage assessment for rapid disaster response with a deep object-based semantic change detection framework: From natural disasters to manmade disasters,” Remote Sensing of Environment, vol. 265, p. 112636, 2021. [39] M. Hu, C. Wu, L. Zhang, and B. Du, “Hyperspectral anomaly change detection based on autoencoder,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 14, pp. 3750– 3762, 2021. [40] X. Li, G. Zhang, H. Cui, S. Hou, S. Wang, X. Li, Y. Chen, Z. Li, and L. Zhang, “MCANet: A joint semantic segmentation framework of optical and SAR images for land use classification,” International Journal of Applied Earth Observation and Geoinformation, vol. 106, p. 102638, 2022. [41] S. Tian, A. Ma, Z. Zheng, Y. Zhong, X. Tan, and L. Zhang, “Large-scale deep learning based binary and semantic change detection in ultra high resolution remote sensing imagery: From benchmark datasets to urban application,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 193, pp. 164–186, 2022. [42] P. Yuan, Q. Zhao, X. Zhao, X. Wang, X. Long, and Y. Zheng, “A transformer-based siamese network and an open optical dataset for semantic change detection of remote sensing images,” International Journal of Digital Earth, vol. 15, no. 1, pp. 1506–1525, 2022. [43] C. Pang, J. Wu, J. Ding, C. Song, and G.-S. Xia, “Detecting building changes with off-nadir aerial images,” Science China Information Sciences, vol. 66, no. 4, p. 140306, 2023. [44] V. Marsocci, V. Coletta, R. Ravanelli, S. Scardapane, and M. Crespi, “Inferring 3D change detection from bitemporal optical images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 196, pp. 325– 339, 2023. [45] X. He, S. Zhang, B. Xue, T. Zhao, and T. Wu, “Cross-modal change detection flood extraction based on convolutional neural network,” International Journal of Applied Earth Observation and Geoinformation, vol. 117, p. 103197, 2023. [46] M. Rahnemoonfar, T. Chowdhury, and R. Murphy, “RescueNet: A high resolution UAV semantic segmentation dataset for natural disaster damage assessment,” Scientific Data, vol. 10, no. 1, p. 913, 2023.
[47] M. Liu, S. Lin, Y. Zhong, Q. Shi, and J. Li, “A memory-guided network and a novel dataset for cropland semantic change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–13, 2024. [48] K. Tang, F. Xu, X. Chen, Q. Dong, Y. Yuan, and J. Chen, “The ClearSCD model: Comprehensively leveraging semantics and change relationships for semantic change detection in high spatial resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 211, pp. 299–317, 2024. [49] Y. Sun, Y. Wang, and M. Eineder, “QuickQuakeBuildings: Postearthquake SAR-optical dataset for quick damaged-building detection,” IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024. [50] S. Fang, W. Li, Y. Song, Z. Li, and J. Zhao, “Rethinking semantic change detection from a semantic alignment perspective,” Authorea Preprints, 2025. [51] X. Tan, G. Chen, X. Zhang, T. Wang, J. Wang, K. Wang, and T. Miao, “TripleS: Mitigating multi-task learning conflicts for semantic change detection in high-resolution remote sensing imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 374–401, 2025. [52] H. Chen, J. Song, O. Dietrich, C. Broni-Bediako, W. Xuan, J. Wang, X. Shao, Y. Wei, J. Xia, C. Lan, K. Schindler, and N. Yokoya, “BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response,” Earth System Science Data Discussions, pp. 1–51, 2025. [53] H. Wang, W. He, Z. Li, and N. Yokoya, “Cross-scenario damaged building extraction network: Methodology, application, and efficiency using single-temporal HRRS imagery,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 228, pp. 228–248, 2025. [54] Y. Wang, D. Hong, J. Sha, L. Gao, L. Liu, Y. Zhang, and X. Rong, “Spectral-spatial-temporal transformers for hyperspectral image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022. [55] W. Xie, X. Xu, and Y. Li, “Decentralized federated GAN for hyperspectral change detection in edge computing,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 8863–8874, 2024. [56] W. Dong, J. Ren, S. Xiao, L. Fang, J. Qu, and Y. Li, “Cycle translationbased collaborative training for hyperspectral-RGB multimodal change detection,” IEEE Transactions on Image Processing, vol. 34, pp. 6347– 6360, 2025. [57] X. Li, Z. Du, Y. Huang, and Z. Tan, “A deep translation (GAN) based change detection network for optical and SAR remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 179, pp. 14–34, 2021. [58] J. Alatalo, T. Sipola, and M. Rantonen, “Improved difference images for change detection classifiers in SAR imagery using deep learning,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023. [59] Z. Liu, J. Zhang, W. Wang, and Y. Gu, “M2CD: A unified multimodal framework for optical-SAR change detection with mixture of experts and self-distillation,” IEEE Geoscience and Remote Sensing Letters, vol. 22, pp. 1–5, 2025. [60] Y. Shen, S. Yao, Z. Qiang, and G. Pei, “SD-Mamba: A lightweight synthetic-decompression network for cross-modal flood change detection,” International Journal of Applied Earth Observation and Geoinformation, vol. 136, p. 104409, 2025. [61] L. Ding, H. Guo, S. Liu, L. Mou, J. Zhang, and L. Bruzzone, “Bitemporal semantic reasoning for the semantic change detection in HR remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022. [62] H. Jhamtani and T. Berg-Kirkpatrick, “Learning to describe differences between pairs of similar images,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2018, pp. 4024–4034. [63] D. H. Park, T. Darrell, and A. Rohrbach, “Robust change captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 4624–4633. [64] G. Hoxha, S. Chouaf, F. Melgani, and Y. Smara, “Change captioning: A new paradigm for multitemporal remote sensing image analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–14, 2022. [65] C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–20, 2022.
9
[66] C. Liu, K. Chen, B. Zhang, Z. Qi, Z. Zou, and Z. Shi, “Change detection as visual question answering,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–18, 2024. [67] X. Li et al., “SECOND-CC: A large-scale change captioning dataset with real-world challenges,” arXiv preprint arXiv:2501.10075, 2025. [68] Z. Chen et al., “CCExpert: Advancing RSCC via an expert-level large vision-language model,” arXiv preprint arXiv:2411.11360, 2024. [69] Z. Li et al., “Towards real-world change captioning: A large-scale disaster-focused dataset,” in Advances in Neural Information Processing Systems (NeurIPS), 2025. [70] U. Türker et al., “Multispectral remote sensing image change captioning with Sentinel-2,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 25 410–25 423, 2025. [71] Z. Yuan, L. Mou, Y. Hua, and X. X. Zhu, “Change detection meets visual question answering,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022. [72] K. Li, F. Dong, D. Wang, S. Li, Q. Wang, X. Gao, and T.-S. Chua, “Show me what and where has changed? question answering and grounding for remote sensing change detection,” arXiv preprint arXiv:2410.23828, 2024. [73] P. Wu et al., “ChangeChat: An interactive model for remote sensing change analysis via multimodal instruction tuning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. [74] J. Irvin, A. Gruca et al., “TEOChat: A large vision-language assistant for temporal earth observation data,” in International Conference on Learning Representations (ICLR), 2025. [75] J. Wang, W. Xuan, H. Qi, Z. Liu, K. Liu, Y. Wu, H. Chen, J. Song, J. Xia, Z. Zheng, and N. Yokoya, “DisasterM3: A remote sensing vision-language dataset for disaster damage assessment and response,” arXiv preprint arXiv:2505.21089, 2025. [76] P. Chen, B. Zhang, D. Hong, Z. Chen, X. Yang, and B. Li, “FCCDN: Feature constraint network for VHR image change detection,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 187, pp. 101– 119, 2022. [77] H. Chen, F. Pu, R. Yang, and X. Xu, “RDP-Net: Region detail preserving network for change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–10, 2022. [78] H. Chen, W. Li, S. Chen, and Z. Shi, “Semantic-aware dense representation learning for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–18, 2022. [79] H. Chen, W. Li, and Z. Shi, “Adversarial instance augmentation for building change detection in remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–16, 2021. [80] Z. Li, C. Tang, X. Liu, W. Zhang, J. Dou, and L. Wang, “Lightweight remote sensing change detection with progressive feature aggregation and supervised attention,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–12, 2023. [81] A. Codegoni, G. Lombardi, and A. Ferrari, “TINYCD: A (not so) deep learning model for change detection,” Neural Computing and Applications, vol. 35, no. 11, pp. 8471–8486, 2023. [82] Y. Xing, J. Jiang, J. Xiang, E. Yan, Y. Song, and D. Mo, “LightCDNet: Lightweight change detection network based on VHR images,” IEEE Geoscience and Remote Sensing Letters, vol. 20, pp. 1–5, 2023. [83] T. Lei, X. Geng, H. Ning, Z. Lv, M. Gong, and Y. Jin, “Ultralightweight spatial-spectral feature cooperation network for change detection in remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023. [84] H. Zhang, M. Lin, G. Yang, and L. Zhang, “ESCNet: An end-to-end superpixel-enhanced change detection network for very-high-resolution remote sensing images,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 1, pp. 28–42, 2023. [85] Y. Feng, J. Jiang, H. Xu, and J. Zheng, “Change detection on remote sensing images using dual-branch multilevel intertemporal network,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–15, 2023. [86] Y. Ye, M. Wang, L. Zhou, G. Lei, J. Fan, and Y. Qin, “Adjacent-level feature cross-fusion with 3-D CNN for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–14, 2023. [87] C. Han, C. Wu, H. Guo, M. Hu, and H. Chen, “HANet: A hierarchical attention network for change detection with bitemporal very-highresolution remote sensing images,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 3867– 3878, 2023.
[88] C. Han, C. Wu, H. Guo, M. Hu, J. Li, and H. Chen, “Change guiding network: Incorporating change prior to guide change detection in remote sensing imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 16, pp. 8395–8407, 2023. [89] Y. Cao and X. Huang, “A full-level fused cross-task transfer learning method for building change detection using noise-robust pretrained networks on crowdsourced labels,” Remote Sensing of Environment, vol. 284, p. 113371, 2023. [90] M. Cheng, W. He, Z. Li, G. Yang, and H. Zhang, “Harmony in diversity: Content cleansing change detection framework for very-highresolution remote-sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 218, pp. 1–19, 2024. [91] Y. Zhao, T. Celik, N. Liu, F. Gao, and H.-C. Li, “SSLChange: A selfsupervised change detection framework based on domain adaptation,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, 2024. [92] Z. Zheng, Y. Zhong, J. Zhao, A. Ma, and L. Zhang, “Unifying remote sensing change detection via deep probabilistic change models: From principles, models to applications,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 215, pp. 239–255, 2024. [93] Z. Liu, J. Zhang, W. Wang, and Y. Gu, “A novel multibranch selfdistillation framework for optimizing remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 22 642–22 655, 2025. [94] Y. Sun, L. Lei, Z. Li, G. Kuang, and Q. Yu, “Detecting changes without comparing images: Rules induced change detection in heterogeneous remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 241–257, 2025. [95] L. Sun, M. Jin, J. Yan, i. He, and L. Cao, “Semantic-TemporalNet: A novel urban block change detection method based on semantic coherence analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–16, 2025. [96] Z. Liu, J. Zhang, W. Wang, and Y. Gu, “A novel multibranch selfdistillation framework for optimizing remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 22 642–22 655, 2025. [97] Z. Liu, J. Zhang, Q. Shi, and Y. Gu, “Unlearningcd: Distill to forget in change detection via perturbed knowledge,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, pp. 1–13, 2026. [98] C. Zhang, L. Wang, S. Cheng, and Y. Li, “SwinSUNet: Pure transformer network for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022. [99] Q. Li, R. Zhong, X. Du, and Y. Du, “TransUNetCD: A hybrid transformer network for change detection in optical remote-sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1– 19, 2022. [100] N. Shi, K. Chen, and G. Zhou, “A divided spatial and temporal context network for remote sensing change detection,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 4897–4908, 2022. [101] Z. Zheng, Y. Zhong, S. Tian, A. Ma, and L. Zhang, “ChangeMask: Deep multi-task encoder-transformer-decoder architecture for semantic change detection,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 183, pp. 228–239, 2022. [102] M. Noman, M. Fiaz, H. Cholakkal, S. Narayan, R. M. Anwer, S. Khan, and F. S. Khan, “Remote sensing change detection with transformers trained from scratch,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–14, 2024. [103] S. Fang, K. Li, and Z. Li, “Changer: Feature interaction is what you need for change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–11, 2023. [104] M. Bernhard, N. Strauß, and M. Schubert, “MapFormer: Boosting change detection by using pre-change information,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023, pp. 16 837–16 846. [105] W. Yu, X. Zhang, S. Das, X. X. Zhu, and P. Ghamisi, “MaskCD: A remote sensing change detection network based on mask classification,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1– 16, 2024. [106] H. Zhang, H. Chen, C. Zhou, K. Chen, C. Liu, Z. Zou, and Z. Shi, “BiFA: Remote sensing image change detection with bitemporal feature alignment,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–17, 2024. [107] H. Zhang, H. Guo, K. Chen, H. Chen, Z. Zou, and Z. Shi, “FoBa: A foreground-background co-guided method and new bench-
10
mark for remote sensing semantic change detection,” arXiv preprint arXiv:2509.15788, 2025. [108] D. Zhu, X. Huang, H. Huang, H. Zhou, and Z. Shao, “Change3D: Revisiting change detection and captioning from a video modeling perspective,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025, pp. 24 011–24 022. [109] H. Zhang, K. Chen, C. Liu, H. Chen, Z. Zou, and Z. Shi, “CDMamba: Incorporating local clues into Mamba for remote sensing image binary change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–16, 2025. [110] S. Liu, S. Wang, W. Zhang, T. Zhang, M. Xu, M. Yasir, and S. Wei, “CD-STMamba: Towards remote sensing image change detection with spatio-temporal interaction Mamba model,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 10 471–10 485, 2025. [111] J. Zhao, J. Xie, Y. Zhou, W. Du, R. Yao, and A. El Saddik, “STMamba: Spatio-temporal synergistic model for remote sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–13, 2025. [112] Y. Liu, G. Cheng, Q. Sun, C. Tian, and L. Wang, “CWMamba: Leveraging CNN-Mamba fusion for enhanced change detection in remote sensing images,” IEEE Geoscience and Remote Sensing Letters, vol. 22, pp. 1–5, 2025. [113] X. Liu, C. Dai, L. Ding, Z. Zhang, Y. Li, X. Zuo, M. Li, H. Wang, and Y. Miao, “GSTM-SCD: Graph-enhanced spatio-temporal state space model for semantic change detection in multi-temporal remote sensing images,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 230, pp. 73–91, 2025. [114] Y. Li, W. Liu, E. Li, L. Zhang, and X. Li, “SAM-Mamba: A two-stage change detection network combining the adapting segment anything and Mamba models,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 21 607–21 619, 2025. [115] J. N. Paranjape, C. De Melo, and V. M. Patel, “A Mamba-based siamese network for remote sensing change detection,” in 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson, AZ, USA: IEEE, 2025, pp. 1186–1196. [116] S. Dong, L. Wang, B. Du, and X. Meng, “ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 208, pp. 53–69, 2024. [117] Z. Zheng, Y. Zhong, L. Zhang, and S. Ermon, “Segment any change,” in Advances in Neural Information Processing Systems, 2024, pp. 81 204– 81 224. [118] L. Ding, K. Zhu, D. Peng, H. Tang, K. Yang, and L. Bruzzone, “Adapting segment anything model for change detection in VHR remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–11, 2024. [119] K. Li, X. Cao, and D. Meng, “A new learning paradigm for foundation model-based remote-sensing change detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–12, 2024. [120] X. Tan, G. Chen, T. Wang, J. Wang, and X. Zhang, “Segment change model (SCM) for unsupervised change detection in VHR remote sensing images: A case study of buildings,” in IGARSS 2024 – 2024 IEEE International Geoscience and Remote Sensing Symposium. Athens, Greece: IEEE, 2024, pp. 8577–8580. [121] K. Li, X. Cao, Y. Deng, J. Song, J. Liu, D. Meng, and Z. Wang, “SemiCD-VL: Visual-language model guidance makes better semisupervised change detector,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, pp. 1–13, 2024. [122] J. Huang, J. Bao, M. Xia, and X. Yuan, “SAM-based efficient feature integration network for remote sensing change detection: A case study on Macao sea reclamation,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 16 916–16 928, 2025. [123] Y. Qin, C. Wang, Y. Fan, and C. Pan, “SAM2-CD: Remote sensing image change detection with SAM2,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 18, pp. 24 575–24 587, 2025. [124] K. Li, X. Cao, Y. Deng, C. Pang, Z. Xin, D. Meng, and Z. Wang, “DynamicEarth: How far are we from open-vocabulary change detection?” arXiv preprint arXiv:2501.12931, 2025. [125] F. Jiang, M. Gong, T. Zhan, and X. Fan, “A semisupervised GAN-based multiple change detection framework in multi-spectral images,” IEEE Geoscience and Remote Sensing Letters, vol. 17, no. 7, pp. 1223–1227, 2020.
[126] P. Jian, K. Chen, and W. Cheng, “GAN-based one-class classification for remote-sensing image change detection,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2022. [127] Z. Zheng, S. Tian, A. Ma, L. Zhang, and Y. Zhong, “Scalable multitemporal remote sensing change data generation via simulating stochastic change process,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 21 761–21 770. [128] W. G. C. Bandara, N. G. Nair, and V. M. Patel, “DDPM-CD: Denoising diffusion probabilistic models as feature extractors for change detection,” arXiv preprint arXiv:2206.11892, 2022. [129] Z. Zheng, S. Ermon, D. Kim, L. Zhang, and Y. Zhong, “Changen2: Multi-temporal remote sensing generative change foundation model,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 2, pp. 725–741, 2024. [130] K. Tang and J. Chen, “ChangeAnywhere: Sample generation for remote sensing change detection via semantic latent diffusion model,” arXiv preprint arXiv:2404.08892, 2024. [131] J. Jia, G. Lee, Z. Wang, Z. Lyu, and Y. He, “Siamese meets diffusion network: SMDNet for enhanced change detection in high-resolution RS imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 17, pp. 8189–8202, 2024. [132] P. Han, Y. Gao, G. Chen, B. Zhao, and X. Li, “Hierarchical diffusion model for remote sensing image change detection,” IEEE Transactions on Geoscience and Remote Sensing, 2025. [133] Q. Zang, J. Yang, S. Wang, D. Zhao, W. Yi, and Z. Zhong, “ChangeDiff: A multi-temporal change detection data generator with flexible text prompts via diffusion model,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2025, pp. 9763–9771. [134] Y. Benidir, N. Gonthier, and C. Mallet, “The change you want to detect: Semantic change detection in earth observation with hybrid data generation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025, pp. 2204–2214. [135] C. Liu, K. Chen, H. Zhang, Z. Qi, Z. Zou, and Z. Shi, “Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–16, 2024. [136] W. Xu, Z. Yu, B. Mu, Z. Wei, Y. Zhang, G. Li, and M. Peng, “RSAgent: Automating remote sensing tasks through intelligent agent,” arXiv preprint arXiv:2406.07089, 2024. [137] P. Feng, Z. Lv, J. Ye, X. Wang, X. Huo, J. Yu, W. Xu, W. Zhang, L. Bai, C. He, and W. Li, “Earth-agent: Unlocking the full landscape of earth observation with agents,” arXiv preprint arXiv:2509.23141, 2025. [138] H. Hu, P. Wang, Y. Feng, K. Wei, W. Yin, W. Diao, M. Wang, H. Bi, K. Kang, T. Ling, K. Fu, and X. Sun, “RingMo-Agent: A unified remote sensing foundation model for multi-platform and multi-modal reasoning,” arXiv preprint arXiv:2507.20776, 2025.