ConceptioArchivearXiv CS
arXiv CSopen access

ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

1

ChangeQuery: Advancing Remote Sensing Change Analysis for Natural and Human-Induced Disasters from Visual Detection to Semantic Understanding

arXiv:2604.22333v1 [cs.CV] 24 Apr 2026

Dongwei Sun, Jing Yao, Senior Member, IEEE, Kan Wei, Xiangyong Cao, Chen Wu, Member, IEEE, Zhenghui Zhao, Pedram Ghamisi, Senior Member, IEEE, Jun Zhou, Fellow, IEEE, Jón Atli Benediktsson, Life Fellow, IEEE

Abstract—Rapid situational awareness is critical in postdisaster response. While remote sensing damage assessment is evolving from pixel-level change detection to high-level semantic analysis, existing vision-language methodologies still struggle to provide actionable intelligence for complex strategic queries. They remain severely constrained by unimodal optical dependence, a prevailing bias towards natural disasters, and a fundamental lack of grounded interactivity. To address these limitations, we present ChangeQuery, a unified multimodal framework designed for comprehensive, all-weather disaster situation awareness. To overcome modality constraints and scenario biases, we construct the Disaster-Induced Change Query (DICQ) dataset, a large-scale benchmark coupling pre-event optical semantics with post-event SAR structural features across a balanced distribution of natural catastrophes and armed conflicts. Furthermore, to provide the high-quality supervision required for interactive reasoning, we propose a novel Automated Semantic Annotation Pipeline. Adhering to a “statistics-first, generation-later” paradigm, this engine automatically transforms raw segmentation masks into grounded, hierarchical instruction sets, effectively equipping the model with fine-grained spatial and quantitative awareness. Trained on this structured data, the ChangeQuery architecture operates as an interactive disaster analyst. It supports multi-task reasoning driven by diverse user queries, delivering precise damage quantification, regionspecific descriptions, and holistic post-disaster summaries. Extensive experiments demonstrate that ChangeQuery establishes a new state-of-the-art, providing a robust and interpretable solution for complex disaster monitoring. The code is available at https://sundongwei.github.io/changequery/. Index Terms—Remote Sensing Change Detection; PostDisaster Damage Assessment; Optical–SAR Data Fusion; Vision Language Models

I. I NTRODUCTION D. Sun and X. Cao are with the School of Computer Science and Technology and the Ministry of Education Key Lab for Intelligent Networks and Network Security, Xi’an Jiaotong University, Xi’an 710049, China. J. Yao and K. Wei are with the State Key Laboratory of Remote Sensing and Digital Earth, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing, 100094, China. C. Wu and Z. Zhao are with the State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, China. P. Ghamisi is with the Helmholtz-Zentrum Dresden-Rossendorf, 09599 Freiberg, Germany; the Lancaster Environment Centre, Lancaster University, LA1 4YR Lancaster, U.K.; and the Faculty of Electrical and Computer Engineering, University of Iceland, Reykjavı́k, Iceland. J. Zhou is with the School of Information and Communication Technology, Griffith University, Nathan, QLD 4111, Australia. J. A. Benediktsson is with the Faculty of Electrical and Computer Engineering, University of Iceland, 101 Reykjavı́k, Iceland.

Change Detection (CD)

Change Captioning (CC)

Pixel-level only Prediction Task

Static & Generic description Task

Change Model Pre-disaster

Pre-disaster

Post-disaster

Lack of semantic captioning

Post-disaster

Caption Model The post disaster show some buildings are destroyed in the image… Unimodal &Weather-affected

(a) Conventional Paradigms User: Describe conditions in left/Evaluation/Summary? All-Weather

ChangeQuery Model Pre-disaster

Post-disaster

Model: Region Analysis⇒ Left: The buildings in the left have been demolished… Level Analysis⇒ Evaluation: The Level 3,Severe Damage… Global Analysis ⇒ Summary: An overview of the zones indicates the top and bottom areas …

(b) ChangeQuery Paradigm

Fig. 1: Conceptual comparison of ChangeQuery. (a) CD yields pixel-level masks without semantics, while RSICC produces generic captions from unimodal (Optical-Optical) inputs. (b) ChangeQuery adopts a multimodal paradigm (Optical-SAR) for all-weather, multi-granular analysis, including region-level descriptions and hierarchical damage assessment.

T

HE escalating frequency of catastrophic events, whether triggered by natural forces such as earthquakes and floods or anthropogenic crises like armed conflicts, has underscored the urgency for resilient monitoring systems [1]–[3]. For emergency responders, the primary challenge is not merely data acquisition, but rapid situational awareness. Earth Observation (EO) has become the de facto standard for monitoring denied areas [4] [5]. However, the operational pipeline remains heavily fragmented. As conceptually illustrated in Fig. 1(a), traditional CD algorithms typically treat damage assessment as a dense pixel classification problem, outputting binary or categorical masks [6]–[8]. While visually indicative, these masks represent raw data rather than actionable intelligence. They inherently lack the semantic capacity to directly answer highlevel questions, such as “What is the distribution of destroyed structures in the northern sector?” or “Does the damage pattern suggest localized strikes or widespread inundation?” To inject semantic understanding into this process, early research efforts pivoted towards Remote Sensing Image Change Captioning (RSICC) [9]–[11]. Yet, as depicted in the conventional captioning paradigm of Fig. 1(a), these models merely generate static, generic sentences. Consequently, they still fail to address the aforementioned complex strategic queries. While modern Visual Language Models (VLMs) [12]–[14] possess the reasoning potential to answer such interactive questions, existing remote sensing VLM approaches exhibit

2

three critical limitations that hinder their deployment in realworld crises: (1) Modality Constraint: Most remote sensing VLMs are trained exclusively on optical imagery [15]. This reliance renders them ineffective during the critical 72-hour golden rescue window, which is often plagued by adverse weather, cloud cover, or smoke. SAR, with its active sensing capability, offers a viable solution, yet aligning its non-intuitive backscatter features with semantic language models remains an open challenge [16]. (2) Scenario Bias: Current datasets predominantly focus on natural disasters, neglecting the distinct morphological signatures of man-made destruction [17], [18]. Unlike the contiguous damage patterns of floods or fires, conflict-induced damage is often discrete, structural, and highly irregular. A robust system must generalize across both natural and anthropogenic domains. (3) Lack of Grounded Interactivity: Current models lack the granularity to engage in multi-turn reasoning based on physical evidence. They are largely unable to perform precise object counting or oriented localization, which are essential quantitative metrics for logistical planning. To address these challenges, we introduce ChangeQuery, a unified framework that redefines disaster analysis as a multimodal, instruction-following task, as illustrated in Fig. 1(b). To support this, we construct the DICQ dataset, a large-scale benchmark that uniquely couples pre-event optical semantics with post-event SAR structural features [19]. By synthesizing heterogeneous data sources, DICQ covers a spectrum of scenarios ranging from seismic events to urban conflicts, ensuring all-weather robustness. Distinct from previous works that rely on manual captioning or simple template filling [9], our approach enables the model to generate comprehensive post-disaster change summaries and conduct robust damage evaluation across affected regions, delivering the interactive and granular analysis required for modern disaster response. To summarize, the contributions of this work are four-fold: • Building upon the BRIGHT benchmark [20], we construct the DICQ dataset by integrating independently curated Armed Conflict zones (AC-1 and AC-2) to establish the first heterogeneous Optical-SAR instructiontuning benchmark. Spanning diverse global locations with varying spatial resolutions (0.3m–1.0m), DICQ ensures a balanced distribution of natural and man-made disasters, fostering the learning of scale-invariant and sensoragnostic features to effectively mitigate domain shift in all weather analysis. • We propose a novel Automated Semantic Annotation Pipeline adhering to a “statistics-first, generation-later” paradigm to provide fine-grained supervision for interactive reasoning. By integrating spatially-aware partitioning, PCA-based localization, and logic-driven grading, it transforms raw masks into grounded instruction sets, effectively endowing the model with precise spatial and quantitative awareness while mitigating hallucinations. • We propose ChangeQuery, a unified multimodal vision language framework that supports multitask reasoning. By integrating a Change-Aware Difference Module with a progressive training strategy, ChangeQuery effectively aligns heterogeneous features with high-level semantic

reasoning, establishing a new baseline for interpretable disaster assessment. • Extensive experiments demonstrate that our proposed method achieves State-of-the-Art (SOTA) performance. ChangeQuery consistently surpasses existing generalist VLMs and specialized remote sensing models across multiple metrics, and demonstrates strong cross-domain generalization across heterogeneous disaster types, including both natural and man-made events. This capability establishes a new baseline for interpretable and interactive disaster damage assessment. The remainder of this paper is organized as follows. Section II reviews related work. Section III introduces the DICQ dataset and its annotation pipeline. Section IV presents the proposed ChangeQuery framework. Section V reports experimental results. Finally, Section VI concludes the paper. II. R ELATED W ORKS A. Change Detection: From Algebraic to Deep Learning Change detection constitutes a fundamental task in remote sensing, aiming to identify distinct differences in the state of a phenomenon by observing it at different times. Historically, this field was dominated by algebraic methods applied to homogeneous optical imagery, such as image differencing, ratioing, and Change Vector Analysis [21], [22]. With the rapid advancement of deep learning, the paradigm has shifted towards fully convolutional Siamese networks [23], [24]. Recently, Transformer-based architectures have become the standard, utilizing self-attention mechanisms to model long-range spatio-temporal dependencies, thereby improving robustness against seasonal variations and sensor noise [25], [26]. However, in the context of rapid disaster response, the assumption of homogeneous data availability is often violated. Optical sensors [27]–[29] are frequently rendered ineffective by cloud cover or smoke. Consequently, heterogeneous change detection, particularly the fusion of pre-event optical imagery with post-event SAR data, has emerged as a critical direction. SAR offers all weather imaging capabilities but introduces significant geometric and radiometric distortions compared to optical data [30]. To bridge this domain gap, recent methodologies have employed cycle-consistent adversarial networks for image-to-image translation [31], [32] or graph neural networks to align structural features across modalities [33], [34]. Techniques focusing on feature space alignment have also been developed to correlate these distinct modalities for identifying damage patterns in both natural disasters and anthropogenic conflicts [35], [36]. Despite these advancements, existing change detection methods predominantly formulate the task as a semantic segmentation problem [37], [38]. The final output is typically a binary or categorical mask. While precise in localization, these masks lack high-level semantic interpretability and cannot explicitly describe the nature of damage or quantify severity in a structured manner useful for strategic decision making. B. Change Interpretation: From Mask to Language To address the semantic limitations of binary masks, the field has gradually expanded towards remote sensing im-

3

age captioning [39]. Early works utilized encoder-decoder architectures to generate static descriptions of single-temporal scenes. With the advent of bi-temporal analysis, research has shifted towards Change Captioning, which aims to describe differences between two images in natural language. Pioneering works like the RSICCformer [40] and attentive fusion networks [41]–[44] have demonstrated the feasibility of generating sentences from bi-temporal pairs. However, these approaches typically function as “image to text translators” rather than “intelligent analysts”. They are largely limited to homogeneous optical data and produce fixed captions that lack interactivity. Current models struggle to perform grounded tasks such as counting damaged structures, estimating damage ratios, or analyzing specific subregions upon user request. This creates a bottleneck where downstream experts must still manually interpret the generated text to extract operational metrics.

TABLE I: Overview of disaster events contained in the DICQ dataset. The dataset covers three major categories: Natural Disasters, Man-made Disasters, and Conflicts, with spatial resolutions ranging from 0.3m to 1m. ID

Type

Date

GSD (m)

Natural Disasters 1 Goma, DR Congo 2 Les Cayes, Haiti 3 La Palma, Spain 4 Boulder, USA 5 Turkey 6 Kyaukpyu, Myanmar 7 Maui, USA 8 Morocco 9 Derna, Libya 10 Acapulco, Mexico 11 Noto, Japan

Disaster Area

Volcano Eruption (VE) Earthquake (EQ) Volcano Eruption (VE) Wildfire (WF) Earthquake (EQ) Cyclone (CC) Wildfire (WF) Earthquake (EQ) Flood (FL) Hurricane (HC) Earthquake (EQ)

May 22, 2021 Aug 14, 2021 Sep 19, 2021 Dec 30, 2021 Feb 06, 2023 May 14, 2023 Aug 08, 2023 Sep 08, 2023 Sep 10, 2023 Oct 25, 2023 Jan 01, 2024

0.33 0.48 0.30-0.35 0.60 0.30-0.35 0.60 0.60 0.35-0.40 0.35 0.35-0.80 0.50

Man-made Disasters 12 Beirut, Lebanon 13 Bata, Eq. Guinea

Explosion (EP) Explosion (EP)

Aug 04, 2020 Mar 07, 2021

1.00 0.50

Conflict Scenarios 14 Armed Conflict Zone 1 15 Armed Conflict Zone 2

Armed Conflict (AC) Armed Conflict (AC)

2022 2024

0.60 1.00

C. Multimodal Learning in Complex Scenarios The emergence of General-domain Multimodal Large Language Models (MLLMs), such as LLaVA [45] and MiniGPT-4 [46], has revolutionized visual reasoning. By aligning visual encoders with Large Language Models (LLMs), these systems demonstrate impressive zero-shot generalization. However, applying them directly to remote sensing presents unique challenges due to the bird’s-eye view perspective, small object scales, and multispectral characteristics [47]–[49]. A critical gap in current literature is the handling of complex, nonnatural disaster scenarios. While datasets focusing on natural disasters exist, they often neglect the unique signatures of conflict zones, where damage is discrete and structurally complex. Furthermore, existing multimodal fusion techniques in remote sensing often treat SAR as a secondary auxiliary channel rather than a primary source of structural information [50]. Our work departs from this by treating optical and SAR as equally critical: optical for pre-event semantic context and SAR for post-event structural verification [51] [52], thereby enabling a unified framework for all weather, grounded disaster analysis. III. C ONSTRUCTION OF THE DICQ DATASET A. Data Overview and Composition The DICQ dataset is meticulously curated to address the limitations of existing benchmarks, which are often restricted to single modalities or specific disaster types. DICQ establishes a large scale, heterogeneous benchmark comprising approximately 70,000 bi-temporal image pairs. A core distinguishing feature of our dataset is its comprehensive typological coverage, as shown in Fig. 2 which categorizes disaster events into three distinct classes: Natural Disasters(First Row), Manmade Disasters(Second Row), and Conflict Scenarios(Third and Fourth Row). DICQ classifies disaster scenarios into three hierarchical categories to test the model’s generalization capabilities in Table I: (1) Natural Disasters: This category represents the foundational component of the dataset (IDs 1-11), encompassing a wide spectrum of meteorological and geological events across diverse terrains. It includes: 1) Seismic Events: Earthquakes in

Les Cayes (Haiti), Turkey, Morocco, and Noto (Japan), providing rich samples of structural collapse and debris. 2) Volcanic Eruptions: Events in Goma (DR Congo) and La Palma (Spain), featuring unique damage patterns caused by lava flow and ash accumulation. 3) Hydro-meteorological Disasters: Wildfires in Boulder and Maui (USA); Cyclones and Hurricanes in Myanmar (Kyaukpyu) and Mexico (Acapulco); and Floods in Libya (Derna). These scenarios challenge the model to identify inundation and burn scars amidst complex urban backgrounds. (2) Man-made Disasters: To test the model’s ability to detect sudden, high-intensity anthropogenic destruction, we incorporate industrial accidents (IDs 12-13). Specifically, the dataset includes the port explosion in Beirut, Lebanon and the barracks explosion in Bata, Equatorial Guinea. Unlike widespread natural disasters, these events are characterized by concentrated, radial blast damage, requiring precise localization capabilities. (3) Conflict Scenarios: Uniquely, DICQ extends the scope to armed conflict zones (IDs 14-15). Damage in this category differs fundamentally from natural calamities, exhibiting discrete, irregular, and targeted structural failures rather than contiguous destruction. The inclusion of conflict data fills a critical gap in the remote sensing literature, allowing the model to learn the complex signatures of war-torn infrastructure. To quantitatively assess the contribution of DICQ, we compare it with mainstream remote sensing image-text datasets in Table II. As illustrated, early datasets such as UCM-Captions and RSICD (IDs 1-2) primarily focus on single-image classification with brief descriptions (avg. 12 words). While recent change captioning benchmarks like LEVIR-CC (ID 8) and RSCC (ID 14) have introduced bi-temporal contexts, their textual annotations remain relatively concise (averaging 40 and 72 words, respectively), limiting their utility for complex reasoning tasks. In stark contrast, DICQ establishes a new standard for semantic density. Although XLRS-Bench (ID 13) offers detailed descriptions, it is constrained by a very small scale (only 934 samples). DICQ strikes a critical balance between scale and depth: it maintains a large volume of

4

Marshall-WF

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Pre-disaster

Post-disaster

DamageMask

Bata-EP

Pre-disaster

AC-1 AC-2 Fig. 2: Overview of the DICQ dataset for multi-disaster, multi-source change detection. The dataset covers different geographic locations and disaster types, including wildfire, explosion and armed conflict in Armed Conflict 1(AC-1) and Armed Conflict 2(AC-2) (from top to bottom). For each disaster scenario, the first column shows the pre-event optical image, the second column shows the post-event SAR image, and the third column displays the corresponding change mask used for supervised learning. The change mask encodes building-level damage states, where green denotes intact buildings, blue indicates damaged buildings, and red represents destroyed buildings. TABLE II: Comparison of DICQ with existing remote sensing image-text datasets. DICQ significantly outperforms prior benchmarks in terms of semantic density (Average Length), providing rich instruction following data for complex disaster analysis. ID

Dataset

Year

Image (Pixels)

Captions

Avg. Length

1 2 3 4 5 6 7 8 9 10 11 12 13 14

UCM-Captions RSICD fMoW SpaceNet 7 S2Looking Qfabric SpaceNet 8 LEVIR-CC Dubai-CCD RSICap VRSBench WHU-CDC XLRS-Bench RSCC

2016 2018 2018 2021 2021 2021 2022 2022 2022 2023 2024 2024 2025 2025

2,100 (1.0B) 10,921 (0.5B) 1M (437.0B) 2,389 (2.6B) 5,000 (5.0B) 2,520 (245.1B) 2,576 (3.0B) 20,154 (1.2B) 1,000 (≤ 0.1B) 2,585 (0.6B) 29,614 (7.8B) 14,868 (1.9B) 934 (67.5B) 124,702 (32.7B)

10,500 54,605 N/A N/A N/A N/A N/A 50,385 2,500 2,585 29,614 37,170 934 62,351

12 12 N/A N/A N/A N/A N/A 40 35 60 52 41 379 72

15

DICQ (Ours)

2026

136,672 (20B)

68,336

571

136,672 image pairs while achieving an unprecedented average caption length of 571 words. This figure is nearly 8 times that of RSCC and significantly surpasses all existing benchmarks.

This exceptional textual length is driven by our automated annotation engine, which generates not just a single caption, but a structured set of instructions including global summaries, zone-specific descriptions, and quantitative damage assessments. This rich semantic alignment is essential for training MLLMs to perform granular disaster analysis rather than simple image tagging. We perform a comprehensive statistical analysis to verify the representativeness and physical consistency of the dataset. Regarding the pixel-wise class balance, as illustrated in Fig. 3(a), the dataset exhibits a significant class imbalance inherent to real-world scenarios, where the “Intact” dominates (82.3%– 87.0%) over the minority “Destroyed” (8.1%–13.6%) and “Damaged” (∼5%) categories. Crucially, despite this skew, the distribution pattern remains highly stable across the training, validation, and testing splits. This consistency is vital, ensuring that the test set faithfully reflects the training distribution and provides a rigorous assessment of the model’s generalization capability on rare but high-stakes damage classes. Complementing this pixel-level view, the instance level scale and physics are detailed in Fig. 3(b), covering over 240,000 building instances. The counts follow a long tail distri-

5

bution (NIntact = 192, 774, NDes = 37, 095, NDam = 14, 790) with vast scale variations (102 –105 pixels). A key physical insight emerges from the size distribution: “Damaged” instances exhibit a larger median pixel area than “Destroyed” ones. This aligns with structural physics, where larger, reinforced architectures are more likely to sustain partial structural damage, whereas smaller structures are often completely obliterated in catastrophic events. This multi-scale characteristic requires the model to maintain robust recognition capabilities across varying object resolutions.

Keywords

Fig. 3: (a) Pixel-wise class balance across the training, validation, and testing splits of the DICQ dataset. The consistent distribution across splits ensures a fair evaluation, despite the natural class imbalance inherent in disaster scenarios. (b) Distribution of building object sizes (in pixel area) across damage categories on a logarithmic scale. The violin plots highlight the multi-scale nature of the DICQ dataset, covering structures from small residential units (∼ 102 pixels) to large industrial complexes (∼ 105 pixels). The instance counts (N ) further quantify the class imbalance inherent in disaster scenarios. (c) Conditional probability heatmap of class co-occurrence P (Column|Row). The low probabilities in the first row highlight the spatial isolation of intact zones, while the high values in the second row (54.3% for Destroyed given Damaged) validate the transitional nature of damage gradients, where partial damage frequently cooccurs with total destruction.

change assessed destruction conditions destroy right top central left bottom intact damage components pixels building 0

0.5

1

1.5

2

2.5

Frequency(Times/106)

To further quantify spatial context, we analyze the class co-occurrence probability P (Column | Row) in Fig. 3(c), revealing two distinct dependency patterns. First, “Intact” structures show strong spatial isolation, rarely co-occurring with “Damaged” (9.8%) or “Destroyed” (8.0%) buildings. This confirms that disaster damage is highly localized. Second, partial damage exhibits a transitional nature: “Damaged” buildings frequently co-occur with “Destroyed” (54.3%) and “Intact” (62.8%) structures. This reflects real world damage gradients where partially damaged areas serve as complex boundaries between epicenters and safe zones, posing significant challenges for fine-grained semantic disambiguation. Finally, we assess the linguistic diversity and spatial balance of the instructions via the word frequency analysis visualized in Fig. 4. The vocabulary is highly domain-specific, dominated by terms like “building” and “damage”, ensuring a strict focus on structural integrity. Crucially, the occurrence counts for spatial keywords: “top”, “bottom”, “left”, “right”, and “central” are nearly identical. This uniformity validates our partitioning strategy, ensuring the model receives unbiased supervision across all regions rather than overfitting to the center. Furthermore, the prevalence of fine-grained terms like “component” and “count” confirms that the instructions contain the precise quantitative details necessary for training rigorous counting and localization tasks.

Fig. 4: Linguistic analysis of the DICQ dataset. The word cloud highlights the domain-specific focus on building damage assessment. The bar chart reveals a balanced distribution of spatial terms (e.g., top, bottom, central), validating the unbiased nature of our zone based annotation pipeline.

B. Automated Semantic Annotation Pipeline To systematically bridge the gap between low-level pixel classification and high-level semantic reasoning, we developed an automated processing framework. Let the input segmentation mask be denoted as M ∈ {0, . . . , C}H×W , where C = 3 corresponds to the background, intact, damaged, and destroyed classes, indexed from 0 to 3. As illustrated in Fig. 5, the workflow transforms M into structured knowledge through a coherent methodology. 1) Automated Processing Framework: The pipeline operates in three sequential phases designed to extract, reason and compile disaster intelligence as in Fig. 6. Phase 1: Spatio-Visual Partitioning and PCA-based Localization. The initial phase focuses on geometric parsing. We employ a spatial partitioning function Fzone that maps the image domain Ω into five mutually exclusive semantic zones {Ztop , Zcen , Zbot , Zlef t , Zright }. For instance, the dominant

6

Phase 1: Image loading &Processing

Phase 2: Statistical&LLM Generation

Phase 3: Final Outputs Compilation

1. Mask Data Loading

4. Statistical Derivation

7. Damage Assessment

Operation: Import segmentation map (classes:0-3) Output: 256x256 Mask Tensor

Operation: Calculate pixel & component metrics per zone Output: Zone Statistics Objects.

5. LLM Zone Description

2. Multi-Zone Partitioning Operation: Apply 3-layer mutually exclusive layout Output: 5 Boolean Zone Masks

Operation: Generate detailed descriptive paragraphs Output: Per zone descriptive

3. Instance Extraction (Connected Component Analysis) Operation: Perform Connected Component Analysis Output: Building Instances & Zone IDs

6. LLM Summary Generation Operation: Generate concise summary sentences Output: Summary of whole area

Operation: Apply dual-threshold decision tree logic Output: Damage Level

8. Quantitative Captioning Operation: Convert statistics into natural language counting Output: Quantified Counting Descriptions

9. Final Report Compilation Operation: Aggregate all structured data & generated text Output: Comprehensive Structured Annotation

Fig. 5: Overview of the proposed Automated Semantic Annotation Pipeline. The workflow transforms raw segmentation masks into comprehensive textual reports through three sequential phases: (1) Image Processing, which performs spatial partitioning and instance extraction via Connected Component Analysis; (2) Statistical Analysis & LLM Generation, where pixel-level metrics are converted into zone specific descriptions using Large Language Models; and (3) Output Compilation, which aggregates logic driven damage assessments from Level 0 to Level 4 and quantitative summaries into a final structured annotation.

Zcen is rigorously defined as: Zcen = {(u, v) | 0.25H ≤ v < 0.75H, 0.2W ≤ u < 0.8W } . (1) To resolve ambiguities for building instances spanning these boundaries, we adopt a majority-overlap strategy to maintain object integrity. Specifically, an instance bi is assigned to the zone z ∗ that contains the maximum number of its constituent pixels N (bi , z): z∗ =

arg max

N (bi , z).

(2)

z∈{Ztop ,...,Zright }

Simultaneously, we perform Connected Component Analysis (CCA) to extract the set of building instances. For each instance, we apply Principal Component Analysis (PCA) to derive rotation-aware Oriented Bounding Boxes (OBB). By computing the Singular Value Decomposition (SVD) of the centered coordinates X̄i = UΣVT , we obtain the principal orientation θ = arctan(vy /vx ) from the first right singular vector. This formulation provides robust, orientationinvariant localization (cx, cy, w, h, θ), surpassing traditional axis-aligned methods for structures in chaotic debris fields. Phase 2: Statistical Derivation and Semantic Reasoning. In this phase, we transition from geometric primitives to semantic descriptors. For each zone z ∈ Z, we compute a dense statistics vector sz ∈ RC×2 , containing both pixel wise counts and instance wise counts for each damage class. These vectors serve as grounded prompts for the LLM. To mitigate hallucinations, we enforce a “statistics-first” constraint: the LLM is tasked to reason explicitly based on sz before generating the description. This ensures that the generated text is strictly grounded in physical evidence rather than generative priors. The text generation is modeled as a conditional probability distribution, where the LLM aims to maximize the likelihood of the description Tz given the statistics sz : Tz = arg max PLLM (T | Prompt(sz , Instruction)). T

(3)

Phase 3: Logic-Driven Damage Assessment. The final phase assigns a disaster severity level L ∈ [0, 4]. Given total building pixels Ntotal , we compute the destruction ratio ρdest = Ndestroyed /Ntotal and the damage ratio ρdam = (Ndamaged + Ndestroyed )/Ntotal . To ensure objective and consistent evaluation, the level L is determined by a dualthreshold rule:   4 (Destroyed) if ρdest ≥ 0.6 ∨ ρdam ≥ 0.85,     3 (Severe) if ρdest ≥ 0.3 ∨ ρdam ≥ 0.6,  L = 2 (Moderate) if ρdest ≥ 0.1 ∨ ρdam ≥ 0.35, (4)    1 (Minor) if Ndamaged > 0,    0 (No Damage) otherwise. This rule ensures that the evaluation is objective and consistent across the dataset, regardless of visual variations. 2) Structured Annotation and Supervision Signals: Fig. 7 visualizes a representative sample of the final annotation output, which serves as a dense, multimodal knowledge base. As depicted in the quantitative analysis panel, the system provides precise, instance level geometric primitives. The structure components table breaks down building counts by damage status across specific zones, offering fine grained statistical supervision that prevents the model from ignoring small or peripheral objects. Furthermore, the inclusion of YOLO-OBB annotations derived via PCA provides rotationaware coordinates, including center position, dimensions, and angle. This is particularly critical for disaster scenarios where collapsed structures often exhibit arbitrary orientations that horizontal bounding boxes fail to capture. These mathematical descriptors serve as anchors, ensuring that any subsequent textual generation is strictly grounded in physical reality. Complementing the numerical data, the semantic analysis panel translates these abstract metrics into coherent natural language. The zone descriptions provide spatially localized narratives, forcing the model to attend to specific image

7

Step 2: Spatial Analysis

Step 3: Output

TOP: The top zone predominantly features intact structures, with no visible damage or destruction, highlighting a stable built environment.

SUMMARY: An overview of the zones indicates the top and bottom areas are characterized by ...

Step 1: Input

L: The left zone is characterized by intact structures with no evidence of damage, showcasing a maintained infrastructure. CENTRAL: The central zone contains intact and one destroyed building, indicating a mix of stability and significant damage.

EVALUATION: The result is level 1 (Minor Damage) with 9.6 percent of mapped buildings ...

R: The right zone contains no mapped buildings and appears mostly background; no intact or damaged conditions can be inferred where structures are absent. BOTTOM: In the bottom zone, intact buildings dominate, reflecting a well-maintained area with no signs of damage or destruction.

Fig. 6: Visualization of the Spatio-Visual Partitioning strategy and sample outputs. To mimic the center-bias tendency of human visual attention, the image domain is strictly partitioned into a dominant Central zone flanked by marginal Top, Bottom, Left, and Right zones. The right panel demonstrates the generated hierarchical annotations, including fine-grained descriptions for each specific zone, a holistic summary, and a quantitative damage evaluation derived from the dual-threshold decision logic.

regions such as the top or central zones rather than generating generic captions. Additionally, the counting section explicitly converts the component matrices into natural language quantifiers, directly addressing the numeracy deficit often observed in generalist vision language models. The evaluation module further synthesizes the destruction ratios into a standardized severity rating, simulating the decision making process of a human analyst. By unifying spatial, geometric, and semantic attributes into a structured format, DICQ bridges the semantic gap. This representation enables the model to align visual features with structured logic rather than isolated tokens, providing a foundation for complex reasoning while mitigating hallucinations.

where Z ∈ RNv ×D represents the sequence of visual tokens. 2) Change-Aware Difference Module: Directly concatenating heterogeneous features often leads to suboptimal performance due to the significant domain gap between optical and SAR modalities. To address this, we introduce a ChangeAware Difference Module inserted between the encoder and the projector. This module operates in two steps: Difference Extraction and Feature Enhancement in Fig. 9. First, to explicitly model the semantic disparity, we employ a Cross-Attention mechanism. We treat the optical features as queries (Q) to attend to the SAR features (Keys K and Values V ). This aligns the post-event structural information with the pre-event semantic context, yielding an explicit difference representation Zdif f :

IV. T HE C HANGE Q UERY F RAMEWORK A. Overall Architecture The ChangeQuery architecture is built upon the LLaVA-1.5 paradigm [45], adapted to accommodate bi-temporal heterogeneous inputs. As illustrated in Fig. 8, the pipeline consists of four primary components: a dual stream visual encoder, a change-aware difference module, a modality aligning projector, and a large language model backbone. 1) Heterogeneous Visual Encoding: The input consists of a bi-temporal image pair: the pre-event optical image Xopt ∈ RH×W ×3 and the post-event SAR image Xsar ∈ RH×W ×C . We employ a pre-trained CLIP-ViT as the shared visual encoder Fvis . This encoder extracts high-level semantic features from both modalities independently, preserving the distinctive texture of optical data and the backscatter properties of SAR data. The extracted feature are denoted as: Zopt = Fvis (Xopt ),

Zsar = Fvis (Xsar ),

(5)

Zdif f = CrossAttn(Q = Zopt , K = Zsar , V = Zsar ) − Zopt . (6) Here, the subtraction operation highlights the features in the SAR imagery that deviate from the optical baseline, effectively isolating the change signal. Second, to ensure the visual tokens passed to the LLM are sensitive to these changes, we perform feature injection. We utilize Zdif f to enhance the original modality-specific features via a residual update mechanism. This generates the final change-aware representations Z̃opt and Z̃sar : Z̃opt = Zopt +α·FFN(Zdif f ),

Z̃sar = Zsar +β·FFN(Zdif f ), (7) where α and β are learnable gating factors and FFN is Feed-Forward Network. This step injects the extracted change information back into the contextual stream, ensuring that both modalities are aware of the damage patterns before projection.

8

Pre-disaster

Post-disaster

Damage Mask

Legend

Background (Black) Intact building (Green) Damage building (Blue)

Quantitative Anal.

Destroyed building (Red) • • • • • • •

Report

Zone Semantic Anal.

Loc. & Class

Damage Level & Pixel Counts: Level 1 (Minor Damage) Background (label 0): 51591 pixels Intact Building (label 1): 12609 pixels Destroyed Building (label 3): 1336 pixels Intact Building: 12 components, 12609 pixels total Damaged Building: 0 components, 0 pixels total Destroyed Building: 1 components, 1336 pixels total Minimum component size: 9 pixels

Components by Zone (Components/ Pixels) Zone

Intact

Damaged

Destroyed

Top

1 / 3091

0/0

0/0

Central

5 / 6839

0/0

1 / 1336

Bottom

3 / 1336

0/0

0/0

Left

3 / 1341

0/0

0/0

Right

0/0

0/0

0/0

• • • • •

Counting Section TOP: the top zone has 1 intact buildings. CENTRAL: the central zone has 5 intact buildings and 1 destroyed buildings. BOTTOM: the bottom zone has 3 intact buildings. LEFT: the left zone has 3 intact buildings. RIGHT: the right zone contains no mapped buildings; no building change can be assessed.

Technical Annotation Data (YOLO-OBB) (class cx cy w h angle deg; angle in (-90, 90]): 1 0.155335 0.189068 0.379289 0.245143 89.10; 1 0.496373 0.284849 0.340360 0.262960 -75.33; 1 0.019352 0.310143 0.139877 0.057863 82.29; 1 0.607662 0.309235 0.117089 0.086951 69.48; 1 0.314736 0.379275 0.128908 0.112809 77.56; 1 0.235688 0.537274 0.356718 0.118956 75.59; 1 0.351391 0.513202 0.152848 0.131489 70.02; 1 0.047569 0.511492 0.154888 0.107564 -35.46; 1 0.019858 0.681970 0.126865 0.063205 79.79; 1 0.143530 0.751549 0.108852 0.046402 67.00; 1 0.088559 0.787109 0.027344 0.017119 -90.00; 1 0.041329 0.896608 0.215417 0.135558 -68.57; 3 0.535580 0.580289 0.167841 0.133588 -27.26;

• • • • •

Zone Descriptions TOP: The top zone predominantly features intact structures, with no visible damage or destruction, highlighting a stable built environment. CENTRAL: The central zone contains intact and one destroyed building, indicating a mix of stability and significant damage. BOTTOM: In the bottom zone, intact buildings dominate, reflecting a well-maintained area with no signs of damage or destruction. LEFT: The left zone is characterixed by intact structures with no evidence of damage, showcasing a maintained infrastructure. RIGHT: the right zone contains no mapped buildings and appears mostly background; no intact or damaged conditions can be inferred where structures are absent.

Summary Sentences (5 variant sentences with the same meaning) 1. The image reveals a predominantly intact environment with zones exhibiting various levels of building conditions: the top and bottom zones are intact, central shows significant damage with one destroyed structure, while the left is intact, and the right lacks mapped buildings. 2. An overview of the zones indicates the top and bottom areas are characterized by intact buildings, with the central zone featuring significant damage due to one destroyed building, whereas the left remains intact, and the right zone shows no mapped buildings. …… Evaluation The result is level 1 (Minor Damage) with 9.6 percent of mapped buildings showing damage and 9.6 percent destroyed; the summary remains: The image reveals a predominantly intact environment with zones exhibiting various levels of building conditions: the top and bottom zones are intact, central shows significant damage with one destroyed structure, while the left is intact, and the right lacks mapped buildings.

Fig. 7: Example of structured annotations from the DICQ dataset generated by our automated pipeline, comprising quantitative statistics including pixel counts and YOLO-OBB coordinates, alongside semantic descriptions with logic-based damage levels, enabling grounded and hallucination-free disaster analysis.

3) Visual Language Projection: To bridge the gap between frozen visual features and the language embedding space, we utilize a Multi Layer Perceptron (MLP) as the trainable connector P. Unlike single-image architectures, our projector handles the concatenated sequence of the bi-temporal features. We fuse the optical and SAR features along the sequence dimension to form a unified visual context. Hvis = P([Z̃opt ; Z̃sar ]),

(8)

where [; ] denotes concatenation and Hvis ∈ R2Nv ×Dllm represents the projected visual tokens aligned with the LLM’s word embedding space. This early fusion strategy allows

the model to implicitly learn the correlation and differences between the two time steps via the self attention of the LLM.

4) LLM Reasoning Backbone: For the reasoning core, we adopt Vicuna-7B v1.5, a decoder-only Transformer finetuned for instruction-following tasks. The model processes a multimodal sequence composed of visual tokens Hvis and textual instruction tokens Hinstr derived from the user query (e.g., “Assess the damage in the central zone”). The response generation is formulated as an autoregressive process. Specifically, the probability of generating the target sequence Y is

9

Train Strategy

Generated Answer ChangeQuery Answer 1: “Across the image, the top zone shows an intact building…” ChangeQuery Answer 2: “The buildings in the left have been demolished…”

LLM

Large Language Model ...

...

Tokenizer & Embedding

Visual Projector 2

Text Prompt

+

+ Attention Block

Projector 🔥

...

Visual Projector 1

Question 1: “Provide a scene summary?” Question 2: “Describe conditions in left?”

Attention Block

Attention Block

Attention Block

Difference-Aware Interaction Module

Vison Encoder 1

Difference Module 🔥 Vision Encoder

DICQ Dataset

Attention Block

Vison Encoder 2

Stage 𝟮 LLM

Difference Module 🔥

🔥

Post-disaster

🔥

Projector 🔥

Vision Encoder

Pre-disaster

Vision Encoder

Attention Block

Attention Block

Attention Block

Stage 𝟭

🔥

Vision Encoder

DICQ Dataset

Fig. 8: The overall architecture of ChangeQuery. The model takes a pre-event optical image and a post-event SAR image as inputs. Features are extracted via a shared CLIP-ViT encoder, concatenated, and projected into the language space. The Vicuna-7B LLM then processes these visual tokens alongside the text instruction to generate grounded responses. The right part is to show our train strategy of changequery model.

defined as: p(Y | Xopt , Xsar , Xinstr ) =

L Y

p(yt | Hvis , Hinstr , y<t ).

t=1

(9) During training, the visual encoder is kept frozen to preserve its general representation capability. In contrast, the MLP projector and the LLM backbone are fine-tuned using LoRA, enabling efficient adaptation to the disaster domain. B. Two-Stage Training Strategy Training a Multimodal Large Language Model on heterogeneous remote sensing data presents significant optimization challenges, primarily due to the distinct feature distributions of optical and SAR modalities and the risk of catastrophic forgetting in the pre-trained backbone. To ensure stable convergence and robust instruction following, we adopt a progressive training paradigm, as illustrated in Fig. 8, which decouples feature alignment from semantic reasoning. 1) Stage 1: Modality Alignment and Feature Initialization: In the initial phase, the primary objective is to bridge the semantic gap between the visual encoders and the LLM while

initializing the change detection capability. During this stage, we keep the massive parameters of both the dual stream Vision Encoders and the LLM Backbone frozen. Optimization is restricted exclusively to the Difference Module and the MLP Projector. This strategy is motivated by the need to construct a stable visual language interface before burdening the model with complex reasoning tasks. Since the Difference Module is initialized from scratch, this focused training allows it to learn the fundamental mapping correlations between optical texture and SAR backscatter without perturbing the generalized knowledge inherent in the pre-trained backbones. Effectively, this stage forces the connector layers to act as a translator, converting raw heterogeneous feature disparities into coherent linguistic embeddings. 2) Stage 2: End-to-End Instruction Tuning: Subsequently, the training transitions to a full scale instruction tuning phase designed to master the specific syntax of disaster reporting. In this stage, we unfreeze the Vision Encoders to allow for domain adaptation, enabling the CLIP model to adjust to the nadir perspective of remote sensing imagery. Simultaneously, the LLM Backbone is finetuned using LoRA. Crucially, this stage is driven by the structured multi-turn conversations

10

Projector (MLP) Output

Concat

~ Enhanced Optical (Z opt )

~

Enhanced SAR (Z sar )

Feature Enhancement

+

Feed-Forward Network (FFN)

+

Zdif f (Difference Representaiton)

Difference Extaction

Cross-Attention Module Q

Input

Optical Feature (Zopt ) ​

K

V

SAR Feature (Zsar ) ​

Fig. 9: Structure of the Change-Aware Difference Module. It utilizes cross-attention to highlight semantic disparities between pre-event optical and post-event SAR features.

generated by our pipeline. The training targets are not merely static captions but interactive dialogues requiring hierarchical reasoning. The model is explicitly supervised to perform tasks ranging from global damage grading (e.g., “Level 0”) to fine-grained regional descriptions. This rigorous supervision ensures the model learns to ground its textual outputs in the specific quantitative evidence extracted from the heterogeneous visual inputs, effectively minimizing hallucinations. V. E XPERIMENTS AND A NALYSIS A. Experimental Settings Baselines. To rigorously evaluate the performance of ChangeQuery, we benchmark it against two categories of stateof-the-art vision-language models: (1) Generalist MLLMs: We compare with leading open-source multimodal models that support multi-image or interleaved inputs, including LLaVA-NeXT-Interleave (8B) [53], LLaVA-OneVision (8B) [54], and InternVL 3 (8B) [55]. These models represent the current state-of-the-art in general visual reasoning capabilities. (2) Specialized Remote Sensing Models: We also include domain-specific change captioning models, specifically CCExpert [56] and TEOChat [57], to assess the advantages of our proposed heterogeneous fusion strategy against existing remote sensing solutions. Implementation Details. The ChangeQuery model is implemented using the PyTorch framework. For the visual encoder, we utilize the pre-trained CLIP-ViT, keeping its weights frozen to preserve generalized feature extraction. The MLP projector and the Vicuna-7B v1.5 backbone are fine-tuned using LoRA [58] to efficiently adapt to the disaster assessment domain. We utilize the AdamW optimizer with a cosine learning rate scheduler. To handle the heterogeneous inputs, optical and SAR images are resized to 252×252 pixels before encoding. B. Evaluation Metrics Evaluating the quality of generated disaster reports requires assessing both the textual overlap with ground truth and the

semantic accuracy of the content. We employ two sets of metrics: (1) N-Gram Overlap Metrics. Following standard practices in image captioning, we report ROUGE-L [59] and METEOR [60]. These metrics measure the precision and recall of ngrams between the generated response and the reference text. While they provide a basic assessment of fluency and lexical overlap, they are often insufficient for capturing the complex reasoning and structural details required in long form disaster analysis. (2) Semantic Similarity Metric (ST5-SCS). To address the limitations of n-gram metrics in capturing semantic alignment for long-form descriptions, we adopt Sharpened Cosine Similarity (ST5-SCS). Utilizing the Sentence T5 encoder [61] for embedding extraction, the score is computed as: SCS(u, v) = Cosine(u, v)3 · sgn(Cosine(u, v)),

(10)

where u and v represent the embeddings. Quantitative results in Table III reveal distinct performance gaps among the evaluated methods. Generalist Models, led by InternVL 3 (8B), achieve a respectable semantic consistency (50.46% SCS) due to their robust pre-trained backbones. However, they suffer from suboptimal N-gram scores (e.g., ROUGE ≈ 10%), indicating a failure to grasp the specific lexicon of disaster reporting and the complex backscatter characteristics of SAR imagery. While Remote Sensing(RS) specialized models like CCExpert improve upon this by adapting to aerial views, achieving a competitive SCS of 50.68%, they still lack explicit mechanisms to model the structural disparities between optical and SAR modalities. Consequently, they struggle to generate the precise, fine-grained details required by the DICQ benchmark, limiting their fluency and factual accuracy. In contrast, ChangeQuery significantly outperforms all baselines, establishing a new SOTA. Most notably, it achieves a METEOR score of 22.00%, surpassing the second-best method (RSCC) by a substantial margin of over 8 points. Since METEOR correlates strongly with human judgment on harmonic mean of precision and recall, this leap indicates that our model generates reports with far superior coherence and logical structure. Furthermore, the highest SCS of 54.07% confirms that ChangeQuery effectively bridges the modality gap. We attribute these gains to the Change-Aware Difference Module, which aligns heterogeneous features, and Two-Stage Instruction Tuning, which strictly grounds the textual generation in physical evidence rather than generic hallucinations.

C. Qualitative Visualization & Analysis Fig. 10 presents a visual qualitative comparison between ChangeQuery and state-of-the-art baselines, including CCExpert and RSCC, across diverse disaster scenarios. To facilitate analysis, descriptions aligned with the ground truth are highlighted in green, whereas severe hallucinations or factual errors are marked with strikethrough text, and vague or inaccurate descriptions are highlighted in gray. Three critical observations can be drawn from these results:

11

TABLE III: Quantitative comparison with state-of-the-art methods on the DICQ benchmark. We categorize baselines into Generalist MLLMs (trained on natural images) and Remote Sensing (RS) Specific Models (fine-tuned on aerial data). Evaluation metrics include N-gram overlap (ROUGE, METEOR) for textual fluency and Sharpened Cosine Similarity (SCS) for semantic consistency. The best results are highlighted in bold, demonstrating that ChangeQuery achieves superior performance across all dimensions. Method

Params

Base Model

N-Gram

Reference

Contextual Similarity

ROUGE (%) ↑

METEOR (%) ↑

ST5-SCS (%) ↑

General Model LLaVA-NeXT-Interleave LLaVA-OneVision InternVL 3

8B 8B 8B

Qwen 1.5 Qwen 3 Qwen 2.5

arXiv’24 [53] arXiv’24 [54] arXiv’25 [55]

10.13 9.71 10.62

14.08 13.30 12.19

45.11 46.07 50.46

CCExpert TEOChat RSCC

7B 7B 7B

Qwen 2 Llama 2 Qwen 2.5

arXiv’25 [56] ICLR’25 [57] NeurIPS’25 [62]

10.55 8.16 10.67

13.00 9.21 13.71

50.68 46.22 47.21

ChangeQuery

7B

Vicuna 1.5

Ours

12.91

22.00

54.07

RS Specific Model

Descriptions aligned with GT

Vague content with GT

strikethrough text

Severe hallucinations or factual errors

Question: Describe the building conditions in bottom zone? GT: The bottom zone is characterized by intact buildings with no visible damage or destruction, reflecting a stable structural condition. CCExpert: There is a noticeable diffuse scattering around the smaller structures in the bottom-center. This can sometimes indicate debris or significant surface roughness changes, potentially pointing to partial damage or collapsed outbuildings. RSCC: Based on a comparative analysis of the provided SAR (Synthetic Aperture Radar) and optical images, here is the description of the building conditions in the bottom zone: The buildings in the bottom zone are collapsed. Ours: Across the image, the bottom zone is characterized by intact buildings with no visible damage or destruction, reflecting a stable structural condition.

Question: Describe the building conditions in top zone? GT: The top zone predominantly features destroyed structures, with no intact or damaged buildings present, indicating widespread devastation. CCExpert: Unlike the central areas where the radar return is more diffuse, the top zone maintains sharp geometric reflections. This suggests that the primary buildings in this section have likely avoided catastrophic collapse. RSCC: Based on a comparative analysis of the provided SAR (Synthetic Aperture Radar) and optical images, here is the description of the building conditions in the top zone: The buildings in the top zone are intact. Ours: The top zone predominantly features destroyed structures, with no intact or damaged buildings present, indicating widespread devastation.

Question: Describe the building conditions in left zone? GT: The left zone is mostly background with a few intact building structures, showing no signs of damage or destruction. CCExpert: While there may be non-structural damage (such as roof or window damage) not easily detectable by SAR, the buildings in the left zone do not show signs of total collapse or significant displacement compared to more heavily impacted areas. RSCC: Based on a comparative analysis of the provided SAR (Synthetic Aperture Radar) and optical images, here is the description of the building conditions in the left zone: The buildings in the left zone are intact. Ours: The left zone is mostly background with a few intact buildings showing no signs of damage or destruction.

Question: Assess the post-disaster damage level? GT: The result is level 0 (No Damage) with 0.0 percent of mapped buildings showing damage and 0.0 percent destroyed; the summary remains: The entire image presents an area with exclusively intact building structures across all zones—top, central, bottom, left, and right—demonstrating no evidence of damage or destruction and thus maintaining complete structural integrity throughout. CCExpert: The area requires immediate emergency intervention as the structural integrity of nearly all visible assets has been compromised, shifting from defined man-made features to a randomized distribution of rubble. RSCC: The affected area shows no severe or large-scale destruction. Observed changes are consistent with light to moderate damage, possibly due to partial building damage, surrounding vegetation disturbance, or sensor/viewing differences rather than total structural loss. Ours: The result is no damage level, the image shows exclusively intact buildings, with no evidence of damage or destruction, maintaining complete structural integrity throughout.

Question: Provide a scene summary? GT: The image illustrates that the top zone has one intact building with no damage or destruction, the central zone shows a mix of structures with intact and damaged buildings but none destroyed, while the bottom zone features primarily intact buildings with minimal damage, and the left and right zones lack any mapped buildings. CCExpert: The scene reflects a shift from a functional built environment to a randomized debris field, highly consistent with the impact of a major earthquake or a high-intensity tornado. RSCC: The scene shows a small, compact rural settlement surrounded by open agricultural land. Compared with the surrounding fields, the built-up area exhibits noticeable structural irregularities in the optical image, including disrupted building shapes and uneven roof patterns, indicating post-event alteration of the settlement layout. Ours: Across the image, the top zone shows an intact building with no observed damage or destruction, the central zone contains both intact and damaged buildings but no destroyed structures, the bottom zone is largely intact with minimal damage; and no buildings are present in either the left or right zones.

Question: Provide a scene summary? GT:The image reveals extensive destruction across all zones, with the top and right zones showing predominantly destroyed buildings, while the central zone experiences complete destruction, the bottom zone suffers widespread devastation, and the left zone indicates moderate to severe damage with a prevalence of damaged structures. CCExpert: The scene reflects a total leveling event. The transition from a functional human settlement to a homogenized debris field is characteristic of extreme kinetic energy, most consistent with a major earthquake or a direct strike by a high-intensity tornado. RSCC: Overall, the pair indicates a significant change event, likely involving building destruction, debris, flooding effects, or ground disruption, causing strong radar backscatter variation across the urban zone. Ours: Analyzing the image, the top and right zones are largely dominated by destroyed buildings, the central zone is completely devastated, the bottom zone shows widespread structural damage, and the left zone exhibits moderate to severe damage, primarily consisting of damaged structures.

Fig. 10: Qualitative comparison of diverse instruction-following tasks on the DICQ benchmark. We compare the responses generated by ChangeQuery against state-of-the-art baselines including CCExpert and RSCC. To facilitate visualization, accurate descriptions aligned with the Ground Truth (GT) are highlighted in green, while severe hallucinations or factual errors are marked with strikethrough text, and vague content is highlighted in gray. The examples illustrate that ChangeQuery generates more grounded and spatially precise analysis, successfully avoiding the common issues of disaster bias and crossmodal misalignment observed in baseline methods.

12

(1) Mitigation of “Disaster Bias” and Hallucinations. A prevalent issue in existing remote sensing VLMs is the tendency to hallucinate damage even in intact areas, a phenomenon we term “disaster bias”. This is clearly observable in the bottom-left case which represents a Level 0 scenario with no damage. Both CCExpert and RSCC misinterpret the scene, fabricating descriptions of “immediate emergency intervention” or “light to moderate damage” despite the buildings being structurally sound. This error likely stems from their inability to distinguish SAR speckle noise from actual debris. In contrast, ChangeQuery correctly identifies the scene with “exclusively intact buildings”, demonstrating superior robustness against false positives in complex SAR environments. (2) Accurate Cross Modal Alignment for Destruction. The top middle example focusing on the top zone illustrates a scenario of severe destruction. While the ground truth indicates “widespread devastation”, CCExpert erroneously claims the zone maintains sharp geometric reflections and avoided catastrophic collapse, while RSCC incorrectly labels it as intact. These failures highlight the struggle of baseline models to align pre-event optical semantics with post-event SAR scattering patterns. Empowered by the proposed Difference Module, ChangeQuery accurately captures the loss of structural coherence in the SAR imagery and correctly describes the zone as predominantly featuring destroyed structures. (3) Precise Spatial Reasoning and Granularity. The scene summary tasks presented in the bottom-middle and bottomright panels test the ability of the models to perform zone level spatial reasoning. The baselines often generate generic, holistic descriptions such as CCExpert’s generalized claim of a “randomized debris field” or hallucinate specific disaster types like tornadoes or flooding without visual evidence. Conversely, ChangeQuery exhibits precise spatial grounding. As demonstrated in the bottom-middle instance, it successfully disentangles complex mixed scenarios by accurately identifying intact and damaged buildings in the central zone while noting minimal damage in the bottom zone. This fine-grained descriptive capability validates the effectiveness of our zoneaware instruction tuning strategy. In summary, ChangeQuery not only generates more accurate damage assessments but also significantly reduces the hallucination rate compared to specialized baselines, providing reliable and grounded intelligence for disaster response. VI. C ONCLUSION In this paper, we introduced ChangeQuery, a unified multimodal framework that redefines disaster damage assessment by shifting the paradigm from low level pixel classification to high-level semantic reasoning. Addressing the critical limitations of existing methodologies, we constructed the DICQ dataset. This large scale benchmark uniquely couples pre-event optical semantics with post-event SAR structural features, covering a diverse spectrum of scenarios ranging from natural catastrophes to anthropogenic conflicts. By employing a novel automated semantic annotation pipeline, we transformed raw segmentation masks into grounded, hierarchical instruction sets, enabling the model to learn spatial reasoning and precise quantification without human intervention.

Methodologically, the ChangeQuery framework incorporates a Change-Aware Difference Module and utilizes a progressive two-stage training strategy to effectively bridge the significant domain gap between heterogeneous modalities. Extensive experiments demonstrate that our approach achieves SOTA performance across multiple metrics, significantly outperforming both generalist vision language models and specialized remote sensing baselines. The model exhibits remarkable capabilities in generating spatially precise logic-driven disaster reports while successfully mitigating common issues such as disaster bias and hallucinations. Ultimately, this work narrows the semantic gap between raw remote sensing data and actionable decision support, providing a robust foundation for the next generation of all weather, interactive humanitarian assistance systems. R EFERENCES [1] G. Gidófalvi, “Review of big data and processing frameworks for disaster response applications,” ISPRS Int. J. Geo-Inf., vol. 8, p. 387, 2019. [2] H. Mueller, A. Groeger, J. Hersh, A. Matranga, and J. Serrat, “Monitoring war destruction from space using machine learning,” Proc. Natl. Acad. Sci., vol. 118, no. 23, 2021. [3] S. Holail, T. Saleh, X. Xiao, J. Xiao, G.-S. Xia, Z. Shao, M. Wang, J. Gong, and D. Li, “Time-series satellite remote sensing reveals gradually increasing war damage in the gaza strip,” Natl. Sci. Rev., vol. 11, no. 9, p. 304, 2024. [4] S. Al Shafian and D. Hu, “Integrating machine learning and remote sensing in disaster management: A decadal review of post-disaster building damage assessment,” Buildings, vol. 14, no. 8, p. 2344, 2024. [5] Z. Liu, J. Zhang, W. Wang, and Y. Gu, “M2cd: A unified multimodal framework for optical-sar change detection with mixture of experts and self-distillation,” IEEE Geosci. Remote Sens. Lett., vol. 22, pp. 1–5, 2025. [6] H. Chen, J. Song, C. Han, J. Xia, and N. Yokoya, “Changemamba: Remote sensing change detection with spatiotemporal state space model,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–20, 2024. [7] Y. Feng, J. Jiang, H. Xu, and J. Zheng, “Change detection on remote sensing images using dual-branch multilevel intertemporal network,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–15, 2023. [8] L. Ding, K. Zhu, D. Peng, H. Tang, K. Yang, and L. Bruzzone, “Adapting segment anything model for change detection in vhr remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–11, 2024. [9] G. Hoxha, S. Chouaf, F. Melgani, and Y. Smara, “Change captioning: A new paradigm for multitemporal remote sensing image analysis,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–14, 2022. [10] S. Chang and P. Ghamisi, “Changes to captions: An attentive network for remote sensing change captioning,” IEEE Trans. Image Process., vol. 32, pp. 6047–6060, 2023. [11] C. Liu, J. Zhang, K. Chen, M. Wang, Z. Zou, and Z. Shi, “Remote sensing spatiotemporal vision–language models: A comprehensive survey,” IEEE Geosci. Remote Sens. Mag., pp. 2–42, 2025. [12] H. Guo, X. Su, C. Wu, B. Du, L. Zhang, and D. Li, “Remote sensing chatgpt: Solving remote sensing tasks with chatgpt and visual models,” in IEEE Int. Geosci. Remote Sens. Symp., 2024, pp. 11 474–11 478. [13] X. Guo, J. Lao, B. Dang, Y. Zhang, L. Yu, L. Ru, L. Zhong, Z. Huang, K. Wu, D. Hu, H. He, J. Wang, J. Chen, M. Yang, Y. Zhang, and Y. Li, “Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery,” in Conf. Comput. Vis. Pattern Recognit., 2024, pp. 27 662–27 673. [14] J. Lin, L. Liu, D. Lu, and K. Jia, “Sam-6d: Segment anything model meets zero-shot 6d object pose estimation,” in Conf. Comput. Vis. Pattern Recognit., 2024, pp. 27 906–27 916. [15] W. Zhang, M. Cai, T. Zhang, Y. Zhuang, and X. Mao, “Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,” IEEE Trans. Geosci. Remote Sens., 2024. [16] S. Dong, L. Wang, B. Du, and X. Meng, “Changeclip: Remote sensing change detection with multimodal vision-language representation learning,” ISPRS J. Photogramm. Remote Sens., vol. 208, pp. 53–69, 2024.

13

[17] E. Weber and H. Kané, “Building disaster damage assessment in satellite imagery with multi-temporal fusion,” arXiv:2004.05525, 2020. [18] J. Wang, W. Xuan, H. Qi, Z. Liu, K. Liu, Y. Wu, H. Chen, J. Song, J. Xia, Z. Zheng, and N. Yokoya, “Disasterm3: A remote sensing visionlanguage dataset for disaster damage assessment and response,” Adv. Neural Inf. Process. Syst., 2025. [19] O. L. Stephenson, T. Köhne, E. Zhan, B. E. Cahill, S.-H. Yun, Z. E. Ross, and M. Simons, “Deep learning-based damage mapping with insar coherence time series,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–17, 2022. [20] H. Chen, J. Song, O. Dietrich, C. Broni-Bediako, W. Xuan, J. Wang, X. Shao, Y. Wei, J. Xia, C. Lan, K. Schindler, and N. Yokoya, “B RIGHT: a globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response,” Earth Syst. Sci. Data, vol. 17, no. 11, pp. 6217–6253, 2025. [21] I. A. Listiani, M. Zanetti, and F. Bovolo, “Time series change vector analysis for semisupervised abrupt land cover change detection,” IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–15, 2025. [22] F. Bovolo and L. Bruzzone, “A theoretical framework for unsupervised change detection based on change vector analysis in the polar domain,” IEEE Trans. Geosci. Remote Sens., vol. 45, no. 1, pp. 218–236, 2007. [23] R. Caye Daudt, B. Le Saux, and A. Boulch, “Fully convolutional siamese networks for change detection,” in IEEE Int. Conf. Image Process., 2018, pp. 4063–4067. [24] C. Zhang, P. Yue, D. Tapete, L. Jiang, B. Shangguan, L. Huang, and G. Liu, “A deeply supervised image fusion network for change detection in high resolution bi-temporal remote sensing images,” ISPRS J. Photogramm. Remote Sens., vol. 166, pp. 183–200, 2020. [25] H. Chen, Z. Qi, and Z. Shi, “Remote sensing image change detection with transformers,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–14, 2022. [26] W. G. C. Bandara and V. M. Patel, “A transformer-based siamese network for change detection,” arXiv:2201.01293, 2022. [27] P. Liu, J. Peng, H. Wang, D. Hong, and X. Cao, “Infrared small target detection via joint low rankness and local smoothness prior,” IEEE Trans. Geosci. Remote Sens., 2024. [28] B. Zhang, X. Cao, S. Wang, and D. Meng, “Class-incremental learning for remote sensing scene classification via stable diffusion based data regeneration,” IEEE Trans. Geosci. Remote Sens., pp. 1–1, 2026. [29] P. Liu, L. Pang, J. Peng, Y. Luo, J. Liu, and X. Cao, “Ctvnet: Gradient prior-guided deep unfolding network for infrared small target detection,” IEEE Trans. Geosci. Remote Sens., vol. 63, pp. 1–14, 2025. [30] D. Brunner, G. Lemoine, and L. Bruzzone, “Earthquake damage assessment of buildings using vhr optical and sar imagery,” IEEE Trans. Geosci. Remote Sens., vol. 48, no. 5, pp. 2403–2420, 2010. [31] X. Niu, M. Gong, T. Zhan, and Y. Yang, “A conditional adversarial network for change detection in heterogeneous images,” IEEE Geosci. Remote Sens. Lett., vol. 16, no. 1, pp. 45–49, 2019. [32] X. Li, Z. Du, Y. Huang, and Z. Tan, “A deep translation (gan) based change detection network for optical and sar remote sensing images,” ISPRS J. Photogramm. Remote Sens., vol. 179, pp. 14–34, 2021. [33] Y. Sun, L. Lei, X. Li, X. Tan, and G. Kuang, “Structure consistencybased graph for unsupervised change detection with homogeneous and heterogeneous remote sensing images,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–21, 2022. [34] H. Chen, N. Yokoya, C. Wu, and B. Du, “Unsupervised multimodal change detection based on structural relationship graph representation learning,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–18, 2022. [35] B. Adriano, N. Yokoya, J. Xia, H. Miura, W. Liu, M. Matsuoka, and S. Koshimura, “Learning from multimodal and multitemporal earth observation data for building damage mapping,” ISPRS J. Photogramm. Remote Sens., vol. 175, pp. 132–143, 2021. [36] X. Zeng and Y. Qu, “Building damage mapping through heterogeneous feature consistency and knowledge integration,” in IEEE Int. Geosci. Remote Sens. Symp., 2025, pp. 233–236. [37] Z. Zheng, Y. Zhong, J. Wang, A. Ma, and L. Zhang, “Building damage assessment for rapid disaster response with a deep object-based semantic change detection framework: From natural disasters to manmade disasters,” Remote Sens. Environ., vol. 265, p. 112636, 2021. [38] A. Toker, L. Kondmann, M. Weber, M. Eisenberger, A. Camero, J. Hu, A. P. Hoderlein, c. Şenaras, T. Davis, D. Cremers, G. Marchisio, X. X. Zhu, and L. Leal-Taixé, “Dynamicearthnet: Daily multi-spectral satellite dataset for semantic change segmentation,” in IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 21 158–21 167. [39] X. Lu, B. Wang, X. Zheng, and X. Li, “Exploring models and data for remote sensing image caption generation,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 4, pp. 2183–2195, 2018.

[40] C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–20, 2022. [41] S. Chouaf, G. Hoxha, Y. Smara, and F. Melgani, “Captioning changes in bi-temporal remote sensing images,” in IEEE Int. Geosci. Remote Sens. Symp., 2021, pp. 2891–2894. [42] D. Sun, Y. Bao, J. Liu, and X. Cao, “A lightweight sparse focus transformer for remote sensing image change captioning,” IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens., vol. 17, pp. 18 727–18 738, 2024. [43] D. Sun, J. Yao, W. Xue, C. Zhou, P. Ghamisi, and X. Cao, “Mask approximation net: A novel diffusion model approach for remote sensing change captioning,” IEEE Trans. Geosci. Remote Sens., 2025. [44] D. Sun, Y. Wang, J. Yao, W. Yu, X. Cao, and P. Ghamisi, “Scnet: Lightweight spatial-channel attention network for remote sensing change captioning,” IEEE Trans. Geosci. Remote Sens., 2026. [45] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Adv. Neural Inf. Process. Syst., vol. 36, 2023, pp. 34 892–34 916. [46] D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” arXiv:2304.10592, 2023. [47] F. Liu, D. Chen, Z. Guan, X. Zhou, J. Zhu, Q. Ye, L. Fu, and J. Zhou, “Remoteclip: A vision language foundation model for remote sensing,” arXiv:2306.11029, 2024. [48] J. Wang, Z. Zheng, Z. Chen, A. Ma, and Y. Zhong, “Earthvqa: towards queryable earth via relational reasoning-based remote sensing visual question answering,” Proc. AAAI Conf. Artif. Intell., p. 609, 2024. [49] Y. Hu, J. Yuan, C. Wen, X. Lu, and X. Li, “Rsgpt: A remote sensing vision language model and benchmark,” arXiv:2307.15266, 2023. [50] Y. Huang, X. Li, Z. Du, and H. Shen, “Spatiotemporal enhancement and interlevel fusion network for remote sensing images change detection,” IEEE Trans. Geosci. Remote Sens., vol. 62, pp. 1–14, 2024. [51] G. Mai, W. Huang, J. Sun, S. Song, D. Mishra, N. Liu, S. Gao, T. Liu, G. Cong, Y. Hu, C. Cundy, Z. Li, R. Zhu, and N. Lao, “On the opportunities and challenges of foundation models for geospatial artificial intelligence,” arXiv:2304.06798, 2023. [52] Y. Zhan, Z. Xiong, and Y. Yuan, “Rsvg: Exploring data and models for visual grounding on remote sensing data,” IEEE Trans. Geosci. Remote Sens., vol. 61, pp. 1–13, 2023. [53] F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li, “Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models,” arXiv:2407.07895, 2024. [54] B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y. Li, Z. Liu, and C. Li, “Llava-onevision: Easy visual task transfer,” arXiv:2408.03326, 2024. [55] J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao et al., “Internvl3: Exploring advanced training and testtime recipes for open-source multimodal models,” arXiv:2504.10479, 2025. [56] Z. Wang, M. Wang, S. Xu, Y. Li, and B. Zhang, “Ccexpert: Advancing mllm capability in remote sensing change captioning with differenceaware integration and a foundational dataset,” arXiv:2411.11360, 2024. [57] J. A. Irvin, E. R. Liu, J. C. Chen, I. Dormoy, J. Kim, S. Khanna, Z. Zheng, and S. Ermon, “Teochat: A large vision-language assistant for temporal earth observation data,” in Int. Conf. Learn. Represent., 2025. [58] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” Int. Conf. Learn. Represent., vol. 1, no. 2, p. 3, 2022. [59] C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summ. Branches Out, 2004, pp. 74–81. [60] S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in Proc. ACL Workshop Intrins. Extrins. Eval. Meas. Mach. Transl. Summ., 2005, pp. 65–72. [61] J. Ni, G. Hernandez Abrego, N. Constant, J. Ma, K. Hall, D. Cer, and Y. Yang, “Sentence-t5: Scalable sentence encoders from pre-trained textto-text models,” in Findings Assoc. Comput. Linguist., 2022, pp. 1864– 1874. [62] Z. Chen, C. Wang, N. Zhang, and F. Zhang, “Rscc: A large-scale remote sensing change caption dataset for disaster events,” in Adv. Neural Inf. Process. Syst., 2025.

Record · ID 134594 · SHA-256 49f31c8a1003b274
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.