FusionRS: A Large-Scale RGB–Infrared Remote Sensing Dataset for Dual-Modal Vision–Language Foundation Models
arXiv:2606.17020v1 [cs.CV] 15 Jun 2026
Jiaju Han1 Ben Zhang2 Xuemeng Sun1 Qike Zhang1 Yuxian Dong1 Chengyin Hu1 Fengyu Zhang1 Yiwei Wei3 Jiujiang Guo3 1 China University of Petroleum-Beijing at Karamay 2 University of Electronic Science and Technology of China 3 Tianjin University
1
Abstract
Introduction
Remote sensing imagery (Cheng et al., 2017) is a fundamental source of information for Earth observation, supporting applications such as land-cover mapping, urban monitoring, disaster assessment, agricultural observation, and infrastructure inspection. Recent progress in vision-language learning has extended remote sensing models from fixedcategory recognition to open-ended understanding, including image-text retrieval, caption generation, visual question answering, and instruction following (Liu et al., 2024a; Zhang et al., 2024; Wang et al., 2024b; Kuckreja et al., 2024; Bazi et al., 2024; Wang et al., 2024a). CLIP-style models (Radford et al., 2021; Liu et al., 2024a; Zhang et al., 2024) have shown strong potential for remote sensing image-text alignment, while generative visionlanguage models (VLMs) further enable flexible semantic reasoning. Despite this progress, existing remote sensing vision-language studies remain largely centered on RGB imagery. Large-scale datasets and models such as RemoteCLIP (Liu et al., 2024a), GeoRSCLIP (Zhang et al., 2024), SkyScript (Wang et al., 2024b), GeoChat (Kuckreja et al., 2024), RSLLaVA (Bazi et al., 2024), RingMoGPT (Wang et al., 2024a), and related works mainly use visiblespectrum images as visual input. As summarized in Table 1, existing remote sensing vision-language datasets are mostly RGB-text based, while RGBIR or infrared datasets are typically designed for surveillance, driving, fusion, detection, or segmentation rather than general remote sensing visionlanguage learning. This leaves a clear gap between RGB-only remote sensing VLMs and the multimodal nature of Earth observation. Infrared imagery provides complementary cues, including intensity structures, object outlines, material
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared images provide distinctive cues, including thermal intensity structures, object boundaries, and illumination-invariant scene features, which can enrich visual-language learning beyond conventional RGB observations. However, a large-scale RGB-infrared-text dataset for remote sensing vision-language modeling is still absent. To address this gap, we introduce FusionRS, the first large-scale RGB-infraredtext dataset designed for dual-modal visionlanguage learning in remote sensing. FusionRS is constructed by translating diverse public RGB remote sensing images into infrared-style counterparts, forming aligned RGB-IR image pairs. Each pair is associated with conventional scene captions and IR-aware captions that explicitly describe infrared-specific visual properties while preserving semantic content. Based on FusionRS, we train dual-modal vision-language foundation models for RGBIR joint understanding. We first train CLIPstyle models for RGB-IR-text alignment, and then fine-tune generative VLMs for dual-modal RGB-IR captioning. Experiments show that FusionRS improves RGB-IR alignment, infraredto-text retrieval, and dual-modal captioning over RGB-only and non-IR-aware training settings. Ablation studies further verify that IRaware captions are crucial for strengthening infrared-language alignment, highlighting the importance of modality-specific textual supervision for more scalable RGB-infrared remote sensing vision-language representation learning.
1
Table 1: Comparison with existing RGB–IR, infrared, and remote sensing vision-language datasets. Existing RGB–IR datasets mainly target surveillance, driving, fusion, detection, or segmentation, whereas remote sensing vision-language datasets are largely limited to RGB–text pairs. In contrast, our dataset provides large-scale RGB–IR– text triplets for remote sensing, supporting both RGB–IR contrastive alignment and infrared-aware VLM training. Dataset LLVIP (Jia et al., 2021) M3FD (Liu et al., 2022) MFNet (Ha et al., 2017) FLIR Thermal ADAS (Teledyne FLIR, 2018)
Size 15K pairs 4K pairs 1.6K pairs 9,711 thermal / 9,233 RGB images
Source & Domains
Coverage
Visible-infrared low-light surveillance scenes Multi-scenario visible-infrared street scenes RGB-thermal urban driving scenes Thermal and visible driving scenes
RGB–IR pairing, pedestrian detection, image fusion RGB–IR fusion, object detection RGB-T semantic segmentation Object detection, autonomous driving
VEDAI (Razakarivony and Jurie, 2016) DroneVehicle (Sun et al., 2022) RSICD (Lu et al., 2017) RSITMD (Yuan et al., 2021) RS5M (Zhang et al., 2024)
1.2K images 28K pairs 10.9K images 4.7K images 5M pairs
Aerial vehicle images with multiple spectral bands Drone-based RGB-infrared vehicle images Remote sensing images from Google Earth and map platforms Remote sensing images from RSICD (Lu et al., 2017) and Google Earth Large-scale remote sensing image-text pairs
Remote sensing vehicle detection UAV traffic scenes, vehicle detection RGB–text captioning, image-text retrieval Fine-grained RGB–text retrieval Remote sensing vision-language pretraining
Ours
600K pairs
RGB remote sensing images translated into infrared counterparts
RGB–IR–text triplets, IR-aware captions, RGB–IR CLIP and VLM training
responses, and illumination-insensitive features, which are useful under illumination changes, haze, smoke, shadows, or weak texture. However, RGBinfrared remote sensing vision-language learning remains insufficiently studied, mainly due to the lack of large-scale RGB-infrared-text data.
while preserving scene semantics. This pipeline converts existing RGB remote sensing resources into RGB-IR-text triplets, enabling both contrastive representation learning and generative instruction tuning within a unified data framework. Based on FusionRS, we train and evaluate two types of remote sensing vision-language models: CLIP-style models that align RGB images, infrared images, and language in a shared embedding space for RGB-to-IR, IR-to-RGB, RGB-to-text, and IR-totext retrieval; and generative VLMs fine-tuned for dual-modal RGB-infrared image understanding, including caption generation and instruction following. Through this dual-track design, FusionRS serves as both a dataset and a testbed for integrating RGB-infrared modality information into remote sensing vision-language foundation models. Our contributions are summarized as follows:
Recent works (Cao et al., 2025; Jiang et al., 2024; Moshtaghi et al., 2025; Cao et al., 2026) have begun to explore infrared and visible-infrared visionlanguage learning. In general computer vision, infrared-oriented VLMs and RGB-thermal benchmarks show that multimodal large language models still struggle with infrared images and paired visible-thermal observations. In remote sensing, FireMM-IR (Cao et al., 2026) constructs FireMMInstruct, a paired RGB-IR instruction dataset for forest-fire monitoring, and trains a model for fire description, reasoning, and pixel-level localization. This work demonstrates the potential of RGB-IR multimodal learning in remote sensing. However, it targets a specific disaster-monitoring scenario, with data and tasks centered on forest fires rather than general remote sensing scenes. It also emphasizes infrared-enhanced MLLM understanding and segmentation, instead of constructing a general RGBIR-text dataset and systematically training both contrastive and generative vision-language models. Therefore, a general-purpose RGB-infrared-text dataset and model framework for remote sensing remains missing.
• To the best of our knowledge, we propose the first large-scale RGB-infrared-text dataset and vision-language framework for general remote sensing understanding. FusionRS constructs aligned RGB-IR image pairs from diverse public remote sensing datasets and associates them with scene-level descriptions and IR-aware captions, enabling RGB-IR-text learning beyond RGB-only remote sensing VLMs and task-specific RGB-IR monitoring. • We conduct extensive experiments to train and evaluate both CLIP-style contrastive models and generative VLMs on FusionRS. The results show that FusionRS improves RGB-IR cross-modal alignment, infrared-text retrieval, and dual-modal RGB-infrared remote sensing image understanding compared with RGBonly and non-IR-aware training settings.
In this work, we introduce FusionRS, a large-scale RGB-infrared-text dataset and visionlanguage model framework for remote sensing. As illustrated in Figure 1, FusionRS is built from diverse public RGB remote sensing datasets covering broad Earth observation scenes. We translate RGB images into infrared-style observations to construct aligned RGB-IR image pairs, associate each pair with textual descriptions, and further generate IRaware captions that capture infrared visual cues
• We conduct comprehensive ablation studies on data construction, IR-aware captioning, modality alignment, and training strategies. 2
Figure 1: Overview of the proposed dual-modal RGB-IR-text remote sensing dataset construction pipeline. FusionRS is constructed by collecting large-scale RGB remote sensing images, translating them into infrared-style observations with an RGB-to-IR diffusion translator, and generating IR-aware textual descriptions from RGB images, IR images, and original scene-level text. The resulting RGB images, IR images, and RGB-IR text annotations are integrated into dual-modal remote sensing triplets for RGB-IR contrastive alignment and infrared-aware VLM training.
These analyses show that infrared-aware textual supervision is critical for bridging the semantic gap between infrared visual patterns and language, and further demonstrate the value of RGB-IR paired data for remote sensing vision-language models.
2
VLMs remain limited in modeling non-RGB modalities, especially infrared imagery, which provides complementary intensity, structural, and illumination-insensitive cues. This RGB-centric design restricts their applicability when visible appearance is degraded by shadows, haze, smoke, low illumination, or weak texture.
Related Work
Remote sensing vision-language learning. Vision-language learning has become an important direction in remote sensing, moving beyond closed-set scene classification toward image-text retrieval, captioning, visual question answering, and instruction-following understanding. Early datasets such as RSICD (Lu et al., 2017), RSITMD (Yuan et al., 2021), and RSVQA (Lobry et al., 2020) provide captions or question-answer pairs for limited-scale RGB remote sensing images, establishing basic benchmarks for semantic grounding in Earth observation. Recent large-scale datasets and models, including RS5M (Zhang et al., 2024), SkyScript (Wang et al., 2024b), RemoteCLIP (Liu et al., 2024a), GeoRSCLIP (Zhang et al., 2024), GeoChat (Kuckreja et al., 2024), RS-LLaVA (Bazi et al., 2024), and RingMoGPT (Wang et al., 2024a), further improve image-text alignment and open-ended visual understanding by adapting CLIP-style contrastive learning (Radford et al., 2021) and generative VLM training to remote sensing data. However, most existing works still use visiblespectrum RGB imagery as the primary or only visual input. As a result, current remote sensing
Infrared and RGB-IR multimodal learning. Infrared and RGB-IR learning has been studied in general computer vision for pedestrian detection, autonomous driving, image fusion, and semantic segmentation, with representative datasets such as LLVIP (Jia et al., 2021), M3FD (Liu et al., 2022), MFNet (Ha et al., 2017), and FLIR Thermal ADAS (Teledyne FLIR, 2018). These datasets provide paired visible-infrared or RGB-thermal observations, but their domains are mainly surveillance or street-level driving, and their annotations target detection, fusion, or segmentation rather than vision-language learning. In remote sensing, datasets such as VEDAI (Razakarivony and Jurie, 2016) and DroneVehicle (Sun et al., 2022) explore aerial or UAV-based visible-infrared data, but mainly focus on object detection or scene-specific analysis. Recent infrared-oriented VLMs and RGBthermal benchmarks (Cao et al., 2025; Jiang et al., 2024; Moshtaghi et al., 2025) show that large multimodal models still struggle to interpret infrared patterns and align them with language. FireMMIR (Cao et al., 2026) demonstrates RGB-IR multimodal learning for forest-fire monitoring, but remains specific to disaster monitoring and does not 3
Table 2: Source-wise statistics of FusionRS. The columns indicate the total number of samples and the train/validation/test split.
Table 3: Statistics of the IR-aware caption subset. The validation split reports captions that pass automatic caption-quality filtering.
Total
Train
Val
Test
Split
Total
RS5M
SkyScript
NWPU
RSICD
RSITMD
RS5M (Zhang et al., 2024) SkyScript (Wang et al., 2024b) NWPU (Cheng et al., 2017) RSICD (Lu et al., 2017) RSITMD (Yuan et al., 2021)
488,033 65,266 31,186 10,824 4,691
471,729 63,134 30,122 10,494 4,521
8,166 1,036 549 166 83
8,138 1,096 515 164 87
IR-aware train IR-aware val IR-aware test
49,068 416 10,000
40,934 369 8,138
4,916 0 1,096
2,219 47 515
719 0 164
280 0 87
Total
600,000
580,000
10,000
10,000
Source
datasets enrich scene categories and textual variations. The split is performed at the sample level. Each RGB image is then paired with an infraredstyle image generated from the same visual content, so all splits contain aligned RGB–IR image pairs and corresponding textual annotations.
provide a general large-scale RGB-IR-text dataset for both contrastive alignment and generative VLM training. In contrast, FusionRS constructs largescale RGB-IR-text triplets from diverse remote sensing sources and supports RGB-IR CLIP-style representation learning and infrared-aware VLM instruction tuning under a unified framework.
3
3.2
Dataset Construction
Since existing remote sensing vision-language datasets are mostly RGB-based, directly collecting large-scale paired RGB-infrared remote sensing data is difficult. To address this limitation, we convert RGB remote sensing images into infraredstyle observations using DiffV2IR (Ran et al., 2025), a diffusion-based visible-to-infrared translation model. Specifically, we use the DiffV2IR checkpoint after_phase_2.ckpt to translate each RGB remote sensing image into an infrared-style image while preserving the original scene layout and object-level semantics.
We introduce FusionRS, a large-scale RGBinfrared-text dataset for remote sensing. FusionRS contains 600,000 RGB–IR–text triplets constructed from diverse public RGB remote sensing datasets. Each triplet consists of an RGB remote sensing image, its translated infrared-style counterpart, and textual supervision. The dataset is designed to support both RGB–IR contrastive alignment and infrared-aware vision-language model training. Figure 1 illustrates the overall construction pipeline. Table 2 summarizes the source-wise composition of FusionRS, and Table 3 reports the statistics of the IR-aware caption subset. 3.1
RGB-to-IR Image Translation
All generated infrared-style images are saved as .jpg files with a spatial resolution of 512 × 512. The images are stored in PIL RGB mode rather than single-channel grayscale mode. Visually, they follow a grayscale infrared style, but the three channels are not forced to be identical. This format allows the generated infrared-style images to be directly used by standard CLIP-style encoders and generative VLMs without modifying their image input interfaces.
Data Sources and Splits
FusionRS is built from five public remote sensing data sources: RS5M (Zhang et al., 2024), SkyScript (Wang et al., 2024b), NWPU (Cheng et al., 2017), RSICD (Lu et al., 2017), and RSITMD (Yuan et al., 2021). These datasets cover diverse Earth observation scenes, including airports, roads, residential areas, farmland, industrial regions, and water bodies. RS5M (Zhang et al., 2024) serves as the main source due to its largescale image-text pairs, while the other datasets provide complementary scene diversity and caption styles. We collect 600,000 RGB remote sensing samples and split them into 580,000 training samples, 10,000 validation samples, and 10,000 test samples. As shown in Table 2, RS5M contributes the largest portion of FusionRS, while the remaining
Through this translation process, each original RGB image is paired with an infrared-style counterpart. The resulting RGB–IR pairs preserve scene-level correspondence while introducing modality-specific visual characteristics, such as intensity-dominant appearance, reduced color texture, stronger structural contrast, and infrared-like object boundaries. These paired observations provide the visual foundation for RGB–IR contrastive learning and dual-modal RGB–IR remote sensing understanding. 4
Figure 2: Visual examples from the FusionRS dataset. For each scene, we show the original RGB image (left), the translated infrared-style image (middle), and the generated IR-aware caption (right). The IR-aware captions describe infrared-specific visual cues such as grayscale intensity, high-contrast structures, and object outlines while preserving scene semantics.
3.3
Text Annotation and IR-aware Caption Generation
itative inspection. The IR-aware training captions form an enhanced subset for infrared-aware representation learning and VLM instruction tuning, while the validation and test captions support evaluation under infrared-aware textual supervision.
Each sample in FusionRS is associated with textual supervision. The full 600,000 triplets retain original captions or scene-level textual descriptions from their source datasets. These captions provide semantic descriptions of the remote sensing scenes and serve as the basic language supervision for image-text alignment. However, original RGB captions usually describe visible-spectrum appearance and scene semantics, but rarely mention infrared-specific visual cues. This creates a modality-language gap when training models to understand infrared-style images. To reduce this gap, we additionally generate IR-aware captions using Qwen2.5-VL-72BInstruct (Bai et al., 2025) through the OpenRouter API. The model takes multi-modal seed inputs, including the RGB image, the translated infraredstyle image, and the original text description, and produces a caption that preserves the scene semantics while explicitly describing infrared visual characteristics. The IR-aware captions are designed to emphasize cues such as grayscale intensity distribution, high-contrast structures, bright and dark regions, object outlines, texture reduction, and infraredstyle visual responses. As reported in Table 3, we generate 49,068 IR-aware training captions and 10,000 IR-aware test captions. For validation, 416 generated captions pass automatic caption-quality filtering and are used for model selection and qual-
3.4
Caption Cleaning and Quality Control
We apply rule-based text cleaning and captionquality filtering to improve textual supervision. For original captions, we remove noisy captions containing irrelevant or dataset-artifact strings, such as stock photo, copyright, and Google Earth. We also filter non-English captions, extremely short or overly long captions, generic captions with limited scene information, captions without clear object- or scene-level content, and weak RS5M (Zhang et al., 2024) captions that provide insufficient semantic supervision. When a caption is missing or too weak but the source dataset provides a reliable class label, we use class-name fallback to preserve basic scenelevel supervision. For IR-aware captions, we further conduct automatic quality checks to remove invalid or unreliable outputs, including captions that are too short, contain bad responses, make unsupported physical claims, or provide weak infrared cues. This filtering ensures that generated captions describe infrared-style visual evidence rather than merely repeating original RGB captions. The final dataset therefore contains large-scale RGB–IR–text triplets with cleaned scene-level supervision and an enhanced subset with explicit infrared-aware captions. 5
Figure 3: Overview of the FusionRS construction and training pipeline. RGB remote sensing images are translated into infrared-style counterparts to form RGB–IR pairs, which are further paired with textual supervision for RGB–IR contrastive learning and infrared-aware VLM training.
Overall, FusionRS converts existing RGB remote sensing resources into a dual-modal RGB– IR–text dataset. RGB images provide original visible-spectrum observations, translated IR images introduce infrared-style modality information, and textual annotations connect both modalities to language. This construction enables systematic training and evaluation of RGB–IR CLIP-style models and infrared-aware generative VLMs.
4
vision-language models. Figure 3 illustrates the overall training pipeline. Table 4 reports the CLIP retrieval results, Figure 4 visualizes the CLIP IRaware caption ablation, and Table 5 reports the VLM training ablation. 4.1
RGB–IR–Text CLIP Training
We first train CLIP-style models using RGB–IR– text triplets from FusionRS. Each training sample contains an RGB remote sensing image, an infrared-style image translated from the same RGB image, and a corresponding text caption. The training objective is designed to align three modalities in a shared representation space: RGB images, infrared-style images, and language descriptions. This setting enables the model to learn not only image-text alignment, but also cross-modal visual alignment between RGB and infrared-style observations. Given a mini-batch of N RGB–IR–text triplets, we denote the RGB image, infrared-style image, and text caption of the i-th sample as ri , vi , and ci , respectively. The image encoder and text encoder map them into normalized embeddings:
Model Training and Evaluation Protocol
This section describes the model training and evaluation protocol used for FusionRS. We train and evaluate both CLIP-style contrastive models and generative VLMs to examine whether the proposed RGB–IR–text dataset can support crossmodal representation learning and RGB–IR visuallanguage understanding. For CLIP-style models, we evaluate multiple backbones, including OpenAI CLIP ViT-L/14 (Radford et al., 2021), OpenAI CLIP ViT-B/32 (Radford et al., 2021), OpenCLIP ViT-B/32 (Cherti et al., 2023), RemoteCLIP ViT-B/32 (Liu et al., 2024a), and GeoRSCLIP ViT-B/32 (Zhang et al., 2024). For generative VLMs, we evaluate multiple representative models, including Qwen2.5-VL-7B-Instruct (Bai et al., 2025), InstructBLIP (Dai et al., 2023)/BLIP-2 (Li et al., 2023), LLaVA-1.5 (Liu et al., 2024b), LLaVA-1.6 (Liu et al., 2024c), GeoChat (Kuckreja et al., 2024), and H2RSVLM (Pang et al., 2025). This multi-backbone protocol is designed to test whether FusionRS consistently benefits both general-domain and remote-sensing-oriented
zri = fθ (ri ),
zvi = fθ (vi ),
zci = gϕ (ci ), (1)
where fθ denotes the visual encoder and gϕ denotes the text encoder. Here, zri , zvi , and zci correspond to RGB, infrared-style, and textual embeddings, respectively. For two modalities a and b, we define the symmetric contrastive loss as: La,b = 6
i 1h CE(Sa,b , y) + CE(S⊤ , y) , a,b 2
(2)
Table 4: CLIP retrieval results on the FusionRS test split with original captions. We report Recall@1, Recall@5, and Recall@10 for IR-to-text, text-to-IR, RGB-to-IR, and IR-to-RGB retrieval. Mean R denotes the average recall over all reported retrieval metrics. Model
IR→Text
Mean R
Text→IR
RGB→IR
IR→RGB
R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 OpenAI CLIP ViT-B/32 (Radford et al., 2021) OpenAI CLIP ViT-L/14 (Radford et al., 2021) OpenCLIP ViT-B/32 (Cherti et al., 2023) RemoteCLIP ViT-B/32 (Liu et al., 2024a) GeoRSCLIP ViT-B/32 (Zhang et al., 2024)
55.49 66.22 59.83 57.58 59.32
22.00 30.76 25.86 23.44 26.22
47.41 59.20 52.35 49.11 52.45
59.75 71.08 63.89 60.80 64.54
Table 5: VLM training ablation on Qwen2.5-VL7B (Bai et al., 2025). A-original uses original captions, B-iraware uses generated IR-aware captions, and Cmixed combines original and IR-aware supervision. Setting
52.55 70.32 53.39
1.10 90.04 4.27
19.64 31.05 20.08
1.29 89.70 9.83
40.09 41.41 40.19
45.00 41.10 44.30
where y = {1, . . . , N } denotes the diagonal matching labels in the mini-batch. The similarity matrix Sa,b is computed as: Sa,b = τ · Za Z⊤ b ,
(3)
where τ is the learnable temperature parameter, and Za and Zb are the normalized embedding matrices of the two modalities. The overall RGB–IR–text contrastive objective is: 1 LCLIP = (Lr,c + Lv,c + Lr,v ) . (4) 3 Here, Lr,c aligns RGB images with text, Lv,c aligns infrared-style images with text, and Lr,v explicitly aligns RGB and infrared-style images. The RGB–IR term is particularly important for enforcing cross-modal visual consistency, while the two image-text terms connect both visual modalities to language. Table 4 reports the retrieval performance of different CLIP backbones on FusionRS. OpenAI CLIP ViT-L/14 (Radford et al., 2021) achieves the best overall performance, reaching a mean recall of 66.22. Across all evaluated backbones, RGB–IR retrieval obtains higher recall than IR–text retrieval, suggesting that visual modality alignment is easier than infrared-language alignment. This gap further motivates the use of IR-aware captions to strengthen infrared-text grounding. 4.2
47.57 59.53 52.15 48.75 52.52
59.85 71.45 63.38 60.57 64.44
56.65 68.45 62.17 59.62 59.02
73.46 81.56 77.60 76.05 76.09
79.49 86.17 82.80 81.47 81.54
53.50 69.52 60.42 58.11 57.63
69.13 81.15 73.52 72.71 73.66
74.55 84.89 77.82 76.95 77.75
fine-tuning. The motivation is that captions from RGB remote sensing datasets mainly describe visible-spectrum scene semantics and often lack infrared-specific cues, such as grayscale intensity, bright/dark regions, high-contrast structures, and weakened texture details. Directly using these captions for infrared-style imagery may therefore introduce a modality-language gap. To address this issue, we use the generated IR-aware captions described in Section 3.3 to construct an enhanced fine-tuning subset. The full CLIP pretraining stage uses 600,000 RGB–IR–text triplets with original captions, while the IR-aware fine-tuning stage uses 49,068 generated IR-aware training captions.
Cap. Auto Cap. IR Cap. R-L IR-cue Obj. QA Scene QA
A-original B-iraware C-mixed
22.46 30.89 26.05 23.41 25.97
We consider three supervision settings: Aoriginal, B-iraware, and C-mixed. A-original pairs infrared images with original captions, B-iraware pairs infrared images with generated IR-aware captions, and C-mixed combines original and IR-aware supervision. For CLIP fine-tuning, the mixed setting includes IR–original caption pairs, IR–IRaware caption pairs, RGB–original caption pairs, and RGB–IR image pairs. This design preserves scene-level semantic supervision, injects infraredspecific textual grounding, and maintains RGB–IR visual correspondence, reducing the risk that the model overfits to textual changes alone.
Figure 4 analyzes the effect of IR-aware caption supervision across different CLIP backbones. Cmixed consistently achieves the best performance across all evaluated models, indicating that IRaware captions are most effective when combined with original scene-level supervision and explicit RGB–IR visual alignment. In contrast, A-original and B-iraware alone do not always improve over the 580K-only baseline, suggesting that original and IR-aware captions provide complementary supervision.
IR-aware Caption Enhanced Fine-tuning
After contrastive training with original captions, we further conduct IR-aware caption enhanced 7
4.3
VLM Instruction Tuning
We further evaluate whether FusionRS supports infrared-aware generative VLM training. Unlike CLIP-style training, which focuses on representation alignment and retrieval, VLM instruction tuning aims to improve natural-language generation and question answering over infrared-style remote sensing images. We construct instruction data from the IR-aware caption subset, where each sample contains an infrared-style image, a naturallanguage instruction, and a target response. For captioning, the model is asked to describe the image; for VQA, it answers scene- or object-level questions according to the visual content. The response is derived from either the original caption or the generated IR-aware caption. We consider three settings: A-original, which uses 49,068 infrared-image instructions with original captions; B-iraware, which uses 49,068 instructions with IR-aware captions; and C-mixed, which combines both supervision types into 98,136 samples to test whether scene-level semantics and infrared-specific descriptions are complementary. For VLM fine-tuning, we adopt LoRA (Hu et al., 2022) to reduce training cost while preserving the general multimodal capability of the backbone models. Unless otherwise specified, we set the LoRA rank to 32, alpha to 64, and dropout to 0.05. Models are trained for one epoch with a learning rate of 2× 10−4 , cosine scheduling, a warmup ratio of 0.03, bf16 precision, per-device batch size of 1, and gradient accumulation steps of 8. Qwen2.5VL-72B-Instruct (Bai et al., 2025) is used only for generating IR-aware captions, not as the fine-tuned VLM backbone. Table 5 reports the VLM training ablation on Qwen2.5-VL-7B (Bai et al., 2025). Compared with A-original, B-iraware substantially improves Caption Auto, Caption IR Score, Caption R-L, IRcue QA, and Object QA, showing that IR-aware captions provide stronger supervision for infraredspecific generation and reasoning. A-original achieves the best Scene QA score, suggesting that original captions retain useful scene-level semantics. C-mixed gives intermediate performance, indicating that mixed supervision may require more careful balancing for generative instruction tuning. 4.4
Figure 4: CLIP IR-aware fine-tuning ablation across different backbones. C-mixed consistently achieves the best performance, showing the benefit of combining original captions, IR-aware captions, and RGB–IR alignment.
its contribution to RGB–IR–text vision-language learning. For CLIP-style models, we conduct crossmodal retrieval on the validation and test splits, covering RGB-to-text, IR-to-text, text-to-RGB, textto-IR, and RGB–IR retrieval. We report Recall@1, Recall@5, and Recall@10 to quantify RGB–text semantic alignment, infrared–language grounding, and cross-modal RGB–IR visual correspondence. For generative VLMs, we evaluate infrared-style remote sensing captioning and visual question answering, which measure whether models can generate semantically faithful, infrared-aware descriptions and answer questions based on infrared visual evidence rather than relying only on RGBbiased semantics. We compare original-caption, IR-aware-caption, and mixed-supervision settings to analyze the effect of infrared-specific textual grounding on both discriminative alignment and generative understanding. Overall, this protocol examines whether FusionRS improves RGB–IR– text representation alignment, narrows the infrared– language gap, and generalizes consistently across different CLIP and VLM backbones, thereby validating its usefulness for infrared remote sensing vision-language learning.
5
Conclusion
We introduced FusionRS, a large-scale RGBinfrared-text dataset for remote sensing visionlanguage learning. FusionRS constructs aligned RGB–IR–text triplets from diverse public remote sensing datasets and enriches them with IR-aware captions that preserve scene semantics while explicitly describing infrared visual cues. This design enables systematic learning of RGB–IR correspondence, infrared–language grounding, and
Evaluation Protocol
We evaluate FusionRS from both retrieval and generation perspectives to comprehensively assess 8
cross-modal semantic alignment. Experiments with CLIP-style models and generative VLMs demonstrate improved RGB–IR alignment, infrared-text retrieval, captioning, and VQA, highlighting the value of infrared-aware supervision. These results suggest that FusionRS provides a useful data foundation for advancing infrared remote sensing vision-language models and broader multimodal Earth observation understanding.
thetically generated, they should not be used as authoritative physical thermal measurements or as direct evidence for safety-critical decisions. Misinterpreting synthetic infrared-style images as real sensor observations may lead to incorrect conclusions in high-stakes scenarios such as emergency response, military analysis, environmental monitoring, or public policy. In addition, models trained on FusionRS may inherit biases from the source datasets and from the RGB-to-IR translation process. We encourage users to clearly disclose the synthetic nature of the infrared modality, evaluate model behavior carefully before deployment, and avoid using the dataset or trained models for surveillance, targeting, or other applications that may harm individuals or communities.
Limitations FusionRS has several limitations. First, its infrared images are generated from RGB remote sensing images rather than captured by physically paired infrared sensors. Therefore, they should be regarded as infrared-style observations instead of fully physical infrared measurements. Although the translation model can synthesize plausible infrared appearances, it may introduce modality artifacts, intensity biases, or inaccurate thermal responses, especially for physical properties that cannot be directly inferred from RGB imagery alone. Second, although we apply caption cleaning and quality filtering, the original and IR-aware captions may still contain incomplete descriptions, weak infrared cues, or occasional semantic errors. Such noise may affect both contrastive alignment and generative VLM training. Third, our experiments mainly focus on standard retrieval, captioning, and VQA tasks. More downstream applications, such as disaster monitoring, object detection, change detection, and crossmodal geospatial reasoning, remain to be further explored. Finally, FusionRS is designed for remote sensing vision-language learning, and its effectiveness on real sensor-captured RGB-infrared paired data still requires further validation.
References Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. Qwen2.5-vl technical report. Preprint, arXiv:2502.13923. Yakoub Bazi, Laila Bashmal, Mohamad Mahmoud Al Rahhal, Riccardo Ricci, and Farid Melgani. 2024. Rs-llava: A large vision-language model for joint captioning and question answering in remote sensing imagery. Remote Sensing, 16(9):1477. Jinghao Cao, Xiajun Liu, and Rui Xue. 2026. Firemmir: An infrared-enhanced multi-modal large language model for comprehensive scene understanding in remote sensing forest fire monitoring. Sensors, 26(2):390. Zhe Cao, Jin Zhang, and Ruiheng Zhang. 2025. Irgpt: Understanding real-world infrared image with bicross-modal curriculum on large-scale benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 166–176.
Ethical Considerations FusionRS is constructed from public remote sensing datasets and generated infrared-style images. It does not intentionally include personally identifiable information, and the spatial resolution of common remote sensing imagery generally limits direct identification of individuals. Nevertheless, remote sensing data may still involve sensitive geographic regions, infrastructure, or human activity patterns. Therefore, FusionRS should be used responsibly and in accordance with the licenses, usage restrictions, and ethical requirements of the original datasets. Since the infrared images in FusionRS are syn-
Gong Cheng, Junwei Han, and Xiaoqiang Lu. 2017. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883. Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. 2023. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2818–2829. Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N
9
Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. Advances in neural information processing systems, 36:49250–49267.
the IEEE/CVF conference on computer vision and pattern recognition, pages 5802–5811. Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. 2020. Rsvqa: Visual question answering for remote sensing data. IEEE Transactions on Geoscience and Remote Sensing, 58(12):8555–8566.
Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada. 2017. Mfnet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5108–5115. IEEE.
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, and Xuelong Li. 2017. Exploring models and data for remote sensing image caption generation. IEEE Transactions on Geoscience and Remote Sensing, 56(4):2183–2195.
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adaptation of large language models. Iclr, 1(2):3.
Mehdi Moshtaghi, Siavash H Khajavi, and Joni Pajarinen. 2025. Rgb-th-bench: A dense benchmark for visual-thermal understanding of vision language models. arXiv preprint arXiv:2503.19654.
Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou. 2021. Llvip: A visible-infrared paired dataset for low-light vision. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3496–3504.
Chao Pang, Xingxing Weng, Jiang Wu, Jiayu Li, Yi Liu, Jiaxing Sun, Weijia Li, Shuai Wang, Litong Feng, Gui-Song Xia, and 1 others. 2025. Vhm: Versatile and honest vision language model for remote sensing image analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 6381–6388.
Shixin Jiang, Zerui Chen, Jiafeng Liang, Yanyan Zhao, Ming Liu, and Bing Qin. 2024. Infrared-llava: Enhancing understanding of infrared images in multimodal large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8573–8591.
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and 1 others. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR.
Kartik Kuckreja, Muhammad Sohail Danish, Muzammal Naseer, Abhijit Das, Salman Khan, and Fahad Shahbaz Khan. 2024. Geochat: Grounded large vision-language model for remote sensing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 27831–27840.
Lingyan Ran, Lidong Wang, Guangcong Wang, Peng Wang, and Yanning Zhang. 2025. Diffv2ir: visibleto-infrared diffusion model via vision-language understanding. arXiv preprint arXiv:2503.19012.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pretraining with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR.
Sebastien Razakarivony and Frederic Jurie. 2016. Vehicle detection in aerial imagery: A small target detection benchmark. Journal of Visual Communication and Image Representation, 34:187–203.
Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. 2024a. Remoteclip: A vision language foundation model for remote sensing. IEEE Transactions on Geoscience and Remote Sensing, 62:1–16.
Yiming Sun, Bing Cao, Pengfei Zhu, and Qinghua Hu. 2022. Drone-based rgb-infrared cross-modality vehicle detection via uncertainty-aware learning. IEEE Transactions on Circuits and Systems for Video Technology, 32(10):6700–6713.
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024b. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26296–26306.
Teledyne FLIR. 2018. Teledyne FLIR ADAS Dataset: FLIR Thermal Dataset for Algorithm Training. https://oem.flir.com/en-gb/solutions/ automotive/adas-dataset-form/. Accessed: 2026-05-25.
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024c. Llavanext: Improved reasoning, ocr, and world knowledge.
Peijin Wang, Huiyang Hu, Boyuan Tong, Ziqi Zhang, Fanglong Yao, Yingchao Feng, Zining Zhu, Hao Chang, Wenhui Diao, Qixiang Ye, and 1 others. 2024a. Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks. IEEE Transactions on Geoscience and Remote Sensing, 63:1–20.
Jinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, and Zhongxuan Luo. 2022. Target-aware dual adversarial learning and a multiscenario multi-modality benchmark to fuse infrared and visible for object detection. In Proceedings of
10
Zhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu, and Ram Rajagopal. 2024b. Skyscript: A large and semantically diverse vision-language dataset for remote sensing. In Proceedings of the AAAI Conference on Artificial Intelligence, 6, pages 5805–5813. Zhiqiang Yuan, Wenkai Zhang, Kun Fu, Xuan Li, Chubo Deng, Hongqi Wang, and Xian Sun. 2021. Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval. IEEE Transactions on Geoscience and Remote Sensing, 60:1–19. Zilun Zhang, Tiancheng Zhao, Yulong Guo, and Jianwei Yin. 2024. Rs5m and georsclip: A large-scale vision-language dataset and a large vision-language model for remote sensing. IEEE Transactions on Geoscience and Remote Sensing, 62:1–23.
11
Appendix We provide supplementary details for dataset construction, experimental settings, additional ablations, and qualitative examples.
A
Dataset Details
A.1
Source Statistics
Source
Train
Val
RS5M 488,033 471,729 SkyScript 65,266 63,134 NWPU 31,186 30,122 RSICD 10,824 10,494 RSITMD 4,691 4,521
8,166 1,036 549 166 83
Total
Total
Test IR Train IR Val IR Test 8,138 1,096 515 164 87
40,934 4,916 2,219 719 280
369 0 47 0 0
8,138 1,096 515 164 87
600,000 580,000 10,000 10,000
49,068
416
10,000
Table 6: Complete source-wise statistics of FusionRS. FusionRS is built from RS5M, SkyScript, NWPU, RSICD, and RSITMD. The full split contains 600K RGB–IR–text triplets: 580K train, 10K validation, and 10K test. The IR-aware subset contains 49,068 training captions, 416 filtered validation captions, and 10K test captions.
Stage
Rule
Purpose
Original caption cleaning Original caption cleaning Original caption cleaning RGB-to-IR translation IR-aware caption filtering IR-aware caption filtering Manual inspection
Remove stock-photo, copyright, and map-source text Remove non-English, too-short, too-long, or generic captions Use class-name fallback for missing or weak captions Remove failed, missing, or corrupted translated IR images Remove too-short outputs, bad responses, and unsupported claims Remove captions that repeat RGB captions without IR cues Inspect sampled RGB–IR–caption triplets
Reduce dataset noise Preserve semantic supervision Keep scene-level supervision Ensure valid pairing Keep reliable annotations Reduce modality-language mismatch Check visual correspondence
Table 7: Filtering and quality-control rules. Original captions are cleaned before training, translated IR images are checked for valid RGB–IR pairing, and IR-aware captions are filtered to keep infrared-style visual evidence. Item
Setting
Caption generator Input context Output target Main IR cues IR-aware train captions Filtered IR-aware val captions IR-aware test captions
Qwen2.5-VL-72B-Instruct through OpenRouter API RGB image, translated infrared-style image, and original caption or scene text Preserve scene semantics while emphasizing infrared-style visual cues Grayscale intensity, high contrast, bright structures, dark or low-texture regions, object outlines 49,068 samples 416 samples 10,000 samples
Table 8: IR-aware caption generation setup. Qwen2.5-VL-72B-Instruct is used only for annotation generation. The annotator receives the RGB image, translated IR-style image, and original caption or scene text, then generates a caption preserving scene semantics while adding infrared cues.
12
Table 9: CLIP tri-modal contrastive objective. Each batch contains aligned RGB images, translated infrared-style images, and captions. The objective aligns RGB–text, IR–text, and RGB–IR pairs in one shared embedding space. Item
Definition
Embeddings
c zri , zv i , and zi denote normalized RGB, IR, and text embeddings.
Probability
pa→b = PN i,j
Pair loss
1 La,b = − 2N
Total loss
Ltotal = Lr,c + Lv,c + Lr,v
exp((za )⊤ zb /τ ) i j
k=1
exp((za )⊤ zb /τ ) i k log pa→b + log pb→a i,i i,i i
P
Table 10: Default VLM LoRA fine-tuning hyperparameters. LoRA is used for VLM instruction tuning to reduce training cost. Hyperparameter
Value
LoRA rank LoRA alpha LoRA dropout Learning rate Epochs Per-device batch size Gradient accumulation Effective batch size Warmup ratio Scheduler Precision Gradient checkpointing
32 64 0.05 2 × 10−4
1 1 8 8 0.03 Cosine bf16 Enabled
Table 11: Prompt templates used for annotation and VLM evaluation. Use
Instruction template
IR-aware annotation
Given the RGB image, translated infrared-style image, and original caption or scene text, generate one caption that preserves scene semantics and explicitly describes grayscale intensity, high contrast, bright structures, dark or low-texture regions, and structural outlines. Describe this infrared remote sensing image in detail. What is the main scene type in this infrared remote sensing image? Answer with a concise scene label only. What are the main objects or land-cover elements visible in this infrared remote sensing image? Answer with a concise comma-separated list. What infrared-specific visual cues can be observed in this image? Mention visible grayscale intensity, high contrast, bright structures, low-texture or dark regions, and structural outlines if present.
IR captioning Scene QA Object QA IR-cue QA
Table 12: CLIP fine-tuning settings. 580K-only is the large-scale baseline. A-original uses original captions, B-iraware uses IR-aware captions, and C-mixed combines scene semantics, infrared-aware language, and RGB–IR visual alignment. Setting
Training pairs
Goal
580K-only A-original B-iraware C-mixed
RGB–caption, IR–caption, RGB–IR IR–original caption IR–IR-aware caption IR–original, IR–IR-aware, RGB–original, RGB–IR
Large-scale baseline training Original-caption IR-language tuning IR-specific language grounding Combined semantic, IR-aware, and visual alignment
Table 13: Retrieval task definitions. These tasks evaluate RGB–text, IR–text, and RGB–IR alignment in the shared CLIP embedding space. Task
Query / Gallery
Purpose
IR→Text Text→IR RGB→IR IR→RGB RGB→Text Text→RGB
IR image retrieves text captions Text caption retrieves IR images RGB image retrieves paired IR images IR image retrieves paired RGB images RGB image retrieves text captions Text caption retrieves RGB images
IR-language alignment Language-to-IR grounding RGB-to-IR visual alignment IR-to-RGB visual alignment RGB-language alignment Text-to-RGB retrieval
13
Table 14: CLIP IR-aware fine-tuning ablation across backbones. Mean R is reported on the FusionRS test split with original captions. C-mixed gives the best result across all CLIP backbones.
Model
580K-only A-original B-iraware C-mixed
OpenAI CLIP ViT-B/32 OpenAI CLIP ViT-L/14 OpenCLIP ViT-B/32 RemoteCLIP ViT-B/32 GeoRSCLIP ViT-B/32
55.49 66.22 59.83 57.58 59.32
53.58 64.25 58.22 55.33 57.49
51.40 62.37 56.21 54.27 55.61
61.41 66.78 65.34 63.17 64.41
Table 15: Full VLM training ablation across generative backbones. Caption metrics are evaluated on the 10K test caption split, and QA metrics are evaluated on the 1K validation QA split. A-original uses original captions, B-iraware uses IR-aware captions, and C-mixed combines both.
Model
Setting
A-original B-iraware C-mixed A-original H2RSVLM B-iraware C-mixed A-original InstructBLIP B-iraware C-mixed A-original LLaVA-1.5 B-iraware C-mixed A-original LLaVA-1.6 B-iraware C-mixed A-original Qwen2.5-VL-7B B-iraware C-mixed GeoChat
Cap. Auto Cap. IR Cap. R-L IR-cue QA Obj. QA Scene QA 50.82 62.11 63.55 51.12 70.51 62.46 51.33 70.22 51.40 50.98 60.93 62.26 51.26 62.50 55.19 52.55 70.32 53.39
1.34 33.05 45.38 0.69 94.78 59.34 0.54 86.32 1.22 1.77 26.81 38.33 0.82 31.70 17.62 1.10 90.04 4.27
14
18.77 13.22 17.49 18.67 29.26 25.58 19.55 31.43 19.58 18.76 12.36 16.12 18.67 13.62 17.30 19.64 31.05 20.08
3.89 96.00 93.54 0.74 92.90 89.56 0.49 84.43 0.71 4.49 90.70 96.77 2.71 79.77 94.49 1.29 89.70 9.83
33.15 33.84 34.39 28.48 31.85 32.59 34.02 33.18 33.53 30.82 32.79 34.79 33.33 33.56 33.37 40.09 41.41 40.19
21.80 30.60 27.80 27.60 30.70 31.10 35.00 32.60 35.90 28.30 26.60 28.70 28.30 29.70 31.00 45.00 41.10 44.30
Table 16: Backbone generalization under B-iraware instruction tuning. This table keeps only the B-iraware setting and compares whether IR-aware supervision generalizes across remote-sensing VLMs and general VLMs.
Backbone GeoChat H2RSVLM InstructBLIP LLaVA-1.5 LLaVA-1.6 Qwen2.5-VL-7B
Cap. Auto Cap. IR IR-cue QA Obj. QA Scene QA 62.11 70.51 70.22 60.93 62.50 70.32
33.05 94.78 86.32 26.81 31.70 90.04
96.00 92.90 84.43 90.70 79.77 89.70
33.84 31.85 33.18 32.79 33.56 41.41
30.60 30.70 32.60 26.60 29.70 41.10
Table 17: Generative VLM task definitions. The current benchmark includes IR captioning, scene QA, object QA, and IR-cue QA; grounding, counting, re-identification, and security-oriented QA are not included. Task
Input / Output
IR captioning
IR image / free-form caption
Scene QA Object QA IR-cue QA
Purpose
Evaluate semantic and infrared-aware description generation IR image + scene question / scene label or short Evaluate scene-level understanding answer IR image + object question / object or land-cover list Evaluate object and land-cover recognition IR image + cue question / infrared cue explanation Evaluate grayscale intensity, contrast, dark regions, and outlines
15
Table 18: Qualitative RGB–IR captioning examples, part 1. Each row shows one case with the RGB image, translated IR image, IR-aware caption, and Qwen2.5-VL-7B output after B-iraware instruction tuning. Case
RGB Image
IR Image
IR-aware Caption
VLM Output
RS5M railway
The infrared-style image captures a train riding on a railway bridge, with the train and tracks exhibiting a bright intensity response due to their structural details.
A grayscale aerial view captures the bridge with high-intensity structural outlines against the lower-intensity river, showcasing distinct contrast between the built and.
RS5M harbor
The infrared-style image captures an aerial view of a waterfront area with boats docked in the water, where the boats and buildings exhibit varying grayscale intensities.
The infrared-style aerial view of the harbor reveals boats with varying grayscale intensities, where brighter areas highlight structural outlines and darker regions indicate.
RSICD river
The infrared-style image shows a river with a dark intensity The river appears as a dark, high-contrast line cutting through the response running through the center, contrasting sharply with the lighter, textured regions representing the bare land and trees. brighter surroundings.
SkyScript random fill
A high-contrast infrared-style aerial view reveals a residential The infrared-style image depicts a single-family residential area area with varying grayscale intensities highlighting the structural with varying grayscale intensities. outlines of numerous houses.
NWPU rectangular farmland
The infrared-style image reveals a landscape of rectangular farmlands with varying grayscale intensities, where some areas appear brighter and others darker, indicating differences in.
The image displays various rectangular farmland plots with distinct grayscale intensities, where some areas appear brighter due to higher reflectivity while others show.
RS5M ship
The infrared-style image reveals a fishing boat and wooden pier with distinct grayscale intensity patterns.
A grayscale aerial view shows a boat with distinct bright edges and low-texture interior, tethered to a dock with high-contrast structural outlines against the darker water.
16
Table 19: Qualitative RGB–IR captioning examples, part 2. Each row shows one case with the RGB image, translated IR image, IR-aware caption, and Qwen2.5-VL-7B output after B-iraware instruction tuning. Case
RGB Image
IR Image
IR-aware Caption
VLM Output
RS5M tennis
The infrared-style image reveals a small area with tennis courts, The tennis courts exhibit high contrast with bright white lines where the courts appear as bright intensity regions with clear against darker surfaces, while the surrounding grass appears as a structural outlines. low-texture, medium-gray region.
RS5M industrial
The infrared-style image reveals an aerial view of a factory with The infrared-style aerial view highlights the factory and industrial multiple stacks and cooling towers, where the structures exhibit buildings with varying grayscale intensities, where structures varying grayscale intensities. appear as high-contrast outlines against.
RS5M ship
The infrared-style image captures the cruise ship in St.
A high-contrast infrared-style aerial view highlights the cruise ship’s structural outlines with bright intensity along its deck and darker tones in the surrounding water,.
RS5M ship
The infrared-style image shows a large container ship with a A high-contrast grayscale aerial view shows a container ship with bright intensity response along its structure, indicating the ship’s distinct bright and dark intensity patterns highlighting its prominent features. structure against the smoother, low-texture.
RS5M airport
The infrared-style image of Lakeland Linder Regional Airport showcases a high-contrast scene with the runway and taxiways appearing as lighter intensity patterns against the darker.
RS5M parking
The infrared-style image reveals a dense aerial view of a parking The infrared-style aerial view highlights the parking lot with lot filled with trucks, exhibiting varying grayscale intensities that varying grayscale intensities, where cars appear as distinct bright highlight the structural outlines of the. patches against darker low-texture regions.
17
The infrared-style aerial view of the airport showcases high-contrast structural outlines with varying grayscale intensities, where darker tones highlight low-texture regions.
Table 20: Qualitative RGB–IR captioning examples, part 3. Each row shows one case with the RGB image, translated IR image, IR-aware caption, and Qwen2.5-VL-7B output after B-iraware instruction tuning. Case
RGB Image
IR Image
IR-aware Caption
VLM Output
RS5M bridge
The infrared-style image captures an aerial view of a city adjacent to a large body of water, with buildings and a bridge prominently outlined in bright intensity responses.
The infrared-style view captures a bridge with high contrast edges spanning over a dark, low-texture river, while the surrounding buildings exhibit varying grayscale.
RS5M baseball
The infrared-style image showcases an aerial view of a baseball stadium with varying grayscale intensities.
The grayscale infrared-style images depict the baseball stadium with varying intensity levels, where the field appears as a low-texture region surrounded by brighter.
NWPU basketball court
The infrared-style image shows four basketball courts adjacent to a baseball diamond and a building, with the courts exhibiting a bright intensity response due to their structured.
The grayscale intensity highlights the structural outlines of the basketball court with varying brightness levels, while the surrounding buildings exhibit low-texture regions.
NWPU mountain
The rugged mountain range exhibits a series of bright and dark intensity responses, highlighting the structural outlines and ridges.
The mountain exhibits varying grayscale intensities with darker regions indicating dense vegetation, while brighter areas suggest open spaces or lighter textures.
NWPU beach
The infrared-style image shows a beach with varying grayscale intensities, where the sand appears as a lighter region compared to the darker land covered with plants.
The beach exhibits high-intensity grayscale tones with smooth, low-texture regions indicating sandy areas, while the adjacent water appears darker with subtle ripples.
RS5M storage tank
The infrared-style image highlights the storage tanks at the center and top-left blocks with bright intensity responses, contrasting sharply against the darker surrounding areas.
The grayscale infrared-style image highlights the storage tanks with distinct circular outlines, showing varying intensity levels where the tanks appear brighter compared to.
18
Table 21: Qualitative RGB–IR captioning examples, part 4. Each row shows one case with the RGB image, translated IR image, IR-aware caption, and Qwen2.5-VL-7B output after B-iraware instruction tuning. Case
RGB Image
IR Image
IR-aware Caption
VLM Output
NWPU dense residential
In the infrared-style image, the dense residential area exhibits a pattern of bright intensity responses from the rooftops, contrasting with the darker, low-texture regions.
The dense residential area exhibits high grayscale intensity with distinct structural outlines, contrasting against the darker, low-texture regions of the adjacent trees and.
NWPU railway station
The railway station exhibits a pattern of bright intensity responses The railway station exhibits high-intensity grayscale patterns with along the railway tracks, contrasting with the darker residential distinct structural outlines, contrasting sharply against the area and the river. lower-intensity surrounding buildings and.
NWPU roundabout
The roundabout with four exits and entrances exhibits a distinct pattern of grayscale intensity, where the roads appear as lighter paths contrasting with the darker surrounding areas.
The roundabout exhibits high contrast with its surrounding areas, appearing as a lighter region due to its smooth texture and structural outlines.
RSICD port
The infrared-style image highlights the port with three sea walls, where the walls appear as bright intensity responses against the darker water.
The grayscale intensity highlights the structural outlines of the port with varying brightness levels, while the surrounding buildings exhibit low-texture regions and.
RS5M road
The infrared-style image captures an aerial view of an empty highway with distinct grayscale intensity patterns.
The grayscale infrared-style view captures the contrast between the dark, low-texture road and the brighter, textured fields below, with structural outlines of the road’s.
RS5M forest
The infrared-style image captures an aerial view of an autumn forest, where the grayscale intensity patterns highlight the structural outlines of trees and vegetation.
The infrared-style aerial view of the forest and mountains reveals high-contrast grayscale patterns, with dense tree canopies appearing as bright, textured regions against.
19
Table 22: Qualitative RGB–IR captioning examples, part 5. Each row shows one case with the RGB image, translated IR image, IR-aware caption, and Qwen2.5-VL-7B output after B-iraware instruction tuning. Case
RGB Image
IR Image
IR-aware Caption
VLM Output
RS5M vehicle
The infrared-style image reveals a bustling street scene with a A high-contrast infrared-style aerial view captures the bustling large crowd gathered in the middle, where the dense aggregation airport scene, where numerous vehicles exhibit varying grayscale of people creates a high-intensity grayscale pattern. intensities, creating distinct bright and.
RS5M industrial
The infrared-style image captures an aerial view of power plants The infrared-style image reveals an aerial view of a factory with and factories, where the cooling towers and smoke stacks exhibit distinct grayscale intensity patterns, where the cooling towers appear as bright, high-contrast structures. bright intensity responses, suggesting structural.
RS5M resort beach
The infrared-style image captures an aerial view of a resort and beach, where the buildings exhibit bright intensity responses, suggesting structural outlines and contrast against the.
The infrared-style aerial view of the beach and resort showcases high-contrast structural outlines with bright intensity responses from buildings and low-texture regions in.
NWPU mobile home park
The mobile home park exhibits a pattern of bright intensity responses along the roads and around the mobile homes, contrasting with the darker grassy areas.
The mobile home park exhibits high-contrast grayscale patterns with bright structural outlines against darker low-texture regions, situated adjacent to a road with varying.
RS5M stadium
The infrared-style image of Aloha Stadium in Honolulu, Hawaii, showcases a bright intensity response around the stadium structure, emphasizing its central position.
The grayscale infrared-style view highlights the stadium with distinct bright and dark intensity patterns, emphasizing its structural outlines against the surrounding.
RSITMD multi-storey roads
The interlaced viaducts and multi-storey highways exhibit a complex network of bright and dark intensity responses, highlighting the structural outlines and contrasts.
The overpass exhibits high-intensity grayscale patterns with distinct structural outlines, contrasting sharply against the lower-intensity, textured regions representing.
20