ConceptioArchivearXiv CS
arXiv CSopen access

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

arXiv:2607.28526v1 [cs.CV] 30 Jul 2026

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration Cencen Liu

Wen Yin

Dongyang Zhang

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

Dongmin Li

Shan Zhao

Bing Su

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

Jiigan Technology Chengdu, Sichuan, China [email protected]

Jiigan Technology Chengdu, Sichuan, China [email protected]

Tao He

Jielei Wang

Guoming Lu∗

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

University of Electronic Science and Technology of China Chengdu, Sichuan, China [email protected]

Abstract

CCS Concepts

All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonly encode heterogeneous degradation conditions in a shared latent space, where degradation-related cues and scene content can remain entangled. We characterize the resulting challenge as dual ambiguity: semantic ambiguity in channel-wise modulation and spatial ambiguity in restoration responses, which can lead to content corruption and residual artifacts. To mitigate this issue, we propose DAR-Net, a Dual-Ambiguity Rectification Network for all-in-one image restoration. DAR-Net first introduces a Degradation Archetype Representation (DAR) module to construct a structured degradation state through simplex-constrained archetype mixture modeling. Based on this state, a Semantic Ambiguity Rectification (SeAR) module generates degradation-aware prompts to improve channel-wise conditioning in the decoder. A Spatial Ambiguity Rectification (SpAR) module further regularizes degradation-aware and complementary features toward orthogonal response subspaces, reducing spatial interference between removal and preservation cues. Extensive experiments on standard all-in-one restoration benchmarks show that DAR-Net achieves the best overall performance under both three-degradation and five-degradation settings, improving the average PSNR over the strongest competitor by 0.14 dB and 0.34 dB, respectively; it additionally shows superior performance on CDD-11 and WeatherBench.

• Computing methodologies → Reconstruction.

∗ Corresponding author.

This work is licensed under a Creative Commons Attribution 4.0 International License. MM ’26, Rio de Janeiro, Brazil © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2213-4/2026/11 https://doi.org/10.1145/3767308.3835300

Keywords All-in-one image restoration, prompt-based restoration, degradation modeling, dual-ambiguity rectification ACM Reference Format: Cencen Liu, Wen Yin, Dongyang Zhang, Dongmin Li, Shan Zhao, Bing Su, Tao He, Jielei Wang, and Guoming Lu. 2026. What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration. In Proceedings of the 34th ACM International Conference on Multimedia (MM ’26), November 10–14, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA, 10 pages. https://doi.org/10.1145/3767308.3835300

1

Introduction

Image restoration aims to recover clean visual content from degraded observations affected by factors such as noise, haze, and rain conditions. As a front-end step in real-world multimedia content capture and enhancement pipelines, image restoration is also closely related to multimedia content quality. Most existing restoration methods are developed under a predefined degradation setting, and many representative models [4, 6, 52] are still deployed in a task-specific manner in practice. Such a paradigm leads to considerable computational and storage overhead and limits generalization in complex environments. To address this limitation, all-in-one image restoration (AIR) has emerged as a unified framework for handling diverse degradations within a single model [10, 19, 30]. Early AIR methods often rely on multi-branch designs, where different degradation types are handled by separate branches or task-specific modules. Although such designs improve degradation specialization, their parameter cost typically scales with the number of degradation types, which limits scalability and makes model deployment increasingly inefficient as restoration scenarios become

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

LQ

GT

PromptIR

Liu et al.

DFPIR

Ours

(a) Content Corruption.

LQ

GT

PromptIR

DFPIR

Ours

(b) Residual Degradation.

Figure 1: Typical failure modes caused by degradationcontent entanglement. (a) Content Corruption, where faithful image details are mistakenly suppressed. (b) Residual Degradation, where degradation artifacts are retained.

more diverse [14, 21]. To improve scalability, subsequent studies increasingly adopt shared-backbone conditional restoration, where a unified network is modulated by degradation cues [19, 30, 37]. Within this paradigm, different conditioning mechanisms have been explored. Prompt-based methods provide a representative and efficient solution by deriving lightweight conditioning signals from the input and injecting them into a shared backbone, thereby enabling input-adaptive restoration with limited additional parameters [1, 17, 24]. In parallel, other alternatives such as MoEbased routing enhance model capacity through dynamic expert selection, allowing different restoration patterns to be handled by different experts when necessary [50, 51, 56]. However, although these methods differ in how they introduce degradation-aware conditioning, they still predominantly rely on shared latent representations within a unified restoration pipeline. As a result, a common difficulty remains insufficiently addressed: degradationrelated cues and content-related representations are often encoded in an entangled manner within the shared latent space. Therefore, existing unified restoration models still face difficulty in distinguishing what should be removed from what should be preserved. This challenge is often reflected in two typical failure modes, as shown in Fig. 1: (1) Content Corruption, where faithful image content is mistakenly removed together with degradation, and (2) Residual Degradation, where degradation patterns are not sufficiently suppressed in the restored image. We analyze these failures through two forms of ambiguity. The first is Semantic Ambiguity, namely the channel-wise entanglement between degradation-related cues and content-related representation, which can make degradation-aware modulation less discriminative. The second is Spatial Ambiguity, namely their entanglement in the spatial dimension, which can weaken the spatial selectivity of restoration responses.

Motivated by these observations, we propose a Dual-Ambiguity Rectification Network (DAR-Net) for all-in-one image restoration. DAR-Net reduces degradation-content entanglement in both the channel and spatial dimensions through three components. It first employs a Degradation Archetype Representation (DAR) module to construct an archetype-based degradation state, which serves as a degradation prior for subsequent rectification. Based on this degradation state, the Semantic Ambiguity Rectification (SeAR) module alleviates channel-wise entanglement through an Archetype-Guided Prompt Generator (AGPG) and a DegradationAware Prompt Integrator (DAPI). Specifically, AGPG first generates a base prompt and then refines it via archetype-guided channel routing conditioned on the degradation state, yielding a degradationaware prompt. DAPI subsequently injects this rectified prompt into the restoration process for degradation-conditioned feature modulation. In addition, the Spatial Ambiguity Rectification (SpAR) module reduces spatial entanglement via Orthogonal Subspace Rectification (OSR), which treats the rectified prompt as a degradation-aware representation and derives a complementary content-related feature from the latent representation, encouraging the two to occupy orthogonal subspaces. Together, these designs aim to better distinguish removal-related and preservation-related cues in image restoration. The main contributions are summarized as follows: • Dual-ambiguity rectification framework: We propose DAR-Net for AIR, which mitigates degradation-content entanglement in both the channel and spatial dimensions. • Semantic ambiguity rectification: We introduce DAR to model degradation states and SeAR to generate degradationaware prompts for channel-wise semantic rectification. • Spatial ambiguity rectification: We introduce SpAR with OSR to separate degradation-related and complementary content-related representations in the spatial dimension. • Comprehensive validation and results: Extensive experiments verify that DAR-Net consistently achieves the best overall performance on standard all-in-one restoration benchmarks and generalizes favorably to mixed and realworld degradations.

2 Related Work 2.1 Task-Specific Image Restoration Traditional image restoration methods are typically designed for a single degradation type, such as denoising [6, 22, 32], deblurring [9, 18, 45], deraining [7, 40, 46, 49], and dehazing [3, 8, 33, 34]. Their goal is to learn a direct mapping from degraded images to clean images under a predefined degradation setting. With the development of restoration architectures, many methods have gradually moved from heavily customized task-specific designs toward more general restoration backbones. Representative models such as IPT [4], SwinIR [22], Uformer [43], Restormer [52], NAFNet [6], and MAXIM [38] exemplify this trend by improving restoration quality through stronger feature modeling and broader contextual interaction. Recent image super-resolution methods further explore hybrid Mamba–Transformer modeling to improve efficient long-range interaction [23]. While task-specific image restoration methods are

 

...

Degradation-Aware Prompt Integrator (DAPI)

c

1×1

Transformer Block

SeAR

Transformer Block

c

1×1

Transformer Block

SeAR

1×1 D3×3

1×1 D3×3

1×1 D3×3

Spatial Ambiguity Rectification (SpAR) 1×1

1×1

1×1

1×1 

Semantic Ambiguity Rectification DAPI AGPG SpAR

Transformer Block





   

Archetype-Guided Prompt Generator (AGPG)

/

/

3×3

Transformer Block

Softmax

...

Activation function

3×3

Softmax

FC

Transformer Block

FC

 

GAP

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

c

GAP

Degradation Archetype Representation

CondNet 3×3 3×3 3×3

Transformer Block

  

Softmax 

Sigmoid

Average GAP Global Pooling

c Channel-wise Concatenation Element-wise Addition Matrix Multiplication

FC Linear Layer

Element-wise Product

1×1 1×1 Convolution

  

FC

3×3

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

Archetype-Guided Channel Routing

   

Downsampling Upsampling Loss Supervision

3×3 3×3 Convolution

Depthwise D3×3 3×3 Convolution

Figure 2: Architecture of the Dual-Ambiguity Rectification Network (DAR-Net). Built upon a hierarchical U-shaped Transformer, DAR-Net mitigates degradation-content entanglement through three components: (1) DAR, which constructs a degradation state via simplex-constrained degradation archetype representation; (2) SeAR, which comprises AGPG to generate a degradationaware prompt from the degradation state and DAPI to integrate this prompt into the decoder for degradation-conditioned restoration; and (3) SpAR, which is implemented by OSR to regularize the degradation-aware prompt and the corresponding content feature toward orthogonal spatial subspaces under the LOSR constraint. effective for individual degradation types, their reliance on separate degradation-specific models leads to poor scalability and limits their applicability in unified real-world restoration settings.

better distinguish what should be removed from what should be preserved.

3 Methods 3.1 Mathematical Preliminaries 2.2

All-in-one Image Restoration

To improve scalability in unified real-world restoration scenarios, all-in-one image restoration aims to handle diverse degradations with a single unified model, making it more suitable for practical settings where the degradation type is unknown or mixed. Existing methods mainly differ in how degradation information is incorporated into a shared restoration pipeline. Representative early directions include degradation representation learning, where AirNet learns contrastive degradation representations [19]; prompt-based conditioning, where PromptIR [30], InstructIR [10], and UP-Restorer [24] inject learned prompts or instructions into the restoration network; and multimodal guidance, where DACLIP [25] and MPerceiver [1] leverage large-scale vision-language priors for restoration. Subsequent works further improve unified restoration either by strengthening degradation modeling and shared representation learning [5, 17, 36, 47, 54] or by introducing degradation-specialized experts to better handle diverse degradation patterns [41, 51, 56]. Despite these different designs, most existing methods still rely on degradation cues to condition, organize, or route shared features within a unified restoration network. In contrast, our method focuses on reducing degradation-content entanglement during feature modulation, so that the model can

Simplices and barycentric coordinates. Let V = {𝑣 1, . . . , 𝑣 𝐾 } ⊂ R𝐷 with 𝐾 ≤ 𝐷 + 1. If 𝑣 1, . . . , 𝑣 𝐾 are affinely independent, then their convex hull (𝐾 ) 𝐾 ∑︁ ∑︁ conv(V) = 𝛼𝑘 𝑣𝑘 𝛼𝑘 ≥ 0, 𝛼𝑘 = 1 (1) 𝑘=1

𝑘=1

forms a geometric (𝐾 − 1)-simplex. Equivalently, letting ( ) 𝐾 ∑︁ Δ𝐾 −1 = 𝛼 ∈ R𝐾 𝛼𝑘 ≥ 0, 𝛼𝑘 = 1 ,

(2)

𝑘=1

the affine map 𝑇 : Δ𝐾 −1 → conv(V),

𝑇 (𝛼) =

𝐾 ∑︁

𝛼𝑘 𝑣 𝑘

(3)

𝑘=1

is bijective. Hence every point 𝑧 ∈ conv(V) admits a unique coefficient vector 𝛼 ∈ Δ𝐾 −1 such that 𝑧=

𝐾 ∑︁

𝛼𝑘 𝑣 𝑘 ,

𝑘=1

where 𝛼 is the barycentric coordinate of 𝑧 with respect to V.

(4)

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

Liu et al.

Orthogonal decomposition. Let H be a finite-dimensional innerproduct space with inner product ⟨·, ·⟩. For any subspace S ⊆ H , its orthogonal complement is

descriptor is projected to a 𝐾-dimensional score vector and normalized by a softmax operator to produce the mixture coefficients

S ⊥ = {𝑦 ∈ H | ⟨𝑥, 𝑦⟩ = 0, ∀𝑥 ∈ S} .

(5)

where W ∈ R𝐾 ×𝐶 , b ∈ R𝐾 , and 𝐾 denotes the number of degra-

By the projection theorem, every 𝑧 ∈ H admits a unique orthogonal decomposition

dation archetypes. Here, each entry of 𝜶 quantifies the contribution of one archetype to the degradation mixture, and thus 𝜶 serves as the simplex coordinate vector described in § 3.1. Let A = [a1, a2, . . . , a𝐾 ] ∈ R𝐶 ×𝐾 denote a learnable degradation archetype matrix, where each column a𝑘 ∈ R𝐶 is an archetypal degradation vector. The degradation state is then constructed as

𝑧 = 𝑃 S 𝑧 + 𝑃 S ⊥ 𝑧,

𝑃 S 𝑧 ∈ S, 𝑃 S ⊥ 𝑧 ∈ S ⊥,

(6)

which implies ∥𝑧 ∥ 22 = ∥𝑃 S 𝑧∥ 22 + ∥𝑃 S ⊥ 𝑧∥ 22 .

(7)

More generally, if S𝐴 ⊥ S𝐵 , then for any 𝑥𝐴 ∈ S𝐴 and 𝑥 𝐵 ∈ S𝐵 , ⟨𝑥𝐴 , 𝑥 𝐵 ⟩ = 0,

∥𝑥𝐴 + 𝑥 𝐵 ∥ 22 = ∥𝑥𝐴 ∥ 22 + ∥𝑥 𝐵 ∥ 22 .

(9)

Moreover, 𝐶 ∑︁ 𝐶 ∑︁

⟨𝑥𝑖 , 𝑦 𝑗 ⟩ 2,

(11)

𝑖=1 𝑗=1

which measures the total pairwise interaction between the two subspaces and vanishes exactly under orthogonality.

3.2

𝐾 ∑︁

𝛼𝑘 a𝑘 = A𝜶 ,

Sdeg ∈ R𝐶 ,

(13)

(14)

𝑘=1

Therefore, the row-generated subspaces of 𝑋𝐴 and 𝑋𝐵 are orthogonal if and only if 𝑋𝐴 𝑋𝐵⊤ = 0𝐶 ×𝐶 . (10) ∥𝑋𝐴 𝑋𝐵⊤ ∥ 2𝐹 =

Sdeg =

(8)

In the Euclidean case, let 𝑋𝐴 , 𝑋𝐵 ∈ R𝐶 ×𝑁 , and denote their row vectors by {𝑥𝑖 }𝐶𝑖=1 and {𝑦 𝑗 }𝐶𝑗=1 . Then the (𝑖, 𝑗)-th entry of the crossGram matrix satisfies [𝑋𝐴 𝑋𝐵⊤ ] 𝑖 𝑗 = ⟨𝑥𝑖 , 𝑦 𝑗 ⟩.

𝜶 ∈ R𝐾 ,

𝜶 = Softmax(Wg + b),

Overview

Built upon a hierarchical U-shaped Transformer backbone, DARNet mitigates two ambiguities in AIR (Fig. 2). Specifically, we design a rectification pipeline: (1) Degradation Archetype Representation (DAR) (§ 3.3) extracts a global degradation descriptor and maps it to simplex-constrained mixture coefficients to construct a degradation state; (2) Semantic Ambiguity Rectification (SeAR) (§ 3.4) uses this state to rectify prompt channels, yielding a degradation-aware prompt that is further integrated into the decoder for degradation-conditioned feature modulation; and (3) Spatial Ambiguity Rectification (SpAR) (§ 3.5) regularizes the degradation-aware prompt and the corresponding content feature at the deepest decoder stage toward orthogonal subspaces. Finally, the restored image is reconstructed with a global residual connection, and the training objective is given in § 3.6.

which is a simplex-constrained convex combination of the learned archetypes. As a result, Sdeg lies in the convex hull of the archetypes and serves as the structured degradation representation used in the subsequent rectification modules.

3.4

Semantic Ambiguity Rectification

The SeAR mitigates semantic ambiguity, i.e., the channel-wise entanglement between degradation representation and content representation. SeAR consists of an Archetype-Guided Prompt Generator (AGPG) and a Degradation-Aware Prompt Integrator (DAPI). Specifically, SeAR first uses the degradation state Sdeg to generate a degradation-aware prompt, and then integrates this prompt into the stage-wise decoding process. Before the deepest decoder stage, SeAR further derives a content feature by residual decomposition. 3.4.1 Archetype-Guided Prompt Generator (AGPG). Formally, let 𝐹𝑖𝑛(𝑙 ) ∈ R𝐶𝑙 ×𝐻𝑙 ×𝑊𝑙 denote the input feature before the 𝑙-th decoder stage (𝑙 ∈ {1, 2, 3}). AGPG aims to construct a degradation-aware prompt by combining two sources of information: the currentstage feature, which provides input-adaptive prompt cues, and the degradation state Sdeg , which provides structured degradation prior. To this end, we first synthesize a base prompt from a set of 𝑀 learnable prompt tensors P (𝑙 ) = {𝑃1(𝑙 ) , . . . , 𝑃𝑀(𝑙 ) }, where each 𝑃𝑚(𝑙 ) ∈ R𝐶𝑙 ×𝐻𝑙 ×𝑊𝑙 . Specifically, we predict an input-dependent mixture weight vector from the globally pooled feature and use it to aggregate the prompt tensors:   w (𝑙 ) = Softmax FC𝑝(𝑙 ) (G(𝐹𝑖𝑛(𝑙 ) )) , ! 𝑀 (15) ∑︁ (𝑙 ) (𝑙 ) (𝑙 ) (𝑙 ) 𝐹𝑝 = Conv3×3 𝑤𝑚 𝑃𝑚 . 𝑚=1

3.3

Degradation Archetype Representation

The DAR module implements the simplex-constrained parameterization introduced in § 3.1 and provides a structured degradation representation for subsequent rectification. Given an input image 𝐼𝑖𝑛 ∈ R3×𝐻 ×𝑊 , where 𝐻 and 𝑊 denote the input height and width, respectively, we first employ a lightweight conditioning network Cond(·), implemented by stacked 3 × 3 convolutional layers, to extract degradation-sensitive features: 𝐹𝑐𝑜𝑛𝑑 = Cond(𝐼𝑖𝑛 ),

𝐹𝑐𝑜𝑛𝑑 ∈ R𝐶 ×𝐻 ×𝑊 ,

(12)

where 𝐶 is the channel dimension. We then use global average pooling to 𝐹𝑐𝑜𝑛𝑑 to obtain a degradation descriptor g ∈ R𝐶 . This

We then inject the degradation prior by mapping Sdeg to a channel-wise routing vector and using it to rectify the base prompt:   (𝑙 ) ŝ (𝑙 ) = 𝜎 FC𝑠(𝑙 ) (Sdeg ) , 𝐹𝑑𝑝 = 𝐹𝑝(𝑙 ) ⊙ ŝ (𝑙 ) , (16) where ŝ (𝑙 ) ∈ R𝐶𝑙 ×1×1 and ⊙ denotes broadcast multiplication over spatial dimensions. In this way, channels that are more consistent with the inferred degradation state are emphasized, while prompt responses unrelated to the current degradation are suppressed. At the deepest decoder stage, we further derive a complementary content-related feature by residual subtraction, (1) 𝐹𝑐(1) = 𝐹𝑖𝑛(1) − 𝐹𝑑𝑝 ,

(17)

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration (1) and forward the pair (𝐹𝑑𝑝 , 𝐹𝑐(1) ) to SpAR for subsequent spatial ambiguity rectification.

3.4.2 Degradation-Aware Prompt Integrator (DAPI). DAPI injects the degradation-aware prompt into the decoder through channelwise attention. For a unified formulation, we define ( (1) ( 𝐹˜𝑑𝑝 , 𝑙 = 1, 𝐹˜𝑐(1) , 𝑙 = 1, (𝑙 ) (𝑙 ) 𝐹𝑞 = 𝐹𝑘𝑣 = (18) (𝑙 ) 𝐹𝑑𝑝 , 𝑙 > 1, 𝐹𝑖𝑛(𝑙 ) , 𝑙 > 1. (1) Here, the deepest decoder stage uses the SpAR-rectified pair ( 𝐹˜𝑑𝑝 , (1) 𝐹˜𝑐 ), while later stages directly use the degradation-aware prompt and the current-stage input feature. We then project these inputs into query, key, and value tensors:

𝑄 (𝑙 ) = Φ𝑄(𝑙 ) (𝐹𝑞(𝑙 ) ),

(𝑙 ) 𝐾 (𝑙 ) = Φ𝐾(𝑙 ) (𝐹𝑘𝑣 ),

(𝑙 ) 𝑉 (𝑙 ) = Φ𝑉(𝑙 ) (𝐹𝑘𝑣 ), (19) where Φ𝑄(𝑙 ) , Φ𝐾(𝑙 ) , and Φ𝑉(𝑙 ) are three independent projection blocks, each implemented by a 1 × 1 convolution followed by a depth-wise 3 × 3 convolution. Let 𝑁𝑙 = 𝐻𝑙 𝑊𝑙 denote the number of spatial locations at the 𝑙-th stage. After reshaping 𝑄 (𝑙 ) , 𝐾 (𝑙 ) , and 𝑉 (𝑙 ) to R𝐶𝑙 ×𝑁𝑙 , we compute channel-wise attention as ! ⊤ 𝑄 (𝑙 ) 𝐾 (𝑙 ) , 𝐴 (𝑙 ) ∈ R𝐶𝑙 ×𝐶𝑙 , 𝐴 (𝑙 ) = Softmax (20) 𝜏

where 𝜏 is a learnable temperature parameter and the softmax (𝑙 ) is applied  row-wise.  The stage output is then obtained as 𝐹𝑜𝑢𝑡 = Reshape 𝐴 (𝑙 ) 𝑉 (𝑙 ) .

3.5

Spatial Ambiguity Rectification

(1) SeAR produces a degradation-aware prompt 𝐹𝑑𝑝 and a comple-

mentary content-related feature 𝐹𝑐(1) . Although residual decomposition separates them coarsely, their spatial responses may still remain entangled, leading to spatial ambiguity. To further separate degradation-related and content-related spatial responses before prompt integration, we introduce an Orthogonal Subspace Rectification (OSR) strategy, which encourages the two representations (1) to lie in orthogonal subspaces. Specifically, we first transform 𝐹𝑑𝑝 and 𝐹𝑐(1) with two learnable mappings Ω𝑝 and Ω𝑐 while preserving their spatial resolution: (1) (1) 𝐹˜𝑑𝑝 = Ω𝑝 (𝐹𝑑𝑝 ),

𝐹˜𝑐(1) = Ω𝑐 (𝐹𝑐(1) ),

(21)

(1) ˜ (1) where 𝐹˜𝑑𝑝 , 𝐹𝑐 ∈ R𝐶1 ×𝐻1 ×𝑊1 . In practice, each mapping is implemented by a 1 × 1 convolution, a point-wise nonlinearity, and another 1 × 1 convolution. These learnable transformations allow the model to project the two features into a space where orthogonality can be imposed more effectively. To instantiate the orthogonality constraint, let 𝑁 1 = 𝐻 1𝑊1 , and let m̃𝑑𝑝,𝑘 , m̃𝑐,𝑘 ∈ R𝑁1 denote the flattened spatial maps of the 𝑘(1) th channel of 𝐹˜𝑑𝑝 and 𝐹˜𝑐(1) , respectively. Since directly shrinking feature magnitudes could trivially reduce their interaction, we first normalize each channel vector: m̃𝑑𝑝,𝑘 m̃𝑐,𝑘 m̂𝑑𝑝,𝑘 = , m̂𝑐,𝑘 = , (22) ∥ m̃𝑑𝑝,𝑘 ∥ 2 + 𝜖 ∥ m̃𝑐,𝑘 ∥ 2 + 𝜖

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

where 𝜖 is a constant for numerical stability. We then stack the normalized vectors row-wise into 𝑋𝑑𝑝 , 𝑋𝑐 ∈ R𝐶1 ×𝑁1 . In this form, the row spaces of 𝑋𝑑𝑝 and 𝑋𝑐 represent the spatial response subspaces of the degradation-aware and content features, respectively. According to the orthogonal decomposition in § 3.1, two row-generated subspaces are orthogonal if and only if their cross-Gram matrix vanishes, i.e., 𝑋𝑑𝑝 𝑋𝑐⊤ = 0. We therefore define the OSR loss as LOSR = ∥𝑋𝑑𝑝 𝑋𝑐⊤ ∥ 2𝐹 =

𝐶 1 ∑︁ 𝐶1 ∑︁

⟨m̂𝑑𝑝,𝑖 , m̂𝑐,𝑗 ⟩ 2 .

(23)

𝑖=1 𝑗=1

Minimizing LOSR suppresses all pairwise inner-product interactions between the channel-wise spatial responses of the two features, thereby encouraging their row-generated subspaces to be orthogo(1) nal. The resulting rectified representations 𝐹˜𝑑𝑝 and 𝐹˜𝑐(1) are then fed into DAPI at the deepest decoder stage.

3.6

Training Objective

DAR-Net is trained with a pixel-wise reconstruction loss and the orthogonality regularization introduced in § 3.5. The overall loss is Ltotal = Lrec + 𝜆 LOSR,

(24)

where 𝜆 is a balancing coefficient. We adopt the L1 loss between the restored image 𝐼 out and the ground-truth image 𝐼 gt as the reconstruction loss: 𝑁 1 ∑︁ (𝑖 ) 𝐼 − 𝐼 (𝑖 ) , (25) Lrec = 𝑁 𝑖=1 out gt where 𝑁 denotes the total number of image elements. The term LOSR , defined in Eq. (23), regularizes the degradation-aware prompt and the corresponding content feature at the deepest decoder stage by encouraging their spatial response subspaces to be orthogonal.

4

Experiments

We evaluate DAR-Net under both three-degradation (3D) and fivedegradation (5D) all-in-one restoration settings. Beyond standard evaluation, we further assess its generalization ability on mixed degradations, and real-world images. We compare DAR-Net with representative restoration methods, including Restormer [52], AirNet [19], PromptIR [30], InstructIR [10], DiffUIR [58], AdaIR [11], VLU-Net [53], MoCE-IR [51], DFPIR [37], ClearAIR [57], MIRAGE [31], WeatherDiff [29], WGWS-Net [59], OneRestore [13], TransWeather [39] and Histoformer [35]. We use PSNR and SSIM [42] for pixel-wise fidelity evaluation, and LPIPS [55] and FID [15] for perceptual quality assessment. Unless otherwise specified, the results of the compared methods are taken from their original papers or from the survey [16]. The best and second-best results are highlighted in bold and underlined, respectively.

4.1

Experimental Settings

Datasets. For the 3D setting, we train on BSD400 [2] and WED [26], and evaluate denoising on BSD68 [27] with Gaussian noise levels 𝜎 ∈ {15, 25, 50}. Rain100L [48] and SOTS [20] are used for deraining and dehazing, respectively. For the 5D setting, we further include GoPro [28] for deblurring and LOL [44] for low-light enhancement. For mixed-degradation evaluation, we use CDD-11 [13]. For realworld evaluation, we adopt the WeatherBench [12].

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

Liu et al.

Table 1: Quantitative comparison on the 3D all-in-one restoration benchmark. We reported PSNR/SSIM. Method

Dehazing

Deraining

Denoising on BSD68

SOTS-Outdoor

Rain100L

𝜎 = 15

𝜎 = 25

𝜎 = 50

Venue

Average

Restormer [52] AirNet [19] PromptIR [30] InstructIR [10] DiffUIR [58] AdaIR [11] VLU-Net [53] MoCE-IR [51] DFPIR [37] ClearAIR [57] MIRAGE [31]

CVPR’22 CVPR’22 NeurIPS’23 ECCV’24 CVPR’24 ICLR’25 CVPR’25 CVPR’25 CVPR’25 AAAI’26 ICLR’26

27.78/0.958 27.94/0.962 30.58/0.974 30.22/0.959 30.18/0.973 31.06/0.980 30.71/0.980 31.34/0.979 31.87/0.980 31.08/0.981 31.86/0.981

33.78/0.958 34.90/0.968 36.37/0.972 37.98/0.978 36.78/0.973 38.64/0.983 38.93/0.984 38.57/0.984 38.65/0.982 38.61/0.984 38.94/0.985

33.72/0.930 33.92/0.933 33.98/0.933 34.15/0.933 33.94/0.932 34.12/0.935 31.13/0.935 34.11/0.932 34.12/0.935 34.18/0.935 34.12/0.935

30.67/0.865 31.26/0.888 31.31/0.888 31.52/0.890 31.26/0.887 31.45/0.892 31.48/0.892 31.45/0.888 31.47/0.893 31.50/0.891 31.46/0.891

27.63/0.792 28.00/0.797 28.06/0.799 28.30/0.804 28.04/0.797 28.19/0.802 28.23/0.804 28.18/0.800 28.25/0.806 28.31/0.804 28.19/0.803

30.75/0.901 31.20/0.910 32.06/0.913 32.43/0.913 32.04/0.912 32.69/0.918 32.10/0.919 32.73/0.917 32.88/0.919 32.74/0.919 32.91/0.919

DAR-Net (Ours)

-

31.93/0.984

39.15/0.986

34.21/0.936

31.58/0.895

28.37/0.808

33.05/0.922

Table 2: Quantitative comparison on the 5D all-in-one restoration benchmark. We reported PSNR/SSIM. Venue

Dehazing SOTS-Outdoor

Deraining Rain100L

Denoising BSD68 (𝜎 = 25)

Deblurring GoPro

Low-Light LOL

Average

Restormer [52] AirNet [19] PromptIR [30] InstructIR [10] DiffUIR [58] AdaIR [11] VLU-Net [53] MoCE-IR [51] DFPIR [37] ClearAIR [57] MIRAGE [31]

CVPR’22 CVPR’22 NeurIPS’23 ECCV’24 CVPR’24 ICLR’25 CVPR’25 CVPR’25 CVPR’25 AAAI’26 ICLR’26

24.09/0.927 21.04/0.884 26.54/0.949 27.10/0.956 29.47/0.965 30.53/0.978 30.84/0.980 30.48/0.974 31.64/0.979 30.12/0.978 31.45/0.980

34.81/0.960 32.98/0.951 36.37/0.970 36.84/0.973 35.98/0.968 38.02/0.981 38.54/0.982 38.04/0.982 37.62/0.978 38.20/0.982 38.92/0.982

31.49/0.884 30.91/0.882 31.47/0.886 31.40/0.887 31.02/0.885 31.35/0.889 31.43/0.891 31.34/0.887 31.29/0.889 31.53/0.888 31.41/0.892

27.22/0.829 24.35/0.781 28.71/0.881 29.40/0.886 27.50/0.845 28.12/0.858 27.46/0.840 30.05/0.899 28.82/0.873 29.67/0.887 28.10/0.858

20.41/0.806 18.18/0.735 22.68/0.832 23.00/0.836 22.32/0.826 23.00/0.845 22.29/0.833 23.00/0.852 23.82/0.843 22.83/0.846 23.59/0.858

27.60/0.881 25.49/0.846 29.15/0.904 29.55/0.907 29.25/0.898 30.20/0.910 30.11/0.905 30.58/0.919 30.64/0.913 30.47/0.916 30.68/0.914

DAR-Net (Ours)

-

31.67/0.981

38.34/0.983

31.46/0.892

29.77/0.889

23.86/0.860

31.02/0.921

Method

Table 3: Quantitative Comparison on CDD-11 [13] Dataset. We reported PSNR/SSIM metrics. L

H

R

S

L+H

L+R

Double L+S

H+R

H+S

AirNet [19] PromptIR [30] WeatherDiff [29] WGWS-Net [59] OneRestore [13] AdaIR [11] MoCE-IR [51]

24.83/0.778 26.32/0.805 23.58/0.763 24.39/0.774 26.48/0.826 26.88/0.821 27.26/0.824

24.21/0.951 26.10/0.969 21.99/0.904 27.90/0.982 32.52/0.990 31.60/0.987 32.66/0.990

26.55/0.891 31.56/0.946 24.85/0.885 33.15/0.964 33.40/0.964 33.84/0.962 34.31/0.970

26.79/0.919 31.53/0.960 24.80/0.888 34.43/0.973 34.31/0.973 34.65/0.974 35.91/0.980

23.23/0.779 24.49/0.789 21.83/0.756 24.27/0.800 25.79/0.822 25.69/0.811 26.24/0.817

22.82/0.710 25.05/0.771 22.69/0.730 25.06/0.772 25.58/0.799 25.90/0.793 26.25/0.800

23.29/0.723 24.51/0.761 22.12/0.707 24.60/0.765 25.19/0.789 25.69/0.783 26.04/0.793

22.21/0.868 24.54/0.924 21.25/0.868 27.23/0.955 29.99/0.957 29.38/0.955 29.93/0.964

23.29/0.901 23.70/0.925 21.99/0.868 27.65/0.960 30.21/0.964 28.95/0.961 30.19/0.970

DAR-Net (Ours)

27.54/0.834 34.21/0.991 34.96/0.972 36.43/0.981 26.74/0.830 26.63/0.812 26.58/0.806 31.25/0.968 31.28/0.971 25.53/0.801 25.83/0.799 29.73/0.888

Method

Single

Implementation details. Our model is built on a hierarchical U-shaped Transformer backbone. We use a 4-level encoder-decoder architecture with [4, 6, 6, 8] Transformer blocks from level-1 to level-4. We optimize the network using AdamW with an initial learning rate of 2 × 10−4 and 𝛽 1 = 0.9, 𝛽 2 = 0.99. The learning rate is decayed to 1 × 10−8 using cosine annealing with five cycles. The model is trained for 450K and 650K iterations under the 3D and 5D settings, respectively, with a batch size of 32. During training, input images are randomly cropped into 128 × 128 patches and

Triple L+H+R L+H+S

Average

21.80/0.708 23.74/0.752 21.23/0.716 23.90/0.772 24.78/0.788 24.82/0.778 25.41/0.789

23.75/0.814 25.90/0.850 22.49/0.799 26.96/0.863 28.47/0.878 28.40/0.873 29.05/0.881

22.24/0.725 23.33/0.747 21.04/0.698 23.97/0.771 24.90/0.791 25.04/0.778 25.39/0.790

augmented by random flipping and rotation. The number of degradation archetypes 𝐾 = 16, the temperature parameter 𝜏 = 0.07, and the loss weight 𝜆 = 0.1. All experiments are implemented in PyTorch and conducted on 2 NVIDIA A800 GPUs.

4.2

Main Results

Three-Degradation Evaluation. As shown in Tab. 1, DAR-Net achieves the best overall performance under the 3D setting, with an average PSNR/SSIM of 33.05/0.922. It consistently ranks first

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

Table 4: Quantitative Comparison of different methods on WeatherBench [12] Dataset. Method

Dehazing PSNR↑ SSIM↑ LPIPS↓

FID↓

Deraining PSNR↑ SSIM↑ LPIPS↓

FID↓

Desnowing PSNR↑ SSIM↑ LPIPS↓

FID↓

Average PSNR↑ SSIM↑ LPIPS↓

FID↓

AirNet [19] TransWeather [39] PromptIR [30] WGWS-Net [59] Histoformer [35] AdaIR [11] DiffUIR [58]

19.27 18.13 19.50 11.78 15.82 21.39 20.96

0.645 0.621 0.658 0.532 0.597 0.680 0.695

0.3829 0.3970 0.3751 0.5351 0.4371 0.3506 0.3550

134.09 123.21 113.55 152.76 128.34 110.07 127.54

31.56 28.59 32.51 34.77 28.87 32.81 33.78

0.912 0.880 0.915 0.939 0.876 0.918 0.931

0.2236 0.2638 0.1980 0.1168 0.2785 0.1916 0.1720

125.54 149.66 111.69 60.99 152.42 109.41 86.96

20.58 24.06 26.35 19.39 23.88 26.87 27.87

0.737 0.754 0.804 0.721 0.769 0.806 0.844

0.2912 0.2250 0.1951 0.2481 0.2252 0.1790 0.1619

138.57 102.99 84.12 128.56 105.82 73.48 68.99

23.80 23.59 26.12 21.98 22.86 27.02 27.54

0.764 0.752 0.792 0.731 0.747 0.801 0.823

0.2992 0.2953 0.2561 0.3000 0.3136 0.2404 0.2296

132.73 125.29 103.12 114.10 128.86 97.65 94.50

DAR-Net (Ours)

23.44

0.732

0.3257

108.35

35.48

0.941

0.1663

82.65

29.37

0.872

0.1569

65.28

29.43

0.848

0.2163

85.43

Input

InstructIR

MoCE-IR

DFPIR

DAR-Net

Ground truth

Figure 3: Qualitative comparison of 3D all-in-one restoration results. Table 5: Ablation on the key Table 6: Ablation on content components of DAR-Net. feature construction. DAR

SeAR

SpAR

PSNR

SSIM

✗ ✓ ✓ ✓

✗ ✗ ✓ ✓

✗ ✗ ✗ ✓

29.15 29.22 30.65 31.02

0.904 0.905 0.916 0.921

Method

PSNR

SSIM

No decomposition Gated suppression Residual subtraction

30.78 30.84 31.02

0.917 0.918 0.921

on dehazing, deraining, and all three denoising levels, demonstrating strong and balanced restoration performance across different degradation types. Compared with the second-best method, DARNet improves the average PSNR by 0.14 dB. These results indicate that DAR-Net can more effectively handle degradation-content entanglement in the all-in-one restoration setting, leading to both stronger degradation removal and better content preservation. Five-Degradation Evaluation. DAR-Net achieves the best overall performance under the 5D setting, with an average PSNR/SSIM of 31.02/0.921 (Tab. 2). Compared with the second-best method, DARNet improves the average PSNR by 0.34 dB. Although it is not the best-performing method on every task, DAR-Net achieves the best results on dehazing and low-light enhancement while remaining competitive on deraining, denoising, and deblurring. These results indicate that DAR-Net maintains a strong overall balance across diverse degradation types in the more challenging 5D setting. Mixed-degradation Evaluation. As shown in Tab. 3, DAR-Net achieves the best results on all CDD-11 [13] subsets, covering single, double, and triple degradations. It obtains the highest average PSNR/SSIM of 29.73/0.888, surpassing the second-best method by

0.68 dB in PSNR and 0.007 in SSIM. The consistent gains across increasingly complex degradation combinations verify the effectiveness of DAR-Net for mixed-degradation restoration. Real-world Evaluation. Tab. 4 reports the quantitative comparison on the WeatherBench [12] dataset. DAR-Net achieves the best overall performance, with particularly clear advantages on dehazing and desnowing. On deraining, DAR-Net also attains the best PSNR and SSIM, while remaining competitive in LPIPS and FID. These results demonstrate that DAR-Net generalizes effectively to real-world weather degradations and yields restoration results with improved fidelity and perceptual quality. Qualitative Results. Fig. 3 presents qualitative results under the 3D setting. DAR-Net removes degradations more thoroughly across diverse restoration tasks while better preserving natural image structures. For example, in the deraining case, our result is free of visible rain-streak residue, whereas competing methods still retain noticeable artifacts. In the denoising example with 𝜎 = 25, DARNet suppresses noise effectively without mistakenly removing the cloud structures in the sky.

4.3

Ablation Study

Effect of Key Components. Tab. 5 reports only the average results under the 5D setting for clarity. The full DAR-Net achieves the best performance, validating the effectiveness of the overall design and the complementarity of its three components. DAR provides a structured degradation prior, while SeAR yields more substantial

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

Liu et al.

Average Simplex Coefficients Across Degradation Types Degradation Type

rain 0.03 0.13 0.01 0.03 0.01 0.14 0.02 0.03 0.03 0.02 0.01 0.02 0.03 0.03 0.32 0.14 haze 0.26 0.01 0.03 0.02 0.20 0.03 0.02 0.03 0.26 0.03 0.01 0.03 0.02 0.02 0.03 0.02 blur 0.02 0.02 0.14 0.14 0.31 0.02 0.01 0.02 0.02 0.03 0.14 0.02 0.02 0.03 0.04 0.01

low_light 0.17 0.03 0.02 0.03 0.03 0.02 0.14 0.11 0.17 0.11 0.03 0.03 0.03 0.01 0.03 0.03

noise_25 0.03 0.03 0.04 0.02 0.04 0.03 0.02 0.03 0.03 0.03 0.04 0.14 0.14 0.13 0.22 0.03 a0 a1 a2 a3 a4 a5 a6 a7 a8 a9 a10 a11 a12 a13 a14 a15 Archetype Index

F1

0.30 0.25 0.20 0.15 0.10 0.05

Input

Figure 4: Visualization of the average simplex coefficients learned by DAR for different degradation types. Distinct degradations exhibit different archetype activation patterns, while related degradations still share partial archetypes.

Gap=0.11

Higher is better

Ours

|F3-F1|

F1

Input

W/O Fdp

F2

W/O Fc |F3-F1|

|F3-F2|

Gap=0.16

Ours

Figure 5: Intra-/inter-class similarity distributions of prompt features before and after SeAR. SeAR increases intra-class similarity while decreasing inter-class similarity. gains by alleviating channel-wise semantic ambiguity. SpAR further improves the performance, and the combination of all three components leads to the best overall result. Effect of Content Feature Construction. We analyze how to construct the content feature 𝐹𝑐 in SpAR while keeping DAR, SeAR, and SpAR enabled. Specifically, we compare three variants: no decomposition (𝐹𝑐 = 𝐹𝑖𝑛 ), gated suppression (𝐹𝑐 = 𝐹𝑖𝑛 ⊙ (1 − 𝜎 (𝐹𝑑𝑝 ))), and residual subtraction (𝐹𝑐 = 𝐹𝑖𝑛 − 𝐹𝑑𝑝 ). Here, 𝐹𝑖𝑛 denotes the mixed input feature and 𝐹𝑑𝑝 denotes the degradation-related feature in § 3.5. As shown in Tab. 6, the residual formulation achieves the best performance, suggesting that explicitly subtracting degradationrelated information is more effective for isolating content. Effect of SpAR Placement. Applying SpAR at the deepest decoder stage yields the best performance; detailed placement results are provided in the supplementary material.

4.4

|F3-F2|

Con. Difference Map

Deg. Difference map

F3

Lower is better

W/O Fc

W/O Fdp F3

F2

Analysis

Analysis of DAR. Fig. 4 shows that different degradations activate distinct archetype combinations, while related degradations share partial archetypes, indicating structured yet transferable degradation representations. Analysis of the archetype number 𝐾 is provided in the supplementary material. Analysis of SeAR. As shown in Fig. 5, SeAR increases intra-class prompt similarity from 0.63 to 0.74 and decreases inter-class similarity from 0.58 to 0.42, demonstrating improved degradation discrimination. Analysis of SpAR. As shown in Fig. 6, removing 𝐹𝑑𝑝 leaves residual degradations, whereas removing 𝐹𝑐 damages structural content, confirming their complementary roles in degradation removal and content preservation. Model Complexity and Efficiency. As shown in Tab. 7, DAR-Net has 35.5M parameters and 771G FLOPs, which are comparable to

Deg. Difference map

Con. Difference Map

Figure 6: Analysis of SpAR. F1, F2, and F3 denote the results without the degradation-related feature, without the contentrelated feature, and with the full model, respectively. The difference maps indicate distinct roles of the two features in degradation removal and content preservation. Table 7: Comparison of model complexity and inference efficiency. FLOPs and latency are measured on an input image of size 720 × 480 on a single NVIDIA A800 GPU. Method Params. FLOPs Latency CPU Memory GPU Memory

PromptIR

AdaIR

MoCE-IR

DFPIR

DAR-Net

34.1M 752G 187ms 4454M 3324M

28.8M 786G 239ms 4336M 3117M

25.4M 493G 161ms 4389M 1445M

31M + 63M 885G 193ms 4369M 3459M

35.5M 771G 165ms 4200M 3162M

existing methods. Despite slightly higher complexity than some lightweight baselines, DAR-Net still achieves competitive inference efficiency, with lower latency than PromptIR, AdaIR, and DFPIR. Compared with MoCE-IR, DAR-Net incurs moderate additional overhead while providing stronger restoration performance, demonstrating a favorable efficiency-performance trade-off.

5

Conclusion

In this paper, we presented DAR-Net, a dual-ambiguity rectification network for all-in-one image restoration. We identify that existing unified restoration methods often suffer from semantic ambiguity in channel-wise representations and spatial ambiguity in spatial responses. To address this, we introduced DAR to learn structured degradation states, SeAR to improve channel-wise degradation discrimination, and SpAR to reduce spatial entanglement between degradation and content. Extensive experiments demonstrate that DAR-Net achieves strong and balanced restoration performance across diverse degradations. These results suggest that explicitly modeling what should be removed and what should be preserved is an effective direction for unified image restoration.

What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration

Acknowledgments This research was partially supported by the National Natural Science Foundation of China (NSFC) (62306064) and the Sichuan Science and Technology Program (granted No. 2024ZDZX0011, No. 2026NSFSC1482 and No. 2025ZHCG0002).

References [1] Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, and Ran He. 2024. Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 25432–25444. [2] Pablo Arbelaez, Michael Maire, Charless Fowlkes, and Jitendra Malik. 2010. Contour detection and hierarchical image segmentation. IEEE transactions on pattern analysis and machine intelligence 33, 5 (2010), 898–916. [3] Bolun Cai, Xiangmin Xu, Kui Jia, Chunmei Qing, and Dacheng Tao. 2016. DehazeNet: An End-to-End System for Single Image Haze Removal. IEEE Transactions on Image Processing 25 (2016), 5187–5198. [4] Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2020. Pre-Trained Image Processing Transformer. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 12294–12305. [5] I-Hsiang Chen, Wei-Ting Chen, Yu-Wei Liu, Yuan Chiang, Sy-Yen Kuo, and MingHsuan Yang. 2025. UniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion Prior. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025), 17969–17979. [6] Liangyu Chen, Xiaojie Chu, Xiangyu Zhang, and Jian Sun. 2022. Simple Baselines for Image Restoration. In Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VII (Tel Aviv, Israel). SpringerVerlag, Berlin, Heidelberg, 17–33. [7] Xiang Chen, Hao-Ran Li, Mingqiang Li, and Jin-shan Pan. 2023. Learning A Sparse Transformer Network for Effective Image Deraining. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 5896–5905. [8] Zixuan Chen, Zewei He, and Zhe-ming Lu. 2023. DEA-Net: Single Image Dehazing Based on Detail-Enhanced Convolution and Content-Guided Attention. IEEE Transactions on Image Processing 33 (2023), 1002–1015. [9] Sung-Jin Cho, Seoyoun Ji, Jun-Pyo Hong, Seung-Won Jung, and Sung-Jea Ko. 2021. Rethinking Coarse-to-Fine Approach in Single Image Deblurring. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 4621–4630. [10] Marcos V. Conde, Gregor Geigle, and Radu Timofte. 2024. InstructIR: HighQuality Image Restoration Following Human Instructions. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part XXXVI (Milan, Italy). Springer-Verlag, Berlin, Heidelberg, 1–21. [11] Yuning Cui, Syed Waqas Zamir, Salman Khan, Alois Knoll, Mubarak Shah, and Fahad Shahbaz Khan. 2025. Adair: Adaptive all-in-one image restoration via frequency mining and modulation. In 13th international conference on learning representations, ICLR 2025. International Conference on Learning Representations, ICLR, 57335–57356. [12] Qiyuan Guan, Qianfeng Yang, Xiang Chen, Tianyu Song, Guiyue Jin, and Jiyu Jin. 2025. WeatherBench: A Real-World Benchmark Dataset for All-in-One Adverse Weather Image Restoration. Proceedings of the 33rd ACM International Conference on Multimedia (2025). [13] Yu Guo, Yuan Gao, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, and Shengfeng He. 2024. Onerestore: A universal restoration framework for composite degradation. In European conference on computer vision. Springer, 255–272. [14] Junlin Han, Weihao Li, Pengfei Fang, Chunyi Sun, Jie Hong, Mohammad Ali Armin, Lars Petersson, and Hongdong Li. 2021. Blind Image Decomposition. In European Conference on Computer Vision. [15] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017). [16] Junjun Jiang, Zengyuan Zuo, Gang Wu, Kui Jiang, and Xianming Liu. 2025. A survey on all-in-one image restoration: Taxonomy, evaluation and future trends. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025). [17] Yitong Jiang, Zhaoyang Zhang, Tianfan Xue, and Jinwei Gu. 2024. Autodir: Automatic all-in-one image restoration with latent diffusion. In European Conference on Computer Vision. Springer, 340–359. [18] Lingshun Kong, Jiangxin Dong, Mingqiang Li, Jianjun Ge, and Jin-shan Pan. 2022. Efficient Frequency Domain-based Transformers for High-Quality Image Deblurring. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 5886–5895. [19] Boyun Li, Xiao Liu, Peng Hu, Zhongqin Wu, Jiancheng Lv, and Xiaocui Peng. 2022. All-In-One Image Restoration for Unknown Corruption. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 17431– 17441.

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

[20] Boyi Li, Wenqi Ren, Dengpan Fu, Dacheng Tao, Dan Feng, Wenjun Zeng, and Zhangyang Wang. 2018. Benchmarking single-image dehazing and beyond. IEEE transactions on image processing 28, 1 (2018), 492–505. [21] Ruoteng Li, Robby T. Tan, and Loong Fah Cheong. 2020. All in One Bad Weather Removal Using Architectural Search. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020), 3172–3182. [22] Jingyun Liang, Jie Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. 2021. SwinIR: Image Restoration Using Swin Transformer. 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (2021), 1833–1844. [23] Cencen Liu, Dongyang Zhang, Guoming Lu, Wen Yin, Jielei Wang, and Guangchun Luo. 2025. SRMamba-T: Exploring the hybrid Mamba–Transformer network for single image super-resolution. Neurocomputing 624 (2025), 129488. [24] Minghao Liu, Wenhan Yang, Jinyi Luo, and Jiaying Liu. 2025. Up-restorer: When unrolling meets prompts for unified image restoration. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 5513–5522. [25] Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjölund, and Thomas B Schön. 2024. Controlling Vision-Language Models for Multi-Task Image Restoration. In The Twelfth International Conference on Learning Representations, Vienna, Austria, May 7, 2024. The International Conference on Learning Representations (ICLR). [26] Kede Ma, Zhengfang Duanmu, Qingbo Wu, Zhou Wang, Hongwei Yong, Hongliang Li, and Lei Zhang. 2016. Waterloo exploration database: New challenges for image quality assessment models. IEEE Transactions on Image Processing 26, 2 (2016), 1004–1016. [27] David Martin, Charless Fowlkes, Doron Tal, and Jitendra Malik. 2001. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proceedings eighth IEEE international conference on computer vision. ICCV 2001, Vol. 2. Ieee, 416–423. [28] Seungjun Nah, Tae Hyun Kim, and Kyoung Mu Lee. 2017. Deep multi-scale convolutional neural network for dynamic scene deblurring. In Proceedings of the IEEE conference on computer vision and pattern recognition. 3883–3891. [29] Ozan Özdenizci and Robert A. Legenstein. 2022. Restoring Vision in Adverse Weather Conditions With Patch-Based Denoising Diffusion Models. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2022), 10346–10357. [30] Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan, and Fahad Shahbaz Khan. 2023. PromptIR: prompting for all-in-one blind image restoration. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, Article 3121, 19 pages. [31] Bin Ren, Yawei Li, Xu Zheng, Yuqian Fu, Danda Pani Paudel, Hong Liu, MingHsuan Yang, Luc Van Gool, and Nicu Sebe. 2026. Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization. In The Fourteenth International Conference on Learning Representations. [32] Hao Shen, Zhongliu Zhao, and Wandi Zhang. 2022. Adaptive Dynamic Filtering Network for Image Denoising. In AAAI Conference on Artificial Intelligence. [33] Hao Shen, Zhong-Qiu Zhao, Yulun Zhang, and Zhao Zhang. 2023. Mutual Information-driven Triple Interaction Network for Efficient Image Dehazing. Proceedings of the 31st ACM International Conference on Multimedia (2023). [34] Yuda Song, Zhuqing He, Hui Qian, and Xin Du. 2022. Vision Transformers for Single Image Dehazing. IEEE Transactions on Image Processing 32 (2022), 1927–1941. [35] Shangquan Sun, Wenqi Ren, Xinwei Gao, Rui Wang, and Xiaochun Cao. 2024. Restoring Images in Adverse Weather Conditions via Histogram Transformer. In European Conference on Computer Vision. [36] Xiaole Tang, Xiaoyi He, Jiayi Xu, Xiang Gu, and Jian Sun. 2026. Learning Continuous Wasserstein Barycenter Space for Generalized All-in-One Image Restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence (2026). [37] Xiangpeng Tian, Xiangyu Liao, Xiao Liu, Meng Li, and Chao Ren. 2025. Degradation-Aware Feature Perturbation for All-in-One Image Restoration. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025), 28165–28175. [38] Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Conrad Bovik, and Yinxiao Li. 2022. MAXIM: Multi-Axis MLP for Image Processing. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 5759–5770. [39] Jeya Maria Jose Valanarasu, Rajeev Yasarla, and Vishal M. Patel. 2021. TransWeather: Transformer-based Restoration of Images Degraded by Adverse Weather Conditions. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 2343–2353. [40] Cong Wang, Yutong Wu, Zhixun Su, and Junyang Chen. 2020. Joint Self-Attention and Scale-Aggregation for Self-Calibrated Deraining Network. Proceedings of the 28th ACM International Conference on Multimedia (2020). [41] Yongzhen Wang, Yongjun Li, Zhuoran Zheng, Xiaoping Zhang, and Mingqiang Wei. 2025. M2Restore: Mixture-of-Experts-Based Mamba-CNN Fusion Framework for All-in-One Image Restoration. IEEE Transactions on Image Processing 34 (2025), 8086–8100. [42] Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on

MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil

Image Processing 13, 4 (2004), 600–612. [43] Zhendong Wang, Xiaodong Cun, Jianmin Bao, and Jianzhuang Liu. 2021. Uformer: A General U-Shaped Transformer for Image Restoration. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 17662–17672. [44] Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. 2018. Deep Retinex Decomposition for Low-Light Enhancement. In British Machine Vision Conference 2018. BMVA Press, 155. [45] Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G. Dimakis, and Peyman Milanfar. 2021. Deblurring via Stochastic Refinement. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 16272–16282. [46] Jie Xiao, Xueyang Fu, Aiping Liu, Feng Wu, and Zhengjun Zha. 2022. Image De-Raining Transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2022), 12978–12995. [47] Hao Yang, Liyuan Pan, Yan Yang, and Wei Liang. 2024. Language-driven allin-one adverse weather removal. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 24902–24912. [48] Wenhan Yang, Robby T Tan, Jiashi Feng, Jiaying Liu, Zongming Guo, and Shuicheng Yan. 2017. Deep joint rain detection and removal from a single image. In Proceedings of the IEEE conference on computer vision and pattern recognition. 1357–1366. [49] Qiaosi Yi, Juncheng Li, Qi Dai, Faming Fang, Guixu Zhang, and Tieyong Zeng. 2021. Structure-Preserving Deraining with Residue Channel Prior Guidance. 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (2021), 4218–4227. [50] Xiaoyan Yu, Shen Zhou, Huafeng Li, and Liehuang Zhu. 2024. Multi-Expert Adaptive Selection: Task-Balancing for All-in-One Image Restoration. IEEE Transactions on Circuits and Systems for Video Technology 35 (2024), 4619–4634. [51] Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan, Danda Pani Paudel, Yulun Zhang, and Radu Timofte. 2025. Complexity Experts are Task-Discriminative Learners for Any Image Restoration. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025), 12753–12763. [52] Syed Waqas Zamir, Aditya Arora, Salman Hameed Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. 2022. Restormer: Efficient Transformer

Liu et al.

for High-Resolution Image Restoration. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022), 5718–5729. [53] Haijin Zeng, Xiangming Wang, Yongyong Chen, Jingyong Su, and Jie Liu. 2025. Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025), 7524–7533. [54] Jinghao Zhang, Jie Huang, Mingde Yao, Zizheng Yang, Huikang Yu, Man Zhou, and Fengmei Zhao. 2023. Ingredient-oriented Multi-Degradation Learning for Image Restoration. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 5825–5835. [55] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition. 586–595. [56] Rongyu Zhang, Yulin Luo, Jiaming Liu, Huanrui Yang, Zhen Dong, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, Yuan Du, et al. 2024. Efficient deweahter mixture-of-experts with uncertainty-aware feature-wise linear modulation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 16812–16820. [57] Xu Zhang, Huan Zhang, Guoli Wang, Qian Zhang, Lefei Zhang, and Bo Du. 2026. ClearAIR: A Human-Visual-Perception-Inspired All-in-One Image Restoration. In Proceedings of the AAAI Conference on Artificial Intelligence. [58] Dian Zheng, Xiao-Ming Wu, Shuzhou Yang, Jian Zhang, Jian-Fang Hu, and WeiShi Zheng. 2024. Selective hourglass mapping for universal image restoration based on diffusion model. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 25445–25455. [59] Yurui Zhu, Tianyu Wang, Xueyang Fu, X. Yang, Xin Guo, Jifeng Dai, Yu Qiao, and Xiaowei Hu. 2023. Learning Weather-General and Weather-Specific Features for Image Restoration Under Multiple Adverse Weather Conditions. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 21747– 21758.

Record · ID 414138 · SHA-256 0f53c9958a05d34e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.